Conversion for matrix-based intra prediction in video coding

Matrix-based intra prediction (MIP) with Low Frequency Non-Separable Transform (LFNST) in video coding addresses the inefficiencies of high-resolution media transmission and storage, enhancing compression efficiency and reducing complexity.

JP2025113307AActive Publication Date: 2025-08-01LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025081838
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-16
Filing Date
2025-05-15
Publication Date
2025-08-01
Estimated Expiration
2040-04-16

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, as well as immersive media such as VR and AR content, has led to higher transmission and storage costs due to increased data volume, necessitating more efficient video coding methods.

Method used

The implementation of matrix-based intra prediction (MIP) in video coding, including methods for signaling, deriving, and coding conversion indices, along with the use of Low Frequency Non-Separable Transform (LFNST) to minimize interference and reduce complexity.

Benefits of technology

This approach enhances video compression efficiency, reduces interference between MIP and LFNST, and maintains optimal coding efficiency while minimizing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113307000001_ABST
    Figure 2025113307000001_ABST
Patent Text Reader

Abstract

To provide a video decoding method.SOLUTION: A method includes a step of obtaining intra-prediction type information and information about a conversion factor from a bitstream, a step of deriving a conversion factor for a current block based on the information about the conversion factor, and a step of generating a residual sample for the current block based on LFNST index information from the conversion factor. The intra-prediction type information includes an MIP flag indicating whether MIP is applied to the current block. The information about the conversion factor includes the LFNST index information based on the MIP flag, and the LFNST index information is set in a coding unit syntax based on the MIP flag. Based on the value of the MIP flag being equal to 1, the information about the conversion factor does not include LFNST index information, and based on LFNST index information being absent, the value of LFNST index information is derived as 0.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to video coding technology, and more particularly, to a transform for matrix-based intra prediction in video coding.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003]

[0004] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

Summary of the Invention

Means for Solving the Problems

[0005] According to an embodiment of this document, a method and an apparatus for increasing video / video coding efficiency are provided.

[0006] ​According to one embodiment of the present document, a conversion method and apparatus for a block to which matrix-based intra prediction (MIP) in video coding is applied are provided.

[0007] According to one embodiment of the present document, a method and apparatus for signaling a conversion index for a block to which MIP is applied are provided.

[0008] According to one embodiment of the present document, a method and apparatus for signaling a conversion index for a block to which MIP is not applied are provided.

[0009] According to one embodiment of the present document, a method and apparatus for deriving a conversion index for a block to which MIP is applied are provided.

[0010] According to one embodiment of the present document, a method and apparatus for binary evolution or coding of a conversion index for a block to which MIP is applied are provided.

[0011] According to one embodiment of the present document, a video / video decoding method executed by a decoding apparatus is provided.

[0012] According to one embodiment of the present document, a decoding apparatus for executing video / video decoding is provided.

[0013] According to one embodiment of the present document, a video / video encoding method executed by an encoding apparatus is provided.

[0014] According to one embodiment of the present document, an encoding apparatus for executing video / video encoding is provided.

[0015] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method disclosed in at least one of the embodiments of this document.

[0016] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded information or encoded video / video information for causing a decoding device to execute a video / video decoding method disclosed in at least one of the embodiments of this document.

Advantages of the Invention

[0017] According to this document, the overall video / video compression efficiency can be increased.

[0018] According to this document, the transform index for a block to which MIP (Matrix based Intra Prediction) is applied can be efficiently signaled.

[0019] According to this document, the transform index for a block to which MIP is applied can be efficiently coded.

[0020] According to this document, the transform index for a block to which MIP is applied can be induced without separately signaling it.

[0021] According to this document, when both MIP and LFNST (Low Frequency Non-Separable Transform) are applied, the interference between them can be minimized, the optimal coding efficiency can be maintained, and the complexity can be reduced.

[0022] The effects that can be obtained through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Accordingly, the specific effects of this document can include various effects that can be understood or derived from the technical features of this document, rather than being limited to those explicitly described in this document.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Mode for Carrying Out the Invention

[0024] This document can be modified in various ways and can have various embodiments. Specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used merely to describe specific embodiments and are not intended to limit the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.

[0025] On the one hand, for the convenience of explaining the characteristic functions of each component in the drawings described in this document, each component is independently illustrated, but this does not mean that each component is implemented by separate hardware or separate software. For example, among the components, two or more components can be combined to form one component, and one component can also be divided into multiple components. As long as the embodiments in which the components are integrated and / or separated do not deviate from the essence of this document, they are included in the scope of rights of this document.

[0026] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals are used for the same components in the drawings, and duplicate descriptions of the same components can be omitted.

[0027] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document can be related to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (for example, the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0028] This document presents various embodiments related to video / video coding, and unless otherwise mentioned, the above embodiments can also be executed in combination with each other.

[0029] In this document, "video" can mean a collection of a series of images over time. "Picture" generally means a unit indicating one image in a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles.

[0030] "Pixel" or "pel" can mean the smallest unit that constitutes one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Or, a sample can also mean the pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also mean the conversion coefficient in the frequency domain.

[0031] "Unit" can indicate the basic unit of video processing. A unit can include at least one of a specific area of a picture and information related to that area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0032] In this document, the terms “ / ” and “,” are interpreted to mean “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Additionally, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” also means “at least one of A, B, and / or C.”

[0033] Further, in this document, the term “or” is interpreted to mean “and / or.” For example, “A or B” can mean 1) only “A,” or 2) only “B,” or 3) “A and B.” As another expression, the “or” in this document can mean “additionally or alternatively.”

[0034] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0035] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0036] Also, the parentheses used in this specification can mean "for example". Specifically, when shown as "prediction (Intra prediction)", "Intra prediction" is proposed as an example of "prediction". As another expression, "prediction" in this specification is not limited to "Intra prediction", but "Intra prediction" is proposed as an example of "prediction". Also, when shown as "prediction (i.e., Intra prediction)", "Intra prediction" is proposed as an example of "prediction".

[0037] In this specification, the technical features individually described within one drawing can be embodied individually or simultaneously.

[0038] FIG. 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0039] Referring to FIG. 1, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data in file or streaming form to the receiving device via a digital storage medium or a network.

[0040] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can be composed of a separate device or an external component.

[0041] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like, in which case the video / image capture process can be replaced by a process in which relevant data is generated.

[0042] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0043] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0044] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0045] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0046] Figure 2 is a diagram schematically explaining the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the video encoding device can include the image encoding device.

[0047] As shown in FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0048] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the above-described final coding unit.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0049] The unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can represent a set such as samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or can also represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).

[0050] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0051] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block according to the prediction mode, or can be located remotely. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0052] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0053] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can also execute intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / moving video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0054] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square and can also be applied to a block of variable size that is not square.

[0055] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in units of NAL (network abstraction layer) units in a bitstream form. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.

[0056] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0057] On the other hand, LMCS (luma mapping with chrom ascaling) can also be applied in the picture encoding and / or restoration process.

[0058] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 290, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 290 and output in the form of a bitstream.

[0059] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 280. Through this, when inter prediction is applied, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0060] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0061] FIG. 3 is a diagram schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.

[0062] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The entropy decoding unit 310, the residual processing unit 320, the prediction unit 330, the addition unit 340, and the filtering unit 350 described above can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Further, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0063] If a bitstream including video / image information is input, the decoding apparatus 300 can restore an image corresponding to the process in which the video / image information is processed by the encoding apparatus in FIG. 3. For example, the decoding apparatus 300 can derive a unit / block based on the block division related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using the processing unit applied in the encoding apparatus. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding apparatus 300 can be reproduced via a reproducing apparatus.

[0064] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the decoding target syntax element information adjacent to and the decoding information of the decoding target block or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and executes arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the information regarding the residual for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0065] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0066] In the inverse conversion unit 322, the conversion coefficient is inversely converted to obtain a residual signal (residual block, residual sample array).

[0067] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0068] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can execute intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed to be similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0069] The intra prediction unit 332 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 332 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0070] The inter prediction unit 331 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 331 can construct a motion information candidate list based on the adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0071] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor 330. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0072] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and as will be described later, can also be output after filtering, or can be used for inter prediction of the next picture.

[0073] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0074] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can send the modified restored picture to the memory 60, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0075] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 331. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 331 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 332.

[0076] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding apparatus 300 can be applied so as to be the same or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding apparatus 200, respectively.

[0077] On the other hand, as described above, prediction is performed to improve the compression efficiency in performing video coding. Thereby, a predicted block including prediction samples for the current block which is the block to be coded can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same manner in the encoding apparatus and the decoding apparatus, and the encoding apparatus can improve the image coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block which is not the original sample value of the original block to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.

[0078] The residual information can be generated via conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thereby signal (via a bitstream) the related residual information to a decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block, and generate a restored picture based on this.

[0079] FIG. 4 schematically shows the multiple conversion techniques according to this document.

[0080] Referring to FIG. 4, the conversion unit can correspond to the conversion unit in the encoding device of FIG. 2 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of FIG. 2 or the inverse conversion unit in the decoding device of FIG. 3 described above.

[0081] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S410). Such a primary conversion can be called a core transform. Here, the primary conversion can be based on Multiple Transform Selection (MTS), and when multiple conversions are applied in the primary conversion, it can be called a multiple core transform.

[0082] For example, the multiple core transform can be shown as a method of performing conversion by additionally using Discrete Cosine Transform (DCT) type 2 (DCT-II), Discrete Sine Transform (DST) type 7 (DST-VII), DCT type 8 (DCT-VIII), and / or DST type 1 (DST-I). That is, the multiple core transform can be shown as a conversion method for converting a residual signal (or residual block) in the spatial domain into conversion coefficients (or primary conversion coefficients) in the frequency domain based on a plurality of selected transform kernels from among the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary conversion coefficients can be called temporary conversion coefficients on the conversion unit side.

[0083] That is, when an existing conversion method is applied, a conversion from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2 to generate conversion coefficients. However, in contrast, when the multi-core conversion is applied, a conversion from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or primary conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be called conversion types, conversion kernels, or conversion cores. Such DCT / DST conversion types can be defined based on basis functions.

[0084] When the multi-core conversion is executed, a vertical conversion kernel and / or a horizontal conversion kernel for the target block can be selected from among the conversion kernels, and a vertical conversion for the target block can be executed based on the vertical conversion kernel, and a horizontal conversion for the target block can be executed based on the horizontal conversion kernel. Here, the horizontal conversion can indicate a conversion for the horizontal component of the target block, and the vertical conversion can indicate a conversion for the vertical component of the target block. The vertical conversion kernel / horizontal conversion kernel can be adaptively determined based on the prediction mode and / or conversion index of the target block (CU or sub-block) including the residual block.

[0085] Alternatively, for example, when applying MTS to perform a primary transformation, specific basis functions can be set to predetermined values, and when it is a vertical transformation or a horizontal transformation, the mapping relationship for the transformation kernel can be set by combining which basis functions are applied. For example, when representing the horizontal transformation kernel as trTypeHor and the vertical transformation kernel as trTypeVer, trTypeHor or trTypeVer having a value of 0 can be set to DCT2, and trTypeHor or trTypeVer having a value of 1 can be set to DST7. trTypeHor or trTypeVer having a value of 2 can be set to DCT8.

[0086] Alternatively, for example, in order to indicate any one of a number of transformation kernel sets, an MTS index can be encoded and the MTS index information can be signaled to the decoding device. Here, the MTS index can be represented by the tu_mts_idx syntax element or the mts_idx syntax element. For example, when the MTS index is 0, it can indicate that all trTypeHor and trTypeVer values are 0. When the MTS index is 1, it can indicate that all trTypeHor and trTypeVer values are 1. When the MTS index is 2, it can indicate that the trTypeHor value is 2 and the trTypeVer value is 1. When the MTS index is 3, it can indicate that the trTypeHor value is 1 and the trTypeVer value is 2. When the MTS index is 4, it can indicate that all trTypeHor and trTypeVer values are 2. For example, the transformation kernel sets according to the MTS index can be shown as in the following table.

[0087]

Table 1

[0088] The conversion unit can derive a corrected (secondary) conversion coefficient by performing a secondary conversion based on the (primary) conversion coefficient (S420). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion can be shown to convert to a more compressed representation by utilizing the correlation existing between the (primary) conversion coefficients.

[0089] For example, the secondary conversion can include a non-separable transform. In this case, the secondary conversion can be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform can be shown to be a transform that performs a secondary conversion on the (primary) conversion coefficient derived through the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, the conversion can be applied at once without separating the vertical conversion and the horizontal conversion (or independently applying the horizontal conversion and the vertical conversion) to the (primary) conversion coefficient based on the non-separable transform matrix.

[0090] That is, the non-separable secondary transform can be shown to be a conversion method that, without separating the vertical component and the horizontal component of the (primary) conversion coefficient, for example, rearranges a two-dimensional signal (conversion coefficient) into a one-dimensional signal through a specific determined direction (e.g., row-first direction or column-first direction), and then generates a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix.

[0091] For example, the row-major direction (or order) can be shown as arranging in a column in the order of the first row, the second row, …, the Nth row for an M×N block, and the column-major direction (or order) can be shown as arranging in a column in the order of the first column, the second column, …, the Mth column for an M×N block. Here, M and N can each indicate the width (W) and height (H) of the block, and are all positive integers.

[0092] For example, the non-separable second-order transform can be applied to the top-left region of a block (hereinafter referred to as a transform coefficient block) composed of (first-order) transform coefficients. For example, when the width (W) and height (H) of the transform coefficient block are all 8 or more, an 8×8 non-separable second-order transform can be applied to the top-left 8×8 region of the transform coefficient block. Also, when the width (W) and height (H) of the transform coefficient block are all 4 or more and the width (W) or height (H) of the transform coefficient block is less than 8, a 4×4 non-separable second-order transform can be applied to the top-left min(8, W)×min(8, H) region of the transform coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width (W) or height (H) of the transform coefficient block is all 4 or more is satisfied, a 4×4 non-separable second-order transform can also be applied to the top-left min(8, W)×min(8, H) region of the transform coefficient block.

[0093] Specifically, for example, when a 4×4 input block is used, the non-separable second-order transform can be executed as follows.

[0094] The 4×4 input block X is shown as follows.

[0095]

Number

[0096] For example, the vector form of the above X is shown as follows.

[0097]

Mathematics

[0098] Referring to Equation 2, JPEG2025113307000005.jpg64 can represent vector X, which is shown by rearranging the 2D block of X in Equation 1 into a 1D vector in row-first order.

[0099] In this case, the second non-separable transform can be calculated as follows.

[0100]

Mathematics

[0101] Here, JPEG2025113307000007.jpg64 can represent the transform coefficient vector, and T can represent a 16×16 (non-separable) transform matrix.

[0102] Based on Equation 3, a 16×1 size of JPEG2025113307000008.jpg64 can be derived, and the JPEG2025113307000009.jpg64 can be re-organized into 4×4 blocks through a scan order (such as horizontal, vertical, or diagonal, etc.). However, the above calculation is just an example, and HyGT (Hypercube-Givens Transform) etc. can also be used for the calculation of non-separable second-order transforms to reduce the computational complexity of non-separable second-order transforms.

[0103] On the other hand, for the non-separable second-order transform, a mode-dependent transform kernel (or transform core, transform type) can also be selected. Here, the mode can include an intra prediction mode and / or an inter prediction mode.

[0104] For example, as described above, the non-separable second-order transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. For example, the 8×8 transform can represent a transform that can be applied to an 8×8 area included inside the corresponding transform coefficient block when both W and H are the same as or greater than 8, and the 8×8 area is the upper left 8×8 area inside the corresponding transform coefficient block. Similarly, the 4×4 transform can represent a transform that can be applied to a 4×4 area included inside the corresponding transform coefficient block when both W and H are the same as or greater than 4, and the 4×4 area is the upper left 4×4 area inside the corresponding transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.

[0105] At this time, for mode-based transform kernel selection, two non-separable second-order transform kernels per transform set for non-separable second-order transform can be configured for both the 8×8 transform and the 4×4 transform, and the number of transform sets is four. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.

[0106] However, the size of the transform, the number of sets, and the number of transform kernels in the set are only examples, and sizes other than 8×8 or 4×4 can also be used, or n sets can be configured, and each set can include k transform kernels. Here, n and k are positive integers, respectively.

[0107] For example, the conversion set can be called an NSST set, and the conversion kernels in the NSST set can be called NSST kernels. For example, the selection of a specific set from the conversion set can be performed based on the intra prediction mode of the target block (CU or sub-block).

[0108] For example, the intra prediction mode can include two non-directional or non-angular intra prediction modes and 65 directional or angular intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is only an example, and the embodiments according to this document can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can also indicate the LM (linear model) mode.

[0109] FIG. 5 exemplarily shows the intra-directional modes of 65 prediction directions.

[0110] Referring to FIG. 5, around the 34th intra prediction mode having the upper left diagonal prediction direction, the intra prediction modes having horizontal directionality and the intra prediction modes having vertical directionality can be distinguished. H and V in FIG. 5 can respectively mean horizontal directionality and vertical directionality, and the numbers from -32 to 32 can indicate a displacement in units of 1 / 32 on the sample grid position. This can indicate an offset with respect to the mode index value.

[0111] For example, the 2nd to 33rd intra prediction modes can have a horizontal directionality, and the 34th to 66th intra prediction modes can have a vertical directionality. On the other hand, the 34th intra prediction mode can be considered to have neither a strict horizontal nor vertical directionality, but can be classified as belonging to the horizontal directionality from the perspective of determining the conversion set of the secondary conversion. The reason is that for the vertical direction modes symmetric about the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the input data alignment method for the horizontal direction mode is used. Here, transposing the input data can mean that for the two-dimensional block data M×N, the rows become columns and the columns become rows to form N×M data.

[0112] Also, the 18th intra prediction mode and the 50th intra prediction mode can respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode has a left reference pixel and predicts in the upper right direction, so it can be called an upper right diagonal intra prediction mode. Similarly, the 34th intra prediction mode can be called a lower right diagonal intra prediction mode, and the 66th intra prediction mode can be called a lower left diagonal intra prediction mode.

[0113] Once it is determined that a specific set is used for the non-separable transform, one of the k transform kernels in the specific set can be selected via the non-separable second-order transform index. For example, the encoding device can derive a non-separable second-order transform index indicating a specific transform kernel based on rate-distortion (RD) checking, and can signal the non-separable second-order transform index to the decoding device. For example, the decoding device can select one of the k transform kernels in the specific set based on the non-separable second-order transform index. For example, an NSST index having a value of 0 can indicate the first non-separable second-order transform kernel, an NSST index having a value of 1 can indicate the second non-separable second-order transform kernel, and an NSST index having a value of 2 can indicate the third non-separable second-order transform kernel. Or, an NSST index having a value of 0 can indicate that the first non-separable second-order transform is not applied to the target block, and an NSST index having a value from 1 to 3 can point to the three transform kernels.

[0114] The transform unit can perform the non-separable second-order transform based on the selected transform kernel and obtain the modified (second-order) transform coefficients. As described above, the modified transform coefficients can be derived as the transform coefficients quantized via the quantization unit, encoded, and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0115] On the other hand, as described above, when the second-order transform is omitted, the (first-order) transform coefficients that are the output of the first-order (separable) transform can be derived as the transform coefficients quantized via the quantization unit, encoded, and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0116] Referring again to FIG. 4, the inverse conversion unit can execute a series of procedures in the reverse order of the procedures executed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficients, performs a secondary (inverse) conversion to derive the (primary) conversion coefficients (S450), and can perform a primary (inverse) conversion on the (primary) conversion coefficients to obtain a residual block (residual samples) (S460). Here, the primary conversion coefficients can be referred to as modified conversion coefficients on the inverse conversion unit side. As described above, the encoding device and / or the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.

[0117] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). For example, the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion is NSST or RST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on the secondary conversion flag parsed or obtained from the bitstream. Or, for example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficients of the residual block.

[0118] The secondary inverse conversion determination unit can determine the secondary inverse conversion. At this time, the secondary inverse conversion determination unit can determine the secondary inverse conversion applied to the current block based on the NSST (or RST) conversion set specified by the intra prediction mode. Or, the secondary conversion determination method can be determined depending on the primary conversion determination method. Or, various combinations of the primary conversion and the secondary conversion can be determined by the intra prediction mode. For example, the secondary inverse conversion determination unit can also determine the area to which the secondary inverse conversion is applied based on the size of the current block.

[0119] On the one hand, as described above, when the secondary (inverse) transform is omitted, a residual block (residual sample) can be obtained by receiving the (inverse quantized) transform coefficient and performing the primary (separation) inverse transform. As described above, the encoding device and / or the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.

[0120] On the other hand, in this document, in order to reduce the computational amount and memory requirement amount by non-separable secondary transform, an RST (reduced secondary transform) in which the size of the transform matrix (kernel) is reduced in the concept of NSST can be applied.

[0121] In this document, RST can be meant to be a (simplified) transform performed on the residual sample for the target block based on a transform matrix whose size is reduced by a simplification factor. When this is executed, the amount of computation required at the time of transform can be reduced by reducing the size of the transform matrix. That is, RST can be used to solve the problem of computational complexity that occurs during the transform of a large-sized block or non-separable transform.

[0122] For example, RST can be called by various terms such as reduced transform, reduced secondary transform, reduction transform, simplified transform or simple transform, etc., and the name by which RST is called is not limited to the listed examples. Or, since RST is mainly performed in the low-frequency region including coefficients that are not 0 in the transform block, it can be called LFNST (Low-Frequency Non-Separable Transform).

[0123] On the other hand, when the second inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 can include an inverse RST unit that derives a transformation coefficient corrected based on the inverse RST for the transformation coefficient, and an inverse primary transformation unit that derives a residual sample for the target block based on the inverse primary transformation for the corrected transformation coefficient. The inverse primary transformation means the inverse transformation of the primary transformation applied to the residual. In this document, deriving a transformation coefficient based on a transformation can mean deriving the transformation coefficient by applying the corresponding transformation.

[0124] FIG. 6 and FIG. 7 are diagrams for explaining RST according to an embodiment of this document.

[0125] For example, FIG. 6 is a drawing for explaining that a forward reduced transform is applied, and FIG. 7 is a diagram for explaining that an inverse reduced transform is applied. In this document, the target block can indicate the current block where coding is executed, the residual block, or the transform block.

[0126] For example, in RST, an N-dimensional vector can be mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined. Here, N and R are each positive integers, and R is smaller than N. N can represent the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can represent the R / N value. The simplification factor can be called by various terms such as reduced factor, reduction factor, simplified factor, or simple factor. On the other hand, R can be called a reduced coefficient, but in some cases, the simplification factor can also represent R. Also, in some cases, the simplification factor can represent the N / R value.

[0127] For example, the simplification factor or reduced coefficient can be signaled via a bitstream, but is not limited thereto. For example, there may be cases where pre-defined values for the simplification factor or reduced coefficient are stored in each encoding device 200 and decoding device 300, and in this case, the simplification factor or reduced coefficient is not signaled separately.

[0128] For example, the size (R×N) of the simplified transformation matrix is smaller than the size (N×N) of the normal transformation matrix and can be defined as in the following formula.

[0129]

Equation

[0130] For example, the matrix T in the reduced transform block shown in FIG. 6 is the matrix T of Equation 4 R×N and can be shown. When the simplified transform matrix T R×N is multiplied by the residual samples for the target block as shown in FIG. 6, the transform coefficients for the target block can be derived.

[0131] For example, when the size of the block to which the transform is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6 can be expressed by the matrix operation of the following Equation 5. In this case, the memory and multiplication operations can be reduced to approximately 1 / 4 by the simplification factor.

[0132] In this document, matrix multiplication can be understood as an operation in which a matrix is placed on the left side of a column vector and the matrix and the column vector are multiplied to obtain a column vector.

[0133]

Equation

[0134] In Equation 5, r1 to r 64 can represent the residual samples for the target block. Or, for example, they are the transform coefficients generated by applying a first-order transform. Based on the operation result of Equation 5, the transform coefficients c i for the target block can be derived.

[0135] For example, when R is 16, the transform coefficients c1 to c 16can be derived. If a normal (regular) conversion is applied instead of RST, and a conversion matrix with a size of 64×64 (N×N) is multiplied by a residual sample with a size of 64×1 (N×1), 64 (N) conversion coefficients for the target block are derived. However, since RST is applied, only 16 (R) conversion coefficients for the target block are derived. Since the total number of conversion coefficients for the target block decreases from N to R and the amount of data transmitted from the encoding device 200 to the decoding device 300 decreases, the transmission efficiency between the encoding device 200 and the decoding device 300 can be increased.

[0136] Considering the size aspect of the conversion matrix, the size of the normal conversion matrix is 64×64 (N×N), while the size of the simplified conversion matrix decreases to 16×64 (R×N). Therefore, when compared with the case of performing a normal conversion, the memory usage can be reduced at a ratio of R / N when performing RST. Also, when compared with the number of multiplication operations N×N when using a normal conversion matrix, the number of multiplication operations can be reduced at a ratio of R / N (R×N) when using a simplified conversion matrix.

[0137] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual sample for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300. The inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and can derive the residual sample for the target block based on an inverse primary conversion for the modified conversion coefficients.

[0138] Inverse RST matrix T according to one embodiment N×RThe size of is N×R, which is smaller than the size N×N of the normal inverse transform matrix, and is the simplified transform matrix T shown in Equation 4 R×N is in a transpose relationship with

[0139] The matrix T in the reduced inverse transform block shown in FIG. 7 t is the inverse RST matrix T R×N T can be shown. Here, the superscript T can indicate transpose. When the inverse RST matrix T R×N T is multiplied by the transform coefficients for the target block as shown in FIG. 7, the modified transform coefficients for the target block or the residual samples for the target block can be derived. The inverse RST matrix T R×N T can also be expressed as (T R×N ) T N×R More specifically, when inverse RST is applied in the second-order inverse transform, when the inverse RST matrix T

[0140] is multiplied by the transform coefficients for the target block, the modified transform coefficients for the target block can be derived. On the other hand, inverse RST can be applied in the inverse first-order transform. In this case, when the inverse RST matrix T R×N T is multiplied by the transform coefficients for the target block, the residual samples for the target block can be derived. R×N T In one embodiment, when the size of the block to which the inverse transform is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7 can be expressed by the matrix operation as shown in Equation 6 below.

[0141] In one example, when the size of the block to which the inverse transform is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7 can be expressed by the matrix operation as shown in the following Equation 6.

[0142]

Equation

[0143] In Equation 6, c1 to c 16 can indicate the conversion coefficient for the target block. Based on the calculation result of Equation 6, r j indicating the corrected conversion coefficient for the target block or the residual sample for the target block can be derived. That is, r1 to r N indicating the corrected conversion coefficient for the target block or the residual sample for the target block can be derived.

[0144] Considering from the perspective of the size of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64×64 (N×N), while the size of the simplified inverse transformation matrix is reduced to 64×16 (N×R). Therefore, compared with the case of performing the normal inverse transformation, when performing the inverse RST, the memory usage can be reduced at a ratio of R / N. Also, compared with the number of multiplication operations N×N when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced at a ratio of R / N (N×R).

[0145] On the other hand, a conversion set can also be configured and applied to 8×8 RST. That is, the corresponding 8×8 RST can be applied by the conversion set. Since one conversion set is composed of two or three conversion kernels depending on the prediction mode within the screen, it can be configured to select one from a maximum of four conversions including the case where no second-order conversion is applied. The conversion when no second-order conversion is applied is regarded as the identity matrix being applied. When each of the four conversions is assigned an index of 0, 1, 2, or 3 (for example, the 0th index can be assigned to the identity matrix, i.e., the case where no second-order conversion is applied), a syntax element called the NSST index can be signaled for each conversion coefficient block to specify the conversion to be applied. That is, 8×8 NSST can be specified for the 8×8 upper-left block via the NSST index, and 8×8 RST can be specified in the RST configuration. 8×8 NSST and 8×8 RST can indicate the conversions that can be applied to the 8×8 area contained within the corresponding conversion coefficient block when all of the W and H of the target block to be converted are the same as or larger than 8, and the said 8×8 area is the upper-left 8×8 area within the corresponding conversion coefficient block. Also, similarly, 4×4 NSST and 4×4 RST can indicate the conversions that can be applied to the 4×4 area contained within the corresponding conversion coefficient block when all of the W and H of the target block are the same as or larger than 4, and the said 4×4 area is the upper-left 4×4 area within the corresponding conversion coefficient block.

[0146] One party, for example, an encoding device can derive a bitstream by encoding values of syntax elements or quantized values of transform coefficients related to residuals based on various coding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), or CABAC (context-adaptive binary arithmetic coding). Also, a decoding device can decode the bitstream based on various coding methods such as exponential Golomb coding, CAVLC, or CABAC, and derive values of syntax elements necessary for video restoration or quantized values of transform coefficients related to residuals.

[0147] For example, the coding methods described above can be executed as described in the content below.

[0148] FIG. 8 exemplarily shows CABAC (context-adaptive binary arithmetic coding) for encoding syntax elements.

[0149] For example, in the coding process of CABAC, when the input signal is a syntax element that is not a binary value, the encoding device can binarize the value of the input signal to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), the input signal can be used as it is without performing binarization. Here, each binary number 0 or 1 that constitutes a binary value can be called a bin. For example, when the binary string after binarization is 110, each of 1, 1, and 0 can be represented as one bin. The bin for one syntax element can indicate the value of the syntax element. Such binarization can be based on various binarization methods such as Truncated Rice binarization process or Fixed-length binarization process, and the binarization method for the target syntax element can be defined in advance. The binarization procedure can be executed by a binarization unit within the entropy encoding unit.

[0150] Thereafter, the binarized bin of the syntax element can be input to a regular coding engine or a bypass coding engine. The regular coding engine of the encoding device can assign a context model that reflects a probability value to the corresponding bin, and can encode the corresponding bin based on the assigned context model. The regular coding engine of the encoding device can update the context model for the corresponding bin after performing coding for each bin. The bin coded as described above can be represented as a context-coded bin.

[0151] On the one hand, when the binary bin of the syntax element is input to the bypass coding engine, it can be coded as follows. For example, the bypass coding engine of the encoding device can omit the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after coding. When bypass coding is applied, the encoding device can code the input bin by applying a uniform probability distribution instead of assigning a context model, and through this, the encoding speed can be improved. The bin coded as described above can be represented as a bypass bin.

[0152] Entropy decoding can show the process of executing the same process as the aforementioned entropy encoding in reverse order.

[0153] The decoding device (entropy decoding unit) can decode the encoded video / video information. The video / video information can include partitioning-related information, prediction-related information (e.g., inter / intra prediction partition information, intra prediction mode information, inter prediction mode information, etc.), residual information, or in-loop filtering-related information, or can include various syntax elements related thereto. The entropy coding can be executed in units of syntax elements.

[0154] The decoding device can perform binarization on the target syntax element. Here, the binarization can be based on various binarization methods such as Truncated Rice binarization process or Fixed-length binarization process, and the binarization method for the target syntax element can be defined in advance. The decoding device can derive a usable bin string (bin string candidate) for the usable value of the target syntax element through the binarization procedure. The binarization procedure can be executed by a binarization unit within the entropy decoding unit.

[0155] The decoding device can compare the derived bin string with the usable bin string for the corresponding syntax element while sequentially decoding or parsing each bin for the target syntax element from the input bits in the bitstream. If the derived bin string is the same as one of the usable bin strings, the value corresponding to the bin string is derived as the value of the corresponding syntax element. If not, the next bit in the bitstream can be further parsed and the above-mentioned procedure can be executed again. Through such a process, without using start bits or end bits for specific information (or specific syntax elements) in the bitstream, variable-length bits can be used to signal the corresponding information. Through this, relatively fewer bits can be allocated for lower values, and the overall coding efficiency can be improved.

[0156] The decoding device can decode each bin in the bin string from the bitstream based on an entropy coding technique such as CABAC or CAVLC, either based on a context model or bypass.

[0157] When a syntax element is decoded based on a context model, the decoding apparatus can receive a bin corresponding to the syntax element via a bitstream, and can determine a context model by using the decoding information of the syntax element and the block to be decoded or an adjacent block or the information of symbols / bins decoded in a previous step, and can derive the value of the syntax element by predicting the occurrence probability of the received bin by the determined context model and performing arithmetic decoding of the bin. Thereafter, based on the determined context model, the context model of the next bin to be decoded can be updated.

[0158] The context model can be assigned and updated for each bin to be context-coded (entropy-coded), and the context model can be indicated based on a context index (ctxIdx: context index) or a context index increment (ctxInc: context index increment). The ctxIdx can be derived based on the ctxInc. Specifically, for example, the ctxIdx indicating the context model for each of the bins to be entropy-coded can be derived as the sum of the ctxInc and a context index offset (ctxIdxOffset: context index offset). For example, the ctxInc can be derived to be different for each bin. The ctxIdxOffset is indicated by the lowest value of the ctxIdx. The ctxIdxOffset is generally a value used for distinguishing from context models for other syntax elements, and the context model for one syntax element can be distinguished or derived based on the ctxInc.

[0159] In the entropy encoding procedure, it can be determined whether to perform encoding via the normal encoding engine or via the bypass encoding engine, whereby the encoding path can be switched. Entropy decoding can execute the same process as entropy encoding in reverse order.

[0160] On the other hand, for example, when a syntax element is bypass decoded, the decoding device can receive the bin corresponding to the syntax element via the bitstream and decode the input bin by applying a uniform probability distribution. In this case, the procedure for deriving the context model of the syntax element and the procedure for updating the context model applied to the bin after decoding can be omitted.

[0161] As described above, the residual samples can be derived as quantized transform coefficients through the conversion and quantization processes. The quantized transform coefficients are also called transform coefficients. In this case, the transform coefficients within the block can be signaled in the form of residual information. The residual information can include syntax or syntax elements related to residual coding. For example, the encoding device can encode the residual information and output it in the form of a bitstream, and the decoding device can decode the residual information from the bitstream to derive the residual (quantized) transform coefficients. The residual information can include syntax elements indicating whether a transform has been applied to the corresponding block, where the position of the last valid transform coefficient within the block is, whether there are valid transform coefficients within the sub-block, or what the magnitude / symbol of the valid transform coefficient is, etc., as described later.

[0162] One party, for example, the prediction unit in the encoding device of FIG. 2 or the prediction unit in the decoding device of FIG. 3, can perform intra prediction. A more detailed description of intra prediction is as follows.

[0163] Intra prediction can indicate a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for the intra prediction of the current block can be derived. The adjacent reference samples of the current block can include samples adjacent to the left boundary of the current block of size nW×nH and a total of 2×nH samples adjacent to the bottom-left, samples adjacent to the top boundary of the current block and a total of 2×nW samples adjacent to the top-right, and 1 sample adjacent to the top-left of the current block. Or, the adjacent reference samples of the current block can also include a plurality of rows of upper adjacent samples and a plurality of columns of left adjacent samples. Also, the adjacent reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom-right of the current block.

[0164] However, some of the adjacent reference samples of the current block may not have been decoded yet or may not be available. In this case, the decoder can substitute samples that are not available with available samples and configure the adjacent reference samples to be used for prediction. Or, the adjacent reference samples to be used for prediction can be configured through interpolation of available samples.

[0165] When an adjacent reference sample is derived, (i) a predicted sample can be induced based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can also be induced based on the reference samples that exist in a specific (predicted) direction with respect to the predicted sample among the adjacent reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angular mode, and in the case of (ii), it can be called a directional mode or an angular mode.

[0166] Also, among the adjacent reference samples, based on the predicted sample of the current block, the predicted sample can also be generated through interpolation between the second adjacent sample and the first adjacent sample located in the direction opposite to the prediction direction of the intra prediction mode of the current block. In the above-mentioned case, it can be called linear interpolation intra prediction (LIP). Also, a chroma prediction sample can be generated based on the luma sample using a linear model. In this case, it can be called the LM (linear model) mode. Also, a temporary prediction sample of the current block is derived based on the filtered adjacent reference samples, and a predicted sample of the current block is derived by performing a weighted sum of at least one reference sample derived by the intra prediction mode and the temporary prediction sample among the existing adjacent reference samples, that is, the unfiltered adjacent reference samples. In the above-mentioned case, it can be called PDPC (Position dependent intra prediction). Also, the intra prediction coding can be performed by selecting the reference sample line with the highest prediction accuracy from among the adjacent multiple reference sample lines of the current block and using the reference sample located in the prediction direction on the corresponding line, and indicating (signaling) the used reference sample line to the decoding device. In the above-mentioned case, it can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Also, the current block can be divided into vertical or horizontal sub-partitions, and intra prediction can be performed based on the same intra prediction mode, and adjacent reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is also applied to the sub-partitions, and by deriving and using adjacent reference samples in units of the sub-partitions, the intra prediction performance can be improved in some cases.Such a prediction method can be called intra sub-partitions (ISP) or ISP-based intra prediction.

[0167] The intra prediction method described above can be called an intra prediction type, distinguished from the intra prediction mode. The intra prediction type can be called by various terms, such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or an additional intra prediction mode, etc.) can include at least one of the LIP, PDPC, MRL, and ISP described above. A general intra prediction method excluding specific intra prediction types such as the LIP, PDPC, MRL, or ISP can be called a normal intra prediction type. The normal intra prediction type can be generally applied when the above specific intra prediction types are not applicable, and prediction can be performed based on the intra prediction mode described above. On the other hand, if necessary, post-processing filtering on the derived prediction sample can also be performed.

[0168] That is, the intra prediction procedure can include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. Also, if necessary, a post-processing filtering step on the derived prediction sample can be performed.

[0169] On the one hand, among the aforementioned intra prediction types, ISP can currently divide a block horizontally or vertically and perform intra prediction in units of the divided blocks. That is, ISP can divide the current block horizontally or vertically to derive sub-blocks, and perform intra prediction for each of the sub-blocks. In this case, encoding / decoding can be performed in units of the divided sub-blocks to generate a reconstructed block, and the reconstructed block can then be used as a reference block for the next divided sub-blocks. Here, the sub-blocks are also called Intra Sub-Partitions.

[0170] For example, when ISP is applied, based on the size of the current block, the current block can be divided into 2 or 4 sub-partitions vertically or horizontally.

[0171] For example, for the application of ISP, a flag indicating whether ISP can be applied can be transmitted in block units. When ISP is applied to the current block, a flag indicating whether the split type is horizontal or vertical, that is, whether the split direction is horizontal or vertical, can be encoded / decoded. The flag indicating whether ISP can be applied can be called an ISP flag, and the ISP flag is indicated by the intra_subpartitions_mode_flag syntax element. Also, the flag indicating the split type can be called an ISP split flag, and the ISP split flag is indicated by the intra_subpartitions_split_flag syntax element.

[0172] For example, information indicating that the ISP is not applied to the current block by an ISP flag or an ISP split flag (IntraSubPartitionsSplitType == ISP_NO_SPLIT), information indicating that it is split horizontally (IntraSubPartitionsSplitType == ISP_HOR_SPLIT), or information indicating that it is split vertically (IntraSubPartitionsSplitType == ISP_VER_SPLIT) is shown. For example, the ISP flag or the ISP split flag is also called ISP-related information regarding sub-partitioning of a block.

[0173] On the other hand, in addition to the intra prediction types described above, ALWIP (affine linear weighted intra prediction) can also be used. The ALWIP is also called LWIP (linear weighted intra prediction), MWIP (matrix weighted intra prediction), or MIP (matrix based intra prediction). When the ALWIP is applied to the current block, i) adjacent reference samples for which an averaging procedure has been performed are used, ii) a matrix-vector-multiplication procedure is performed, and iii) if necessary, a horizontal / vertical interpolation procedure is further performed to derive prediction samples for the current block.

[0174] The intra prediction mode used for the ALWIP can be configured to be different from the LIP, PDPC, MRL, or ISP intra prediction described above, or the intra prediction modes used in normal intra prediction. The intra prediction mode for the ALWIP can be called the ALWIP mode. For example, by the intra prediction mode for the ALWIP, the matrix and offset used in the matrix vector multiplication can be set to be different. Here, the matrix can be called an (affine) weighted metric, and the offset can be called an (affine) offset vector or an (affine) bias vector. In this document, the intra prediction mode for ALWIP is also called the ALWIP mode, the ALWIP intra prediction mode, the LWIP mode, the LWIP intra prediction mode, the MWIP mode, the MWIP intra prediction mode, the MIP mode, or the MIP intra prediction mode. Specific ALWIP methods will be described later.

[0175] FIG. 9 is a diagram for explaining MIP for an 8×8 block.

[0176] To predict samples of a rectangular block with width W and height H, MIP can utilize the samples adjacent to the left boundary and the samples adjacent to the upper boundary of the block. Here, the samples adjacent to the left boundary can indicate the samples located in one line adjacent to the left boundary of the block, and can indicate the restored samples. The samples adjacent to the upper boundary can indicate the samples located in one line adjacent to the upper boundary of the block, and can indicate the restored samples.

[0177] For example, if the restored samples are not available, similar to conventional intra prediction, the restored samples can be generated or derived and utilized.

[0178] The predicted signal (or predicted sample) can be generated based on an averaging process, a matrix vector multiplication process, and a (linear) interpolation process.

[0179] For example, the averaging process is a process of extracting samples outside the boundary through averaging. For example, the samples to be extracted are 4 samples when both the width W and the height H are 4, and 8 samples in other cases. For example, in FIG. 9, bdry left and bdry top can respectively indicate the extracted left - hand side samples and upper - hand side samples.

[0180] For example, the matrix vector multiplication process is a process of inputting the averaged samples and performing matrix vector multiplication. Or, an offset can be further added. For example, in FIG. 9, A k can indicate a matrix, b k can indicate an offset, and bdry red is a reduced signal with respect to the samples extracted through the averaging process. Or, bdry red is bdry left and bdry top for the reduced information. The result is a reduced predicted signal (pred red ) for the set of subsampled samples within the original block.

[0181] For example, the (linear) interpolation process is a process in which a predicted signal is generated at the remaining positions from the predicted signal for the set subsampled by linear interpolation. Here, linear interpolation can indicate single linear interpolation in each direction. For example, in FIG. 9, the reduced predicted signal (pred redLinear interpolation can be performed based on the samples at the boundaries and adjacent boundaries, through which all the predicted samples within the block can be derived.

[0182] For example, the matrix (in FIG. 9, A k ) and the offset vector (in FIG. 9, b k ) required to generate the predicted signal (or predicted block, predicted sample) can be obtained from three sets of matrices (S0, S1, and S2). For example, set S0 can be composed of 18 matrices (A0 i , i = 0, 1,..., 17) and 18 offset vectors (b0 i , i = 0, 1,..., 17). Here, each of the 18 matrices can have 16 rows and 4 columns, and each of the 18 offset vectors can have a size of 16. The matrices and offset vectors of set S0 can be used for blocks of size 4×4. For example, set S1 can be composed of 10 matrices (A1 i , i = 0, 1,..., 9) and 10 offset vectors (b1 i , i = 0, 1,..., 9). Here, each of the 10 matrices can have 16 rows and 8 columns, and each of the 10 offset vectors can have a size of 16. The matrices and offset vectors of set S1 can be used for blocks of size 4×8, 8×4, or 8×8. For example, set S2 can be composed of 6 matrices (A2 i , i = 0, 1,..., 5) and 6 offset vectors (b2 i , i = 0, 1,..., 5). Here, each of the 6 matrices can have 64 rows and 8 columns, and each of the 6 offset vectors can have a size of 64. The matrices and offset vectors of set S2 can be used for all the remaining blocks.

[0183] On the one hand, in one embodiment of this document, LFNST index information can be signaled for a block to which MIP is applied. Alternatively, an encoding device can encode LFNST index information for the transformation of a block to which MIP is applied to generate a bitstream, and a decoding device can parse or decode the bitstream to obtain the LFNST index information for the transformation of the block to which MIP is applied.

[0184] For example, the LFNST index information is information for distinguishing this according to the number of transformations constituting the LFNST transformation set. For example, an optimal LFNST kernel can be selected for a block for which intra prediction with MIP applied based on the LFNST index information. For example, the LFNST index information is indicated by an st_idx syntax element or an lfnst_idx syntax element.

[0185] For example, the LFNST index information (or the st_idx syntax element) can be included in the syntax as shown in the following table.

[0186]

Table 2

[0187]

Table 3

[0188]

Table 4

[0189]

Table 5

[0190] Tables 2 to 5 above show one syntax or information continuously.

[0191] For example, in Tables 2 to 5 above, the information or semantics indicated by the intra_mip_flag syntax element, intra_mip_mpm_flag syntax element, intra_mip_mpm_idx syntax element, intra_mip_mpm_remainder syntax element, or st_idx syntax element are as shown in the following table.

[0192]

Table 6

[0193] For example, the intra_mip_flag syntax element can indicate information regarding whether MIP is applied to luma samples or the current block. Or, for example, the intra_mip_mpm_flag syntax element, intra_mip_mpm_idx syntax element, or intra_mip_mpm_remainder syntax element can indicate information regarding the intra prediction mode applied to the current block when MIP is applied. Or, for example, the st_idx syntax element can indicate information regarding the transform kernel (LFNST kernel) applied to the LFNST for the current block. That is, the st_idx syntax element is information indicating one of the transform kernels within the LFNST transform set. Here, the st_idx syntax element can also be represented by the lfnst_idx syntax element or LFNST index information.

[0194] Figure 10 is a flowchart for explaining how MIP and LFNST are applied.

[0195] On the other hand, in other embodiments of this document, the LFNST index information may not be signaled for the blocks to which MIP is applied. Alternatively, the encoding device can encode video information excluding the LFNST index information for the conversion of the blocks to which MIP is applied to generate a bitstream, and the decoding device can parse or decode the bitstream and execute the conversion process of the blocks without the LFNST index information for the conversion of the blocks to which MIP is applied.

[0196] For example, when the LFNST index information is not signaled, the LFNST index information can be derived as a basic value. For example, the LFNST index information derived as a basic value is a value of 0. For example, the LFNST index information with a value of 0 can indicate that LFNST is not applied to the corresponding block. In this case, there is an effect of reducing the amount of bits for coding the LFNST index information by not transmitting the LFNST index information. Also, it is possible to prevent MIP and LFNST from being applied simultaneously to reduce the complexity, and there is also an effect of reducing latency thereby.

[0197] Referring to FIG. 10, first, it can be determined whether MIP is applied to the corresponding block. That is, it can be determined whether the value of the intra_mip_flag syntax element is 1 or 0 (S1000). For example, if the value of the intra_mip_flag syntax element is 1, it can be regarded as true or yes, indicating that MIP is applied to the corresponding block. Therefore, MIP prediction can be executed for the corresponding block (S1010). That is, MIP prediction can be executed to derive a predicted block for the corresponding block. Thereafter, the inverse primary transform procedure can be executed (S1020), and the intra reconstruction procedure can be executed (S1030). That is, an inverse primary transform can be executed on the transform coefficients obtained from the bitstream to derive a residual block, and a reconstructed block can be generated based on the predicted block by the MIP prediction and the residual block. That is, the LFNST index information for the block to which MIP is applied is not included. Or, LFNST is not applied to the block to which MIP is applied.

[0198] Alternatively, for example, when the value of the intra_mip_flag syntax element is 0, it can be regarded as false or no, indicating that MIP is not applied to the corresponding block. That is, conventional intra prediction can be applied to the corresponding block (S1040). That is, a prediction block for the corresponding block can be derived by performing conventional intra prediction. Thereafter, it can be determined whether LFNST is applied to the corresponding block based on the LFNST index information. That is, it can be determined whether the value of the st_idx syntax element is greater than 0 (S1050). For example, when the value of the st_idx syntax element is greater than 0, an inverse LFNST transform procedure can be executed using the transform kernel indicated by the st_idx syntax element (S1060). Alternatively, when the value of the st_idx syntax element is not greater than 0, this can indicate that LFNST is not applied to the corresponding block, and the inverse LFNST transform procedure is not executed. Thereafter, an inverse primary transform procedure can be executed (S1020), and an intra reconstruction procedure can be executed (S1030). That is, an inverse primary transform can be performed on the transform coefficients obtained from the bitstream to derive a residual block, and a reconstructed block can be generated based on the prediction block by the conventional intra prediction and the residual block.

[0199] To summarize, when MIP is applied, a MIP prediction block can be generated without decoding the LFNST index information, and an inverse primary transform can be applied to the received coefficient to generate a final intra reconstruction signal.

[0200] On the contrary, when MIP is not applied, if the LFNST index information is decoded and the value of its flag (or the LFNST index information or the st_idx syntax element) is greater than 0, the inverse LFNST transform and the inverse first-order transform can be applied to the received coefficient to generate the final intra restoration signal.

[0201] For example, for the above-described procedure, the LFNST index information (or the st_idx syntax element) can be included in the syntax or video information based on the information regarding whether MIP is applied (or the intra_mip_flag syntax element), and can be signaled. Alternatively, the LFNST index information (or the st_idx syntax element) can be selectively configured / parsed / signaled / transmitted / received with reference to the information regarding whether MIP is applied (or the intra_mip_flag syntax element). For example, the LFNST index information is indicated by the st_idx syntax element or the lfnst_idx syntax element.

[0202] For example, the LFNST index information (or the st_idx syntax element) can be included as shown in Table 7.

[0203]

Table 7

[0204] For example, referring to Table 7, the st_idx syntax element can be included based on the intra_mip_flag syntax element. That is, when the value of the intra_mip_flag syntax element is 0 (!intra_mip_flag), the st_idx syntax element can be included.

[0205] Alternatively, for example, the LFNST index information (or the lfnst_idx syntax element) can also be included as shown in Table 8.

[0206]

Table 8

[0207] For example, referring to Table 8, the lfnst_idx syntax element can be included based on the intra_mip_flag syntax element. That is, when the value of the intra_mip_flag syntax element is 0 (!intra_mip_flag), the lfnst_idx syntax element can be included.

[0208] For example, referring to Table 7 or Table 8, the st_idx syntax element or the lfnst_idx syntax element can also be included based on the ISP (Intra Sub-Partitions) related information regarding the sub-partitioning of the block. For example, the ISP related information can include an ISP flag or an ISP split flag, and can indicate information regarding whether sub-partitioning is performed on the block through this. For example, the information regarding whether sub-partitioning is performed is indicated by IntraSubPartitionsSplitType, and ISP_NO_SPLIT can indicate that sub-partitioning is not performed, ISP_HOR_SPLIT can indicate that sub-partitioning is performed in the horizontal direction, and ISP_VER_SPLIT can indicate that sub-partitioning is performed in the vertical direction.

[0209] The residual related information includes the LFNST index information based on the MIP flag and the ISP related information.

[0210] On the other hand, other embodiments of this document can induce the LFNST index information for the block to which MIP is applied without separately signaling it. Alternatively, the encoding device can encode video information excluding the LFNST index information for the conversion of the block to which MIP is applied to generate a bitstream, and the decoding device can parse or decode the bitstream, induce and obtain the LFNST index information for the conversion of the block to which MIP is applied, and based on this, execute the conversion process of the block.

[0211] That is, it is possible to determine an index that divides the conversions that make up the LFNST conversion set through an induction process without decoding the LFNST index information for the corresponding block. Alternatively, it is also possible to determine that a separate optimized conversion kernel is used for the block to which MIP is applied through the induction process. In this case, while selecting the optimal LFNST kernel for the block to which MIP is applied, it is possible to have the effect of reducing the amount of bits for coding this.

[0212] For example, the LFNST index information can be induced based on at least one of the reference line index information for intra prediction, the intra prediction mode information, the block size information, or the MIP applicability information.

[0213] On the other hand, other embodiments of this document can be binary-coded to signal the LFNST index information for the block to which MIP is applied. For example, depending on whether MIP is applied to the current block, the number of applicable LFNST conversions is different, and for this reason, the binary-coding method for the LFNST index information can be selectively switched.

[0214] For example, one LFNST kernel can be used for the block to which MIP is applied, and this kernel is one of the LFNST kernels applied to the blocks to which MIP is not applied. Alternatively, without using the LFNST kernel that has been used for the block to which MIP is applied, a separate kernel optimized for the block to which MIP is applied can be defined and used.

[0215] In this case, for the block to which MIP is applied, by using a reduced number of LFNST kernels compared to the other blocks, the overhead due to signaling the LFNST index information can be reduced, and there is an effect of reducing the complexity.

[0216] For example, the following binary method can be used for the LFNST index information as shown in the table below.

[0217]

Table 9

[0218] Referring to Table 9, for example, the st_idx syntax element can be binary-coded into TR (Truncated Rice) when MIP is not applied to the corresponding block, when intra_mip_flag[][] == false, or when the value of the intra_mip_flag syntax element is 0. For example, in this case, the input parameter cMax can have a value of 2, and cRiceParam can have a value of 0.

[0219] Alternatively, for example, the st_idx syntax element can be binary-coded into FL (Fixed-Length) when MIP is applied to the corresponding block, when intra_mip_flag[][] == true, or when the value of the intra_mip_flag syntax element is 1. For example, in this case, the input parameter cMax can have a value of 1.

[0220] Here, the st_idx syntax element can indicate LFNST index information and can also be represented by the lfnst_idx syntax element.

[0221] On the other hand, other embodiments of this document can signal LFNST-related information for the blocks to which MIP is applied.

[0222] For example, the LFNST index information can include one syntax element and can indicate information on whether LFNST is applied based on one syntax element and information on the type of conversion kernel used for LFNST. In this case, the LFNST index information can be represented by, for example, the st_idx syntax element or the lfnst_idx syntax element.

[0223] Alternatively, for example, the LFNST index information can include one or more syntax elements, and can also indicate information regarding whether the LFNST is applied based on the one or more syntax elements and information regarding the type of conversion kernel used for the LFNST. For example, the LFNST index information can include two syntax elements. In this case, the LFNST index information can include a syntax element indicating information regarding whether the LFNST is applied and a syntax element indicating information regarding the type of conversion kernel used for the LFNST. For example, the information regarding whether the LFNST is applied can be indicated as an LFNST flag and can also be represented by an st_flag syntax element or an lfnst_flag syntax element. Alternatively, for example, the information regarding the type of conversion kernel used for the LFNST can be indicated as a conversion kernel index flag and can also be represented by an st_idx_flag syntax element, an st_kernel_flag syntax element, an lfnst_idx_flag syntax element, or an lfnst_kernel_flag syntax element. For example, when the LFNST index information includes one or more syntax elements as described above, the LFNST index information is also referred to as LFNST-related information.

[0224] For example, the LFNST-related information (e.g., an st_flag syntax element or an st_idx_flag syntax element) can be included as shown in Table 10.

[0225]

Table 10

[0226] On one hand, the blocks to which MIP is applied can use a different number of LFNST conversions (kernels) from the blocks to which MIP is not applied. For example, the blocks to which MIP is applied can also use only one LFNST conversion kernel. For example, the one LFNST conversion kernel is one of the LFNST kernels applied to the blocks to which MIP is not applied. Or, without using the LFNST kernel that has been used for the blocks to which MIP is applied, a separate kernel optimized for the blocks to which MIP is applied can be defined and used.

[0227] In this case, among the LFNST-related information, the information regarding the type of conversion kernel used for LFNST (e.g., conversion kernel index flag) can be selectively signaled depending on whether MIP is applied, and the LFNST-related information at this time can be included, for example, as shown in Table 11.

[0228] [Table 11]

[0229] That is, referring to Table 11, the information regarding the type of conversion kernel used for LFNST (or the st_idx_flag syntax element) can be included based on the information regarding whether MIP is applied to the corresponding block (or the intra_mip_flag syntax element). Or, for example, the st_idx_flag syntax element can be signaled when MIP is not applied to the corresponding block (!intra_mip_flag).

[0230] For example, in Table 10 or Table 11, the information or semantics indicated by the st_flag syntax element or the st_idx_flag syntax element are as follows in the following table.

[0231] [Table 12]

[0232] For example, the st_flag syntax element can indicate information regarding whether a secondary transformation is applied. For example, when the value of the st_flag syntax element is 0, it can indicate that the secondary transformation is not applied, and when it is 1, it can indicate that the secondary transformation is applied. For example, the st_idx_flag syntax element can indicate information regarding the secondary transformation kernel to be applied among two candidate kernels within the selected transformation set.

[0233] For example, the following binary method as shown in the table below can be used for the LFNST related information.

[0234] [Table 13]

[0235] Referring to Table 13, for example, the st_flag syntax element can be binary evolved to FL. For example, in this case, the input parameter cMax can have a value of 1. Or, for example, the st_idx_flag syntax element can be binary evolved to FL. For example, in this case, the input parameter cMax can have a value of 1.

[0236] For example, referring to Table 10 or Table 11, the descriptor of the st_flag syntax element or the st_idx_flag syntax element is ae(v). Here, ae(v) can indicate context-adaptive arithmetic entropy-coding. Or, the syntax element whose descriptor is ae(v) is a context-adaptive arithmetic entropy-coded syntax element. That is, the LFNST-related information (e.g., the st_flag syntax element or the st_idx_flag syntax element) can have context-adaptive arithmetic entropy-coding applied to it. Or, the LFNST-related information (e.g., the st_flag syntax element or the st_idx_flag syntax element) is information or a syntax element to which context-adaptive arithmetic entropy-coding has been applied. Or, the LFNST-related information (e.g., the bin of the bin string of the st_flag syntax element or the st_idx_flag syntax element) can be encoded / decoded based on the aforementioned CABAC, etc. Here, context-adaptive arithmetic entropy-coding can also be referred to as context model-based coding, context coding, or regular coding.

[0237] For example, the context index increment (ctxInc) of LFNST-related information (e.g., st_flag syntax element or st_idx_flag syntax element) or the ctxInc based on the bin position of the st_flag syntax element or st_idx_flag syntax element can be assigned or determined as shown in Table 14. Or, as shown in Table 14, the context model can be selected based on the bin position of the st_flag syntax element or st_idx_flag syntax element. Or, as shown in Table 14, the context model can be selected based on the ctxInc based on the bin position of the st_flag syntax element or st_idx_flag syntax element that is assigned or determined.

[0238]

Table 14

[0239] Referring to Table 14, for example, the st_flag syntax element (the bin or the first bin of the bin string) can use two context models (or ctxIdx), and the context model can be selected based on the ctxInc having a value of 0 or 1. Or, for example, the st_idx_flag syntax element (the bin or the first bin of the bin string) can have bypass coding applied. Or, a uniform probability distribution can be applied for coding.

[0240] For example, the ctxInc of the st_flag syntax element (the bin or the first bin of the bin string) can be determined based on Table 15 below.

[0241]

Table 15

[0242] Referring to Table 15, for example, the ctxInc of the st_flag syntax element (the bin or the first bin of the bin string) can be determined based on the MTS index (or the tu_mts_idx syntax element) or the tree type information (treeType). For example, ctxInc can be derived as 1 when the value of the MTS index is 0 and the tree type is not a single tree. Or ctxInc can be derived as 0 when the value of the MTS index is not 0 or the tree type is a single tree.

[0243] In this case, for the block to which MIP is applied, by using a reduced number of LFNST kernels compared to the other blocks, the overhead caused by signaling the LFNST index information can be reduced, which has the effect of reducing the complexity.

[0244] On the other hand, in other embodiments of this document, the LFNST kernel can be induced and used for the block to which MIP is applied. That is, it can be induced without separately signaling the information regarding the LFNST kernel. Or the encoding device can encode video information excluding the LFNST index information for the conversion of the block to which MIP is applied or the information regarding the conversion kernel (type) used for LFNST to generate a bitstream, and the decoding device can parse or decode the bitstream, induce and obtain the LFNST index information for the conversion of the block to which MIP is applied or the information regarding the conversion kernel used for LFNST, and based on this, execute the conversion process of the block.

[0245] That is, it is possible to determine an index that classifies the conversions that constitute the LFNST conversion set through an induction process without decrypting the LFNST index information for the corresponding block or the information on the conversion kernel used for LFNST. Alternatively, it is also possible to determine that a separate optimized conversion kernel is used for the block to which MIP is applied through the induction process. In this case, while selecting the optimal LFNST kernel for the block to which MIP is applied, it is possible to have the effect of reducing the amount of bits for coding this.

[0246] For example, the LFNST index information or the information on the conversion kernel used for LFNST can be induced based on at least one of the reference line index information for intra prediction, the intra prediction mode information, the block size information, or the MIP applicability information.

[0247] In the embodiments of the present document described above, FL (Fixed-Length) binary evolution can indicate a method of evolving into a fixed length such as a specific number of bits. The specific number of bits can be defined in advance or can be indicated based on cMax. TU (Truncated Unary) binary evolution can indicate a method of evolving into a variable length that uses as many 1s as the number of symbols to be represented and one 0, and when the number of symbols to be represented is the same as the maximum length, no 0 is added. The maximum length can be indicated based on cMax. TR (Truncated Rice) binary evolution can indicate a method of evolving in a form where a prefix and a suffix are concatenated like TU + FL, and uses the maximum length and shift information. When the shift information has a value of 0, it is the same as TU. Here, the maximum length can be indicated based on cMax, and the shift information can be indicated based on cRiceParam.

[0248] FIG. 11 and FIG. 12 schematically show an example of a video / video encoding method and related components according to an embodiment of the present document.

[0249] The method disclosed in FIG. 11 can be executed by the encoding device disclosed in FIG. 2 or FIG. 12. Specifically, for example, S1100 to S1120 in FIG. 11 can be executed by the prediction unit 220 of the encoding device in FIG. 12, S1130 to S1150 in FIG. 11 can be executed by the residual processing unit 230 of the encoding device in FIG. 12, and S1160 in FIG. 11 can be executed by the entropy encoding unit 240 of the encoding device in FIG. 12. Also, although not shown in FIG. 11, in FIG. 12, the prediction unit 220 of the encoding device can derive a prediction sample or prediction-related information, the residual processing unit 230 of the encoding device can derive residual information from the original sample or the prediction sample, and the entropy encoding unit 240 of the encoding device can generate a bitstream from the residual information or the prediction-related information. The method disclosed in FIG. 11 can include the embodiments detailed in this document.

[0250] Referring to FIG. 11, the encoding device can determine the intra prediction type for the current block (S1100) and generate intra prediction type information for the current block based on the intra prediction type (S1110). For example, the encoding device can determine the intra prediction type for the current block in consideration of the rate distortion (RD) cost. The intra prediction type information can indicate information regarding the applicability of a normal intra prediction type using a reference line adjacent to the current block, a multi-reference line (MRL) using a reference line not adjacent to the current block, an intra sub-partitions (ISP) that performs sub-partitioning on the current block, or a matrix-based intra prediction (MIP) that utilizes a matrix.

[0251] For example, the intra prediction type information may include a MIP flag indicating whether MIP is applicable to the current block. Or, for example, the intra prediction type information may include ISP (Intra Sub-Partitions) related information regarding sub-partitioning of ISP for the current block. For example, the ISP related information may include an ISP flag indicating whether ISP is applicable to the current block or an ISP split flag indicating the split direction. Or, for example, the intra prediction type information may also include the MIP flag and the ISP related information. For example, the MIP flag may indicate an intra_mip_flag syntax element. Or, for example, the ISP flag may indicate an intra_subpartitions_mode_flag syntax element, and the ISP split flag may indicate an intra_subpartitions_split_flag syntax element.

[0252] Also, although not shown in FIG. 11, for example, the encoding device can determine an intra prediction mode for the current block and generate intra prediction information for the current block based on the intra prediction mode. For example, the encoding device can determine the intra prediction mode in consideration of the RD cost. The intra prediction mode information can indicate the intra prediction mode applied to the current block among the intra prediction modes. For example, the intra prediction mode can include intra prediction modes numbered from 0 to 66. For example, the 0th intra prediction mode can indicate a planar mode, and the 1st intra prediction mode can indicate a DC mode. Also, the 2nd to 66th intra prediction modes can be indicated as directional or angular intra prediction modes and can indicate the reference direction. Or, the 0th and 1st intra prediction modes can be indicated as non-directional or non-angular intra prediction modes. A detailed description thereof has been elaborated together with FIG. 5.

[0253] For example, the encoding device can generate prediction-related information for the current block, and the prediction-related information can also include intra prediction mode information and / or intra prediction type information.

[0254] The encoding device can derive prediction samples for the current block based on the intra prediction type (S1120). Or, for example, the encoding device can generate the prediction samples based on the intra prediction mode and / or the intra prediction type. Or, the prediction samples can be generated based on the prediction-related information.

[0255] The encoding device can generate residual samples for the current block based on the prediction samples (S1130). For example, the encoding device can generate the residual samples based on the original samples (e.g., the input video signal) and the prediction samples. Or, for example, the encoding device can generate the residual samples based on the difference between the original samples and the prediction samples.

[0256] The encoding device can derive transform coefficients for the current block based on the residual samples (S1140). For example, the encoding device can perform a first-order transform based on the residual samples to derive the transform coefficients. Or, for example, the encoding device can perform a first-order transform based on the residual samples to derive temporary transform coefficients, and can also derive the transform coefficients by applying LFNST to the temporary transform coefficients. For example, when the LFNST is applied, the encoding device can generate the LFNST index information. That is, the LFNST index information can be generated based on the transform kernel used for the derivation of the transform coefficients.

[0257] The encoding device can generate residual-related information based on the conversion coefficient (S1150). For example, the encoding device can perform quantization based on the conversion coefficient to derive the quantized conversion coefficient. Further, the encoding device can generate information regarding the quantized conversion coefficient based on the quantized conversion coefficient. Further, the residual-related information can include information regarding the quantized conversion coefficient.

[0258] The encoding device can encode the intra prediction type information and the residual-related information (S1160). For example, as described above, the residual-related information can include information regarding the quantized conversion coefficient. Also for example, the residual-related information can also include the LFNST index information. Or, for example, the residual-related information may not include the LFNST index information.

[0259] For example, the residual related information may include LFNST index information indicating information related to non-separable conversion for the low-frequency conversion coefficients of the current block based on the MIP flag. Or, for example, the residual related information may include LFNST index information based on the MIP flag or the size of the current block. Or, for example, the residual related information may include LFNST index information based on the MIP flag or information related to the current block. Here, the information related to the current block may include at least one of the size of the current block, tree structure information indicating a single tree or a dual tree, an LFNST available flag, or ISP related information. For example, the MIP flag is one of a plurality of conditions for determining whether the residual related information includes LFNST index information, and the residual related information may also include LFNST index information according to other conditions such as the size of the current block in addition to the MIP flag. However, the following will mainly explain with the MIP flag as the center. Here, the LFNST index information may also be referred to as conversion index information. Or, the LFNST index information may also be represented by the st_idx syntax element or the lfnst_idx syntax element.

[0260] For example, the residual related information can include the LFNST index information based on the MIP flag indicating that the MIP is not applied. Or, for example, the residual related information does not include the LFNST index information based on the MIP flag indicating that the MIP is applied. That is, when the MIP flag indicates that the MIP is applied to the current block (for example, when the value of the intra_mip_flag syntax element is 1), the residual related information may not include the LFNST index information, and when the MIP flag indicates that the MIP is not applied to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the residual related information can include the LFNST index information.

[0261] Or, for example, the residual related information can also include the LFNST index information based on the MIP flag and the ISP related information. For example, when the MIP flag indicates that the MIP is not applied to the current block (for example, when the value of the intra_mip_flag syntax element is 0), referring to the ISP related information (IntraSubPartitionsSplitType), the residual related information can include the LFNST index information. Here, IntraSubPartitionsSplitType can indicate that the ISP is not applied (ISP_NO_SPLIT), applied horizontally (ISP_HOR_SPLIT), or applied vertically (ISP_VER_SPLIT), which can be derived based on the ISP flag or the ISP split flag.

[0262] For example, the LFNST index information can be induced or derived and used by indicating that the MIP is applied to the current block by the MIP flag. In this case, the residual related information does not include the LFNST index information. That is, the encoding device does not signal the LFNST index information. For example, the LFNST index information can be induced or derived and used based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.

[0263] Or, for example, the LFNST index information can also include an LFNST flag indicating whether a non-separable transform is applied to the low-frequency transform coefficients of the current block and / or a transform kernel index flag indicating the transform kernel applied to the current block among the transform kernel candidates. That is, the LFNST index information can indicate information regarding the non-separable transform for the low-frequency transform coefficients of the current block based on one syntax element or one piece of information, but can also be indicated based on two syntax elements or two pieces of information. For example, the LFNST flag can also be represented by the st_flag syntax element or the lfnst_flag syntax element, and the transform kernel index flag can also be represented by the st_idx_flag syntax element, the st_kernel_flag syntax element, the lfnst_idx_flag syntax element, or the lfnst_kernel_flag syntax element. Here, the transform kernel index flag can also be included in the LFNST index information based on the LFNST flag indicating that the non-separable transform is applied and the MIP flag indicating that the MIP is not applied. That is, when the LFNST flag indicates that the non-separable transform is applied and the MIP flag indicates that the MIP is applied, the LFNST index information can include the transform kernel index flag.

[0264] For example, the MIP flag can be used by inducing or deriving the LFNST flag and the transform kernel index flag by indicating that the MIP is applied to the current block. In this case, the residual related information does not include the LFNST flag and the transform kernel index flag. That is, the encoding device does not signal the LFNST flag and the transform kernel index flag. For example, the LFNST flag and the transform kernel index flag can be induced or derived and used based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.

[0265] For example, when the residual related information includes the LFNST index information, the LFNST index information is indicated via binary evolution. For example, based on the MIP flag indicating that the MIP is not applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via Truncated Rice (TR)-based binary evolution, and based on the MIP flag indicating that the MIP is applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via Fixed Length (FL)-based binary evolution. That is, when the MIP flag indicates that the MIP is not applied to the current block (e.g., when the intra_mip_flag syntax element is 0 or false), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via TR-based binary evolution, and when the MIP flag indicates that the MIP is applied to the current block (e.g., when the intra_mip_flag syntax element is 1 or true), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via FL-based binary evolution.

[0266] Alternatively, for example, when the residual related information includes the LFNST index information and the LFNST index information includes the LFNST flag and the transform kernel index flag, the LFNST flag and the transform kernel index flag are indicated via Fixed Length (FL)-based binary evolution.

[0267] For example, the LFNST index information is indicated as a bin string (of bin) via binary evolution as described above, and this can be coded to generate bits, a bit string, or a bit stream.

[0268] For example, the (first) bin of the bin string of the LFNST flag can be coded based on context coding, and the context coding can be executed based on the value of the context index increase or decrease for the LFNST flag. Here, context coding is coding executed based on a context model, and is also called regular coding. Further, the context model is indicated by a context index (ctxIdx), and the context index is indicated based on a context index increase or decrease (ctxInc) and a context index offset (ctxIdxOffset). For example, the value of the context index increase or decrease is indicated by one of the candidates including 0 and 1. For example, the value of the context index increase or decrease can be determined based on an MTS index (for example, an mts_idx syntax element or a tu_mts_idx syntax element) indicating the set of transform kernels used for the current block among the set of transform kernels and tree type information indicating the split structure of the current block. Here, the tree type information can indicate a single tree indicating that the split structures of the luma component and the chroma component of the current block are the same, or a dual tree indicating that the split structures of the luma component and the chroma component of the current block are different from each other.

[0269] For example, the (first) bin of the bin string of the transform kernel index flag can be coded based on bypass coding. Here, bypass coding can also indicate executing context coding based on a uniform probability distribution, and the coding efficiency can be improved by omitting procedures such as the update procedure of context coding.

[0270] Alternatively, although not shown in FIG. 11, for example, the encoding device can also generate a restored sample based on the residual sample and the predicted sample. Further, a restored block and a restored picture can also be derived based on the restored sample.

[0271] For example, the encoding device can encode video information including all or part of the aforementioned information (or syntax elements) to generate a bitstream or encoded information. Alternatively, it can be output in the form of a bitstream. Further, the bitstream or the encoded information can be transmitted to the decoding device via a network or a storage medium. Alternatively, the bitstream or the encoded information can be stored in a computer-readable storage medium, and the bitstream or the encoded information can be generated by the aforementioned video encoding method.

[0272] FIGS. 13 and 14 schematically show an example of a video / video decoding method and related components according to an embodiment of this document.

[0273] The method disclosed in FIG. 13 can be executed by the decoding device disclosed in FIG. 3 or FIG. 14. Specifically, for example, S1300 in FIG. 13 can be executed by the entropy decoding unit 310 of the decoding device in FIG. 14, and S1310 and S1320 in FIG. 13 can be executed by the residual processing unit 320 of the decoding device in FIG. 14. Also, although not shown in FIG. 13, in FIG. 14, prediction-related information or residual information can be derived from the bitstream by the entropy decoding unit 310 of the decoding device, residual samples can be derived from the residual information by the residual processing unit 320 of the decoding device, prediction samples can be derived from the prediction-related information by the prediction unit 330 of the decoding device, and a restored block or a restored picture can be derived from the residual samples or the prediction samples by the addition unit 340 of the decoding device. The method disclosed in FIG. 13 can include the embodiments detailed in this document.

[0274] Referring to FIG. 13, the decoding device can obtain intra prediction type information and residual-related information for the current block from the bitstream (S1300). For example, the decoding device can parse or decode the bitstream to obtain the intra prediction type information or the residual-related information. Here, the bitstream is also referred to as encoded (video) information.

[0275] For example, the decoding device can obtain prediction-related information from the bitstream, and the prediction-related information can also include intra prediction mode information and / or intra prediction type information. For example, the decoding device can generate prediction samples for the current block based on the prediction-related information.

[0276] The intra prediction mode information can indicate the intra prediction mode applied to the current block among the intra prediction modes. For example, the intra prediction mode can include intra prediction modes numbered from 0 to 66. For example, the 0th intra prediction mode can indicate the planar mode, and the 1st intra prediction mode can indicate the DC mode. Also, the 2nd to 66th intra prediction modes can be indicated as directional or angular intra prediction modes and can indicate the reference direction. Or, the 0th and 1st intra prediction modes can be indicated as non-directional or non-angular intra prediction modes. A detailed explanation for this was described in detail together with FIG. 5.

[0277] Also, the intra prediction type information can indicate information regarding the applicability of a normal intra prediction type that uses a reference line adjacent to the current block, an MRL (Multi-Reference Line) that uses a reference line not adjacent to the current block, an ISP (Intra Sub-Partitions) that performs sub-partitioning on the current block, or an MIP (Matrix based Intra Prediction) that utilizes a matrix.

[0278] For example, the decoding device can obtain residual related information from the bitstream. Here, the residual related information can indicate the information used to derive the residual samples and can include information regarding the residual samples, (inverse) transform related information, and / or (inverse) quantization related information. For example, the residual related information can include information regarding the quantized transform coefficients.

[0279] For example, the intra prediction type information can include a MIP flag indicating whether MIP is applied to the current block. Or, for example, the intra prediction type information can include ISP (Intra Sub-Partitions) related information regarding sub-partitioning of ISP for the current block. For example, the ISP related information can include an ISP flag indicating whether ISP is applied to the current block or an ISP split flag indicating the split direction. Or, for example, the intra prediction type information can also include the MIP flag and the ISP related information. For example, the MIP flag can indicate an intra_mip_flag syntax element. Or, for example, the ISP flag can indicate an intra_subpartitions_mode_flag syntax element, and the ISP split flag can indicate an intra_subpartitions_split_flag syntax element.

[0280] For example, the residual related information may include Low Frequency Non Separable Transform (LFNST) index information indicating information related to non-separable transform for the low-frequency transform coefficients of the current block based on the MIP flag. Or, for example, the residual related information may include LFNST index information based on the MIP flag or the size of the current block. Or, for example, the residual related information may include LFNST index information based on the MIP flag or information related to the current block. Here, the information related to the current block may include at least one of the size of the current block, tree structure information indicating a single tree or a dual tree, an LFNST available flag, or ISP related information. For example, the MIP flag is one of a plurality of conditions for determining whether the residual related information includes LFNST index information, and the residual related information may also include LFNST index information according to other conditions such as the size of the current block in addition to the MIP flag. However, the following will mainly explain based on the MIP flag. Here, the LFNST index information may also be referred to as transform index information. Or, the LFNST index information may also be represented by the st_idx syntax element or the lfnst_idx syntax element.

[0281] For example, the residual related information may include the LFNST index information based on the MIP flag indicating that the MIP is not applicable. Or, for example, the residual related information may not include the LFNST index information based on the MIP flag indicating that the MIP is applicable. That is, when the MIP flag indicates that the MIP is applicable to the current block (for example, when the value of the intra_mip_flag syntax element is 1), the residual related information may not include the LFNST index information. When the MIP flag indicates that the MIP is not applicable to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the residual related information may include the LFNST index information.

[0282] Or, for example, the residual related information may also include the LFNST index information based on the MIP flag and the ISP related information. For example, when the MIP flag indicates that the MIP is not applicable to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the residual related information can include the LFNST index information by referring to the ISP related information (IntraSubPartitionsSplitType). Here, IntraSubPartitionsSplitType can indicate that the ISP is not applicable (ISP_NO_SPLIT), is applied horizontally (ISP_HOR_SPLIT), or is applied vertically (ISP_VER_SPLIT), which can be derived based on the ISP flag or the ISP split flag.

[0283] For example, when the MIP flag indicates that the MIP is applied to the current block and the residual related information does not include the LFNST index information, that is, when the LFNST index information is not signaled, the LFNST index information can be derived or deduced. For example, the LFNST index information can be derived based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.

[0284] Alternatively, for example, the LFNST index information can also include an LFNST flag indicating whether a non-separable transform is applied to the low-frequency transform coefficients of the current block and / or a transform kernel index flag indicating the transform kernel applied to the current block among the transform kernel candidates. That is, the LFNST index information can indicate information regarding the non-separable transform for the low-frequency transform coefficients of the current block based on one syntax element or one piece of information, but can also be indicated based on two syntax elements or two pieces of information. For example, the LFNST flag can also be represented by the st_flag syntax element or the lfnst_flag syntax element, and the transform kernel index flag can also be represented by the st_idx_flag syntax element, the st_kernel_flag syntax element, the lfnst_idx_flag syntax element, or the lfnst_kernel_flag syntax element. Here, the transform kernel index flag can also be included in the LFNST index information based on the LFNST flag indicating that the non-separable transform is applied and the MIP flag indicating that the MIP is not applied. That is, when the LFNST flag indicates that the non-separable transform is applied and the MIP flag indicates that the MIP is applied, the LFNST index information can include the transform kernel index flag.

[0285] For example, when the MIP flag indicates that the MIP is applied to the current block, and the residual related information does not include the LFNST flag and the conversion kernel index flag, that is, when the LFNST flag and the conversion kernel index flag are not signaled, the LFNST flag and the conversion kernel index flag can be derived. For example, the LFNST flag and the conversion kernel index flag can be derived based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.

[0286] For example, when the residual related information includes the LFNST index information, the LFNST index information can be derived through binarization. For example, based on the MIP flag indicating that the MIP is not applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through Truncated Rice (TR)-based binarization, and based on the MIP flag indicating that the MIP is applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through Fixed Length (FL)-based binarization. That is, when the MIP flag indicates that the MIP is not applied to the current block (e.g., when the intra_mip_flag syntax element is 0 or false), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through TR-based binarization, and when the MIP flag indicates that the MIP is applied to the current block (e.g., when the intra_mip_flag syntax element is 1 or true), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through FL-based binarization.

[0287] Or, for example, when the residual related information includes the LFNST index information and the LFNST index information includes the LFNST flag and the transform kernel index flag, the LFNST flag and the transform kernel index flag can be derived through Fixed Length (FL)-based binarization.

[0288] For example, the LFNST index information can be derived through the above-described binary evolution, and the bin obtained by parsing or decoding the bitstream can be compared with the candidates, and through this, the LFNST index information can be obtained.

[0289] For example, the (first) bin of the bin string of the LFNST flag can be derived based on context coding, and the context coding can be executed based on the value of the context index increase or decrease for the LFNST flag. Here, context coding is coding executed based on a context model, and is also called regular coding. Further, the context model is indicated by a context index (ctxIdx), and the context index can be derived based on a context index increase or decrease (ctxInc) and a context index offset (ctxIdxOffset). For example, the value of the context index increase or decrease can be derived as one of candidates including 0 and 1. For example, the value of the context index increase or decrease can be derived based on an MTS index (for example, an mts_idx syntax element or a tu_mts_idx syntax element) indicating the set of transform kernels used for the current block among the set of transform kernel sets and tree type information indicating the split structure of the current block. Here, the tree type information can indicate a single tree indicating that the split structures of the luma component and the chroma component of the current block are the same or a dual tree indicating that the split structures of the luma component and the chroma component of the current block are different from each other.

[0290] For example, the (first) bin of the bin string of the conversion kernel index flag can be derived based on bypass coding. Here, bypass coding can indicate that context coding is performed based on a uniform probability distribution, and the coding efficiency can be improved by omitting the update procedure of context coding and the like.

[0291] The decoding device can derive the conversion coefficient for the current block based on the residual related information (S1310). For example, the residual related information can include information about the quantized conversion coefficient, and the decoding device can derive the quantized conversion coefficient for the current block based on the information about the quantized conversion coefficient. For example, the decoding device can perform inverse quantization on the quantized conversion coefficient to derive the conversion coefficient for the current block.

[0292] The decoding device can generate the residual sample of the current block based on the conversion coefficient (S1320). For example, the decoding device can generate the residual sample from the conversion coefficient based on the LFNST index information. For example, when the LFNST index information is included in the residual related information or when the LFNST index information is induced or derived, LFNST can be performed on the conversion coefficient by the LFNST index information, and the corrected conversion coefficient can be derived. Thereafter, the decoding device can generate the residual sample based on the corrected conversion coefficient. Or, for example, when the LFNST index information is not included in the residual related information or indicates that LFNST is not performed, the decoding device can generate the residual sample based on the conversion coefficient without performing LFNST on the conversion coefficient.

[0293] Although not shown in FIG. 13, for example, the decoding device can generate a restored sample based on the prediction sample and the residual sample. Also, for example, a restored block and a restored picture can be derived based on the restored sample.

[0294] For example, the decoding device can decode a bitstream or encoded information to obtain video information including all or part of the aforementioned information (or syntax elements). Also, the bitstream or encoded information can be stored in a computer-readable storage medium, and the aforementioned decoding method can be executed.

[0295] In the foregoing embodiments, the method is described based on a flowchart in a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.

[0296] The method according to the foregoing embodiments of this document can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, display devices, etc.

[0297] In this document, when an embodiment is implemented by software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (for example, information on instructions) or an algorithm can be stored in a digital storage medium.

[0298] In addition, the decoding device and encoding device to which the embodiments of this document are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video dialogue device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a picture phone video device, a transportation means terminal (for example, a vehicle terminal (including an autonomous driving vehicle), an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process video signals or data signals. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0299] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the embodiments of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0300] In addition, the embodiments of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0301] FIG. 15 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0302] Referring to FIG. 15, the content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0303] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate a bitstream, the encoding server can be omitted.

[0304] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0305] The streaming server transmits multimedia data to the user device based on a user request via a web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.

[0306] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0307] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, and a digital signage.

[0308] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.

[0309] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and embodied in a device, and the technical features of the device claims in this specification can be combined and embodied in a method. Also, the technical features of the method claims and the technical features of the device claims in this specification can be combined and embodied in a device, and the technical features of the method claims and the technical features of the device claims in this specification can be combined and embodied in a method.

Claims

1. In a video decoding method executed by a decoding apparatus, obtaining prediction-related information and information on transform coefficients from a bitstream; deriving a predicted sample of a current block based on the prediction-related information; deriving transform coefficients for the current block based on the information on the transform coefficients; generating a residual sample of the current block from the transform coefficients based on LFNSST (low frequency non-separable transform) index information; generating a restored sample of the current block based on the predicted sample and the residual sample, and including: the prediction-related information includes intra prediction type information; the intra prediction type information includes an MIP flag indicating whether MIP (Matrix based Intra Prediction) is applied to the current block; based on the value of the MIP flag not being equal to 1, the information on the transform coefficients includes the LFNSST index information related to a transform kernel applied to LFNSST; based on the value of the MIP flag being equal to 1, it is determined that the LFNSST is not applied to derive the residual sample of the current block; the LFNSST index information is set within the coding unit syntax based on the MIP flag, a method.

2. In a video encoding method executed by an encoding apparatus, determining an intra prediction type for a current block; generating prediction-related information including intra prediction type information based on the intra prediction type for the current block; deriving a predicted sample of the current block based on the intra prediction type; generating a residual sample of the current block based on the predicted sample; deriving transform coefficients for the current block based on the residual sample; generating information on the transform coefficients; encoding the prediction-related information and the information on the transform coefficients, and including: The intra prediction type information includes an MIP flag indicating whether MIP (Matrix based Intra Prediction) is applied to the current block, Based on the value of the MIP flag not being equal to 1, the information regarding the transform coefficients includes LFST index information related to a transform kernel applied to LFST (low frequency non-separable transform), Based on the value of the MIP flag being equal to 1, the LFST is not applied to derive the transform coefficients for the current block, The method in which the LFST index information is set within the coding unit syntax based on the MIP flag.

3. A method for transmitting data for video, A step of obtaining a bitstream for the video, where the bitstream, A step of determining an intra prediction type for a current block, A step of generating prediction-related information including intra prediction type information based on the intra prediction type for the current block, A step of deriving prediction samples for the current block based on the intra prediction type, A step of generating residual samples for the current block based on the prediction samples, A step of deriving transform coefficients for the current block based on the residual samples, A step of generating information regarding the transform coefficients, A step generated based on the steps of encoding the prediction-related information and the information regarding the transform coefficients, A step of transmitting the data including the bitstream, and includes, The intra prediction type information includes an MIP flag indicating whether MIP (Matrix based Intra Prediction) is applied to the current block, Based on the value of the MIP flag not being equal to 1, the information regarding the transform coefficients includes LFST index information related to a transform kernel applied to LFST (low frequency non-separable transform), Based on the value of the MIP flag being equal to 1, the LFST is not applied to derive the transform coefficients for the current block, The method in which the LFNS T index information is set in the coding unit syntax based on the MIP flag.

Citation Information

Patent Citations

  • Transform coding based on matrix-based intra prediction

    WO2020207493A1

  • An encoder, a decoder and corresponding methods harmonzting matrix-based intra prediction and secoundary transform core selection

    WO2020211765A1

  • Encoding device, decoding device, encoding method, and decoding method

    WO2020213677A1