Video Coding Using a Conversion Index
The method enhances video coding efficiency by using transform index information and signaling conversion indices for blocks with MIP, addressing the challenges of high-resolution video compression and diverse image characteristics.
Patent Information
- Application Number
- JP2024117878
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-16
- Filing Date
- 2024-07-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-04-16
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, such as 4K or UHD, has led to higher data transmission and storage costs, as well as the need for efficient compression technologies to handle diverse image characteristics, including those from immersive media like VR and AR.
A method and apparatus for improving video coding efficiency using transform index information, specifically by signaling and coding conversion index information for blocks using matrix-based intra prediction (MIP), and deriving this information to minimize interference between MIP and Low Frequency Non-Separable Transform (LFNST).
This approach enhances overall video compression efficiency, allows efficient signaling of transform indices, and reduces complexity and interference between MIP and LFNST, maintaining optimal coding efficiency.
Smart Images

Figure 0007684491000026 
Figure 0007684491000027 
Figure 0007684491000028
Abstract
Description
Technical Field
[0001] This document relates to video coding technology, and more particularly, to video coding using a transform index.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.
[0004] Accordingly, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.
Summary of the Invention
Means for Solving the Problems
[0005] According to an embodiment of this document, a method and apparatus for increasing video / video coding efficiency are provided.
[0006] According to an embodiment of this document, a video coding method and apparatus using transform index information are provided.
[0007] According to one embodiment of this document, a conversion method and apparatus for a block to which matrix-based intra prediction (MIP) is applied are provided.
[0008] According to one embodiment of this document, a method and apparatus for signaling conversion index information to convert a block to which MIP is applied are provided.
[0009] According to one embodiment of this document, a method and apparatus for signaling conversion index information only for blocks to which MIP is not applied are provided.
[0010] According to one embodiment of this document, a method and apparatus for deriving conversion index information to convert a block to which MIP is applied are provided.
[0011] According to one embodiment of this document, a method and apparatus for binary conversion or entropy coding for conversion index information are provided.
[0012] According to one embodiment of this document, a video / video decoding method executed by a decoding apparatus is provided.
[0013] According to one embodiment of this document, a decoding apparatus for executing video / video decoding is provided.
[0014] According to one embodiment of this document, a video / video encoding method executed by an encoding apparatus is provided.
[0015] According to one embodiment of this document, an encoding apparatus for executing video / video encoding is provided.
[0016] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method disclosed in at least one of the embodiments of this document.
[0017] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded information for causing a decoding device to execute a video / video decoding method disclosed in at least one of the embodiments of this document or storing encoded video / video information.
Advantages of the Invention
[0018] According to this document, the overall video / video compression efficiency can be increased.
[0019] According to this document, the transform index can be efficiently signaled to efficiently (inversely) transform the block to which MIP (Matrix based Intra Prediction) is applied.
[0020] According to this document, the transform index for the block to which MIP is applied can be efficiently coded.
[0021] According to this document, the transform index for the block to which MIP is applied can be induced without separately signaling it.
[0022] According to this document, when both MIP and LFNST (Low Frequency Non-Separable Transform) are applied, the interference between them can be minimized, the optimal coding efficiency can be maintained, and the complexity can be reduced.
[0023] The effects that can be obtained through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that a person having ordinary skill in the related art can understand or induce from this document. Accordingly, the specific effects of this document can include various effects that can be understood or induced from the technical features of this document, rather than being limited to those explicitly described in this document.
Brief Description of Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Best Mode for Carrying Out the Invention
[0025] This document can be modified in various ways, can have various embodiments, and specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the existence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc. is not precluded in advance.
[0026] On the one hand, each component in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each component is implemented by separate hardware or separate software. For example, among the components, two or more components can be combined to form one component, and one component can also be divided into multiple components. As long as the embodiments in which the components are integrated and / or separated do not deviate from the essence of this document, they are included in the scope of rights of this document.
[0027] In this document, “ / ” and “,” are interpreted as “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Additionally, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” also means “at least one of A, B, and / or C.” (In this document,the term “ / ” and “,” should be interpreted to indicate “and / or.” For instance,the expression “A / B” may mean “A and / or B.”Further,“A,B” may mean “A and / or B.”Further,“A / B / C” may mean “at least one of A,B,and / or C.”Also,“A / B / C” may mean “at least one of A,B,and / or C.”)
[0028] Further, in this document, the term “or” should be interpreted to indicate “and / or.” For instance, the expression “A or B” may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term “or” in this document should be interpreted to indicate “additionally or alternatively.”
[0029] In this document, the expression “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Also, in this document, expressions such as “at least one of A or B” and “at least one of A and / or B” may be interpreted in the same way as “at least one of A and B.”
[0030] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0031] Also, the parentheses used in this specification can mean "for example". Specifically, when displayed as "prediction (intra prediction)", "intra prediction" is proposed as an example of "prediction". As another expression, "prediction" in this specification is not limited to "intra prediction", but "intra prediction" is proposed as an example of "prediction". Also, when displayed as "prediction (i.e., intra prediction)", "intra prediction" is proposed as an example of "prediction".
[0032] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.
[0033] In this specification, the technical features individually described in one drawing can be embodied individually or simultaneously.
[0034] FIG. 1 schematically shows an example of a video / coding system to which this document can be applied.
[0035] Referring to FIG. 1, a video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data in a file or streaming form to the receiving device via a digital storage medium or a network.
[0036] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can be composed of a separate device or an external component.
[0037] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like, in which case the video / image capture process can be replaced by a process in which relevant data is generated.
[0038] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0039] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0040] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0041] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0042] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document can be related to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).
[0043] This document presents various embodiments related to video / video coding. Unless otherwise stated, the embodiments can be executed in combination with each other.
[0044] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles.
[0045] A pixel or pel can mean the smallest unit that makes up a picture (or video). Also, the term "sample" can be used as the corresponding term for a pixel. A sample can generally indicate a pixel or a pixel value, can indicate only the pixel / pixel value of the luma component, or can indicate only the pixel / pixel value of the chroma component. Or, a sample can also mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also mean a conversion coefficient in the frequency domain.
[0046] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific region of a picture and information related to that region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In a general case, an M×N block can include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.
[0047] FIG. 2 is a diagram schematically explaining the configuration of a video / video encoding apparatus to which this document can be applied. Hereinafter, a video encoding apparatus can include a video encoding apparatus.
[0048] As shown in FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0049] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied thereafter. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can be divided or partitioned from the final coding unit described above, respectively.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0050] The unit can, in some cases, be used interchangeably with terms such as a block or an area. In general, an M×N block can represent a set such as samples or transform coefficients consisting of M columns and N rows. Samples can generally represent pixels or pixel values, and can represent only the pixels / pixel values of the luma component, or can also represent only the pixels / pixel values of the chroma component. Samples can be used as terms corresponding to pixels or pels for one picture (or image).
[0051] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform a prediction on a processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0052] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or remotely located depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the level of detail of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the adjacent blocks.
[0053] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0054] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of a single block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can also execute an intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.
[0055] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square and can also be applied to a block of a variable size that is not square.
[0056] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in units of NAL (network abstraction layer) units in bitstream form. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.
[0057] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0058] On the other hand, LMCS (luma mapping with chrom ascaling) can also be applied during the picture encoding and / or reconstruction process.
[0059] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 290, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 290 and output in the form of a bitstream.
[0060] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 280. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.
[0061] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.
[0062] FIG. 3 is a diagram schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.
[0063] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoding unit 310, residual processing unit 320, prediction unit 330, addition unit 340, and filtering unit 350 can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.
[0064] If a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device in FIG. 3. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a playback device.
[0065] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the decoding target syntax element information adjacent to and the decoding information of the decoding target block or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and executes arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the information regarding the residual for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.
[0066] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.
[0067] In the inverse conversion unit 322, the conversion coefficient is inversely converted to obtain a residual signal (residual block, residual sample array).
[0068] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0069] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can execute intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in a manner similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.
[0070] The Intra prediction unit 332 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to the current block or at a distance depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The Intra prediction unit 332 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to an adjacent block.
[0071] The Inter prediction unit 331 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks or samples based on the correlation of the motion information between an adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the Inter prediction unit 331 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.
[0072] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor 330. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0073] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and as will be described later, can also be output after filtering, or can be used for inter prediction of the next picture.
[0074] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.
[0075] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0076] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 331. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the picture that has already been restored. The stored motion information can be transmitted to the inter prediction unit 331 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 332.
[0077] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding apparatus 300 can be applied to be the same or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding apparatus 200, respectively.
[0078] On the other hand, as described above, prediction is performed to improve the compression efficiency in performing video coding. Thereby, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same manner in the encoding apparatus and the decoding apparatus, and the encoding apparatus can improve the image coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.
[0079] The residual information can be generated via a conversion and quantization procedure. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thereby signal (via a bitstream) the related residual information to a decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. A decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). A decoding device can generate a restored picture based on the predicted block and the residual block. Also, an encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block, and generate a restored picture based thereon.
[0080] FIG. 4 schematically shows a multiple conversion technique according to this document.
[0081] Referring to FIG. 4, the conversion unit can correspond to the conversion unit in the encoding device of FIG. 2 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of FIG. 2 or the inverse conversion unit in the decoding device of FIG. 3 described above.
[0082] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) within the residual block to derive (primary) conversion coefficients (S410). Such a primary conversion can be called a core transform. Here, the primary conversion can be based on Multiple Transform Selection (MTS), and when multiple conversion is applied in the primary conversion, it can be called a multiple core transform.
[0083] For example, the multiple core transform can be shown as a method of performing conversion by additionally using Discrete Cosine Transform (DCT) type 2 (DCT-II), Discrete Sine Transform (DST) type 7 (DST-VII), DCT type 8 (DCT-VIII), and / or DST type 1 (DST-I). That is, the multiple core transform can be shown as a conversion method for converting a residual signal (or residual block) in the spatial domain into conversion coefficients (or primary conversion coefficients) in the frequency domain based on a plurality of conversion kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary conversion coefficients can be called temporary conversion coefficients on the conversion unit side.
[0084] That is, when an existing conversion method is applied, a conversion from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2 to generate conversion coefficients. However, in contrast, when the multi-core conversion is applied, a conversion from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or primary conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be called conversion types, conversion kernels, or conversion cores. Such DCT / DST conversion types can be defined based on basis functions.
[0085] When the multi-core conversion is executed, a vertical conversion kernel and / or a horizontal conversion kernel for the target block can be selected from among the conversion kernels, and a vertical conversion for the target block can be executed based on the vertical conversion kernel, and a horizontal conversion for the target block can be executed based on the horizontal conversion kernel. Here, the horizontal conversion can indicate a conversion for the horizontal component of the target block, and the vertical conversion can indicate a conversion for the vertical component of the target block. The vertical conversion kernel / horizontal conversion kernel can be adaptively determined based on the prediction mode and / or conversion index of the target block (CU or sub-block) including the residual block.
[0086] Alternatively, for example, when applying MTS to perform a primary transformation, specific basis functions can be set to predetermined values, and when it is a vertical transformation or a horizontal transformation, a mapping relationship for the transformation kernel can be set by combining which basis functions are applied. For example, when representing the horizontal transformation kernel as trTypeHor and the vertical transformation kernel as trTypeVer, trTypeHor or trTypeVer having a value of 0 can be set to DCT2, and trTypeHor or trTypeVer having a value of 1 can be set to DST7. trTypeHor or trTypeVer having a value of 2 can be set to DCT8.
[0087] Alternatively, for example, in order to indicate any one of a number of transformation kernel sets, an MTS index can be encoded and the MTS index information can be signaled to the decoding device. Here, the MTS index can be represented by the tu_mts_idx syntax element or the mts_idx syntax element. For example, when the MTS index is 0, it can indicate that all trTypeHor and trTypeVer values are 0. When the MTS index is 1, it can indicate that all trTypeHor and trTypeVer values are 1. When the MTS index is 2, it can indicate that the trTypeHor value is 2 and the trTypeVer value is 1. When the MTS index is 3, it can indicate that the trTypeHor value is 1 and the trTypeVer value is 2. When the MTS index is 4, it can indicate that all trTypeHor and trTypeVer values are 2. For example, the transformation kernel sets according to the MTS index can be shown as in the following table.
[0088]
Table 1
[0089] The conversion unit can derive a corrected (secondary) conversion coefficient by performing a secondary conversion based on the (primary) conversion coefficient (S420). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion can be shown to convert to a more compressed representation using the correlation existing between the (primary) conversion coefficients.
[0090] For example, the secondary conversion can include a non-separable transform. In this case, the secondary conversion can be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform can be shown to perform a secondary conversion on the (primary) conversion coefficient derived through the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, the conversion can be applied at once without separating the vertical conversion and the horizontal conversion (or independently applying the horizontal conversion and the vertical conversion) to the (primary) conversion coefficient based on the non-separable transform matrix.
[0091] That is, the non-separable secondary transform can be shown to be a conversion method that, without separating the vertical component and the horizontal component of the (primary) conversion coefficient, for example, rearranges a two-dimensional signal (conversion coefficient) into a one-dimensional signal through a specified direction (e.g., row-first direction or column-first direction), and then generates a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix.
[0092] For example, the row-major direction (or order) can be shown by arranging in a column in the order of the first row, the second row, …, the Nth row for an M×N block, and the column-major direction (or order) can be shown by arranging in a column in the order of the first column, the second column, …, the Mth column for an M×N block. Here, M and N can each indicate the width (W) and height (H) of the block, and are all positive integers.
[0093] For example, the non-separable second-order transform can be applied to the top-left region of a block (hereinafter referred to as a transform coefficient block) composed of (first-order) transform coefficients. For example, when the width (W) and height (H) of the transform coefficient block are both 8 or more, an 8×8 non-separable second-order transform can be applied to the top-left 8×8 region of the transform coefficient block. Also, when the width (W) and height (H) of the transform coefficient block are both 4 or more and the width (W) or height (H) of the transform coefficient block is less than 8, a 4×4 non-separable second-order transform can be applied to the top-left min(8, W)×min(8, H) region of the transform coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width (W) or height (H) of the transform coefficient block is 4 or more is satisfied, a 4×4 non-separable second-order transform can also be applied to the top-left min(8, W)×min(8, H) region of the transform coefficient block.
[0094] Specifically, for example, when a 4×4 input block is used, the non-separable second-order transform can be executed as follows.
[0095] The 4×4 input block X is shown as follows.
[0096] [Number]
[0097] For example, the vector form of the said X is shown as follows.
[0098]
Mathematics
[0099] Referring to Equation 2, TIFF0007684491000004.tif64 can represent vector X, and is shown by rearranging the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.
[0100] In this case, the quadratic non-separable transform can be calculated as follows.
[0101]
Mathematics
[0102] Here, TIFF0007684491000006.tif64 can represent the transformation coefficient vector, and T can represent a 16×16 (non-separable) transformation matrix.
[0103] Based on Equation 3, a 16×1-sized TIFF0007684491000007.tif64 can be derived, and the TIFF0007684491000008.tif64 can be re-organized in 4×4 blocks through a scan order (such as horizontal, vertical, or diagonal, etc.). However, the above calculation is only an example, and HyGT (Hypercube-Givens Transform) etc. can also be used for the calculation of non-separable quadratic transforms to reduce the computational complexity of non-separable quadratic transforms.
[0104] On the other hand, for the non-separable quadratic transform, a mode-dependent transformation kernel (or transformation core, transformation type) can also be selected. Here, the mode can include an intra prediction mode and / or an inter prediction mode.
[0105] For example, as described above, the non-separable second-order transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. For example, the 8×8 transform can indicate a transform that can be applied to an 8×8 region included inside the corresponding transform coefficient block when both W and H are the same as or greater than 8, and the 8×8 region is the upper left 8×8 region inside the corresponding transform coefficient block. Similarly, the 4×4 transform can indicate a transform that can be applied to a 4×4 region included inside the corresponding transform coefficient block when both W and H are the same as or greater than 4, and the 4×4 region is the upper left 4×4 region inside the corresponding transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0106] At this time, for mode-based transform kernel selection, two non-separable second-order transform kernels per transform set for non-separable second-order transform can be configured for both the 8×8 transform and the 4×4 transform, and the number of transform sets is four. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.
[0107] However, the size of the transform, the number of sets, and the number of transform kernels in a set are only examples, and sizes other than 8×8 or 4×4 can also be used, or n sets can be configured, and each set can include k transform kernels. Here, n and k are positive integers, respectively.
[0108] For example, the conversion set can be called an NSST set, and the conversion kernels in the NSST set can be called NSST kernels. For example, the selection of a specific set from the conversion set can be performed based on the intra prediction mode of the target block (CU or sub-block).
[0109] For example, the intra prediction mode can include two non-directional or non-angular intra prediction modes and 65 directional or angular intra prediction modes. The non-directional intra prediction mode can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction mode can include 65 intra prediction modes numbered from 2 to 66. However, this is only an example, and the embodiments according to this document can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can also indicate the LM (linear model) mode.
[0110] FIG. 5 exemplarily shows the intra-directional modes of 65 prediction directions.
[0111] Referring to FIG. 5, around the 34th intra prediction mode having the upper left diagonal prediction direction, the intra prediction modes having horizontal directionality and the intra prediction modes having vertical directionality can be distinguished. H and V in FIG. 5 can respectively mean horizontal directionality and vertical directionality, and the numbers from -32 to 32 can indicate a displacement of 1 / 32 unit on the sample grid position. This can indicate an offset with respect to the mode index value.
[0112] For example, the intra prediction modes from 2 to 33 can have a horizontal directionality, and the intra prediction modes from 34 to 66 can have a vertical directionality. On the other hand, the 34th intra prediction mode can be regarded as having neither a strict horizontal nor vertical directionality, but can be classified as belonging to the horizontal directionality from the perspective of determining the conversion set of the secondary conversion. The reason is that for the vertical direction modes symmetric about the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the input data alignment method for the horizontal direction mode is used. Here, transposing the input data can mean that for the two-dimensional block data M×N, the rows become columns and the columns become rows to form N×M data.
[0113] Also, the 18th intra prediction mode and the 50th intra prediction mode can respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode has a left reference pixel and predicts in the upper right direction, so it can be called an upper right diagonal intra prediction mode. Similarly, the 34th intra prediction mode can be called a lower right diagonal intra prediction mode, and the 66th intra prediction mode can be called a lower left diagonal intra prediction mode.
[0114] On the one hand, when it is determined that a specific set is used for the non-separable transform, one of the k transform kernels in the specific set can be selected via the non-separable second-order transform index. For example, the encoding device can derive a non-separable second-order transform index indicating a specific transform kernel based on rate-distortion (RD) checking, and can signal the non-separable second-order transform index to the decoding device. For example, the decoding device can select one of the k transform kernels in the specific set based on the non-separable second-order transform index. For example, an NSST index having a value of 0 can indicate the first non-separable second-order transform kernel, an NSST index having a value of 1 can indicate the second non-separable second-order transform kernel, and an NSST index having a value of 2 can indicate the third non-separable second-order transform kernel. Or, an NSST index having a value of 0 can indicate that the first non-separable second-order transform is not applied to the target block, and an NSST index having a value from 1 to 3 can point to the three transform kernels.
[0115] The transform unit can perform the non-separable second-order transform based on the selected transform kernel and obtain modified (second-order) transform coefficients. As described above, the modified transform coefficients can be derived as the transform coefficients quantized via the quantization unit, and can be encoded and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.
[0116] On the other hand, as described above, when the second-order transform is omitted, the (first-order) transform coefficients that are the output of the first-order (separable) transform can be derived as the transform coefficients quantized via the quantization unit as described above, and can be encoded and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.
[0117] Referring back to FIG. 4, the inverse conversion unit can execute a series of procedures in the reverse order of the procedures executed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficients, performs a secondary (inverse) conversion to derive the (primary) conversion coefficients (S450), and can perform a primary (inverse) conversion on the (primary) conversion coefficients to obtain a residual block (residual samples) (S460). Here, the primary conversion coefficients can be referred to as modified conversion coefficients on the inverse conversion unit side. As described above, the encoding device and / or the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.
[0118] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). For example, the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion is NSST or RST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on the secondary conversion flag parsed or obtained from the bitstream. Or, for example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficients of the residual block.
[0119] The secondary inverse conversion determination unit can determine the secondary inverse conversion. At this time, the secondary inverse conversion determination unit can determine the secondary inverse conversion applied to the current block based on the NSST (or RST) conversion set specified by the intra prediction mode. Or, the secondary conversion determination method can be determined dependently on the primary conversion determination method. Or, various combinations of the primary conversion and the secondary conversion can be determined by the intra prediction mode. For example, the secondary inverse conversion determination unit can also determine the area to which the secondary inverse conversion is applied based on the size of the current block.
[0120] On the one hand, as described above, when the secondary (inverse) transform is omitted, a residual block (residual sample) can be obtained by receiving the (inverse quantized) transform coefficient and performing the primary (separation) inverse transform. As described above, the encoding device and / or the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.
[0121] On the other hand, in this document, in order to reduce the computational amount and memory requirement amount by the non-separable secondary transform, an RST (reduced secondary transform) in which the size of the transform matrix (kernel) is reduced in the concept of NSST can be applied.
[0122] In this document, RST can mean a (simplified) transform performed on the residual sample for the target block based on a transform matrix whose size is reduced by a simplification factor. When this is executed, the amount of computation required at the time of transform can be reduced by reducing the size of the transform matrix. That is, RST can be used to solve the problem of computational complexity that occurs during the transform of a large block or non-separable transform.
[0123] For example, RST can be called by various terms such as reduced transform, reduced secondary transform, reduction transform, simplified transform, or simple transform, and the name by which RST is called is not limited to the listed examples. Or, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it can be called LFNST (Low-Frequency Non-Separable Transform).
[0124] On the other hand, when the second inverse transform is performed based on RST, the inverse transform unit 235 of the encoding device 200 and the inverse transform unit 322 of the decoding device 300 can include an inverse RST unit that derives a modified transform coefficient based on the inverse RST for the transform coefficient, and an inverse primary transform unit that derives a residual sample for the target block based on the inverse primary transform for the modified transform coefficient. The inverse primary transform means the inverse transform of the primary transform applied to the residual. In this document, deriving a transform coefficient based on a transform can mean deriving the transform coefficient by applying the corresponding transform.
[0125] FIG. 6 and FIG. 7 are diagrams for explaining RST according to an embodiment of this document.
[0126] For example, FIG. 6 is a drawing for explaining that a forward reduced transform is applied, and FIG. 7 is a diagram for explaining that an inverse reduced transform is applied. In this document, the target block can indicate the current block where coding is performed, the residual block, or the transform block.
[0127] For example, in RST, an N-dimensional vector can be mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined. Here, N and R are each positive integers, and R is smaller than N. N can mean the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can mean the R / N value. The simplification factor can be called by various terms such as reduced factor, reduction factor, simplified factor, or simple factor. On the other hand, R can be called a reduced coefficient, but in some cases, the simplification factor can also mean R. Also, in some cases, the simplification factor can mean the N / R value.
[0128] For example, the simplification factor or the reduced coefficient can be signaled via a bitstream, but is not limited thereto. For example, there may be cases where predefined values for the simplification factor or the reduced coefficient are stored in each encoding device 200 and decoding device 300, and in this case, the simplification factor or the reduced coefficient is not signaled separately.
[0129] For example, the size (R×N) of the simplified transformation matrix is smaller than the size (N×N) of the normal transformation matrix and can be defined as in the following formula.
[0130]
Equation
[0131] For example, the matrix T within the reduced transform block shown in FIG. 6 can be the matrix T of Equation 4. R×N As shown in FIG. 6, when the simplified transform matrix T R×N is multiplied by the residual samples for the target block, the transform coefficients for the target block can be derived.
[0132] For example, when the size of the block to which the transform is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6 can be expressed by the matrix operation of the following Equation 5. In this case, the memory and multiplication operations can be reduced by approximately 1 / 4 by the simplification factor.
[0133] In this document, matrix operation can be understood as an operation in which a matrix is placed on the left side of a column vector and the matrix and the column vector are multiplied to obtain a column vector.
[0134]
Number
[0135] In Equation 5, r 1 to r 64 can represent the residual samples for the target block. Or, for example, they are the transform coefficients generated by applying a first-order transform. Based on the operation result of Equation 5, the transform coefficients c i for the target block can be derived.
[0136] For example, when R is 16, the transform coefficients c 1 to c 16can be derived. If a normal (regular) conversion is applied instead of RST, and a conversion matrix with a size of 64×64 (N×N) is multiplied by a residual sample with a size of 64×1 (N×1), 64 (N) conversion coefficients for the target block are derived. However, since RST is applied, only 16 (R) conversion coefficients for the target block are derived. Since the total number of conversion coefficients for the target block decreases from N to R and the amount of data transmitted from the encoding device 200 to the decoding device 300 decreases, the transmission efficiency between the encoding device 200 and the decoding device 300 can be increased.
[0137] Considering the size aspect of the conversion matrix, the size of the normal conversion matrix is 64×64 (N×N), while the size of the simplified conversion matrix decreases to 16×64 (R×N). Therefore, when compared with the case of performing a normal conversion, the memory usage can be reduced at a ratio of R / N when performing RST. Also, when compared with the number of multiplication operations N×N when using a normal conversion matrix, when using a simplified conversion matrix, the number of multiplication operations can be reduced at a ratio of R / N (R×N).
[0138] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual sample for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300. The inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and can derive a residual sample for the target block based on an inverse primary conversion for the modified conversion coefficients.
[0139] Inverse RST matrix T according to one embodiment N×RThe size of is N×R, which is smaller than the size N×N of the normal inverse transformation matrix, and is the simplified transformation matrix T shown in Equation 4 R×N is in a transpose relationship with
[0140] The matrix T in the reduced inverse transform block shown in FIG. 7 t is the inverse RST matrix T R×N T can be shown. Here, the superscript T can indicate the transpose. When the inverse RST matrix T R×N T is multiplied by the transformation coefficients for the target block as shown in FIG. 7, the corrected transformation coefficients for the target block or the residual samples for the target block can be derived. The inverse RST matrix T R×N T can also be expressed as (T R×N ) T N×R More specifically, when inverse RST is applied in the second-order inverse transformation, when the inverse RST matrix T
[0141] is multiplied by the transformation coefficients for the target block, the corrected transformation coefficients for the target block can be derived. On the other hand, inverse RST can be applied in the first-order inverse transformation. In this case, when the inverse RST matrix T R×N T is multiplied by the transformation coefficients for the target block, the residual samples for the target block can be derived. R×N T In one embodiment, when the size of the block to which the inverse transformation is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7 can be expressed by the matrix operation as shown in Equation 6 below.
[0142] In one embodiment, when the size of the block to which the inverse transformation is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7 can be expressed by the matrix operation as shown in Equation 6 below.
[0143] [Number]
[0144] In Equation 6, c 1 to c 16 can indicate the conversion coefficient for the target block. Based on the calculation result of Equation 6, r j indicating the corrected conversion coefficient for the target block or the residual sample for the target block can be derived. That is, r 1 to r N indicating the corrected conversion coefficient for the target block or the residual sample for the target block can be derived.
[0145] Considering the size aspect of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64×64 (N×N), while the size of the simplified inverse transformation matrix is reduced to 64×16 (N×R). Therefore, compared with the case of performing the normal inverse transformation, when performing the inverse RST, the memory usage can be reduced at a ratio of R / N. Also, compared with the number of multiplication operations N×N when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced at a ratio of R / N (N×R).
[0146] On the other hand, a conversion set can also be configured and applied to 8×8 RST. That is, the corresponding 8×8 RST can be applied by the conversion set. Since one conversion set is composed of two or three conversion kernels depending on the prediction mode within the screen, it can be configured to select one from a maximum of four conversions including the case where no second-order conversion is applied. The conversion when no second-order conversion is applied is regarded as the identity matrix being applied. When each of the four conversions is assigned an index of 0, 1, 2, or 3 (for example, the 0th index can be assigned to the identity matrix, i.e., the case where no second-order conversion is applied), a syntax element called the NSST index can be signaled for each conversion coefficient block to specify the conversion to be applied. That is, 8×8 NSST can be specified for the 8×8 upper left block via the NSST index, and 8×8 RST can be specified in the RST configuration. 8×8 NSST and 8×8 RST can indicate the conversion that can be applied to the 8×8 region included inside the corresponding conversion coefficient block when all of the W and H of the target block to be converted are the same as or larger than 8, and the 8×8 region is the upper left 8×8 region inside the corresponding conversion coefficient block. Similarly, 4×4 NSST and 4×4 RST can indicate the conversion that can be applied to the 4×4 region included inside the corresponding conversion coefficient block when all of the W and H of the target block are the same as or larger than 4, and the 4×4 region is the upper left 4×4 region inside the corresponding conversion coefficient block.
[0147] One party, for example, the encoding device can derive a bitstream by encoding the value of a syntax element or the quantized value of a transform coefficient related to a residual based on various coding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), or CABAC (context-adaptive binary arithmetic coding). Also, the decoding device can decode the bitstream based on various coding methods such as exponential Golomb coding, CAVLC, or CABAC, and derive the value of a syntax element necessary for video restoration or the quantized value of a transform coefficient related to a residual, etc.
[0148] For example, the aforementioned coding method can be executed as described in the content below.
[0149] FIG. 8 exemplarily shows CABAC (context-adaptive binary arithmetic coding) for encoding a syntax element.
[0150] For example, in the coding process of CABAC, when the input signal is a syntax element that is not a binary value, the encoding device can binarize the value of the input signal to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), the input signal can be used as it is without performing binarization. Here, each binary number 0 or 1 that constitutes a binary value can be called a bin. For example, when the binary string after binarization is 110, each of 1, 1, and 0 can be represented as one bin. The bin for one syntax element can indicate the value of the syntax element. Such binarization can be based on various binarization methods, such as Truncated Rice binarization process or Fixed-length binarization process, and the binarization method for the target syntax element can be defined in advance. The binarization procedure can be executed by a binarization unit within the entropy encoding unit.
[0151] Thereafter, the binarized bin of the syntax element can be input to a regular coding engine or a bypass coding engine. The regular coding engine of the encoding device can assign a context model that reflects a probability value to the corresponding bin, and can encode the corresponding bin based on the assigned context model. The regular coding engine of the encoding device can update the context model for the corresponding bin after performing coding for each bin. The bin coded as described above can be represented as a context-coded bin.
[0152] On the one hand, when the binary bin of the syntax element is input to the bypass coding engine, it can be coded as follows. For example, the bypass coding engine of the encoding device can omit the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after coding. When bypass coding is applied, the encoding device can code the input bin by applying a uniform probability distribution instead of assigning a context model, and through this, the encoding speed can be improved. The bin coded as described above can be represented as a bypass bin.
[0153] Entropy decoding can show the process of executing the same process as the aforementioned entropy encoding in reverse order.
[0154] The decoding device (entropy decoding unit) can decode the encoded video / video information. The video / video information can include partitioning related information, prediction related information (e.g., inter / intra prediction partition information, intra prediction mode information, inter prediction mode information, etc.), residual information, or in-loop filtering related information, etc., or can include various syntax elements related thereto. The entropy coding can be executed in units of syntax elements.
[0155] The decoding device can perform binarization on the target syntax element. Here, the binarization can be based on various binarization methods such as Truncated Rice binarization process or Fixed - length binarization process, and the binarization method for the target syntax element can be defined in advance. The decoding device can derive a usable bin string (bin string candidate) for the usable values of the target syntax element through the binarization procedure. The binarization procedure can be executed by a binarization unit within the entropy decoding unit.
[0156] The decoding device can compare the derived bin string with the usable bin string for the corresponding syntax element while sequentially decoding or parsing each bin for the target syntax element from the input bits in the bitstream. If the derived bin string is the same as one of the usable bin strings, the value corresponding to the bin string is derived as the value of the corresponding syntax element. If not, the next bit in the bitstream can be further parsed and the above - mentioned procedure can be executed again. Through such a process, without using start bits or end bits for specific information (or specific syntax elements) in the bitstream, variable - length bits can be used to signal the corresponding information. Through this, relatively fewer bits can be allocated for lower values, and the overall coding efficiency can be improved.
[0157] The decoding device can decode each bin in the bin string from the bitstream based on an entropy coding technique such as CABAC or CAVLC, either based on a context model or bypass.
[0158] When a syntax element is decoded based on a context model, the decoding apparatus can receive a bin corresponding to the syntax element via a bitstream, and can determine a context model by using the decoding information of the syntax element and the decoding target block or adjacent blocks or the information of symbols / bins decoded in the previous step, and predict the occurrence probability of the received bin by the determined context model and execute arithmetic decoding of the bin to derive the value of the syntax element. Thereafter, based on the determined context model, the context model of the next bin to be decoded can be updated.
[0159] The context model can be allocated and updated for each bin to be context-coded (entropy-coded), and the context model can be indicated based on a context index (ctxIdx: context index) or a context index increment (ctxInc: context index increment). The ctxIdx can be derived based on the ctxInc. Specifically, for example, the ctxIdx indicating the context model for each of the bins to be entropy-coded can be derived as the sum of the ctxInc and a context index offset (ctxIdxOffset: context index offset). For example, the ctxInc can be derived to be different for each bin. The ctxIdxOffset is indicated by the lowest value of the ctxIdx. The ctxIdxOffset is generally a value used for distinguishing from context models for other syntax elements, and the context model for one syntax element can be distinguished or derived based on the ctxInc.
[0160] In the entropy encoding procedure, it can be determined whether to perform encoding via the normal encoding engine or via the bypass encoding engine, whereby the encoding path can be switched. Entropy decoding can execute the same process as entropy encoding in reverse order.
[0161] On the other hand, for example, when a syntax element is bypass decoded, the decoding device can receive the bin corresponding to the syntax element via the bitstream and can decode the input bin by applying a uniform probability distribution. In this case, the procedure for deriving the context model of the syntax element and the procedure for updating the context model applied to the bin after decoding can be omitted.
[0162] As described above, the residual samples can be derived as quantized transform coefficients through the conversion and quantization processes. The quantized transform coefficients are also called transform coefficients. In this case, the transform coefficients within the block can be signaled in the form of residual information. The residual information can include syntax or syntax elements related to residual coding. For example, the encoding device can encode the residual information and output it in the form of a bitstream, and the decoding device can decode the residual information from the bitstream to derive the residual (quantized) transform coefficients. The residual information can include syntax elements indicating whether a transform has been applied to the corresponding block, where the position of the last valid transform coefficient within the block is, whether there are valid transform coefficients within the sub-block, or what the magnitude / symbol of the valid transform coefficients is, etc., as described later.
[0163] One side, for example, the prediction unit in the encoding device of FIG. 2 or the prediction unit in the decoding device of FIG. 3 can perform intra prediction. A more detailed description of the intra prediction is as follows.
[0164] Intra prediction can indicate a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for the intra prediction of the current block can be derived. The adjacent reference samples of the current block can include samples adjacent to the left boundary of the current block of size nW×nH and a total of 2×nH samples adjacent to the bottom - left, samples adjacent to the top boundary of the current block and a total of 2×nW samples adjacent to the top - right, and 1 sample adjacent to the top - left of the current block. Or, the adjacent reference samples of the current block can also include a plurality of columns of upper - adjacent samples and a plurality of rows of left - adjacent samples. Also, the adjacent reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom - right of the current block.
[0165] However, some of the adjacent reference samples of the current block may not have been decoded yet or may not be available. In this case, the decoder can substitute the samples that are not available with samples that are available and configure the adjacent reference samples to be used for prediction. Or, the adjacent reference samples to be used for prediction can be configured through interpolation of the available samples.
[0166] When an adjacent reference sample is derived, (i) a predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can also be derived based on the reference samples that exist in a specific (predicted) direction with respect to the predicted sample among the adjacent reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angular mode, and in the case of (ii), it can be called a directional mode or an angular mode.
[0167] Also, among the adjacent reference samples, based on the predicted sample of the current block, the predicted sample can also be generated through interpolation between the second adjacent sample and the first adjacent sample located in the opposite direction of the prediction direction of the intra prediction mode of the current block. In the case described above, it can be called linear interpolation intra prediction (LIP). Also, a chroma prediction sample can be generated based on luma samples using a linear model. In this case, it can be called the LM (linear model) mode. Also, a temporary prediction sample of the current block is derived based on the filtered adjacent reference samples, and at least one reference sample derived by the intra prediction mode among the existing adjacent reference samples, that is, the non-filtered adjacent reference samples, and the temporary prediction sample are weighted and summed to derive the predicted sample of the current block. In the case described above, it can be called PDPC (Position dependent intra prediction). Also, the reference sample line with the highest prediction accuracy is selected from the adjacent multiple reference sample lines of the current block, and the predicted sample is derived using the reference sample located in the prediction direction on the corresponding line. At this time, the intra prediction coding can be executed by a method of instructing (signaling) the used reference sample line to the decoding device. In the case described above, it can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Also, the current block is divided into vertical or horizontal sub-partitions, and intra prediction is executed based on the same intra prediction mode. The adjacent reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is also applied to the sub-partitions, and by deriving and using the adjacent reference samples in units of the sub-partitions, in some cases, the intra prediction performance can be improved.Such a prediction method can be called intra sub - partitions (ISP) or ISP - based intra prediction.
[0168] The intra prediction method described above can be called an intra prediction type, distinguished from the intra prediction mode. The intra prediction type can be called by various terms, such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or an additional intra prediction mode, etc.) can include at least one of the LIP, PDPC, MRL, and ISP described above. A general intra prediction method excluding specific intra prediction types such as the LIP, PDPC, MRL, or ISP can be called a normal intra prediction type. The normal intra prediction type can be generally applied when the above - mentioned specific intra prediction types are not applicable, and prediction can be performed based on the intra prediction mode described above. On the other hand, if necessary, post - processing filtering for the derived prediction samples can also be performed.
[0169] That is, the intra prediction procedure can include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and an intra prediction mode / type - based prediction sample derivation step. Also, if necessary, a post - processing filtering step for the derived prediction samples can be performed.
[0170] FIG. 9 and FIG. 10 show an example in which a block to which ISP is applied is divided into sub - blocks based on the size of the block.
[0171] On the one hand, among the aforementioned intra prediction types, ISP can currently divide a current block in the horizontal or vertical direction and perform intra prediction in units of the divided blocks. That is, ISP can divide the current block in the horizontal or vertical direction to derive sub-blocks and perform intra prediction for each of the sub-blocks. In this case, encoding / decoding can be performed in units of the divided sub-blocks to generate a reconstructed block, and the reconstructed block can then be used as a reference block for the next divided sub-blocks. Here, the sub-blocks are also called Intra Sub-Partitions.
[0172] For example, when ISP is applied, based on the size of the current block, the current block can be divided into two or four sub-partitions in the vertical or horizontal direction.
[0173] Referring to FIG. 9, when the size of the current block is 4×8 or 8×4, the current block can be divided into two sub-blocks. Also, referring to FIG. 10, when the size of the current block is other than 4×4, 4×8, and 8×4 (i.e., when it is larger than 4×8 or 8×4), the current block can be divided into four sub-blocks.
[0174] For example, for the application of the ISP, the flag indicating whether the ISP can be applied can be transmitted in block units, and when the ISP is currently applied to the block, a flag indicating whether the split type is horizontal or vertical, i.e., whether the split direction is horizontal or vertical, can be encoded / decoded. The flag indicating whether the ISP can be applied can be called the ISP flag, and the ISP flag is indicated by the intra_subpartitions_mode_flag syntax element. Also, the flag indicating the split type can be called the ISP split flag, and the ISP split flag is indicated by the intra_subpartitions_split_flag syntax element.
[0175] For example, information indicating that the ISP is not applied to the current block by the ISP flag or the ISP split flag (IntraSubPartitionsSplitType == ISP_NO_SPLIT), information indicating that it is split horizontally (IntraSubPartitionsSplitType == ISP_HOR_SPLIT), or information indicating that it is split vertically (IntraSubPartitionsSplitType == ISP_VER_SPLIT) is indicated. For example, the ISP flag or the ISP split flag is also called ISP-related information regarding sub-partitioning of the block.
[0176] On the other hand, in addition to the above-described intra prediction types, ALWIP (affine linear weighted intra prediction) can also be used. The ALWIP is also referred to as LWIP (linear weighted intra prediction), MWIP (matrix weighted intra prediction), or MIP (matrix based intra prediction). When the ALWIP is applied to the current block, i) adjacent reference samples for which an averaging procedure has been executed are utilized, ii) a matrix-vector-multiplication procedure is executed, and iii) if necessary, a horizontal / vertical interpolation procedure is further executed to derive prediction samples for the current block.
[0177] The intra prediction mode used for the ALWIP can be configured to be different from the LIP, PDPC, MRL, or ISP intra prediction described above, or the intra prediction modes used in normal intra prediction. The intra prediction mode for the ALWIP can be referred to as the ALWIP mode. For example, depending on the intra prediction mode for the ALWIP, the matrix and offset used in the matrix-vector-multiplication can be set to be different. Here, the matrix can be referred to as an (affine) weighted value matrix, and the offset can be referred to as an (affine) offset vector or an (affine) bias vector. In this document, the intra prediction mode for ALWIP is also referred to as the ALWIP mode, the ALWIP intra prediction mode, the LWIP mode, the LWIP intra prediction mode, the MWIP mode, the MWIP intra prediction mode, the MIP mode, or the MIP intra prediction mode. Specific ALWIP methods will be described later.
[0178] FIG. 11 is a diagram for explaining MIP for an 8×8 block.
[0179] To predict a sample of a rectangular block with width W and height H, MIP can utilize the samples adjacent to the left boundary and the samples adjacent to the upper boundary of the block. Here, the samples adjacent to the left boundary can represent the samples located on one line adjacent to the left boundary of the block and can represent the restored samples. The samples adjacent to the upper boundary can represent the samples located on one line adjacent to the upper boundary of the block and can represent the restored samples.
[0180] For example, if the restored samples are not available, similar to conventional intra prediction, the restored samples can be generated or derived and utilized.
[0181] The prediction signal (or prediction sample) can be generated based on an averaging process, a matrix vector multiplication process, and a (linear) interpolation process.
[0182] For example, the averaging process is a process of extracting samples outside the boundary through averaging. For example, the samples to be extracted are 4 samples when both the width W and the height H are 4, and 8 samples in other cases. For example, in FIG. 11, bdry left and bdry top can respectively represent the extracted left samples and upper samples.
[0183] For example, the matrix vector multiplication process is a process of inputting the averaged samples and performing matrix vector multiplication. Or, an offset can be further added. For example, in FIG. 11, A k can represent the matrix, and b kcan indicate an offset, bdry red is a reduced signal for samples extracted through an averaging process. Or, bdry red is bdry left and bdry top is reduced information for. The result is a reduced prediction signal (pred red ) for a set of subsampled samples within the original block.
[0184] For example, the (linear) interpolation process is a process in which a prediction signal is generated at the remaining positions from the prediction signal for the set subsampled by linear interpolation. Here, linear interpolation can indicate single linear interpolation in each direction. For example, in FIG. 11, a reduced prediction signal (pred red ) displayed in gray within the block and adjacent boundary samples can be used to perform linear interpolation, through which all prediction samples within the block can be derived.
[0185] For example, the matrix (in FIG. 11, A k ) and offset vector (in FIG. 11, b k ) required to generate a prediction signal (or, prediction block, prediction sample) can be obtained from three sets of matrices (S 0 , S 1 , and S 2 ). For example, set S 0 can be composed of 18 matrices (A 0 i , i = 0, 1,..., 17) and 18 offset vectors (b 0 i , i = 0, 1,..., 17). Here, each of the 18 matrices can have 16 rows and 4 columns, and the 18 offset vectors can each have a size of 16. The matrices and offset vectors of set S 0 can be used for 4×4 sized blocks. For example, set S 1consists of 10 matrices (A 1 i , where i = 0, 1, ..., 9) and 10 offset vectors (b 1 i , where i = 0, 1, ..., 9). Here, each of the 10 matrices can have 16 rows and 8 columns, and each of the 10 offset vectors can have a size of 16. The matrices and offset vectors in the set S 1 can be used in blocks of size 4×8, 8×4, or 8×8. For example, the set S 2 can consist of 6 matrices (A 2 i , where i = 0, 1, ..., 5) and 6 offset vectors (b 2 i , where i = 0, 1, ..., 5). Here, each of the 6 matrices can have 64 rows and 8 columns, and each of the 6 offset vectors can have a size of 64. The matrices and offset vectors in the set S 2 can be used in all the remaining blocks.
[0186] On the other hand, one embodiment of this document can signal LFNST index information for the blocks to which MIP is applied. Alternatively, the encoding device can encode the LFNST index information for the conversion of the blocks to which MIP is applied to generate a bitstream, and the decoding device can parse or decode the bitstream to obtain the LFNST index information for the conversion of the blocks to which MIP is applied.
[0187] For example, the LFNST index information is information for classifying this according to the number of conversions that make up the LFNST conversion set. For example, an optimal LFNST kernel can be selected for a block in which intra prediction to which MIP is applied is executed based on the LFNST index information. For example, the LFNST index information is indicated by the st_idx syntax element or the lfnst_idx syntax element.
[0188] For example, the LFNST index information (or the st_idx syntax element) can be included in the syntax as shown in the following table.
[0189] [Table 2]
[0190] [Table 3]
[0191] [Table 4]
[0192] [Table 5]
[0193] Tables 2 to 5 above show one syntax or information continuously.
[0194] For example, in Tables 2 to 5 above, the information or semantics indicated by the intra_mip_flag syntax element, the intra_mip_mpm_flag syntax element, the intra_mip_mpm_idx syntax element, the intra_mip_mpm_remainder syntax element, or the st_idx syntax element are as shown in the following table.
[0195]
Table 6
[0196] For example, the intra_mip_flag syntax element can indicate information on whether MIP is applied to the luma sample or the current block. Or, for example, the intra_mip_mpm_flag syntax element, the intra_mip_mpm_idx syntax element, or the intra_mip_mpm_remainder syntax element can indicate information on the intra prediction mode applied to the current block when MIP is applied. Or, for example, the st_idx syntax element can indicate information on the transform kernel (LFNST kernel) applied to the LFNST for the current block. That is, the st_idx syntax element is information indicating one of the transform kernels within the LFNST transform set. Here, the st_idx syntax element can also be represented by the lfnst_idx syntax element or the LFNST index information.
[0197] FIG. 12 is a flowchart for explaining how MIP and LFNST are applied.
[0198] On the other hand, in other embodiments of this document, the LFNST index information may not be signaled for the blocks to which MIP is applied. Or, the encoding device can encode video information excluding the LFNST index information for the transformation of the blocks to which MIP is applied to generate a bitstream, and the decoding device can parse or decode the bitstream and execute the transformation process of the blocks without the LFNST index information for the transformation of the blocks to which MIP is applied.
[0199] For example, when the LFNST index information is not signaled, the LFNST index information can be derived as a basic value. For example, the LFNST index information derived as a basic value is a 0 value. For example, the LFNST index information with a value of 0 can indicate that LFNST is not applied to the corresponding block. In this case, not transmitting the LFNST index information has the effect of reducing the amount of bits for coding the LFNST index information. Also, it is possible to prevent MIP and LFNST from being applied simultaneously and reduce the complexity, which also has the effect of reducing latency.
[0200] Referring to FIG. 12, first, it can be determined whether MIP is applied to the corresponding block. That is, it can be determined whether the value of the intra_mip_flag syntax element is 1 or 0 (S1200). For example, when the value of the intra_mip_flag syntax element is 1, it can be regarded as true or yes, indicating that MIP is applied to the corresponding block. Therefore, MIP prediction can be performed on the corresponding block (S1210). That is, MIP prediction can be performed to derive a predicted block for the corresponding block. Thereafter, an inverse primary transform procedure can be executed (S1220), and an intra reconstruction procedure can be executed (S1230). That is, an inverse primary transform can be performed on the transform coefficients obtained from the bitstream to derive a residual block, and a reconstructed block can be generated based on the predicted block by the MIP prediction and the residual block. That is, the LFNST index information for the block to which MIP is applied is not included. Or, LFNST is not applied to the block to which MIP is applied.
[0201] Alternatively, for example, when the value of the intra_mip_flag syntax element is 0, it can be regarded as false or no, indicating that MIP is not applied to the corresponding block. That is, conventional intra prediction can be applied to the corresponding block (S1240). That is, a prediction block for the corresponding block can be derived by performing conventional intra prediction. Thereafter, it can be determined whether LFNST is applied to the corresponding block based on the LFNST index information for the corresponding block. That is, it can be determined whether the value of the st_idx syntax element is greater than 0 (S1250). For example, when the value of the st_idx syntax element is greater than 0, an inverse LFNST transform procedure can be executed using the transform kernel indicated by the st_idx syntax element (S1260). Alternatively, when the value of the st_idx syntax element is not greater than 0, this can indicate that LFNST is not applied to the corresponding block, and the inverse LFNST transform procedure is not executed. Thereafter, an inverse primary transform procedure can be executed (S1220), and an intra reconstruction procedure can be executed (S1230). That is, an inverse primary transform can be performed on the transform coefficients obtained from the bitstream to derive a residual block, and a reconstructed block can be generated based on the prediction block by the conventional intra prediction and the residual block.
[0202] To summarize, when MIP is applied, a MIP prediction block can be generated without decoding the LFNST index information, and an inverse primary transform can be applied to the received coefficient to generate a final intra reconstruction signal.
[0203] On the contrary, when MIP is not applied, if the LFNST index information is decoded and the value of its flag (or the LFNST index information or the st_idx syntax element) is greater than 0, an inverse LFNST transform and an inverse first-order transform can be applied to the received coefficient to generate a final intra restoration signal.
[0204] For example, for the above-described procedure, the LFNST index information (or the st_idx syntax element) can be included in the syntax or video information based on the information on whether MIP is applied (or the intra_mip_flag syntax element) and can be signaled. Alternatively, the LFNST index information (or the st_idx syntax element) can be selectively configured / parsed / signaled / transmitted / received with reference to the information on whether MIP is applied (or the intra_mip_flag syntax element). For example, the LFNST index information is indicated by the st_idx syntax element or the lfnst_idx syntax element.
[0205] For example, the LFNST index information (or the st_idx syntax element) can be included as shown in Table 7.
[0206]
Table 7
[0207] For example, referring to Table 7, the st_idx syntax element can be included based on the intra_mip_flag syntax element. That is, when the value of the intra_mip_flag syntax element is 0 (!intra_mip_flag), the st_idx syntax element can be included.
[0208] Alternatively, for example, the LFNST index information (or the lfnst_idx syntax element) can also be included as shown in Table 8.
[0209]
Table 8
[0210] For example, referring to Table 8, the lfnst_idx syntax element can be included based on the intra_mip_flag syntax element. That is, when the value of the intra_mip_flag syntax element is 0 (!intra_mip_flag), the lfnst_idx syntax element can be included.
[0211] For example, referring to Table 7 or Table 8, the st_idx syntax element or the lfnst_idx syntax element can also be included based on the ISP (Intra Sub-Partitions) related information regarding the sub-partitioning of the block. For example, the ISP related information can include an ISP flag or an ISP split flag, and can indicate information regarding whether sub-partitioning is performed on the block through this. For example, the information regarding whether sub-partitioning is performed is indicated by IntraSubPartitionsSplitType, and ISP_NO_SPLIT can indicate that sub-partitioning is not performed, ISP_HOR_SPLIT can indicate that sub-partitioning is performed in the horizontal direction, and ISP_VER_SPLIT can indicate that sub-partitioning is performed in the vertical direction.
[0212] The residual related information includes the LFNST index information based on the MIP flag and the ISP related information.
[0213] On the other hand, other embodiments of this document can derive the LFNST index information for the blocks to which MIP is applied without separately signaling it. Alternatively, the encoding device can encode video information excluding the LFNST index information for the conversion of the blocks to which MIP is applied to generate a bitstream, and the decoding device can parse or decode the bitstream, derive and obtain the LFNST index information for the conversion of the blocks to which MIP is applied, and execute the conversion process of the blocks based on this.
[0214] That is, it is possible to determine an index that divides the conversions that make up the LFNST transform set through an induction process without decoding the LFNST index information for the corresponding blocks. Alternatively, it is also possible to determine that a separate optimized transform kernel is used for the blocks to which MIP is applied through the induction process. In this case, while selecting the optimal LFNST kernel for the blocks to which MIP is applied, it is possible to have the effect of reducing the amount of bits for coding this.
[0215] For example, the LFNST index information can be derived based on at least one of the reference line index information for intra prediction, the intra prediction mode information, the block size information, or the MIP applicability information.
[0216] On the other hand, other embodiments of this document can be binarized to signal the LFNST index information for the blocks to which MIP is applied. For example, depending on whether MIP is applied to the current block, the number of applicable LFNST transforms is different, and for this reason, the binarization method for the LFNST index information can be selectively switched.
[0217] For example, one LFNST kernel can be used for a block to which MIP is applied, and this kernel is one of the LFNST kernels applied to blocks to which MIP is not applied. Or, instead of using the LFNST kernel that has been used for the block to which MIP is applied, a separate kernel optimized for the block to which MIP is applied can be defined and used.
[0218] In this case, for the block to which MIP is applied, by using a reduced number of LFNST kernels compared to the other blocks, the overhead due to signaling the LFNST index information can be reduced, and there is an effect of reducing the complexity.
[0219] For example, the following binary method can be used for the LFNST index information as shown in the table below.
[0220]
Table 9
[0221] Referring to Table 9, for example, the st_idx syntax element can be binary-coded to TR (Truncated Rice) when MIP is not applied to the corresponding block, when intra_mip_flag[][] == false, or when the value of the intra_mip_flag syntax element is 0. For example, in this case, the input parameter cMax can have a value of 2, and cRiceParam can have a value of 0.
[0222] Or, for example, the st_idx syntax element can be binary-coded to FL (Fixed-Length) when MIP is applied to the corresponding block, when intra_mip_flag[][] == true, or when the value of the intra_mip_flag syntax element is 1. For example, in this case, the input parameter cMax can have a value of 1.
[0223] Here, the st_idx syntax element can indicate LFNST index information and can also be represented by the lfnst_idx syntax element.
[0224] On the other hand, other embodiments of this document can signal LFNST-related information for the blocks to which MIP is applied.
[0225] For example, the LFNST index information can include one syntax element and can indicate information on whether LFNST is applied based on one syntax element and information on the type of conversion kernel used for LFNST. In this case, the LFNST index information can be represented by, for example, the st_idx syntax element or the lfnst_idx syntax element.
[0226] Alternatively, for example, the LFNST index information can include one or more syntax elements, and can also indicate information regarding whether the LFNST is applied based on the one or more syntax elements and information regarding the type of conversion kernel used for the LFNST. For example, the LFNST index information can include two syntax elements. In this case, the LFNST index information can include a syntax element indicating information regarding whether the LFNST is applied and a syntax element indicating information regarding the type of conversion kernel used for the LFNST. For example, the information regarding whether the LFNST is applied can be indicated as an LFNST flag, and can also be represented by an st_flag syntax element or an lfnst_flag syntax element. Alternatively, for example, the information regarding the type of conversion kernel used for the LFNST can be indicated as a conversion kernel index flag, and can also be represented by an st_idx_flag syntax element, an st_kernel_flag syntax element, an lfnst_idx_flag syntax element, or an lfnst_kernel_flag syntax element. For example, when the LFNST index information includes one or more syntax elements as described above, the LFNST index information is also referred to as LFNST-related information.
[0227] For example, the LFNST-related information (e.g., the st_flag syntax element or the st_idx_flag syntax element) can be included as shown in Table 10.
[0228]
Table 10
[0229] On one hand, the number of LFNST transforms (kernels) used in the blocks to which MIP is applied can be different from that in the blocks to which MIP is not applied. For example, in the blocks to which MIP is applied, only one LFNST transform kernel can be used. For example, the one LFNST transform kernel is one of the LFNST kernels applied to the blocks to which MIP is not applied. Or, without using the LFNST kernel that has been used for the blocks to which MIP is applied, a separate kernel optimized for the blocks to which MIP is applied can be defined and used.
[0230] In this case, among the LFNST-related information, the information regarding the type of transform kernel used for LFNST (e.g., transform kernel index flag) can be selectively signaled depending on whether MIP is applied, and the LFNST-related information at this time can be included as shown in Table 11, for example.
[0231]
Table 11
[0232] That is, referring to Table 11, the information regarding the type of transform kernel used for LFNST (or the st_idx_flag syntax element) can be included based on the information regarding whether MIP is applied to the corresponding block (or the intra_mip_flag syntax element). Or, for example, the st_idx_flag syntax element can be signaled when MIP is not applied to the corresponding block (!intra_mip_flag).
[0233] For example, in Table 10 or Table 11, the information or semantics indicated by the st_flag syntax element or the st_idx_flag syntax element are as shown in the following table.
[0234]
Table 12
[0235] For example, the st_flag syntax element can indicate information regarding whether a secondary transformation is applied. For example, when the value of the st_flag syntax element is 0, it can indicate that the secondary transformation is not applied, and when it is 1, it can indicate that the secondary transformation is applied. For example, the st_idx_flag syntax element can indicate information regarding the secondary transformation kernel to be applied among two candidate kernels within the selected transformation set.
[0236] For example, the LFNST-related information can utilize a binary evolution method as shown in the following table.
[0237]
Table 13
[0238] Referring to Table 13, for example, the st_flag syntax element can be binary evolved to FL. For example, in this case, the input parameter cMax can have a value of 1. Or, for example, the st_idx_flag syntax element can be binary evolved to FL. For example, in this case, the input parameter cMax can have a value of 1.
[0239] For example, referring to Table 10 or Table 11, the descriptor of the st_flag syntax element or the st_idx_flag syntax element is ae(v). Here, ae(v) can indicate context-adaptive arithmetic entropy-coding. Or, the syntax element whose descriptor is ae(v) is a context-adaptive arithmetic entropy-coded syntax element. That is, the LFNST-related information (for example, the st_flag syntax element or the st_idx_flag syntax element) can have context-adaptive arithmetic entropy-coding applied. Or, the LFNST-related information (for example, the st_flag syntax element or the st_idx_flag syntax element) is information or a syntax element to which context-adaptive arithmetic entropy-coding has been applied. Or, the LFNST-related information (for example, the bin of the bin string of the st_flag syntax element or the st_idx_flag syntax element) can be encoded / decoded based on the aforementioned CABAC, etc. Here, context-adaptive arithmetic entropy-coding can also be indicated as context-model-based coding, context coding, or regular coding.
[0240] For example, the context index increment (ctxInc) of LFNST-related information (e.g., st_flag syntax element or st_idx_flag syntax element) or the ctxInc based on the bin position of the st_flag syntax element or st_idx_flag syntax element can be assigned or determined as shown in Table 14. Or, as shown in Table 14, the context model can be selected based on the bin position of the st_flag syntax element or st_idx_flag syntax element. Or, as shown in Table 14, the context model can be selected based on the ctxInc according to the bin position of the st_flag syntax element or st_idx_flag syntax element that is assigned or determined.
[0241]
Table 14
[0242] Referring to Table 14, for example, the st_flag syntax element (the bin or the first bin of the bin string) can use two context models (or ctxIdx), and the context model can be selected based on the ctxInc having a value of 0 or 1. Or, for example, the st_idx_flag syntax element (the bin or the first bin of the bin string) can have bypass coding applied. Or, a uniform probability distribution can be applied for coding.
[0243] For example, the ctxInc of the st_flag syntax element (the bin or the first bin of the bin string) can be determined based on Table 15 below.
[0244]
Table 15
[0245] Referring to Table 15, for example, the ctxInc of the st_flag syntax element (the bin or the first bin of the bin string) can be determined based on the MTS index (or the tu_mts_idx syntax element) or the tree type information (treeType). For example, ctxInc can be derived as 1 when the value of the MTS index is 0 and the tree type is not a single tree. Or ctxInc can be derived as 0 when the value of the MTS index is not 0 or the tree type is a single tree.
[0246] In this case, for the block to which MIP is applied, by using a reduced number of LFNST kernels compared to the other blocks, the overhead caused by signaling the LFNST index information can be reduced, and there is an effect of reducing the complexity.
[0247] On the other hand, in other embodiments of this document, the LFNST kernel can be induced and used for the block to which MIP is applied. That is, it can be induced without separately signaling the information regarding the LFNST kernel. Or the encoding device can encode video information excluding the LFNST index information for the conversion of the block to which MIP is applied or the information regarding the type of conversion kernel used for LFNST to generate a bitstream, and the decoding device can parse or decode the bitstream, induce and obtain the LFNST index information for the conversion of the block to which MIP is applied or the information regarding the conversion kernel used for LFNST, and based on this, execute the conversion process of the block.
[0248] That is, it is possible to determine an index that classifies the conversions that constitute the LFNST conversion set through an induction process without decrypting the LFNST index information for the corresponding block or the information for the conversion kernel used for the LFNST. Alternatively, it is also possible to determine that a separate optimized conversion kernel is used for the block to which MIP is applied through the induction process. In this case, while selecting the optimal LFNST kernel for the block to which MIP is applied, it is possible to have the effect of reducing the amount of bits for coding this.
[0249] For example, the LFNST index information or the information for the conversion kernel used for the LFNST can be induced based on at least one of the reference line index information for intra prediction, the intra prediction mode information, the block size information, or the MIP applicability information.
[0250] In the embodiments of the present document described above, FL (Fixed-Length) binary evolution can indicate a method of evolving with a length fixed like a specific number of bits, and the specific number of bits can be defined in advance or can be indicated based on cMax. TU (Truncated Unary) binary evolution can indicate a method of evolving with a variable length that uses as many 1s as the number of symbols to be represented and one 0, and when the number of symbols to be represented is the same as the maximum length, no 0 is added, and the maximum length can be indicated based on cMax. TR (Truncated Rice) binary evolution can indicate a method of evolving in a form where a prefix and a suffix are concatenated like TU + FL, and uses the maximum length and shift information, and when the shift information has a value of 0, it is the same as TU. Here, the maximum length can be indicated based on cMax, and the shift information can be indicated based on cRiceParam.
[0251] FIG. 13 and FIG. 14 schematically show an example of a video / video encoding method and related components according to an embodiment of the present document.
[0252] The method disclosed in FIG. 13 can be executed by the encoding device disclosed in FIG. 2 or FIG. 14. Specifically, for example, S1300 to S1310 in FIG. 13 can be executed by the prediction unit 220 of the encoding device in FIG. 14, S1320 to S1330 in FIG. 13 can be executed by the residual processing unit 230 of the encoding device in FIG. 14, and S1340 in FIG. 13 can be executed by the entropy encoding unit 240 of the encoding device in FIG. 14. Also, although not shown in FIG. 13, in FIG. 14, the prediction unit 220 of the encoding device can derive a prediction sample or prediction-related information, the residual processing unit 230 of the encoding device can derive residual information from the original sample or the prediction sample, and the entropy encoding unit 240 of the encoding device can generate a bitstream from the residual information or the prediction-related information. The method disclosed in FIG. 13 can include the embodiments detailed in this document.
[0253] Referring to FIG. 13, the encoding device can perform an intra prediction on the current block to generate a prediction sample of the current block (S1300), and can generate intra prediction type information for the current block based on the performed intra prediction (S1310). For example, the encoding device can determine the intra prediction mode and / or the intra prediction type for the current block in consideration of the rate distortion (RD) cost. The intra prediction mode information is information indicating the determined intra prediction mode, and the intra prediction type information is information indicating the determined intra prediction type. That is, the encoding device can generate intra prediction mode information based on the determined intra prediction mode. Alternatively, the encoding device can generate intra prediction type information based on the determined intra prediction type.
[0254] The intra prediction type information can indicate information regarding applicability such as a normal intra prediction type that uses a reference line adjacent to the current block, an MRL (Multi-Reference Line) that uses a reference line not adjacent to the current block, an ISP (Intra Sub-Partitions) that performs sub-partitioning on the current block, or an MIP (Matrix based Intra Prediction) that utilizes a matrix.
[0255] For example, the intra prediction type information can include an MIP flag indicating whether MIP is applicable to the current block. Or, for example, the intra prediction type information can include ISP (Intra Sub-Partitions) related information regarding sub-partitioning of the ISP for the current block. For example, the ISP related information can include an ISP flag indicating whether ISP is applicable to the current block or an ISP split flag indicating the direction of splitting. Or, for example, the intra prediction type information can also include the MIP flag and the ISP related information. For example, the MIP flag can indicate an intra_mip_flag syntax element. Or, for example, the ISP flag can indicate an intra_subpartitions_mode_flag syntax element, and the ISP split flag can indicate an intra_subpartitions_split_flag syntax element.
[0256] The intra prediction mode information can indicate the intra prediction mode applied to the current block among the intra prediction modes. For example, the intra prediction mode can include intra prediction modes from No. 0 to No. 66. For example, the intra prediction mode No. 0 can indicate the planar mode, and the intra prediction mode No. 1 can indicate the DC mode. Also, the intra prediction modes from No. 2 to No. 66 can be indicated as directional or angular intra prediction modes and can indicate the reference direction. Or, the intra prediction modes No. 0 and No. 1 can be indicated as non-directional or non-angular intra prediction modes. A detailed description thereof was elaborated together with FIG. 5.
[0257] For example, the encoding device can generate prediction-related information for the current block, and the prediction-related information can also include intra prediction mode information and / or intra prediction type information. For example, the encoding device can generate the prediction sample based on the intra prediction mode and / or the intra prediction type. Or, the prediction sample can be generated based on the prediction-related information.
[0258] The encoding device can generate a residual sample of the current block based on the prediction sample (S1320). For example, the encoding device can generate the residual sample based on the original sample (for example, the input video signal) and the prediction sample. Or, for example, the encoding device can generate the residual sample based on the difference between the original sample and the prediction sample.
[0259] The encoding device can generate residual information including information on the conversion coefficients for the current block based on the residual samples (S1330). For example, the encoding device can perform a first - order conversion based on the residual samples to derive the conversion coefficients. Or, for example, the encoding device can perform a first - order conversion based on the residual samples to derive temporary conversion coefficients, and can also derive the conversion coefficients by applying LFNST to the temporary conversion coefficients. For example, when the LFNST is applied, the encoding device can generate the LFNST index information. That is, the LFNST index information can be generated based on the conversion kernel used for generating the information on the conversion coefficients.
[0260] For example, the encoding device can perform quantization based on the conversion coefficients to derive the quantized conversion coefficients. Also, the encoding device can generate information on the quantized conversion coefficients based on the quantized conversion coefficients. Also, the residual information can include information on the quantized conversion coefficients. Here, the information on the quantized conversion coefficients is also simply called information on the conversion coefficients.
[0261] The encoding device can encode video information including intra - prediction type information and residual information (S1340). For example, the residual information can include information on (quantized) conversion coefficients as described above. Also, for example, the video information can include the LFNST index information. Or, for example, the video information may not include the LFNST index information in the residual information.
[0262] For example, the video information may include LFNST index information indicating information related to non-separable conversion for the low-frequency conversion coefficients of the current block based on the MIP flag. Or, for example, the residual related information may include LFNST index information based on the MIP flag or the size of the current block. Or, for example, the residual related information may include LFNST index information based on the MIP flag or information related to the current block. Here, the information related to the current block may include at least one of the size of the current block, tree structure information indicating a single tree or a dual tree, an LFNST available flag, or ISP related information. For example, the MIP flag is one of a plurality of conditions for determining whether the residual related information includes LFNST index information, and in addition to the MIP flag, the residual related information may also include LFNST index information according to other conditions such as the size of the current block. However, the following will focus on the explanation of the MIP flag. Here, the LFNST index information can also be referred to as conversion index information. Or, the LFNST index information can also be represented by the st_idx syntax element or the lfnst_idx syntax element.
[0263] For example, the video information may include the LFNST index information based on the MIP flag indicating that the MIP is not applied. Or, for example, the video information does not include the LFNST index information based on the MIP flag indicating that the MIP is applied. That is, when the MIP flag indicates that the MIP is applied to the current block (for example, when the value of the intra_mip_flag syntax element is 1), the video information may not include the LFNST index information, and when the MIP flag indicates that the MIP is not applied to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the video information may include the LFNST index information.
[0264] Alternatively, for example, the video information may also include the LFNST index information based on the MIP flag and the ISP-related information. For example, when the MIP flag indicates that MIP is not applied to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the residual-related information can include the LFNST index information by referring to the ISP-related information (IntraSubPartitionsSplitType). Here, IntraSubPartitionsSplitType can indicate that ISP is not applied (ISP_NO_SPLIT), applied horizontally (ISP_HOR_SPLIT), or applied vertically (ISP_VER_SPLIT), which can be derived based on the ISP flag or the ISP split flag.
[0265] For example, the LFNST index information can be induced or derived and used by indicating that the MIP is applied to the current block by the MIP flag. In this case, the video information does not include the LFNST index information. That is, the encoding device does not signal the LFNST index information. For example, the LFNST index information can be induced or derived and used based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.
[0266] Alternatively, for example, the LFNST index information may include an LFNST flag indicating whether non-separable transform is applied to the low-frequency transform coefficients of the current block and / or a transform kernel index flag indicating the transform kernel applied to the current block among transform kernel candidates. That is, the LFNST index information can indicate information regarding non-separable transform for the low-frequency transform coefficients of the current block based on one syntax element or one piece of information, but can also be indicated based on two syntax elements or two pieces of information. For example, the LFNST flag can also be represented by a st_flag syntax element or an lfnst_flag syntax element, and the transform kernel index flag can also be represented by a st_idx_flag syntax element, a st_kernel_flag syntax element, an lfnst_idx_flag syntax element, or an lfnst_kernel_flag syntax element. Here, the transform kernel index flag can also be included in the LFNST index information based on the LFNST flag indicating that non-separable transform is applied and the MIP flag indicating that MIP is not applied. That is, when the LFNST flag indicates that non-separable transform is applied and the MIP flag indicates that MIP is applied, the LFNST index information can include the transform kernel index flag.
[0267] For example, by indicating that the MIP flag is applied to the current block, the LFNST flag and the transform kernel index flag can be induced or derived and used. In this case, the video information does not include the LFNST flag and the transform kernel index flag. That is, the encoding device does not signal the LFNST flag and the transform kernel index flag. For example, the LFNST flag and the transform kernel index flag can be induced or derived and used based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.
[0268] For example, when the video information includes the LFNST index information, the LFNST index information is indicated via binary evolution. For example, based on the MIP flag indicating that the MIP is not applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via Truncated Rice (TR)-based binary evolution, and based on the MIP flag indicating that the MIP is applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via Fixed Length (FL)-based binary evolution. That is, when the MIP flag indicates that the MIP is not applied to the current block (e.g., when the intra_mip_flag syntax element is 0 or false), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via TR-based binary evolution, and when the MIP flag indicates that the MIP is applied to the current block (e.g., when the intra_mip_flag syntax element is 1 or true), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) is indicated via FL-based binary evolution.
[0269] Alternatively, for example, when the video information includes the LFNST index information and the LFNST index information includes the LFNST flag and the conversion kernel index flag, the LFNST flag and the conversion kernel index flag are indicated via FL (Fixed Length)-based binary evolution.
[0270] For example, the LFNST index information is indicated by a bin string (of bins) via the binary evolution as described above, and this can be coded to generate bits, a bit string, or a bit stream.
[0271] For example, the (first) bin of the bin string of the LFNST flag can be coded based on context coding, and the context coding can be executed based on the value of the context index increase or decrease for the LFNST flag. Here, context coding is coding executed based on a context model, and is also called regular coding. Also, the context model is indicated by a context index (ctxIdx), and the context index is indicated based on a context index increase or decrease (ctxInc) and a context index offset (ctxIdxOffset). For example, the value of the context index increase or decrease is indicated by one of the candidates including 0 and 1. For example, the value of the context index increase or decrease can be determined based on an MTS index (for example, an mts_idx syntax element or a tu_mts_idx syntax element) indicating the conversion kernel set used for the current block among the conversion kernel sets and tree type information indicating the split structure of the current block. Here, the tree type information can indicate a single tree indicating that the split structures of the luma component and the chroma component of the current block are the same or a dual tree indicating that the split structures of the luma component and the chroma component of the current block are different from each other.
[0272] For example, the (first) bin of the bin string of the conversion kernel index flag can be coded based on bypass coding. Here, bypass coding can also indicate performing context coding based on a uniform probability distribution, and the coding efficiency can be improved by omitting procedures such as the update procedure of context coding.
[0273] Or, although not shown in FIG. 13, for example, the encoding device can also generate a restored sample based on the residual sample and the predicted sample. Also, a restored block and a restored picture can be derived based on the restored sample.
[0274] For example, the encoding device can encode video information including all or part of the aforementioned information (or syntax elements) to generate a bitstream or encoded information. Or, it can be output in the form of a bitstream. Also, the bitstream or the encoded information can be transmitted to a decoding device via a network or a storage medium. Or, the bitstream or the encoded information can be stored in a computer-readable storage medium, and the bitstream or the encoded information can be generated by the aforementioned video encoding method.
[0275] FIGS. 15 and 16 schematically show an example of a video / video decoding method and related components according to an embodiment of this document.
[0276] The method disclosed in FIG. 15 can be executed by the decoding device disclosed in FIG. 3 or FIG. 16. Specifically, for example, S1500 in FIG. 15 can be executed by the entropy decoding unit 310 of the decoding device in FIG. 16, S1510 in FIG. 15 can be executed by the residual processing unit 320 of the decoding device in FIG. 16, and S1520 in FIG. 15 can be executed by the addition unit 340 of the decoding device in FIG. 16. Although not shown in FIG. 15, in FIG. 16, prediction-related information or residual information can be derived from the bitstream by the entropy decoding unit 310 of the decoding device, residual samples can be derived from the residual information by the residual processing unit 320 of the decoding device, prediction samples can be derived from the prediction-related information by the prediction unit 330 of the decoding device, and a restored block or a restored picture can be derived from the residual samples or the prediction samples by the addition unit 340 of the decoding device. The method disclosed in FIG. 15 can include the embodiments detailed in this document.
[0277] Referring to FIG. 15, the decoding device can receive video information including residual information for the current block (S1500). For example, the video information can further include intra prediction type information. For example, the decoding device can pass or decode the bitstream to obtain intra prediction type information or residual-related information. Here, the bitstream is also referred to as encoded (video) information.
[0278] For example, the decoding device can obtain prediction-related information from the bitstream, and the prediction-related information can include intra prediction mode information and / or intra prediction type information. For example, the decoding device can generate prediction samples for the current block based on the prediction-related information.
[0279] The intra prediction mode information can indicate the intra prediction mode applied to the current block among the intra prediction modes. For example, the intra prediction mode can include intra prediction modes numbered from 0 to 66. For example, the intra prediction mode numbered 0 can indicate the planar mode, and the intra prediction mode numbered 1 can indicate the DC mode. Also, the intra prediction modes numbered from 2 to 66 can be indicated as directional or angular intra prediction modes and can indicate the reference direction. Alternatively, the intra prediction modes numbered 0 and 1 can be indicated as non-directional or non-angular intra prediction modes. A detailed description thereof has been elaborated together with FIG. 5.
[0280] Also, the intra prediction type information can indicate information regarding the applicability of, for example, a normal intra prediction type that uses a reference line adjacent to the current block, an MRL (Multi-Reference Line) that uses a reference line not adjacent to the current block, an ISP (Intra Sub-Partitions) that performs sub-partitioning on the current block, or an MIP (Matrix based Intra Prediction) that utilizes a matrix.
[0281] For example, the decoding device can obtain residual information from the bitstream. Here, the residual information can indicate the information used to derive residual samples and can include information regarding the residual samples, (inverse) transform-related information, and / or (inverse) quantization-related information. For example, the residual information can include information regarding (quantized) transform coefficients.
[0282] For example, the intra prediction type information may include a MIP flag indicating whether MIP is applied to the current block. Or, for example, the intra prediction type information may include ISP (Intra Sub-Partitions) related information regarding sub-partitioning of ISP for the current block. For example, the ISP related information may include an ISP flag indicating whether ISP is applied to the current block or an ISP split flag indicating the split direction. Or, for example, the intra prediction type information may also include the MIP flag and the ISP related information. For example, the MIP flag may indicate an intra_mip_flag syntax element. Or, for example, the ISP flag may indicate an intra_subpartitions_mode_flag syntax element, and the ISP split flag may indicate an intra_subpartitions_split_flag syntax element.
[0283] For example, the video information may include Low Frequency Non Separable Transform (LFNST) index information indicating information related to non-separable transform for the low-frequency transform coefficients of the current block based on the MIP flag. Or, for example, the residual related information may include LFNST index information based on the MIP flag or the size of the current block. Or, for example, the residual related information may include LFNST index information based on the MIP flag or information related to the current block. Here, the information related to the current block may include at least one of the size of the current block, tree structure information indicating a single tree or a dual tree, an LFNST available flag, or ISP related information. For example, the MIP flag is one of a plurality of conditions for determining whether the residual related information includes LFNST index information, and the residual related information may also include LFNST index information according to other conditions such as the size of the current block in addition to the MIP flag. However, the following description will focus on the MIP flag. Here, the LFNST index information may also be referred to as transform index information. Or, the LFNST index information may also be represented by the st_idx syntax element or the lfnst_idx syntax element.
[0284] For example, the video information can include the LFNST index information based on the MIP flag indicating that the MIP is not applied. Or, for example, the video information does not include the LFNST index information based on the MIP flag indicating that the MIP is applied. That is, when the MIP flag indicates that the MIP is applied to the current block (for example, when the value of the intra_mip_flag syntax element is 1), the video information may not include the LFNST index information, and when the MIP flag indicates that the MIP is not applied to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the video information can include the LFNST index information.
[0285] Or, for example, the video information can also include the LFNST index information based on the MIP flag and the ISP-related information. For example, when the MIP flag indicates that the MIP is not applied to the current block (for example, when the value of the intra_mip_flag syntax element is 0), the video information can include the LFNST index information by referring to the ISP-related information (IntraSubPartitionsSplitType). Here, IntraSubPartitionsSplitType can indicate that the ISP is not applied (ISP_NO_SPLIT), applied horizontally (ISP_HOR_SPLIT), or applied vertically (ISP_VER_SPLIT), which can be derived based on the ISP flag or the ISP split flag.
[0286] For example, when the MIP flag indicates that the MIP is applied to the current block, if the video information does not include the LFNST index information, that is, when the LFNST index information is not signaled, the LFNST index information can be derived or deduced. For example, the LFNST index information can be derived based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.
[0287] Or, for example, the LFNST index information can also include an LFNST flag indicating whether a non-separable conversion is applied to the low-frequency transform coefficients of the current block and / or a transform kernel index flag indicating the transform kernel applied to the current block among the transform kernel candidates. That is, the LFNST index information can indicate information regarding the non-separable conversion for the low-frequency transform coefficients of the current block based on one syntax element or one piece of information, but can also be indicated based on two syntax elements or two pieces of information. For example, the LFNST flag can also be represented by the st_flag syntax element or the lfnst_flag syntax element, and the transform kernel index flag can also be represented by the st_idx_flag syntax element, the st_kernel_flag syntax element, the lfnst_idx_flag syntax element, or the lfnst_kernel_flag syntax element. Here, the transform kernel index flag can also be included in the LFNST index information based on the LFNST flag indicating that the non-separable conversion is applied and the MIP flag indicating that the MIP is not applied. That is, when the LFNST flag indicates that the non-separable conversion is applied and the MIP flag indicates that the MIP is applied, the LFNST index information can include the transform kernel index flag.
[0288] For example, when the MIP flag indicates that the MIP is applied to the current block, and the video information does not include the LFNST flag and the conversion kernel index flag, that is, when the LFNST flag and the conversion kernel index flag are not signaled, the LFNST flag and the conversion kernel index flag can be induced or derived. For example, the LFNST flag and the conversion kernel index flag can be derived based on at least one of the reference line index information for the current block, the intra prediction mode information of the current block, the size information of the current block, and the MIP flag.
[0289] For example, when the video information includes the LFNST index information, the LFNST index information can be derived through binary evolution. For example, based on the MIP flag indicating that the MIP is not applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through Truncated Rice (TR)-based binary evolution, and based on the MIP flag indicating that the MIP is applied, the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through Fixed Length (FL)-based binary evolution. That is, when the MIP flag indicates that the MIP is not applied to the current block (e.g., when the intra_mip_flag syntax element is 0 or false), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through TR-based binary evolution, and when the MIP flag indicates that the MIP is applied to the current block (e.g., when the intra_mip_flag syntax element is 1 or true), the LFNST index information (e.g., the st_idx syntax element or the lfnst_idx syntax element) can be derived through FL-based binary evolution.
[0290] Alternatively, for example, when the video information includes the LFNST index information and the LFNST index information includes the LFNST flag and the transform kernel index flag, the LFNST flag and the transform kernel index flag can be derived through Fixed Length (FL)-based binary evolution.
[0291] For example, the LFNST index information can derive candidates through the binary evolution as described above, and compare the bins obtained by passing or decoding the bitstream with the candidates, through which the LFNST index information can be acquired.
[0292] For example, the (first) bin of the bin string of the LFNST flag can be derived based on context coding, and the context coding can be executed based on the value of the context index increase or decrease for the LFNST flag. Here, context coding is coding executed based on a context model, and is also called regular coding. Further, the context model is indicated by a context index (ctxIdx), and the context index can be derived based on a context index increase or decrease (ctxInc) and a context index offset (ctxIdxOffset). For example, the value of the context index increase or decrease can be derived as one of candidates including 0 and 1. For example, the value of the context index increase or decrease can be derived based on an MTS index (e.g., mts_idx syntax element or tu_mts_idx syntax element) indicating the set of transform kernels used for the current block among the set of transform kernel sets and tree type information indicating the split structure of the current block. Here, the tree type information can indicate a single tree indicating that the split structures of the luma component and the chroma component of the current block are the same, or a dual tree indicating that the split structures of the luma component and the chroma component of the current block are different from each other.
[0293] For example, the (first) bin of the bin string of the conversion kernel index flag can be derived based on bypass coding. Here, bypass coding can indicate that context coding is performed based on a uniform probability distribution, and the coding efficiency can be improved by omitting procedures such as the update procedure of context coding.
[0294] For example, the decoding device can generate the residual sample from the information regarding the conversion coefficient based on the LFNST index information. For example, the residual information can include information regarding the conversion coefficient of the current block. Or, for example, the residual related information can include information regarding the quantized conversion coefficient, and the decoding device can derive the quantized conversion coefficient for the current block based on the information regarding the quantized conversion coefficient. For example, the decoding device can perform inverse quantization on the quantized conversion coefficient to derive the conversion coefficient for the current block. Or, for example, the decoding device can generate the residual sample using the LFNST index information from the derived conversion coefficient.
[0295] For example, when the LFNST index information is included in the video information or when the LFNST index information is induced or derived, the LFNST can be performed on the conversion coefficient by the LFNST index information, and the corrected conversion coefficient can be derived. Thereafter, the decoding device can generate the residual sample based on the corrected conversion coefficient. Or, for example, when the LFNST index information is not included in the video information or when it indicates not to perform LFNST, the decoding device can generate the residual sample based on the conversion coefficient without performing LFNST on the conversion coefficient.
[0296] The decoding device can generate restored samples for the current block based on the residual samples (S1520). For example, the decoding device can generate restored samples based on the prediction samples and the residual samples. Also, for example, a restored block and a restored picture can be derived based on the restored samples.
[0297] For example, the decoding device can decode a bitstream or encoded information to obtain video information including all or part of the aforementioned information (or syntax elements). Also, the bitstream or encoded information can be stored in a computer-readable storage medium, and the aforementioned decoding method can be executed.
[0298] In the foregoing embodiments, the method is described based on a flowchart in a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.
[0299] The method according to the foregoing embodiments of this document can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in a device that executes video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0300] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (for example, information on instructions) or an algorithm can be stored in a digital storage medium.
[0301] In addition, the decoding device and the encoding device to which the embodiments of this document are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video dialogue device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a picture phone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process video signals or data signals. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0302] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the embodiments of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network.
[0303] In addition, the embodiments of this document can be embodied in a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.
[0304] FIG. 17 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.
[0305] Referring to FIG. 17, the content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0306] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted.
[0307] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0308] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.
[0309] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0310] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, and a digital signage.
[0311] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.
[0312] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.
Claims
1. 1. A video decoding method performed by a decoding device, comprising: receiving image information including prediction mode information and information regarding transformation coefficients for a current block; deriving a prediction sample of the current block by performing intra prediction based on the prediction mode information; generating a residual sample of the current block based on the information about the transform coefficients; generating a reconstructed sample for the current block based on the predicted sample and the residual sample; the image information includes intra-prediction type information for the current block, The intra prediction type information includes a matrix based intra prediction (MIP) flag indicating whether matrix based intra prediction (MIP) is applied to the current block based on whether a tree type of the current block is single tree or dual tree luma; In response to the MIP being applied to the current block because the value of the MIP flag is equal to 1, the predicted samples of the current block are derived by: decreasing neighboring reference samples of the current block; performing matrix multiplication on the decreased neighboring reference samples; and adding an offset to the matrix multiplied values; The image information includes a Low Frequency Non Separable Transform (LFNST) index information associated with one of transform kernels in a LFNST transform set for the current block, based on the fact that intra sub-partitions (ISP) are not applied to the current block, the width and height of the current block being 4 or more, the prediction mode information indicating an intra mode, and the MIP flag indicating that the MIP is not applied; A method according to claim 1, wherein the residual samples are generated based on the information regarding the transform coefficients by using the LFNST index information.
2. 1. A video encoding method performed by an encoding device, comprising: performing intra prediction on a current block to generate a predicted sample of the current block; generating prediction mode information and intra-prediction type information for the current block based on the performed intra-prediction; generating a residual sample of the current block based on the predicted sample; generating information about transform coefficients based on the residual samples; encoding video information including the prediction mode information, the intra-prediction type information, and the information regarding the transform coefficients; The intra prediction type information includes a matrix based intra prediction (MIP) flag indicating whether matrix based intra prediction (MIP) is applied to the current block based on whether a tree type of the current block is single tree or dual tree luma; In response to the value of the MIP flag being equal to 1 because the MIP is applied to the current block, the predicted samples of the current block are derived by: decreasing neighboring reference samples of the current block; performing matrix multiplication on the decreased neighboring reference samples; and adding an offset to the matrix multiplied values; A method, wherein, based on the fact that ISP (intra sub-partitions) is not applied to the current block, the width and height of the current block are 4 or more, the prediction mode information indicating an intra mode, and the MIP flag indicating that the MIP is not applied, the image information includes LFNST index information associated with one of the transform kernels in an LFNST (Low Frequency Non Separable Transform) transform set for the current block.
3. A method for transmitting data for video, comprising the steps of: obtaining a bitstream for the video, the bitstream comprising: performing intra prediction on a current block to generate a predicted sample of the current block; generating prediction mode information and intra-prediction type information for the current block based on the performed intra-prediction; generating a residual sample of the current block based on the predicted sample; generating information about transform coefficients based on the residual samples; encoding video information including the prediction mode information, the intra-prediction type information, and the information about the transform coefficients; transmitting the data including the bitstream; The intra prediction type information includes a matrix based intra prediction (MIP) flag indicating whether matrix based intra prediction (MIP) is applied to the current block based on whether a tree type of the current block is single tree or dual tree luma; In response to the value of the MIP flag being equal to 1 because the MIP is applied to the current block, the predicted samples of the current block are derived by: decreasing neighboring reference samples of the current block; performing matrix multiplication on the decreased neighboring reference samples; and adding an offset to the matrix multiplied values; A method, wherein, based on the fact that ISP (intra sub-partitions) is not applied to the current block, the width and height of the current block are 4 or more, the prediction mode information indicating an intra mode, and the MIP flag indicating that the MIP is not applied, the image information includes LFNST index information associated with one of the transform kernels in an LFNST (Low Frequency Non Separable Transform) transform set for the current block.
Citation Information
Patent Citations
Transform coding based on matrix-based intra prediction
WO2020207493A1
An encoder, a decoder and corresponding methods harmonzting matrix-based intra prediction and secoundary transform core selection
WO2020211765A1
Encoding device, decoding device, encoding method, and decoding method
WO2020213677A1