Coding of information about transform kernel set

By implementing a method for coding an MTS index and context coding in image/video compression, the solution addresses the need for efficient compression of high-resolution and immersive media, enhancing transmission and storage efficiency.

JP2025100588AActive Publication Date: 2025-07-03LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025062442
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-19
Filing Date
2025-04-04
Publication Date
2025-07-03
Estimated Expiration
2040-06-11

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, including immersive media such as VR and AR content, has led to a need for highly efficient image/video compression technologies to reduce transmission and storage costs while maintaining quality.

Method used

A method and apparatus for coding an MTS index in image coding, signaling MTS index information, and context coding or bypass coding for bins of the MTS index, along with a video/image decoding method and encoding device to improve image/video compression efficiency.

Benefits of technology

Enhances overall image/video compression efficiency by efficiently signaling and coding MTS index information, reducing the complexity of the coding system and improving transmission and storage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100588000001_ABST
    Figure 2025100588000001_ABST
Patent Text Reader

Abstract

To provide an image decoding method to be performed by a decoding apparatus.SOLUTION: A method includes the steps of: acquiring image information including MTS index and residual information from a bitstream; deriving quantized transform coefficients for a current block on the basis of the residual information; and deriving the transform coefficients for the current block by executing inverse quantization on the quantized transform coefficients; and generating residual samples of the current block by executing inverse transform on the transform coefficients on the basis of a transform kernel set related to the MTS index. The MTS index indicates a transform kernel set to be applied to the current block among transform kernel set candidates. The transform kernel set includes a transform coefficient to be applied to the current block in a horizontal direction, and a transform coefficient to be applied to the transform block in a vertical direction.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to image coding technology, and more particularly, to coding for information on a set of transform kernels.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Accordingly, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce information on high-resolution and high-quality images / videos having various characteristics as described above.

Summary of the Invention

Means for Solving the Problems

[0005] According to an embodiment of this document, a method and apparatus for increasing image / video coding efficiency are provided.

[0006] According to an embodiment of this document, a method and apparatus for coding an MTS index in image coding are provided.

[0007] According to one embodiment of this document, a method and apparatus for signaling MTS index information are provided.

[0008] According to one embodiment of this document, a method and apparatus for signaling information representing a conversion kernel set applied to a current block among a plurality of conversion kernel sets are provided.

[0009] According to one embodiment of this document, a method and apparatus for context coding or bypass coding for bins of an MTS index are provided.

[0010] According to one embodiment of this document, a video / image decoding method performed by a decoding device is provided.

[0011] According to one embodiment of this document, a decoding device for performing video / image decoding is provided.

[0012] According to one embodiment of this document, a video / image encoding method performed by an encoding device is provided.

[0013] According to one embodiment of this document, an encoding device for performing video / image encoding is provided.

[0014] According to one embodiment of this document, a computer-readable digital storage medium storing encoded video / image information generated by a video / image encoding method disclosed in at least one of the embodiments of this document is provided.

[0015] According to one embodiment of this document, a computer-readable digital storage medium storing encoded information or encoded video / image information for causing a decoding device to perform a video / image decoding method disclosed in at least one of the embodiments of this document is provided.

Advantages of the Invention

[0016] According to this document, the overall image / video compression efficiency can be improved.

[0017] According to this document, MTS index information can be efficiently signaled.

[0018] According to this document, MTS index information can be efficiently coded to reduce the complexity of the coding system.

[0019] The effects obtainable through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Accordingly, the specific effects of this document can include various effects that can be understood or derived from the technical features of this document, rather than being limited to those explicitly described in this document.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Modes for Carrying Out the Invention

[0021] This document can be modified in various ways, can have various embodiments, and specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used merely to explain specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the existence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc. is not precluded in advance.

[0022] On the one hand, each component in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each component is implemented by separate hardware or separate software. For example, among the components, two or more components can be combined to form one component, and one component can also be divided into multiple components. As long as the embodiments in which the components are integrated and / or separated do not deviate from the essence of this document, they are included in the scope of rights of this document.

[0023] Hereinafter, with reference to the accompanying drawings, the preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals are used for the same components in the drawings, and the repeated descriptions for the same components can be omitted.

[0024] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be related to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (for example, the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0025] This document presents various embodiments related to video / image coding, and unless otherwise mentioned, the embodiments can also be combined with each other.

[0026] In this document, "video" can mean a collection of a series of "images" over time. "Picture" generally means a unit representing one image at a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile can contain one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can contain one or more tiles.

[0027] "Pixel" or "pel" can mean the smallest unit that constitutes one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Or, a sample can also mean the pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also mean the conversion coefficient in the frequency domain.

[0028] "Unit" can indicate the basic unit of image processing. A unit can contain at least one of a specific region of a picture and the information related to that region. One unit can contain one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In a general case, an M×N block can contain a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0029] In this document, the terms “ / ” and “,” shall be construed to mean “and / or.” For example, “A / B” shall be construed to mean “A and / or B,” and “A, B” shall be construed to mean “A and / or B.” Additionally, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” also means “at least one of A, B, and / or C.”

[0030] Further, in this document, the term “or” shall be construed to mean “and / or.” For example, “A or B” can mean 1) only “A,” or 2) only “B,” or 3) “A and B.” As another expression, the “or” in this document can mean “additionally or alternatively.”

[0031] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0032] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0033] Also, the parentheses used in this specification can mean "for example". Specifically, when presented as "prediction (intra prediction)", "intra prediction" is proposed as an example of "prediction". As another expression, "prediction" in this specification is not limited to "intra prediction", but "intra prediction" is proposed as an example of "prediction". Also, when presented as "prediction (i.e., intra prediction)", "intra prediction" is proposed as an example of "prediction".

[0034] In this specification, the technical features individually described within one drawing can be embodied individually or simultaneously.

[0035] FIG. 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0036] Referring to FIG. 1, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in a file or streaming form via a digital storage medium or a network.

[0037] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.

[0038] The video source can obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.

[0039] The encoding device can encode input video / images. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0040] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a pre-determined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0041] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0042] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0043] Figure 2 is a diagram schematically explaining the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the video encoding device can include the image encoding device.

[0044] As shown in FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0045] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can be divided or partitioned from the above-described final coding unit, respectively.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0046] The unit can, in some cases, be used interchangeably with terms such as a block or an area. In general, an M×N block can represent a set of samples or transform coefficients, etc., consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or can also represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).

[0047] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input image signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform a prediction on the processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0048] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block according to the prediction mode, or can also be located remotely. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the adjacent blocks.

[0049] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0050] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can also perform intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content images / moving image coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in a manner similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0051] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square and can also be applied to a block of variable size that is not square.

[0052] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of NAL (network abstraction layer) units in the form of a bitstream. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document can be encoded through the above-described encoding procedure and included in the bitstream. The bitstream can be transmitted via a network or stored in a digital storage medium.Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.

[0053] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next processing target block within the current picture, and as will be described later, it can also be used for inter prediction of the next picture after passing through filtering.

[0054] On the other hand, LMCS (luma mapping with chrom ascaling) can also be applied during the picture encoding and / or restoration process.

[0055] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 290, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 290 and output in the form of a bit stream.

[0056] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 280. Through this, when inter prediction is applied, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0057] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0058] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which this document can be applied.

[0059] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0060] If a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device of FIG. 3. For example, the decoding device 300 can derive units / blocks based on block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad-tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a reproducing device.

[0061] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or, picture restoration). The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded and the information of the adjacent and decoded blocks of the block to be decoded or the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, and executes arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the information regarding the residual for which entropy decoding has been performed by the entropy decoding unit 310, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0062] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain transform coefficients.

[0063] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).

[0064] The prediction unit can perform a prediction on the current block and generate a predicted block that includes a prediction sample for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0065] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can perform intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content images / moving images coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in a manner similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0066] The intra prediction unit 332 can predict the current block by referring to samples within the current picture. The samples referred to can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The intra prediction unit 332 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent block.

[0067] The inter prediction unit 331 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 331 can construct a motion information candidate list based on adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0068] The addition unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit 330. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0069] The addition unit 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and as will be described later, can also be output after filtering, or can be used for inter prediction of the next picture.

[0070] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.

[0071] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be sent to the memory 60, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0072] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 331. The memory 360 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 331 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 332.

[0073] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transformation unit 322, and filtering unit 350, etc. of the decoding apparatus 300 can be applied to the prediction unit 220, inverse quantization unit 234, inverse transformation unit 235, and filtering unit 260, etc. of the encoding apparatus 200 so as to be the same or corresponding respectively.

[0074] On the one hand, as described above, prediction is performed to improve the compression efficiency in video coding. By doing so, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in the encoding device and the decoding device. The encoding device can improve the image coding efficiency by signaling information (residual information) regarding the residual between the original block, which is not the original sample value of the original block itself, and the predicted block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.

[0075] The residual information can be generated through the conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thus signal the related residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.

[0076] FIG. 4 schematically shows the multiple conversion techniques according to this document.

[0077] Referring to FIG. 4, the conversion unit can correspond to the conversion unit in the encoding device of FIG. 2 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of FIG. 2 or the inverse conversion unit in the decoding device of FIG. 3 described above.

[0078] The conversion unit can perform a primary conversion based on residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S410). Such a primary conversion can be called a core transform. Here, the primary conversion can be based on Multiple Transform Selection (MTS), and when multiple conversion is applied in the primary conversion, it can be called a multiple core transform.

[0079] For example, the multiple core transform can be shown as a method of performing conversion by additionally using Discrete Cosine Transform (DCT) type 2 (DCT-II), Discrete Sine Transform (DST) type 7 (DST-VII), DCT type 8 (DCT-VIII), and / or DST type 1 (DST-I). That is, the multiple core transform can be shown as a conversion method for converting a residual signal (or residual block) in the spatial domain into conversion coefficients (or primary conversion coefficients) in the frequency domain based on a plurality of conversion kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary conversion coefficients can be called temporary conversion coefficients on the conversion unit side.

[0080] That is, when an existing conversion method is applied, a conversion from the spatial domain to the frequency domain for the residual signal (or residual block) based on DCT type 2 can be applied to generate conversion coefficients. However, in contrast, when the multi-core conversion is applied, a conversion from the spatial domain to the frequency domain for the residual signal (or residual block) can be applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or primary conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be referred to as conversion types, conversion kernels, or conversion cores. Such DCT / DST conversion types can be defined based on basis functions.

[0081] When the multi-core conversion is performed, a vertical conversion kernel and / or a horizontal conversion kernel for the target block can be selected from among the conversion kernels, a vertical conversion for the target block is executed based on the vertical conversion kernel, and a horizontal conversion for the target block can be performed based on the horizontal conversion kernel. Here, the horizontal conversion can indicate a conversion for the horizontal component of the target block, and the vertical conversion can indicate a conversion for the vertical component of the target block. The vertical conversion kernel / horizontal conversion kernel can be adaptively determined based on the prediction mode and / or conversion index of the target block (CU or sub-block) including the residual block.

[0082] Alternatively, for example, when performing a primary transformation by applying MTS, specific basis functions can be set to predetermined values, and when it is a vertical transformation or a horizontal transformation, the mapping relationship for the transformation kernel can be set by combining which basis functions are applied. For example, when the horizontal transformation kernel is represented by trTypeHor and the vertical transformation kernel is represented by trTypeVer, trTypeHor or trTypeVer having a value of 0 can be set to DCT2, and trTypeHor or trTypeVer having a value of 1 can be set to DST7. trTypeHor or trTypeVer having a value of 2 can be set to DCT8.

[0083] Alternatively, for example, in order to indicate any one of a plurality of conversion kernel sets, an MTS index can be encoded and the MTS index information can be signaled to a decoding device. Here, the MTS index can be represented by a tu_mts_idx syntax element or an mts_idx syntax element. For example, when the MTS index is 0, it can indicate that both the trTypeHor and trTypeVer values are 0, and (trTypeHor, trTypeVer) can be (DCT2, DCT2). When the MTS index is 1, it can indicate that both the trTypeHor and trTypeVer values are 1, and (trTypeHor, trTypeVer) can be (DST7, DST7). When the MTS index is 2, it can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, and (trTypeHor, trTypeVer) can be (DCT8, DST7). When the MTS index is 3, it can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and (trTypeHor, trTypeVer) can be (DST7, DCT8). When the MTS index is 4, it can indicate that both the trTypeHor and trTypeVer values are 2, and (trTypeHor, trTypeVer) can be (DCT8, DCT8). For example, the conversion kernel set according to the MTS index can be represented as in the following table.

[0084]

Table 1

[0085] The conversion unit can derive a corrected (secondary) conversion coefficient by performing a secondary conversion based on the (primary) conversion coefficient (S420). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion can be shown to convert to a more compressed representation by utilizing the correlation existing between the (primary) conversion coefficients.

[0086] For example, the secondary conversion can include a non-separable transform. In this case, the secondary conversion can be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform can be shown to be a transform that performs a secondary conversion on the (primary) conversion coefficient derived through the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, the conversion can be applied at once without separating (or independently applying) the vertical conversion and the horizontal conversion to the (primary) conversion coefficient based on the non-separable transform matrix.

[0087] That is, the non-separable secondary transform can be shown to be a conversion method that, without separating the vertical component and the horizontal component of the (primary) conversion coefficient, for example, rearranges a two-dimensional signal (conversion coefficient) into a one-dimensional signal through a specified direction (e.g., row-first direction or column-first direction), and then generates a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix.

[0088] For example, the row-major direction (or order) can be shown by arranging in a column in the order of the first row, the second row, ..., the Nth row for the M×N block, and the column-major direction (or order) can be shown by arranging in a column in the order of the first column, the second column, ..., the Mth column for the M×N block. Here, M and N can each indicate the width (W) and height (H) of the block, and are all positive integers.

[0089] For example, the non-separable second-order transform can be applied to the top-left region of a block (hereinafter referred to as a transform coefficient block) composed of (first-order) transform coefficients. For example, when the width (W) and height (H) of the transform coefficient block are both 8 or more, an 8×8 non-separable second-order transform can be applied to the top-left 8×8 region of the transform coefficient block. Also, when the width (W) and height (H) of the transform coefficient block are both 4 or more and the width (W) or height (H) of the transform coefficient block is less than 8, a 4×4 non-separable second-order transform can be applied to the top-left min(8, W)×min(8, H) region of the transform coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width (W) or height (H) of the transform coefficient block is 4 or more is satisfied, a 4×4 non-separable second-order transform can also be applied to the top-left min(8, W)×min(8, H) region of the transform coefficient block.

[0090] Specifically, for example, when a 4×4 input block is used, the non-separable second-order transform can be performed as follows.

[0091] The 4×4 input block X is shown as follows.

[0092]

Equation

[0093] For example, the vector form of the above X is shown as follows.

[0094]

Mathematics

[0095] Referring to Equation 2, JPEG2025100588000005.jpg97 can represent vector X, and is shown by rearranging the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.

[0096] In this case, the second non-separable transform can be calculated as follows.

[0097]

Mathematics

[0098] Here, JPEG2025100588000007.jpg76 can represent the transform coefficient vector, and T can represent a 16×16 (non-separable) transform matrix.

[0099] Based on Equation 3, a 16×1 size of JPEG2025100588000008.jpg76 can be derived, and the JPEG2025100588000009.jpg86 can be re-organized in 4×4 blocks through a scan order (such as horizontal, vertical or diagonal, etc.). However, the above calculation is only an example, and HyGT (Hypercube-Givens Transform) etc. can also be used for the calculation of non-separable second-order transforms to reduce the computational complexity of non-separable second-order transforms.

[0100] On the other hand, for the non-separable second-order transformation, a mode-dependent transformation kernel (or transformation core, transformation type) can also be selected. Here, the mode can include an intra prediction mode and / or an inter prediction mode.

[0101] For example, as described above, the non-separable second-order transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. For example, the 8×8 transformation can indicate a transformation that can be applied to an 8×8 region included inside the corresponding transformation coefficient block when both W and H are the same as or greater than 8, and the 8×8 region is the upper-left 8×8 region inside the corresponding transformation coefficient block. Similarly, the 4×4 transformation can indicate a transformation that can be applied to a 4×4 region included inside the corresponding transformation coefficient block when both W and H are the same as or greater than 4, and the 4×4 region is the upper-left 4×4 region inside the corresponding transformation coefficient block. For example, the 8×8 transformation kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transformation kernel matrix can be a 16×16 / 8×16 matrix.

[0102] At this time, for mode-based transformation kernel selection, two non-separable second-order transformation kernels per transformation set for non-separable second-order transformation can be configured for both the 8×8 transformation and the 4×4 transformation, and the number of transformation sets is four. That is, four transformation sets can be configured for the 8×8 transformation, and four transformation sets can be configured for the 4×4 transformation. In this case, each of the four transformation sets for the 8×8 transformation can include two 8×8 transformation kernels, and each of the four transformation sets for the 4×4 transformation can include two 4×4 transformation kernels.

[0103] However, the size of the conversion, the number of sets, and the number of conversion kernels within a set are merely examples, and sizes other than 8×8 or 4×4 can also be used, or n sets can be configured, and each set can contain k conversion kernels. Here, n and k are positive integers, respectively.

[0104] For example, the conversion set can be called an NSST set, and the conversion kernels within the NSST set can be called NSST kernels. For example, the selection of a specific set from among the conversion sets can be performed based on the intra prediction mode of the target block (CU or sub-block).

[0105] For example, the intra prediction mode can include two non-directional or non-angular intra prediction modes and 65 directional or angular intra prediction modes. The non-directional intra prediction mode can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction mode can include 65 intra prediction modes numbered from 2 to 66. However, this is merely an example, and the embodiments according to this document can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can also indicate the LM (linear model) mode.

[0106] FIG. 5 exemplarily shows the intra-directional mode of 65 prediction directions.

[0107] Referring to FIG. 5, around the 34th intra prediction mode having a diagonal prediction direction upward to the left, an intra prediction mode having horizontal directionality and an intra prediction mode having vertical directionality can be distinguished. H and V in FIG. 5 can respectively mean horizontal directionality and vertical directionality, and the numbers from -32 to 32 can indicate a displacement in units of 1 / 32 on the sample grid position. This can indicate an offset with respect to the mode index value.

[0108] For example, the 2nd to 33rd intra prediction modes can have horizontal directionality, and the 34th to 66th intra prediction modes can have vertical directionality. On the other hand, the 34th intra prediction mode can be considered to have neither strictly horizontal nor vertical directionality, but can be classified as belonging to the horizontal directionality from the viewpoint of determining the conversion set of the secondary conversion. The reason is that for the vertical mode symmetric about the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the input data alignment method for the horizontal mode is used. Here, transposing the input data can mean that for the two-dimensional block data M×N, the rows become columns and the columns become rows to form N×M data.

[0109] Also, the 18th intra prediction mode and the 50th intra prediction mode can respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. Since the 2nd intra prediction mode has a left reference pixel and predicts in the upward right direction, it can be called an upward right diagonal intra prediction mode. Similarly, the 34th intra prediction mode can be called a downward right diagonal intra prediction mode, and the 66th intra prediction mode can be called a downward left diagonal intra prediction mode.

[0110] On the one hand, when it is determined that a specific set is used for the non-separable transform, one of the k transform kernels in the specific set can be selected via the non-separable second-order transform index. For example, the encoding device can derive a non-separable second-order transform index indicating a specific transform kernel based on rate-distortion (RD) checking, and can signal the non-separable second-order transform index to the decoding device. For example, the decoding device can select one of the k transform kernels in the specific set based on the non-separable second-order transform index. For example, an NSST index having a value of 0 can indicate the first non-separable second-order transform kernel, an NSST index having a value of 1 can indicate the second non-separable second-order transform kernel, and an NSST index having a value of 2 can indicate the third non-separable second-order transform kernel. Or, an NSST index having a value of 0 can indicate that the first non-separable second-order transform is not applied to the target block, and an NSST index having a value from 1 to 3 can point to the three transform kernels.

[0111] The transform unit can perform the non-separable second-order transform based on the selected transform kernel and obtain modified (second-order) transform coefficients. As described above, the modified transform coefficients can be derived as the transform coefficients quantized via the quantization unit, and can be encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0112] On the other hand, as described above, when the second-order transform is omitted, the (first-order) transform coefficients that are the output of the first-order (separable) transform can be derived as the transform coefficients quantized via the quantization unit, and can be encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0113] Referring again to FIG. 4, the inverse conversion unit can perform a series of procedures in the reverse order of the procedures executed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficients, performs a secondary (inverse) conversion to derive the (primary) conversion coefficients (S450), and can perform a primary (inverse) conversion on the (primary) conversion coefficients to obtain a residual block (residual samples) (S460). Here, the primary conversion coefficients can be referred to as modified conversion coefficients on the inverse conversion unit side. As described above, the encoding device and / or the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.

[0114] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). For example, the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion is NSST or RST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on a secondary conversion flag parsed or obtained from the bitstream. Or, for example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficients of the residual block.

[0115] The secondary inverse conversion determination unit can determine the secondary inverse conversion. At this time, the secondary inverse conversion determination unit can determine the secondary inverse conversion applied to the current block based on the NSST (or RST) conversion set specified by the intra prediction mode. Or, the secondary conversion determination method can be determined depending on the primary conversion determination method. Or, various combinations of the primary conversion and the secondary conversion can be determined by the intra prediction mode. For example, the secondary inverse conversion determination unit can also determine the area to which the secondary inverse conversion is applied based on the size of the current block.

[0116] On the one hand, as described above, when the secondary (inverse) transformation is omitted, a residual block (residual sample) can be obtained by receiving the (inverse quantized) transformation coefficient and performing the primary (separation) inverse transformation. As described above, the encoding device and / or the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.

[0117] On the other hand, in this document, in order to reduce the computational amount and memory requirement by non-separable secondary transformation, an RST (reduced secondary transform) with a reduced size of the transform matrix (kernel) in the concept of NSST can be applied.

[0118] In this document, RST can be meant to be a (simplified) transformation performed on the residual sample for the target block based on a transform matrix whose size is reduced by a simplification factor. When this is done, the amount of computation required at the time of transformation can be reduced by reducing the size of the transform matrix. That is, RST can be used to solve the problem of computational complexity that occurs during the transformation of a block with a large size or non-separable transformation.

[0119] For example, RST can be called by various terms such as reduced transform, reduced secondary transform, reduction transform, simplified transform or simple transform, etc., and the name by which RST is called is not limited to the listed examples. Or, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it can be called LFNST (Low-Frequency Non-Separable Transform).

[0120] On the other hand, when the second inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 may include an inverse RST unit that derives a transformation coefficient corrected based on the inverse RST for the transformation coefficient, and an inverse primary transformation unit that derives a residual sample for the target block based on the inverse primary transformation for the corrected transformation coefficient. The inverse primary transformation means the inverse transformation of the primary transformation applied to the residual. In this document, deriving a transformation coefficient based on a transformation can mean deriving the transformation coefficient by applying the corresponding transformation.

[0121] FIG. 6 and FIG. 7 are diagrams for explaining RST according to an embodiment of this document.

[0122] For example, FIG. 6 is a drawing for explaining that a forward reduced transform is applied, and FIG. 7 is a diagram for explaining that an inverse reduced transform is applied. In this document, the target block can indicate the current block where coding is performed, the residual block, or the transform block.

[0123] For example, in RST, an N-dimensional vector can be mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined. Here, N and R are positive integers, respectively, and R is smaller than N. N can represent the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can represent the R / N value. The simplification factor can be called by various terms such as reduced factor, reduction factor, simplified factor, or simple factor. On the other hand, R can be called a reduced coefficient, but in some cases, the simplification factor can also represent R. Also, in some cases, the simplification factor can represent the N / R value.

[0124] For example, the simplification factor or the reduced coefficient can be signaled via a bitstream, but is not limited thereto. For example, there may be cases where pre-defined values for the simplification factor or the reduced coefficient are stored in each encoding device 200 and decoding device 300. In this case, the simplification factor or the reduced coefficient is not signaled separately.

[0125] For example, the size (R×N) of the simplified transformation matrix is smaller than the size (N×N) of the normal transformation matrix and can be defined as in the following formula.

[0126]

Equation

[0127] For example, the matrix T in the reduced transform block shown in FIG. 6 is the matrix T of Equation 4R×N can be shown. As shown in FIG. 6, for the residual sample with respect to the target block, the simplified conversion matrix T R×N is multiplied, and the conversion coefficient for the target block can be derived.

[0128] For example, when the size of the block to which the conversion is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6 can be expressed by the matrix operation of the following Equation 5. In this case, the memory and multiplication operations can be reduced by approximately 1 / 4 by the simplification factor.

[0129] In this document, the matrix operation can be understood as an operation in which a matrix is placed on the left side of a column vector, and the matrix and the column vector are multiplied to obtain a column vector.

[0130]

Equation

[0131] In Equation 5, r1 to r 64 can represent the residual sample with respect to the target block. Or, for example, they are the conversion coefficients generated by applying a first-order conversion. Based on the operation result of Equation 5, the conversion coefficient c i for the target block can be derived.

[0132] For example, when R is 16, the conversion coefficients c1 to c 16It can be derived. If a normal (regular) conversion is applied instead of RST, and a conversion matrix with a size of 64×64 (N×N) is multiplied by a residual sample with a size of 64×1 (N×1), 64 (N) conversion coefficients for the target block are derived. However, since RST is applied, only 16 (R) conversion coefficients for the target block are derived. Since the total number of conversion coefficients for the target block decreases from N to R and the amount of data transmitted from the encoding device 200 to the decoding device 300 decreases, the transmission efficiency between the encoding device 200 and the decoding device 300 can be increased.

[0133] Considering the size aspect of the conversion matrix, the size of the normal conversion matrix is 64×64 (N×N), while the size of the simplified conversion matrix is reduced to 16×64 (R×N). Therefore, when compared with performing normal conversion, the memory usage can be reduced at a ratio of R / N when performing RST. Also, when compared with the number of multiplication operations N×N when using the normal conversion matrix, when using the simplified conversion matrix, the number of multiplication operations can be reduced at a ratio of R / N (R×N).

[0134] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for the target block by performing a primary conversion and an RST-based secondary conversion on the residual sample for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300. The inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and can derive the residual sample for the target block based on the inverse primary conversion for the modified conversion coefficients.

[0135] Inverse RST matrix T according to one embodiment N×R has a size of N×R, which is smaller than the size N×N of the normal inverse conversion matrix. The simplified conversion matrix T shown in Equation 4 R×Nis in a transpose relationship.

[0136] The matrix T within the reduced inverse transform block shown in FIG. 7 t is the inverse RST matrix T R×N T can be shown. Here, the superscript T can indicate transpose. When the inverse RST matrix T R×N T is multiplied by the transform coefficients for the target block as shown in FIG. 7, the corrected transform coefficients for the target block or the residual samples for the target block can be derived. The inverse RST matrix T R×N T can also be expressed as (T R×N ) T N×R in another way.

[0137] More specifically, when inverse RST is applied in the second - order inverse transform, when the inverse RST matrix T R×N T is multiplied by the transform coefficients for the target block, the corrected transform coefficients for the target block can be derived. On the other hand, inverse RST can be applied in the inverse first - order transform. In this case, when the inverse RST matrix T R×N T is multiplied by the transform coefficients for the target block, the residual samples for the target block can be derived.

[0138] In one embodiment, when the size of the block to which the inverse transform is applied is 8×8 and R is 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7 can be expressed by the matrix operation of the following Equation 6.

[0139]

Equation

[0140] In Equation 6, c1 to c16 can indicate the conversion coefficient for the target block. r indicating the corrected conversion coefficient for the target block or the residual sample for the target block based on the calculation result of Equation 6 j can be derived. That is, r1 to r indicating the corrected conversion coefficient for the target block or the residual sample for the target block N can be derived.

[0141] Considering the size aspect of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64×64 (N×N), while the size of the simplified inverse transformation matrix is reduced to 64×16 (N×R). Therefore, compared with the time of performing the normal inverse transformation, when performing the inverse RST, the memory usage can be reduced at a ratio of R / N. Also, compared with the number of multiplication operations N×N when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced at a ratio of R / N (N×R).

[0142] On the other hand, a conversion set can also be configured and applied to 8×8 RST. That is, the corresponding 8×8 RST can be applied by the conversion set. Since one conversion set is composed of two or three conversion kernels depending on the prediction mode within the screen, it can be configured to select one from a maximum of four conversions including the case where no second-order conversion is applied. The conversion when no second-order conversion is applied is regarded as the identity matrix being applied. When each of the four conversions is assigned an index of 0, 1, 2, or 3 (for example, the 0th index can be assigned to the identity matrix, i.e., the case where no second-order conversion is applied), a syntax element called the NSST index can be signaled for each conversion coefficient block to specify the conversion to be applied. That is, 8×8 NSST can be specified for the 8×8 upper left block via the NSST index, and 8×8 RST can be specified in the RST configuration. 8×8 NSST and 8×8 RST can indicate the conversion that can be applied to the 8×8 region included inside the corresponding conversion coefficient block when all of the W and H of the target block to be converted are the same as or larger than 8, and the 8×8 region is the upper left 8×8 region inside the corresponding conversion coefficient block. Similarly, 4×4 NSST and 4×4 RST can indicate the conversion that can be applied to the 4×4 region included inside the corresponding conversion coefficient block when all of the W and H of the target block are the same as or larger than 4, and the 4×4 region is the upper left 4×4 region inside the corresponding conversion coefficient block.

[0143] One party, for example, the encoding device, can encode values of syntax elements or quantized values of transform coefficients related to residuals based on various coding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC) to derive a bitstream. Also, the decoding device can decode the bitstream based on various coding methods such as exponential Golomb coding, CAVLC, or CABAC, and derive values of syntax elements necessary for image restoration or quantized values of transform coefficients related to residuals.

[0144] For example, the above-described coding method can be performed as described in the following content.

[0145] FIG. 8 exemplarily shows context-adaptive binary arithmetic coding (CABAC) for encoding syntax elements.

[0146] For example, in the coding process of CABAC, when the input signal is a syntax element that is not a binary value, the encoder can binarize the value of the input signal to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), the input signal can be used as it is without performing binarization. Here, each binary number 0 or 1 that constitutes a binary value can be referred to as a bin. For example, when the binary string after binarization is 110, each of 1, 1, and 0 can be represented as one bin. The bin for one syntax element can indicate the value of the syntax element. Such binarization can be based on various binarization methods such as Truncated Rice binarization process or Fixed-length binarization process, and the binarization method for the target syntax element can be defined in advance. The binarization procedure can be performed by a binarization unit within the entropy encoding unit.

[0147] Thereafter, the binarized bin of the syntax element can be input to a regular coding engine or a bypass coding engine. The regular coding engine of the encoder can assign a context model that reflects a probability value to the corresponding bin, and can encode the corresponding bin based on the assigned context model. The regular coding engine of the encoder can update the context model for the corresponding bin after performing coding for each bin. The bin coded as described above can be referred to as a context-coded bin.

[0148] On the one hand, when the binary bin of the syntax element is input to the bypass coding engine, it can be coded as follows. For example, the bypass coding engine of the encoding device can omit the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after coding. When bypass coding is applied, the encoding device can code the input bin by applying a uniform probability distribution instead of assigning a context model, and through this, the encoding speed can be improved. The bin coded as described above can be represented as a bypass bin.

[0149] Entropy decoding can show the process of performing the same process as the above-mentioned entropy encoding in reverse order.

[0150] The decoding device (entropy decoding unit) can decode the encoded image / video information. The image / video information can include partitioning-related information, prediction-related information (e.g., inter / intra prediction partition information, intra prediction mode information, inter prediction mode information, etc.), residual information, or in-loop filtering-related information, or can include various syntax elements related thereto. The entropy coding can be performed on a syntax element-by-syntax element basis.

[0151] The decoding device can perform binarization on the target syntax element. Here, the binarization can be based on various binarization methods such as Truncated Rice binarization process or Fixed-length binarization process, and the binarization method for the target syntax element can be defined in advance. The decoding device can derive an available bin string (bin string candidate) for the available values of the target syntax element through the binarization procedure. The binarization procedure can be performed by a binarization unit within the entropy decoding unit.

[0152] The decoding device can compare the derived bin string with the available bin string for the corresponding syntax element while sequentially decoding or parsing each bin for the target syntax element from the input bits in the bitstream. If the derived bin string is the same as one of the available bin strings, the value corresponding to the bin string is derived as the value of the corresponding syntax element. If not, the next bit in the bitstream can be further parsed and the above-described procedure can be performed again. Through such a process, the relevant information can be signaled using variable-length bits without using start bits or end bits for specific information (or specific syntax elements) in the bitstream. Through this, relatively fewer bits can be allocated for lower values, and the overall coding efficiency can be improved.

[0153] The decoding device can decode each bin in the bin string from the bitstream based on a context model or bypass based on an entropy coding technique such as CABAC or CAVLC.

[0154] When a syntax element is decoded based on a context model, the decoding apparatus can receive a bin corresponding to the syntax element via a bitstream, and can determine a context model using the syntax element and decoding information of a decoding target block or an adjacent block or information of symbols / bins decoded in a previous step, and can derive a value of the syntax element by predicting a generation probability of the received bin by the determined context model and performing arithmetic decoding of the bin. Thereafter, based on the determined context model, a context model of a bin to be decoded next can be updated.

[0155] The context model can be assigned and updated for each bin to be context-coded (entropy-coded), and the context model can be indicated based on a context index (ctxIdx: context index) or a context index increment (ctxInc: context index increment). The ctxIdx can be derived based on the ctxInc. Specifically, for example, the ctxIdx indicating a context model for each of the bins to be entropy-coded can be derived as a sum of the ctxInc and a context index offset (ctxIdxOffset: context index offset). For example, the ctxInc can be derived to be different for each bin. The ctxIdxOffset is indicated by the lowest value of the ctxIdx. The ctxIdxOffset is generally a value used for distinction from context models for other syntax elements, and a context model for one syntax element can be distinguished or derived based on the ctxInc.

[0156] In the entropy encoding procedure, it can be determined whether to perform encoding via the normal coding engine or via the bypass coding engine, whereby the coding path can be switched. Entropy decoding can perform the same process as entropy encoding in reverse order.

[0157] On the other hand, for example, when a syntax element is bypass decoded, the decoding device can receive the bin corresponding to the syntax element via the bit stream and decode the input bin by applying a uniform probability distribution. In this case, the procedures of deriving the context model of the syntax element and updating the context model applied to the bin after decoding can be omitted.

[0158] On the one hand, one embodiment of this document can propose a solution for signaling the MTS index. Here, as described above, the MTS index can represent any one of a plurality of conversion kernel sets. The MTS index can be encoded and the MTS index information can be signaled to the decoding device. The decoding device can decode the MTS index information to obtain the MTS index, and can determine the conversion kernel set to be applied based on the MTS index. The MTS index can also be represented by the tu_mts_idx syntax element or the mts_idx syntax element. For example, the MTS index can be binary-coded using the Rice-Golomb parameter order 0, but can also be binary-coded based on Truncated Rice. When binary-coded based on Truncated Rice, the input parameter cMax can have a value of 4, and cRiceParam can have a value of 0. For example, the encoding device can binary-code the MTS index to derive a bin (etc.) for the MTS index, encode the derived bin (etc.), and derive bits (etc.) for the MTS index information (for the MTS index), and the MTS index information can be signaled to the decoding device. The decoding device can decode the MTS index information to derive a bin (etc.) for the MTS index, and compare the derived bin (etc.) for the MTS index with candidate bins (etc.) for the MTS index to derive the MTS index.

[0159] For example, the MTS index (e.g., the tu_mts_idx syntax element or the mts_idx syntax element) can be context-coded for all bins based on a context model or a context index. In this case, the context index increment (ctxInc: context index increment) or ctxInc based on the bin position for the context coding of the MTS index can be assigned or determined as shown in Table 2. Alternatively, as shown in Table 2, a context model can be selected based on the bin position.

[0160]

Table 2

[0161] Referring to Table 2, the ctxInc for bin 0 (the first bin) can be assigned based on cqtDepth. Here, cqtDepth can represent the quad-tree depth for the current block and can be derived to one value among 0 to 5. That is, the ctxInc for bin 0 can be assigned to one value among 0 to 5 by cqtDepth. Also, the ctxInc for bin 1 (the second bin) can be assigned 6, the ctxInc for bin 2 (the third bin) can be assigned 7, and the ctxInc for bin 3 (the fourth bin) can be assigned 8. That is, bins 0 to 3 can be assigned different values of ctxInc. Here, different ctxInc values can represent different context models, and in this case, the number of context models for the coding of the MTS index can be nine.

[0162] Alternatively, for example, the MTS index (e.g., the tu_mts_idx syntax element or the mts_idx syntax element) can also be bypass-coded for all bins as shown in Table 3. In this case, the number of context models for coding the MTS index can be zero.

[0163]

Table 3

[0164] Alternatively, for example, the MTS index (e.g., the tu_mts_idx syntax element or the mts_idx syntax element) can be context-coded based on a context model or a context index for bin 0 (the first bin) as shown in Table 4, and can also be bypass-coded for the remaining bins. That is, ctxInc for bin 0 (the first bin) can be assigned 0. In this case, the number of context models for coding the MTS index can be one.

[0165]

Table 4

[0166] Alternatively, for example, the MTS index (e.g., the tu_mts_idx syntax element or the mts_idx syntax element) can be context-coded based on a context model or a context index for bin 0 (the first bin) and bin 1 (the second bin) as shown in Table 5, and can also be bypass-coded for the remaining bins. That is, ctxInc for bin 0 (the first bin) can be assigned 0, and ctxInc for bin 1 (the second bin) can be assigned 1. In this case, the number of context models for coding the MTS index can be two.

[0167]

Table 5

[0168] Alternatively, for example, the MTS index (e.g., tu_mts_idx syntax element or mts_idx syntax element) can be context-coded for all bins based on a context model or context index as shown in Table 6, and one ctxInc can be assigned to each bin. That is, the ctxInc for bin 0 (the first bin) can be assigned 0, the ctxInc for bin 1 (the second bin) can be assigned 1, the ctxInc for bin 2 (the third bin) can be assigned 2, and the ctxInc for bin 3 (the fourth bin) can be assigned 2. In this case, the context model for coding the MTS index can be four.

[0169]

Table 6

[0170] As described above, in one embodiment, whether bypass coding is applied to all or part of the bins of the MTS index or context coding is applied, applying a specific value to ctxInc to reduce the number of context models can reduce the complexity and may have the effect of increasing the output power of the decoder. Also, in one embodiment, when using a context model as described above, the initial value and / or the size of the multiple windows can also be variable based on the occurrence statistics for the position of each bin.

[0171] FIG. 9 and FIG. 10 schematically show an example of a video / image encoding method and related components according to an embodiment (etc.) of this document.

[0172] The method disclosed in FIG. 9 can be performed by the encoding device disclosed in FIG. 2 or FIG. 10. Specifically, for example, S900 to S920 in FIG. 9 can be performed by the residual processing unit 230 of the encoding device in FIG. 10, and S930 in FIG. 9 can be performed by the entropy encoding unit 240 of the encoding device in FIG. 10. Although not shown in FIG. 9, a prediction sample or prediction-related information can be derived by the prediction unit 220 of the encoding device in FIG. 10, residual information can be derived from the original sample or the prediction sample by the residual processing unit 230 of the encoding device, and a bitstream can be generated from the residual information or the prediction-related information by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 9 can include the embodiments described above in this document.

[0173] As shown in FIG. 9, the encoding device derives a residual sample for the current block (S900). For example, the encoding device can derive a residual sample based on the prediction sample and the original sample. Although not shown in FIG. 9, the encoding device can perform intra prediction or inter prediction on the current block considering the rate distortion (RD) cost to generate a prediction sample for the current block, and can generate prediction-related information including prediction mode / type information.

[0174] The encoding device derives a conversion coefficient for the current block based on the residual samples (S910). For example, the encoding device can perform a conversion on the residual samples to derive the conversion coefficient. Here, the conversion can be performed based on a conversion kernel or a set of conversion kernels. For example, the set of conversion kernels can include a horizontal direction conversion kernel and a vertical direction conversion kernel. For example, the encoding device can perform a first-order conversion on the residual samples to derive the conversion coefficient. Or, for example, the encoding device can perform a first-order conversion on the residual samples to derive a temporary conversion coefficient, and then perform a second-order conversion on the temporary conversion coefficient to derive the conversion coefficient. For example, the conversion performed based on the set of conversion kernels can represent a first-order conversion.

[0175] The encoding device generates an MTS index and residual information based on the conversion coefficient (S920). In other words, the encoding device can generate an MTS index and / or residual information based on the conversion coefficient.

[0176] The MTS index can represent the set of conversion kernels applied to the current block (conversion coefficient) among the set of conversion kernel candidates. Here, the MTS index can also be represented as a tu_mts_idx syntax element or an mts_idx syntax element. As described above, the set of conversion kernels can include a horizontal direction conversion kernel and a vertical direction conversion kernel. The horizontal direction conversion kernel can be represented as trTypeHor, and the vertical direction conversion kernel can be represented as trTypeVer.

[0177] For example, the values of trTypeHor and trTypeVer can be represented by the horizontal direction conversion kernel and the vertical direction conversion kernel applied to the current block (conversion coefficient), and the MTS index can be represented by one of the candidates including 0 to 4 according to the values of trTypeHor and trTypeVer.

[0178] For example, when the MTS index is 0, it can represent that both trTypeHor and trTypeVer are 0. Or, when the MTS index is 1, it can represent that both trTypeHor and trTypeVer are 1. Or, when the MTS index is 2, it can represent that trTypeHor is 2 and trTypeVer is 1. When the MTS index is 3, it can represent that trTypeHor is 1 and trTypeVer is 1. Or, when the MTS index is 4, it can represent that both trTypeHor and trTypeVer are 2. For example, when the value of trTypeHor or trTypeVer is 0, it can represent that DCT2 has been applied horizontally or vertically to the current block (transformation coefficient). When it is 1, it can represent that DST7 has been applied. When it is 2, it can represent that DCT8 has been applied. That is, each of the transformation kernel applied horizontally and the transformation kernel applied vertically can be represented by one of the candidates including DCT2, DST7, and DCT8 based on the MTS index.

[0179] The MTS index can be represented based on the bins of the bin string of the MTS index. In other words, the MTS index can be binary-coded and represented by the bins of the bin string of the MTS index, and the bin string of the MTS index (bins) can be entropy-encoded.

[0180] In other words, among the bins of the bin string of the MTS index, at least one bin can be represented based on context coding. Here, the context coding can be performed based on the value of the context index increment (ctxInc). Or, the context coding can be performed based on the context index (ctxIdx) or the context model. Here, the context index can be represented based on the value of the context index increment. Or, the context index can also be represented based on the value of the context index increment and the context index offset (ctxIdxOffset).

[0181] For example, all bins of the bin string of the MTS index can be represented based on context coding. For example, for the first bin or bin 0 of the bin string of the MTS index, ctxInc can be represented based on cqtDepth. Here, cqtDepth can represent the quad-tree depth for the current block and can be represented by one value among 0 to 5. Also, ctxInc for the second bin or bin 1 can be represented by 6, ctxInc for the third bin or bin 2 can be represented by 7, and ctxInc for the fourth bin or bin 3 can be represented by 8. Or, for example, for the first bin or bin 0 of the bin string of the MTS index, ctxInc can be represented by 0, ctxInc for the second bin or bin 1 can be represented by 1, ctxInc for the third bin or bin 2 can be represented by 2, and ctxInc for the fourth bin or bin 3 can be represented by 3. That is, the number of values of the context index increment that can be used for context coding of the first bin among the bins of the bin string can be one.

[0182] Alternatively, for example, some of the bins in the bin string of the MTS index can be represented based on context coding, and the rest can be represented based on bypass coding. For example, for the first bin or bin 0 in the bin string of the MTS index, ctxInc can be represented as 0, and the remaining bins can be represented based on bypass coding. Or, for example, for the first bin or bin 0 in the bin string of the MTS index, ctxInc can be represented as 0, for the second bin or bin 1, ctxInc can be represented as 1, and the remaining bins can be represented based on bypass coding. That is, the number of values of the context index increment that can be used for context coding of the first bin among the bins of the bin string can be one.

[0183] Alternatively, all of the bins in the bin string of the MTS index can be represented based on bypass coding. Here, bypass coding can also represent performing context coding based on a uniform probability distribution, and by omitting procedures such as the update procedure of context coding, the coding efficiency can be improved.

[0184] Residual information can represent information used to derive residual samples. Also, for example, an encoding device can perform quantization on transform coefficients, and the residual information can include information regarding residual samples, transform-related information, and / or quantization-related information. For example, the residual information can include information regarding quantized transform coefficients.

[0185] The encoding device encodes image information including an MTS index and residual information (S930). For example, the image information can further include prediction-related information. For example, the encoding device can encode the image information to generate a bitstream. The bitstream can also be called encoded (image) information.

[0186] Alternatively, although not shown in FIG. 9, for example, the encoding device can also generate a restored sample based on the residual sample and the prediction sample. Also, a restored block and a restored picture can be derived based on the restored sample.

[0187] For example, the encoding device can encode image information including all or part of the above-described information (or syntax elements) to generate a bitstream or encoded information. Alternatively, it can be output in bitstream form. Also, the bitstream or encoded information can be transmitted to a decoding device via a network or a storage medium. Alternatively, the bitstream or encoded information can be stored in a computer-readable storage medium, and the bitstream or the encoded information can be generated by the above-described image encoding method.

[0188] FIGS. 11 and 12 schematically show an example of a video / image decoding method and related components according to an embodiment (etc.) of this document.

[0189] The method disclosed in FIG. 11 can be performed by the decoding device disclosed in FIG. 3 or FIG. 12. Specifically, for example, S1100 in FIG. 11 can be performed by the entropy decoding unit 310 of the decoding device in FIG. 12, and S1110 and S1120 in FIG. 11 can be performed by the residual processing unit 320 of the decoding device in FIG. 12. Although not shown in FIG. 11, prediction-related information or residual information can be derived from the bitstream by the entropy decoding unit 310 of the decoding device in FIG. 12, residual samples can be derived from the residual information by the residual processing unit 320 of the decoding device, prediction samples can be derived from the prediction-related information by the prediction unit 330 of the decoding device, and a restored block or a restored picture can be derived from the residual samples or the prediction samples by the addition unit 340 of the decoding device. The method disclosed in FIG. 11 can include the embodiments described above in this document.

[0190] As shown in FIG. 11, the decoding device acquires an MTS index and residual information from the bitstream (S1100). For example, the decoding device can parse or decode the bitstream to acquire the MTS index and / or the residual information. Here, the bitstream can also be called encoded (image) information.

[0191] The MTS index can represent a set of transform kernels to be applied to the current block among the set of transform kernel candidates. Here, the MTS index can also be represented as a tu_mts_idx syntax element or an mts_idx syntax element. The set of transform kernels can include a transform kernel applied horizontally to the current block and a transform kernel applied vertically to the current block. Here, the transform kernel applied horizontally can be represented as trTypeHor, and the transform kernel applied vertically can be represented as trTypeVer.

[0192] For example, the MTS index can be derived from one of the candidates including 0 to 4, and by the MTS index, trTypeHor and trTypeVer can each be derived from one of 0 to 2. For example, when the MTS index is 0, both trTypeHor and trTypeVer can be 0. Or, when the MTS index is 1, both trTypeHor and trTypeVer can be 1. Or, when the MTS index is 2, trTypeHor can be 2 and trTypeVer can be 1. When the MTS index is 3, trTypeHor can be 1 and trTypeVer can be 1. Or, when the MTS index is 4, both trTypeHor and trTypeVer can be 2. For example, the value of trTypeHor or trTypeVer can represent a transform kernel, and when it is 0, it can represent DCT2, when it is 1, it can represent DST7, and when it is 2, it can represent DCT8. That is, each of the transform kernel applied in the horizontal direction and the transform kernel applied in the vertical direction can be derived from one of the candidates including DCT2, DST7, and DCT8 based on the MTS index.

[0193] The MTS index can be derived based on the bins of the bin string of the MTS index. In other words, the MTS index information can be derived from the entropy decoded and binary MTS index, and the binary MTS index can be represented by the bin string (of the bins) of the MTS index.

[0194] In other words, among the bins of the bin string of the MTS index, at least one bin can be derived based on context coding. Here, the context coding can be performed based on the value of the context index increment (ctxInc). Alternatively, the context coding can be performed based on the context index (ctxIdx) or the context model. Here, the context index can be derived based on the value of the context index increment. Alternatively, the context index can also be derived based on the value of the context index increment and the context index offset (ctxIdxOffset).

[0195] For example, all bins of the bin string of the MTS index can be derived based on context coding. For example, for the first bin or bin 0 of the bin string of the MTS index, ctxInc can be assigned based on cqtDepth. Here, cqtDepth can represent the quad-tree depth for the current block and can be derived to one value among 0 to 5. Also, for the second bin or bin 1, ctxInc can be assigned 6, for the third bin or bin 2, ctxInc can be assigned 7, and for the fourth bin or bin 3, ctxInc can be assigned 8. Or, for example, for the first bin or bin 0 of the bin string of the MTS index, ctxInc can be assigned 0, for the second bin or bin 1, ctxInc can be assigned 1, for the third bin or bin 2, ctxInc can be assigned 2, and for the fourth bin or bin 3, ctxInc can be assigned 3. That is, the number of values of the context index increment that can be used for context coding of the first bin among the bins of the bin string can be one.

[0196] Alternatively, for example, some of the bins in the bin string of the MTS index can be derived based on context coding, and the rest can be derived based on bypass coding. For example, for the first bin or bin 0 in the bin string of the MTS index, ctxInc can be assigned 0, and the remaining bins can be derived based on bypass coding. Or, for example, for the first bin or bin 0 in the bin string of the MTS index, ctxInc can be assigned 0, for the second bin or bin 1, ctxInc can be assigned 1, and the remaining bins can be derived based on bypass coding. That is, the number of values of the context index increment that can be used for context coding of the first bin among the bins of the bin string can be one.

[0197] Alternatively, all of the bins in the bin string of the MTS index can be derived based on bypass coding. Here, bypass coding can also represent performing context coding based on a uniform probability distribution, and by omitting procedures such as the update procedure of context coding, the coding efficiency can be improved.

[0198] Residual information can represent the information used to derive residual samples, and can include information related to residual samples, inverse transformation related information, and / or inverse quantization related information. For example, residual information can include information related to quantized transform coefficients.

[0199] The decoding device derives the transform coefficients for the current block based on the residual information (S1110). For example, the decoding device can derive the quantized transform coefficients for the current block based on the information regarding the quantized transform coefficients included in the residual information. For example, the decoding device can perform inverse quantization on the quantized transform coefficients to derive the transform coefficients for the current block.

[0200] The decoding device generates the residual samples of the current block based on the MTS index and the transform coefficients (S1120). For example, the residual samples can be generated based on the transform kernel set represented by the transform coefficients and the MTS index. That is, the decoding device can generate the residual samples from the transform coefficients through inverse transformation using the transform kernel set represented by the MTS index. Here, the inverse transformation using the transform kernel set represented by the MTS index can be included in the first-order inverse transformation. Alternatively, when generating the residual samples from the transform coefficients, the decoding device can utilize not only the first-order inverse transformation but also the second-order inverse transformation. In this case, the decoding device can also perform a second-order inverse transformation on the transform coefficients to derive the corrected transform coefficients, and then perform a first-order inverse transformation on the corrected transform coefficients to generate the residual samples.

[0201] Although not illustrated in FIG. 11, for example, the decoding device can obtain prediction-related information including prediction mode / type information from the bitstream, and perform intra prediction or inter prediction based on the prediction mode / type information to generate prediction samples for the current block. Also, for example, the decoding device can generate restored samples based on the prediction samples and the residual samples. Also, for example, a restored block or a restored picture can be derived based on the restored samples.

[0202] For example, the decoding device can decode a bitstream or encoded information to obtain image information including all or part of the above-described information (or syntax elements). Further, the bitstream or encoded information can be stored in a computer-readable storage medium, and the above-described decoding method can be performed.

[0203] In the above-described embodiments, the method is described based on a flowchart in a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with a step different from the above. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.

[0204] The method according to the above-described embodiments of this document can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in devices that perform image processing such as, for example, TVs, computers, smartphones, set-top boxes, display devices, and the like.

[0205] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and performed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.

[0206] In addition, the decoding device and encoding device to which the embodiments of this document are applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video intercom devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, pay-per-view (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, picture phone video devices, transportation means terminals (e.g., vehicle terminals including autonomous driving vehicles, airplane terminals, ship terminals, etc.), and medical video devices, etc., and can be used to process video signals or data signals. For example, as an over-the-top (OTT) video device, it can include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0207] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the embodiments of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0208] In addition, the embodiments of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0209] FIG. 13 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0210] As shown in FIG. 13, the content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0211] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and plays the role of sending this to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate a bitstream, the encoding server can be omitted.

[0212] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.

[0213] The streaming server sends multimedia data to the user device based on a user request via a web server, and the web server plays the role of a medium for informing the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server sends multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.

[0214] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0215] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, and a digital signage.

[0216] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.

[0217] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.

Claims

1. In an image decoding method performed by a decoding device, the steps of obtaining an MTS (Multiple Transform Selection) index and residual information from a bitstream, deriving transform coefficients for a current block based on the residual information, generating residual samples of the current block based on the MTS index and the transform coefficients, comprising: the MTS index represents a set of transform kernels applied to the current block among a set of transform kernel candidates, bins of a bin string of the MTS index are derived based on context coding, the context coding is performed using a context model represented by a context index for each of the bins with respect to the MTS index, the context index for each of the bins with respect to the MTS index is derived as a sum of a value of a context index increment for each of the bins with respect to the MTS index and a value of a context index offset for each of the bins with respect to the MTS index, the bins with respect to the MTS index include a first bin, a second bin, a third bin, and a fourth bin, the value of the context index increment for the first bin is fixed to 0, the value of the context index increment for the second bin is fixed to 1, the value of the context index increment for the third bin is fixed to 2, the value of the context index increment for the fourth bin is fixed to 3, an image decoding method.

2. In an image encoding method performed by an encoding device, the step of deriving residual samples for a current block, the step of deriving transform coefficients for the current block based on the residual samples, the step of generating an MTS (Multiple Transform Selection) index and residual information based on the transform coefficients, the step of encoding image information including the MTS index and the residual information, comprising: The MTS index represents a set of transform kernels applied to the current block among the transform kernel set candidates. The bins of the bin string of the MTS index are represented based on context coding. The context coding is performed using a context model represented by a context index for each of the bins with respect to the MTS index. The context index for each of the bins with respect to the MTS index is derived as the sum of the value of the context index increment for each of the bins with respect to the MTS index and the value of the context index offset for each of the bins with respect to the MTS index. The bins with respect to the MTS index include a first bin, a second bin, a third bin, and a fourth bin. The value of the context index increment for the first bin is fixed to 0. The value of the context index increment for the second bin is fixed to 1. The value of the context index increment for the third bin is fixed to 2. The value of the context index increment for the fourth bin is fixed to 3, an image encoding method. [

3. ] A method for transmitting data for an image, comprising: obtaining a bitstream for the image, wherein the bitstream is deriving a residual sample for the current block; deriving a transform coefficient for the current block based on the residual sample; generating an MTS (Multiple Transform Selection) index and residual information based on the transform coefficient; encoding image information including the MTS index and the residual information; and transmitting the data including the bitstream. The method includes: The MTS index represents a set of transform kernels applied to the current block among the transform kernel set candidates. The bins of the bin string of the MTS index are represented based on context coding. The context coding is performed using a context model represented by a context index for each of the bins with respect to the MTS index. The context index for each of the bins with respect to the MTS index is derived as the sum of a value of a context index increment for each of the bins with respect to the MTS index and a value of a context index offset for each of the bins with respect to the MTS index. The bins with respect to the MTS index include a first bin, a second bin, a third bin, and a fourth bin. The value of the context index increment for the first bin is fixed to 0. The value of the context index increment for the second bin is fixed to 1. The value of the context index increment for the third bin is fixed to 2. The value of the context index increment for the fourth bin is fixed to 3. A transmission method.

Citation Information

Patent Citations

  • Video encoding device and video decoding device

    JP2020053924A

  • Coding for information about the transformation kernel set

    JP7303335B2

  • Binarizing secondary transform index

    WO2017192705A1