Video image coding method based on secondary conversion and device of the same

The video coding method addresses the challenges of high-resolution video data by employing a reduced secondary transform and conversion set within the video coding apparatus, resulting in improved coding efficiency and compression performance.

JP2025090835AActive Publication Date: 2025-06-17LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025044400
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-06
Filing Date
2025-03-19
Publication Date
2025-06-17
Estimated Expiration
2039-12-05

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video data, such as 4K or UHD, poses challenges in efficient compression, transmission, storage, and reproduction, leading to higher costs and reduced efficiency in existing video coding technologies.

Method used

A video coding method and apparatus that utilizes a reduced secondary transform (RST) and a conversion set to enhance coding efficiency, involving the derivation of quantized conversion coefficients, inverse quantization, inverse RST, and inverse primary conversion to generate restored pictures.

Benefits of technology

The proposed method improves the efficiency of video coding, enhances the conversion efficiency, and optimizes the secondary conversion process, leading to better compression and transmission of high-resolution video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090835000001_ABST
    Figure 2025090835000001_ABST
Patent Text Reader

Abstract

To provide a video image coding method for enhancing an efficiency of a video image coding and a device.SOLUTION: A video image coding method contains: a step of introducing a conversion coefficient via an inverse quantization on the basis of the conversion coefficient to be quantized to an object block; a step of introducing the conversion coefficient to be corrected on the basis of an inverse RST to the conversion coefficient; and a step of generating a decoding picture on the basis of a residual sample to the object block on the basis of an inverse primary conversion to the conversion coefficient to be corrected. The inverse RST is performed on the basis of a conversion kernel matrix selected from a conversion set to be determined on the basis of a mapping relation by an intra-prediction mode adopted to the object block and two conversion kernel matrixes contained to the conversion set, and is performed on the basis of whether or not the inverse RST is adopted or a conversion index for instructing any one of the conversion kernel matrixes contained to the conversion set.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to video coding technology, and more particularly, to a video coding method based on transform in a video coding system and an apparatus therefor.

Background Art

[0002] Recently, the demand for high-resolution and high-quality video / video such as 4K or UHD (Ultra High Definition) video / video of 8K or higher has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to conventional video / video data. Therefore, when transmitting video data using a medium such as a conventional wired or wireless broadband line or storing video / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, recently, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / video having video characteristics different from those of real-world video such as game video has been increasing.

[0004] Therefore, in order to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality video / video having various characteristics as described above, a highly efficient video / video compression technology is required.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and an apparatus for increasing the coding efficiency of video.

[0006] Another technical problem of this document is to provide a method and an apparatus for increasing the conversion efficiency.

[0007] Another technical problem of this document is to provide a method and an apparatus for improving the efficiency of the secondary conversion through coding of the conversion index.

[0008] Another technical problem of this document is to provide a video coding method and an apparatus based on RST (reduced secondary transform).

[0009] Another technical problem of this document is to provide a video coding method and an apparatus based on a conversion set that can increase the coding efficiency.

Means for Solving the Problem

[0010] According to an embodiment of this document, a video decoding method performed by a decoding apparatus is provided. The method includes: deriving a quantized conversion coefficient for a target block from a bitstream; deriving a conversion coefficient through inverse quantization based on the quantized conversion coefficient for the target block; deriving a corrected conversion coefficient based on an inverse RST (reduced secondary transform) for the conversion coefficient; deriving a residual sample for the target block based on an inverse primary conversion for the corrected conversion coefficient; and generating a restored picture based on the residual sample for the target block. The inverse RST is performed based on a conversion set determined based on a mapping relationship according to an intra prediction mode applied to the target block, and a selected conversion kernel matrix among two conversion kernel matrices included in each of the conversion sets, and can be performed based on whether the inverse RST is applied and a conversion index indicating any one of the conversion kernel matrices included in the conversion set.

[0011] According to another embodiment of the present document, a decoding apparatus for performing video decoding is provided. The decoding apparatus includes an entropy decoding unit that derives quantized transform coefficients for a target block and information for prediction from a bitstream, a prediction unit that generates prediction samples for the target block based on the information for prediction, an inverse quantization unit that derives transform coefficients via inverse quantization based on the quantized transform coefficients for the target block, an inverse RST (reduced secondary transform) unit that derives corrected transform coefficients based on an inverse RST for the transform coefficients, an inverse transform unit including an inverse primary transform unit that derives residual samples for the target block based on an inverse primary transform for the corrected transform coefficients, and an addition unit that generates restored samples based on the residual samples and the prediction samples. The inverse RST is performed based on a conversion set determined based on a mapping relationship according to an intra prediction mode applied to the target block, and a selected conversion kernel matrix among two conversion kernel matrices included in each of the conversion sets, and is performed based on whether the inverse RST is applied and a conversion index indicating any one of the conversion kernel matrices included in the conversion set.

[0012] According to one embodiment of the present document, a video encoding method performed by an encoding device is provided. The method includes: deriving a prediction sample based on an intra prediction mode applied to a target block; deriving a residual sample for the target block based on the prediction sample; deriving a transform coefficient for the target block based on a first-order transform of the residual sample; deriving a modified transform coefficient based on an RST (reduced secondary transform) of the transform coefficient, where the inverse RST is performed based on a transform set determined based on a mapping relationship according to the intra prediction mode applied to the target block, and a selected transform kernel matrix among two transform kernel matrices included in each of the transform sets, quantizing based on the modified transform coefficient to derive a quantized transform coefficient; and generating a transform index indicating whether the RST is applied and any one of the transform kernel matrices included in the transform set.

[0013] According to another embodiment of the present document, a digital storage medium storing video data including encoded video information generated by a video encoding method performed by an encoding device can be provided.

[0014] According to another embodiment of the present document, a digital storage medium storing video data including encoded video information that causes a decoding device to perform the video decoding method can be provided.

Advantages of the Invention

[0015] According to the present document, the efficiency of general video / video compression can be improved.

[0016] According to the present document, the efficiency of secondary transform can be improved through coding of the transform index.

[0017] According to this document, video coding can be performed based on a conversion set, and the efficiency of video coding can be improved.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Modes for Carrying Out the Invention

[0019] This document can be modified in various ways and can have various embodiments. Therefore, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used only for explaining specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates a different meaning. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should not be understood as precluding the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0020] On the other hand, each configuration on the drawings described in this document is shown independently for the convenience of explaining different characteristic functions, and it does not mean that each configuration is realized by different hardware or different software. For example, among the configurations, two or more configurations may be combined to form one configuration, or one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0021] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components on the drawings, and duplicate descriptions of the same components will be omitted.

[0022] This document relates to video / video coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0023] This document presents various examples related to video / video coding, and unless otherwise stated, the examples may be combined with each other.

[0024] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles.

[0025] A pixel or pel can mean the smallest unit that constitutes one picture (or video). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Or, a sample can mean the pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also mean the conversion coefficient in the frequency domain.

[0026] A unit can represent the basic unit of video processing. A unit can include at least one of a specific area of a picture and information regarding the corresponding area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. The term “unit” may be used interchangeably with terms such as “block” or “area” as the case may be. In general, an MxN block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0027] In this document, the symbols “ / ” and “,” are interpreted to mean “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Further, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” means “at least one of A, B, and / or C.”

[0028] Further, in the document, the term “or” should be interpreted to indicate “and / or.” For instance, the expression “A or B” may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term “or” in this document should be interpreted to indicate “additionally or alternatively.”

[0029] FIG. 1 schematically shows an example of a video / image coding system to which this document is applicable.

[0030] Referring to FIG. 1, the video / image coding system may include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or a stream.

[0031] The source device can include a video source, an encoding device, and a transmission unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of another device or an external component.

[0032] The video source can obtain video / video through processes such as video / video capture, synthesis, or generation. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive containing previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated through a computer or the like, and in this case, the process of generating related data can be substituted for the video / video capture process.

[0033] The encoding device can encode the input video / video. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for the sake of compression and coding efficiency. The encoded data (encoded video / video information) can be output in the form of a bitstream.

[0034] The transmitting unit can transmit the encoded video / video information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0035] The decoding device can perform a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device to decode the video / video.

[0036] The renderer can render the decoded video / video. The rendered video / video can be displayed via the display unit.

[0037] Figure 2 is a diagram schematically explaining the configuration of a video / video encoding device to which this document is applicable. Hereinafter, the video encoding device can include the video encoding device.

[0038] Referring to FIG. 2, the encoding apparatus 200 may be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The aforementioned image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 may include a DPB (decoded picture buffer) and may also be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.

[0039] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or the largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary-tree structure may be applied later. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to this document is performed. In this case, based on the coding efficiency, etc. according to the characteristics of the video, the largest coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively divided into coding units with a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can be divided or partitioned from the aforementioned final coding unit respectively.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit that derives a conversion coefficient and / or a unit that derives a residual signal from the conversion coefficient.

[0040] The unit may, depending on the case, be used interchangeably with terms such as a block or an area. In a general case, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).

[0041] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. As will be described later in the description of each prediction mode, the prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0042] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.

[0043] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can configure a candidate list of motion information based on the peripheral blocks, and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction is performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0044] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply intra prediction or inter prediction for predicting a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) for predicting a block. The intra block copy can be used for coding content videos / movies such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but is performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0045] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform), etc. Here, GBT means the conversion obtained from this graph when expressing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block having the same size of a square or may be applied to a block of variable size that is not a square.

[0046] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. The quantization unit 233 can reorder the quantized transform coefficients in the form of a block into a one-dimensional vector based on the scan order of the coefficients, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the form of the one-dimensional vector. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document can be encoded through the above-described encoding procedure and can be included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200 by a transmitting unit (not shown) for transmission and / or a storing unit (not shown) for storage, or the transmitting unit may be included in the entropy encoding unit 240.

[0047] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0048] On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0049] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 290 as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 290 and output in the form of a bitstream.

[0050] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 280. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and also improve the encoding efficiency.

[0051] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0052] FIG. 3 is a diagram schematically illustrating the configuration of a video / video decoding apparatus to which this document is applicable.

[0053] Referring to FIG. 3, the decoding apparatus 300 can include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter predictor 331 and an intra predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0054] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information is processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the information regarding block partitioning obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Further, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.

[0055] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / video information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and the block to be decoded, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin according to the determined context model, performs arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element.At this time, in the CABAC entropy decoding method, after determining the context model, the context model can be updated using the information of the decoded symbol / bin for the context model of the next symbol / bin. Among the information related to prediction in the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the information regarding the residual for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the inverse quantization unit 321. Also, the information related to filtering among the information decoded by the entropy decoding unit 310 can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document may be called a video / video / picture decoding device, and the decoding device may be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0056] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in the form of a two-dimensional block. In this case, the reordering can be performed based on the scan order of the coefficients performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0057] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0058] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0059] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Further, the prediction unit can also perform intra block copy (IBC) for the prediction of a block. The intra block copy can be used for the coding of content videos / movies such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but is performed in the same manner as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0060] The intra prediction unit 332 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 332 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.

[0061] The inter prediction unit 331 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 331 can configure a candidate list of motion information based on the neighboring blocks, and derive the motion vector and / or the reference picture index of the current block based on the received candidate selection information. Inter prediction is performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0062] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor 330. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0063] The adder 340 may be referred to as a restoration unit or a restored block generation unit. The generated restored signal may be used for intra prediction of the next block to be processed in the current picture, may be output after filtering as described later, or may be used for inter prediction of the next picture.

[0064] On the other hand, LMCS (luma mapping with chroma scaling) may be applied in the process of decoding a picture.

[0065] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0066] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 331. The memory 360 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 331 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 332.

[0067] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding apparatus 300 can be applied to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding apparatus 200 in the same or corresponding manner.

[0068] As described above, in performing video coding, prediction is performed to improve the compression efficiency. Through this, a predicted block including prediction samples for the current block which is the block to be coded can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in the encoding apparatus and the decoding apparatus, and the encoding apparatus can improve the efficiency of video coding by signaling information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values of the original block, to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0069] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, perform a transformation procedure on the residual samples (residual sample array) included in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the related residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transformation coefficients. A decoding device can perform an inverse quantization / inverse transformation procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. The encoding device can also inverse-quantize / inverse-transform the quantized transformation coefficients for reference in inter-prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.

[0070] Figure 4 schematically shows the multiple transformation techniques according to this document.

[0071] Referring to Figure 4, the transformation unit can correspond to the transformation unit in the encoding device of Figure 2 described above, and the inverse transformation unit can correspond to the inverse transformation unit in the encoding device of Figure 2 or the inverse transformation unit in the decoding device of Figure 3 described above.

[0072] The transformation unit can perform a primary transformation based on the residual samples (residual sample array) in the residual block to derive (primary) transformation coefficients (S410). Such a primary transformation can be referred to as a core transform. Here, the primary transformation can be based on Multiple Transform Selection (MTS), and when a multiple transformation is applied as the primary transformation, it can be referred to as a multiple core transform.

[0073] The multi-core transform can be shown as a method of further using DCT (Discrete Cosine Transform) type 2, DST (Discrete Sine Transform) type 7, DCT type 8, and / or DST type 1 for transformation. That is, the multi-core transform can be shown as a transformation method for converting a residual signal (or residual block) in the spatial domain into a transform coefficient (or primary transform coefficient) in the frequency domain based on a plurality of selected transform kernels among the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary transform coefficient can be called a temporary transform coefficient from the perspective of the transform unit.

[0074] In other words, when an existing transform method is applied, the conversion from the spatial domain to the frequency domain for the residual signal (or residual block) is applied based on DCT type 2, and transform coefficients can be generated. Different from this, when the multi-core transform is applied, the conversion from the spatial domain to the frequency domain for the residual signal (or residual block) is applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., and transform coefficients (or primary transform coefficients) can be generated. Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. can be called transform types, transform kernels, or transform cores.

[0075] For reference, the DCT / DST transform type can be defined based on a basis function, and the basis function can be shown as follows in the following table.

[0076]

Table 1

[0077] When the multi-core transformation is performed, a vertical transformation kernel and a horizontal transformation kernel for a target block among the transformation kernels can be selected. A vertical transformation for the target block can be performed based on the vertical transformation kernel, and a horizontal transformation for the target block can be performed based on the horizontal transformation kernel. Here, the horizontal transformation may indicate a transformation for a horizontal component of the target block, and the vertical transformation may indicate a transformation for a vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on a prediction mode and / or a transformation index of a target block (CU or sub-block) including a residual block.

[0078] Also, according to an example, when applying MTS to perform a primary transformation, specific basis functions are set to predetermined values, and a mapping relationship for the transformation kernel can be set by combining which basis functions are applied when it is a vertical transformation or a horizontal transformation. For example, when the horizontal transformation kernel is indicated by trTypeHor and the vertical transformation kernel is indicated by trTypeVer, the value 0 of trTypeHor or trTypeVer can be set to DCT2, the value 1 of trTypeHor or trTypeVer can be set to DST7, and the value 2 of trTypeHor or trTypeVer can be set to DCT8.

[0079] In this case, in order to indicate any one of a number of conversion kernel sets, the MTS index information can be encoded and signaled to the decoding device. For example, if the MTS index is 0, it indicates that all the values of trTypeHor and trTypeVer are 0; if the MTS index is 1, it indicates that all the values of trTypeHor and trTypeVer are 1; if the MTS index is 2, it indicates that the value of trTypeHor is 2 and the value of trTypeVer is 1; if the MTS index is 3, it indicates that the value of trTypeHor is 1 and the value of trTypeVer is 2; if the MTS index is 4, it can indicate that all the values of trTypeHor and trTypeVer are 2.

[0080] The conversion unit can perform a secondary conversion based on the (primary) conversion coefficient and derive a corrected (secondary) conversion coefficient (S420). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion means converting to a more compressed representation using the correlation existing between the (primary) conversion coefficients. The secondary conversion can include a non-separable transform. In this case, the secondary conversion can be called a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The non-separable secondary transform can be shown as a transform that performs a secondary conversion on the (primary) conversion coefficient derived through the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, based on the non-separable transform matrix, the vertical conversion and the horizontal conversion cannot be separated (or the horizontal and vertical conversions are independent) and applied to the (primary) conversion coefficient, and the conversion can be applied at once. In other words, the non-separable secondary transform does not separate the vertical component and the horizontal component of the (primary) conversion coefficient. For example, after rearranging a two-dimensional signal (conversion coefficient) into a one-dimensional signal through a specific defined direction (e.g., row-first direction or column-first direction), it can be shown as a conversion method that generates a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix. For example, the row-first order arranges the first row, the second row,..., the Nth row of the MxN block in a column in sequence, and the column-first order arranges the first column, the second column,..., the Mth column of the MxN block in a column in sequence. The non-separable secondary transform can be applied to the top-left region of a block composed of (primary) conversion coefficients (hereinafter, may be called a conversion coefficient block).For example, when both the width (W) and height (H) of the conversion coefficient block are 8 or more, an 8×8 non-separable second-order conversion can be applied to the upper left 8×8 region of the conversion coefficient block. Also, when both the width (W) and height (H) of the conversion coefficient block are 4 or more and either the width (W) or height (H) of the conversion coefficient block is less than 8, a 4×4 non-separable second-order conversion can be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that all of the width (W) or height (H) of the conversion coefficient block is 4 or more is satisfied, a 4×4 non-separable second-order conversion can be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block.

[0081] Specifically, for example, when a 4×4 input block is used, the non-separable second-order conversion is performed as follows.

[0082] The 4×4 input block X can be shown as follows.

[0083]

Number

[0084] When showing X in the form of a vector, the vector JPEG2025090835000004.jpg84 can be shown as follows.

[0085]

Number

[0086] As in Equation 2, the vector JPEG2025090835000006.jpg84 rearranges the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.

[0087] In this case, the second non-separable conversion can be calculated as follows.

[0088]

Number

[0089] Here, JPEG2025090835000008.jpg75 represents the vector of conversion coefficients, and T represents a 16×16 (non-separable) conversion matrix.

[0090] Through the above formula 3, a 16×1 vector of conversion coefficients JPEG2025090835000009.jpg75 can be derived, and the JPEG2025090835000010.jpg75 can be re-organized in 4×4 blocks through the scan order (horizontal, vertical, diagonal, etc.). However, the above calculations are examples, and for reducing the complexity of the non-separable quadratic transformation calculation, HyGT (Hypercube-Givens Transform), etc. may also be used for the non-separable quadratic transformation calculation.

[0091] On the other hand, for the non-separable quadratic transformation, the conversion kernel (or conversion core, conversion type) can be selected as mode dependent. Here, the mode can include the intra prediction mode and / or the inter prediction mode.

[0092] As described above, the non-separable second-order transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 region included inside the corresponding transform coefficient block when both W and H are equal to or greater than 8, and the corresponding 8×8 region may be the upper left 8×8 region inside the corresponding transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 region included inside the corresponding transform coefficient block when both W and H are equal to or greater than 4, and the corresponding 4×4 region may be the upper left 4×4 region inside the corresponding transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.

[0093] At this time, for the selection of the mode-based transform kernel, three non-separable second-order transform kernels can be configured for each transform set for the non-separable second-order transform for both the 8×8 transform and the 4×4 transform, and the number of transform sets can be 35. That is, 35 transform sets can be configured for the 8×8 transform, and 35 transform sets can be configured for the 4×4 transform. In this case, each of the 35 transform sets for the 8×8 transform may include three 8×8 transform kernels, and in this case, each of the 35 transform sets for the 4×4 transform may include three 4×4 transform kernels. However, the size of the transform, the number of sets, and the number of transform kernels in the set are examples, and sizes other than 8×8 or 4×4 may be used, or n sets may be configured, and each set may include k transform kernels.

[0094] The transform set may be referred to as an NSST set, and the transform kernels in the NSST set may be referred to as NSST kernels. The selection of a specific set among the transform sets can be performed based on, for example, the intra prediction mode of the target block (CU or sub-block).

[0095] For reference, for example, the intra prediction mode can include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction mode can include the 0th planar intra prediction mode and the 1st DC intra prediction mode, and the directional intra prediction mode can include 65 intra prediction modes from the 2nd to the 66th. However, this is an example, and this document can also be applied when the number of intra prediction modes is different. On the other hand, optionally, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can indicate the LM (linear model) mode.

[0096] FIG. 5 exemplarily shows the intra-directional mode of 65 prediction directions.

[0097] Referring to FIG. 5, intra prediction modes having horizontal directionality and intra prediction modes having vertical directionality can be classified centering on the 34th intra prediction mode having a prediction direction of the upper left diagonal. H and V in FIG. 5 respectively represent horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in 1 / 32 units on the sample grid position. This can indicate an offset with respect to the mode index value. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. On the other hand, strictly speaking, the 34th intra prediction mode can be regarded as neither horizontal directionality nor vertical directionality, but from the viewpoint of determining the conversion set of the secondary conversion, it can be classified as belonging to horizontal directionality. This is because for the vertical mode symmetric with respect to the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the alignment method of the input data for the horizontal mode is used. Transposing the input data means that for the data MxN of the two-dimensional block, the rows become columns, the columns become rows, and data of NxM is configured. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. Since the 2nd intra prediction mode has a reference pixel on the left and predicts in the upper right direction, it can be called an upper right diagonal intra prediction mode. Along the same line, the 34th intra prediction mode can be called a lower right diagonal intra prediction mode, and the 66th intra prediction mode can be called a lower left diagonal intra prediction mode.

[0098] In this case, the mapping between the 35 conversion sets and the intra prediction mode can be shown, for example, as in the following table. For reference, when the LM mode is applied to the target block, the second-order conversion may not be applied to the target block.

[0099]

Table 2

[0100] On the other hand, once it is determined that a specific set is to be used, one of the k conversion kernels within the specific set can be selected through the index of the non-separable second-order conversion. The encoding device can derive the index of the non-separable second-order conversion that refers to a specific conversion kernel based on rate-distortion (RD) checking, and can signal the index of the non-separable second-order conversion to the decoding device. The decoding device can select one of the k conversion kernels within the specific set based on the index of the non-separable second-order conversion. For example, the index value 0 of the NSST can refer to the first non-separable second-order conversion kernel, the index value 1 of the NSST can refer to the second non-separable second-order conversion kernel, and the index value 2 of the NSST can refer to the third non-separable second-order conversion kernel. Or, the index value 0 of the NSST can indicate that the first non-separable second-order conversion is not applied to the target block, and the index values 1 to 3 of the NSST can refer to the 3 conversion kernels.

[0101] Referring to FIG. 4 again, the conversion unit can perform the non-separable second-order conversion based on the selected conversion kernel and obtain the modified (second-order) conversion coefficients. The modified conversion coefficients can be derived as the conversion coefficients quantized through the quantization unit as described above, and can be encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse conversion unit within the encoding device.

[0102] On the one hand, when the secondary conversion is omitted as described above, the (primary) conversion coefficients that are the output of the primary (separation) conversion can be derived as the quantization coefficients quantized through the quantization unit as described above, encoded, and signaled to the decoding device and transmitted to the inverse quantization / inverse conversion unit in the encoding device.

[0103] The inverse conversion unit can perform a series of procedures in the reverse order of the procedures performed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficients, performs a secondary (inverse) conversion to derive the (primary) conversion coefficients (S450), and performs a primary (inverse) conversion on the (primary) conversion coefficients to obtain a residual block (residual samples) (S460). Here, the primary conversion coefficients can be called modified conversion coefficients from the perspective of the inverse conversion unit. As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0104] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). The secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion can be NSST or RST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on the secondary conversion flag parsed from the bitstream. As another example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficients of the residual block.

[0105] The second inverse transform determination unit can determine the second inverse transform. At this time, the second inverse transform determination unit can determine the second inverse transform applied to the current block based on the NSST (or RST) transform set specified by the intra prediction mode. Also, as an example, the second transform determination method can be determined depending on the first transform determination method. Various combinations of the first transform and the second transform can be determined by the intra prediction mode. Also, as an example, the second inverse transform determination unit can also determine the area to which the second inverse transform is applied based on the size of the current block.

[0106] On the other hand, as described above, when the second (inverse) transform is omitted, the (inverse quantized) transform coefficients can be received and the first (separation) inverse transform can be performed to obtain a residual block (residual sample). As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and can generate a restored picture based on this.

[0107] On the other hand, in this document, in order to reduce the computational amount and memory requirement associated with the non-separable second transform, an RST (reduced secondary transform) in which the size of the transform matrix (kernel) is reduced with the concept of NSST can be applied.

[0108] On the one hand, the conversion kernel, conversion matrix, and coefficients constituting the conversion kernel matrix described in this document, that is, the kernel coefficients or matrix coefficients, can be represented in 8 bits. This can be a condition for implementation in a decoding device and an encoding device, and while there is a reasonably acceptable performance degradation compared to existing 9-bit or 10-bit, the memory requirement for storing the conversion kernel can be reduced. Also, by representing the kernel matrix in 8 bits, a small multiplier can be used, which may be suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.

[0109] In this specification, RST can be meant to refer to a conversion performed on residual samples for a target block based on a transform matrix whose size has been reduced by a simplification factor. When performing a simplified conversion, the amount of computation required during conversion can be reduced due to the reduction in the size of the transform matrix. That is, RST can be used to address the issue of computational complexity that occurs during the conversion of large-sized blocks or non-separable conversions.

[0110] RST can be referred to by various terms such as reduced conversion, reduced transform, reduced secondary transform, reduction transform, simplified transform, simple transform, etc., and the names by which RST can be referred are not limited to the listed examples. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in the conversion block, it may also be referred to as LFNST (Low-Frequency Non-Separable Transform).

[0111] On the other hand, when the second inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 can include an inverse RST unit that derives a transformation coefficient corrected based on the inverse RST for the transformation coefficient, and an inverse primary transformation unit that derives a residual sample for the target block based on the inverse primary transformation for the corrected transformation coefficient. The inverse primary transformation means the inverse transformation of the primary transformation applied to the residual. In this document, deriving a transformation coefficient based on a transformation can mean deriving the transformation coefficient by applying the corresponding transformation.

[0112] FIG. 6 is a diagram for explaining RST according to an embodiment of this document.

[0113] In this specification, the "target block" can mean the current block or the residual block on which coding is performed.

[0114] In RST according to an embodiment, an N-dimensional vector can be mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined, where R is smaller than N. N can mean the square of the length of one side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can mean the R / N value. The simplification factor can be referred to by various terms such as a reduced factor, a reduction factor, a reduced factor, a reduction factor, a simplified factor, a simple factor, etc. On the other hand, R can be referred to as a reduced coefficient, but in some cases, the simplification factor may mean R. Also, in some cases, the simplification factor may mean the N / R value.

[0115] In one embodiment, the simplification factor or coefficient can be signaled via a bitstream, but the embodiments are not limited thereto. For example, a predefined value for the simplification factor or coefficient may be stored in each encoding device 200 and decoding device 300, in which case the simplification factor or coefficient may not need to be signaled separately.

[0116] The size of the simplification transform matrix according to one embodiment is RxN, which is smaller than the size NxN of the normal transform matrix, and can be defined as in Equation 4 below.

[0117]

Equation

[0118] The matrix T in the reduced transform block shown in FIG. 6(a) can mean the matrix T of Equation 4. RxN As shown in FIG. 6(a), when the simplification transform matrix T RxN is multiplied by the residual samples for the target block, the transform coefficients for the target block can be derived.

[0119] In one embodiment, when the size of the block to which the transform is applied is 8x8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6(a) can be expressed by the matrix operation as in Equation 5 below. In this case, the memory and multiplication operations can be reduced by approximately 1 / 4 due to the simplification factor.

[0120]

Equation

[0121] In Equation 5, r1 to r 64can indicate the residual sample for the target block, and more specifically, can be the conversion coefficient generated by applying the first-order conversion. The calculation result of Equation 5, the conversion coefficient c for the target block i can be derived, and c i is derived as shown in Equation 6.

[0122]

Equation

[0123] The calculation result of Equation 6, the conversion coefficients c1 to c for the target block R can be derived. That is, when R = 16, the conversion coefficients c1 to c for the target block 16 can be derived. If a normal (regular) conversion is applied instead of RST, and a conversion matrix with a size of 64x64 (NxN) is multiplied by a residual sample with a size of 64x1 (Nx1), 64 (N) conversion coefficients for the target block should be derived. However, since RST is applied, only 16 (R) conversion coefficients for the target block are derived. Since the total number of conversion coefficients for the target block decreases from N to R, and the amount of data transmitted from the encoding device 200 to the decoding device 300 decreases, the transmission efficiency between the encoding device 200 and the decoding device 300 can be increased.

[0124] From the perspective of the size of the conversion matrix, the size of the normal conversion matrix is 64x64 (NxN), but the size of the simplified conversion matrix decreases to 16x64 (RxN). Therefore, when performing RST compared to performing a normal conversion, the use of memory can be reduced at a ratio of R / N. Also, compared to the number of multiplication operations NxN when using a normal conversion matrix, if a simplified conversion matrix is used, the number of multiplication operations can be reduced at a ratio of R / N (RxN).

[0125] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual samples for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300. The inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and based on an inverse primary conversion for the modified conversion coefficients, can derive residual samples for the target block.

[0126] Inverse RST matrix T according to one embodiment NxR has a size of NxR, which is smaller than the size NxN of the normal inverse conversion matrix, and is related to the simplified conversion matrix T shown in Equation 4 RxN in a transpose relationship.

[0127] The matrix T in the reduced Inv. Transform block shown in (b) of FIG. 6 t can mean the inverse RST matrix T RxN T (the superscript T means transpose). As shown in (b) of FIG. 6, when the inverse RST matrix T RxN T is multiplied by the conversion coefficients for the target block, modified conversion coefficients for the target block or residual samples for the target block can be derived. The inverse RST matrix T RxN T can also be expressed as (T RxN ) T NxR .

[0128] More specifically, when inverse RST is applied as the secondary inverse conversion, the inverse RST matrix T RxN TWhen applied, a modified conversion coefficient for the target block can be derived. On the other hand, inverse RST can be applied as an inverse linear transformation, and in this case, the inverse RST matrix T RxN T is multiplied to derive the residual samples for the target block.

[0129] In one embodiment, when the size of the block to which the inverse transformation is applied is 8x8 and R = 16 (i.e., when R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6(b) can be expressed by the matrix operation of the following Equation 7.

[0130]

Equation

[0131] In Equation 7, c1 to c 16 can indicate the conversion coefficients for the target block. The operation result of Equation 7, r j indicating the modified conversion coefficients for the target block or the residual samples for the target block can be derived, and the derivation process of r j is as shown in Equation 8.

[0132]

Equation

[0133] The operation results of Equation 8, r1 to r Ncan be derived. From the perspective of the size of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64x64 (NxN), but the size of the simplified inverse transformation matrix is reduced to 64x16 (NxR). Therefore, when performing inverse RST compared to performing normal inverse transformation, the memory usage can be reduced at a ratio of R / N. Also, compared to the number of multiplication operations NxN when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced at a ratio of R / N (NxR).

[0134] On the other hand, for 8x8 RST as well, the configuration of the conversion set as shown in Table 2 can be applied. That is, the corresponding 8x8 RST can be applied by the conversion set in Table 2. Since one conversion set is composed of two or three conversions (kernels) depending on the prediction mode within the screen, it can be configured to select one out of a maximum of four conversions including the case where no secondary conversion is applied. The conversion when no secondary conversion is applied can be regarded as the one to which the identity matrix is applied. When indices 0, 1, 2, and 3 are assigned to the four conversions respectively (for example, the 0th index can be assigned to the identity matrix, i.e., the case where no secondary conversion is applied), the syntax element called the NSST index can be signaled for each conversion coefficient block to specify the conversion to be applied. That is, 8x8 NSST can be specified for the upper left 8x8 block of 8x8 via the NSST index, and 8x8 RST can be specified in the RST configuration. 8x8 NSST and 8x8 RST refer to the conversions that can be applied to the 8x8 region included inside the corresponding conversion coefficient block when both W and H of the target block to be converted are equal to or greater than 8, and the corresponding 8x8 region can be the upper left 8x8 region inside the corresponding conversion coefficient block. Similarly, 4x4 NSST and 4x4 RST refer to the conversions that can be applied to the 4x4 region included inside the corresponding conversion coefficient block when both W and H of the target block are equal to or greater than 4, and the corresponding 4x4 region can be the upper left 4x4 region inside the corresponding conversion coefficient block.

[0135] On the one hand, when applying an (forward) 8x8 RST such as Equation 4, 16 valid conversion coefficients are generated. Thus, it can be seen that the 64 input data constituting the 8x8 region are reduced to 16 output data. From the perspective of the two-dimensional region, only the conversion coefficients valid for 1 / 4 of the region are filled. Therefore, the 16 output data obtained by applying the forward 8x8 RST can be filled in the upper left stage region (the 1st to 16th conversion coefficients) of the block as shown in FIG. 7.

[0136] FIG. 7 is a diagram showing the scanning order of conversion coefficients according to an embodiment of this document. As described above, when the forward scanning order starts from No. 1, the reverse scanning can be performed in the arrow direction and order shown in FIG. 7 from the 64th to the 17th in the forward scanning order.

[0137] In FIG. 7, the upper left 4x4 region is the ROI (Region Of Interest) region filled with valid conversion coefficients, and the remaining regions will be left empty. The empty regions can be filled with 0 values as the default.

[0138] That is, when applying an 8x8 RST with a forward conversion matrix form of 16x64 to an 8x8 region, the output conversion coefficients are arranged in the upper left 4x4 region, and the regions where the output conversion coefficients do not exist can be filled with 0 according to the scanning order in FIG. 7 (from the 64th to the 17th).

[0139] If a non-zero valid conversion coefficient is found outside the ROI region in Fig. 7, it is certain that the 8x8 RST is not applied, so the index coding of the corresponding NSST can be omitted. Conversely, if no non-zero conversion coefficient is found outside the ROI region in Fig. 7 (for example, when the 8x8 RST is applied and the conversion coefficients for regions outside the ROI are set to 0), since there is a possibility that the 8x8 RST is applied, the index of the NSST can be coded. Such conditional NSST index coding has to check for the presence or absence of non-zero conversion coefficients, so it can be performed after the residual coding process.

[0140] This document deals with the design of the RST applicable to 4x4 blocks from the RST structure described in this embodiment and related optimization methods. Naturally, for some concepts, they can be applied not only to the 4x4 RST but also to the 8x8 RST or other forms of transformation.

[0141] Fig. 8 is a flowchart showing the inverse RST process according to an embodiment of this document.

[0142] Each step disclosed in Fig. 8 can be performed by the decoding device 300 disclosed in Fig. 3. More specifically, S800 can be performed by the inverse quantization unit 321 disclosed in Fig. 3, and S810 and S820 can be performed by the inverse transformation unit 322 disclosed in Fig. 3. Therefore, specific content that overlaps with the content described above in Fig. 3 will be omitted from the description or simplified. On the other hand, in this document, RST is applied to forward transformation, and inverse RST can mean transformation applied in the inverse direction.

[0143] In one embodiment, the detailed operation by inverse RST is only the reverse of the order of the detailed operation by RST, and the detailed operation by RST and the detailed operation by inverse RST may be substantially similar. Therefore, an ordinary technician in the art can easily understand that the descriptions of S800 to S820 for inverse RST described below can be applied identically or similarly to RST as well.

[0144] A decoding device 300 according to one embodiment can perform inverse quantization on the quantized transform coefficients for a target block to derive the transform coefficients (S800).

[0145] On the other hand, the decoding device 300 can determine whether to apply the inverse secondary transform after the inverse primary transform and before the inverse secondary transform. For example, the inverse secondary transform can be NSST or RST. As an example, the decoding device can determine whether to apply the inverse secondary transform based on the flag of the secondary transform parsed from the bitstream. As another example, the decoding device can also determine whether to apply the inverse secondary transform based on the transform coefficients of the residual block.

[0146] Also, the decoding device 300 can determine the inverse secondary transform. At this time, the decoding device 300 can determine the inverse secondary transform applied to the current block based on the NSST (or RST) transform set specified by the intra prediction mode. Further, as an example, the secondary transform determination method can be determined depending on the primary transform determination method. For example, it can be determined that RST or LFNST is applied only when DCT-2 is applied as the transform kernel in the primary transform. Or, various combinations of the primary transform and the secondary transform can be determined by the intra prediction mode.

[0147] Also, as an example, prior to the step of determining the inverse secondary transform, the decoding device 300 can also determine the region to which the inverse secondary transform is applied based on the size of the current block.

[0148] A decoding device 300 according to an embodiment can select a transform kernel (S810). More specifically, the decoding device 300 can select a transform kernel based on at least one of a transform index, the width and height of the region to which the transform is applied, the intra prediction mode used in video decoding, and information regarding the color component of the target block. However, the embodiment is not limited thereto. For example, the transform kernel may be predefined, and separate information for selecting the transform kernel may not be signaled.

[0149] In one example, information regarding the color component of the target block can be indicated via CIdx. When the target block is a luma block, CIdx can indicate 0, and when the target block is a chroma block, for example, a Cb block or a Cr block, CIdx can indicate a value other than 0 (e.g., 1).

[0150] A decoding device 300 according to an embodiment can apply an inverse RST to the transform coefficients based on the selected transform kernel and a reduced factor (S820).

[0151] Hereinafter, according to an embodiment of this document, a method for determining a secondary NSST set, that is, a secondary transform set or a transform set, considering the intra prediction mode and the block size will be proposed.

[0152] As an embodiment, based on the aforementioned intra prediction mode, by configuring a set for the current transform block, a transform set composed of transform kernels of various sizes can be applied to the transform block. When the transform sets in Table 3 are represented from 0 to 3, it is as shown in Table 4.

[0153]

Table 3

[0154]

Table 4

[0155] The indexes 0, 2, 18, 34 shown in Table 3 respectively correspond to 0, 1, 2, 3 in Table 4. In Tables 3 and 4, instead of 35 conversion sets, only 4 conversion sets are used, which can significantly reduce the memory space.

[0156] Also, the various numbers of conversion kernel matrices that can be included in each conversion set can be set as shown in the following table.

[0157]

Table 5

[0158]

Table 6

[0159]

Table 7

[0160] Table 5 shows that two available conversion kernels are used for each conversion set, and thus the conversion index will have a range from 0 to 2.

[0161] According to Table 6, for conversion set 0, that is, the conversion set for the DC mode and the planar mode among the intra prediction modes, two available conversion kernels are used, and for the remaining conversion sets, one conversion kernel is used respectively. At this time, the available conversion index for conversion set 1 is from 0 to 2, and the conversion indexes for the remaining conversion sets 1 to 3 are from 0 to 1.

[0162] In Table 7, one available conversion kernel is used for each conversion set, whereby the conversion index will have a range from 0 to 1.

[0163] On the other hand, in the mapping of the conversion sets in Table 3, all four conversion sets can be used and the four conversion sets can be rearranged as shown in Table 4 so as to be classified into indexes of 0, 1, 2, and 3. Tables 8 and 9 below exemplarily show the four conversion sets available for the second conversion. Table 8 presents a conversion kernel matrix applicable to an 8x8 block, and Table 9 presents a conversion kernel matrix applicable to a 4x4 block. Tables 8 and 9 are each composed of two conversion kernel matrices per conversion set, and two conversion kernel matrices can be applied for all intra prediction modes as in Table 5.

[0164]

Table 8-1

[0165]

Table 8-2

[0166]

Table 8-3

[0167]

Table 8-4

[0168]

Table 8-5

[0169]

Table 8-6

[0170]

Table 8-7

[0171]

Table 8-8

[0172]

Table 9-1

[0173]

Table 9-2

[0174]

Table 9-3

[0175]

Table 9-4

[0176]

Table 9-5

[0177]

Table 9-6

[0178]

Table 9-7

[0179]

Table 9-8

[0180] The exemplary conversion kernel matrices presented in Table 8 are all conversion kernel matrices multiplied by 128 as the scaling value. In the g_aiNsst8x8[N1][N2]

[16]

[64] array that appears in the matrix array of Table 8, N1 indicates the number of conversion sets (N1 is divided into 4 or 35, indices 0, 1, …, N1 - 1), N2 indicates the number of conversion kernel matrices that make up each conversion set (1 or 2), and

[16]

[64] indicates a 16x64 Reduced Secondary Transform (RST).

[0181] As in Tables 3 and 4, when a certain conversion set is composed of one conversion kernel matrix, in Table 8, either the first or the second conversion kernel matrix can be used for the corresponding conversion set.

[0182] When the corresponding RST is applied, 16 conversion coefficients are output. However, if only the mx64 part of the 16x64 matrix is applied, it can be configured to output only m conversion coefficients. For example, when m = 8, instead of multiplying only the top 8x64 matrix from the top and outputting only 8 conversion coefficients, the computational complexity can be reduced by half. To reduce the computational complexity in the worst case, an 8x64 matrix can be applied to an 8x8 transformation unit (TU).

[0183] The exemplary conversion kernel matrices presented in Table 9 that can be applied to a 4x4 region are all conversion kernel matrices multiplied by 128 as the scaling value. In the g_aiNsst4x4[N1][N2]

[16]

[64] array that appears in the matrix array of Table 9, N1 indicates the number of transform sets (N1 is divided into 4 or 35, indices 0, 1, …, N1 - 1), N2 indicates the number of conversion kernel matrices that make up each transform set (1 or 2), and

[16]

[16] indicates a 16x16 transformation.

[0184] As shown in Tables 3 and 4, when a certain conversion set is composed of one conversion kernel matrix, in Table 9, either the first or the second conversion kernel matrix for the corresponding conversion set can be used.

[0185] Similar to the case of 8x8 RST, when only the mx16 part of the 16x16 matrix is used, it can be configured to output only m conversion coefficients. For example, when m = 8, instead of multiplying only the top 8x16 matrix from the top and outputting only 8 conversion coefficients, the computational complexity can be reduced by half. To reduce the computational complexity in the worst case, an 8x16 matrix can be applied to a 4x4 transformation unit (TU).

[0186] Basically, the conversion kernel matrix applicable to the 4x4 area presented in Table 9 can be applied to a 4x4 TU, a 4xM TU, or an Mx4 TU (in the case of 4xM TU and Mx4 TU, either apply the conversion kernel matrix specified for each divided into 4x4 areas, or apply it only to the largest upper-left 4x8 or 8x4 area), or it can be applied only to the upper-left 4x4 area. When the secondary conversion is configured to be applied only to the upper-left 4x4 area, the conversion kernel matrix that can be applied to the 8x8 area presented in Table 8 may become unnecessary.

[0187] On the other hand, in order to reduce the computational complexity for the worst case, the following embodiments can be proposed. Hereinafter, a matrix composed of M rows and N columns is denoted as an MxN matrix, and the MxN matrix means a conversion matrix applied when performing a forward conversion, that is, a conversion (RST) in an encoding device. Therefore, in the inverse conversion (inverse RST) performed by the decoding device, an NxM matrix obtained by taking the transpose of the MxN matrix can be used.

[0188] 1) For a block (e.g., a conversion unit) with width W and height H, when W≥8 and H≥8, apply the conversion kernel matrix applicable to an 8x8 area to the upper left 8x8 area of the block. When W = 8 and H = 8, only the 8x64 part of the 16x64 matrix can be applied. That is, 8 conversion coefficients can be generated.

[0189] 2) For a block (e.g., a conversion unit) with width W and height H, when one of W and H is smaller than 8, that is, when one of W and H is 4, apply the conversion kernel matrix applicable to a 4x4 area to the upper left of the block. When W = 4 and H = 4, only the 8x16 part of the 16x16 matrix can be applied, and in this case, 8 conversion coefficients are generated.

[0190] If (W, H) = (4, 8) or (8, 4), apply only the second-order conversion to the upper left 4x4 area. When W or H is greater than 8, that is, when W or H is equal to or greater than 16 and the other is 4, apply the second-order conversion only up to the two upper left 4x4 blocks. That is, the conversion kernel matrix specified by dividing into two 4x4 blocks can be applied only up to the maximum upper left 4x8 or 8x4 area.

[0191] 3) For a block (e.g., a conversion unit) with width W and height H, when both W and H are 4, the second-order conversion may not be applied.

[0192] 4) For a block (e.g., a conversion unit) with width W and height H, the number of coefficients generated by applying the second-order conversion can be configured to be maintained at 1 / 4 or less with respect to the area of the conversion unit (i.e., the total number of pixels constituting the conversion unit = WxH). For example, when both W and H are 4, the uppermost 4x16 matrix of the 16x16 matrix can be applied so that 4 conversion coefficients are generated.

[0193] When applying the second-order transformation only to the largest upper-left 8x8 region among all transformation units (TUs), for a 4x8 transformation unit or an 8x4 transformation unit, since no more than 8 coefficients need to be generated, for the upper-left 4x4 region, it can be configured to apply the uppermost 8x16 matrix among the 16x16 matrices. For an 8x8 transformation unit, it can be applied up to a maximum of 16x64 matrix (up to 16 coefficients can be generated), and for a 4xN or Nx4 (N≥16) transformation unit, the 16x16 matrix can be applied to the upper-left 4x4 block, or the uppermost 8x16 matrix among the 16x16 matrices can be applied to two 4x4 blocks located in the upper left. In a similar way, for a 4x8 transformation unit or an 8x4 transformation unit, the uppermost 4x16 matrix among the 16x16 matrices can be applied to two 4x4 blocks located in the upper left respectively, and all 8 transformation coefficients can be generated.

[0194] 5) The maximum size of the second-order transformation applied to the 4x4 region can be limited to 8x16. In this case, the amount of memory required to store the transformation kernel matrix applied to the 4x4 region can be reduced to half compared to the 16x16 matrix.

[0195] For example, for all the transformation kernel matrices presented in Table 9, only the uppermost 8x16 matrix among the 16x16 matrices can be extracted respectively, the maximum size can be limited to 8x16, and it can be realized to store only the corresponding 8x16 matrix of the transformation kernel matrix in the actual video coding system.

[0196] If the maximum applicable transformation size is 8x16 and the maximum number of multiplications required to generate one coefficient is limited to 8, for a 4x4 block, the maximum 8x16 matrix can be applied, and for a 4xN block or an Nx4 block (N≥8, N = 2 n 、n≥3), the maximum 8x16 matrix can be applied to the two largest upper-left 4x4 blocks that make up the inside respectively. For example, for a 4xN block or an Nx4 block (N≥8, N = 2 nFor n ≥ 3), an 8x16 matrix can be stored for one 4x4 block in the upper left segment.

[0197] According to one embodiment, when coding an index specifying a secondary transformation to be applied to the luma component, more specifically, when one set of transformations is composed of two transformation kernel matrices, it is necessary to specify whether to apply the secondary transformation and, if applicable, which transformation kernel matrix to apply. For example, when not applying the secondary transformation, the transformation index can be coded as 0, and when applying it, the transformation indices for the two sets of transformations can be coded as 1 and 2, respectively.

[0198] In this case, when coding the transformation index, truncated unary coding can be used. For example, binary codes 0, 10, 11 can be assigned to the transformation indices 0, 1, 2, respectively, for coding.

[0199] Also, when coded in the truncated unary manner, different CABAC contexts can be assigned to each bin. When coding the transformation indices 0, 10, 11 according to the above-described example, two CABAC contexts can be used.

[0200] On the other hand, when coding a transformation index specifying a secondary transformation to be applied to the chrominance component, more specifically, when one set of transformations is composed of two transformation kernel matrices, it is necessary to specify whether to apply the secondary transformation and, if applicable, which transformation kernel matrix to apply, in the same way as when coding the transformation index for the secondary transformation for the luma component. For example, when not applying the secondary transformation, the transformation index can be coded as 0, and when applying it, the transformation indices for the two sets of transformations can be coded as 1 and 2, respectively.

[0201] In this case, when coding the conversion index, truncated unary coding can be used. For example, binary codes 0, 10, and 11 can be assigned to conversion indices 0, 1, and 2 respectively for coding.

[0202] Also, when coded in the truncated unary method, different CABAC contexts can be assigned to each bin. According to the above examples, when coding conversion indices 0, 10, and 11, two CABAC contexts can be used.

[0203] Moreover, according to one embodiment, different CABAC context sets can be assigned according to the chroma intra prediction mode. For example, when distinguishing between non-directional modes such as in the planar mode or DC mode and other directional modes (i.e., when dividing into two groups), when coding 0, 10, and 11 as in the above examples, the corresponding CABAC context sets (each composed of two contexts) can be assigned for each group.

[0204] Thus, when splitting the chroma intra prediction mode into several groups and assigning the corresponding CABAC context set, the chroma intra prediction mode value should be found before the transform index coding for the second transform. However, in the case of the Chroma direct mode (DM), since the luma intra prediction mode value is used as it is, the intra prediction mode value for the luma component should also be found. Therefore, when coding information for the chrominance component, data dependency may occur for the information of the luma component. Thus, in the case of the chroma DM mode, when performing transform index coding for the second transform without information on the intra prediction mode, it can be mapped to a specific group to remove the aforementioned data dependency. For example, if the chroma intra prediction mode is the chroma DM mode, it can be regarded as the planar mode or the DC mode, use the corresponding CABAC context set, perform the corresponding transform index coding, or regard it as another directional mode and apply the corresponding CABAC context set.

[0205] FIG. 9 is a flowchart showing the operation of a video decoding apparatus according to an embodiment of the present document.

[0206] Each step disclosed in FIG. 9 can be performed by the decoding apparatus 300 disclosed in FIG. 3. More specifically, S910 can be performed by the entropy decoding unit 310 disclosed in FIG. 3, S920 can be performed by the inverse quantization unit 321 disclosed in FIG. 3, S930 and S940 can be performed by the inverse transform unit 322 disclosed in FIG. 3, and S950 can be performed by the addition unit 340 disclosed in FIG. 3. Also, the operations according to S910 to S950 are based on a part of the content described above with reference to FIGS. 4 to 8. Therefore, specific content that overlaps with the content described above with reference to FIGS. 3 to 8 will be omitted or simplified in the description.

[0207] A decoding device 300 according to an embodiment can derive quantized transform coefficients for a target block from a bitstream (S910). More specifically, the decoding device 300 can decode information regarding the quantized transform coefficients for the target block from the bitstream, and based on the information regarding the quantized transform coefficients for the target block, can derive the quantized transform coefficients for the target block. The information regarding the quantized transform coefficients for the target block can be included in an SPS (Sequence Parameter Set) or a slice header, and can include at least one of information regarding whether a reduced transform (RST) is applied, information regarding a reduction factor, information regarding the minimum transform size to which the reduced transform is applied, information regarding the maximum transform size to which the reduced transform is applied, a reduced inverse transform size, and information regarding a transform index indicating any one of transform kernel matrices included in a transform set.

[0208] A decoding device 300 according to an embodiment can perform inverse quantization on the quantized transform coefficients for the target block to derive transform coefficients (S920).

[0209] A decoding device 300 according to an embodiment can derive modified transform coefficients based on an inverse RST (reduced secondary transform) for the transform coefficients (S930).

[0210] In one example, the inverse RST can be performed based on an inverse RST matrix, and the inverse RST matrix can be a non-square matrix in which the number of columns is less than the number of rows.

[0211] In one embodiment, S930 may include steps of decoding a conversion index, determining whether a condition for applying an inverse RST is met based on the conversion index, selecting a conversion kernel matrix, and when the condition for applying the inverse RST is met, applying the inverse RST to the conversion coefficients based on the selected conversion kernel matrix and / or a simplification factor. At this time, the size of the simplified inverse transformation matrix can be determined based on the simplification factor.

[0212] The decoding apparatus 300 according to one embodiment can derive residual samples for a target block based on an inverse transformation for the modified conversion coefficients (S940).

[0213] The decoding apparatus 300 can perform an inverse first-order transformation on the modified conversion coefficients for the target block. At this time, the inverse first-order transformation may apply a simplified inverse transformation or use a normal separable transformation.

[0214] The decoding apparatus 300 according to one embodiment can generate restored samples based on the residual samples for the target block and the predicted samples for the target block (S950).

[0215] Referring to S930, it can be confirmed that residual samples for the target block are derived based on the inverse RST for the conversion coefficient for the target block. Considering from the perspective of the size of the inverse transform matrix, the size of the normal inverse transform matrix is NxN, but the size of the inverse RST matrix decreases to NxR. Therefore, compared with the case of performing normal conversion, the memory usage can be reduced at a ratio of R / N when performing inverse RST. Also, compared with the number of multiplication operations NxN when using the normal inverse transform matrix, when using the inverse RST matrix, the number of multiplication operations can be reduced (NxR) at a ratio of R / N. Further, when applying inverse RST, only R conversion coefficients need to be decoded. Therefore, compared with the case where N conversion coefficients must be decoded when normal inverse conversion is applied, the total number of conversion coefficients for the target block decreases from N to R, and the decoding efficiency can be increased. In summary, according to S930, the (inverse) conversion efficiency and decoding efficiency of the decoding apparatus 300 can be increased via inverse RST.

[0216] FIG. 10 is a control flowchart for explaining inverse RST according to an embodiment of the present document.

[0217] The decoding apparatus 300 receives the conversion index and information for the intra prediction mode from the bit stream (S1000).

[0218] Such information is received as syntax information, and the syntax information is received as a binary bit string including 0 and 1.

[0219] On the other hand, the entropy decoding unit 310 can derive binary information for the syntax element of the conversion index.

[0220] This generates a candidate set for the binary values that the syntax elements of the received transform index can have. When following this embodiment, the syntax elements of the transform index can be binary-coded in a truncated unary code system.

[0221] The syntax elements of the transform index according to this embodiment can indicate whether inverse RST is applied and any one of the transform kernel matrices included in the transform set. When the transform set includes two transform kernel matrices, the value of the syntax element of the transform index can be three.

[0222] That is, according to one embodiment, the syntax element values for the transform index can include 0 indicating the case where inverse RST is not applied to the target block, 1 indicating the first transform kernel matrix among the transform kernel matrices, and 2 indicating the second transform kernel matrix among the transform kernel matrices.

[0223] In this case, the syntax element values for the three transform indexes can be coded as 0, 10, 11 by the truncated unary code system. That is, the value 0 for the syntax element can be binary-coded as "0", the value 1 for the syntax element can be binary-coded as "10", and the value 2 for the syntax element can be binary-coded as "11".

[0224] The entropy decoding unit 310 derives context information, that is, a context model, for the bin string of the transform index (S1010), and can decode the bins of the bin string of the syntax element based on the context information (S1020).

[0225] In summary, the entropy decoding unit 310 receives a bin string binary-coded in the truncated unary code system, and decodes the syntax elements of the transform index through the candidate set for the corresponding binary value.

[0226] According to this embodiment, for two bins of the conversion index, different context information, that is, probability models, can be applied respectively. That is, the two bins of the conversion index can both be decoded in a context mode rather than a bypass mode. Among the bins of the syntax elements for the conversion index, the first bin is decoded based on the first context information, and the second bin among the bins of the syntax elements for the conversion index can be decoded based on the second context information.

[0227] Among the binary values that the syntax elements of the conversion index may have by such context information-based decoding, the value of the syntax element for the conversion index applied to the target block can be derived (S1030).

[0228] That is, it can be derived whether any one of the conversion indexes 0, 1, and 2 is applied to the current target block.

[0229] The inverse conversion unit 332 of the decoding device 300 determines a conversion set based on the mapping relationship according to the intra prediction mode applied to the target block (S1040), and can perform inverse RST based on the conversion set and the value of the syntax element for the conversion index (S1050).

[0230] As described above, a plurality of conversion sets can be determined according to the intra prediction mode of the conversion block to be converted, and the inverse RST can be performed based on any one of the conversion kernel matrices included in the conversion set indicated by the conversion index.

[0231] FIG. 11 is a flowchart showing the operation of a video encoding device according to an embodiment of this document.

[0232] Each step disclosed in FIG. 11 can be performed by the encoding device 200 disclosed in FIG. 2. More specifically, S1110 can be performed by the prediction unit 220 disclosed in FIG. 2, S1120 can be performed by the subtraction unit 231 disclosed in FIG. 2, S1130 and S1140 can be performed by the conversion unit 232 disclosed in FIG. 2, and S1150 can be performed by the quantization unit 233 and the entropy encoding unit 240 disclosed in FIG. 2. Also, the operations according to S1110 to S1150 are based on a part of the content described above with reference to FIGS. 4 to 8. Therefore, the specific content overlapping with the content described above with reference to FIGS. 2 and 4 to 8 will be omitted or simplified in the description.

[0233] An encoding device 200 according to an embodiment can derive a prediction sample based on an intra prediction mode applied to a target block (S1110).

[0234] An encoding device 200 according to an embodiment can derive a residual sample for a target block (S1120).

[0235] An encoding device 200 according to an embodiment can derive a conversion coefficient for the target block based on a primary conversion of the residual sample (S1130). The primary conversion can be performed through a plurality of conversion kernels, and in this case, the conversion kernel can be selected based on the intra prediction mode.

[0236] The decoding device 300 can perform a secondary conversion, specifically NSST, on the conversion coefficient for the target block. At this time, NSST can be performed based on a simplified transform (RST) or without being based on RST. When NSST is performed based on RST, it can correspond to the operation according to S1140.

[0237] An encoding device 200 according to an embodiment can derive a modified transform coefficient for a target block based on an RST for the transform coefficient (S1140). In one example, the RST can be performed based on a simplified transform matrix or a transform kernel matrix, and the simplified transform matrix can be a non-square matrix in which the number of rows is less than the number of columns.

[0238] In one embodiment, S1140 can include a step of determining whether the condition for applying the RST is met, a step of generating and encoding a transform index based on the determination, a step of selecting a transform kernel matrix, and a step of applying the RST to the residual samples based on the selected transform kernel matrix and / or a simplification factor when the condition for applying the RST is met. At this time, the size of the simplified transform kernel matrix can be determined based on the simplification factor.

[0239] An encoding device 200 according to an embodiment can perform quantization based on the modified transform coefficient for the target block, derive a quantized transform coefficient, and encode information regarding the quantized transform coefficient (S1160).

[0240] More specifically, the encoding device 200 can generate information regarding the quantized transform coefficient and encode the generated information regarding the quantized transform coefficient.

[0241] In one example, the information regarding the quantized transform coefficient can include at least one of information regarding whether the RST is applied, information regarding the simplification factor, information regarding the minimum transform size to which the RST is applied, and information regarding the maximum transform size to which the RST is applied.

[0242] Referring to S1140, it can be confirmed that the conversion coefficient for the target block is derived based on the RST for the residual sample. Considering from the perspective of the size of the conversion kernel matrix, the size of the normal conversion kernel matrix is NxN, but the size of the simplified conversion matrix is reduced to RxN. Therefore, when compared with performing normal conversion, the memory usage when performing RST can be reduced at a ratio of R / N. Also, when compared with the number of multiplication operations NxN when using the normal conversion kernel matrix, when using the simplified conversion kernel matrix, the number of multiplication operations can be reduced (RxN) at a ratio of R / N. Furthermore, since only R conversion coefficients are derived when RST is applied, when normal conversion is applied, compared with the derivation of N conversion coefficients, the total number of conversion coefficients for the target block is reduced from N to R, and the amount of data transmitted from the encoding device 200 to the decoding device 300 can be reduced. In summary, according to S1140, the conversion efficiency and coding efficiency of the encoding device 200 can be increased through RST.

[0243] FIG. 12 is a control flowchart for explaining RST according to an embodiment of this document.

[0244] First, the encoding device 200 can determine a conversion set based on the mapping relationship according to the intra prediction mode applied to the target block (S1200).

[0245] Thereafter, the conversion unit 232 can derive the conversion coefficient by performing RST based on any one of the conversion kernel matrices included in the conversion set (S1210).

[0246] In this embodiment, the conversion coefficient is a modified conversion coefficient after the secondary conversion is performed after the primary conversion, and each conversion set may include two conversion kernel matrices.

[0247] When RST is performed in this way, information regarding RST can be encoded by the entropy encoding unit 240.

[0248] First, the entropy encoding unit 240 can derive a syntax element value for a transform index that indicates any one of the transform kernel matrices included in the transform set (S1220).

[0249] The syntax element of the transform index according to this embodiment can indicate whether (inverse) RST is applied and any one of the transform kernel matrices included in the transform set. When the transform set includes two transform kernel matrices, the value of the syntax element of the transform index can be three.

[0250] According to one embodiment, the syntax element value for the transform index can be derived as 0 indicating that (inverse) RST is not applied to the target block, 1 indicating the first transform kernel matrix among the transform kernel matrices, and 2 indicating the second transform kernel matrix among the transform kernel matrices.

[0251] Thereafter, the entropy encoding unit 240 can binaryize the derived syntax element value for the transform index (S1230).

[0252] The entropy encoding unit 240 can binaryize the syntax element values for the three transform indexes by the truncated unary code method as 0, 10, and 11. That is, the value 0 for the syntax element can be binaryized as "0", the value 1 for the syntax element can be binaryized as "10", and the value 2 for the syntax element can be binaryized as "11". The entropy encoding unit 240 can binaryize the derived syntax element for the transform index with any one of "0", "10", and "11".

[0253] The entropy encoding unit 240 derives context information for the bin string of the conversion index, that is, a context model (S1240), and can encode the bins of the bin string of the syntax element based on the context information (S1250).

[0254] According to this embodiment, different context information can be applied to two bins of the conversion index respectively. That is, the two bins of the conversion index can all be encoded in a context method that is not the bypass method. Among the bins of the syntax element for the conversion index, the first bin can be encoded based on the first context information, and among the bins of the syntax element for the conversion index, the second bin can be encoded based on the second context information.

[0255] The encoded bin string of the syntax element can be output to the decoding device 300 or externally in the form of a bit stream.

[0256] In the foregoing embodiments, the method is described based on a flowchart as a series of steps or blocks. However, this document is not limited to the order of the steps. A certain step can occur in a different order from the steps described above, or simultaneously. Also, those skilled in the art should understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps of the flowchart can be deleted without affecting the scope of this document.

[0257] The method according to the foregoing document can be implemented in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0258] When the embodiments in this document are implemented by software, the above-described methods can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each figure can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip.

[0259] In addition, the decoding device and the encoding device to which this document is applied can be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a videophone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, the OTT video (Over the top video) device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.

[0260] In addition, the processing method to which this document is applied can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by a computer. The multimedia data having the data structure according to the present invention can also be stored in a recording medium readable by a computer. The recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by a computer is stored. The recording medium readable by the computer may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Further, the recording medium readable by the computer includes a medium realized in the form of a carrier wave (for example, transmission via the Internet). Also, a bit stream generated by an encoding method can be stored in a recording medium readable by a computer or transmitted via a wired or wireless communication network. Further, the embodiments of this document can be realized by a computer program product with program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0261] FIG. 13 exemplarily shows the structure of a content streaming system to which this document is applied.

[0262] In addition, the content streaming system to which this document is applied can include a large encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0263] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, camcorders, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, camcorders, etc. directly generate a bitstream, the encoding server can be omitted. The bitstream can be generated by the encoding method or the bitstream generation method to which this document applies, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0264] The streaming server transmits multimedia data to the user device based on the user's request via the web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a service desired by the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device in the content streaming system.

[0265] The streaming server can receive content from the media storage and / or the encoding server. For example, when it is to receive content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0266] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, a smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like. Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.

Claims

1. 1. A video decoding method performed by a decoding device, comprising: receiving residual information for a current block from a bitstream; deriving transform coefficients for the current block based on the residual information; deriving modified transform coefficients based on an inverse secondary transform on the transform coefficients; and deriving residual samples for the current block based on an inverse linear transform on the modified transform coefficients; The step of deriving the modified transform coefficients comprises: deriving a transformation kernel matrix; performing a matrix operation on the transformation coefficients and the transformation kernel matrix; The transformation kernel matrix is ​​derived based on a transformation index and a transformation set; The transformation index represents at least one of first index information indicating that the inverse secondary transformation is not applied to the current block, second index information indicating a first transformation kernel matrix as a transformation kernel matrix for the inverse secondary transformation, and third index information indicating a secondary transformation kernel matrix as the transformation kernel matrix for the inverse secondary transformation; The inverse secondary transformation includes performing a matrix operation on input data for an upper left region of the target block and deriving output data that is greater than the input data through the transformation kernel matrix; 11. A method for decoding an image, comprising: the input data for the top left region is in a top left 4x4 region of the target block; and the output data for the top left region is in a top left 4x4 region of the target block or a top left 8x8 region of the target block.

2. A video encoding method performed by an encoding device, comprising: deriving a residual sample for a current block; deriving transform coefficients for the current block based on a linear transform on the residual samples; deriving modified transform coefficients based on a quadratic transformation of the transform coefficients; encoding residual information associated with the modified transform coefficients and a transform index associated with a transform kernel matrix; The step of deriving the modified transform coefficients comprises: determining a transformation set and the transformation kernel matrix; performing a matrix operation on the transformation coefficients and the transformation kernel matrix; the transformation index represents at least one of first index information indicating that the secondary transformation is not applied to the current block, second index information indicating a first transformation kernel matrix as a transformation kernel matrix for the secondary transformation, and third index information indicating a secondary transformation kernel matrix as the transformation kernel matrix for the secondary transformation; The secondary transformation includes performing a matrix operation on input data for an upper left region of the target block and deriving output data that is smaller than the input data through the transformation kernel matrix; 11. The method of claim 10, wherein the input data for the top left region is in a top left 4x4 region of the target block or a top left 8x8 region of the target block, and the output data for the top left region is in a top left 4x4 region of the target block.

3. A method for transmitting data for a video, comprising the steps of: obtaining a bitstream for the video, the bitstream being generated based on: deriving residual samples for a current block, deriving transform coefficients for the current block based on a linear transformation of the residual samples, deriving modified transform coefficients based on a secondary transformation of the transform coefficients, and encoding residual information associated with the modified transform coefficients and transform indices associated with a transform kernel matrix to output the bitstream; transmitting the data including the bitstream; Deriving the modified transform coefficients comprises: determining a transformation set and the transformation kernel matrix; performing a matrix operation on the transform coefficients and the transform kernel matrix; the transformation index represents at least one of first index information indicating that the secondary transformation is not applied to the current block, second index information indicating a first transformation kernel matrix as a transformation kernel matrix for the secondary transformation, and third index information indicating a secondary transformation kernel matrix as the transformation kernel matrix for the secondary transformation; The secondary transformation is a process of performing a matrix operation on input data for the upper left region of the target block and deriving output data that is smaller than the input data through the transformation kernel matrix.

13. A method for transmitting data, wherein the input data for the upper left region is in a 4x4 region at the top left of the target block or a 8x8 region at the top left of the target block, and the output data for the upper left region is in a 4x4 region at the top left of the target block.

Citation Information

Patent Citations

  • Non-separable secondary transform for video coding

    WO2017058614A1

  • Transform method in image coding system and apparatus for same

    WO2018174402A1

  • Method and apparatus for video coding

    WO2019173522A1

  • Method for encoding / decoding video signal, and apparatus therefor

    WO2020050665A1