Image coding method based on transformation and device therefor

The use of reduced and simplified transform matrices in video coding addresses the inefficiencies of high-resolution image data transmission and storage by enhancing compression efficiency and reducing data volume.

JP2025164943APending Publication Date: 2025-10-30LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025144598
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-12-15
Filing Date
2025-09-01
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to a significant increase in the amount of data, resulting in higher transmission and storage costs due to inefficient video coding techniques.

Method used

A video coding method and apparatus utilizing reduced transform matrices, where the inverse transform matrix has fewer columns than rows, and simplified transform matrices with fewer rows than columns, to improve efficiency in residual coding.

Benefits of technology

This approach enhances image/video compression efficiency by reducing data transmission and improving residual coding efficiency, concentrating non-zero transform coefficients in low frequency components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025164943000001_ABST
    Figure 2025164943000001_ABST
Patent Text Reader

Abstract

To provide an image decoding method performed by means of a decoding device.SOLUTION: An image decoding method performed by means of a decoding device according to the present invention includes the steps of: deriving quantized transform coefficients with respect to a target block from a bitstream; performing inverse quantization with respect to the quantized transform coefficients with respect to the target block and deriving transform coefficients; deriving residual samples with respect to the target block based on reduced inverse transform with respect to the transform coefficients; and generating a reconstructed picture based on the residual samples with respect to the target block and prediction samples with respect to the target block. The reduced inverse transform is performed based on a reduced inverse transform matrix, and the reduced inverse transform matrix is a non-square matrix of which the number of columns is smaller than the number of rows.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video coding technology, and more particularly to a transform-based video coding method and apparatus in a video coding system. [Background technology]

[0002] Recently, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the resolution and quality of image data increases, the amount of information or bits to be transmitted increases relatively compared to existing image data. Therefore, when image data is transmitted using a medium such as an existing wired or wireless broadband line or when image data is stored using an existing storage medium, transmission costs and storage costs increase.

[0003] This requires highly efficient image compression techniques to effectively transmit, store, and reproduce high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] SUMMARY OF THE INVENTION A technical object of the present invention is to provide a method and apparatus for improving video coding efficiency.

[0005] Another technical object of the present invention is to provide a method and apparatus for increasing the conversion efficiency.

[0006] Another technical object of the present invention is to provide a method and apparatus for increasing the efficiency of residual coding through transformation.

[0007] Another technical object of the present invention is to provide a video coding method and apparatus based on a reduced transform. [Means for solving the problem]

[0008] According to an embodiment of the present invention, there is provided a video decoding method executed by a decoding device, the method including the steps of: deriving quantized transform coefficients for a current block from a bitstream; deriving transform coefficients by performing inverse quantization on the quantized transform coefficients for the current block; deriving residual samples for the current block based on a reduced inverse transform of the transform coefficients; and generating a reconstructed picture based on the residual samples for the current block and predicted samples for the current block, wherein the reduced inverse transform is performed based on a reduced inverse transform matrix, and the reduced inverse transform matrix is ​​a non-square matrix having a number of columns less than a number of rows.

[0009] According to another embodiment of the present invention, there is provided a video encoding method performed by an encoding apparatus, the method including the steps of: deriving residual samples for a current block; deriving transform coefficients for the current block based on a reduced transform of the residual samples; performing quantization on the transform coefficients for the current block to derive quantized transform coefficients; and encoding information on the quantized transform coefficients, wherein the simplified transform is performed based on a simplified transform matrix, and the simplified transform matrix is ​​a non-square matrix having a number of rows that is less than the number of columns.

[0010] According to another embodiment of the present invention, there is provided a decoding device for performing video decoding, the decoding device including: an entropy decoding unit for deriving quantized transform coefficients for a current block from a bitstream; an inverse quantization unit for performing inverse quantization on the quantized transform coefficients for the current block to derive transform coefficients; an inverse transform unit for deriving residual samples for the current block based on a simplified inverse transform of the transform coefficients; and an adder for generating a reconstructed picture based on the residual samples for the current block and predicted samples for the current block, wherein the simplified inverse transform is performed based on a simplified inverse transform matrix, and the simplified inverse transform matrix is ​​a non-square matrix having a number of columns less than a number of rows.

[0011] According to another embodiment of the present invention, there is provided an encoding apparatus for performing video encoding, the encoding apparatus including: a subtraction unit for deriving residual samples for a current block; a transformation unit for deriving transform coefficients for the current block based on a reduced transform of the residual samples; a quantization unit for performing quantization on the transform coefficients for the current block to derive quantized transform coefficients; and an entropy encoding unit for encoding information on the quantized transform coefficients, wherein the simplified transformation is performed based on a simplified transformation matrix, and the simplified transformation matrix is ​​a non-square matrix having a number of rows that is less than the number of columns. [Effects of the Invention]

[0012] According to the present invention, the overall image / video compression efficiency can be improved.

[0013] According to the present invention, the amount of data to be transmitted for residual processing can be reduced through efficient conversion, thereby improving residual coding efficiency.

[0014] According to the present invention, non-zero transform coefficients can be concentrated in low frequency components through a quadratic transform in the frequency domain.

[0015] According to the present invention, video coding can be performed based on simplified transformation to improve video coding efficiency. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a diagram illustrating a configuration of a video / image encoding device to which the present invention can be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image decoding device to which the present invention can be applied. [Figure 3] 1 illustrates a schematic diagram of a multiple conversion technique according to one embodiment. [Figure 4] 65 intra-directional modes of prediction directions are shown exemplarily. [Figure 5a] 1 is a flowchart illustrating a non-separable quadratic transformation process according to an embodiment. [Figure 5b] 1 is a flowchart illustrating a non-separable quadratic transformation process according to an embodiment. [Figure 5c] 1 is a flowchart illustrating a non-separable quadratic transformation process according to an embodiment. [Figure 6] 1 is a diagram illustrating a simplified conversion according to an embodiment of the present invention; [Figure 7] 1 is a flow chart illustrating a simplified conversion process according to an embodiment of the present invention. [Figure 8] 10 is a flowchart illustrating a simplified conversion process according to another embodiment of the present invention. [Figure 9] 1 is a flowchart illustrating a simplified transformation process based on a non-separable quadratic transformation according to an embodiment of the present invention. [Figure 10] 3 illustrates a block to which a simplifying transformation is applied according to one embodiment of the present invention. [Figure 11] 4 is a flowchart illustrating the operation of a video encoding apparatus according to an embodiment of the present invention. [Figure 12]1 is a flowchart showing the operation of a video decoding device according to an embodiment of the present invention.

[0017] According to an embodiment of the present disclosure, there is provided a video decoding method executed by a decoding device, the method including the steps of: deriving quantized transform coefficients for a current block from a bitstream; deriving transform coefficients by performing inverse quantization on the quantized transform coefficients for the current block; deriving residual samples for the current block based on a reduced inverse transform of the transform coefficients; and generating a reconstructed picture based on the residual samples for the current block and predicted samples for the current block, wherein the reduced inverse transform is performed based on a reduced inverse transform matrix, and the reduced inverse transform matrix is ​​a non-square matrix having a number of columns less than a number of rows. DETAILED DESCRIPTION OF THE INVENTION

[0018] The present invention may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments. The terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical spirit of the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0019] Meanwhile, each component in the drawings described in the present invention is illustrated independently for the convenience of explaining different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0020] The following description may be applied to technical fields dealing with video, images, or pictures. For example, the methods or embodiments disclosed in the following description may be related to the launch of the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), a next-generation video / image coding standard after VVC, or a standard before VVC (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265)).

[0021] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. In the following, the same reference numerals are used to designate the same components in the drawings, and redundant description of the same components will be omitted.

[0022] In this specification, "video" refers to a collection of a series of images over time. "Picture" generally refers to a unit representing one image at a specific time period, and "slice" refers to a unit constituting a part of a picture in coding. One picture may be composed of multiple slices, and pictures and slices may be used interchangeably as needed.

[0023] A pixel or a pel may refer to the smallest unit constituting one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample generally refers to a pixel or a pixel value, and may refer to only the value of a pixel / pixel of a luminance (luma) component, or may refer to only the value of a pixel / pixel of a chroma component.

[0024] A unit refers to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information about the region. The term unit may be used interchangeably with terms such as block or area. In general, an MxN block may refer to a set of samples or transform coefficients consisting of M columns and N rows.

[0025] 1 is a diagram illustrating the configuration of a video / image encoding apparatus to which the present invention can be applied. Hereinafter, the encoding apparatus may include a video encoding apparatus and / or an image encoding apparatus, and the video encoding apparatus may also be used as a concept including an image encoding apparatus.

[0026] 1, the video encoding apparatus 100 may include a picture partitioning module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filtering module 150, and a memory 160. The residual processing module 120 may include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, a dequantization module 125, and an inverse transform module 126.

[0027] The picture division unit 105 can divide an input picture into at least one processing unit.

[0028] For example, the processing unit is called a coding unit (CU). In this case, the coding units may be recursively divided from the largest coding unit (LCU) according to a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and the ternary tree structure. Alternatively, the binary tree structure / ternary tree structure may be applied first. The coding procedure according to the present invention may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, conversion, and restoration, which will be described later.

[0029] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding units may be split into coding units of deeper depths using a quadtree structure, starting from the largest coding unit (LCU). In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to video characteristics, or the coding unit may be recursively split into coding units of lower depths as needed, and the coding unit of the optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that serves as the basis for partitioning or division into prediction units or transform units. A prediction unit is a unit that is partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into subblocks. A transform unit can be divided from a coding unit according to a quadtree structure and is a unit that derives transform coefficients and / or a unit that derives a residual signal from the transform coefficients. Hereinafter, a coding unit is also referred to as a coding block (CB), a prediction unit is also referred to as a prediction block (PB), and a transform unit is also referred to as a transform block (TB). A prediction block or a prediction unit refers to a specific region in a block form within a picture and can include an array of prediction samples.Also, a transform block or transform unit refers to a specific region in a picture in the form of a block, and can include an array of transform coefficients or residual samples.

[0030] The prediction unit 110 performs prediction on a current block (hereinafter, may refer to a current block or a residual block) and generates a predicted block including prediction samples for the current block. The prediction unit 110 performs prediction on a coding block, a transform block, or a prediction block.

[0031] The predictor 110 may determine whether intra prediction or inter prediction is applied to the current block. For example, the predictor 110 may determine whether intra prediction or inter prediction is applied to each CU.

[0032] In intra prediction, the predictor 110 may derive a prediction sample for a current block based on a reference sample outside the current block within a picture to which the current block belongs (hereinafter, the current picture). In this case, the predictor 110 may (i) derive a prediction sample based on an average or interpolation of neighboring reference samples of the current block, or (ii) derive a prediction sample based on a reference sample present in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. (i) is referred to as a non-directional mode or a non-angular mode, and (ii) is referred to as a directional mode or an angular mode. Prediction modes in intra prediction may include, for example, 33 directional prediction modes and at least two or more non-directional modes. Non-directional modes may include a DC prediction mode and a planar mode. The predictor 110 may also determine a prediction mode to be applied to the current block using a prediction mode applied to a neighboring block.

[0033] In the case of inter prediction, the predictor 110 may derive a predicted sample for the current block based on a sample identified by a motion vector on a reference picture. The predictor 110 may derive a predicted sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the skip mode and the merge mode, the predictor 110 may use motion information of a neighboring block as motion information of the current block. In the skip mode, unlike the merge mode, a difference (residual) between a predicted sample and an original sample is not transmitted. In the MVP mode, the motion vector of the current block may be derived by using the motion vector of the neighboring block as a motion vector predictor.

[0034] In the case of inter-prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. The reference picture including the temporal neighboring blocks is also called a collocated picture (colPic). Motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and output in the form of a bitstream.

[0035] When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture. Reference pictures included in the reference picture list can be sorted based on the POC (Picture Order Count) difference between the current picture and the corresponding reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.

[0036] The subtractor 121 generates residual samples, which are the differences between the original samples and the predicted samples. When the skip mode is applied, the residual samples are not generated as described above.

[0037] The transform unit 122 transforms residual samples in units of transform blocks to generate transform coefficients. The transform unit 122 may perform the transform according to the size of the corresponding transform block and a prediction mode applied to a coding block or a prediction block spatially overlapping with the corresponding transform block. For example, if intra prediction is applied to the coding block or the prediction block overlapping with the transform block and the transform block is a 4x4 residual array, the residual samples may be transformed using a Discrete Sine Transform (DST) transform kernel; otherwise, the residual samples may be transformed using a Discrete Cosine Transform (DCT) transform kernel.

[0038] The quantization unit 123 can quantize the transform coefficients to generate quantized transform coefficients.

[0039] The rearrangement unit 124 rearranges the quantized transform coefficients. The rearrangement unit 124 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form through a coefficient scanning method. Here, the rearrangement unit 124 has been described as a separate component, but it may also be a part of the quantization unit 123.

[0040] The entropy encoding unit 130 may perform entropy encoding on the quantized transform coefficients. Entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also encode information required for video restoration (e.g., syntax element values) in addition to the quantized transform coefficients using entropy encoding or a preset method, either together with or separately from the quantized transform coefficients. The encoded information may be transmitted or stored in network abstraction layer (NAL) units in the form of a bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. The network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0041] The inverse quantization unit 125 inversely quantizes the values ​​(quantized transformation coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inversely transforms the values ​​inversely quantized by the inverse quantization unit 125 to generate residual samples.

[0042] The adder 140 reconstructs a picture by adding residual samples and predicted samples. The residual samples and predicted samples may be added in block units to generate reconstructed blocks. Although the adder 140 has been described as a separate component, it may be part of the prediction unit 110. Meanwhile, the adder 140 may also be referred to as a reconstruction module or a reconstructed block generator.

[0043] The filter unit 150 may apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through the deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process may be corrected. The sample adaptive offset may be applied on a sample-by-sample basis and may be applied after the deblocking filtering process is completed. The filter unit 150 may also apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filtering and / or the sample adaptive offset have been applied.

[0044] The memory 160 may store a reconstructed picture (a decoded picture) or information necessary for encoding / decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 150. The stored reconstructed picture may be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 may store (reference) pictures used for inter prediction. In this case, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list.

[0045] 2 is a diagram illustrating a configuration of a video / image decoding apparatus to which the present invention can be applied. Hereinafter, the term "video decoding apparatus" may include an image decoding apparatus.

[0046] 2, the video decoding apparatus 200 may include an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filtering module 250, and a memory 260. Here, the residual processing module 220 may include a rearrangement module 221, a dequantization module 222, and an inverse transform module 223. Although not shown, the video decoding apparatus 200 may also include a receiver that receives a bitstream including video information. The receiver may be configured as a separate module or may be included in the entropy decoding module 210.

[0047] When a bitstream including video / picture information is input, the video decoding apparatus 200 can restore the video / picture / image corresponding to the process in which the video / picture information was processed in the video encoding apparatus.

[0048] For example, the video decoding apparatus 200 may perform video decoding using a processing unit applied in a video encoding apparatus. Accordingly, a processing unit block for video decoding may be a coding unit, for example, or a coding unit, a prediction unit, or a transform unit, for example. The coding unit may be divided from the largest coding unit into a quad tree structure, a binary tree structure, and / or a ternary tree structure.

[0049] A prediction unit and a transform unit may also be used in some cases. In this case, a prediction block is a block derived or partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into sub-blocks. A transform unit may be divided from a coding unit using a quadtree structure and is a unit that derives transform coefficients or a unit that derives a residual signal from the transform coefficients.

[0050] The entropy decoding unit 210 may parse the bitstream and output information necessary for video or picture reconstruction. For example, the entropy decoding unit 210 may decode information in the bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements necessary for video reconstruction and quantized values ​​of transform coefficients for residuals.

[0051] More specifically, the CABAC entropy decoding method receives BINs corresponding to each syntax element from a bitstream, determines a context model using information on the syntax element to be decoded and decoding information on adjacent and target blocks or information on symbols / BINs decoded in a previous step, predicts the occurrence probability of BINs according to the determined context model, and performs arithmetic decoding of the BINs to generate symbols corresponding to the values ​​of each syntax element. In this case, the CABAC entropy decoding method can update the context model using information on the decoded symbols / BINs for the context model of the next symbol / BIN after determining the context model.

[0052] Among the information decoded by the entropy decoding unit 210, information regarding prediction is provided to the prediction unit 230, and the residual values, i.e., the quantized transform coefficients, on which entropy decoding is performed by the entropy decoding unit 210 can be input to the reordering unit 221.

[0053] The rearrangement unit 221 may rearrange the quantized transform coefficients in a two-dimensional block format. The rearrangement unit 221 may perform rearrangement in response to coefficient scanning performed in the encoding apparatus. Here, although the rearrangement unit 221 has been described as a separate component, it may also be a part of the inverse quantization unit 222.

[0054] The inverse quantization unit 222 may inversely quantize the quantized transform coefficients based on the (inverse) quantization parameter and output the transform coefficients. In this case, information for deriving the quantization parameter may be signaled from the encoding apparatus.

[0055] The inverse transform unit 223 can inversely transform the transform coefficients to derive residual samples.

[0056] The prediction unit 230 may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit 230 performs prediction on a coding block, a transform block, or a prediction block.

[0057] The prediction unit 230 may determine whether to apply intra prediction or inter prediction based on the information regarding the prediction. In this case, the unit for determining whether to apply intra prediction or inter prediction differs from the unit for generating prediction samples. In addition, the unit for generating prediction samples differs between inter prediction and intra prediction. For example, whether to apply inter prediction or intra prediction may be determined on a CU basis. Furthermore, for example, in inter prediction, a prediction mode may be determined on a PU basis to generate prediction samples, and in intra prediction, a prediction mode may be determined on a PU basis to generate prediction samples on a TU basis.

[0058] In the case of intra prediction, the prediction unit 230 may derive a prediction sample for the current block based on neighboring reference samples in the current picture. The prediction unit 230 may derive a prediction sample for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined using the intra prediction mode of the neighboring block.

[0059] In the case of inter prediction, the predictor 230 may derive a prediction sample for the current block based on a sample identified on the reference picture by a motion vector on the reference picture. The predictor 230 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP mode. In this case, motion information required for inter prediction of the current block provided by the video encoding apparatus, such as information on a motion vector and a reference picture index, may be obtained or induced based on the information on the prediction.

[0060] In the skip mode and merge mode, motion information of neighboring blocks can be used as motion information of the current block, where the neighboring blocks can include spatial neighboring blocks and temporal neighboring blocks.

[0061] The predictor 230 constructs a merge candidate list using motion information of available neighboring blocks and can use information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. The motion information can include a motion vector and a reference picture. When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture.

[0062] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.

[0063] In the MVP mode, the motion vector of the current block can be derived using the motion vector of a neighboring block as a motion vector predictor, where the neighboring block can include a spatial neighboring block and a temporal neighboring block.

[0064] For example, when a merge mode is applied, a merge candidate list may be generated using the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vector corresponding to the Col block, which is a temporally adjacent block. In the merge mode, the motion vector of a candidate block selected from the merge candidate list is used as the motion vector of the current block. The prediction information may include a merge index indicating a candidate block having an optimal motion vector selected from the candidate blocks included in the merge candidate list. In this case, the prediction unit 230 may derive the motion vector of the current block using the merge index.

[0065] As another example, when a Motion Vector Prediction (MVP) mode is applied, a motion vector predictor candidate list may be generated using the motion vector of a reconstructed spatially neighboring block and / or the motion vector corresponding to a Col block, which is a temporally neighboring block. That is, the motion vector of a reconstructed spatially neighboring block and / or the motion vector corresponding to a Col block, which is a temporally neighboring block, may be used as a motion vector candidate. The prediction information may include a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list. In this case, the predictor 230 may select a predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. A predictor of the encoding apparatus may obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode it, and output it in the form of a bitstream. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction unit 230 may obtain a motion vector differential included in the information for the prediction and derive the motion vector of the current block by adding the motion vector differential and the motion vector predictor. In addition, the prediction unit may obtain or induce a reference picture index indicating a reference picture from the information for the prediction.

[0066] The adder 240 may reconstruct a current block or a current picture by adding residual samples and predicted samples. The adder 240 may also reconstruct a current picture by adding residual samples and predicted samples in block units. When a skip mode is applied, residuals are not transmitted, and therefore predicted samples may become reconstructed samples. Although the adder 240 has been described as a separate component, it may also be part of the prediction unit 230. Meanwhile, the adder 240 may also be referred to as a reconstruction module or a reconstructed block generator.

[0067] The filter unit 250 may apply deblocking filtering, sample adaptive offset, and / or ALF to the reconstructed picture. In this case, the sample adaptive offset may be applied in sample units or may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or sample adaptive offset.

[0068] The memory 260 may store a reconstructed picture (a decoded picture) or information necessary for decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 250. For example, the memory 260 may store a picture used for inter prediction. In this case, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for another picture. The memory 260 may also output the reconstructed picture in an output order.

[0069] Meanwhile, as described above, prediction is performed to improve compression efficiency during video coding. Accordingly, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device, and the encoding device signals information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values ​​of the original block, to the decoding device, thereby improving video coding efficiency. The decoding device derives a residual block including residual samples based on the residual information, and adds the residual block and the predicted block to generate a reconstructed block including reconstructed samples, thereby generating a reconstructed picture including the reconstructed block.

[0070] The residual information may be generated through a transform and quantization procedure. For example, an encoding apparatus may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information (via a bitstream) to a decoding apparatus. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding apparatus may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding apparatus may generate a reconstructed picture based on the predicted block and the residual block. The encoding apparatus may also derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.

[0071] FIG. 3 illustrates a schematic diagram of the multiple conversion technique according to the present invention.

[0072] Referring to Figure 3, the transform unit may correspond to the transform unit in the encoding device of Figure 1 described above, and the inverse transform unit may correspond to the inverse transform unit in the encoding device of Figure 1 described above or the inverse transform unit in the decoding device of Figure 2.

[0073] The transform unit may derive (first-order) transform coefficients by performing a first-order transform based on the residual samples (residual sample array) in the residual block (S310). Here, the first-order transform may include a multiple transform set (MTS). The multiple transform set may also be referred to as an adaptive multiple core transform.

[0074] The adaptive multi-kernel transform may refer to a transform method that additionally uses a DCT (Discrete Cosine Transform) Type 2 and a DST (Discrete Sine Transform) Type 7, DCT Type 8, and / or DST Type 1. That is, the multi-kernel transform may refer to a transform method that transforms a spatial domain residual signal (or a residual block) into frequency domain transform coefficients (or primary transform coefficients) based on a plurality of transform kernels selected from the DCT Type 2, the DST Type 7, the DCT Type 8, and the DST Type 1. Here, the primary transform coefficients are called temporary transform coefficients in the transform unit.

[0075] That is, when an existing transform method is applied, a transform coefficient is generated by applying a spatial domain to a frequency domain to a residual signal (or a residual block) based on DCT type 2. In contrast, when the adaptive multi-kernel transform is applied, a transform coefficient (or a primary transform coefficient) is generated by applying a spatial domain to a frequency domain to a residual signal (or a residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. Here, DCT type 2, DST type 7, DCT type 8, DST type 1, etc. are referred to as transform types, transform kernels, or transform cores.

[0076] For reference, the DCT / DST transform type can be defined based on the basis functions, which are shown in the table below.

[0077] [Table 1]

[0078] When the adaptive multi-kernel transform is performed, a vertical transform kernel and a horizontal transform kernel for a current block may be selected from the transform kernels, and a vertical transform for the current block may be performed based on the vertical transform kernel, and a horizontal transform for the current block may be performed based on the horizontal transform kernel. Here, the horizontal transform may indicate a transform for a horizontal component of the current block, and the vertical transform may indicate a transform for a vertical component of the current block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on a transform index indicating a prediction mode and / or a transform subset of a current block (CU or sub-block) surrounding a residual block.

[0079] The transform unit may derive (secondary) transform coefficients by performing a secondary transform based on the (primary) transform coefficients (S320). If the primary transform is a transform from the spatial domain to the frequency domain, the secondary transform may be considered a transform from the frequency domain to the frequency domain. The secondary transform may include a non-separable transform. In this case, the secondary transform is referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may refer to a transform that generates transform coefficients (or secondary transform coefficients) for a residual signal by performing a secondary transform on the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix. Here, a transform may be applied to the (primary) transform coefficients at once based on the non-separable transform matrix without separately applying a vertical transform and a horizontal transform (or independently applying a horizontal-vertical transform). That is, the non-separable quadratic transform may refer to a transform method in which vertical and horizontal components of the (primary) transform coefficients are transformed together without being separated based on the non-separable transform matrix to generate transform coefficients (or secondary transform coefficients). The non-separable quadratic transform may be applied to the top-left region of a block (hereinafter referred to as a transform coefficient block) composed of the (primary) transform coefficients. For example, if the width (W) and height (H) of the transform coefficient block are both 8 or greater, an 8x8 non-separable quadratic transform may be applied to the top-left 8x8 region of the transform coefficient block. Also, if the width (W) and height (H) of the transform coefficient block are both 4 or greater and the width (W) or height (H) of the transform coefficient block is less than 8, a 4x4 non-separable quadratic transform may be applied to the top-left min(8,W)xmin(8,H) region of the transform coefficient block.However, the embodiment is not limited to this. For example, even if the only condition that the width (W) or height (H) of the transform coefficient block is less than 8 is met, a 4x4 non-separable quadratic transform can be applied to the upper left min(8,W)xmin(8,H) area of ​​the transform coefficient block.

[0080] Specifically, for example, if a 4x4 input block is used, a non-separable quadratic transform can be performed as follows:

[0081] The 4x4 input block X is given as follows:

[0082]

number

[0083] When X is expressed in vector form, the vector JPEG2025164943000004.jpg84 is shown as follows:

[0084]

number

[0085] In this case, the second-order non-separable transform can be calculated as follows:

[0086]

number

[0087] where: JPEG2025164943000007.jpg75 denotes the transform coefficient vector, and T denotes the 16x16 (non-separable) transform matrix.

[0088] 16×1 transform coefficient vector through Equation 3 JPEG2025164943000008.jpg75 can be derived, JPEG2025164943000009.jpg75 can be re-organized into 4x4 blocks via scan order (horizontal, vertical, diagonal, etc.). However, the above calculations are merely examples, and in order to reduce the computational complexity of non-separable quadratic transforms, HyGT (Hypercube-Givens Transform) or the like can also be used to calculate the non-separable quadratic transform.

[0089] Meanwhile, the non-separable quadratic transform may be a mode-dependent transform kernel (or transform core, transform type), where the mode may include an intra-prediction mode and / or an inter-prediction mode.

[0090] As described above, the non-separable quadratic transform may be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. That is, the non-separable quadratic transform may be performed based on an 8×8 sub-block size or a 4×4 sub-block size. For example, for the mode-based transform kernel selection, 35 sets of three non-separable quadratic transform kernels for the non-separable quadratic transform may be configured for both the 8×8 sub-block size and the 4×4 sub-block size. That is, 35 transform sets may be configured for the 8×8 sub-block size, and 35 transform sets may be configured for the 4×4 sub-block size. In this case, each of the 35 transform sets for the 8×8 sub-block size may include three 8×8 transform kernels, and each of the 35 transform sets for the 4×4 sub-block size may include three 4×4 transform kernels. However, the transform sub-block size, the number of sets, and the number of transform kernels in a set are merely examples, and sizes other than 8x8 or 4x4 can be used, or n sets can be constructed and each set can contain k transform kernels.

[0091] The transform set is referred to as an NSST set, and the transform kernels in the NSST set are referred to as NSST kernels. Selection of a particular set from the transform set may be performed based on, for example, the intra prediction mode of the current block (CU or sub-block).

[0092] For reference, for example, the intra prediction modes may include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes may include a planar intra prediction mode numbered 0 and a DC intra prediction mode numbered 1, and the directional intra prediction modes may include 65 intra prediction modes numbered 2 to 66. However, this is merely an example, and the present invention may be applied to cases where the number of intra prediction modes is different. Meanwhile, in some cases, a 67th intra prediction mode may also be used, and the 67th intra prediction mode may indicate a linear model (LM) mode.

[0093] FIG. 4 exemplarily shows the intra-directional modes of 65 prediction directions.

[0094] Referring to FIG. 4, intra prediction modes with horizontal directionality and intra prediction modes with vertical directionality can be distinguished from each other, with the 34th intra prediction mode having a left-up diagonal prediction direction. H and V in FIG. 3 represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate displacements in 1 / 32 units on the sample grid position. The 2nd to 33rd intra prediction modes have horizontal directionality, while the 34th to 66th intra prediction modes have vertical directionality. The 18th and 50th intra prediction modes are horizontal and vertical intra prediction modes, respectively. The 2nd intra prediction mode is referred to as a left-down diagonal intra prediction mode, the 34th intra prediction mode is referred to as a left-up diagonal intra prediction mode, and the 66th intra prediction mode is referred to as a right-up diagonal intra prediction mode.

[0095] In this case, the mapping between the 35 transform sets and the intra prediction modes is shown in the following table: For reference, when the LM mode is applied to the current block, no secondary transform is applied to the current block.

[0096] [Table 2]

[0097] On the other hand, if it is determined that a specific set is to be used, one of the k transform kernels in the specific set may be selected through a non-separable secondary transform index. The encoding device may derive a non-separable secondary transform index that points to a specific transform kernel based on a rate-distortion (RD) check and signal the non-separable secondary transform index to a decoding device. The decoding device may select one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, an NSST index value of 0 may point to the first non-separable secondary transform kernel, an NSST index value of 1 may point to the second non-separable secondary transform kernel, and an NSST index value of 2 may point to the third non-separable secondary transform kernel. Alternatively, an NSST index value of 0 may indicate that the first non-separable secondary transform is not applied to the current block, and NSST index values ​​of 1 to 3 may point to the three transform kernels.

[0098] 3, the transform unit may perform the non-separable quadratic transform based on the selected transform kernel to obtain (quadratic) transform coefficients. The transform coefficients may be derived as quantized transform coefficients through the quantizer, encoded, and signaled in the decoding device and transmitted to an inverse quantization / inverse transform unit in the encoding device.

[0099] On the other hand, when the second-order transform is omitted as described above, the (first-order) transform coefficients, which are the output of the first-order (separate) transform, can be derived as quantized transform coefficients through the quantization unit as described above, encoded, and signaled in the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0100] The inverse transform unit may perform a series of steps in the reverse order of the steps performed by the transform unit described above. The inverse transform unit may receive (dequantized) transform coefficients, perform a secondary (inverse) transform to derive (primary) transform coefficients (S350), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain residual blocks (residual samples). Here, the primary transform coefficients are called modified transform coefficients from the inverse transform unit's perspective. As described above, the encoding and decoding devices may generate reconstructed blocks based on the residual blocks and predicted blocks, and generate reconstructed pictures based on the reconstructed blocks.

[0101] On the other hand, as described above, if the secondary (inverse) transform is omitted, a residual block (residual sample) can be obtained by receiving (dequantized) transform coefficients and performing the primary (separate) transform. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and a predicted block, and generate a reconstructed picture based on the reconstructed block.

[0102] 5a to 5c are diagrams illustrating a simplified transformation according to an embodiment of the present invention.

[0103] As described above in FIG. 3, in the non-separable quadratic transform (NSST), block data of transform coefficients obtained by applying a linear transform is divided into M×M blocks, and then M 2 ×M 2 NSST can be performed, where M is, for example, but not limited to, 4 or 8.

[0104] M 2 ×M 2NSST can be applied in the form of matrix multiplication, but to reduce the amount of calculation and memory required, the Hypercube-Givens Transform (HyGT) described in Figure 3 can be used to calculate NSST. HyGT is an orthogonal transform, and is a Givens rotation G defined by an orthogonal matrix G(m,n,θ). i、j (m,n) can be included as a basic component. Givens rotation G i、j (m, n) is as shown in the following formula 4.

[0105]

number

[0106] The Givens rotation based on Equation 4 can be shown in Figure 5a. Referring to Equation 4 and Figure 5a, it can be seen that one Givens rotation can be described by only one angle (θ).

[0107] Figure 5b shows an example of one round of a 16x16 NSST. More specifically, HyGT can be performed by combining Givens rotations in a hypercube arrangement. The HyGT flow for 16 elements can be represented in butterfly form as shown in Figure 5b. As shown in Figure 5b, one round consists of four Givens rotation layers, and each Givens rotation layer consists of eight Givens rotations. Each Givens rotation can be structured to select two input data, apply a rotation transformation to them, and then output them directly to the selected position, as shown in the concatenated configuration shown in Figure 5b. The 16x16 NSST can sequentially apply round 2 and permutation layer 1, allowing for arbitrary mixing of 16 data through the corresponding permutation layer. All two rounds can be concatenated as shown in Figure 5b, but the Givens rotation layers for the two rounds can be different.

[0108] The 64x64 NSST consists of a Givens rotation layer with 64 inputs and outputs. Similar to the 16x16 NSST, at least one round can be applied, and each round can consist of six Givens rotation layers connected in a manner similar to Figure 5b. In one example, the 64x64 NSST can apply four rounds, followed by a permutation layer to arbitrarily mix the 64 data. The Givens rotation layers for each of the four rounds can be different from each other.

[0109] Figure 5b shows the rounds applied to the inverse transform. When applying the inverse transform, the inverse permutation layer is first applied, and then the corresponding Givens rotation is applied in the order from the last round to the first round, from bottom to top in Figure 5b. The angle corresponding to each Givens rotation of the inverse NSST can be a value obtained by applying a negative sign to the corresponding inverse angle.

[0110] To increase coding efficiency, one or more HyGT rounds can be used. As shown in FIG. 5c, NSST can be composed of R HyGT rounds and can additionally include a sorting pass. The sorting pass can also be interpreted as an optional permutation pass, which can sort transform coefficients based on variance. As an example, 2-round HyGT can be applied to 16×16 NSST, and 4-round HyGT can be applied to 64×64 NSST.

[0111] FIG. 6 is a diagram illustrating a simplified conversion according to an embodiment of the present invention.

[0112] In this specification, the term "current block" refers to the current block or residual block on which coding is performed.

[0113] In this specification, the term "simplified transform" refers to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced by a simplification factor. When a simplified transform is performed, the amount of calculation required during the transform can be reduced by reducing the size of the transform matrix. That is, a simplified transform can be used to solve the problem of computational complexity that occurs when transforming a large block or performing a non-separable transform. A simplified transform can be used for any type of transform, such as a linear transform (also called a core transform. Linear transforms include, for example, DCT, DST, etc.) or a quadratic transform (for example, NSST).

[0114] The simplified transform may be referred to by various terms such as reduced transform, reduced transform, reduced secondary transform, reduction transform, simplified transform, simple transform, RTS, RST, etc., and the names referred to as simplified transforms are not limited to the examples listed.

[0115] In a simplified transformation according to one embodiment, an N-dimensional vector may be mapped to an R-dimensional vector located in another space to determine a simplified transformation matrix, where R is smaller than N. N refers to the square of the length of one side of a block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor refers to the value R / N. The simplification factor may also be referred to by various terms such as a reduced factor, reduction factor, simplified factor, or simple factor. Meanwhile, R is referred to as a reduced coefficient, but in some cases the simplification factor may refer to R. In other cases, the simplification factor may refer to the value N / R.

[0116] In one embodiment, the simplification factors or simplification coefficients may be signaled via a bitstream, but the embodiment is not limited thereto. For example, predefined values ​​for the simplification factors or simplification coefficients may be stored in each of the encoding apparatus 100 and the decoding apparatus 200, in which case the simplification factors or simplification coefficients are not separately signaled.

[0117] The size of the simplified transformation matrix according to an embodiment is R×N, which is smaller than the size N×N of the normal transformation matrix, and can be defined as Equation 5 below.

[0118]

number

[0119] The matrix T in the Reduced Transform block shown in (a) of FIG. 6 is the matrix T in Equation 5. RxNAs shown in Fig. 6(a), the simplified transformation matrix T RxN When multiplied by , the transform coefficients for the current block can be derived.

[0120] In one embodiment, if the size of the block to which the transformation is applied is 8x8, R=16 (i.e., R / N=16 / 64=1 / 4), and the size of the target block is 64x64, the simplified transformation according to (a) of Figure 6 can be expressed by a matrix operation such as the following Equation 6.

[0121]

number

[0122] In Equation 6, r1 to r 64 can represent the residual sample for the target block. As a result of the calculation of Equation 6, the transform coefficient c for the target block i can be derived, and c i The derivation process is as shown in Equation 7.

[0123]

number

[0124] As a result of the calculation of Equation 7, the transform coefficients c1 to c2 for the target block are R That is, when R=16, the transform coefficients c1 to c2 for the current block can be derived. 16can be derived. If a regular transform, rather than a simplified transform, is applied and a transform matrix of size 64×64 (N×N) is multiplied by a matrix containing residual samples of size 64×1 (N×1), 64 (N) transform coefficients for the current block are derived. However, because the simplified transform is applied, only 16 (R) transform coefficients for the current block are derived. Since the total number of transform coefficients for the current block is reduced from N to R, the amount of data transmitted from the encoding apparatus 100 to the decoding apparatus 200 is reduced, thereby improving the transmission efficiency between the encoding apparatus 100 and the decoding apparatus 200.

[0125] Considering the size of the transformation matrix, the size of a normal transformation matrix is ​​64x64 (NxN), while the size of a simplified transformation matrix is ​​reduced to 16x64 (RxN), so compared to when performing a normal transformation, memory usage can be reduced by a ratio of R / N when performing a simplified transformation. Also, compared to the number of multiplication operations NxN when using a normal transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N when using a simplified transformation matrix (RxN).

[0126] In one embodiment, the transform unit 122 of the encoding apparatus 100 may transform residual samples for a current block to derive transform coefficients for the current block, and the transform coefficients for the current block may be transmitted to an inverse transform unit of the decoding apparatus 200, which may then inversely transform the transform coefficients for the current block. Residual samples for the current block may be derived based on the inverse transform performed on the transform coefficients for the current block. That is, the detailed operations based on the (simplified) inverse transform are opposite in order to the detailed operations based on the (simplified) transform, whereas the detailed operations based on the (simplified) inverse transform and the detailed operations based on the (simplified) transform are substantially similar.

[0127] A simplified inverse transformation matrix T according to one embodimentNxR The size of the simplified transformation matrix T RxN It is in a transpose relationship with

[0128] The matrix T in the Reduced Inverse Transform block shown in Figure 6(b) t is the simplified inverse transformation matrix T NxR As shown in (b) of FIG. 6, the simplified inverse transform matrix T NxR When multiplied by , a primary transform coefficient for the current block or a residual sample for the current block can be derived.

[0129] More specifically, when a simplified inverse transform is applied based on a quadratic inverse transform, a simplified inverse transform matrix T NxR On the other hand, when a simplified inverse transform is applied based on the primary inverse transform, the simplified inverse transform matrix T NxR When multiplied by , the residual sample for the current block can be derived.

[0130] In one embodiment, if the size of the block to which the inverse transform is applied is 8x8, R=16 (i.e., R / N=16 / 64=1 / 4), and the size of the target block is 64x64, the simplified inverse transform according to (b) of Figure 6 can be expressed by a matrix operation such as the following Equation 8.

[0131]

number

[0132] In Equation 8, c1 to c 16can represent the transform coefficients for the target block. As a result of the calculation of Equation 8, r represents the primary transform coefficients for the target block or the residual samples for the target block. j can be derived, and r j The derivation process is as shown in Equation 9.

[0133]

number

[0134] The calculation result of Equation 9 is r1 to r2, which indicate the primary transform coefficients for the target block or the residual samples for the target block. N can be derived. Considering the size of the inverse transformation matrix, the size of a normal inverse transformation matrix is ​​64 x 64 (N x N), while the size of a simplified inverse transformation matrix is ​​reduced to 64 x 16 (N x R). Therefore, compared to performing a normal inverse transformation, memory usage can be reduced by a ratio of R / N when performing a simplified inverse transformation. Also, compared to the number of multiplication operations (N x N) when using a normal inverse transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N (N x R) when using a simplified inverse transformation matrix.

[0135] FIG. 7 is a flow chart illustrating a simplified conversion process according to one embodiment of the present invention.

[0136] Each step disclosed in Fig. 7 may be performed by the decoding apparatus 200 disclosed in Fig. 2. More specifically, S700 may be performed by the inverse quantization unit 222 disclosed in Fig. 2, and S710 and S720 may be performed by the inverse transform unit 223 disclosed in Fig. 2. Therefore, the description of the details that overlap with the details described above in Fig. 2 will be omitted or simplified.

[0137] 6, the detailed operations of the (simplified) transform are opposite in order to the detailed operations of the (simplified) inverse transform, whereas the detailed operations of the (simplified) transform and the detailed operations of the (simplified) inverse transform are substantially similar. Therefore, those skilled in the art will readily understand that the description of steps S700 to S720 for the simplified inverse transform described below can be applied in the same or similar manner to the simplified transform.

[0138] The decoding apparatus 200 according to an embodiment may derive transform coefficients by performing inverse quantization on quantized transform coefficients for a current block (S700).

[0139] According to an embodiment, the decoding apparatus 200 may select a transform kernel (S710). More specifically, the decoding apparatus 200 may select the transform kernel based on at least one of a transform index, a width and height of a region to which the transform is applied, an intra prediction mode used in video decoding, and information on a color component of a target block. However, the embodiment is not limited thereto. For example, the transform kernel may be predefined, and separate information for selecting the transform kernel may not be signaled.

[0140] In one example, information on the hue component of a target block may be signaled via CIdx. If the target block is a luma block, CIdx may indicate 0. If the target block is a chroma block, e.g., a Cb block or a Cr block, CIdx may indicate a non-zero value (e.g., 1).

[0141] The decoding apparatus 200 according to an embodiment may apply a simplified inverse transform to the transform coefficients based on the selected transform kernel and the reduced factor (S720).

[0142] FIG. 8 is a flow chart showing a simplified conversion process according to another embodiment of the present invention.

[0143] Each step disclosed in Fig. 8 may be performed by the decoding apparatus 200 disclosed in Fig. 2. More specifically, S800 may be performed by the inverse quantization unit 222 disclosed in Fig. 2, and S810 to S860 may be performed by the inverse transform unit 223 disclosed in Fig. 2. Therefore, the description of the details overlapping with the details described above in Fig. 2 will be omitted or simplified.

[0144] 6, the detailed operations of the (simplified) transform are opposite in order to the detailed operations of the (simplified) inverse transform, whereas the detailed operations of the (simplified) transform and the detailed operations of the (simplified) inverse transform are substantially similar. Therefore, those skilled in the art will readily understand that the description of S800 to S860 for the simplified inverse transform described below can be applied in the same or similar manner to the simplified transform.

[0145] According to an embodiment, the decoding apparatus 200 may perform inverse quantization on quantized coefficients for a current block (S800). If a transform has been performed in the encoding apparatus 100, the decoding apparatus 200 may perform inverse quantization on quantized transform coefficients for the current block to derive transform coefficients for the current block in S800. On the other hand, if a transform has not been performed in the encoding apparatus 100, the decoding apparatus 200 may perform inverse quantization on quantized residual samples for the current block to derive residual samples for the current block in S800.

[0146] The decoding apparatus 200 according to an embodiment may determine whether a transform has been performed on a residual sample for a current block in the encoding apparatus 100 (S810), and if it is determined that a transform has been performed, may parse a transform index (or decode from the bitstream) (S820). The transform index may include a horizontal transform index for a horizontal transform and a vertical transform index for a vertical transform.

[0147] In one example, the transform index may include a primary transform index, a core transform index, an NSST index, etc. The transform index may be represented by, for example, Transform_idx, and the NSST index may be represented by, for example, NSST_idx. Also, the horizontal transform index may be represented by Transform_idx_h, and the vertical transform index may be represented by Transform_idx_v.

[0148] If it is determined in step S810 that the encoding apparatus 100 has not performed transformation on the residual samples for the current block, the decoding apparatus 200 according to an embodiment may omit operations S820 to S860.

[0149] A decoding device 200 according to one embodiment may select a transform kernel based on at least one of a transform index, the width and height of the region to which the transform is applied, an intra prediction mode used in video decoding, and information on the color component of the target block (S830).

[0150] The decoding apparatus 200 according to an embodiment may determine whether a condition for performing a simplified inverse transform on transform coefficients of a current block is met (S840).

[0151] In one example, if the width and height of the region to which the simplified inverse transform is applied are greater than the first coefficient, the decoding apparatus 200 may determine that a condition exists for performing the simplified inverse transform on the transform coefficients of the current block.

[0152] In another example, if the product of the width and height of the area to which the simplified inverse transform is applied is greater than the second coefficient and the smaller of the width and height of the area to which the simplified inverse transform is applied is greater than the third coefficient, the decoding device 200 may determine that the conditions for performing a simplified inverse transform on the transform coefficients of the target block are met.

[0153] In another example, if the width and height of the region to which the simplified inverse transform is applied are smaller than or equal to the fourth coefficient, respectively, the decoding apparatus 200 may determine that the conditions for performing the simplified inverse transform on the transform coefficients of the target block are met.

[0154] In another example, if the product of the width and height of the area to which the simplified inverse transform is applied is smaller than or equal to the fifth coefficient and the smaller of the width and height of the area to which the simplified inverse transform is applied is smaller than or equal to the sixth coefficient, the decoding device 200 may determine that the conditions for performing a simplified inverse transform on the transform coefficients of the target block are met.

[0155] In another example, if at least one of the following conditions is satisfied: the width and height of the region to which the simplified inverse transform is applied are each greater than a first coefficient; the product of the width and height of the region to which the simplified inverse transform is applied is greater than a second coefficient and the smaller of the width and height of the region to which the simplified inverse transform is applied is greater than a third coefficient; the width and height of the region to which the simplified inverse transform is applied are each smaller than or equal to a fourth coefficient; and the product of the width and height of the region to which the simplified inverse transform is applied is smaller than or equal to a fifth coefficient and the smaller of the width and height of the region to which the simplified inverse transform is applied is smaller than or equal to a sixth coefficient, the decoding device 200 may determine that the conditions for performing a simplified inverse transform on the transform coefficients of the current block are met.

[0156] In the above example, the first to sixth coefficients are any predefined positive integers, for example, 4, 8, 16, or 32.

[0157] According to one embodiment, the simplified inverse transform may be applied to a square region included in a target block (i.e., when the width and height of the region to which the simplified inverse transform is applied are the same), and in some cases, the width and height of the region to which the simplified inverse transform is applied may be fixed to predefined coefficient values ​​(e.g., 4, 8, 16, 32, etc.). Meanwhile, the region to which the simplified inverse transform is applied is not limited to a square region, and the simplified inverse transform may also be applied to a rectangular region or a non-rectangular region. A more detailed description of the region to which the simplified inverse transform is applied will be provided below with reference to FIG. 10.

[0158] In one example, whether a condition for performing a simplified inverse transform is met may be determined based on a transform index, i.e., the transform index may indicate what transform has been performed on the current block.

[0159] If it is determined in step S840 that the conditions for performing a simplified inverse transform are not met, the decoding apparatus 200 according to an embodiment may perform a (regular) inverse transform on the transform coefficients of the current block. As described above with reference to FIG. 3, the (inverse) transform may include, but is not limited to, DCT2, DCT4, DCT5, DCT7, DCT8, DST1, DST4, DST7, NSST, JEM-NSST (HyGT), etc.

[0160] If it is determined in step S840 that a condition for performing a simplified inverse transform is met, the decoding apparatus 200 according to an embodiment may perform a simplified inverse transform on transform coefficients of the current block (step S860).

[0161] FIG. 9 is a flow chart illustrating a simplified transformation process based on a non-separable quadratic transformation according to an embodiment of the present invention.

[0162] Each step disclosed in Figure 9 may be performed by the decoding apparatus 200 disclosed in Figure 2. More specifically, S900 may be performed by the inverse quantization unit 222 disclosed in Figure 2, and S910 to S980 may be performed by the inverse transform unit 223 disclosed in Figure 2. In addition, S900 in Figure 9 may correspond to S800 in Figure 8, S940 in Figure 9 may correspond to S830 in Figure 8, and S950 in Figure 9 may correspond to S840 in Figure 8. Therefore, detailed content that overlaps with the content described above in Figures 2 and 8 will be omitted or simplified.

[0163] 6, the detailed operations of the (simplified) transform are the exact opposite of the detailed operations of the (simplified) inverse transform, whereas the detailed operations of the (simplified) transform and the detailed operations of the (simplified) inverse transform are substantially similar. Therefore, those skilled in the art will readily understand that the description of steps S900 to S980 for the simplified inverse transform described below can be applied in the same or similar manner to the simplified transform.

[0164] The decoding apparatus 200 according to an embodiment may perform inverse quantization on the quantized coefficients of the current block (S900).

[0165] The decoding device 200 according to one embodiment can determine whether NSST has been performed on the residual samples for the current block in the encoding device 100 (S910), and if it is determined that NSST has been performed, can parse the NSST index (or decode it from the bitstream) (S920).

[0166] A decoding device 200 according to one embodiment can determine whether the NSST index is greater than 0 (S930), and if it is determined that the NSST index is greater than 0, can select a transform kernel based on at least one of the NSST index, the width and height of the area to which the NSST is applied, the intra prediction mode, and information on the hue component of the target block (S940).

[0167] The decoding apparatus 200 according to an embodiment may determine whether a condition for performing a simplified inverse transform on transform coefficients of a current block is met (S950).

[0168] In one embodiment, if it is determined in S950 that the conditions for performing a simplified inverse transform are not met, the decoding device 200 may perform a (normal) inverse inverse transform that is not based on the simplified inverse transform on the transform coefficients for the target block.

[0169] If it is determined in step S950 that the conditions for performing a simplified inverse transform are met, the decoding apparatus 200 according to one embodiment may perform an inverse NSST based on the simplified inverse transform on the transform coefficients of the current block.

[0170] If it is determined in step S910 that the encoding apparatus 100 has not performed NSST on the residual samples for the current block, the decoding apparatus 200 according to an embodiment may omit operations S920 to S970.

[0171] If it is determined in S930 that the NSST index is not greater than 0, the decoding apparatus 200 according to an embodiment may omit operations S940 to S970.

[0172] The decoding apparatus 200 according to an embodiment may perform a first inverse transform on first transform coefficients of a current block derived by applying the inverse NSST. After the first inverse transform is performed on the first transform coefficients, residual samples of the current block may be derived.

[0173] FIG. 10 illustrates a block to which a simplifying transformation is applied according to one embodiment of the present invention.

[0174] As described above with reference to FIG. 8, the area to which the simplified (inverse) transform is applied within the target block is not limited to a square area, but the simplified transform can also be applied to a rectangular area or a non-rectangular area.

[0175] FIG. 10 shows an example in which a simplified transformation is applied to a non-rectangular region within a target block 1000 having a size of 16×16. Ten shaded blocks 1010 in FIG. 10 indicate the region within the target block 1000 to which the simplified transformation is applied. Since each minimum block has a size of 4×4, the simplified transformation is applied to ten 4×4 pixels (i.e., the simplified transformation is applied to 160 pixels) according to the example of FIG. 10. When R=16, the size of the simplified transformation matrix can be 16×160.

[0176] Meanwhile, those skilled in the art can easily understand that the arrangement of atomic blocks 1010 included in the area to which the simplified transformation is applied shown in Fig. 10 is merely one of many examples. For example, the atomic blocks included in the area to which the simplified transformation is applied may not be adjacent to each other, or may be in a relationship in which they share only one vertex.

[0177] FIG. 11 is a flowchart showing the operation of a video encoding apparatus according to an embodiment of the present invention.

[0178] The steps disclosed in Fig. 11 may be performed by the encoding apparatus 100 disclosed in Fig. 1. More specifically, S1100 may be performed by the subtraction unit 121 disclosed in Fig. 1, S1110 may be performed by the transformation unit 122 disclosed in Fig. 1, S1120 may be performed by the quantization unit 123 disclosed in Fig. 1, and S1130 may be performed by the entropy encoding unit 130 disclosed in Fig. 1. In addition, the operations of S1100 to S1130 are performed based on some of the content described above with reference to Figs. 6 to 10. Therefore, detailed content that overlaps with the content described above with reference to Figs. 1 and 6 to 10 will be omitted or simplified.

[0179] The encoding apparatus 100 according to an embodiment may derive residual samples for a current block (S1100).

[0180] The encoding apparatus 100 according to an embodiment may derive transform coefficients for the current block based on a simplified transform of the residual samples (S1110). In one example, the simplified transform may be performed based on a simplified transform matrix, which is a non-square matrix having fewer rows than columns.

[0181] In one embodiment, step S1110 may include determining whether a condition for applying a simplified transform is met, generating and encoding a transform index based on the determination, selecting a transform kernel, and if the condition for applying a simplified transform is met, applying the simplified transform to the residual samples based on the selected transform kernel and a simplification factor, where the size of the simplified transform matrix may be determined based on the simplification factor.

[0182] If the simplified transform performed in S1110 is based on a first-order transform, first-order transform coefficients for the current block may be derived by performing the simplified transform on residual samples for the current block. The decoding apparatus 200 may perform NSST on the first-order transform coefficients for the current block, and in this case, the NSST may be performed based on the simplified transform or may not be performed based on the simplified transform. When the NSST is performed based on the simplified transform, it may correspond to the operation in S1110.

[0183] The encoding apparatus 100 according to an embodiment may perform quantization based on the transform coefficients of the current block and derive quantized transform coefficients (S1120).

[0184] The encoding apparatus 100 according to an embodiment may encode information about quantized transform coefficients (S1130). More specifically, the encoding apparatus 100 may generate information about quantized transform coefficients and encode the generated information about quantized transform coefficients. The information about quantized transform coefficients may include residual information.

[0185] In one example, the information on the quantized transform coefficients may include at least one of information on whether a simplified transform is applied, information on a simplification factor, information on a minimum transform size to which the simplified transform is applied, and information on a maximum transform size to which the simplified transform is applied. A more detailed description of the information on the quantized transform coefficients will be provided below with reference to FIG. 12.

[0186] Referring to S1110, it can be seen that transform coefficients for a current block are derived based on a simplified transform of residual samples. Considering the size of the transform matrix, the size of a normal transform matrix is ​​N×N, whereas the size of the simplified transform matrix is ​​reduced to R×N. Therefore, compared to performing a normal transform, memory usage can be reduced by a ratio of R / N when performing the simplified transform. Furthermore, compared to the number of multiplication operations required when using a normal transform matrix (N×N), the number of multiplication operations can be reduced by a ratio of R / N (R×N) when using the simplified transform matrix. Furthermore, since only R transform coefficients are derived when the simplified transform is applied, the total number of transform coefficients for the current block is reduced from N to R, compared to the number of transform coefficients required when using a normal transform (N), thereby reducing the amount of data transmitted from the encoding apparatus 100 to the decoding apparatus 200. In summary, according to S1110, the simplified transform can improve the transform efficiency and coding efficiency of the encoding apparatus 100.

[0187] FIG. 12 is a flowchart showing the operation of a video decoding apparatus according to an embodiment of the present invention.

[0188] Each step disclosed in Fig. 12 may be performed by the decoding apparatus 200 disclosed in Fig. 2. More specifically, S1200 may be performed by the entropy decoding unit 210 disclosed in Fig. 2, S1210 may be performed by the inverse quantization unit 222 disclosed in Fig. 2, S1220 may be performed by the inverse transform unit 223 disclosed in Fig. 2, and S1230 may be performed by the addition unit 240 disclosed in Fig. 2. In addition, the operations of S1200 to S1230 are performed based on some of the content described above with reference to Figs. 6 to 10. Therefore, detailed content that overlaps with the content described above with reference to Figs. 2 and 6 to 10 will be omitted or simplified.

[0189] A decoding apparatus 200 according to an embodiment may derive quantized transform coefficients for a current block from a bitstream (S1200). More specifically, the decoding apparatus 200 may decode information on the quantized transform coefficients for the current block from the bitstream, and may derive quantized transform coefficients for the current block based on the information on the quantized transform coefficients for the current block. The information on the quantized transform coefficients for the current block may be included in a Sequence Parameter Set (SPS) or a slice header, and may include at least one of information on whether a simplified transform is applied, information on a simplification factor, information on a minimum transform size for applying the simplified transform, information on a maximum transform size for applying the simplified transform, and information on a simplified inverse transform size.

[0190] More specifically, information on whether a simplified transform is applied can be indicated via an availability flag, information on a simplification factor can be indicated via a simplification factor value, information on a minimum transform size for applying a simplified inverse transform can be indicated via a minimum transform size value, information on a maximum transform size for applying a simplified inverse transform can be indicated via a maximum transform size value, and information on a simplified inverse transform size can be indicated via a size value of the simplified inverse transform. In this case, the availability flag can be signaled via a first syntax element, the simplification factor value can be signaled via a second syntax element, the minimum transform size value can be signaled via a third syntax element, the maximum transform size value can be signaled via a fourth syntax element, and the simplified inverse transform size value can be signaled via a fifth syntax element.

[0191] In one example, the first syntax element may be represented by a syntax element Reduced_transform_enabled_flag. If a simplified transform is applied, the syntax element Reduced_transform_enabled_flag may indicate 1, and if a simplified transform is not applied, the syntax element Reduced_transform_enabled_flag may indicate 0. If the syntax element Reduced_transform_enabled_flag is not signaled, the value of the syntax element Reduced_transform_enabled_flag may be inferred to be 0.

[0192] Also, the second syntax element may be expressed as a syntax element Reduced_transform_factor. The syntax element Reduced_transform_factor may indicate a value of R / N, where N is the square of the length of one side of the block to which the transform is applied or the total number of transform coefficients corresponding to the block to which the transform is applied. R indicates a simplification factor smaller than N. However, examples are not limited thereto, and for example, Reduced_transform_factor may indicate R instead of R / N. Considered from the perspective of a simplified inverse transform matrix, R indicates the number of columns of the simplified inverse transform matrix, and N indicates the number of rows of the simplified inverse transform matrix, where the number of columns of the simplified inverse transform matrix must be less than the number of rows. R may be, for example, a value of 8, 16, 32, etc., but is not limited thereto. If the syntax element Reduced_transform_factor is not signaled, the value of Reduced_transform_factor can be estimated to be R / N (or R).

[0193] Furthermore, the third syntax element may be expressed as the syntax element min_reduced_transform_size. If the syntax element min_reduced_transform_size is not signaled, the value of min_reduced_transform_size may be inferred to 0.

[0194] Furthermore, the fourth syntax element may be expressed as the syntax element max_reduced_transform_size. If the syntax element max_reduced_transform_size is not signaled, the value of max_reduced_transform_size may be assumed to be 0.

[0195] Also, the fifth syntax element may be expressed as the syntax element "reduced_transform_size." The size value of the simplified inverse transform signaled in the syntax element "reduced_transform_size" may indicate, but is not limited to, the size of the region to which the simplified inverse transform is applied or the size of the simplified transform matrix. If the syntax element "reduced_transform_size" is not signaled, the value of "reduced_transform_size" may be assumed to be 0.

[0196] An example in which information on quantized transform coefficients for a current block is included in the SPS and signaled is shown in Table 3 below.

[0197] [Table 3]

[0198] The decoding apparatus 200 according to an embodiment may derive transform coefficients by performing inverse quantization on the quantized transform coefficients for the current block (S1210).

[0199] The decoding apparatus 200 according to an embodiment may derive residual samples for the current block based on a simplified inverse transform of the transform coefficients (S1220). In one example, the simplified inverse transform may be performed based on a simplified inverse transform matrix, which is a non-square matrix having fewer columns than rows.

[0200] In one embodiment, step S1220 may include the steps of decoding a transform index, determining whether a condition for applying a simplified inverse transform is met based on the transform index, selecting a transform kernel, and, if the condition for applying the simplified inverse transform is met, applying a simplified inverse transform to the transform coefficients based on the selected transform kernel and a simplification factor. In this case, the size of the simplified inverse transform matrix may be determined based on the simplification factor.

[0201] If the simplified inverse transform performed in S1220 is based on the inverse NSST, first-order transform coefficients for the current block may be derived by performing the simplified inverse transform on the transform coefficients for the current block. The decoding apparatus 200 may perform a first-order inverse transform on the first-order transform coefficients for the current block, and in this case, the first-order inverse transform may be performed based on the simplified inverse transform or may not be performed based on the simplified inverse transform.

[0202] Alternatively, if the simplified inverse transform performed by S1220 is based on a linear inverse transform, residual samples for the current block may be derived by performing the simplified inverse transform on the transform coefficients for the current block.

[0203] The decoding apparatus 200 according to an embodiment may generate a reconstructed picture based on the residual sample for the current block and the predicted sample for the current block (S1230).

[0204] Referring to S1220, it can be seen that residual samples for a target block are derived based on a simplified inverse transform of transform coefficients for the target block. Considering the size of the inverse transform matrix, the size of a typical inverse transform matrix is ​​N×N, whereas the size of the simplified inverse transform matrix is ​​reduced to N×R. Therefore, compared to performing a typical transform, memory usage can be reduced by a ratio of R / N when performing a simplified transform. Furthermore, compared to the number of multiplication operations (N×N) required when using a typical inverse transform matrix, the number of multiplication operations can be reduced by a ratio of R / N (N×R) when using a simplified inverse transform matrix. Furthermore, since only R transform coefficients need to be decoded when applying the simplified inverse transform, compared to the number of transform coefficients required when applying a typical inverse transform (N), the total number of transform coefficients for the target block is reduced from N to R, thereby improving decoding efficiency. In summary, S1220 improves the (inverse) transform efficiency and coding efficiency of the decoding device 200 through the simplified inverse transform.

[0205] The internal components of the device described above may be a processor that executes a series of executable processes stored in memory, or other hardware components that are configured with hardware, and may be located inside or outside the device.

[0206] The aforementioned modules may be omitted depending on the embodiment or may be replaced by other modules performing similar / same operations.

[0207] The above-described method according to the present invention can be implemented in software form, and the encoding device and / or decoding device according to the present invention can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.

[0208] In the above-described embodiments, the method is described based on a flowchart with a series of steps or blocks, but the present invention is not limited to the order of steps, and some steps may occur in a different order or simultaneously with other steps than those described. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps in the flowcharts may be omitted without affecting the scope of the present invention.

[0209] When an embodiment of the present invention is implemented in software, the above-described methods may be implemented with modules (processes, functions, etc.) that perform the above-described functions. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices.

Claims

1. A video decoding method performed by a decoding device, comprising: obtaining information about quantized transform coefficients from the bitstream; deriving quantized transform coefficients for a current block based on the information about the quantized transform coefficients; performing inverse quantization on the quantized transform coefficients for the current block to derive transform coefficients; performing a quadratic inverse transform on the transform coefficients based on an inverse transform matrix, the quadratic inverse transform being based on a non-separable transform; deriving primary transform coefficients for the current block based on the result of the secondary inverse transform; deriving residual samples for the current block by performing an inverse linear transform on the linear transform coefficients; generating a reconstructed picture based on the residual samples for the current block; The size of the inverse transform matrix is ​​NxR, where R is equal to the number of elements in the input domain of the quadratic inverse transform based on the non-separable transform, R is less than N, and each of N and R is a positive integer.

2. A video encoding method performed by an encoding device, comprising: deriving a residual sample for the current block; performing a linear transform on the residual samples to derive linear transform coefficients for the current block; performing a secondary transform on the primary transform coefficients based on a transformation matrix, the secondary transform being based on a non-separable transform; deriving transform coefficients for the current block based on the results of the secondary transform; performing quantization on the transform coefficients to derive quantized transform coefficients for the current block; generating information about the quantized transform coefficients; encoding the information about the quantized transform coefficients to output a bitstream; The size of the transformation matrix is ​​RxN, where N is equal to the number of elements in the input domain of the quadratic transform based on the non-separable transform, R is less than N, and each of N and R is a positive integer.

3. A method for transmitting data for video, comprising: obtaining a bitstream for the video, the bitstream comprising: deriving a residual sample for the target block; deriving linear transform coefficients for the current block by performing a linear transform on the residual samples; performing a secondary transform on the primary transform coefficients based on a transform matrix, the secondary transform being based on a non-separable transform; deriving transform coefficients for the current block based on the results of the secondary transform; performing quantization on the transform coefficients to derive quantized transform coefficients for the current block; generating information about the quantized transform coefficients; and generating the information about the quantized transform coefficients based on encoding the information about the quantized transform coefficients. transmitting the data including the bitstream; The size of the transformation matrix is ​​RxN, where N is equal to the number of elements in the input domain of the quadratic transform based on the non-separable transform, R is less than N, and each of N and R is a positive integer.

Citation Information

Patent Citations

  • Transform-based video coding method and apparatus

    JP7240476B2

  • Method and device for the transformation and method and device for the reverse transformation of images

    US20130195177A1

  • Reduced size inverse transform for decoding and encoding

    US20170034530A1

  • Video decoder memory optimization

    US20170230677A1

  • Non-separable secondary transform for video coding

    WO2017058614A1