Video coding method and apparatus based on selective transformation

The video decoding method employs a modified transform matrix with selected basis vectors to enhance coding efficiency by reducing data transfer and computational complexity, addressing the increased costs of high-resolution video transmission and storage.

JP7867115B2Active Publication Date: 2026-05-28LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality videos leads to higher transmission and storage costs due to increased video data, necessitating improved video coding efficiency and conversion efficiency.

Method used

A video decoding method and apparatus utilize a modified transform matrix with selected basis vectors to perform selective transforms, reducing the amount of data transferred and computational complexity for residual processing.

Benefits of technology

This approach enhances residual coding efficiency by minimizing data transfer and memory load, improving overall coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007867115000015
    Figure 0007867115000015
  • Figure 0007867115000016
    Figure 0007867115000016
  • Figure 0007867115000017
    Figure 0007867115000017
Patent Text Reader

Abstract

To achieve a method and a device for increasing the efficiency of picture coding.SOLUTION: The method for picture decoding by a decoder device according to the present invention includes deriving a residual sample of a target block on the basis of a non-separation conversion of a conversion coefficient based on a bit stream. A prediction sample is derived on the basis of an intra-prediction, one of 67 intra-prediction modes is used to derive a prediction sample, and 67 intra-prediction modes have 2 non-angular prediction mode and 65 angular prediction modes. A non-separation conversion is performed on the basis of a conversion matrix with a corrected bottom vector, and the corrected bottom vector has less elements than N number of elements. N is equal to the number of conversion coefficients located in a region to which a non-separation conversion is applied in a target block. A region to which a non-separation conversion is applied is 8x8 upper-left stage target regions of the target block and N is equal to 64.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video coding technology, and more particularly, to a video decoding method and apparatus according to selective conversion in a video coding system.

Background Art

[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As the video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to the existing video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video data using an existing storage medium, the transmission cost and storage cost increase.

[0003]

Summary of the Invention

Problems to be Solved by the Invention

[0004] The technical problem of the present invention is to provide a method and apparatus for increasing video coding efficiency.

[0005] Another technical problem of the present invention is to provide a method and apparatus for increasing conversion efficiency.

[0006]

[0007] Yet another technical problem of the present invention is to provide a method and apparatus for increasing the efficiency of residual coding through conversion. <x

Means for Solving the Problem

[0008] According to an embodiment of the present invention, a video decoding method performed by a decoding device is provided. The method includes: deriving transform coefficients of a target block from a bitstream; deriving residual samples for the target block based on a selective transform on the transform coefficients; and generating a reconstructed picture based on the residual samples for the target block and prediction samples for the target block. The selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix having modified basis vectors, and the modified basis vectors have a selected specific number of elements out of N elements.

[0009] According to another embodiment of the present invention, a decoding device for performing video decoding is provided. The decoding device includes: an entropy decoding unit that derives transform coefficients of a target block from a bitstream; an inverse transform unit that derives residual samples for the target block based on a selective transform on the transform coefficients; and an addition unit that generates a reconstructed picture based on the residual samples for the target block and prediction samples for the target block. The selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix having modified basis vectors, and the modified basis vectors have a selected specific number of elements out of N elements.

[0010] According to yet another embodiment of the present invention, a video encoding method performed by an encoding device is provided. The method comprises the steps of: deriving residual samples of a target block; deriving transformation coefficients of the target block based on a selective transform of the residual samples; and encoding information regarding the transformation coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix being a matrix having modified basis vectors, and the modified basis vectors having a specific number of elements selected from N elements.

[0011] According to yet another embodiment of the present invention, a video encoding device is provided. The encoding device comprises an additive unit for deriving residual samples of a target block, a transforming unit for deriving transform coefficients of the target block based on a selective transform of the residual samples, and an entropy encoding unit for encoding information about the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is ​​a matrix having modified basis vectors, and the modified basis vectors have a specific number of elements selected from N elements. [Effects of the Invention]

[0012] According to the present invention, the amount of data that must be transferred for residual processing can be reduced through efficient conversion, thereby improving residual coding efficiency.

[0013] According to the present invention, an inseparable transformation can be performed based on a transformation matrix composed of basis vectors containing a specific number of selected elements. This reduces the memory load and computational complexity for the inseparable transformation, thereby improving residual coding efficiency.

[0014] According to the present invention, non-separable transformations can be performed based on a simplified transformation matrix, thereby reducing the amount of data that must be transferred (transmitted) for residual processing and improving residual coding efficiency. [Brief explanation of the drawing]

[0015] [Figure 1] This figure schematically illustrates the configuration of a video encoding device to which the present invention can be applied. [Figure 2] This figure schematically illustrates the configuration of a video decoding device to which the present invention can be applied. [Figure 3] This figure schematically illustrates the multiple transform technique according to the present invention. [Figure 4] This diagram illustrates 65 intra-directional modes for predicting directions. [Figure 5A] This figure illustrates a selective transform according to one embodiment of the present invention. [Figure 5B] This figure illustrates a selective transform according to one embodiment of the present invention. [Figure 5C] This figure illustrates a selective transform according to one embodiment of the present invention. [Figure 6] This diagram schematically illustrates a multiple transformation technique that applies the above selective transformation to a quadratic transformation. [Figure 7] This figure illustrates the arrangement of conversion coefficients based on a target block according to an embodiment of the present invention. [Figure 8] This figure shows an example of deriving transformation coefficients through a transformation that combines the above simplification transformation and the above selective transformation. [Figure 9] This figure shows an example of deriving conversion coefficients through the selective transformation described above. [Figure 10] This figure shows an example of deriving the above-mentioned association vector and performing a selective transformation based on two factors applied to the association vector. [Figure 11] This figure schematically illustrates a video encoding method using an encoding device according to the present invention. [Figure 12] This figure schematically shows an encoding device that performs a video encoding method according to the present invention. [Figure 13] This figure schematically illustrates a video decoding method using a decoding device according to the present invention. [Figure 14] This figure schematically shows a decoding device that performs a video decoding method according to the present invention. [Modes for carrying out the invention]

[0016] The present invention can be modified in various ways and has many embodiments, and specific embodiments are illustrated and described in detail in the drawings. However, this does not limit the present invention to any particular embodiment. The terms used herein are used solely to describe specific embodiments and are not intended to limit the technical idea of ​​the present invention. Singular expressions include plural expressions unless they are clearly different in context. In this specification, terms such as “includes” or “has” indicate the presence of features, figures, steps, actions, components, parts or combinations thereof described in the specification, and should be understood not to preclude the presence or addition of one or more other features, figures, steps, actions, components, parts or combinations thereof.

[0017] On the other hand, each component shown in the drawings described in this invention is illustrated independently for the convenience of explaining its distinct characteristic functions, and does not mean that each component is embodied in separate hardware or separate software. For example, two or more components may be combined to form a single component, and a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the invention, as long as they do not deviate from the essence of the invention.

[0018] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings.Hereinafter, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.

[0019] On the other hand, the present invention relates to video / image coding. For example, the methods / examples disclosed in the present invention can be applied to methods disclosed in the VVC (Versatile Video Coding) standard or next-generation video / image coding.

[0020] In this specification, "picture" generally refers to a unit representing a single image at a specific time period, and "slice" is a unit that constitutes a part of a picture in coding. A single picture may consist of multiple slices, and pictures and slices may be mixed together as needed.

[0021] A pixel or pel can refer to the smallest unit that makes up a picture (or image). The term "sample" can also be used as a counterpart to "pixel." A sample generally refers to a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance (luma) component, or only the pixel / pixel value of the chroma component.

[0022] A unit represents a basic unit of image processing. A unit can contain at least one of the following: a specific region of a picture and information about that region. A unit may sometimes be used interchangeably with terms such as block or area. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows.

[0023] Figure 1 is a schematic diagram illustrating the configuration of a video encoding device to which the present invention can be applied.

[0024] Referring to Figure 1, the video encoding device 100 may include a picture splitting unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a conversion unit 122, a quantization unit 123, a realignment unit 124, an inverse quantization unit 125, and an inverse conversion unit 126.

[0025] The picture splitting unit 105 can split the input picture into at least one processing unit.

[0026] For example, a processing unit is called a coding unit (CU). In this case, a coding unit can be recursively divided from the largest coding unit (LCU) by a quad-tree binary-tree (QTBT) structure. For example, a single coding unit can be divided into multiple deeper coding units based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to the present invention can be executed. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.

[0027] As another example, a processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). A coding unit can be split from the largest coding unit (LCU) into deeper coding units using a quadtree structure. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively split into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. If a smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to the underlying coding unit that is partitioned or split into prediction units or transform units. A prediction unit is a unit partitioned from a coding unit and is a unit for sample prediction. In this case, the prediction unit can also be divided into subblocks. The transformation unit can be separated from the coding unit by a quadtree structure and is a unit that derives (induces) transformation coefficients and / or a unit that derives a residual signal from the transformation coefficients. Hereafter, the coding unit will also be called the coding block (CB), the prediction unit the prediction block (PB), and the transformation unit the transformation block (TB). A prediction block or prediction unit refers to a specific region in block form within a picture and may contain an array of prediction samples.Furthermore, a conversion block or conversion unit refers to a specific region in block form within a picture and may include an array of conversion coefficients or residual samples.

[0028] The prediction unit 110 performs a prediction on the block to be processed (hereinafter referred to as the current block) and can generate a predicted block that includes prediction samples for the current block. The unit of prediction performed by the prediction unit 110 is a coding block, or a transformation block, or a prediction block.

[0029] The prediction unit 110 can determine whether intra-prediction or inter-prediction is applied to the current block. For example, the prediction unit 110 can determine whether intra-prediction or inter-prediction is applied on a CU basis.

[0030] In the case of intra-prediction, the prediction unit 110 can derive prediction samples for the current block based on reference samples outside the current block within the picture to which the current block belongs (hereinafter referred to as the current picture). At this time, the prediction unit 110 can (i) derive prediction samples based on the average or interpolation of neighboring reference samples of the current block, or (ii) derive the above prediction samples based on reference samples among the neighboring reference samples of the current block that exist in a specific (prediction) direction relative to the prediction sample. Case (i) is called a non-directional mode or non-angular mode, and case (ii) is called a directional mode or angular mode. The prediction modes in intra-prediction can, for example, have 33 directional prediction modes and at least 2 or more non-directional modes. Non-directional modes can include DC prediction modes and planar modes. The prediction unit 110 can also use the prediction modes applied to neighboring blocks to determine the prediction mode to be applied to the current block.

[0031] In interpretation, the prediction unit 110 can derive predicted samples for the current block based on samples identified by motion vectors on a reference picture. The prediction unit 110 can derive predicted samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP (Motion Vector Prediction) mode. In skip mode and merge mode, the prediction unit 110 can use motion information from adjacent blocks as motion information for the current block. In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted. In MVP mode, the motion vectors of adjacent blocks can be used as motion vector predictors and used as motion vector predictors for the current block to derive the motion vectors of the current block.

[0032] In interpretation, adjacent blocks can include spatially adjacent blocks within the current picture and temporally adjacent blocks within the reference picture. A reference picture containing the aforementioned temporally adjacent blocks is also called a collocated picture (colPic). Motion information can include motion vectors and reference picture indices. Information such as prediction mode information and motion information can be (entropy) encoded and output in bitstream format.

[0033] In skip mode and merge mode, when motion information for temporally adjacent blocks is used, the top-level picture in the reference picture list can also be used as the reference picture. Reference pictures included in the Picture Order Count can be sorted based on the difference in Picture Order Count (POC) between the current picture and the corresponding reference picture. The POC corresponds to the display order of the pictures and can be distinguished from the coding order.

[0034] The subtraction unit 121 generates a residual sample, which is the difference between the original sample and the predicted sample. When skip mode is applied, a residual sample is not generated, as described above.

[0035] The transformation unit 122 transforms residual samples in units of transformation blocks to generate transformation coefficients. The transformation unit 122 can perform the transformation depending on the size of the transformation block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the transformation block. For example, if intraprediction is applied to the coding block or prediction block that overlaps with the transformation block, and the transformation block is a 4x4 residual array, the residual samples are transformed using a DST (Discrete Sine Transform) transformation kernel. Otherwise, the residual samples can be transformed using a DCT (Discrete Cosine Transform) transformation kernel.

[0036] The quantization unit 123 can quantize the conversion coefficients and generate the quantized conversion coefficients.

[0037] The realignment unit 124 realigns the quantized transformation coefficients. The realignment unit 124 can realign the block-shaped quantized transformation coefficients into a one-dimensional vector form via a coefficient scanning method. Here, although the realignment unit 124 is described in a separate configuration, it may also be part of the quantization unit 123.

[0038] The entropy encoding unit 130 can perform entropy encoding on the quantized conversion coefficients. Entropy encoding can include encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding). In addition to the quantized conversion coefficients, the entropy encoding unit 130 can also encode information necessary for video restoration (e.g., the values ​​of syntax elements) together with or separately by entropy encoding or by a pre-configured method. The encoded information can be transmitted or stored in bitstream form in units of NAL (Network Abstraction Layer) units.

[0039] The inverse quantization unit 125 inversely quantizes the values ​​(quantized conversion coefficients) quantized by the quantization unit 123, and the inverse transformation unit 126 inversely transforms the values ​​inversely quantized by the inverse quantization unit 125 to generate residual samples.

[0040] The addition unit 140 reconstructs the picture by adding the residual sample and the predicted sample. The residual sample and the predicted sample are added in block units to generate a reconstructed block. Here, although the addition unit 140 is described in a separate configuration, it may also be part of the prediction unit 110. On the other hand, the addition unit 140 is also called the reconstruction module or the reconstructed block generation unit.

[0041] The filter unit 150 can apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries and distortions during the quantization process within the reconstructed picture can be corrected. The sample adaptive offset can be applied on a sample-by-sample basis and can be applied after the deblocking filtering process is complete. The filter unit 150 can also apply an Adaptive Loop Filter (ALF) to the reconstructed picture. The ALF can be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset have been applied.

[0042] Memory 160 can store a restored picture (decoded picture) or information necessary for encoding / decoding. Here, the restored picture is the restored picture after the filtering procedure by the filter unit 150 has been completed. The stored restored picture can be used as a reference picture for (inter)prediction of other pictures. For example, memory 160 can store a (reference) picture used for interprediction. In this case, the picture used for interprediction can be specified by a reference picture set or a reference picture list.

[0043] Figure 2 is a schematic diagram illustrating the configuration of a video decoding device to which the present invention can be applied.

[0044] Referring to Figure 2, the video decoding device 200 may include an entropy decoding unit 210, a residual processing unit 220, a prediction unit 230, an addition unit 240, a filter unit 250, and a memory 260. Here, the residual processing unit 220 may include a realignment unit 221, an inverse quantization unit 222, and an inverse transformation unit 223.

[0045] When a bitstream containing video information is input, the video decoder 200 can reconstruct the video in accordance with the process by which the video information was processed by the video encoder.

[0046] For example, the video decoding device 200 can perform video decoding using the processing units applied in the video encoding device. Therefore, a processing unit block for video decoding is, as an example, a coding unit, and as another example, a coding unit, a prediction unit, or a transformation unit. A coding unit can be divided from a maximum coding unit by a quadtree structure and / or a binary tree structure.

[0047] Prediction units and transformation units may be used further, in which case the prediction block is a block derived or partitioned from the coding unit and is a unit of sample prediction. In this case, the prediction unit may also be divided into subblocks. Transformation units may be partitioned from the coding unit by a quadtree structure and are units that derive transformation coefficients or units that derive residual signals from transformation coefficients.

[0048] The entropy decoding unit 210 can purge the bitstream and output information necessary for video or picture restoration. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of the syntax elements necessary for video restoration and the quantized values ​​of the conversion coefficients for the residuals.

[0049] More specifically, the CABAC entropy decoding method receives a BIN corresponding to each syntax element in a bitstream, determines a context model using the information of the syntax element to be decoded, as well as the decoding information of adjacent and decoded blocks, or the information of symbols / BINs decoded in a previous step, predicts the probability of BIN occurrence based on the determined context model, and performs arithmetic decoding of the BIN to generate symbols corresponding to the values ​​of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model for the next symbol / BIN using the information of the decoded symbol / BIN.

[0050] Information related to predictions from the information decoded by the entropy decode unit 210 is provided to the prediction unit 230, and the residual values ​​after entropy decoding is performed in the entropy decode unit 210, i.e., the quantized conversion coefficients, can be input to the realignment unit 221.

[0051] The realignment unit 221 can realign the quantized transformation coefficients in a two-dimensional block form. The realignment unit 221 can perform realignment in response to the coefficient scan performed by the encoding device. Here, although the realignment unit 221 has been described in a separate configuration, it may also be part of the inverse quantization unit 222.

[0052] The inverse quantization unit 222 can inverse quantize the quantized conversion coefficients based on (inverse) quantization parameters and output the conversion coefficients. At this time, information for deriving the quantization parameters can be signaled from the encoding device.

[0053] The inverse transformation unit 223 can derive the residual sample by inversely transforming the transformation coefficients.

[0054] The prediction unit 230 can perform predictions on the current block and generate a predicted block that includes prediction samples for the current block. The unit of prediction performed by the prediction unit 230 is a coding block, or a transformation block, or a prediction block.

[0055] The prediction unit 230 can decide whether to apply intra-prediction or inter-prediction based on the prediction information. In this case, the unit for deciding which of intra-prediction and inter-prediction to apply is different from the unit for generating prediction samples. In addition, the unit for generating prediction samples is also different for inter-prediction and intra-prediction. For example, the decision on whether to apply inter-prediction or intra-prediction can be made at the CU (Unit of Reference) level. Also, for example, in inter-prediction, the prediction mode can be determined at the PU (Phone Unit) level and prediction samples can be generated, and in intra-prediction, the prediction mode can be determined at the PU level and prediction samples can be generated at the TU (Unit of Reference) level.

[0056] In the case of intra-prediction, the prediction unit 230 can derive prediction samples for the current block based on adjacent reference samples in the current picture. The prediction unit 230 can derive prediction samples for the current block by applying a directional mode or a non-directional mode based on adjacent reference samples in the current block. At this time, the prediction mode to be applied to the current block can also be determined by utilizing the intra-prediction mode of the adjacent block.

[0057] In the case of interpretation, the prediction unit 230 can derive predicted samples for the current block based on samples identified on the reference picture by motion vectors on the reference picture. The prediction unit 230 can derive predicted samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP mode. At this time, motion information necessary for interpretation of the current block provided by the video encoding device, such as motion vectors and reference picture indices, can be obtained or derived based on the above prediction information.

[0058] In skip mode and merge mode, the movement information of adjacent blocks can be used as the movement information of the current block. In this case, adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.

[0059] The prediction unit 230 constructs a merge candidate list using motion information of available adjacent blocks, and can use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. Motion information can include motion vectors and reference pictures. When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top-level picture on the reference picture list can be used as the reference picture.

[0060] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.

[0061] In MVP mode, the motion vector of the current block can be derived by using the motion vector of the adjacent block as a motion vector predictor. In this case, adjacent blocks can include both spatially adjacent blocks and temporally adjacent blocks.

[0062] For example, when merge mode is applied, a merge candidate list can be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col. In merge mode, the motion vectors of the candidate blocks selected from the merge candidate list are used as the motion vectors of the current block. The information regarding the above prediction may include a merge index that indicates the candidate block with the optimal motion vector selected from among the candidate blocks included in the merge candidate list. In this case, the prediction unit 230 can use the merge index to derive the motion vector of the current block.

[0063] As another example, when the MVP (Motion Vector Prediction) mode is applied, a list of motion vector predictor candidates can be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col. That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col, can be used as motion vector candidates. The prediction information may include a prediction motion vector index that indicates the optimal motion vector selected from the motion vector candidates included in the list. At this time, the prediction unit 230 can use the motion vector index to select the predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list. The prediction unit of the encoding device can calculate the difference in motion vectors (MVD) between the motion vector of the current block and the motion vector predictor, and can encode this and output it in bitstream format. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit 230 can obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. Furthermore, the prediction unit can obtain or derive a reference picture index that indicates a reference picture, etc., from the information related to the prediction.

[0064] The addition unit 240 can reconstruct the current block or current picture by adding the residual sample and the predicted sample. The addition unit 240 can also reconstruct the current picture by adding the residual sample and the predicted sample in block units. When skip mode is applied, the residual is not transmitted, so the predicted sample can become the reconstructed sample. Here, although the addition unit 240 has been described in a separate configuration, it may also be part of the prediction unit 230. On the other hand, the addition unit 240 is also called the reconstruction module or the reconstructed block generation unit.

[0065] The filter unit 250 can apply deblocking filtering, sample-adaptive offset, and / or ALF to the restored picture. The sample-adaptive offset can be applied on a sample-by-sample basis and can also be applied after deblocking filtering. ALF can also be applied after deblocking filtering and / or sample-adaptive offset.

[0066] Memory 260 can store a restored picture (decoded picture) or information necessary for decoding. Here, the restored picture is a restored picture after the filtering procedure by the filter unit 250 has been completed. For example, memory 260 can store a picture used for inter prediction. In this case, the picture used for inter prediction can also be specified by a reference picture set or a reference picture list. The restored picture can be used as a reference picture for other pictures. Memory 260 can also output the restored pictures in the order they are output.

[0067] On the other hand, through the aforementioned transformation, a lower frequency transformation coefficient can be derived for the residual block of the current block, and a zero tail can be derived at the end of the residual block.

[0068] Specifically, the above transformation can consist of two main processes, the main processes of which may include a core transform and a secondary transform. A transformation including the core transform and the secondary transform may be referred to as a multiple transformation technique.

[0069] Figure 3 schematically shows the multiplexing technique according to the present invention.

[0070] Referring to Figure 3, the conversion unit can correspond to the conversion unit in the encoding device shown in Figure 1, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device shown in Figure 1 or the inverse conversion unit in the decoding device shown in Figure 2.

[0071] The transformation unit can perform a linear transformation based on the residual samples (residual sample array) in the residual block to derive (linear) transformation coefficients (S310). Here, the linear transformation may include an Adaptive Multiple Core Transform (AMT). The Adaptive Multiple Core Transform may be denoted as an MTS (Multiple Transform Set).

[0072] The above adaptive multicore transformation can also be described as a method of transformation that additionally uses DCT (Discrete Cosine Transform) type 2 and DST (Discrete Sine Transform) type 7, DCT type 8, and / or DST type 1. That is, the above adaptive multicore transformation can be described as a method of transformation that transforms a residual signal (or residual block) in the spatial domain into a transformation coefficient (or first-order transformation coefficient) based on a selection of transformation kernels from DCT type 2, DST type 7, DCT type 8, and DST type 1. Here, the first-order transformation coefficient can also be called a temporary (provisional) transformation coefficient from the perspective of the transformation unit.

[0073] In other words, when an existing conversion method is applied, a spatial-to-frequency domain conversion of the residual signal (or residual block) can be applied based on DCT type 2 to generate conversion coefficients. In contrast, when the adaptive multicore conversion described above is applied, a spatial-to-frequency domain conversion of the residual signal (or residual block) can be applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or first-order conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be referred to as conversion types, conversion kernels, or conversion cores.

[0074] For reference, the above DCT / DST conversion types can be defined based on basis functions, which are shown in the following table.

[0075] [Table 1]

[0076] When the above adaptive multicore transformation is performed, a vertical transformation kernel and a horizontal transformation kernel can be selected from the transformation kernels for the target block, and a vertical transformation can be performed on the target block based on the vertical transformation kernel, and a horizontal transformation can be performed on the target block based on the horizontal transformation kernel. Here, the horizontal transformation can represent a transformation on the horizontal component of the target block, and the vertical transformation can represent a transformation on the vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on a transformation index that points to the prediction mode and / or transformation subset of the target block (CU or subblock) encompassing the residual block.

[0077] For example, the adaptive multicore transformation described above can be applied when the width and height of the target block are both less than or equal to 64, and whether or not the adaptive multicore transformation described above can be applied to the target block can be determined based on the CU level flag. Specifically, if the CU level flag is 0, the existing transformation method described above can be applied. That is, if the CU level flag is 0, a spatial-to-frequency domain transformation can be applied to the residual signal (or residual block) based on the DCT type 2 to generate transformation coefficients, and these transformation coefficients can be encoded. On the other hand, the target block here can be a CU. If the CU level flag is 0, the adaptive multicore transformation described above can be applied to the target block.

[0078] Furthermore, in the case of a Luma block of the target block to which the above adaptive multicore transformation is applied, two additional flags can be signaled, and a vertical and horizontal transformation kernel can be selected based on these flags. The flag relating to the vertical transformation kernel can be denoted as the AMT vertical flag, where AMT_TU_vertical_flag (or EMT_TU_vertical_flag) represents the syntax element of the AMT vertical flag. The flag relating to the horizontal transformation kernel can be denoted as the AMT horizontal flag, where AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) represents the syntax element of the AMT horizontal flag. The AMT vertical flag can point to one of the transformation kernel candidates included in the transformation subset for the vertical transformation kernel, and the transformation kernel candidate pointed to by the AMT vertical flag can be derived as the vertical transformation kernel for the target block. Furthermore, the AMT horizontal flag can point to one of the candidate transformation kernels included in the transformation subset for the horizontal transformation kernel, and the candidate transformation kernel pointed to by the AMT horizontal flag can be derived as the horizontal transformation kernel for the target block. On the other hand, the AMT vertical flag can also be represented as the MTS vertical flag, and the AMT horizontal flag can also be represented as the MTS horizontal flag.

[0079] On the other hand, three transformation subsets can be pre-configured, and one of these transformation subsets can be derived as a transformation subset for the vertical transformation kernel based on the intra-prediction mode applied to the target block. Also, one of these transformation subsets can be derived as a transformation subset for the horizontal transformation kernel based on the intra-prediction mode applied to the target block. For example, the pre-configured transformation subsets can be derived as shown in the following table.

[0080] [Table 2]

[0081] Referring to Table 2, a transformation subset with an index value of 0 can represent a transformation subset that includes DST type 7 and DCT type 8 as transformation kernel candidates, a transformation subset with an index value of 1 can represent a transformation subset that includes DST type 7 and DST type 1 as transformation kernel candidates, and a transformation subset with an index value of 2 can represent a transformation subset that includes DST type 7 and DCT type 8 as transformation kernel candidates.

[0082] The transformation subsets for the vertical transformation kernel and the horizontal transformation kernel, derived based on the intra-prediction mode applied to the target block described above, can be derived as shown in the following table.

[0083] [Table 3]

[0084] Here, V represents a subset of the transformations for the vertical transformation kernel, and H represents a subset of the transformations for the horizontal transformation kernel.

[0085] If the value of the AMT flag (or EMT_CU_flag) for the target block is 1, then, as shown in Table 3, a transformation subset for the vertical transformation kernel and a transformation subset for the horizontal transformation kernel can be derived based on the intra-prediction mode of the target block. Subsequently, among the transformation kernel candidates included in the transformation subset for the vertical transformation kernel, the transformation kernel candidate pointed to by the AMT vertical flag of the target block can be derived as the vertical transformation kernel of the target block, and among the transformation kernel candidates included in the transformation subset for the horizontal transformation kernel, the transformation kernel candidate pointed to by the AMT horizontal flag of the target block can be derived as the horizontal transformation kernel of the target block. On the other hand, the AMT flag can also be referred to as the MTS flag.

[0086] For reference, for example, an intra-prediction mode may include two non-directional (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. The non-directional intra-prediction modes may include a planar intra-prediction mode (number 0) and a DC intra-prediction mode (number 1), and the directional intra-prediction modes may include 65 intra-prediction modes (numbers 2 through 66). However, this is merely an example, and the present invention can also be applied when the number of intra-prediction modes differs. On the other hand, a 67th intra-prediction mode may be used in some cases, and this 67th intra-prediction mode may represent an LM (Linear Model) mode.

[0087] Figure 4 illustrates 65 intradirectional modes for predicting directions.

[0088] Referring to Figure 4, intra-prediction modes can be divided into those with horizontal directionality and those with vertical directionality, centered around intra-prediction mode 34, which has a diagonal prediction direction pointing upward to the left. In Figure 4, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 units on the sample grid position. Intra-prediction modes 2 through 33 have horizontal directionality, while intra-prediction modes 34 through 66 have vertical directionality. Intra-prediction modes 18 and 50 represent horizontal intra-prediction modes and vertical intra-prediction modes, respectively. Intra-prediction mode 2 can be called a diagonal intra-prediction mode pointing downward to the left, intra-prediction mode 34 a diagonal intra-prediction mode pointing upward to the left, and intra-prediction mode 66 a diagonal intra-prediction mode pointing upward to the right.

[0089] The transformation unit can derive (secondary) transformation coefficients by performing a quadratic transformation based on the above (primary) transformation coefficients (S320). If the above primary transformation was a transformation from the spatial domain to the frequency domain, then the above quadratic transformation can be seen as a transformation from the frequency domain to the frequency domain. The above quadratic transformation may include a non-separable transform. In this case, the above quadratic transformation may be called a Non-Separable Secondary Transform (NSST) or MDNSST (Mode-Dependent Non-Separable Secondary Transform). The above non-separable quadratic transformation can represent a transformation that generates transformation coefficients (or quadratic transformation coefficients) for the residual signal by performing a quadratic transformation on the (primary) transformation coefficients derived through the above primary transformation based on a non-separable transform matrix. Here, the vertical and horizontal transformations are not applied separately (or horizontal and vertical transformations are applied independently) to the above (primary) transformation coefficients based on the non-separable transform matrix, but the transformation can be applied all at once. In other words, the above inseparable quadratic transformation can be described as a transformation method that generates transformation coefficients (or quadratic transformation coefficients) by transforming both the vertical and horizontal components of the (linear) transformation coefficients together, without separating them, based on the above inseparable transformation matrix. The above inseparable quadratic transformation can be applied to the top-left region of a block composed of (linear) transformation coefficients (hereinafter referred to as a transformation coefficient block or target block). For example, if both the width (W) and height (H) of the above transformation coefficient block are 8 or greater, an 8x8 inseparable quadratic transformation can be applied to the top-left 8x8 region of the above transformation coefficient block (hereinafter referred to as the top-left target region). Also, if both the width (W) and height (H) of the above transformation coefficient block are 4 or greater, and either the width (W) or height (H) of the above transformation coefficient block is less than 8, a 4x4 inseparable quadratic transformation can be applied to the top-left min(8,W)×min(8,H) region of the above transformation coefficient block.

[0090] Specifically, for example, when a 4x4 input block is used, the unseparated quadratic transform can be performed as follows:

[0091] The above 4x4 input block X can be represented as follows:

[0092] <Formula 1>

number

[0093] When the above X is shown in vector form, TIFF0007867115000005.tif84 can be represented as follows:

[0094] <Formula 2>

number

[0095] In this case, the above quadratic inseparable transform can be calculated as follows:

[0096] <Formula 3>

number

[0097] Here, TIFF0007867115000008.tif75 represents the transformation coefficient vector, and T represents the 16×16 (inseparable) transformation matrix.

[0098] Through the above formula 3, a 16 × 1 transformation coefficient vector TIFF0007867115000009.tif75 can be derived, and the above TIFF0007867115000010.tif75 can be reorganized into 4x4 blocks through a scan order (horizontal, vertical, diagonal, etc.). However, the calculation described above is just an example, and methods such as HyGT (Hypercube-Givens Transform) can also be used to calculate inseparable quadratic transforms in order to reduce the computational complexity of the inseparable quadratic transform.

[0099] On the other hand, in the above non-separable quadratic transformation, the transformation kernel (or transformation core, transformation type) can be selected to be mode-dependent. Here, the mode can include intra-predictive mode and / or inter-predictive mode.

[0100] As mentioned above, the non-separable quadratic transformation can be performed based on an 8x8 transformation or a 4x4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. That is, the non-separable quadratic transformation can be performed based on an 8x8 subblock size or a 4x4 subblock size. For example, for the mode-dependent transformation kernel selection, 35 sets of 3 non-separable quadratic transformation kernels can be configured for both the 8x8 subblock size and the 4x4 subblock size. That is, 35 transformation sets can be configured for the 8x8 subblock size, and 35 transformation sets can be configured for the 4x4 subblock size. In this case, each of the 35 transformation sets for the 8x8 subblock size may contain 3 8x8 transformation kernels, and each of the 35 transformation sets for the 4x4 subblock size may contain 3 4x4 transformation kernels. However, the above-mentioned conversion subblock size, the number of sets, and the number of conversion kernels within a set can be sizes other than, for example, 8x8 or 4x4, or there can be n sets, each containing k conversion kernels.

[0101] The above transformation set may be referred to as an NSST set, and the transformation kernel within the above NSST set may be referred to as an NSST kernel. The selection of a specific set from the above transformation sets can be performed, for example, based on the intra-prediction mode of the target block (CU or subblock).

[0102] In this case, the mapping between the 35 transformation sets and the intra-prediction modes can be shown, for example, as in the following table. For reference, if the LM mode is applied to the target block, the quadratic transformation may not be applied to that block.

[0103] [Table 4]

[0104] On the other hand, if it is determined that a specific set is to be used, one of the k transformation kernels within that specific set can be selected via an inseparable quadratic transformation index. The encoding device can derive an inseparable quadratic transformation index that points to a specific transformation kernel based on an RD (Rate-Distortion) check, and can signal the decoding device to the inseparable quadratic transformation index. The decoding device can select one of the k transformation kernels within the specific set based on the inseparable quadratic transformation index. For example, NSST index value 0 can point to the first inseparable quadratic transformation kernel, NSST index value 1 can point to the second inseparable quadratic transformation kernel, and NSST index value 2 can point to the third inseparable quadratic transformation kernel. Alternatively, NSST index value 0 can indicate that the first inseparable quadratic transformation is not applied to the target block, and NSST index values ​​1 to 3 can point to the three transformation kernels mentioned above.

[0105] Furthermore, referring to Figure 3, the transformation unit can perform the above-mentioned non-separable quadratic transformation based on the selected transformation kernel to obtain (quadratic) transformation coefficients. As mentioned above, these transformation coefficients can be derived as quantized transformation coefficients through the quantization unit, encoded and signaled to the decoding unit, and transmitted to the inverse quantization / inverse transformation unit within the encoding unit.

[0106] On the other hand, if the quadratic transformation is omitted, the (primary) transformation coefficients, which are the output of the above-mentioned primary (separated) transformation, can be derived as quantized transformation coefficients through the quantization unit, as described above, and can be encoded and signaled to the decoding unit and transmitted to the inverse quantization / inverse transformation unit within the encoding unit.

[0107] The inverse transform unit can perform a series of procedures in the reverse order of the procedures performed by the transform unit described above. The inverse transform unit can receive the (inversely quantized) transform coefficients, perform a quadratic (inverse) transform to derive (primary) transform coefficients (S350), and perform a primary (inverse) transform on the above (primary) transform coefficients to obtain residual blocks (residual samples) (S360). Here, the above primary transform coefficients can be called modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the predicted block, and a reconstructed picture can be generated based on this.

[0108] On the other hand, as mentioned above, if the quadratic (inverse) transformation is omitted, the (inversely quantized) transformation coefficients can be received and the above-mentioned linear (separated) transformation can be performed to obtain residual blocks (residual samples). As mentioned above, the encoding and decoding devices can generate reconstructed blocks based on the residual blocks and predicted blocks, and a reconstructed picture can be generated based on these.

[0109] Figures 5A to 5C illustrate a selective transform according to one embodiment of the present invention.

[0110] In this specification, “target block” may mean the current block or residual block on which coding is performed.

[0111] Figure 5A shows an example of deriving conversion coefficients through a transformation.

[0112] In video coding, transformation can be described as the process of transforming an input vector R based on a transformation matrix to generate a transformation coefficient vector C for the input vector R, as illustrated in Figure 5A. The input vector R can also represent a linear transformation coefficient. Alternatively, the input vector R can represent a residual vector, i.e., a residual sample. On the other hand, the transformation coefficient vector C can also be described as an output vector C.

[0113] Figure 5B shows a concrete example of deriving the transformation coefficients through the transformation. Figure 5B specifically shows the transformation process illustrated in Figure 5A. As mentioned above in Figure 3, in the non-separable quadratic transformation (hereinafter referred to as 'NSST'), the block data of the transformation coefficients obtained by applying the linear transformation is divided into MxM blocks, and then M is applied to each MxM block. 2 xM 2 NSST can be performed. M can be, for example, 4 or 8, but is not limited to these. The above M 2 It can be N. In this case, as shown in Figure 5B, the above input vector R is a (linear) transformation coefficient r1 to r N It can be a (1xN) dimensional vector containing, and the above transformation coefficient vector C is a transformation coefficient c1 to c N It can be an (Nx1) dimensional vector containing the following: That is, the above input vector R has N (linear) transformation coefficients r1 to r N It can include and the size of the above input vector R can be 1xN. Also, the above transformation coefficient vector C has N transformation coefficients c1 to c N It can include, and the size of the above transformation coefficient vector C can be Nx1.

[0114] To derive the conversion coefficient vector C, the input vector R can pass through (go through) the conversion matrix. That is, the input vector R can be derived as the conversion coefficient vector C based on the conversion matrix.

[0115] On the other hand, the conversion matrix can include N basis vectors B1 to B N As shown in FIG. 5B, the basis vectors B1 to B N can be (1xN)-dimensional vectors. That is, the size of the basis vectors B1 to B N can be 1xN. The conversion coefficient vector C can be generated based on the (primary) conversion coefficients of the input vector R and each of the basis vectors of the conversion matrix. For example, the inner product of the input vector and each basis vector can be derived as the conversion coefficient vector C.

[0116] On the other hand, two main (main) issues (issues) occur in the aforementioned conversion. Specifically, high computational complexity and memory requirements for storing the generated coefficients can occur as main issues in relation to the number of multiplications and additions required to generate the output vector.

[0117] For example, the computational complexity and memory requirements required for separable transform and non-separable transform can be derived as shown in the following table.

[0118]

Table 5

[0119] Referring to Table 5, the memory for storing the coefficients generated through the separable transform is N2 This can be requested, and the number of calculations is 2N. 3 This is possible. The number of calculations mentioned above indicates the computational complexity. Furthermore, the memory required to store the coefficients generated through the non-separable transformation may be N4, and the number of calculations mentioned above may be N4. The number of calculations mentioned above indicates the computational complexity. In other words, a larger number of calculations may indicate higher computational complexity, and a smaller number of calculations may indicate lower computational complexity.

[0120] As illustrated in Table 5 above, the memory requirements and number of calculations for the non-separable transformation can be increased significantly compared to the separable transformation. Furthermore, as the size of the target block on which the non-separable transformation is performed increases, that is, as N increases, the discrepancy between the memory requirements and number of calculations for the separable transformation and the memory requirements and number of calculations for the non-separable transformation can increase.

[0121] The above non-separated transformation provides better coding gain compared to the separated transformation. However, as shown in Table 5 above, due to the computational complexity of the non-separated transformation, it is not used in existing video coding standards. Furthermore, since the computational complexity of the separated transformation also increases with the size of the target block, the existing HEVC standard restricts the use of the separated transformation to target blocks that are 32x32 or smaller.

[0122] Herein, the present invention proposes a selective transformation. This selective transformation can significantly reduce computational complexity and memory requirements, thereby increasing the efficiency of computationally intensive transformation blocks and improving coding efficiency. In other words, this selective transformation can be used to resolve the problem of computational complexity that arises when transforming large blocks or during non-separable transformations. This selective transformation can be used for any type of transformation, such as a linear transformation (or core transformation) or a quadratic transformation. For example, this selective transformation can be applied to the core transformation for an encoding / decoding device, resulting in a significant reduction in encoding / decoding time.

[0123] Figure 5C shows an example of deriving transformation coefficients through the selective transformation described above. The selective transformation can be defined as a transformation performed on a target block based on a transformation matrix containing basis vectors with a selective number of elements.

[0124] The simplified transformation is a method proposed with the motivation that among the N elements of the basis vectors of the transformation matrix, there may be duplicate or insignificant elements, and that excluding these elements can reduce computational complexity and memory requirements. For example, referring to Figure 5C, among the N elements of basis vector B1, Z0 elements may be insignificant, in which case a truncated basis vector B1 containing only N1 elements can be derived. Here, N1 can be N-Z0. The truncated basis vector B1 can be referred to as the modified basis vector B1.

[0125] Referring to Figure 5C, when the modified basis vector B1 is applied to the input vector R as part of the transformation matrix, the transformation coefficient C1 can be derived. From the experimental results, it is observed that the transformation coefficient C1 is the same value as the transformation coefficient C1 derived when the existing basis vector B1 is applied to the input vector R as part of the transformation matrix. That is, assuming that a small number of elements in each basis vector are 0 to derive the result means that the number of multiplications required can be greatly reduced without a large difference in the result. Furthermore, the number of elements that must be stored for this operation (i.e., elements of the transformation matrix) can be reduced.

[0126] To define the positions of non-important (or non-meaningful) elements of a basis vector, association vectors are proposed. To derive the above N1-dimensional modified basis vector B1, a (1xN)-dimensional association vector A1 can be considered. That is, to derive a 1xN1-sized modified basis vector B1 (i.e., a modified basis vector B1 containing N1 elements), a 1xN-sized association vector A1 can be considered.

[0127] Referring to Figure 5C, the above input vector R is connected to the above basis vectors B1 to B N The values ​​derived by applying the associated vector to each of the elements can be transmitted to the basis vector. Through this, only some elements of the input vector R can be computed as elements of the basis vector. Specifically, the associated vector can contain 0 and 1, and the selected elements of the input vector R are multiplied by 1, and the unselected elements are multiplied by 0, so that only the selected elements pass through and are transmitted to the basis vector.

[0128] For example, the above-mentioned related vector A1 can be applied to the above-mentioned input vector R, and only N1 elements of the above-mentioned input vector R specified by the above-mentioned related vector A1 can be used to calculate the inner product with the above-mentioned basis vector B1. The above inner product can represent C1 of the above-mentioned transformation coefficient vector C. Here, the above-mentioned basis vector B1 can contain N1 elements, and the above-mentioned N1 can be less than or equal to N. The above-mentioned operation between the input vector R, related vector A1 and basis vector B1 is performed using related vectors A2 to A N and basis vectors B2 to B N This can also be done against them.

[0129] Since the above association vector contains only 0 and / or 1 and binary values, there may be advantages in storing the above association vector. A 0 in the above association vector can indicate that the element of the input vector R corresponding to the 0 is not transmitted to the transformation matrix for the dot product calculation, and a 1 in the above association vector can indicate that the element of the input vector R corresponding to the 1 is transmitted to the transformation matrix for the dot product calculation. For example, an association vector A of size 1xN. k is, A k1 Or A kN It can include the above related vector A. k A kn If 0, then the input vector R's r n It may not be able to pass through. That is, the r of the above input vector R n This is the above transformation vector B n It may not be transmitted. Also, as mentioned above, A kn If r is 1, then the input vector R above is r n It can pass through. That is, the r of the above input vector R n This is the above transformation vector B n It can be transmitted to the above conversion coefficient vector C of c n It can be used in calculations to derive [the result].

[0130] Figure 6 schematically shows a multiple transformation technique in which the above selective transformation is applied to a quadratic transformation.

[0131] Referring to Figure 6, the conversion unit can correspond to the conversion unit in the encoding device shown in Figure 1, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device shown in Figure 1 or the inverse conversion unit in the decoding device shown in Figure 2.

[0132] The conversion unit can perform a linear transformation based on the residual samples (residual sample array) within the residual block to derive (linear) transformation coefficients (S610). Here, the above linear transformation may include the AMT described above.

[0133] When the above adaptive multicore conversion is applied, a spatial-to-frequency domain conversion is applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or first-order conversion coefficients). Here, the above first-order conversion coefficients can be called primary conversion coefficients from the perspective of the conversion unit. Also, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. can be called conversion types, conversion kernels, or conversion cores. For reference, the above DCT / DST conversion types can be defined based on basis functions, and the above basis functions can be shown as in Table 1 above. Specifically, the process of deriving the above first-order conversion coefficients by applying the above adaptive multicore conversion is as described above.

[0134] The transformation unit can perform a selective transformation based on the above (linear) transformation coefficients to derive (quadratic) transformation coefficients (S620). The selective transformation can mean a transformation performed on the above (linear) transformation coefficients for the target block based on a transformation matrix containing the modified basis vectors and an association matrix containing the associated vectors for the above basis vectors. The above modified basis vectors can represent basis vectors containing N or fewer elements. That is, the above modified basis vectors can represent basis vectors containing a specific number of elements selected from N elements. For example, modified basis vector B n is, (1xN n It can be a )-dimensional vector, and the above N n This can be less than or equal to the above N. That is, the modified basis vector B n The size is (1xN n ) may be the size, as per the above N n This can be less than or equal to the above N, where N can be the product of the height and width of the upper left target region of the target block to which the selected transformation is applied. Alternatively, N can be the total number of transformation coefficients of the upper left target region of the target block to which the selected transformation is applied. On the other hand, the transformation matrix containing the modified basis vectors can be shown as a modified transformation matrix. The transformation matrix can also be shown as a Transform Bases Block (TBB), and the association matrix can be shown as an Association Vectors Block (AVB).

[0135] The transformation unit can perform the selective transformation based on the modified transformation matrix and the associated matrix to obtain (quadratic) transformation coefficients. These transformation coefficients can be derived as quantized transformation coefficients through the quantization unit, as described above, and can be encoded and transmitted to the decoder for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.

[0136] The inverse transform unit can perform a series of procedures in the reverse order of the procedures performed by the transform unit described above. The inverse transform unit can receive the (inversely quantized) transform coefficients, perform a selective (inverse) transform to derive (first-order) transform coefficients (S650), and perform a first-order (inverse) transform on the above (first-order) transform coefficients to obtain residual blocks (residual samples) (S660). Here, the above first-order transform coefficients can be called modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the predicted block, and a reconstructed picture can be generated based on this.

[0137] On the other hand, the present invention proposes a selective transformation combined with a simplified transformation in one embodiment of the selective transformation described above.

[0138] In this specification, “simplification transformation” can mean a transformation performed on the residual sample for the target block based on a transform matrix whose size has been reduced by a simplification factor.

[0139] In one embodiment of the simplified transformation, an N-dimensional vector can be mapped to an R-dimensional vector located in another space to determine a simplified transformation matrix, where R is less than N. That is, the simplified transformation can be said to mean a transformation performed on residual samples for a target block based on a reduced transformation matrix containing R basis vectors. Here, N can mean the square of the length of one side of the block (or target region) to which the transformation is applied, or the total number of transformation coefficients corresponding to the block (or target region) to which the transformation is applied, and the simplified factor can mean the R / N value. The simplified factor can be referred to by various terms such as reduced factor, reduction factor, simplified factor, and simple factor. On the other hand, R can be called the reduced coefficient, but in some cases the simplified factor can mean R. Also, in some cases the simplified factor can mean the N / R value.

[0140] The size of the simplified transformation matrix according to one embodiment is RxN, which is smaller than the size NxN of a normal transformation matrix, and can be defined as shown in the following equation 4.

[0141] <Formula 4>

number

[0142] The simplified transformation matrix T is applied to the transformation coefficients to which the linear transformation of the target block has been applied. RxN When multiplied by this, the (quadratic) transformation coefficient for the above target block can be derived.

[0143] When the above RST is applied, a simplified transformation matrix of size RxN is applied to the quadratic transformation, so the transformation coefficients from R+1 to N can implicitly be 0. In other words, if the transformation coefficients of the target block are derived by applying the above RST, the values ​​of the transformation coefficients from R+1 to N can be 0. Here, the transformation coefficients from R+1 to N can represent the transformation coefficients from the R+1th to the Nth among the transformation coefficients. Specifically, the array of transformation coefficients of the target block can be described as follows.

[0144] Figure 7 illustrates the arrangement of transformation coefficients based on a target block according to an embodiment of the present invention. The explanation of the transformation described later in Figure 7 can also be applied to the inverse transformation. A linear transformation and an NSST based on a simplified transformation can be performed on the target block (or residual block) 700. In one example, the 16x16 block shown in Figure 7 represents the target block 700, and the 4x4 blocks denoted A to P represent subgroups of the target block 700. The linear transformation can be performed over the entire range of the target block 700, and after the linear transformation is performed, the NSST can be applied to the 8x8 block (hereinafter, the upper left target region) composed of subgroups A, B, E, and F. In this case, if an NSST based on a simplified transformation is performed, only R NSST transformation coefficients (where R is the simplified coefficient, and R is less than N) are derived, so the NSST transformation coefficients in the range from the R+1th to the Nth can each be determined to be 0. For example, if R is 16, the 16 transformation coefficients derived by NSST based on the simplified transformation can be assigned to each block in subgroup A, which is the upper left 4x4 block contained in the upper left target area of ​​the target block 700, and a transformation coefficient of 0 can be assigned to each of the NR blocks, i.e., 64-16=48 blocks, contained in subgroups B, E, and F. The linear transformation coefficients for which NSST based on the simplified transformation is not performed can be assigned to each block contained in subgroups C, D, G, H, I, J, K, L, M, N, O, and P.

[0145] Figure 8 shows an example of deriving transformation coefficients through a transformation that combines the simplified transformation and the selective transformation. Referring to Figure 8, the transformation matrix can contain R basis vectors, and the association matrix can contain R association vectors. Here, the transformation matrix containing R basis vectors can be denoted as a reduced transform matrix, and the association matrix containing R association vectors can be denoted as a reduced association matrix.

[0146] Furthermore, each of the above basis vectors can contain only selected elements from among the N elements. For example, referring to Figure 8, basis vector B1 can be a 1xN1-dimensional vector containing N1 elements, basis vector B2 can be a 1xN2-dimensional vector containing N2 elements, and basis vector B R is, N R 1xN containing a number of elements R It can be a dimensional vector. N1, N2, and N above R This can be a value less than or equal to the above N. The transformation obtained by combining the above simplified transformation and the above selective transformation can be used for any type of transformation, such as quadratic and linear transformations.

[0147] Referring to Figure 8, the encoding and decoding devices can apply a transformation that combines the simplified transformation and the selective transformation to a quadratic transformation. For example, the selective transformation can be performed based on the simplified transformation matrix and simplified association matrix containing the modified basis vectors, and (quadratic) transformation coefficients can be obtained. In addition, to reduce the computational complexity of the selective transformation, a HyGT (Hypercube-Givens Transform) or the like can be used for the calculation of the selective transformation.

[0148] Figure 9 shows an example of deriving transformation coefficients through the selective transformation described above. In one embodiment of the selective transformation described above, the association vectors A1, A2, ...A of the association matrix are shown. NThere may be no patterns for this, and the association vectors may be derived in different forms from each other. Alternatively, in other embodiments of the selective transformation, the association vectors A1, A2, ...A of the association matrix may be N These can be derived as identical forms.

[0149] Specifically, for example, the above association vector can contain the same number of 1s. For example, if the number of 1s is M, the above association vector can contain M 1s and NM 0s. In this case, M of the (linear) transformation coefficients of the above input vector R can be transmitted to the basis vector. Therefore, the length of the above basis vector can also be M. That is, as illustrated in Figure 9, the above basis vector can contain M elements and can be a (1xM) dimensional vector, and N1=N2=...=N N This can be derived as =M. The related matrix and the modified transformation matrix structure shown in Figure 9 can be described as a symmetric architecture for the above selective transformation.

[0150] Furthermore, in other examples, there may be specific patterns of elements with a value of 1, and these patterns can be iterated, rotated, and / or translated in any manner to derive the associated vectors.

[0151] On the other hand, the aforementioned selective transformation can be applied not only to the simplified transformation and / or HyGT described above, but also to other transformation techniques.

[0152] Furthermore, the present invention proposes a method for simplifying the association vector in the selective transformation described above. By simplifying the association vector, the storage of information for performing the selective transformation and the handling of the selective transformation can be improved. In other words, the memory load for performing the selective transformation can be reduced, and the handling capability of the selective transformation can be improved.

[0153] The effect of simplifying the association vector described above becomes even more apparent when the non-zero elements in the association vector have a continuous distribution. For example, association vector A k This can include a continuous string of ones. In this case, A above k This involves two factors A ks and A kL It can be expressed through the above factor A. ks and A kL Based on the above related vector A k The following can be derived. Here, the above A ks This can be a factor indicating the starting point of a non-zero element (e.g., 1), as described above. kL This can be a factor indicating the length of a non-zero element. The above related vector A is shown based on the above factor. k This can be derived as shown in the following table.

[0154] [Table 6]

[0155] Referring to Table 6, the above related vector A k It can contain 16 elements. That is, the above related vector A k This can be a 1x16-dimensional vector. The related vector A above. k Factor A indicates the starting point of non-zero elements. ks The value of can be derived as 0, in which case the above factor A ks The starting point of the non-zero element is the associated vector A. k It can be pointed to as the first element of the above related vector A. k Factor A indicates the length of the non-zero elements. kL The value of can be derived as 8, in which case the above factor A kLThis can be interpreted as having a length of 8 non-zero elements. Therefore, the above A k Based on the above factors, it can be derived as a vector in which the first to eighth elements are 1 and the remaining elements are 0, as shown in Table 6.

[0156] Figure 10 shows an example of deriving the above-mentioned association vectors based on two factors for the association vectors and performing a selective transformation. Referring to Figure 10, an association matrix can be derived based on the factors for each association vector, and the transformation coefficients for the target block can be derived based on the above-mentioned association matrix and the above-mentioned modified transformation matrix.

[0157] On the other hand, the starting point and number of non-zero elements of each non-zero element in the association vector can be derived as fixed values, or the starting point and number of non-zero elements of each non-zero element in the association vector can be derived in various ways.

[0158] Furthermore, for example, the starting point and number of non-zero elements of the related vector can be derived based on the size of the upper-left target region where the transformation is performed. Here, the size of the upper-left target region can represent the number of transformation coefficients of the upper-left target region, or the product of the height and width of the upper-left target region. In another example, the starting point and number of non-zero elements of the related vector can also be derived based on the intra-prediction mode of the target block. Specifically, for example, the starting point and number of non-zero elements of the related vector can be derived based on whether the intra-prediction mode of the target block is a non-directional intra-prediction mode.

[0159] Alternatively, for example, the starting point and number of non-zero elements in the association vector can be predetermined. Alternatively, for example, information indicating the starting point and number of non-zero elements in the association vector can be signaled, and the association vector can be derived based on this information. Alternatively, other information can be used instead of the information indicating the starting point of the non-zero elements. For example, information indicating the last position of the non-zero elements can be used instead of the information indicating the starting point of the non-zero elements, and the association vector can be derived based on this information.

[0160] On the other hand, the method for deriving the above-mentioned related vector based on the above-mentioned factor can be applied to both separable transformations and non-separable transformations such as the simplified transformation HyGT.

[0161] Figure 11 schematically shows a video encoding method using an encoding device according to the present invention. The method disclosed in Figure 11 can be performed using the encoding device disclosed in Figure 1. Specifically, for example, S1100 in Figure 11 can be performed by the subtraction unit of the encoding device, S1110 by the conversion unit of the encoding device, and S1120 by the entropy encoding unit of the encoding device. Although not shown, the process of deriving prediction samples can be performed by the prediction unit of the encoding device.

[0162] The encoding device derives residual samples for the target block (S1100). For example, the encoding device can decide whether to perform interpretation or intrapretation for the target block, and can determine a specific interpretation mode or specific intrapretation mode based on the RD cost. Depending on the determined mode, the encoding device can derive predicted samples for the target block, and can derive residual samples by adding the original samples and predicted samples for the target block.

[0163] The encoding device derives the transformation coefficients of the target block based on a selective transform applied to the residual samples (S1110). The selective transform can be performed based on a modified transform matrix, which is a matrix containing modified basis vectors, and which can contain a specific number of elements selected from N elements. The selective transform can also be performed on the upper left target region of the target block, where N is the number of residual samples located in the upper left target region. Alternatively, N is the product of the width and height of the upper left target region. For example, N can be 16 or 64.

[0164] The encoding device can perform a core transformation on the residual sample to derive corrected transformation coefficients, and based on the corrected transformation matrix and the association matrix which includes association vectors for the corrected basis vectors, it can perform the selective transformation on the corrected transformation coefficients located in the upper left target region of the target block to derive the transformation coefficients of the target block.

[0165] Specifically, the core transformation for the residual sample can be performed as follows: The encoding device can decide whether or not to apply an Adaptive Multiple Core Transform (AMT) to the target block. In this case, an AMT flag can be generated to indicate whether or not the Adaptive Multiple Core Transform is applied to the target block. If the AMT is not applied to the target block, the encoding device can derive DCT type 2 as the transformation kernel for the target block, and perform the transformation on the residual sample based on DCT type 2 to derive the modified transformation coefficients.

[0166] When the above AMT is applied to the target block, the encoding device can configure a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, derive the horizontal conversion kernel and the vertical conversion kernel based on the conversion subset, and perform a conversion on the residual samples based on the horizontal conversion kernel and the vertical conversion kernel to derive the corrected conversion coefficients. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel may include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Furthermore, conversion index information can be generated, and this conversion index information may include an AMT horizontal flag pointing to the horizontal conversion kernel and an AMT vertical flag pointing to the vertical conversion kernel. On the other hand, the conversion kernel may be referred to as a conversion type or conversion core.

[0167] If the above-mentioned modified transformation coefficients are derived, the encoding device can perform the above-mentioned selective transformation on the modified transformation coefficients located in the upper left target region of the target block, based on the above-mentioned modified transformation matrix and the association matrix which includes the association vectors for the above-mentioned modified basis vectors, and derive the above-mentioned transformation coefficients of the target block. Modified transformation coefficients other than the above-mentioned modified transformation coefficients located in the upper left region of the target block can be derived as the above-mentioned transformation coefficients of the target block as they are.

[0168] Specifically, among the modified transformation coefficients located in the upper left target area, the modified transformation coefficients for the elements of the associated vector that are 1 can be derived, and the transformation coefficients of the target block can be derived based on the derived modified transformation coefficients and the modified basis vector. Here, the associated vector for the modified basis vector can contain N elements, and the N elements can contain elements of 1 and / or elements of 0, and the number of elements of 1 can be A. Also, the modified basis vector can contain A elements.

[0169] On the other hand, in one example, the modified transformation matrix may contain N modified basis vectors, and the association matrix may contain N association vectors. The association vectors may contain the same number of elements of 1, and all the modified basis vectors may contain the same number of elements. Alternatively, the association vectors may not contain the same number of elements of 1, and all the modified basis vectors may not contain the same number of elements.

[0170] Alternatively, in other examples, the modified transformation matrix may contain R modified basis vectors, and the association matrix may contain R association vectors. R may be a reduced coefficient, and R may be less than N. The association vectors may contain the same number of elements of 1, and all modified basis vectors may contain the same number of elements. Alternatively, the association vectors may not contain the same number of elements of 1, and all modified basis vectors may not contain the same number of elements.

[0171] On the other hand, the association vector can be configured to have consecutive elements of 1. In this case, for example, information about the association vector can be entropi-encoded. For example, the information about the association vector may include information indicating the starting point of the elements of 1 and information indicating the number of elements of 1. Alternatively, for example, the information about the association vector may include information indicating the last position of the elements of 1 and information indicating the number of elements of 1.

[0172] Furthermore, in other examples, the associated vector can be derived based on the size of the upper left target region. For example, the starting point of the 1 element and the number of 1 elements of the associated vector can be derived based on the size of the upper left target region.

[0173] Alternatively, in other examples, the association vector can be derived based on the intra-prediction mode of the target block. For example, the starting point of one element and the number of elements of the association vector can be derived based on the intra-prediction mode. Also, for example, the starting point of one element and the number of elements of the association vector can be derived based on whether the intra-prediction mode is a non-directional intra-prediction mode.

[0174] The encoding device encodes information regarding the conversion coefficients (S1330). The information regarding the conversion coefficients may include information regarding the size and position of the conversion coefficients. Furthermore, as mentioned above, the information regarding the association vector can be entropi-encoded. For example, the information regarding the association vector may include information indicating the starting point of the 1 element and information indicating the number of 1 elements. Alternatively, for example, the information regarding the association vector may include information indicating the last position of the 1 element and information indicating the number of 1 elements.

[0175] The video information, including the information regarding the above-mentioned conversion coefficients and / or the information regarding the above-mentioned related vectors, can be output in bitstream format. The video information may further include prediction information. This prediction information may include, as information related to the prediction procedure, prediction mode information and motion information (for example, if interpretation is applied).

[0176] The output bitstream can be transmitted to a decoding device via a storage medium or network.

[0177] Figure 12 schematically shows an encoding apparatus that performs a video encoding method according to the present invention. The method disclosed in Figure 11 can be performed by the encoding apparatus disclosed in Figure 12. Specifically, for example, the addition unit of the encoding apparatus in Figure 12 can perform S1100 in Figure 11, the conversion unit of the encoding apparatus can perform S1110, and the entropy encoding unit of the encoding apparatus can perform S1120 to S1130. Although not shown, the process of deriving prediction samples can be performed by the prediction unit of the encoding apparatus.

[0178] Figure 13 schematically shows a video decoding method using a decoding device according to the present invention. The method disclosed in Figure 13 can be performed using the decoding device disclosed in Figure 2. Specifically, for example, steps S1300 to S1310 in Figure 13 can be performed by the entropy decoding unit of the decoding device, S1320 by the inverse transformation unit of the decoding device, and S1330 by the addition unit of the decoding device. Although not shown, the process of deriving the predicted sample can be performed by the prediction unit of the decoding device.

[0179] The decoding device derives the conversion coefficients of the target block from the bitstream (S1300). The decoding device can derive the conversion coefficients of the target block by decoding the information regarding the conversion coefficients of the target block received through the bitstream. The received information regarding the conversion coefficients of the target block can be shown as residual information.

[0180] The decoding device derives residual samples for the target block based on a selective transform applied to the transformation coefficients (S1310). The selective transform can be performed based on a modified transform matrix, which is a matrix containing modified basis vectors, and which can contain a specific number of elements selected from N elements. The selective transform can also be performed on transformation coefficients located in the upper left target region of the target block, where N is the number of transformation coefficients located in the upper left target region. Alternatively, N may be the product of the width and height of the upper left target region. For example, N may be 16 or 64.

[0181] The decoding device can derive the modified transformation coefficients by performing the selective transformation on the transformation coefficients located in the upper left target region of the target block, based on the modified transformation matrix and the association matrix which includes the association vectors for the modified basis vectors.

[0182] Specifically, the transformation coefficients located in the upper left target region above can be derived for the elements of the associated vector that are set to 1. Based on the derived transformation coefficients and the modified basis vector, the modified transformation coefficients can be derived. Here, the associated vector for the modified basis vector can contain N elements, and these N elements can contain elements that are set to 1 and / or elements that are set to 0. The number of elements that are set to 1 can be A. The modified basis vector can also contain A elements.

[0183] On the other hand, in one example, the modified transformation matrix may contain N modified basis vectors, and the association matrix may contain N association vectors. The association vectors may each contain the same number of elements of 1, and all the modified basis vectors may each contain the same number of elements. Alternatively, the association vectors may not each contain the same number of elements of 1, and all the modified basis vectors may not each contain the same number of elements.

[0184] Alternatively, in other examples, the modified transformation matrix may contain R modified basis vectors, and the association matrix may contain R association vectors. R may be a reduced coefficient, and R may be less than N. The association vectors may contain the same number of elements of 1, and all modified basis vectors may contain the same number of elements. Alternatively, the association vectors may not contain the same number of elements of 1, and all modified basis vectors may not contain the same number of elements.

[0185] On the other hand, the association vector can be configured to have consecutive elements of 1. In this case, for example, information about the association vector can be obtained from the bitstream, and the association vector can be derived based on this information. For example, the information about the association vector may include information indicating the starting point of the elements of 1 and information indicating the number of elements of 1. Alternatively, for example, the information about the association vector may include information indicating the last position of the elements of 1 and information indicating the number of elements of 1.

[0186] Alternatively, in other examples, the association vector can be derived based on the size of the upper left target region. For example, the starting point of the 1 element and the number of 1 elements of the association vector can be derived based on the size of the upper left target region.

[0187] Alternatively, in other examples, the association vector can be derived based on the intra-prediction mode of the target block. For example, the starting point of one element and the number of elements of the association vector can be derived based on the intra-prediction mode. Also, for example, the starting point of one element and the number of elements of the association vector can be derived based on whether the intra-prediction mode is a non-directional intra-prediction mode.

[0188] Once the above-mentioned modified conversion coefficients are derived, the decoding device can perform a core transformation on the target block including the above-mentioned modified conversion coefficients to derive the above-mentioned residual sample.

[0189] The core transformation for the target block can be performed as follows: The decoding device can obtain an AMT flag from the bitstream indicating whether or not an Adaptive Multiple Core Transform (AMT) is applied. If the value of the AMT flag is 0, the decoding device can derive DCT type 2 as the transformation kernel for the target block. Based on DCT type 2, the decoding device can perform an inverse transformation on the target block including the modified transformation coefficients to derive the residual samples.

[0190] If the value of the above AMT flag is 1, the decoding device can configure a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, and can derive the horizontal conversion kernel and the vertical conversion kernel based on the conversion index information obtained from the bitstream and the above conversion subset, and can derive the residual sample by performing an inverse conversion on the target block including the modified conversion coefficients based on the above horizontal conversion kernel and the above vertical conversion kernel. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel may include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Furthermore, the conversion index information may include an AMT horizontal flag pointing to one of the candidates included in the conversion subset for the horizontal conversion kernel and an AMT vertical flag pointing to one of the candidates included in the conversion subset for the vertical conversion kernel. On the other hand, the above conversion kernel may be referred to as a conversion type or conversion core.

[0191] The decoding device generates a reconstructed picture based on the residual samples (S1320). The decoding device can generate a reconstructed picture based on the residual samples. For example, the decoding device can perform inter-prediction or intra-prediction for the target block based on prediction information received through the bitstream, derive prediction samples, and generate the reconstructed picture by adding the prediction samples and the residual samples. As previously mentioned, thereafter, if necessary, in-loop filtering procedures such as deblock filtering, SAO and / or ALF procedures can be applied to the reconstructed picture to improve subjective / objective image quality.

[0192] Figure 14 schematically shows a decoding device that performs the video decoding method according to the present invention. The method disclosed in Figure 13 can be performed by the decoding device disclosed in Figure 14. Specifically, for example, the entropy decoding unit of the decoding device in Figure 14 can perform S1300 in Figure 13, the inverse conversion unit of the decoding device in Figure 14 can perform S1310 in Figure 13, and the addition unit of the decoding device in Figure 16 can perform S1320 in Figure 15. In addition, although not shown, the process of deriving the predicted sample can be performed by the prediction unit of the decoding device in Figure 14.

[0193] According to the present invention described above, the amount of data that must be transferred for residual processing can be reduced through efficient conversion, thereby improving residual coding efficiency.

[0194] Furthermore, according to the present invention, an inseparable transformation can be performed based on a transformation matrix composed of basis vectors containing a selected specific number of elements, thereby reducing the memory load and computational complexity for the inseparable transformation and improving residual coding efficiency.

[0195] Furthermore, according to the present invention, non-separable transformations can be performed based on a simplified transformation matrix, thereby reducing the amount of data that must be transferred for residual processing and improving residual coding efficiency.

[0196] In the embodiments described above, the method is explained based on a sequence of steps or blocks (flowchart), but the present invention is not limited to the order of the steps, and some steps may occur in different steps and in different orders, or simultaneously, than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the present invention.

[0197] The method according to the present invention described above can be embodied in software form, and the encoding and / or decoding device according to the present invention can be included in, for example, video processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0198] In the present invention, when embodiments are embodied in software, the methods described above can be embodied in modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by a variety of well-known means. The processor may include an ASIC (Application-Specific Integrated Circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (Read-Only Memory), RAM (Random Access Memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in the present invention can be embodied and performed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing can be embodied and performed on a computer, processor, microprocessor, controller, or chip.

[0199] Furthermore, the decoding and encoding devices to which the present invention applies can include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, OTT video (Over the Top video) devices, internet streaming service providers, 3D video devices, image-phone video devices, and medical video devices, and can be used to process video signals or data signals. For example, OTT video (Over The Top video) devices can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), and the like.

[0200] Furthermore, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and stored on a computer-readable recording medium. Multimedia data having a data structure according to the present invention can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), Universal Serial Bus (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wireless communication network. Embodiments of the present invention can also be embodied in a computer program product in the form of program code, and the program code can be executed by a computer according to embodiments of the present invention. The program code can be stored on a computer-readable carrier.

[0201] Furthermore, a content streaming system to which the present invention applies may broadly include an encoding server, a streaming server, a web server, a media storage unit, a user device, and a multimedia input device.

[0202] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transferring it to the streaming server. In other cases, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by the encoding method or bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transferring or receiving it.

[0203] The streaming server transfers multimedia data to the user's device based on user requests via the web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, which then transfers the multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server is responsible for controlling the commands and responses between the devices within the content streaming system.

[0204] The above streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the above encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the above streaming server can store the above bitstream for a certain period of time.

[0205] Examples of user devices mentioned above include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation systems, slate PCs, tablet PCs, ultrabooks (ULTRABOOK®), wearable devices (e.g., smartwatches, smart glasses, HMDs (Head Mounted Displays)), digital TVs, desktop computers, and digital signage. Each server within the above content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

Claims

1. A video decoding method performed by a decoding device, The steps include: deriving the conversion coefficients of the target block from the bitstream, The steps include: deriving a residual sample for the target block based on the non-separable transformation of the transformation coefficients; The step of generating a reconstructed picture based on the residual sample for the target block and the predicted sample for the target block, The aforementioned non-separable transformation is performed based on the transformation matrix, The transformation matrix includes the modified basis vectors, Each of the aforementioned modified basis vectors contains fewer than N elements, The aforementioned N is equal to the number of transformation coefficients located in the region to which the non-separable transformation is applied in the target block, The region to which the non-separable transformation is applied is the 8x8 upper left region within the target block, where N is equal to 64. The number of the modified basis vectors is less than N, according to the method.

2. A video encoding method performed by an encoding device, The steps include: deriving the residual sample of the target block, The steps include: deriving the conversion coefficient of the target block based on the non-separated conversion of the residual sample; The step of encoding information regarding the conversion coefficients, The aforementioned non-separable transformation is performed based on the transformation matrix, The transformation matrix includes the modified basis vectors, Each of the aforementioned modified basis vectors contains fewer than N elements, N is equal to the number of transformation coefficients located in the region to which the non-separable transformation is applied in the target block. The region to which the non-separable transformation is applied is the 8x8 upper left region within the target block, where N is equal to 64. The number of the modified basis vectors is less than N, according to the method.

3. A method for transmitting data including a bitstream relating to video, A step of obtaining the bitstream relating to the video, wherein the bitstream is The steps include: deriving the residual sample of the target block, The steps include: deriving the conversion coefficient of the target block based on the non-separated conversion of the residual sample; The steps include: encoding information regarding the conversion coefficients to generate the bitstream, and generating the bitstream by performing these steps, The step of transmitting the data, which includes the bitstream, The aforementioned non-separable transformation is performed based on the transformation matrix, The transformation matrix includes the modified basis vectors, Each of the aforementioned modified basis vectors contains fewer than N elements, N is equal to the number of transformation coefficients located in the region to which the non-separable transformation is applied in the target block. The region to which the non-separable transformation is applied is the 8x8 upper left region within the target block, where N is equal to 64. The number of the modified basis vectors is less than N, according to the method.

Citation Information

Patent Citations

  • Method and device for the transformation and method and device for the reverse transformation of images

    US20130195177A1

  • Reduced size inverse transform for decoding and encoding

    US20170034530A1

  • Non-separable secondary transform for video coding

    US20170094313A1

  • Primary transform and secondary transform in video coding

    US20180103252A1