Video coding method and apparatus based on selective conversion
The video decoding method employs a modified transform matrix with selectively chosen basis vectors to enhance coding efficiency and reduce data transfer and computational complexity, addressing the high-cost challenge of high-resolution video transmission and storage.
Patent Information
- Application Number
- JP2024177601
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-21
- Filing Date
- 2024-10-10
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2038-12-21
AI Technical Summary
The increasing demand for high-resolution and high-quality videos has led to a surge in video data transmission and storage costs due to the higher amount of information required, necessitating improved video coding efficiency and residual coding efficiency.
A video decoding method and apparatus that utilize a modified transform matrix with selectively chosen basis vectors to perform efficient selective transforms, reducing data transfer and computational complexity.
This approach enhances residual coding efficiency by reducing data transfer and computational load through non-separable transforms with a simplified structure, improving overall video coding efficiency.
Smart Images

Figure 0007715904000015 
Figure 0007715904000016 
Figure 0007715904000017
Abstract
Description
Technical Field
[0001] The present invention relates to video coding technology, and more particularly, to a video decoding method and apparatus that follow selective conversion in a video coding system.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] Accordingly, a highly efficient video compression technology is required to effectively transmit, store, and reproduce high-resolution and high-quality video information.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technical problem of the present invention is to provide a method and apparatus for increasing video coding efficiency.
[0005] Another technical problem of the present invention is to provide a method and apparatus for increasing conversion efficiency.
[0006]
[0007] Still another technical problem of the present invention is to provide a method and apparatus for increasing the efficiency of residual coding through conversion.
Means for Solving the Problem
[0008] According to an embodiment of the present invention, a video decoding method performed by a decoding device is provided. The method includes: deriving a transform coefficient of a target block from a bitstream; deriving a residual sample for the target block based on a selective transform on the transform coefficient; and generating a restored picture based on the residual sample for the target block and a prediction sample for the target block. The selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix having modified basis vectors, and the modified basis vectors have a selected specific number of elements out of N elements.
[0009] According to another embodiment of the present invention, a decoding device for performing video decoding is provided. The decoding device includes: an entropy decoding unit that derives a transform coefficient of a target block from a bitstream; an inverse transform unit that derives a residual sample for the target block based on a selective transform on the transform coefficient; and an addition unit that generates a restored picture based on the residual sample for the target block and a prediction sample for the target block. The selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix having modified basis vectors, and the modified basis vectors have a selected specific number of elements out of N elements.
[0010] According to still another embodiment of the present invention, there is provided a video encoding method performed by an encoding apparatus. The method includes: deriving a residual sample of a target block; deriving a transform coefficient of the target block based on a selective transform on the residual sample; and encoding information regarding the transform coefficient. The selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix having modified basis vectors, and the modified basis vectors have a selected specific number of elements out of N elements.
[0011] According to still another embodiment of the present invention, there is provided a video encoding apparatus. The encoding apparatus includes: an addition unit that derives a residual sample of a target block; a transform unit that derives a transform coefficient of the target block based on a selective transform on the residual sample; and an entropy encoding unit that encodes information regarding the transform coefficient. The selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix having modified basis vectors, and the modified basis vectors have a selected specific number of elements out of N elements.
Advantages of the Invention
[0012] According to the present invention, it is possible to reduce the amount of data that must be transferred for residual processing through efficient transformation, and to improve the residual coding efficiency.
[0013] According to the present invention, it is possible to perform a non-separable transform based on a transform matrix composed of basis vectors including a selected specific number of elements, thereby reducing the memory load (load) and computational complexity for the non-separable transform, and improving the residual coding efficiency.
[0014] According to the present invention, non-separable conversion can be performed based on a conversion matrix with a simplified structure, thereby reducing the amount of data that must be transferred (sent) for residual processing and enhancing residual coding efficiency.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0016] The present invention can be subjected to various modifications and can have various embodiments. Specific embodiments are illustrated in the drawings and described in detail. However, this does not limit the present invention to specific embodiments. The terms used in this specification are merely used to explain specific embodiments and are not intended to limit the technical idea of the present invention. Singular expressions include plural expressions unless the context clearly indicates a different meaning. In this specification, terms such as "including" or "having" are used to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should not be understood that the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is precluded in advance.
[0017] On the other hand, each configuration in the drawings described in the present invention is independently illustrated for the convenience of explaining different characteristic functions from each other, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, two or more of the configurations may be combined to form one configuration, or one configuration may be divided into a plurality of configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0018] Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present invention will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.
[0019] On the other hand, the present invention relates to video / video coding. For example, the methods / embodiments disclosed in the present invention can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard or next-generation video / image coding.
[0020] In this specification, a "picture" generally means a unit indicating one video in a specific time period, and a "slice" is a unit constituting a part of a picture in coding. One picture may be composed of a plurality of slices, and if necessary, the picture and the slice may be used interchangeably.
[0021] A "pixel" or "pel" can mean the smallest unit constituting one picture (or video). Also, the term "sample" can be used as a term corresponding to a pixel. A sample generally indicates a pixel or the value of a pixel, and may indicate only the pixel / pixel value of the luminance (luma) component, or may indicate only the pixel / pixel value of the chrominance (chroma) component.
[0022] A "unit" indicates a basic unit of video processing. The unit can include at least one of a specific region of a picture and information related to the corresponding region. The unit may be used interchangeably with terms such as "block" or "area" in some cases. In general, an MxN block can indicate a set of samples or transform coefficients consisting of M columns and N rows.
[0023] FIG. 1 is a drawing schematically explaining the configuration of a video encoding apparatus to which the present invention can be applied.
[0024] Referring to FIG. 1, the video encoding apparatus 100 may include a picture division unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a conversion unit 122, a quantization unit 123, a reordering unit 124, an inverse quantization unit 125, and an inverse conversion unit 126.
[0025] The picture division unit 105 can divide the input picture into at least one processing unit.
[0026] As an example, the processing unit is called a Coding Unit (CU). In this case, the coding unit can be recursively divided from the Largest Coding Unit (LCU) by a Quad-Tree Binary-Tree (QTBT) structure. For example, one coding unit can be divided into a plurality of deeper coding units based on a quadtree structure and / or a binary tree structure. In this case, for example, the quadtree structure can be applied first and the binary tree structure can be applied later. Alternatively, the binary tree structure can also be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to the present invention can be executed. In this case, based on the coding efficiency according to video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with deeper depths so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and restoration described later.
[0027] As another example, the processing unit may also include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit can be split from the largest coding unit (LCU) into deeper coding units by a quadtree structure. In this case, based on coding efficiency according to video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively split into coding units with deeper depths, and the coding unit with the optimal size can be used as the final coding unit. When the smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit means the coding unit that serves as the basis for partitioning or splitting into a prediction unit or a transform unit. The prediction unit is a unit that is partitioned from the coding unit and is a unit for sample prediction. At this time, the prediction unit can also be divided into subblocks. The transform unit can be split from the coding unit by a quadtree structure and is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit is also called a coding block (CB), the prediction unit is also called a prediction block (PB), and the transform unit is also called a transform block (TB). The prediction block or the prediction unit means a specific area in the form of a block within a picture and can include an array of prediction samples.Also, a transform block or a transform unit means a specific area in a block form within a picture and can include an array of transform coefficients or residual samples.
[0028] The prediction unit 110 can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block including prediction samples for the current block. The unit of prediction executed by the prediction unit 110 is a coding block, or a transform block, or a prediction block.
[0029] The prediction unit 110 can determine whether intra prediction or inter prediction is applied to the current block. As an example, the prediction unit 110 can determine whether intra prediction or inter prediction is applied in terms of a CU unit.
[0030] In the case of intra prediction, the prediction unit 110 can derive prediction samples for the current block based on reference samples outside the current block within the picture (hereinafter referred to as the current picture) to which the current block belongs. At this time, the prediction unit 110 can (i) derive prediction samples based on the average or interpolation of neighboring reference samples of the current block, and (ii) also derive the prediction samples based on reference samples existing in a specific (prediction) direction with respect to the prediction samples among the neighboring reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angle mode, and in the case of (ii), it is called a directional mode or an angular mode. The prediction modes in intra prediction can have, for example, 33 directional prediction modes and at least two or more non-directional modes. The non-directional modes can include a DC prediction mode and a planar mode (Planar mode). The prediction unit 110 can also use the prediction modes applied to neighboring blocks to determine the prediction modes to be applied to the current block.
[0031] In the case of inter prediction, the prediction unit 110 can derive a prediction sample for the current block based on samples specified by motion vectors on the reference picture. The prediction unit 110 can apply any one of a skip mode, a merge mode, and an MVP (Motion Vector Prediction) mode to derive a prediction sample for the current block. In the case of the skip mode and the merge mode, the prediction unit 110 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, the motion vector of an adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be derived by using it as the motion vector predictor of the current block.
[0032] In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the temporal neighboring blocks is also called a collocated picture (colPic). The motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and output in the form of a bitstream.
[0033] When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top picture on the reference picture list can also be used as a reference picture. The reference pictures included in the reference picture list (Picture Order Count) can be sorted based on the difference in POC (Picture Order Count) between the current picture and the corresponding reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0034] The subtraction unit 121 generates a residual sample that is the difference between the original sample and the predicted sample. When the skip mode is applied, no residual sample is generated as described above.
[0035] The conversion unit 122 converts the residual samples in units of conversion blocks to generate transform coefficients. The conversion unit 122 can perform the conversion according to the size of the corresponding conversion block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the corresponding conversion block. For example, when intra prediction is applied to the coding block or the prediction block that overlaps with the conversion block, and the conversion block is a 4×4 residual array, the residual samples are converted using a DST (Discrete Sine Transform) conversion kernel, and in other cases, the residual samples can be converted using a DCT (Discrete Cosine Transform) conversion kernel.
[0036] The quantization unit 123 can quantize the transform coefficients to generate quantized transform coefficients.
[0037] The reordering unit 124 reorders the quantized transform coefficients. The reordering unit 124 can reorder the quantized transform coefficients in block form into a one-dimensional vector form via a coefficient scanning method. Here, although the reordering unit 124 has been described as a separate configuration, it may also be part of the quantization unit 123.
[0038] The entropy encoding unit 130 can perform entropy encoding on the quantized transform coefficients. Entropy encoding can include encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding), etc. The entropy encoding unit 130 can also encode information necessary for video restoration (e.g., values of syntax elements, etc.) other than the quantized transform coefficients, either together or separately, by entropy encoding or a pre-set method. The encoded information can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer) units.
[0039] The inverse quantization unit 125 inverse-quantizes the values (quantized transform coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inverse-transforms the values inverse-quantized by the inverse quantization unit 125 to generate residual samples.
[0040] The addition unit 140 adds the residual samples and the prediction samples to restore the picture. The residual samples and the prediction samples can be added in block units to generate a restored block. Here, although the addition unit 140 has been described as a separate configuration, it may also be part of the prediction unit 110. On the other hand, the addition unit 140 is also called a reconstruction module or a restored block generation unit.
[0041] For the reconstructed picture, the filter unit 150 can apply a deblocking filter and / or a sample adaptive offset. Through deblock filtering and / or sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process can be corrected. The sample adaptive offset can be applied on a sample-by-sample basis and can be applied after the deblock filtering process is completed. The filter unit 150 can also apply an ALF (Adaptive Loop Filter) to the reconstructed picture. The ALF can be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset have been applied.
[0042] The memory 160 can store the reconstructed picture (decoded picture) or information necessary for encoding / decoding. Here, the reconstructed picture is the one after the filtering procedure by the filter unit 150 is completed. The stored reconstructed picture can be utilized as a reference picture for (inter) prediction of other pictures. For example, the memory 160 can store the (reference) picture used for inter prediction. At this time, the picture used for inter prediction can be specified by a reference picture set or a reference picture list.
[0043] FIG. 2 is a drawing schematically explaining the configuration of a video decoding apparatus to which the present invention can be applied.
[0044] Referring to FIG. 2, the video decoding apparatus 200 can include an entropy decoding unit 210, a residual processing unit 220, a prediction unit 230, an addition unit 240, a filter unit 250, and a memory 260. Here, the residual processing unit 220 can include a reordering unit 221, an inverse quantization unit 222, and an inverse transform unit 223.
[0045] When a bitstream including video information is input, the video decoding apparatus 200 can restore the video corresponding to the process in which the video information is processed by the video encoding apparatus.
[0046] For example, the video decoding apparatus 200 can execute video decoding using the processing units applied by the video encoding apparatus. Therefore, the processing unit block for video decoding is, as an example, a coding unit, and as another example, a coding unit, a prediction unit, or a conversion unit. The coding unit can be divided from the maximum coding unit by a quadtree structure and / or a binary tree structure.
[0047] The prediction unit and the conversion unit can be further used as the case may be. In this case, the prediction block is a block derived from or partitioned from the coding unit and is a unit for sample prediction. At this time, the prediction unit can also be divided into sub-blocks. The conversion unit can be divided from the coding unit by a quadtree structure and is a unit for deriving conversion coefficients or a unit for deriving a residual signal from the conversion coefficients.
[0048] The entropy decoding unit 210 can parse the bitstream and output information necessary for video restoration or picture restoration. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the conversion coefficient for the residual.
[0049] More specifically, the CABAC entropy decoding method receives the BIN corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of adjacent and blocks to be decoded, or the information of symbols / BINs decoded in the previous step, predicts the occurrence probability of the BIN based on the determined context model, and executes arithmetic decoding (decoding) of the BIN, thereby generating a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / BIN for the context model of the next symbol / BIN.
[0050] Among the information related to prediction in the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction unit 230, and the residual value obtained by performing entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficient, can be input to the rearrangement unit 221.
[0051] The rearrangement unit 221 can rearrange the quantized transform coefficients in a two-dimensional block form. The rearrangement unit 221 can perform rearrangement corresponding to the coefficient scan executed in the encoding device. Here, although the rearrangement unit 221 is described separately, it may be a part of the inverse quantization unit 222.
[0052] The inverse quantization unit 222 can inverse-quantize the quantized transform coefficients based on the (inverse) quantization parameters and output the transform coefficients. At this time, the information for deriving the quantization parameters can be signaled from the encoding device.
[0053] The inverse transform unit 223 can inverse-transform the transform coefficients to derive residual samples.
[0054] The prediction unit 230 can perform a prediction on the current block and generate a predicted block that includes a prediction sample for the current block. The unit of prediction executed by the prediction unit 230 is a coding block, or a transform block, or a prediction block.
[0055] Based on the information regarding the prediction, the prediction unit 230 can determine whether to apply intra prediction or inter prediction. At this time, the unit for determining which of intra prediction and inter prediction to apply is different from the unit for generating the prediction sample. Also, in inter prediction and intra prediction, the units for generating the prediction sample are different. For example, which of inter prediction and intra prediction to apply can be determined in CU units. Also, for example, in inter prediction, the prediction mode can be determined in PU units to generate a prediction sample, and in intra prediction, the prediction mode can be determined in PU units and the prediction sample can be generated in TU units.
[0056] In the case of intra prediction, the prediction unit 230 can derive a prediction sample for the current block based on adjacent reference samples within the current picture. The prediction unit 230 can apply a directional mode or a non - directional mode based on the adjacent reference samples of the current block to derive a prediction sample for the current block. At this time, the intra - prediction mode of the adjacent block can also be used to determine the prediction mode to be applied to the current block.
[0057] In the case of inter prediction, the prediction unit 230 can derive a prediction sample for the current block based on samples specified on the reference picture by motion vectors on the reference picture. The prediction unit 230 can apply any one of the skip mode, the merge mode, and the MVP mode to derive a prediction sample for the current block. At this time, motion information necessary for the inter prediction of the current block provided in the video encoding device, such as information regarding motion vectors, reference picture indexes, etc., can be obtained or derived based on the information regarding the above prediction.
[0058] In the case of the skip mode and the merge mode, the motion information of adjacent blocks can be used as the motion information of the current block. At this time, the adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.
[0059] The prediction unit 230 can construct a merge candidate list with the motion information of available adjacent blocks, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. The motion information can include motion vectors and reference pictures. When the motion information of temporally adjacent blocks is used in the skip mode and the merge mode, the top picture on the reference picture list can be used as the reference picture.
[0060] In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted.
[0061] In the case of the MVP mode, the motion vector of the current block can be derived by using the motion vectors of adjacent blocks as motion vector predictors. At this time, the adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.
[0062] As an example, when the merge mode is applied, a merge candidate list can be generated by using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks that are temporally adjacent blocks. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The information related to the prediction can include a merge index indicating a candidate block having an optimal motion vector selected from among the candidate blocks included in the merge candidate list. At this time, the prediction unit 230 can derive the motion vector of the current block by using the merge index.
[0063] As another example, when the MVP (Motion Vector Prediction) mode is applied, a motion vector predictor candidate list can be generated by using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks which are temporally adjacent blocks. That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks which are temporally adjacent blocks can be used as motion vector candidates. The information related to the prediction can include a predicted motion vector index indicating the optimal motion vector selected from among the motion vector candidates included in the above list. At this time, the prediction unit 230 can use the above motion vector index to select the predicted motion vector of the current block from among the motion vector candidates included in the motion vector candidate list. The prediction unit of the encoding device can obtain the difference (MVD) of the motion vectors between the motion vector of the current block and the motion vector predictor, and can encode this and output it in the form of a bit stream. That is, the MVD is obtained as the value obtained by subtracting the above motion vector predictor from the motion vector of the current block. At this time, the prediction unit 230 can obtain the motion vector difference included in the above information related to the prediction, and can derive the above motion vector of the current block through the addition of the above motion vector difference and the above motion vector predictor. Also, the prediction unit can obtain or derive from the above information related to the prediction a reference picture index indicating a reference picture and the like.
[0064] The addition unit 240 can add the residual samples and the prediction samples to restore the current block or the current picture. The addition unit 240 can also add the residual samples and the prediction samples in block units to restore the current picture. When the skip mode is applied, since the residual is not transmitted, the prediction samples can become the restored samples. Here, although the addition unit 240 has been described in a separate configuration, it may also be a part of the prediction unit 230. On the other hand, the addition unit 240 is also called a reconstruction module or a reconstructed block generation unit.
[0065] The filter unit 250 can apply a deblocking filter sample adaptive offset, and / or an ALF, etc. to the restored picture. At this time, the sample adaptive offset can be applied in units of samples, and can also be applied after deblocking filtering. The ALF can also be applied after deblocking filtering and / or the sample adaptive offset.
[0066] The memory 260 can store the restored picture (decoded picture) or the information necessary for decoding. Here, the restored picture is the restored picture after the filtering procedure by the above filter unit 250. For example, the memory 260 can store the picture used for inter prediction. At this time, the picture used for inter prediction can also be specified by a reference picture set or a reference picture list. The restored picture can be used as a reference picture for other pictures. Also, the memory 260 can output the restored pictures in the output order.
[0067] On the other hand, through the above-described transformation, the low-frequency conversion coefficients for the residual block of the current block can be derived, and a zero tail can be derived at the end of the residual block.
[0068] Specifically, the above transformation can be composed of two main processes, where the main processes can include a core transform and a secondary transform. The transform including the core transform and the secondary transform can be referred to as a multiple transform technique.
[0069] FIG. 3 schematically shows the multiple transform technique according to the present invention.
[0070] Referring to FIG. 3, the conversion unit can correspond to the conversion unit in the encoding device of FIG. 1 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of FIG. 1 or the inverse conversion unit in the decoding device of FIG. 2 described above.
[0071] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S310). Here, the primary conversion can include an Adaptive Multiple core Transform (AMT). The Adaptive Multiple core Transform can be denoted as MTS (Multiple Transform Set).
[0072] The Adaptive Multiple core Transform can be shown as a method of additionally using DCT (Discrete Cosine Transform) type 2 and DST (Discrete Sine Transform) type 7, DCT type 8, and / or DST type 1 for conversion. That is, the Adaptive Multiple core Transform is a conversion method that converts a residual signal (or a residual block) in the spatial domain into conversion coefficients (or primary conversion coefficients) in the frequency domain based on a plurality of selected conversion kernels among the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary conversion coefficients can also be referred to as temporary conversion coefficients from the perspective of the conversion unit.
[0073] In other words, when an existing conversion method is applied, a conversion from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2 to generate conversion coefficients. In contrast, when the adaptive multi-core conversion is applied, a conversion from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or primary conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be referred to as conversion types, conversion kernels, or conversion cores.
[0074] For reference, the DCT / DST conversion types can be defined based on basis functions, and the basis functions can be shown as follows in the following table.
[0075]
Table 1
[0076] When the adaptive multi-core conversion is performed, among the conversion kernels, a vertical conversion kernel and a horizontal conversion kernel for the target block can be selected, a vertical conversion for the target block can be performed based on the vertical conversion kernel, and a horizontal conversion for the target block can be performed based on the horizontal conversion kernel. Here, the horizontal conversion can indicate a conversion for the horizontal component of the target block, and the vertical conversion can indicate a conversion for the vertical component of the target block. The vertical conversion kernel / horizontal conversion kernel can be adaptively determined based on the prediction mode of the target block (CU or sub-block) encompassing the residual block and / or the conversion index indicating the conversion subset.
[0077] For example, the adaptive multi-core transform can be applied when the width and height of the target block are both smaller than or equal to 64. Whether the adaptive multi-core transform of the target block is applied or not can be determined based on the CU level flag. Specifically, when the CU level flag is 0, the existing conversion method described above can be applied. That is, when the CU level flag is 0, the conversion from the spatial domain to the frequency domain for the residual signal (or residual block) based on the DCT type 2 can be applied to generate conversion coefficients, and the conversion coefficients can be encoded. On the other hand, here the target block can be a CU. When the CU level flag is 0, the adaptive multi-core transform can be applied to the target block.
[0078] Also, in the case of the luma block of the target block to which the adaptive multi-core conversion is applied, two additional flags can be signaled, and based on the above flags, the vertical conversion kernel and the horizontal conversion kernel can be selected. The flag related to the vertical conversion kernel can be indicated as the AMT vertical flag, and AMT_TU_vertical_flag (or EMT_TU_vertical_flag) can indicate the syntax element of the AMT vertical flag. The flag related to the horizontal conversion kernel can be indicated as the AMT horizontal flag, and AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) can indicate the syntax element of the AMT horizontal flag. The AMT vertical flag can indicate one of the conversion kernel candidates included in the conversion subset for the vertical conversion kernel, and the conversion kernel candidate indicated by the AMT vertical flag can be derived as the vertical conversion kernel for the target block. Also, the AMT horizontal flag can indicate one of the conversion kernel candidates included in the conversion subset for the horizontal conversion kernel, and the conversion kernel candidate indicated by the AMT horizontal flag can be derived as the horizontal conversion kernel for the target block. On the other hand, the AMT vertical flag can also be indicated as the MTS vertical flag, and the AMT horizontal flag can also be indicated as the MTS horizontal flag.
[0079] On the other hand, three conversion subsets can be preset, and based on the intra prediction mode applied to the target block, one of the conversion subsets can be derived as the conversion subset for the vertical conversion kernel. Also, based on the intra prediction mode applied to the target block, one of the conversion subsets can be derived as the conversion subset for the horizontal conversion kernel. For example, the preset conversion subsets can be derived as follows in the following table.
[0080]
Table 2
[0081] Referring to Table 2, the conversion subset with an index value of 0 can indicate a conversion subset that includes DST type 7 and DCT type 8 as conversion kernel candidates, the conversion subset with an index value of 1 can indicate a conversion subset that includes DST type 7 and DST type 1 as conversion kernel candidates, and the conversion subset with an index value of 2 can indicate a conversion subset that includes DST type 7 and DCT type 8 as conversion kernel candidates.
[0082] The conversion subset for the vertical conversion kernel and the conversion subset for the horizontal conversion kernel derived based on the intra prediction mode applied to the target block can be derived as shown in the following table.
[0083]
Table 3
[0084] Here, V indicates the conversion subset for the vertical conversion kernel, and H indicates the conversion subset for the horizontal conversion kernel.
[0085] When the value of the AMT flag (or EMT_CU_flag) for the target block is 1, as illustrated in Table 3, a conversion subset for the vertical conversion kernel and a conversion subset for the horizontal conversion kernel can be derived based on the intra prediction mode of the target block. After that, among the conversion kernel candidates included in the conversion subset for the vertical conversion kernel, the conversion kernel candidate indicated by the AMT vertical flag of the target block can be derived as the vertical conversion kernel of the target block, and among the conversion kernel candidates included in the conversion subset for the horizontal conversion kernel, the conversion kernel candidate indicated by the AMT horizontal flag of the target block can be derived as the horizontal conversion kernel of the target block. On the other hand, the AMT flag can also be denoted as the MTS flag.
[0086] For reference, for example, the intra prediction mode can include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is an example, and the present invention can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can indicate the LM (Linear Model) mode.
[0087] FIG. 4 exemplarily shows the 65 intra-directional modes of the prediction directions.
[0088] Referring to FIG. 4, it is possible to distinguish an intra prediction mode having horizontal directionality and an intra prediction mode having vertical directionality, centered around the 34th intra prediction mode having a diagonal prediction direction upward to the left. H and V in FIG. 4 respectively mean horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in 1 / 32 units on the sample grid position. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode can be called a diagonal intra prediction mode downward to the left, the 34th intra prediction mode can be called a diagonal intra prediction mode upward to the left, and the 66th intra prediction mode can be called a diagonal intra prediction mode upward to the right.
[0089] The conversion unit can perform a secondary conversion based on the above (primary) conversion coefficients to derive (secondary) conversion coefficients (S320). If the above primary conversion is a conversion from the spatial domain to the frequency domain, the above secondary conversion can be regarded as a conversion from the frequency domain to the frequency domain. The above secondary conversion can include a non-separable transform. In this case, the above secondary conversion can be referred to as a Non-Separable Secondary Transform (NSST) or a Mode-Dependent Non-Separable Secondary Transform (MDNSST). The above non-separable secondary conversion can be shown as a conversion that performs a secondary conversion on the (primary) conversion coefficients derived through the above primary conversion based on a non-separable transform matrix to generate conversion coefficients (or secondary conversion coefficients) for the residual signal. Here, the vertical conversion and the horizontal conversion cannot be separated (or the horizontal and vertical conversions cannot be applied independently) for the above (primary) conversion coefficients based on the above non-separable transform matrix, and the conversion can be applied at once. In other words, the above non-separable secondary conversion can be shown as a conversion method that does not separate the vertical component and the horizontal component of the above (primary) conversion coefficients based on the above non-separable transform matrix, but converts them together to generate conversion coefficients (or secondary conversion coefficients). The above non-separable secondary conversion can be applied to the top-left region of a block composed of (primary) conversion coefficients (hereinafter, which can be referred to as a conversion coefficient block or a target block). For example, when both the width (W) and the height (H) of the above conversion coefficient block are 8 or more, an 8×8 non-separable secondary conversion can be applied to the top-left 8×8 region (hereinafter, the top-left target region) of the above conversion coefficient block. Also, when both the width (W) and the height (H) of the above conversion coefficient block are 4 or more and the width (W) or the height (H) of the above conversion coefficient block is less than 8, a 4×4 non-separable secondary conversion can be applied to the top-left min(8, W)×min(8, H) region of the above conversion coefficient block.
[0090] Specifically, for example, when a 4×4 input block is used, the non-separable quadratic transformation can be performed as follows.
[0091] The above 4×4 input block X can be shown as follows.
[0092] <Equation 1>
Number
[0093] When the above X is shown in vector form, the vector TIFF0007715904000005.tif84 can be shown as follows.
[0094] <Equation 2>
Number
[0095] In this case, the above quadratic non-separable transformation can be calculated as follows.
[0096] <Equation 3>
Number
[0097] Here, TIFF0007715904000008.tif75 represents the transformation coefficient vector, and T represents the 16×16 (non-separable) transformation matrix.
[0098] Through the above Equation 3, the 16×1 transformation coefficient vector TIFF0007715904000009.tif75 can be derived, and the above TIFF0007715904000010.tif75 can be re-organized into 4×4 blocks through the scanning order (horizontal, vertical, diagonal, etc.). However, the above calculations are examples, and for reducing the computational complexity of non-separable second-order transforms, HyGT (Hypercube-Givens Transform) etc. can also be used for the calculation of non-separable second-order transforms.
[0099] On the other hand, in the above non-separable second-order transform, the transformation kernel (or transformation core, transformation type) can be selected in a mode dependent manner. Here, the mode can include the intra prediction mode and / or the inter prediction mode.
[0100] As described above, the above non-separable second-order transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the above transformation coefficient block. That is, the above non-separable second-order transform can be performed based on an 8×8 sub-block size or a 4×4 sub-block size. For example, for the above mode-dependent transformation kernel selection, for both the 8×8 sub-block size and the 4×4 sub-block size, 35 sets of non-separable second-order transformation kernels for non-separable second-order transforms can be configured, three for each. That is, 35 sets of transformations can be configured for the 8×8 sub-block size, and 35 sets of transformations can be configured for the 4×4 sub-block size. In this case, each of the 35 sets of transformations for the 8×8 sub-block size can include three 8×8 transformation kernels, and in this case, each of the 35 sets of transformations for the 4×4 sub-block size can include three 4×4 transformation kernels. However, for the above transformation sub-block size, the number of sets, and the number of transformation kernels in the set, sizes other than 8×8 or 4×4 can be used as examples, or n sets can be configured, and each set can include k transformation kernels.
[0101] The above conversion set can be referred to as an NSST set, and the conversion kernels within the NSST set can be referred to as NSST kernels. Among the above conversion sets, the selection of a specific set can be performed, for example, based on the intra prediction mode of a target block (CU or sub-block).
[0102] In this case, the mapping between the above 35 conversion sets and the intra prediction mode can be shown, for example, as in the following table. For reference, when the LM mode is applied to a target block, no secondary conversion may be applied to the target block.
[0103]
Table 4
[0104] On the other hand, once it is determined that a specific set is to be used, one of the k conversion kernels within the specific set can be selected through a non-separable secondary conversion index. The encoding device can derive a non-separable secondary conversion index indicating a specific conversion kernel based on an RD (Rate-Distortion) check, and can signal the non-separable secondary conversion index to the decoding device. The decoding device can select one of the k conversion kernels within the specific set based on the non-separable secondary conversion index. For example, an NSST index value of 0 can indicate the first non-separable secondary conversion kernel, an NSST index value of 1 can indicate the second non-separable secondary conversion kernel, and an NSST index value of 2 can indicate the third non-separable secondary conversion kernel. Alternatively, an NSST index value of 0 can indicate that no first non-separable secondary conversion is applied to the target block, and NSST index values of 1 to 3 can indicate the above three conversion kernels.
[0105] Referring to FIG. 3, the conversion unit can perform the above non-separable second-order conversion based on the selected conversion kernel to obtain (second-order) conversion coefficients. As described above, the conversion coefficients can be derived as the quantization coefficients quantized through the quantization unit, encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse conversion unit in the encoding device.
[0106] On the other hand, when the second-order conversion is omitted, the (first-order) conversion coefficients, which are the output of the above first-order (separable) conversion, can be derived as the quantization coefficients quantized through the quantization unit as described above, encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse conversion unit in the encoding device.
[0107] The inverse conversion unit can perform a series of procedures in the reverse order of the procedures performed by the above conversion unit. The inverse conversion unit receives the (inverse quantized) conversion coefficients, performs a second-order (inverse) conversion to derive the (first-order) conversion coefficients (S350), and performs a first-order (inverse) conversion on the above (first-order) conversion coefficients to obtain a residual block (residual samples) (S360). Here, the above first-order conversion coefficients can be referred to as the (modified) conversion coefficients corrected from the perspective of the inverse conversion unit. As described above, the encoding device and the decoding device can generate a restored block based on the above residual block and the predicted block, and a restored picture can be generated based on this.
[0108] On the other hand, as described above, when the second-order (inverse) conversion is omitted, the (inverse quantized) conversion coefficients can be received and the above first-order (separable) conversion can be performed to obtain a residual block (residual samples). As described above, the encoding device and the decoding device can generate a restored block based on the above residual block and the predicted block, and a restored picture can be generated based on this.
[0109] FIGS. 5A to 5C are diagrams for explaining selective transform according to an embodiment of the present invention.
[0110] In this specification, "target block" can mean the current block or the residual block on which coding is performed.
[0111] FIG. 5A shows an example of deriving transform coefficients through transformation.
[0112] In video coding, transformation, as illustrated in FIG. 5A, can show a process of transforming an input vector R based on a transform matrix to generate a transform coefficient vector C for the input vector R. The above input vector R can also represent first-order transform coefficients. Alternatively, the above input vector R can also represent a residual vector, that is, residual samples. On the other hand, the above transform coefficient vector C can also be denoted as output vector C.
[0113] FIG. 5B shows a specific example of deriving transform coefficients through transformation. FIG. 5B specifically shows the transformation process illustrated in FIG. 5A. As previously described in FIG. 3, in non-separable second-order transformation (hereinafter referred to as 'NSST'), after dividing the block data of the transform coefficients obtained by applying the first-order transformation into MxM blocks, for each MxM block, M 2 xM 2 NSST can be performed. M can be, for example, 4 or 8, but is not limited thereto. The above M 2 can be N. In this case, as illustrated in FIG. 5B, the above input vector R can be a (1xN)-dimensional vector including (first-order) transform coefficients r1 to r N and the above transform coefficient vector C can be an (Nx1)-dimensional vector including transform coefficients c1 to c N . That is, the above input vector R can include N (first-order) transform coefficients r1 to r N and the size of the above input vector R can be 1xN. Also, the above transform coefficient vector C can include N transform coefficients c1 to c N and the size of the above transform coefficient vector C can be Nx1.
[0114] In order to derive the above conversion coefficient vector C, the above input vector R can go through the above conversion matrix. That is, the above input vector R can be derived as the above conversion coefficient vector C based on the above conversion matrix.
[0115] On the other hand, the above conversion matrix can include N basis vectors B1 to B N As shown in FIG. 5B, the above basis vectors B1 to B N can be (1xN)-dimensional vectors. That is, the size of the above basis vectors B1 to B N can be 1xN. The conversion coefficient vector C can be generated based on the (primary) conversion coefficients of the above input vector R and each of the above basis vectors of the above conversion matrix. For example, the inner product of the above input vector and each basis vector can be derived as the above conversion coefficient vector C.
[0116] On the other hand, two main issues occur in the above-described conversion. Specifically, in relation to the number of multiplications and additions required to generate the output vector, high computational complexity and memory requirements for storing the generated coefficients can occur as main problems.
[0117] For example, the computational complexity and memory requirements required for separable transform and non-separable transform can be derived as shown in the following table.
[0118]
Table 5
[0119] Referring to Table 5, the memory for storing the coefficients generated through the above separable transform is N2 can be required, and the number of calculations is 2N 3 can be. The number of calculations indicates the computational complexity. Also, the memory for storing (storing) the coefficients generated through the non-separable transform can be required to be N4, and the number of calculations can be N4. The number of calculations indicates the computational complexity. That is, the higher the number of calculations, the higher the computational complexity may be, and the lower the number of calculations, the lower the computational complexity may be.
[0120] As shown in Table 5 above, the memory requirements and the number of calculations for the non-separable transform can increase significantly compared to the separable transform. Also, the larger the size of the target block for which the non-separable transform is performed, that is, the larger the N, the greater the discrepancy between the memory requirements and the number of calculations for the separable transform and the memory requirements and the number of calculations for the non-separable transform can be.
[0121] The non-separable transform provides a better coding gain compared to the separable transform, but as shown in Table 5 above, due to the computational complexity of the non-separable transform, in existing video coding standards, the non-separable transform is not used, and also, since the computational complexity of the separable transform also increases as the size of the target block increases, in the existing HEVC standard, the separable transform is restricted to be used only for target blocks with a size of 32x32 or less.
[0122] Here, the present invention proposes a selective transformation. The above selective transformation can significantly reduce the computational complexity and memory requirements, thereby increasing the efficiency of computationally intensive transformation blocks and generating the effect of improving coding efficiency. That is, the above selective transformation can be used to solve the problem of computational complexity that occurs during the transformation of large-sized blocks or non-separable transformations. The above selective transformation can be used for any type of transformation, such as a primary transformation (or can be referred to as a core transform), a secondary transformation, etc. For example, the above selective transformation can be applied to the above core transformation for an encoding device / decoding device, and can generate the effect of significantly reducing the encoding time / decoding time.
[0123] FIG. 5C shows an example of deriving transform coefficients through the above selective transformation. The above selective transformation can be meant as a transformation performed on a target block based on a transformation matrix including basis vectors including a selective number of elements.
[0124] The simplified transformation is a method proposed through the motivation that the N elements of the basis vector of the transformation matrix may include overlapping or unimportant elements, and the above elements can be excluded to reduce the computational complexity and memory requirements. For example, referring to FIG. 5C, among the N elements of the basis vector B1, Z0 elements may be unimportant elements. In this case, a truncated basis vector B1 including only N1 elements can be derived. Here, the above N1 can be N - Z0. The above truncated basis vector B1 can be shown as a modified basis vector B1.
[0125] Referring to FIG. 5C, when the corrected basis vector B1 is applied as part of the transformation matrix to the input vector R, the transformation coefficient C1 can be derived. From the experimental results, it is observed that the transformation coefficient C1 is the same value as the transformation coefficient C1 derived when the existing basis vector B1 is applied as part of the transformation matrix to the input vector R. That is, it means that assuming that the minute elements of each basis vector are 0, the number of multiplications required can be significantly reduced without a large difference in the results. Also, next, the number of elements (i.e., the elements of the transformation matrix) that have to be stored for this operation can be reduced.
[0126] An association vector is proposed to define the positions of the elements of the basis vector that are not important (or meaningful). A (1xN) - dimensional association vector A1 can be considered to derive the corrected basis vector B1 of N1 dimensions. That is, an association vector A1 of size 1xN can be considered to derive a corrected basis vector B1 of size 1xN1 (i.e., the corrected basis vector B1 containing N1 elements).
[0127] Referring to FIG. 5C, the values derived by applying the association vectors for each of the basis vectors B1 to B N can be transmitted to the basis vectors. Through this, only some elements of the input vector R can be calculated with the elements of the basis vectors. Specifically, the association vector can include 0 and 1, and is calculated such that among the elements of the input vector R, the selected elements are multiplied by 1 and the unselected elements are multiplied by 0, so that only the selected elements can pass through and be transmitted to the basis vectors.
[0128] For example, the above related vector A1 can be applied to the above input vector R, and among the elements of the above input vector R, only N1 elements specified by the above related vector A1 can be used to calculate the inner product with the above basis vector B1. The above inner product can indicate C1 of the above conversion coefficient vector C. Here, the above basis vector B1 can include N1 elements, and the above N1 can be N or less. The operations of the above input vector R with the related vector A1 and the basis vector B1 described above are for the related vectors A2 to A N and the basis vectors B2 to B N can also be performed.
[0129] Since the above related vector includes only 0 and / or 1, binary values, there can be an advantage in storing the above related vector. 0 of the above related vector can indicate that the element of the above input vector R corresponding to 0 is not transmitted to the conversion matrix for the above inner product calculation, and 1 of the above related vector can indicate that the element of the above input vector R corresponding to 1 is transmitted to the above conversion matrix for the above inner product calculation. For example, the related vector A of size 1xN k can include A k1 to A kN . When A k of the above related vector A kn is 0, r n of the above input vector R may not be able to pass through. That is, r n of the above input vector R may not be transmitted to the above conversion vector B n . Also, when the above A kn is 1, r n of the above input vector R can pass through. That is, r n of the above input vector R can be transmitted to the above conversion vector B n and can be used in the calculation to derive c n of the above conversion coefficient vector C.
[0130] FIG. 6 schematically shows a multiple conversion technique in which the above selective conversion is applied to a second-order conversion.
[0131] Referring to FIG. 6, the conversion unit can correspond to the conversion unit in the encoding device of FIG. 1 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of FIG. 1 or the inverse conversion unit in the decoding device of FIG. 2 described above.
[0132] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S610). Here, the above primary conversion can include the AMT described above.
[0133] When the above adaptive multi-core conversion is applied, a conversion from the spatial domain to the frequency domain for the residual signal (or, the residual block) is applied based on, for example, DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., and conversion coefficients (or, primary conversion coefficients) can be generated. Here, the above primary conversion coefficients can be referred to as temporary conversion coefficients from the perspective of the conversion unit. Also, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. can be referred to as conversion types, conversion kernels or conversion cores. For reference, the above DCT / DST conversion types can be defined based on basis functions, and the above basis functions can be shown as described in Table 1 above. Specifically, the process of applying the above adaptive multi-core conversion to derive the above primary conversion coefficients is as described above.
[0134] The conversion unit can perform selective conversion based on the above (primary) conversion coefficients to derive (secondary) conversion coefficients (S620). The selective conversion can be meant as a conversion performed on the above (primary) conversion coefficients for a target block based on a conversion matrix including modified basis vectors and an association matrix including associated vectors for the above basis vectors. The above modified basis vectors can indicate basis vectors including elements of N or less. That is, the above modified basis vectors can indicate basis vectors including a selected specific number of elements out of N elements. For example, the modified basis vector B n can be a (1xN n )-dimensional vector, and the above N n may be less than or equal to the above N. That is, the size of the modified basis vector B n can be of (1xN n ) size, and the above N n may be less than or equal to the above N. Here, the above N can be the product of the height and width of the upper left target area of the target block to which the above selected conversion is applied. Alternatively, the above N can be the total number of conversion coefficients of the upper left target area of the target block to which the above selected conversion is applied. On the other hand, the above conversion matrix including the modified basis vectors can be shown as a modified conversion matrix. Also, the above conversion matrix can be shown as a Transform Bases Block (TBB), and the above association matrix can be shown as an Association Vectors Block (AVB).
[0135] The conversion unit can perform the above selective conversion based on the above modified conversion matrix and the above association matrix to obtain (secondary) conversion coefficients. The above conversion coefficients can be derived as conversion coefficients quantized through the quantization unit as described above, and can be encoded and signaled to the decoding device and transmitted to the inverse quantization unit / inverse conversion unit in the encoding device.
[0136] The inverse conversion unit can perform a series of procedures in the reverse order of the procedures performed by the above-described conversion unit. The inverse conversion unit receives (inverse quantized) conversion coefficients, performs a selective (inverse) conversion to derive (primary) conversion coefficients (S650), and can perform a primary (inverse) conversion on the above (primary) conversion coefficients to obtain a residual block (residual samples) (S660). Here, the above primary conversion coefficients can be referred to as modified conversion coefficients from the perspective of the inverse conversion unit. As described above, the encoding device and the decoding device can generate a restored block based on the above residual block and the predicted block, and a restored picture can be generated based on this.
[0137] On the other hand, in one embodiment of the above selective conversion, the present invention proposes a selective conversion combined with a simplified conversion.
[0138] In this specification, "simplified conversion" can be meant to be a conversion performed on the residual samples for the target block based on a transform matrix whose size has been reduced by a simplification factor.
[0139] In a simplification transformation according to an embodiment, an N-dimensional vector can be mapped to an R-dimensional vector located in another space to determine a simplification transformation matrix, where R is smaller than N. That is, the above simplification transformation can mean a transformation performed on the residual samples for the target block based on a reduced transform matrix including R basis vectors. Here, N can mean the square of the length of one side of the block (or target region) to which the transformation is applied or the total number of transformation coefficients corresponding to the block (or target region) to which the transformation is applied, and the simplification factor can mean the R / N value. The simplification factor can be referred to by various terms such as a reduced factor, a reduction factor, a reduced factor, a reduction factor, a simplified factor, a simple factor, etc. On the other hand, R can be referred to as a reduced coefficient, but in some cases, the simplification factor can mean R. Also, in some cases, the simplification factor can also mean the N / R value.
[0140] The size of the above simplification transformation matrix according to an embodiment is RxN, which is smaller than the size NxN of the normal transformation matrix, and can be defined as shown in Equation 4 below.
[0141] <Equation 4>
Number
[0142] When the simplification transformation matrix T is multiplied by the transformation coefficients to which the first transformation of the target block is applied, the (second) transformation coefficients for the target block can be derived. RxN When multiplied, the (second) transformation coefficients for the target block can be derived.
[0143] When the above RST is applied, since a simplified transformation matrix of size RxN is applied to the secondary transformation, the transformation coefficients from R+1 to N can implicitly become 0. In other words, when the transformation coefficients of the target block are derived with the above RST applied, the values of the transformation coefficients from R+1 to N can be 0. Here, the transformation coefficients from R+1 to N can indicate the (R+1)-th to N-th transformation coefficients among the transformation coefficients. Specifically, the array of the transformation coefficients of the target block can be described as follows.
[0144] FIG. 7 is a diagram for explaining an array of transformation coefficients based on a target block according to an embodiment of the present invention. Hereinafter, the description regarding the transformation described later with reference to FIG. 7 can be similarly applied to the inverse transformation. For the target block (or residual block) 700, NSST based on the primary transformation and the simplified transformation can be performed. In one example, the 16x16 block illustrated in FIG. 7 represents the target block 700, and the 4x4 blocks denoted by A to P can represent subgroups of the target block 700. The primary transformation can be performed over the entire range of the target block 700. After the primary transformation is performed, NSST can be applied to the 8x8 block (hereinafter, the upper left target region) constituted by the subgroups A, B, E, and F. At this time, if NSST based on the simplified transformation is performed, only R (here, R means the simplification coefficient and R is smaller than N) NSST transformation coefficients are derived, so the NSST transformation coefficients in the range from the (R+1)-th to the N-th can each be determined to be 0. For example, when R is 16, the 16 transformation coefficients derived by performing NSST based on the simplified transformation can be assigned to each block included in the subgroup A, which is the upper left 4x4 block included in the upper left target region of the target block 700, and for each of the N-R, that is, 64-16 = 48 blocks included in the subgroups B, E, and F, the transformation coefficient 0 can be assigned. The primary transformation coefficients for which NSST based on the simplified transformation is not performed can be assigned to each block included in the subgroups C, D, G, H, I, J, K, L, M, N, O, and P.
[0145] FIG. 8 shows an example of deriving a conversion coefficient through a conversion in which the above-mentioned simplification conversion and the above-mentioned selective conversion are combined. Referring to FIG. 8, the conversion matrix can include R basis vectors, and the above-mentioned correlation matrix can include R correlation vectors. Here, the conversion matrix including the R basis vectors can be shown as a reduced transform matrix, and the correlation matrix including the R correlation vectors can be shown as a reduced association matrix.
[0146] Also, each of the above-mentioned basis vectors can include only selected elements out of N elements. For example, referring to FIG. 8, the basis vector B1 can be a 1xN1-dimensional vector including N1 elements, the basis vector B2 can be a 1xN2-dimensional vector including N2 elements, and the basis vector B R can be a 1xN R -dimensional vector including N R elements. The above N1, N2, and N R can be values less than or equal to the above N. The conversion in which the above-mentioned simplification conversion and the above-mentioned selective conversion are combined can be used for any type of conversion such as a secondary conversion and a primary conversion.
[0147] Referring to FIG. 8, an encoding device and a decoding device can apply a conversion in which the above-mentioned simplification conversion and the above-mentioned selective conversion are combined to a secondary conversion. For example, the above-mentioned selective conversion can be performed based on the reduced conversion matrix and the reduced association matrix including the above-mentioned modified basis vectors, and a (secondary) conversion coefficient can be obtained. Also, HyGT (Hypercube-Givens Transform) or the like can be used for the calculation of the above-mentioned selective conversion in order to reduce the computational complexity of the above-mentioned selective conversion.
[0148] FIG. 9 shows an example of deriving a conversion coefficient through the above-mentioned selective conversion. In one embodiment of the above-mentioned selective conversion, the correlation vectors A1, A2,... A of the above-mentioned correlation matrix NPatterns for may not exist, and the related vectors can be derived in different forms from each other. Alternatively, in other embodiments for the above selective conversion, the related vectors A1, A2,... A of the related matrix N can be derived in the same form.
[0149] Specifically, for example, the related vectors can include the same number of 1s. For example, when the number of 1s is M, the related vectors can include M 1s and N - M 0s. In this case, among the (primary) conversion coefficients of the input vector R, M conversion coefficients can be transmitted to the basis vectors. Therefore, the length of the basis vectors can also be M. That is, as shown in FIG. 9, the basis vectors can include M elements, can be (1xM)-dimensional vectors, and N1 = N2 =... = N N can be derived as = M. The related matrix and the modified conversion matrix structure shown in FIG. 9 can be shown as a symmetric architecture for the above selective conversion.
[0150] Also, in other examples, a specific pattern of elements with a value of 1 can exist, and the related vectors can be derived by repeating, rotating, and / or translating the pattern in any manner.
[0151] On the other hand, the above-described selective conversion can be applied not only with the above simplification conversion and / or the above HyGT but also with other conversion techniques.
[0152] Further, the present invention proposes a method for simplifying the related vectors in the above-described selective conversion. By simplifying the related vectors, the storage of information for performing the above selective conversion and the handling of the above selective conversion can be further improved. That is, the memory load for performing the above selective conversion can be reduced, and the handling ability of the above selective conversion can be further improved.
[0153] The effect of the simplification of the above-mentioned correlation vector can be more clearly shown when the non-zero elements among the elements included in the above-mentioned correlation vector have a continuous distribution. For example, the correlation vector A k can include a continuous string of 1s. In this case, the above A k can be expressed through two factors A ks and A kL . That is, the above-mentioned correlation vector A ks and A kL can be derived based on the above. Here, the above A k can be a factor indicating the starting point of non-zero elements (for example, 1), and the above A ks can be a factor indicating the length of non-zero elements. The above-mentioned correlation vector A kL shown based on the above factors can be derived as follows in the following table. k is derived as follows.
[0154]
Table 6
[0155] Referring to Table 6, the above-mentioned correlation vector A k can include 16 elements. That is, the above-mentioned correlation vector A k can be a 1x16 dimensional vector. The value of the factor A k indicating the starting point of non-zero elements of the above-mentioned correlation vector A ks can be derived as 0, and in this case, the above factor A ks can indicate the starting point of the non-zero elements as the first element of the above-mentioned correlation vector A k . Also, the value of the factor A k indicating the length of non-zero elements of the above-mentioned correlation vector A kL can be derived as 8, and in this case, the above factor A kLcan indicate that the length of the non-zero elements is eight. Therefore, the above A k can be derived as a vector in which, based on the above factors, as illustrated in Table 6, the first to eighth elements are 1 and the remaining elements are 0.
[0156] FIG. 10 shows an example of deriving the related vector based on two factors for the related vector and performing selective conversion. Referring to FIG. 10, a related matrix can be derived based on the factors for each related vector, and a conversion coefficient for the target block can be derived based on the related matrix and the modified conversion matrix.
[0157] On the other hand, the start point of the non-zero elements and the number of non-zero elements of each of the related vectors can be derived as fixed values, or the start point of the non-zero elements and the number of non-zero elements of each of the related vectors can also be derived in various ways.
[0158] Also, for example, the start point of the non-zero elements and the number of non-zero elements of the related vector can also be derived based on the size of the upper left target region where the conversion is performed. Here, the size of the upper left target region can indicate the number of conversion coefficients of the upper left target region, or can indicate the product of the height and width of the upper left target region. Also, in another example, the start point of the non-zero elements and the number of non-zero elements of the related vector can also be derived based on the intra prediction mode of the target block. Specifically, for example, the start point of the non-zero elements and the number of non-zero elements of the related vector can be derived based on whether the intra prediction mode of the target block is a non-directional intra prediction mode.
[0159] Alternatively, for example, the start point of non-zero elements of the correlation vector and the number of non-zero elements can be preset. Alternatively, for example, information indicating the start point of non-zero elements of the correlation vector and information indicating the number of non-zero elements can be signaled, and the correlation vector can be derived based on the information indicating the start point of non-zero elements and the information indicating the number of non-zero elements. Alternatively, other information can be used instead of the information indicating the start point of non-zero elements. For example, instead of the information indicating the start point of non-zero elements, information indicating the last position of non-zero elements can be used, and the correlation vector can be derived based on the information indicating the last position of non-zero elements.
[0160] On the other hand, the method for deriving the correlation vector based on the factor can be applied to separable transformation and non-separable transformation such as simplified transformation HyGT.
[0161] FIG. 11 schematically shows a video encoding method by an encoding apparatus according to the present invention. The method disclosed in FIG. 11 can be performed by the encoding apparatus disclosed in FIG. 1. Specifically, for example, S1100 in FIG. 11 can be performed by the subtraction unit of the encoding apparatus, S1110 can be performed by the conversion unit of the encoding apparatus, and S1120 can be performed by the entropy encoding unit of the encoding apparatus. Also, although not shown, the process of deriving the prediction sample can be performed by the prediction unit of the encoding apparatus.
[0162] The encoding device derives the residual samples of the target block (S1100). For example, the encoding device can determine whether to perform inter prediction or intra prediction on the target block, and can determine a specific inter prediction mode or a specific intra prediction mode based on the RD cost. According to the determined mode, the encoding device can derive the prediction samples for the target block, and can derive the residual samples through the addition of the original samples and the prediction samples for the target block.
[0163] The encoding device derives the transform coefficients of the target block based on a selective transform on the residual samples (S1110). The selective transform can be performed based on a modified transform matrix, and the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors can include a selected specific number of elements out of N elements. Also, the selective transform can be performed on the upper left target region of the target block, and N can be the number of residual samples located in the upper left target region. Alternatively, N can be the value obtained by multiplying the width and height of the upper left target region. For example, N can be 16 or 64.
[0164] The encoding device can perform a core transform on the residual samples to derive modified transform coefficients, and based on an association matrix including an association vector for the modified transform matrix and the modified basis vectors, perform the selective transform on the modified transform coefficients located in the upper left target region of the target block to derive the transform coefficients of the target block.
[0165] Specifically, the core conversion for the above residual samples can be performed as follows. The encoding device can determine whether to apply an Adaptive Multiple core Transform (AMT) to the above target block. In this case, an AMT flag indicating whether the adaptive multi-core transform of the above target block is applied can be generated. When the AMT is not applied to the above target block, the encoding device can derive DCT type 2 as a conversion kernel for the above target block, and perform a conversion on the above residual samples based on the DCT type 2 to derive the above modified conversion coefficients.
[0166] When the AMT is applied to the above target block, the encoding device can configure a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, derive the horizontal conversion kernel and the vertical conversion kernel based on the above conversion subsets, and perform a conversion on the above residual samples based on the above horizontal conversion kernel and the above vertical conversion kernel to derive the modified conversion coefficients. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel can include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Also, conversion index information can be generated, and the conversion index information can include an AMT horizontal flag indicating the above horizontal conversion kernel and an AMT vertical flag indicating the above vertical conversion kernel. On the other hand, the above conversion kernel can be referred to as a conversion type or a conversion core.
[0167] When the above-mentioned corrected conversion coefficient is derived, the encoding device can perform the selective conversion on the corrected conversion coefficient located in the upper left target region of the target block based on an association matrix including the above-mentioned corrected conversion matrix and an association vector for the above-mentioned corrected basis vector, so as to derive the conversion coefficient of the above-mentioned target block. The corrected conversion coefficients other than the corrected conversion coefficient located in the upper left region of the above-mentioned target block can be directly derived as the conversion coefficients of the above-mentioned target block.
[0168] Specifically, among the corrected conversion coefficients located in the above-mentioned upper left target region, the corrected conversion coefficients for the elements that are 1 in the above-mentioned association vector can be derived, and the conversion coefficients of the above-mentioned target block can be derived based on the derived corrected conversion coefficients and the above-mentioned corrected basis vectors. Here, the above-mentioned association vector for the above-mentioned corrected basis vector can include N elements, the above-mentioned N elements can include elements that are 1 and / or elements that are 0, and the number of the above-mentioned elements that are 1 can be A. Also, the above-mentioned corrected basis vector can include the above-mentioned A elements.
[0169] On the other hand, in one example, the above-mentioned corrected conversion matrix can include N corrected basis vectors, and the above-mentioned association matrix can include N association vectors. The above-mentioned association vectors can include the same number of elements that are 1, and the above-mentioned corrected basis vectors can all include the same number of elements. Alternatively, the above-mentioned association vectors may not include the same number of elements that are 1, and the above-mentioned corrected basis vectors may not all include the same number of elements.
[0170] Alternatively, in other examples, the modified transformation matrix can include R modified basis vectors, and the correlation matrix can include R correlation vectors. The R can be a reduced coefficient, and the R can be smaller than N. The correlation vectors can include the same number of elements of 1, and the modified basis vectors can all include the same number of elements. Alternatively, the correlation vectors may not include the same number of elements of 1, and the modified basis vectors may not all include the same number of elements.
[0171] On the other hand, the correlation vectors can be configured such that the elements of 1 are continuous. In this case, in one example, the information regarding the correlation vectors can be entropy encoded. For example, the information regarding the correlation vectors can include information indicating the starting point of the elements of 1 and information indicating the number of elements of 1. Alternatively, for example, the information regarding the correlation vectors can include information indicating the last position of the elements of 1 and information indicating the number of elements of 1.
[0172] Also, in other examples, the correlation vectors can be derived based on the size of the upper left target region. For example, based on the size of the upper left target region, the starting point of the elements of 1 of the correlation vectors and the number of the elements of 1 can be derived.
[0173] Alternatively, in other examples, the correlation vectors can be derived based on the intra prediction mode of the target block. For example, based on the intra prediction mode, the starting point of the elements of 1 of the correlation vectors and the number of the elements of 1 can be derived. Also, for example, based on whether the intra prediction mode is a non-directional intra prediction mode, the starting point of the elements of 1 of the correlation vectors and the number of the elements of 1 can be derived.
[0174] The encoding device encodes information regarding the conversion coefficients (S1330). The information regarding the conversion coefficients can include information regarding the size, position, etc. of the conversion coefficients. Also, as described above, the information regarding the correlation vector can be entropy-encoded. For example, the information regarding the correlation vector can include information indicating the start point of the elements of 1 and information indicating the number of elements of 1. Alternatively, for example, the information regarding the correlation vector can include information indicating the last position of the elements of 1 and information indicating the number of elements of 1.
[0175] The video information including the information regarding the conversion coefficients and / or the information regarding the correlation vector can be output in the form of a bitstream. Also, the video information can further include prediction information. The prediction information can include prediction mode information and information regarding motion information (for example, when inter prediction is applied) as information related to the prediction procedure.
[0176] The output bitstream can be transmitted to the decoding device via a storage medium or a network.
[0177] FIG. 12 schematically shows an encoding device that performs the video encoding method according to the present invention. The method disclosed in FIG. 11 can be performed by the encoding device disclosed in FIG. 12. Specifically, for example, the addition unit of the encoding device in FIG. 12 can perform S1100 in FIG. 11, the conversion unit of the encoding device can perform S1110, and the entropy encoding unit of the encoding device can perform S1120 to S1130. Also, although not shown, the process of deriving the prediction sample can be performed by the prediction unit of the encoding device.
[0178] Figure 13 schematically shows a video decoding method by a decoding apparatus according to the present invention. The method disclosed in Figure 13 can be performed by the decoding apparatus disclosed in Figure 2. Specifically, for example, S1300 to S1310 in Figure 13 can be performed by the entropy decoding unit of the decoding apparatus, S1320 can be performed by the inverse transform unit of the decoding apparatus, and S1330 can be performed by the addition unit of the decoding apparatus. Further, although not shown, the process of deriving a prediction sample can be performed by the prediction unit of the decoding apparatus.
[0179] The decoding apparatus derives the transform coefficients of the target block from the bitstream (S1300). The decoding apparatus can decode the information regarding the transform coefficients of the target block received through the bitstream to derive the transform coefficients of the target block. The received information regarding the transform coefficients of the target block can be shown as residual information.
[0180] The decoding apparatus derives residual samples for the target block based on a selective transform on the transform coefficients (S1310). The selective transform can be performed based on a modified transform matrix, and the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors can include a selected specific number of elements out of N elements. Further, the selective transform can be performed on the transform coefficients located in the upper left target region of the target block, and the N can be the number of transform coefficients located in the upper left target region or the like. Alternatively, the N can be the value obtained by multiplying the width and height of the upper left target region. For example, the N can be 16 or 64.
[0181] The decoding device can perform the selective conversion on the conversion coefficients located in the upper left target region of the target block based on an association matrix including the corrected transformation matrix and an association vector for the corrected basis vectors, and derive corrected conversion coefficients.
[0182] Specifically, conversion coefficients for the elements of 1 of the association vectors among the conversion coefficients located in the upper left target region can be derived, and corrected conversion coefficients can be derived based on the derived conversion coefficients and the corrected basis vectors. Here, the association vector for the corrected basis vectors can include N elements, the N elements can include elements of 1 and / or elements of 0, and the number of elements of 1 can be A. Also, the corrected basis vectors can include the A elements.
[0183] On the other hand, in one example, the corrected transformation matrix can include N corrected basis vectors, and the association matrix can include N association vectors. The association vectors can include the same number of elements of 1, and the corrected basis vectors can all include the same number of elements. Alternatively, the association vectors may not include the same number of elements of 1, and the corrected basis vectors may not all include the same number of elements.
[0184] Alternatively, in other examples, the modified transformation matrix can include R modified basis vectors, and the correlation matrix can include R correlation vectors. The R can be a reduced coefficient, and the R may be smaller than N. The correlation vectors can include the same number of elements of 1, and the modified basis vectors can all include the same number of elements. Alternatively, the correlation vectors may not include the same number of elements of 1, and the modified basis vectors may not all include the same number of elements.
[0185] On the other hand, the correlation vectors can be configured such that the elements of 1 are continuous. In this case, in one example, information regarding the correlation vectors can be obtained from a bitstream, and based on the information regarding the correlation vectors, the correlation vectors can be derived. For example, the information regarding the correlation vectors can include information indicating the starting point of the elements of 1 and information indicating the number of elements of 1. Alternatively, for example, the information regarding the correlation vectors can include information indicating the last position of the elements of 1 and information indicating the number of elements of 1.
[0186] Alternatively, in other examples, based on the size of the upper left segment target region, the correlation vectors can be derived. For example, based on the size of the upper left segment target region, the starting point of the elements of 1 of the correlation vectors and the number of the elements of 1 can be derived.
[0187] Alternatively, in other examples, based on the intra prediction mode of the target block, the correlation vectors can be derived. For example, based on the intra prediction mode, the starting point of the elements of 1 of the correlation vectors and the number of the elements of 1 can be derived. Also, for example, based on whether the intra prediction mode is a non-directional intra prediction mode, the starting point of the elements of 1 of the correlation vectors and the number of the elements of 1 can be derived.
[0188] When the above-mentioned corrected conversion coefficient is derived, the decoding device can perform a core conversion on the target block including the above-mentioned corrected conversion coefficient to derive the residual sample.
[0189] The core conversion on the target block can be performed as follows. The decoding device can obtain an AMT flag indicating whether Adaptive Multiple core Transform (AMT) is applied from the bitstream. When the value of the AMT flag is 0, the decoding device can derive DCT type 2 as the conversion kernel for the target block, and perform an inverse conversion on the target block including the above-mentioned corrected conversion coefficient based on the DCT type 2 to derive the residual sample.
[0190] When the value of the AMT flag is 1, the decoding device can configure a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, and based on the conversion index information obtained from the bitstream and the above-mentioned conversion subsets, derive the horizontal conversion kernel and the vertical conversion kernel. Based on the horizontal conversion kernel and the vertical conversion kernel, perform an inverse conversion on the target block including the above-mentioned corrected conversion coefficient to derive the residual sample. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel can include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Also, the conversion index information can include an AMT horizontal flag indicating one of the candidates included in the conversion subset for the horizontal conversion kernel and an AMT vertical flag indicating one of the candidates included in the conversion subset for the vertical conversion kernel. On the other hand, the conversion kernel can be referred to as a conversion type or a core.
[0191] The decoding device generates a restored picture based on the residual samples (S1320). The decoding device can generate a restored picture based on the residual samples. For example, the decoding device can perform inter prediction or intra prediction on a target block based on prediction information received through a bitstream to derive prediction samples, and can generate the restored picture through addition of the prediction samples and the residual samples. Thereafter, as described above, if necessary, in-loop filtering procedures such as deblock filtering, SAO, and / or ALF procedures can be applied to the restored picture to improve subjective / objective picture quality.
[0192] FIG. 14 schematically shows a decoding device that performs a video decoding method according to the present invention. The method disclosed in FIG. 13 can be performed by the decoding device disclosed in FIG. 14. Specifically, for example, the entropy decoding unit of the decoding device in FIG. 14 can perform S1300 in FIG. 13, the inverse transformation unit of the decoding device in FIG. 14 can perform S1310 in FIG. 13, and the addition unit of the decoding device in FIG. 16 can perform S1320 in FIG. 15. Also, although not shown, the process of deriving prediction samples can be performed by the prediction unit of the decoding device in FIG. 14.
[0193] According to the present invention described above, it is possible to reduce the amount of data that must be transferred for residual processing through efficient transformation, and to improve residual coding efficiency.
[0194] Also, according to the present invention, non-separable transformation can be performed based on a transformation matrix composed of basis vectors including a selected specific number of elements, through which the memory load and computational complexity for non-separable transformation can be reduced, and residual coding efficiency can be improved.
[0195] Further, according to the present invention, non-separable conversion can be performed based on a conversion matrix having a simplified structure, thereby reducing the amount of data that must be transferred for residual processing and increasing the residual coding efficiency.
[0196] In the above-described embodiments, the method has been described based on an order (flowchart) in a series of steps or blocks, but the present invention is not limited to the order of the steps, and a certain step can occur in a different step and a different order from those described above, or simultaneously. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps of the flowchart can be deleted without affecting the scope of the present invention.
[0197] The method according to the present invention described above can be embodied in software form, and the encoding device and / or decoding device according to the present invention can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, or the like.
[0198] In the present invention, when an embodiment is implemented by software, the aforementioned method can be implemented by modules (processes, functions, etc.) that perform the aforementioned functions. The modules can be stored in a memory and executed by a processor. The memory may be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (Application-Specific Integrated Circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (Read-Only Memory), a RAM (Random Access Memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present invention can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0199] In addition, the decoding device and the encoding device to which the present invention is applied can be included in a multimedia broadcast transceiver device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video intercom device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a picture phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, an OTT video (Over The Top video) device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0200] In addition, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the present invention can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial (general-purpose serial) Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy (registered trademark) disk, and an optical data storage device. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transfer through the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transferred via a wired or wireless communication network. Also, embodiments of the present invention can be embodied in a computer program product by program code, and the program code can be performed by a computer according to embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0201] In addition, the content streaming system to which the present invention is applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0202] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and is responsible for transferring this to the streaming server. In other examples, when multimedia input devices such as smartphones, cameras, and camcorders directly generate a bitstream, the encoding server can be omitted. The bitstream can be generated by an encoding method or a bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream in the process of transferring or receiving the bitstream.
[0203] The streaming server transfers multimedia data to a user device based on a user request (invitation) through a web server, and the web server serves as a medium for informing the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transfers multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server is responsible for controlling commands / responses between each device within the content streaming system.
[0204] The streaming server can receive content from a media storage and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0205] Examples of the user device may include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigators, slate PCs, tablet PCs, ULTRABOOKs (registered trademark), wearable devices (e.g., smartwatches, smart glasses, HMDs (Head Mounted Displays)), digital TVs, desktop computers, digital signage, and the like. Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.
Claims
**Claim 1** A video decoding method performed by a decoding apparatus, comprising: deriving inverse quantized coefficients of a target block from a bitstream; deriving modified transform coefficients by performing a secondary inverse transform according to a non-separable transform on the inverse quantized coefficients; deriving residual samples for the target block by performing a primary inverse transform on the modified transform coefficients; generating a restored picture by adding the residual samples for the target block and prediction samples for the target block; wherein the secondary inverse transform according to the non-separable transform is performed by using a transform matrix; the transform matrix includes basis vectors; each of the basis vectors includes M elements, and M is less than N; the number of the basis vectors is R, and R is less than N; N is the number of transform coefficients located in an upper left target region to which the non-separable transform is applied in the target block; in the case of the upper left target region having a size of 8×8, the N is equal to 64; the number of the inverse quantized coefficients to which the secondary inverse transform is applied is the same as the number R of the basis vectors. **Claim 2** A video encoding method performed by an encoding apparatus, comprising: deriving residual samples of a target block; deriving modified transform coefficients of the target block by performing a primary transform on the residual samples; deriving transform coefficients of the target block by performing a secondary transform according to a non-separable transform on the modified transform coefficients; encoding information regarding the transform coefficients. wherein the secondary transform according to the non-separable transform is performed by using a transform matrix; the transform matrix includes basis vectors; each of the basis vectors includes M elements, and M is less than N; the number of the basis vectors is R, and R is less than N; N is the number of modified transform coefficients located in an upper left target region to which the non-separable transform is applied in the target block; in the case of the upper left target region having a size of 8×8, the N is equal to 64; the number of the transform coefficients derived by performing the secondary transform is the same as the number R of the basis vectors in the transform matrix. **Claim 3** A transmission method for data including a bitstream related to an image, The step of generating the bitstream related to the image, wherein the bitstream is The step of deriving residual samples of a target block; The step of deriving modified transform coefficients of the target block by performing a primary transform on the residual samples; The step of deriving transform coefficients of the target block by performing a secondary transform according to a non-separable transform on the modified transform coefficients; The step of encoding information regarding the transform coefficients, and is generated by performing; The step of transmitting the data including the bitstream, and includes The secondary transform according to the non-separable transform is performed by using a transform matrix, The transform matrix includes basis vectors, Each of the basis vectors includes M elements, and M is less than N, The number of the basis vectors is R, and R is less than N, N is the number of modified transform coefficients located in the upper left target region to which the non-separable transform is applied in the target block, In the case of the upper left target region of size 8×8, N is equal to 64, The number of the transform coefficients derived by performing the secondary transform is the same as the number R of the basis vectors in the transform matrix.
Citation Information
Patent Citations
Image conversion method and device, and image inverse conversion method and device
JP2013542664A
Reduced size inverse transform for decoding and encoding
US20170034530A1
Non-separable secondary transform for video coding
WO2017058614A1