Image decoding and encoding methods, medium for storing bitstreams, and data transmission method
Through the selective transformation method, the modified transformation matrix and basis vector are used to solve the problem of inefficiency in high-resolution image coding, and the data volume reduction and encoding efficiency improvement are achieved.
Patent Information
- Application Number
- CN202211470299.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-21
- Filing Date
- 2018-12-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2038-12-21
AI Technical Summary
The prior art When transmitting and storing high-resolution and high-quality images, the increase in the amount of information leads to an increase in costs, and it is necessary to improve image encoding efficiency and transformation efficiency.
The selective transformation method is adopted to decode images through the modified transformation matrix and the base vector, including selecting a specific number of elements from N elements for transformation, reducing the residual processing data volume and improving coding efficiency.
By simplifying the transformation matrix and selective transformation of the structure, the storage burden and calculation complexity of inseparable transformation are reduced, and the residual coding efficiency is improved.
Smart Images

Figure CN115941940B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 201880086807.8 (International Application No.: PCT / KR2018 / 016437, Application Date: December 21, 2018, Invention Title: Image Coding Method Based on Selective Transformation and Apparatus for the Method). Technical Field
[0002] The present disclosure relates to image coding technology, and more particularly, to an image decoding method performed according to selective transformation in an image coding system and an apparatus for the image decoding method. Background Art
[0003] In various fields, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images is increasing. Since image data has high resolution and high quality, the amount of information or bits to be transmitted increases compared to conventional image data. Therefore, when transmitting image data using a medium such as a conventional wired / wireless broadband line or storing image data using an existing storage medium, the transmission cost and storage cost increase.
[0004] Therefore, there is a need for an efficient image compression technology for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images. Summary of the Invention
[0005] Technical Tasks
[0006] The technical problem to be solved by the present disclosure is to provide a method and apparatus for improving image coding efficiency.
[0007] Another technical problem to be solved by the present disclosure is to provide a method and apparatus for improving transformation efficiency.
[0008] Another technical problem to be solved by the present disclosure is to provide a method and apparatus for improving residual coding efficiency through transformation.
[0009] Another technical problem to be solved by the present disclosure is to provide an image coding method and apparatus based on selective transformation.
[0010] Solutions
[0011] According to an example of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes the following steps: deriving transform coefficients of a target block from a bitstream; deriving residual samples of the target block based on a selective transform of the transform coefficients; and generating a reconstructed picture based on the residual samples of the target block and prediction samples of the target block, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from among N elements.
[0012] According to another example of the present disclosure, a decoding device for performing image decoding is provided. The image decoding device includes: an entropy decoder that derives transform coefficients of a target block from a bitstream; an inverse transformer that derives residual samples of the target block based on a selective transform of the transform coefficients; and an adder that generates a reconstructed picture based on the residual samples of the target block and prediction samples of the target block, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from among N elements.
[0013] According to another example of the present disclosure, a video encoding method performed by an encoding device is provided. The method includes the following steps: deriving residual samples of a target block; deriving transform coefficients of the target block based on a selective transform of the residual samples; and encoding information about the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from among N elements.
[0014] According to yet another example of the present disclosure, a video encoding device is provided. The encoding device includes: an adder that derives residual samples of a target block; a transformer that derives transform coefficients of the target block based on a selective transform of the residual samples; and an entropy encoder that encodes information about the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from among N elements.
[0015] Technical effects
[0016] According to the present disclosure, through an efficient transform, it is possible to reduce the amount of data that must be transmitted for residual processing and increase the residual coding efficiency.
[0017] According to the present disclosure, a non-separable transform can be performed based on a transform matrix composed of basis vectors including a specific number of selected elements. Accordingly, the storage burden and computational complexity of the non-separable transform can be reduced and the residual coding efficiency can be increased.
[0018] According to the present disclosure, a non-separable transform can be performed based on a transform matrix having a simplified structure. Accordingly, the amount of data that must be transmitted for residual processing can be reduced and the residual coding efficiency can be increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic diagram illustrating the configuration of a video encoding device to which the present disclosure is applicable.
[0020] Figure 2 is a schematic diagram illustrating the configuration of a video decoding device to which the present disclosure is applicable.
[0021] Figure 3 Schematically represents a multiple transform technique according to the present disclosure.
[0022] Figure 4 Exemplarily shows intra prediction modes with 65 prediction directions.
[0023] Figures 5a to 5c is a diagram for explaining selective transform according to an example of the present disclosure.
[0024] Figure 6 Schematically represents a multiple transform technique in which selective transform is applied as a secondary transform.
[0025] Figure 7 is a diagram for explaining the arrangement of transform coefficients based on a target block according to an example of the present disclosure.
[0026] Figure 8 Represents an example of deriving transform coefficients by combining a simplified transform and a selective transform with each other.
[0027] Figure 9 Represents an example of deriving transform coefficients by selective transform.
[0028] Figure 10 Represents an example of performing selective transform by deriving a correlation vector based on two factors of a correlation vector.
[0029] Figure 11 Schematically represents an image encoding method performed by an encoding device according to the present disclosure.
[0030] Figure 12 Schematically represents an encoding device that performs an image encoding method according to the present disclosure.
[0031] Figure 13Schematically represents an image decoding method performed by a decoding device according to the present disclosure.
[0032] Figure 14 Schematically represents a decoding device that performs an image decoding method according to the present disclosure. Detailed implementation manners
[0033] The present disclosure can be modified in various forms, and its specific implementation manners will be described and illustrated in the accompanying drawings. However, these implementation manners are not intended to limit the present disclosure. The terms used in the following description are only for describing specific implementation manners and are not intended to limit the present disclosure. Singular expressions include plural expressions as long as they are clearly and differently understood. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus it should be understood that there is no exclusion of the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0034] In addition, for the purpose of conveniently explaining different specific functions, the elements in the drawings described in the present disclosure are independently drawn, which does not mean that these elements are implemented by independent hardware or independent software. For example, two or more of these elements can be combined to form a single element, or one element can be divided into multiple elements. Without departing from the concept of the present disclosure, the implementation manners in which elements are combined and / or divided belong to the present disclosure.
[0035] Hereinafter, the implementation manners of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.
[0036] In addition, the present disclosure relates to video / image coding. For example, the methods / implementation manners disclosed in the present disclosure can be applied to the methods disclosed in the Versatile Video Coding (VVC) standard or the next-generation video / image coding standard.
[0037] In the present disclosure, generally, a picture means a unit representing an image in a specific time slot, and a slice is a unit that forms a part of a picture in coding. A picture can include multiple slices, and in some cases, pictures and slices can be used in a mixed manner.
[0038] A pixel or pel can mean the smallest unit that constitutes a picture (or image). In addition, the term "sample" can be used corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0039] A unit represents a basic unit of image processing. The unit may include at least one of a specific area of a picture and information related to the corresponding area. In some cases, the unit may be used in a mixed manner with blocks or regions. In general, an M×N block may represent a set of samples or transform coefficients including M columns and N rows.
[0040] Figure 1 is a diagram briefly illustrating a video encoding device to which the present disclosure is applicable.
[0041] Referring to Figure 1 , the video encoding device 100 may include a picture splitter 105, a predictor 110, a residual processor 120, an entropy encoder 130, an adder 140, a filter 150, and a memory 160. The residual processor 120 may include a subtractor 121, a transformer 122, a quantizer 123, a rearranger 124, an inverse quantizer 125, and an inverse transformer 126.
[0042] The picture splitter 105 may split an input picture into at least one processing unit.
[0043] For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBT) structure. For example, a coding unit may be split into multiple coding units with a deeper depth based on a quadtree structure and / or a binary tree structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the finally non-split coding unit. In this case, the largest coding unit may be used as the final coding unit based on the encoding efficiency according to the image characteristics, or the coding unit may be recursively split into coding units with a deeper depth when necessary, and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include processes of prediction, transformation, and reconstruction to be described later.
[0044] In another example, the processing unit may include an encoding unit (CU), a prediction unit (PU), or a transform unit (TU). The encoding unit may be divided from a largest coding unit (LCU) into deeper coding units according to a quadtree structure. In this case, the largest coding unit may be directly used as a final coding unit based on encoding efficiency or the like according to image characteristics, or the coding unit may be recursively divided into deeper coding units when necessary, and a coding unit having an optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit may not be divided into coding units smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that is divided or partitioned into a prediction unit or a transform unit. The prediction unit is a unit divided from the coding unit and may be a unit for sample prediction. Here, the prediction unit may be divided into sub-blocks. The transform unit may be separated from the coding unit according to a quadtree structure, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transform unit may be referred to as a transform block (TB). The prediction block or prediction unit may refer to a specific region in the form of a block in a picture and includes an array of prediction samples. In addition, the transform block or transform unit may refer to a specific region in the form of a block in a picture and includes an array of transform coefficients or residual samples.
[0045] The predictor 110 may perform prediction on a processing target block (hereinafter, the current block) and may generate a prediction block including prediction samples for the current block. The unit for performing prediction in the predictor 110 may be a coding block, or may be a transform block, or may be a prediction block.
[0046] The predictor 110 may determine whether to apply intra prediction or inter prediction to the current block. For example, the predictor 110 may determine whether to apply intra prediction or inter prediction in units of CUs.
[0047] In the case of intra prediction, the predictor 110 may derive a prediction sample of a current block based on reference samples outside the current block in a picture (hereinafter, the current picture) to which the current block belongs. In this case, the predictor 110 may derive a prediction sample based on the average or interpolation of neighboring reference samples of the current block (case (i)), or may derive a prediction sample based on a reference sample in which a prediction sample among the neighboring reference samples of the current block exists in a specific (prediction) direction (case (ii)). Case (i) may be referred to as a non-directional mode or a non-angular mode, and case (ii) may be referred to as a directional mode or an angular mode. In intra prediction, as an example, the prediction mode may include 33 directional modes and at least two non-directional modes. The non-directional modes may include a DC mode and a planar mode. The predictor 110 may determine a prediction mode to be applied to the current block by using a prediction mode applied to a neighboring block.
[0048] In the case of inter prediction, the predictor 110 may derive a prediction sample of the current block based on samples specified by a motion vector on a reference picture. The predictor 110 may derive a prediction sample of the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the case of the skip mode and the merge mode, the predictor 110 may use motion information of a neighboring block as motion information of the current block. In the case of the skip mode, different from the merge mode, a difference (residual) between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, a motion vector of a neighboring block is used as a motion vector predictor, and thus is used as a motion vector predictor of the current block to derive a motion vector of the current block.
[0049] In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in a reference picture. The reference picture including the temporal neighboring blocks may also be referred to as a collocated picture (colPic). Motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and then output in the form of a bitstream.
[0050] When using motion information of a temporal neighboring block in the skip mode and the merge mode, the highest picture in the reference picture list may be used as a reference picture. The reference pictures included in the reference picture list may be aligned based on a picture order count (POC) difference between the current picture and the corresponding reference picture. The POC corresponds to a display order and may be distinguished from an encoding order.
[0051] The subtractor 121 generates a residual sample, which is a difference between the original sample and the prediction sample. If the skip mode is applied, the residual sample may not be generated as described above.
[0052] The transformer 122 transforms residual samples in units of transform blocks to generate transform coefficients. The transformer 122 may perform the transformation based on the size of the corresponding transform block and the prediction mode applied to a prediction block or a coded block that spatially overlaps with the transform block. For example, if intra prediction is applied to a prediction block or a coded block that overlaps with the transform block, a discrete sine transform (DST) transform kernel may be used to transform the residual samples, the transform block being a 4×4 residual array, and in other cases a discrete cosine transform (DCT) transform kernel is used to transform it.
[0053] The quantizer 123 may quantize the transform coefficients to generate quantized transform coefficients.
[0054] The re-arranger 124 re-arranges the quantized transform coefficients. The re-arranger 124 may re-arrange the quantized transform coefficients in block form into a one-dimensional vector by a coefficient scanning method. Although the re-arranger 124 is described as a separate component, the re-arranger 124 may be part of the quantizer 123.
[0055] The entropy encoder 130 may perform entropy coding on the quantized transform coefficients. The entropy coding may include coding methods such as (for example) exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 130 may code, together or separately, information necessary for video reconstruction other than the quantized transform coefficients (for example, the values of syntax elements, etc.). The entropy-coded information may be sent or stored in the form of a bitstream in units of NAL (network abstraction layer).
[0056] The de-quantizer 125 de-quantizes the values (transform coefficients) quantized by the quantizer 123, and the inverse transformer 126 inverse-transforms the values de-quantized by the de-quantizer 125 to generate residual samples.
[0057] The adder 140 adds the residual samples and the prediction samples to reconstruct a picture. The residual samples and the prediction samples may be added in units of blocks to generate a reconstructed block. Although the adder 140 is described as a separate component, the adder 140 may be part of the predictor 110. In addition, the adder 140 may be referred to as a reconstructor or a reconstructed block generator.
[0058] Filter 150 may apply deblocking filtering and / or sample adaptive offset to the reconstructed picture. Artifacts at block boundaries in the reconstructed picture or distortions in quantization may be corrected by deblocking filtering and / or sample adaptive offset. After deblocking filtering is completed, sample adaptive offset may be applied on a sample-by-sample basis. Filter 150 may apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture to which deblocking filtering and / or sample adaptive offset have been applied.
[0059] Memory 160 may store the reconstructed picture (decoded picture) or information required for encoding / decoding. Here, the reconstructed picture may be the reconstructed picture filtered by Filter 150. The stored reconstructed picture may be used as a reference picture for (inter-frame) prediction of other pictures. For example, Memory 160 may store a (reference) picture for inter-frame prediction. Here, the picture for inter-frame prediction may be specified according to a reference picture set or a reference picture list.
[0060] Figure 2 is a schematic diagram illustrating the configuration of a video decoding device to which the present disclosure is applicable.
[0061] Refer to Figure 2 , video decoding device 200 includes image decoder 210, residual processor 220, predictor 230, adder 240, filter 250, and memory 260. Here, residual processor 220 may include reorderer 221, dequantizer 222, and inverse transformer 223.
[0062] When a bitstream including video information is input, video decoding device 200 may reconstruct video corresponding to the processing of the video information in the video encoding device.
[0063] For example, video decoding device 200 may perform video decoding using the processors applied in the video encoding device. Therefore, the processing unit blocks for video decoding may be, for example, coding units, or may be, for another example, coding units, prediction units, or transform units. Coding units may be split from the largest coding unit according to a quadtree structure and / or a binary tree structure.
[0064] In some cases, prediction units and transform units may also be used, and in such cases, the prediction block is a block derived or split from the coding unit and may be a unit for sample prediction. Here, the prediction unit may be divided into sub-blocks. The transform unit may be split from the coding unit according to a quadtree structure, and the transform unit may be a unit for deriving transform coefficients or a unit for deriving a residual signal from the transform coefficients.
[0065] The entropy decoder 210 can parse the bitstream to output the information required for video reconstruction or picture reconstruction. For example, the entropy decoder 210 can decode the information in the bitstream based on coding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for video reconstruction and the quantization values of the transform coefficients for the residuals.
[0066] More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, after determining the context model, the CABAC entropy decoding method can use the information of the symbols / bins decoded by the context model for the next symbol / bin to update the context model.
[0067] The information for prediction among the information decoded in the entropy decoder 210 can be provided to the predictor 230, and the residual values, i.e., the quantized transform coefficients for which the entropy decoder 210 has performed entropy decoding, can be input to the reorderer 221.
[0068] The reorderer 221 can reorder the quantized transform coefficients into a two-dimensional block form. The reorderer 221 can perform a reordering corresponding to the coefficient scan performed by the encoding device. Although the reorderer 221 is described as a separate component, the reorderer 221 can be a part of the inverse quantizer 222.
[0069] The inverse quantizer 222 can inverse-quantize the quantized transform coefficients based on the (inverse) quantization parameters to output the transform coefficients. In this case, the information for deriving the quantization parameters can be signaled from the encoding device.
[0070] The inverse transformer 223 can inverse-transform the transform coefficients to derive the residual samples.
[0071] The predictor 230 can perform prediction on the current block and can generate a prediction block including the prediction samples of the current block. The unit of prediction performed in the predictor 230 can be an encoding block or can be a transform block or can be a prediction block.
[0072] The predictor 230 may determine whether to apply intra prediction or inter prediction based on the information for prediction. In this case, the unit for determining which of intra prediction and inter prediction will be used may be different from the unit for generating prediction samples. Additionally, in both inter prediction and intra prediction, the units for generating prediction samples may also be different. For example, it may be determined which of inter prediction and intra prediction will be applied on a CU-by-CU basis. Additionally, for example, in inter prediction, prediction samples may be generated by determining a prediction mode on a PU-by-PU basis, while in intra prediction, prediction samples may be generated on a TU-by-TU basis by determining a prediction mode on a PU-by-PU basis.
[0073] In the case of intra prediction, the predictor 230 may derive the prediction samples of the current block based on neighboring reference samples in the current picture. The predictor 230 may derive the prediction samples of the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined by using the intra prediction mode of neighboring blocks.
[0074] In the case of inter prediction, the predictor 230 may derive the prediction samples of the current block based on the samples specified in the reference picture according to the motion vector. The predictor 230 may use one of the skip mode, the merge mode, and the MVP mode to derive the prediction samples of the current block. Here, the motion information required for inter prediction of the current block provided by the video coding device, such as the motion vector and the information for the reference picture index, may be obtained or derived based on the information for prediction.
[0075] In the skip mode and the merge mode, the motion information of neighboring blocks may be used as the motion information of the current block. Here, the neighboring blocks may include spatially neighboring blocks and temporally neighboring blocks.
[0076] The predictor 230 may use the motion information of available neighboring blocks to construct a merge candidate list, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled by the coding device. The motion information may include the motion vector and the reference picture. When using the motion information of temporally neighboring blocks in the skip mode and the merge mode, the highest picture in the reference picture list may be used as the reference picture.
[0077] In the case of the skip mode, different from the merge mode, the difference (residual) between the prediction samples and the original samples is not sent.
[0078] In the case of the MVP mode, the motion vector of neighboring blocks may be used as a motion vector predictor to derive the motion vector of the current block. Here, the neighboring blocks may include spatially neighboring blocks and temporally neighboring blocks.
[0079] When applying the merge mode, for example, the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col blocks as temporally neighboring blocks can be used to generate a merge candidate list. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The above-mentioned information for prediction can include a merge index, which indicates the candidate block with the best motion vector selected from the candidate blocks included in the merge candidate list. Here, the predictor 230 can use the merge index to derive the motion vector of the current block.
[0080] When applying the MVP (Motion Vector Prediction) mode as another example, the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col blocks as temporally neighboring blocks can be used to generate a motion vector predictor candidate list. That is, the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col blocks as temporally neighboring blocks can be used as motion vector candidates. The above-mentioned information for prediction can include a predicted motion vector index indicating the best motion vector selected from the motion vector candidates included in the list. Here, the predictor 230 can use the motion vector index to select the predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list. The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode the MVD, and output the encoded MVD in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the predictor 230 can obtain the motion vector included in the information for prediction, and derive the motion vector of the current block by adding the motion vector difference to the motion vector predictor. Additionally, the predictor can obtain or derive a reference picture index indicating the reference picture from the above-mentioned information for prediction.
[0081] The adder 240 can add the residual samples and the predicted samples to reconstruct the current block or the current picture. The adder 240 can reconstruct the current picture by adding the residual samples and the predicted samples in block units. When applying the skip mode, no residual is sent, so the predicted samples can become the reconstructed samples. Although the adder 240 is described as a separate component, the adder 240 can be a part of the predictor 230. In addition, the adder 240 can be referred to as a reconstructor or a reconstructed block generator.
[0082] The filter 250 can apply deblocking filter, sample adaptive offset, and / or ALF to the reconstructed picture. Here, after the deblocking filter, the sample adaptive offset can be applied on a sample-by-sample basis. ALF can be applied after the deblocking filter and / or the application of the sample adaptive offset.
[0083] The memory 260 may store the reconstructed picture (decoded picture) or the information required for decoding. Here, the reconstructed picture may be the reconstructed picture filtered by the filter 250. For example, the memory 260 may store the picture for inter prediction. Here, the picture for inter prediction may be specified according to the reference picture set or the reference picture list. The reconstructed picture may be used as a reference picture for other pictures. The memory 260 may output the reconstructed pictures in the output order.
[0084] In addition, the lower-frequency transform coefficients of the residual block of the current block may be derived through the above-mentioned transform, and zero tails may be derived at the end of the residual block.
[0085] Specifically, the transform may consist of two main processes, and these main processes may include a kernel transform and a secondary transform. The transform including the kernel transform and the secondary transform may be expressed as a multiple transform technique.
[0086] Figure 3 Schematically represents the multiple transform technique according to the present disclosure.
[0087] Referring to Figure 3 , the transformer may correspond to the transformer in the encoding device above Figure 1 , and the inverse transformer may correspond to the inverse transformer in the encoding device above Figure 1 or Figure 2 the inverse transformer in the decoding device above.
[0088] The transformer may derive (primary) transform coefficients (S310) by performing a primary transform based on the residual samples (residual sample array) in the residual block. In this regard, the primary transform may include an adaptive multi-kernel transform (AMT). The adaptive multi-kernel transform may be expressed as a multi-transform set (MTS).
[0089] The adaptive multi-kernel transform may represent a method of additionally using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1 for transformation. That is, the adaptive multi-kernel transform may represent a transform method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this regard, from the perspective of the transformer, the primary transform coefficients may be referred to as temporary transform coefficients.
[0090] In other words, if the existing transformation method is applied, the transform coefficients can be generated by applying a transformation from the spatial domain to the frequency domain based on DCT type 2 for the residual signal (or residual block). In contrast, if the adaptive multi-core transformation is applied, the transform coefficients (or primary transform coefficients) can be generated by applying a transformation from the spatial domain to the frequency domain based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1 for the residual signal (or residual block). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transformation types, transform kernels, or transform cores.
[0091] As a reference, the DCT / DST transform types can be defined based on basis functions, and the basis functions can be represented as in the following table.
[0092] [Table 1]
[0093]
[0094] If the adaptive multi-core transformation is performed, the vertical transform kernel and the horizontal transform kernel for the target block can be selected from among the transform kernels, the vertical transform for the target block can be performed based on the vertical transform kernel, and the horizontal transform for the target block can be performed based on the horizontal transform kernel. Here, the horizontal transform can represent the transformation of the horizontal component of the target block, and the vertical transform can represent the transformation of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the transform index indicating the prediction mode and / or transform subset of the target block (CU or sub-block) including the residual block.
[0095] For example, when both the width and height of the target block are less than or equal to 64, the adaptive multi-core transformation can be applied, and it can be determined based on the CU-level flag related to whether to apply the adaptive multi-core transformation to the target block. Specifically, if the CU-level flag is 0, the above-described existing transformation method can be applied. That is, if the CU-level flag is 0, the transform coefficients can be generated by applying a transformation from the spatial domain to the frequency domain based on DCT type 2 for the residual signal (or residual block), and the transform coefficients can be encoded. Further, here, the target block can be a CU. If the CU-level flag is 0, the adaptive multi-core transformation can be applied to the target block.
[0096] Alternatively, if the target block to which the adaptive multi-core transform is applied is a luma block, two additional flags may be signaled, and the vertical transform kernel and the horizontal transform kernel may be selected based on the flags. The flag for the vertical transform kernel may be represented as the AMT vertical flag, and AMT_TU_vertical_flag (or EMT_TU_vertical_flag) may represent the syntax element of the AMT vertical flag. The flag for the horizontal transform kernel may be represented as the AMT horizontal flag, and AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) may represent the syntax element of the AMT horizontal flag. The AMT vertical flag may indicate one of the transform kernel candidates included in the transform subset of the vertical transform kernel, and the transform kernel candidate indicated by the AMT vertical flag may be derived as the vertical transform kernel of the target block. Additionally, the AMT horizontal flag may indicate one of the transform kernel candidates included in the transform subset of the horizontal transform kernel, and the transform kernel candidate indicated by the AMT horizontal flag may be derived as the horizontal transform kernel of the target block. Furthermore, the AMT vertical flag may be represented as the MTS vertical flag, and the AMT horizontal flag may be represented as the MTS horizontal flag.
[0097] In addition, three transform subsets may be predefined, and based on the intra prediction mode applied to the target block, one of the transform subsets may be derived as the transform subset of the vertical transform kernel. Additionally, based on the intra prediction mode applied to the target block, one of the transform subsets may be derived as the transform subset of the horizontal transform kernel. For example, the predefined transform subsets may be derived as shown in the following table.
[0098] [Table 2]
[0099] Transformation set Transformation candidate 0 DST-VII, DCT-VIII 1 DST-VII, DST-I 2 DST-VII, DCT-VIII
[0100] Referring to Figure 2 , the transform subset with an index value of 0 may represent the transform subset including DST type 7 and DCT type 8 as transform kernel candidates, the transform subset with an index value of 1 may represent the transform subset including DST type 7 and DST type 1 as transform kernel candidates, and the transform subset with an index value of 2 may represent the transform subset including DST type 7 and DCT type 8 as transform kernel candidates.
[0101] The transform subset of the vertical transform kernel and the transform subset of the horizontal transform kernel derived based on the intra prediction mode applied to the target block may be derived as shown in the following table.
[0102] [Table 3]
[0103]
[0104] Here, V represents a transform subset of the vertical transform kernel, and H represents a transform subset of the horizontal transform kernel.
[0105] If the value of the AMT flag (or EMT_CU_flag) of the target block is 1, then as shown in Table 3, the transform subset of the vertical transform kernel and the transform subset of the horizontal transform kernel can be derived based on the intra prediction mode of the target block. Thereafter, among the transform kernel candidates included in the transform subset of the vertical transform kernel indicated by the AMT vertical flag of the target block, the transform kernel candidate can be derived as the vertical transform kernel of the target block, and among the transform kernel candidates included in the transform subset of the horizontal transform kernel indicated by the AMT horizontal flag of the target block, the transform kernel candidate can be derived as the horizontal transform kernel of the target block. In addition, the AMT flag can be represented as the MTS flag.
[0106] As a reference, in the example, the intra prediction mode may include two non - directional (or non - angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non - directional intra prediction modes may include the 0 - th plane intra prediction mode and the 1 - st DC intra prediction mode, and the directional intra prediction modes may include 65 intra prediction modes between the 2 - nd intra prediction mode and the 66 - th intra prediction mode. However, this is an example, and the present disclosure can be applied to cases where there are different numbers of intra prediction modes. In addition, depending on the situation, the 67 - th intra prediction mode may also be used, and the 67 - th intra prediction mode may represent a linear model (LM) mode.
[0107] Figure 4 Illustratively represent the 65 prediction - direction intra - directional modes.
[0108] Refer to Figure 4 , based on the 34 - th intra prediction mode with the upper - left diagonal prediction direction, the intra prediction modes with horizontal directivity and the intra prediction modes with vertical directivity can be classified. Figure 4 In, H and V respectively mean horizontal directivity and vertical directivity, and the numbers - 32 to 32 indicate the shift in units of 1 / 32 at the sample grid position. The 2 - nd to 33 - rd intra prediction modes have horizontal directivity, and the 34 - th to 66 - th intra prediction modes have vertical directivity. The 18 - th intra prediction mode and the 50 - th intra prediction mode respectively represent the horizontal intra prediction mode and the vertical intra prediction mode, and the 2 - nd intra prediction mode can be called the lower - left diagonal intra prediction mode; the 34 - th intra prediction mode can be called the upper - left diagonal intra prediction mode; and the 66 - th intra prediction mode can be called the upper - right diagonal intra prediction mode.
[0109] The transformer can derive the (secondary) transformation coefficients by performing a secondary transformation based on the (primary) transformation coefficients (S320). If the primary transformation is a transformation from the spatial domain to the frequency domain, the secondary transformation can be regarded as a transformation from the frequency domain to the frequency domain. The secondary transformation can include an inseparable transformation. In this case, the secondary transformation can be referred to as an inseparable secondary transformation (NSST) or a mode-dependent inseparable secondary transformation (MDNSST). The inseparable secondary transformation can represent a transformation for generating transformation coefficients (or secondary transformation coefficients) for the residual signal by performing a secondary transformation on the (primary) transformation coefficients derived by the primary transformation based on an inseparable transformation matrix. At this time, instead of applying vertical and horizontal transformations to the (primary) transformation coefficients separately (or not applying horizontal and vertical transformations independently), the transformation can be applied at once based on the inseparable transformation matrix. In other words, the inseparable secondary transformation can represent a transformation method for generating transformation coefficients (or secondary transformation coefficients) not by separating the vertical and horizontal components of the (primary) transformation coefficients but by transforming them together based on the inseparable transformation matrix. The inseparable secondary transformation can be applied to the upper left region of a block configured with the (primary) transformation coefficients (hereinafter, may be referred to as a transformation coefficient block or a target block). For example, if both the width (W) and height (H) of the transformation coefficient block are equal to or greater than 8, an 8×8 inseparable secondary transformation can be applied to the upper left 8×8 region of the transformation coefficient block (hereinafter, referred to as the upper left target region). Additionally, if both the width (W) and height (H) of the transformation coefficient block are equal to or greater than 4 and the width (W) and height (H) of the transformation coefficient block are less than 8, a 4×4 inseparable secondary transformation can be applied to the upper left min(8, W)×min(8, H) region of the transformation coefficient block.
[0110] Specifically, for example, if a 4×4 input block is used, the inseparable secondary transformation can be performed as follows.
[0111] The 4×4 input block X can be represented as follows.
[0112] [Equation 1]
[0113]
[0114] If X is represented in vector form, the vector can be represented as follows
[0115] [Equation 2]
[0116]
[0117] In this case, the secondary inseparable transformation can be calculated as follows.
[0118] [Equation 3]
[0119]
[0120] Among them, represents the transform coefficient vector, and T represents a 16×16 (non-separable) transform matrix.
[0121] Through Equation 3 above, a 16×1 transform coefficient vector can be derived and can be reorganized into 4×4 blocks in a scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is an example, and hypercube-Givens transform (HyGT), etc. can be used to calculate the non-separable quadratic transform in order to reduce the computational complexity of the non-separable quadratic transform.
[0122] In addition, in the non-separable quadratic transform, a transform kernel (or transform core, transform type) can be selected such that it can be pattern-dependent. In this case, the pattern can include an intra prediction pattern and / or an inter prediction pattern.
[0123] As described above, the non-separable quadratic transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. That is, the non-separable quadratic transform can be performed based on an 8×8 sub-block size or a 4×4 sub-block size. For example, in order to select a pattern-dependent transform kernel, 35 sets of 3 non-separable quadratic transform kernels for the non-separable quadratic transform can be configured for both the 8×8 sub-block size and the 4×4 sub-block size. That is, 35 transform sets can be configured for the 8×8 sub-block size, and 35 transform sets can be configured for the 4×4 sub-block size. In this case, each of the 35 transform sets for the 8×8 sub-block size can contain three 8×8 transform kernels, and in this case, each of the 35 transform sets for the 4×4 sub-block size can include three 4×4 transform kernels. However, the transform sub-block size, the number of sets, and the number of transform kernels in the sets are examples, and any other size other than 8×8 or 4×4 can be used, or n sets can be configured, and each set can include k kernels.
[0124] The transform set can be referred to as an NSST set, and the transform kernel in the NSST set can be referred to as an NSST kernel. For example, the selection of a specific set among the transform sets can be performed based on the intra prediction pattern of the target block (CU or sub-block).
[0125] In this case, for example, the mapping between the 35 transform sets and the intra prediction pattern can be represented as in the following table. For reference, if the LM mode is applied to the target block, the quadratic transform may not be applied to the target block.
[0126] [Table 4]
[0127]
[0128] In addition, if it is determined to use a specific set, one of the k transform kernels in the specific set can be selected by the non-separable quadratic transform index. The encoding device can derive the non-separable quadratic transform index indicating the specific transform kernel based on rate distortion (RD) checking, and can signal the non-separable quadratic transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable quadratic transform index. For example, the NSST index value 0 can indicate the first non-separable quadratic transform kernel, the NSST index value 1 can indicate the second non-separable quadratic transform kernel, and the NSST index value 2 can indicate the third non-separable quadratic transform kernel. Alternatively, the NSST index value 0 can indicate not applying the first non-separable quadratic transform to the target block, and the NSST index values 1 to 3 can indicate three transform kernels.
[0129] Return reference Figure 3 , and the transformer can perform a non-separable quadratic transform based on the selected transform kernel and can obtain (quadratic) transform coefficients. As described above, the transform coefficients can be derived as the transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0130] In addition, if the quadratic transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as the transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device and transmitted to the dequantizer / inverse transformer in the encoding device.
[0131] The inverse transformer can perform a series of processes in an order opposite to the order in which they have been performed in the above-mentioned transformer. The inverse transformer can receive the (dequantized) transform coefficients, and derive the (primary) transform coefficients by performing a quadratic (inverse) transform (step S350), and can obtain the residual block (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S360). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0132] In addition, as described above, if the quadratic (inverse) transform is omitted, the (dequantized) transform coefficients can be received, the primary (separable) transform can be performed, and the residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0133] Figures 5a to 5c It is a diagram for explaining a selective transformation according to an example of the present disclosure.
[0134] In this specification, the term "target block" may mean a current block or a residual block on which encoding is performed.
[0135] Figure 5a It shows an example of deriving transform coefficients through a transformation.
[0136] The transformation in video coding may represent a process of generating a transform coefficient vector C of an input vector R by transforming the input vector R based on a transformation matrix as shown in Figure 5a . The input vector R may represent primary transform coefficients. Alternatively, the input vector R may represent a residual vector, i.e., residual samples. In addition, the transform coefficient vector C may be represented as an output vector C.
[0137] Figure 5b It shows a specific example of deriving transform coefficients through a transformation. Figure 5b Specifically, it represents Figure 5a the transformation process shown in Figure 3 . As described above in 2 ×M 2 NSST, after dividing an M×M block into block data of transform coefficients obtained by applying a primary transformation, an M 2 ×M Figure 5b NSST can be performed on each M×M block. For example, M may be 4 or 8, but is not limited thereto. M N may be N. In this case, as shown in N , the input vector R may be a (1×N)-dimensional vector including (primary) transform coefficients r1 to r N , and the transform coefficient vector C may be an (N×1)-dimensional vector including transform coefficients c1 to c N . That is, the input vector R may include N (primary) transform coefficients r1 to r
[0138]
[0139] Figure 5b N N N Figure 5b N N may be (1×N)-dimensional vectors. That is, the basis vectors B1 to B NThe size can be (1×N). A transform coefficient vector C can be generated based on each of the (primary) transform coefficients of the input vector R and the basis vectors of the transform matrix. For example, the inner product between the input vector and each of the basis vectors can be derived as the transform coefficient vector C.
[0140] In addition, in the above transformation, two main problems arise. Specifically, high computational complexity related to the number of multiplications and additions required to generate the output vector and the storage requirements for storing the generated coefficients may arise as the main problems.
[0141] For example, the computational complexity and storage requirements for separable and non-separable transforms can be derived as shown in the following table.
[0142] [Table 5]
[0143] Separable transform (NxN) Non-separable transform (NxN) Storage <![CDATA[N 2 > <![CDATA[N 4 > Multiplication <![CDATA[2N 3 > <![CDATA[N 4 >
[0144] Referring to Table 5, the memory required to store the coefficients generated by the separable transform can be N 2 , and the number of calculations can be 2N 3 . The number of calculations indicates the computational complexity. In addition, the memory required to store the coefficients generated by the non-separable transform can be N 4 , and the number of calculations can be N 4 . The number of calculations indicates the computational complexity. That is, the more the number of calculations, the higher the computational complexity, and the fewer the number of calculations, the lower the computational complexity.
[0145] As shown in Table 5, compared with the separable transform, the storage requirements and the number of calculations for the non-separable transform increase significantly. In addition, as the size of the target block for which the non-separable transform is performed increases, that is, as N gets larger, the difference between the storage requirements and the number of calculations for the separable transform and the storage requirements and the number of calculations for the non-separable transform becomes larger.
[0146] The non-separable transform provides a better coding gain compared to the separable transform, but as shown in Table 5, due to the computational complexity of the non-separable transform, the non-separable transform is not used in existing video coding standards, and since the computational complexity of the separable transform also increases as the size of the target block increases, in the existing HEVC standard, it is proposed to use the separable transform only in target blocks with a size of 32×32 or smaller.
[0147] Therefore, the present disclosure proposes selective transformation. Selective transformation can significantly reduce the computational complexity and storage requirements, thereby achieving effects such as increased efficiency of computationally intensive transform blocks and improved coding efficiency. That is, selective transformation can be used to solve the computational complexity problems that occur during non-separable transformation or transformation of large-sized blocks. Selective transformation can be used for any type of transformation such as primary transformation (or can be referred to as kernel transformation), secondary transformation, etc. For example, selective transformation can be applied as the kernel transformation of an encoding device / decoding device and can have the effect of significantly reducing the encoding time / decoding time.
[0148] Figure 5c An example of deriving transform coefficients through selective transformation is shown. Selective transformation can mean a transformation performed on a target block based on a transformation matrix including basis vectors with a selective number of elements.
[0149] Simplified transformation is a method as follows: Since there may be redundant or unimportant elements among the N elements of the basis vectors that can include a transformation matrix, this method is proposed with the motivation of being able to reduce the computational complexity and storage requirements by excluding such elements. For example, referring to Figure 5c , among the N elements of basis vector B1, Z0 elements may not be important elements, and in this case, a truncated basis vector B1 including only N1 elements can be derived. Here, N1 can be N - Z0. The truncated basis vector B1 can be represented as the modified basis vector B1.
[0150] Referring to Figure 5c , if the modified basis vector B1 is applied to the input vector R as part of a transformation matrix, the transform coefficient C1 can be derived. According to experimental results, it is observed that the transform coefficient C1 is the same value as the transform coefficient C1 derived by applying the existing basis vector B1, which is part of the transformation matrix, to the input vector R. That is, by deriving the result assuming that the unimportant elements of each basis vector are 0, the number of unnecessary multiplications can be significantly reduced, and the difference in the results is small. Additionally, then, the number of elements (i.e., elements of the transformation matrix) that must be stored for this calculation can be reduced.
[0151] To define the positions of the elements among the elements of the unimportant (or meaningful) basis vectors, an association vector is proposed. To derive the modified N1 - dimensional basis vector B1, a (1×N) - dimensional association vector A1 can be considered. That is, to derive the modified basis vector B1 of size 1×N1 (i.e., the modified basis vector B1 including N1 elements), a 1×N - dimensional association vector A1 can be considered.
[0152] Referring to Figure 5c , by applying to the input vector R for basis vectors B1 to BN Values derived from the associated vectors of each of these can be transmitted to the basis vectors. Thus, elements of the basis vectors can be calculated using only some of the elements of the input vector R. Specifically, the associated vectors can include 0s and 1s, and an operation is performed such that elements selected from among the elements of the input vector R are multiplied by 1, and elements not selected therefrom are multiplied by 0, so that only the selected elements can pass through it and be transmitted to the basis vectors.
[0153] For example, the associated vector A1 can be applied to the input vector R, and the inner product with the basis vector B1 can be calculated using only the N1 elements of the input vector R that have been specified by the associated vector A1. The inner product can represent C1 of the transformation coefficient vector C. In this regard, the basis vector B1 can include N1 elements, and N1 can be N or less. The above operations of the associated vector A2 to A N and the basis vectors B2 to B N can be performed on the input vector R with the associated vector A1 and the basis vector B1.
[0154] Since the associated vectors include only binary values such as 0 and / or 1, there are advantages in storing the associated vectors. A 0 in the associated vector can indicate that an element of the input vector R for 0 is not transmitted to the transformation matrix for inner product calculation, and a 1 in the associated vector can indicate that an element of the input vector R for 1 is transmitted to the transformation matrix for inner product calculation. For example, the 1×N-sized associated vector A k can include A k1 to A kn . If A k of the associated vector A kn is 0, the r n of the input vector R is not allowed to pass through. That is, the r n of the input vector R may not be transmitted to the transformation vector B n . Additionally, if A kn is 1, the r n of the input vector R is allowed to pass through. That is, the r n of the input vector R can be transmitted to the transformation vector B n , and can be used in the calculation of c n of the derived transformation coefficient vector C.
[0155] Figure 6 Schematically shows a multiple transformation technique in which a selective transformation is applied as a secondary transformation.
[0156] Referring to Figure 6 , the transformer can correspond to the transformer in the encoding device above Figure 1 , and the inverse transformer can correspond to the inverse transformer in the encoding device above Figure 1 or the aboveFigure 2 The inverse transformer in the decoding device.
[0157] The transformer may derive (primary) transform coefficients (S610) by performing a primary transform based on residual samples (residual sample array) in the residual block. Here, the first transform may include the above AMT.
[0158] If an adaptive multi-core transform is applied, a transform from the spatial domain to the frequency domain may be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1 to generate transform coefficients (or primary transform coefficients). In this regard, from the perspective of the transformer, the primary transform coefficients may be referred to as temporary transform coefficients. Additionally, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. For reference, the DCT / DST transform types may be defined based on basis functions, and the basis functions may be represented as in Table 1 above. Specifically, the process of deriving the primary transform coefficients by applying the adaptive multi-core transform is as described above.
[0159] The transformer may derive (secondary) transform coefficients (S620) by performing a selective transform based on the (primary) transform coefficients. The selective transform may mean a transform performed on the (primary) transform coefficients of the target block based on a transform matrix including modified basis vectors and an association matrix including association vectors of the basis vectors. The modified basis vector may represent a basis vector including N or fewer elements. That is, the modified basis vector may represent a basis vector including a specific number of elements selected from among N elements. For example, the modified basis vector B n may be a (1×N n )-dimensional vector, and N n may be less than or equal to N. That is, the modified basis vector B n may be of size (1×N n ), and N n may be less than or equal to N. Here, N may be the product of the height and width of the upper left target region of the target block to which the selective transform is applied. Alternatively, N may be the total number of transform coefficients of the upper left target region of the target block to which the selective transform is applied. In addition, the transform matrix including the modified basis vectors may be represented as a modified transform matrix. Additionally, the transform matrix may be represented as a transform basis block (TBB), and the association matrix may be represented as an association vector block (AVB).
[0160] The transformer can perform a selective transformation based on the modified transformation matrix and the correlation matrix, and obtain (secondary) transformation coefficients. As described above, the transformation coefficients can be derived as the transformation coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device and transmitted to the inverse quantizer / inverse transformer in the encoding device.
[0161] The inverse transformers can perform a series of processes in an order opposite to the order in which they have been performed in the above-mentioned transformers. The inverse transformers can receive the (inverse quantized) transformation coefficients, and derive the (primary) transformation coefficients by performing a selective (inverse) transformation (step S650), and can obtain the residual block (residual samples) by performing a primary (inverse) transformation on the (primary) transformation coefficients (S660). In this regard, from the perspective of the inverse transformers, the primary transformation coefficients can be referred to as the modified transformation coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0162] In addition, the present disclosure proposes a selective transformation combined with a simplified transformation as an example of the selective transformation.
[0163] In this specification, the term "simplified transformation" may mean a transformation performed on the residual samples of a target block based on a transformation matrix whose size is reduced according to a simplification factor.
[0164] In the simplified transformation according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a simplified transformation matrix can be determined, where R is less than N. That is, the simplified transformation may mean a transformation performed on the residual samples of a target block based on a simplified transformation matrix including R basis vectors. Here, N may mean the square of the length of the side of the block (or target region) to which the transformation is applied, or the total number of transformation coefficients corresponding to the block (or target region) to which the transformation is applied, and the simplification factor may mean the value of R / N. The simplification factor may be referred to by various terms such as a factor of simplification, a reduction factor, a factor of simplification, a simple factor, and other various terms. In addition, R may be referred to as a simplified coefficient, but depending on the situation, the simplification factor may mean R. Additionally, depending on the situation, the simplification factor may mean the value of N / R.
[0165] The size of the simplified transformation matrix according to the example can be R×N, which is smaller than the size N×N of the conventional transformation matrix, and can be defined as in Equation 4 below.
[0166] [Equation 4]
[0167]
[0168] If the simplified transformation matrix T RxNBy multiplying the transform coefficients of the primary transform applied to the target block, the (secondary) transform coefficients of the target block can be derived.
[0169] If RST is applied, since the reduced transform matrix of size R×N is applied to the secondary transform, the transform coefficients from R + 1 to N can implicitly become 0. In other words, if the transform coefficients of the target block are derived by applying RST, the values of the transform coefficients from R + 1 to N can be 0. Here, the transform coefficients from R + 1 to N can represent the (R + 1)-th to N-th transform coefficients among the transform coefficients. Specifically, the arrangement of the transform coefficients of the target block can be described as follows.
[0170] Figure 7 is a diagram for explaining the arrangement of transform coefficients of a target block according to an example of the present disclosure. Hereinafter, the explanations of the transforms to be described later in Figure 7 can be similarly applied to the inverse transform. The NSST based on the primary transform and the reduced transform can be performed on the target block (or residual block) 700. In the example, Figure 7 the 16×16 block shown in can represent the target block 700, and the 4×4 blocks labeled A to P can represent the sub-blocks of the target block 700. The primary transform can be performed on the entire range of the target block 700, and after the primary transform has been performed, the NSST can be applied to the 8×8 block (hereinafter, the upper left target region) composed of the subgroups A, B, E, and F. At this time, if the NSST based on the reduced transform is performed, since only R (here, R means the reduction coefficient and R is less than N) NSST transform coefficients are derived, the (R + 1)-th to N-th NSST transform coefficients can be determined to be 0. For example, if R is 16, the 16 transform coefficients derived by performing the NSST based on the reduced transform can be assigned to each block included in the subgroup A, which is the upper left 4×4 block included in the upper left target region of the target block 700, and the transform coefficient 0 can be assigned to each of the N - R (that is, 64 - 16 = 48) blocks included in the subgroups B, E, and F. The primary transform coefficients for which the NSST based on the reduced transform has not been performed can be assigned to each block included in the subgroups C, D, G, H, I, J, K, L, M, N, O, and P.
[0171] Figure 8 represents an example of deriving transform coefficients by combining the reduced transform and the selective transform with each other. Referring to Figure 8 the transform matrix can include R basis vectors, and the correlation matrix can include R correlation vectors. Here, the transform matrix including R basis vectors can be represented as the reduced transform matrix, and the correlation matrix including R correlation vectors can be represented as the reduced correlation matrix.
[0172] In addition, each of the basis vectors may include only elements selected from among N elements. For example, referring to Figure 8 , the basis vector B1 may be a 1×N1 dimensional vector including N1 elements, the basis vector B2 may be a 1×N2 dimensional vector including N2 elements, and the basis vector B R may be a 1×N R dimensional vector including N R elements. N1, N2, and N R may be values equal to or less than N. A transformation in which the simplified transformation and the selective transformation are combined with each other can be used for any type of transformation such as a secondary transformation, a primary transformation, etc.
[0173] Referring to Figure 8 , an encoding device and a decoding device may apply a transformation in which the simplified transformation and the selective transformation are combined with each other as a secondary transformation. For example, a selective transformation may be performed based on a simplified transformation matrix including a modified basis vector and a simplified correlation matrix, and (secondary) transformation coefficients may be obtained. In addition, in order to reduce the computational complexity of the selective transformation, a hypercube-Givens transform (HyGT) or the like may be used for the calculation of the selective transformation.
[0174] Figure 9 shows an example of deriving transformation coefficients by selective transformation. In the example regarding the selective transformation, there may be no patterns for the correlation vectors A1, A2,... A N of the correlation matrix, and the correlation vectors may be derived in different forms from each other. Alternatively, in another example regarding the selective transformation, the correlation vectors A1, A2,... A N of the correlation matrix may be derived in the same form.
[0175] Specifically, for example, each correlation vector may have the same number of 1s. For example, if the number of 1s is M, the correlation vector may have M 1s and (N−M) 0s. In this case, M transformation coefficients among the (primary) transformation coefficients of the input vector R may be transmitted to the basis vector. Thus, the length of the basis vector may also be M. That is, as Figure 9 shows, each basis vector may include M elements, may be a (1×M) dimensional vector, and may be derived such that N1 = N2 =... = N N = M. Figure 9 The correlation matrix and the modified transformation matrix architecture shown in
[0176] may be represented as a symmetric architecture using the selective transformation.
[0177] In addition, the above selective transformation can be applied together with other transformation techniques and the simplification transformation and / or HyGT.
[0178] Furthermore, the present disclosure proposes a method for simplifying the correlation vector in the above selective transformation. By simplifying the correlation vector, the storage of information for performing the selective transformation and the processing of the selective transformation can be further improved. That is, the storage burden for performing the selective transformation can be reduced, and the processing ability of the selective transformation can be further enhanced.
[0179] If the non-zero elements among the elements included in the correlation vector have a continuous distribution, the effect brought by simplifying the correlation vector can be more clearly demonstrated. For example, the correlation vector A k may include a continuous string of 1s. In this case, the correlation vector A ks and A kL can be used to represent A k . That is, the correlation vector A ks and A kL can be used to derive the correlation vector A k . Here, A ks can be a factor representing the starting point of non-zero elements (e.g., 1), and A kL can be a factor representing the length of non-zero elements. The correlation vector A k represented based on these factors can be derived as shown in the following table.
[0180] [Table 6]
[0181]
[0182] Referring to Table 6, the correlation vector A k may include 16 elements. That is, the correlation vector A k can be a 1×16 dimensional vector. The value of the factor A k representing the starting point of non-zero elements in the correlation vector A ks can be derived as 0, and in this case, the factor A ks can indicate the first element of the correlation vector A k as the starting point of non-zero elements. In addition, the value of the factor A k representing the length of non-zero elements in the correlation vector A kL can be derived as 8, and in this case, the factor A kL can indicate that the length of non-zero elements is 8. Therefore, as shown in Table 6, A k can be derived as a vector with the first to eighth elements being 1 and the remaining elements being 0 based on this factor.
[0183] Figure 10An example of performing selective transformation by deriving a correlation vector based on two factors of a correlation vector is shown. Refer to Figure 10 , a correlation matrix can be derived based on the factors of each of the correlation vectors, and the transform coefficients of the target block can be derived based on the correlation matrix and the modified transform matrix.
[0184] In addition, the starting point of the non-zero elements and the number of non-zero elements in each of the correlation vectors can be derived as fixed values, or the starting point of the non-zero elements and the number of non-zero elements in each of the correlation vectors can be derived in various ways.
[0185] Alternatively, for example, the starting point of the non-zero elements and the number of non-zero elements in the correlation vector can be derived based on the size of the upper-left target region that has been transformed. Here, the size of the upper-left target region can represent the number of transform coefficients of the upper-left target region, or can represent the product of the height and width of the upper-left target region. Alternatively, in another example, the starting point of the non-zero elements and the number of non-zero elements in the correlation vector can be derived based on the intra prediction mode of the target block. Specifically, for example, the starting point of the non-zero elements in the correlation vector can be derived based on whether the intra prediction mode of the target block is a non-directional intra prediction mode.
[0186] Alternatively, for example, the starting point of the non-zero elements and the number of non-zero elements in the correlation vector can be predetermined. Alternatively, for example, information indicating the starting point of the non-zero elements and information indicating the number of non-zero elements can be signaled, and the correlation vector can be derived based on the information indicating the starting point of the non-zero elements and the information indicating the number of non-zero elements. Alternatively, other information can be used instead of the information indicating the starting point of the non-zero elements. For example, as an alternative to the information indicating the starting point of the non-zero elements, information indicating the last position of the non-zero elements can be used, and the correlation vector can be derived based on the information indicating the last position of the non-zero elements.
[0187] In addition, the method of deriving the correlation vector based on these factors can also be applied to separable transforms and non-separable transforms such as the simplified transform HyGT.
[0188] Figure 11 Schematically represents an image encoding method performed by an encoding device according to the present disclosure. Figure 11 The method disclosed in Figure 1 can be performed by the encoding device disclosed in Figure 11 . Specifically, for example, S1100 of can be performed by the subtractor of the encoding device; S1110 can be performed by the transformer of the encoding device; and S1120 can be performed by the entropy encoder of the encoding device. In addition, although not shown, the process of deriving the predicted sample can be performed by the predictor of the encoding device.
[0189] The encoding device derives the residual samples of the target block (step S1100). For example, the encoding device can determine whether to perform inter-frame prediction or intra-frame prediction on the target block, and can determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. According to the determined mode, the encoding device can derive the prediction samples of the target block, and can derive the residual samples by adding the original samples of the target block and the prediction samples.
[0190] The encoding device derives the transform coefficients of the target block based on the selective transform of the residual samples (step S1110). The selective transform can be performed based on a modified transform matrix, which is a matrix including modified basis vectors, and the modified basis vectors can include a specific number of elements selected from among N elements. Additionally, the selective transform can be performed on the upper-left target region of the target block, and N can be the number of residual samples located in the upper-left target region. Additionally, N can be a value obtained by multiplying the width and height of the upper-left target region. For example, N can be 16 or 64.
[0191] The encoding device can derive the modified transform coefficients by performing a kernel transform on the residual samples, and can derive the transform coefficients of the target block by performing a selective transform on the modified transform coefficients located in the upper-left target region of the target block based on an association matrix of association vectors including the modified basis vectors and the modified transform matrix.
[0192] Specifically, the kernel transform on the residual samples can be performed as follows. The encoding device can determine whether to apply adaptive multi-kernel transform (AMT) to the target block. In this case, an AMT flag indicating whether to apply the adaptive multi-kernel transform to the target block can be generated. If AMT is not applied to the target block, the encoding device can derive DCT type 2 as the transform kernel for the target block, and can derive the modified transform coefficients by performing a transform on the residual samples based on DCT type 2.
[0193] If AMT is applied to the target block, the encoding device can configure a transform subset of the horizontal transform kernel and a transform subset of the vertical transform kernel, can derive the horizontal transform kernel and the vertical transform kernel based on the transform subsets, and can derive the modified transform coefficients by performing a transform on the residual samples based on the horizontal transform kernel and the vertical transform kernel. At this point, the transform subsets of the horizontal transform kernel and the vertical transform kernel can include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Additionally, transform index information can be generated, and the transform index information can include an AMT level flag indicating the AMT level of the horizontal transform kernel and an AMT vertical flag indicating the AMT vertical level of the vertical transform kernel. Furthermore, the transform kernel can be referred to as a transform type or a transform core.
[0194] If the modified transform coefficients are derived, the encoding device may derive the transform coefficients of the target block by performing selective transformation on the modified transform coefficients located in the upper left target region of the target block based on the correlation matrix of the correlation vectors including the modified basis vectors and the modified transform matrix. Other modified transform coefficients other than the modified transform coefficients located in the upper left region of the target block may be derived as the transform coefficients of the target block without any change.
[0195] Specifically, the modified transform coefficients of the elements having a value of 1 in the correlation vectors among the modified transform coefficients located in the upper left target region may be derived, and the transform coefficients of the target block may be derived based on the modified basis vectors and the derived modified transform coefficients. In this regard, the correlation vectors of the modified basis vectors may include N elements, the N elements may include elements having a value of 1 and / or elements having a value of 0, and the number of elements having a value of 1 may be A. In addition, the modified basis vectors may include A elements.
[0196] In addition, in one example, the modified transform matrix may include N modified basis vectors, and the correlation matrix may include N correlation vectors. The correlation vectors may include the same number of elements having a value of 1, and all the modified basis vectors may include the same number of elements. Alternatively, the correlation vectors may not include the same number of elements having a value of 1, and the modified basis vectors may not all include the same number of elements.
[0197] Alternatively, in another example, the modified transform matrix may include R modified basis vectors, and the correlation matrix may include R correlation vectors. R may be a reduction coefficient, and R may be less than N. The correlation vectors may include the same number of elements having a value of 1, and all the modified basis vectors may include the same number of elements. Alternatively, the correlation vectors may not include the same number of elements having a value of 1, and the modified basis vectors may not all include the same number of elements.
[0198] In addition, the correlation vectors may be configured such that the elements having a value of 1 are arranged continuously. In this case, in one example, the information about the correlation vectors may be entropy encoded. For example, the information about the correlation vectors may include the information indicating the starting point of the elements having a value of 1 and the information indicating the number of the elements having a value of 1. Alternatively, for example, the information about the correlation vectors may include the information indicating the last position of the elements having a value of 1 and the information indicating the number of the elements having a value of 1.
[0199] In addition, in another example, the correlation vector can be derived based on the size of the upper left target region. For example, the starting point of the elements with a value of 1 and the number of elements with a value of 1 in the correlation vector can be derived based on the size of the upper left target region.
[0200] Alternatively, in another example, the correlation vector can be derived based on the intra prediction mode of the target block. For example, the starting point of the elements with a value of 1 and the number of elements with a value of 1 in the correlation vector can be derived based on the intra prediction mode. Additionally, for example, the starting point of the elements with a value of 1 and the number of elements with a value of 1 in the correlation vector can be derived based on whether the intra prediction mode is a non - directional intra prediction mode.
[0201] The encoding device encodes information about the transform coefficients (S1330). The information about the transform coefficients can include information about the magnitude, position, etc. of the transform coefficients. Additionally, as described above, the information about the correlation vector can be entropy - encoded. For example, the information about the correlation vector can include information representing the starting point of the elements with a value of 1 and information representing the number of elements with a value of 1. Alternatively, for example, the information about the correlation vector can include information representing the last position of the elements with a value of 1 and information representing the number of elements with a value of 1.
[0202] The image information including information about the transform coefficients and / or information about the correlation vector can be output in the form of a bitstream. Additionally, the image information can further include prediction information. The prediction information can include information about motion information (e.g., when inter prediction is applied) and prediction mode information as multiple pieces of information related to the prediction process.
[0203] The output bitstream can be sent to the decoding device via a storage medium or a network.
[0204] Figure 12 Schematically represents an encoding device that executes an image encoding method according to the present disclosure. Figure 11 The method disclosed in Figure 12 can be executed by the encoding device disclosed in Figure 12 Specifically, for example, the adder of the encoding device in Figure 11 can execute step S1100, the transformer of the encoding device can execute S1110, and the entropy encoder of the encoding device can execute S1120 to S1130. Additionally, although not shown, the process of deriving the predicted samples can be executed by the predictor of the encoding device.
[0205] Figure 13 Schematically represents an image decoding method performed by a decoding device according to the present disclosure. Figure 13 The method disclosed in Figure 2The decoding device disclosed in [reference] performs. Specifically, for example, S1300 to S1310 in [reference] can be performed by the entropy decoder of the decoding device; S1320 can be performed by the inverse transformer of the decoding device; and S1330 can be performed by the adder of the decoding device. Additionally, although not shown, the process of deriving predicted samples can be performed by the predictor of the decoding device. Figure 13 in [reference].
[0206] The decoding device derives the transform coefficients of the target block from the bitstream (S1300). The decoding device can derive the transform coefficients of the target block by decoding the information on the transform coefficients of the target block received through the bitstream. The received information on the transform coefficients of the target block can be represented as residual information.
[0207] The decoding device derives the residual samples of the target block based on the selective transformation of the transform coefficients (S1310). The selective transformation can be performed based on a modified transform matrix, which is a matrix including modified basis vectors, and the modified basis vectors can include a specific number of elements selected from among N elements. Additionally, the selective transformation can be performed on the transform coefficients located in the upper-left target region of the target block, and N can be the number of transform coefficients located in the upper-left target region. Additionally, N can be a value obtained by multiplying the width and height of the upper-left target region. For example, N can be 16 or 64.
[0208] The decoding device can derive the modified transform coefficients by performing selective transformation on the transform coefficients located in the upper-left target region of the target block based on the modified transform matrix and an association matrix including association vectors including the modified basis vectors.
[0209] Specifically, the transform coefficients of the elements among the transform coefficients located in the upper-left target region whose association vector values are 1 can be derived, and the modified transform coefficients can be derived based on the derived transform coefficients and the modified basis vectors. In this regard, the association vectors of the modified basis vectors can include N elements, which can include elements with a value of 1 and / or elements with a value of 0, and the number of elements with a value of 1 can be A. Additionally, the modified basis vectors can include A elements.
[0210] Furthermore, in one example, the modified transform matrix can include N modified basis vectors, and the association matrix can include N association vectors. The association vectors can include the same number of elements with a value of 1, and all the modified basis vectors can include the same number of elements. Alternatively, the association vectors can not include the same number of elements with a value of 1, and the modified basis vectors can not all include the same number of elements.
[0211] Alternatively, in another example, the modified transform matrix may include R modified basis vectors, and the correlation matrix may include R correlation vectors. R may be a reduction coefficient, and R may be less than N. The correlation vectors may include the same number of elements with a value of 1, and all of the modified basis vectors may include the same number of elements. Alternatively, the correlation vectors may not include the same number of elements with a value of 1, and all of the modified basis vectors may not include the same number of elements.
[0212] In addition, the correlation vectors may be configured such that the elements with a value of 1 are arranged consecutively. In this case, in one example, information about the correlation vectors may be obtained from the bitstream, and the correlation vectors may be derived based on the information about the correlation vectors. For example, the information about the correlation vectors may include information indicating the starting point of the elements with a value of 1 and information indicating the number of the elements with a value of 1. Alternatively, for example, the information about the correlation vectors may include information indicating the last position of the elements with a value of 1 and information indicating the number of the elements with a value of 1.
[0213] Alternatively, in another example, the correlation vectors may be derived based on the size of the upper left target region. For example, the starting point of the elements with a value of 1 and the number of the elements with a value of 1 in the correlation vectors may be derived based on the size of the upper left target region.
[0214] Alternatively, in another example, the correlation vectors may be derived based on the intra prediction mode of the target block. For example, the starting point of the elements with a value of 1 and the number of the elements with a value of 1 in the correlation vectors may be derived based on the intra prediction mode. Additionally, for example, the starting point of the elements with a value of 1 and the number of the elements with a value of 1 in the correlation vectors may be derived based on whether the intra prediction mode is a non-directional intra prediction mode.
[0215] If the modified transform coefficients are derived, the decoding device may derive the residual samples by performing a kernel transform on the target block including the modified transform coefficients.
[0216] The kernel transform on the target block may be performed as follows. The decoding device may obtain an AMT flag indicating whether adaptive multi-kernel transform (AMT) is applied from the bitstream, and if the value of the AMT flag is 0, the decoding device may derive the DCT type 2 as the transform kernel of the target block, and may derive the residual samples by performing an inverse transform on the target block including the modified transform coefficients based on the DCT type 2.
[0217] If the value of the AMT flag is 1, the decoding device can configure the transform subsets of the horizontal transform kernel and the vertical transform kernel, can derive the horizontal transform kernel and the vertical transform kernel based on the transform subsets and the transform index information obtained from the bitstream, and can derive the residual samples by performing an inverse transform on a target block including modified transform coefficients based on the horizontal transform kernel and the vertical transform kernel. In this regard, the transform subsets of the horizontal transform kernel and the vertical transform kernel can include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Additionally, the transform index information can include an AMT horizontal flag indicating one of the candidates included in the transform subset of the horizontal transform kernel and an AMT vertical flag indicating one of the candidates included in the transform subset of the vertical transform kernel. Furthermore, the transform kernel can be referred to as a transform type or a transform core.
[0218] The decoding device generates a reconstructed picture based on the residual samples (S1320). The decoding device can generate a reconstructed picture based on the residual samples. For example, the decoding device can perform inter prediction or intra prediction on a target block based on prediction information received through the bitstream, can derive prediction samples, and can generate a reconstructed picture by adding the prediction samples and the residual samples. Thereafter, as described above, loop filtering processes such as the ALF process, SAO, and / or deblocking filtering can be applied to the reconstructed picture as needed to improve the subjective / objective video quality.
[0219] Figure 14 Schematically represents a decoding device that executes an image decoding method according to the present disclosure. Figure 13 The method disclosed in Figure 14 can be executed by the Figure 14 entropy decoder of the decoding device in Figure 13 ; Figure 14 the inverse transformer of the decoding device in Figure 13 can execute Figure 14 ; and Figure 13 the adder of the decoding device in Figure 14 can execute
[0220] According to the above present disclosure, through efficient transformation, it is possible to reduce the amount of data that must be transmitted for residual processing and to increase the residual coding efficiency.
[0221] Additionally, according to the present disclosure, it is possible to perform a non-separable transform based on a transform matrix composed of basis vectors including a specific number of selected elements. Accordingly, it is possible to reduce the storage burden and computational complexity of the non-separable transform and to increase the residual coding efficiency.
[0222] In addition, according to the present disclosure, an inseparable transform can be performed based on a transform matrix of a simplified architecture, whereby the amount of data that must be transmitted for residual processing can be reduced, and the residual coding efficiency can be increased.
[0223] In the above-described embodiments, the method is explained based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may occur in an order or steps different from the above-described order or steps, or may occur simultaneously with another step. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and one or more steps that can be incorporated may be removed without affecting the scope of the present disclosure.
[0224] The above method according to the present disclosure can be implemented in software form, and the encoding device and / or decoding device according to the present disclosure can be included in a device for image processing such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0225] When the embodiments in the present disclosure are implemented by software, the above method can be implemented as a module (process, function, etc.) to perform the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor via various well-known devices. The processor can include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller, or a chip. In addition, the functional units shown in each figure can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0226] In addition, the decoding device and the encoding device applying the present disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and can be used to process video signals or data signals. For example, an over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0227] In addition, the processing method applying the present disclosure can be generated in the form of a program executable by a computer and stored in a computer-readable recording medium. The multimedia data having a data structure according to the present disclosure can also be stored in the computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bitstream generated by an encoding method can be stored in the computer-readable recording medium or transmitted through a wired or wireless communication network. In addition, the embodiments of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed on a computer by the embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.
[0228] In addition, the content streaming system applying the present disclosure mainly includes an encoding server, a streaming server, a network server, a media memory, a user device, and a multimedia input device.
[0229] The encoding server is used to compress the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream and send it to the streaming server. As another example, in the case where the multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or the bitstream generation method of the present disclosure. And, the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0230] The streaming server sends the multimedia data to the user device through the network server based on the user's request, and the network server serves as an instrument for informing the user of what services exist. When the user requests the service he or she wants, the network server transmits it to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server is used to control the commands / responses between various devices in the content streaming system.
[0231] The streaming server can receive the content from the media memory and / or the encoding server. For example, in the case of receiving the content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide the streaming service.
[0232] For example, the user equipment may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigator, a touchscreen PC, a tablet PC, a superbook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc. Each of the servers in the content streaming system may operate as a distributed server, and in this case, the data received by each server may be processed in a distributed manner.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Derive the inverse-quantized transform coefficients of a target block from a bitstream; Derive the primary transform coefficients of the target block based on an inseparable transform of the inverse-quantized transform coefficients; Derive the residual samples of the target block based on a separable transform of the primary transform coefficients; and Generate a reconstructed picture based on the residual samples of the target block and the prediction samples of the target block, wherein the inseparable transform is performed based on a transform matrix, wherein the transform matrix includes modified basis vectors, wherein each of the modified basis vectors includes fewer than N elements, wherein the N is equal to the number of inverse-quantized transform coefficients in the region of the target block where the inseparable transform is applied, wherein the region where the inseparable transform is applied is the 8×8 upper-left target region in the target block, and the N is equal to 64, and wherein the number of the modified basis vectors is fewer than N.
2. The image decoding method according to claim 1, wherein, Perform the inseparable transform on the inverse-quantized transform coefficients in the 8×8 upper-left target region based on the transform matrix and an incidence matrix including incidence vectors including modified basis vectors, wherein the incidence vector sequentially includes M elements with a value of 1 and N-M elements with a value of 0, wherein the M elements among the N transform coefficients in the 8×8 upper-left target region are derived by multiplying the elements by the value 1 in the incidence vector, and wherein the primary transform coefficients are derived by performing the inseparable transform on the M elements among the N transform coefficients in the 8×8 upper-left target region.
3. The image decoding method according to claim 2, wherein, The modified basis vector includes M elements.
4. The image decoding method according to claim 2, wherein, The transform matrix includes N modified basis vectors, and the incidence matrix includes N incidence vectors.
5. The image decoding method according to claim 4, wherein, The incidence vectors include the same number of elements with a value of 1, and all the modified basis vectors include the same number of elements.
6. The image decoding method according to claim 2, wherein, The transform matrix includes R modified basis vectors, and the incidence matrix includes R incidence vectors, and the R is less than N.
7. The image decoding method according to claim 2, wherein, Obtain information about the incidence vector from the bitstream, and derive the incidence vector based on the information about the incidence vector, and wherein the information about the incidence vector includes information indicating the starting point of the elements with a value of 1 and information indicating the number of the elements with a value of 1.
8. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Derive the residual samples of a target block; Derive the primary transform coefficients of the target block based on a separable transform of the residual samples; Derive the secondary transform coefficients of the target block based on an inseparable transform of the primary transform coefficients; Derive the quantized transform coefficients of the target block based on a secondary transform; and Encode information about the quantized transform coefficients, wherein the inseparable transform is performed based on a transform matrix, wherein the transform matrix includes modified basis vectors, Each of the modified basis vectors includes fewer than N elements, wherein the N is equal to the number of primary transform coefficients in the region of the target block where the non-separable transform is applied, wherein the region where the non-separable transform is applied is the 8×8 upper left target region in the target block, and the N is equal to 64, and wherein the number of the modified basis vectors is fewer than N.
9. The image encoding method according to claim 8, wherein, Perform the non-separable transform on the primary transform coefficients in the 8×8 upper left target region of the target block based on the transform matrix and the correlation matrix including the correlation vectors including the modified basis vectors, wherein the correlation vector sequentially includes M elements with a value of 1 and N-M elements with a value of 0, wherein the M elements among the N transform coefficients in the 8×8 upper left target region are derived by multiplying the elements by the value 1 in the correlation vector, and wherein the secondary transform coefficients are derived by performing the non-separable transform on the M elements among the N transform coefficients in the 8×8 upper left target region.
10. The image encoding method according to claim 9, wherein, The transform matrix includes R modified basis vectors, and the correlation matrix includes R correlation vectors, and the R is less than N.
11. The image encoding method according to claim 9, wherein, The information about the correlation vector is entropy encoded, and the information about the correlation vector includes the information representing the start point of the elements with a value of 1 and the information representing the number of the elements with a value of 1.
12. A method for sending data including a bitstream for an image, the sending method comprising the following steps: Obtain a bitstream for an image, wherein the bitstream is generated by the following steps: deriving residual samples of a target block, deriving primary transform coefficients of the target block by a separable transform based on the residual samples, deriving secondary transform coefficients of the target block by a non-separable transform based on the primary transform coefficients, deriving quantized transform coefficients of the target block based on a secondary transform, and encoding information about the quantized transform coefficients; and Send the data including the bitstream, wherein the non-separable transform is performed based on a transform matrix, wherein the transform matrix includes modified basis vectors, wherein each of the modified basis vectors includes fewer than N elements, wherein the N is equal to the number of primary transform coefficients in the region of the target block where the non-separable transform is applied, wherein the region where the non-separable transform is applied is the 8×8 upper left target region in the target block, and the N is equal to 64, and wherein the number of the modified basis vectors is fewer than N.
Citation Information
Patent Citations
Method of decoding motion vector
CN103152562A
Non-separable secondary transform for video coding
US20170094313A1