Decoding and encoding devices, media for storing bit streams, and data transmission devices
Through the image decoding method of selective transformation matrix and basis vector, the problem of high-resolution image transmission and storage costs is solved, and the encoding efficiency and transformation efficiency are improved.
Patent Information
- Application Number
- CN202211470298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-21
- Filing Date
- 2018-12-21
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2038-12-21
AI Technical Summary
The prior art when transmitting and storing high-resolution and high-quality images, the increase in the amount of information leads to high costs, and it is necessary to improve image encoding efficiency and transformation efficiency.
The selective transformation method is adopted to decode the image through the modified transformation matrix and the base vector, including selecting a specific number of elements from N elements for transformation.
The data volume and calculation complexity of residual processing are reduced, the residual coding efficiency is improved, and the storage and transmission costs are reduced.
Smart Images

Figure CN115834876B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application with application number 201880086807.8 (International application number: PCT / KR2018 / 016437, application date: December 21, 2018, invention name: Image coding method based on selective transformation and device for this method). Technical Field
[0002] The present disclosure relates to an image encoding technology, and more particularly, to an image decoding method according to selective transformation in an image encoding system and a device used for the image decoding method. Background Art
[0003] Across various fields, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition), is growing. Because image data has high resolution and high quality, the amount of information, or bits, to be transmitted increases compared to conventional image data. Consequently, when image data is transmitted using media such as conventional wired / wireless broadband lines or stored using existing storage media, transmission and storage costs increase.
[0004] Therefore, there is a need for efficient image compression technology for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images. Summary of the Invention
[0005] Technical tasks
[0006] The technical problem to be solved by the present disclosure is to provide a method and device for improving image coding efficiency.
[0007] Another technical problem to be solved by the present disclosure is to provide a method and device for improving conversion efficiency.
[0008] Another technical problem to be solved by the present disclosure is to provide a method and apparatus for improving residual coding efficiency through transformation.
[0009] Another technical problem to be solved by the present disclosure is to provide an image encoding method and apparatus based on selective transformation.
[0010] Solution
[0011] According to an example of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes the following steps: deriving transform coefficients of a target block from a bitstream; deriving residual samples of the target block based on a selective transform of the transform coefficients; and generating a reconstructed picture based on the residual samples of the target block and predicted samples of the target block, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from N elements.
[0012] According to another example of the present disclosure, a decoding device for performing image decoding is provided. The image decoding device includes: an entropy decoder that derives transform coefficients of a target block from a bitstream; an inverse transformer that derives residual samples of the target block based on a selective transform of the transform coefficients; and an adder that generates a reconstructed picture based on the residual samples of the target block and predicted samples of the target block, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from N elements.
[0013] According to another example of the present disclosure, a video encoding method performed by an encoding device is provided. The method includes the following steps: deriving residual samples of a target block; deriving transform coefficients of the target block based on a selective transform of the residual samples; and encoding information about the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from N elements.
[0014] According to another example of the present disclosure, a video encoding device is provided. The encoding device includes: an adder that derives residual samples of a target block; a transformer that derives transform coefficients of the target block based on a selective transform of the residual samples; and an entropy encoder that encodes information about the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors include a specific number of elements selected from N elements.
[0015] Technical Effects
[0016] According to the present disclosure, through efficient transformation, it is possible to reduce the amount of data that must be transmitted for residual processing and increase residual coding efficiency.
[0017] According to the present disclosure, it is possible to perform an inseparable transform based on a transform matrix composed of basis vectors including a specific number of selected elements, thereby reducing the storage burden and computational complexity of the inseparable transform and increasing residual coding efficiency.
[0018] According to the present disclosure, non-separable transform can be performed based on a transform matrix of a simplified structure, thereby being able to reduce the amount of data that must be transmitted for residual processing and increase residual encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. 1 is a schematic diagram illustrating a configuration of a video encoding device to which the present disclosure is applicable.
[0020] Figure 2 FIG. 1 is a schematic diagram illustrating a configuration of a video decoding device to which the present disclosure is applicable.
[0021] Figure 3 Schematic representation of the multiple transformation technique according to the present disclosure.
[0022] Figure 4 65 intra-frame directional modes for prediction directions are exemplarily shown.
[0023] Figures 5a to 5c is a diagram for explaining selective transformation according to an example of the present disclosure.
[0024] Figure 6 Schematic representation of a multiple transform technique where selective transforms are applied as secondary transforms.
[0025] Figure 7 is a diagram for explaining arrangement of transform coefficients based on a target block according to an example of the present disclosure.
[0026] Figure 8 This shows an example of deriving transform coefficients through transforms in which simplified transform and selective transform are combined with each other.
[0027] Figure 9 An example of deriving transform coefficients through selective transform is shown.
[0028] Figure 10 An example of performing selective transformation by deriving a correlation vector based on two factors of the correlation vector is shown.
[0029] Figure 11 The following schematically illustrates an image encoding method performed by an encoding device according to the present disclosure.
[0030] Figure 12 The figure schematically shows an encoding device for performing an image encoding method according to the present disclosure.
[0031] Figure 13The figure schematically shows an image decoding method performed by a decoding device according to the present disclosure.
[0032] Figure 14 A decoding device for performing an image decoding method according to the present disclosure is schematically shown. DETAILED DESCRIPTION
[0033] The present disclosure can be modified in various forms, and its specific embodiments will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are used to describe only specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions as long as they are clearly understood differently. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, so it should be understood that there is no exclusion of the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0034] In addition, for the purpose of conveniently illustrating different specific functions, the elements in the figures described in this disclosure are drawn independently, which does not mean that these elements are implemented by independent hardware or independent software. For example, two or more of these elements can be combined to form a single element, or an element can be divided into multiple elements. Without departing from the concept of the present disclosure, embodiments of combining and / or dividing elements belong to the present disclosure.
[0035] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, in the entire drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.
[0036] Furthermore, the present disclosure relates to video / image coding. For example, the methods / implementations disclosed in the present disclosure may be applied to methods disclosed in the Versatile Video Coding (VVC) standard or next-generation video / image coding standards.
[0037] In this disclosure, generally, a picture refers to a unit representing an image in a specific time slot, and a slice is a unit constituting a part of a picture in coding. A picture can include multiple slices, and in some cases, pictures and slices can be used in a mixed manner.
[0038] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, the term "sample" may be used in conjunction with a pixel. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of a luma component or only the pixel / pixel value of a chroma component.
[0039] A unit represents the basic unit of image processing. A unit can include at least one of a specific region of an image and information related to the corresponding region. In some cases, units can be used in a mixed manner with blocks or regions. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows.
[0040] Figure 1 FIG2 is a diagram briefly illustrating a video encoding device to which the present disclosure is applicable.
[0041] Reference Figure 1 , the video encoding apparatus 100 may include a picture splitter 105, a predictor 110, a residual processor 120, an entropy encoder 130, an adder 140, a filter 150, and a memory 160. The residual processor 120 may include a subtractor 121, a transformer 122, a quantizer 123, a rearranger 124, an inverse quantizer 125, and an inverse transformer 126.
[0042] The picture divider 105 may divide the input picture into at least one processing unit.
[0043] For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from the largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBT) structure. For example, one coding unit may be split into a plurality of coding units of a deeper depth based on a quadtree structure and / or a binary tree structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on a final coding unit that is no longer split. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively split into coding units of a deeper depth when necessary, and a coding unit of an optimal size may be used as the final coding unit. Here, the encoding process may include the processes of prediction, transformation, and reconstruction to be described later.
[0044] In another example, a processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). A coding unit may be split from a maximum coding unit (LCU) into deeper coding units according to a quadtree structure. In this case, the maximum coding unit may be directly used as the final coding unit based on coding efficiency, etc., according to image characteristics, or the coding unit may be recursively split into deeper coding units when necessary, and the coding unit with the optimal size may be used as the final coding unit. When a minimum coding unit (SCU) is set, the coding unit may not be split into coding units smaller than the minimum coding unit. Here, the final coding unit refers to a coding unit that is split or divided into prediction units or transform units. A prediction unit is a unit split from a coding unit and may be a unit for sample prediction. Here, a prediction unit may be divided into sub-blocks. A transform unit may be separated from a coding unit according to a quadtree structure, and may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients. Hereinafter, a coding unit may be referred to as a coding block (CB), a prediction unit may be referred to as a prediction block (PB), and a transform unit may be referred to as a transform block (TB). A prediction block or prediction unit may refer to a specific region in the form of a block in a picture and include an array of prediction samples. Additionally, a transform block or transform unit may refer to a specific region in the form of a block in a picture and include an array of transform coefficients or residual samples.
[0045] The predictor 110 may perform prediction on a processing target block (hereinafter, a current block) and may generate a prediction block including prediction samples for the current block. The unit of prediction performed in the predictor 110 may be a coding block, a transform block, or a prediction block.
[0046] The predictor 110 may determine whether to apply intra prediction or inter prediction to the current block. For example, the predictor 110 may determine whether to apply intra prediction or inter prediction in units of CUs.
[0047] In the case of intra-frame prediction, the predictor 110 can derive the prediction sample of the current block based on a reference sample outside the current block in the picture to which the current block belongs (hereinafter, the current picture). In this case, the predictor 110 can derive the prediction sample based on the average or interpolation of the neighboring reference samples of the current block (case (i)), or can derive the prediction sample based on the reference sample existing in a specific (prediction) direction among the neighboring reference samples of the current block (case (ii)). Case (i) can be called a non-directional mode or a non-angle mode, and case (ii) can be called a directional mode or an angle mode. In intra-frame prediction, as an example, the prediction mode can include 33 directional modes and at least two non-directional modes. The non-directional mode can include a DC mode and a planar mode. The predictor 110 can determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring block.
[0048] In the case of inter-frame prediction, the predictor 110 can derive the prediction samples of the current block based on the samples specified by the motion vector on the reference picture. The predictor 110 can derive the prediction samples of the current block by applying any one of the skip mode, merge mode, and motion vector prediction (MVP) mode. In the case of skip mode and merge mode, the predictor 110 can use the motion information of the neighboring block as the motion information of the current block. In the case of skip mode, unlike merge mode, the difference (residual) between the prediction sample and the original sample is not sent. In the case of MVP mode, the motion vector of the neighboring block is used as a motion vector predictor and is therefore used as the motion vector predictor of the current block to derive the motion vector of the current block.
[0049] In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the temporal neighboring blocks may also be referred to as a collocated picture (colPic). Motion information may include motion vectors and reference picture indexes. Information such as prediction mode information and motion information may be (entropy) encoded and then output as a bitstream.
[0050] When using motion information of temporally neighboring blocks in skip mode and merge mode, the highest picture in the reference picture list can be used as the reference picture. The reference pictures included in the reference picture list can be aligned based on the picture order (POC) difference between the current picture and the corresponding reference picture. The POC corresponds to the display order and can be distinguished from the coding order.
[0051] The subtractor 121 generates residual samples, which are differences between original samples and predicted samples. If the skip mode is applied, the residual samples may not be generated as described above.
[0052] The transformer 122 transforms the residual samples in units of transform blocks to generate transform coefficients. The transformer 122 may perform the transform based on the size of the corresponding transform block and the prediction mode applied to the prediction block or coding block that spatially overlaps with the transform block. For example, if intra prediction is applied to the prediction block or coding block that overlaps with the transform block, the residual samples may be transformed using a discrete sine transform (DST) transform kernel, the transform block being a 4×4 residual array, and in other cases, a discrete cosine transform (DCT) transform kernel may be used to transform it.
[0053] The quantizer 123 may quantize the transform coefficient to generate a quantized transform coefficient.
[0054] The rearranger 124 rearranges the quantized transform coefficients. The rearranger 124 may rearrange the quantized transform coefficients in block form into a one-dimensional vector using a coefficient scanning method. Although the rearranger 124 is described as a separate component, the rearranger 124 may be part of the quantizer 123.
[0055] The entropy encoder 130 may perform entropy coding on the quantized transform coefficients. Entropy coding may include coding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 130 may encode information necessary for video reconstruction (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients, either together or separately. The entropy-coded information may be transmitted or stored in units of NALs (Network Abstraction Layers) in the form of a bitstream.
[0056] The dequantizer 125 dequantizes the value (transform coefficient) quantized by the quantizer 123 , and the inverse transformer 126 inversely transforms the value dequantized by the dequantizer 125 to generate residual samples.
[0057] Adder 140 adds the residual samples to the prediction samples to reconstruct the picture. The residual samples and the prediction samples can be added in units of blocks to generate a reconstructed block. Although adder 140 is described as a separate component, adder 140 can be part of predictor 110. In addition, adder 140 can be referred to as a reconstructor or a reconstructed block generator.
[0058] The filter 150 may apply deblocking filtering and / or sample adaptive offset to the reconstructed picture. Deblocking filtering and / or sample adaptive offset may be used to correct artifacts at block boundaries or distortion in quantization in the reconstructed picture. After deblocking filtering is completed, sample adaptive offset may be applied on a sample-by-sample basis. The filter 150 may apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture to which deblocking filtering and / or sample adaptive offset have been applied.
[0059] The memory 160 may store reconstructed pictures (decoded pictures) or information required for encoding / decoding. Here, the reconstructed pictures may be reconstructed pictures filtered by the filter 150. The stored reconstructed pictures may be used as reference pictures for (inter) prediction of other pictures. For example, the memory 160 may store (reference) pictures used for inter prediction. Here, pictures used for inter prediction may be specified based on a reference picture set or a reference picture list.
[0060] Figure 2 FIG. 1 is a schematic diagram illustrating a configuration of a video decoding device to which the present disclosure is applicable.
[0061] Reference Figure 2 , the video decoding apparatus 200 includes an image decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. Here, the residual processor 220 may include a rearranger 221, an inverse quantizer 222, and an inverse transformer 223.
[0062] When a bitstream including video information is input, the video decoding apparatus 200 may reconstruct a video corresponding to a process of processing the video information in the video encoding apparatus.
[0063] For example, the video decoding device 200 may use a processor used in a video encoding device to perform video decoding. Therefore, the processing unit block of video decoding may be, for example, a coding unit, or may be a coding unit, a prediction unit, or a transform unit. Coding units may be divided from the maximum coding unit according to a quadtree structure and / or a binary tree structure.
[0064] In some cases, prediction units and transform units may also be used. In this case, a prediction block is a block derived or partitioned from a coding unit and may be a unit for sample prediction. Here, a prediction unit may be divided into subblocks. A transform unit may be partitioned from a coding unit according to a quadtree structure, and may be a unit for deriving transform coefficients or a unit for deriving a residual signal from the transform coefficients.
[0065] The entropy decoder 210 can parse the bitstream to output information required for video reconstruction or picture reconstruction. For example, the entropy decoder 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of syntax elements required for video reconstruction and the quantized values of the transform coefficients used for the residual.
[0066] More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the adjacent and decoding target blocks or the information of the symbol / bin decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding on the bin to generate a symbol corresponding to each syntax element value. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded by the context model for the next symbol / bin.
[0067] Information used for prediction among the information decoded in the entropy decoder 210 may be provided to the predictor 230 , and residual values, ie, quantized transform coefficients on which the entropy decoder 210 has performed entropy decoding, may be input to the rearranger 221 .
[0068] The rearranger 221 may rearrange the quantized transform coefficients into a two-dimensional block form. The rearranger 221 may perform rearrangement corresponding to the coefficient scanning performed by the encoding device. Although the rearranger 221 is described as a separate component, the rearranger 221 may be part of the inverse quantizer 222.
[0069] The inverse quantizer 222 may inversely quantize the quantized transform coefficient based on the (inverse) quantization parameter to output the transform coefficient. In this case, information for deriving the quantization parameter may be signaled from the encoding device.
[0070] The inverse transformer 223 may inversely transform the transform coefficients to derive residual samples.
[0071] The predictor 230 may perform prediction on the current block and may generate a prediction block including prediction samples of the current block. A unit of prediction performed in the predictor 230 may be a coding block, a transform block, or a prediction block.
[0072] The predictor 230 can determine whether to apply intra-frame prediction or inter-frame prediction based on the information used for prediction. In this case, the unit used to determine which of intra-frame prediction and inter-frame prediction to use may be different from the unit used to generate prediction samples. In addition, the unit used to generate prediction samples may also be different in inter-frame prediction and intra-frame prediction. For example, it can be determined on a CU basis whether to apply inter-frame prediction or intra-frame prediction. In addition, for example, in inter-frame prediction, prediction samples can be generated by determining a prediction mode on a PU basis, while in intra-frame prediction, prediction samples can be generated on a TU basis by determining a prediction mode on a PU basis.
[0073] In the case of intra prediction, the predictor 230 may derive prediction samples for the current block based on neighboring reference samples in the current picture. The predictor 230 may derive prediction samples for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined by using the intra prediction mode of the neighboring block.
[0074] In the case of inter-frame prediction, the predictor 230 can derive the prediction sample of the current block based on the sample specified in the reference picture according to the motion vector. The predictor 230 can derive the prediction sample of the current block using one of the skip mode, merge mode, and MVP mode. Here, the motion information required for inter-frame prediction of the current block provided by the video encoding device, such as the motion vector and information for the reference picture index, can be acquired or derived based on the information used for prediction.
[0075] In skip mode and merge mode, the motion information of the neighboring blocks can be used as the motion information of the current block. Here, the neighboring blocks can include spatial neighboring blocks and temporal neighboring blocks.
[0076] The predictor 230 can use the motion information of available neighboring blocks to construct a merge candidate list and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled by the encoding device. The motion information can include a motion vector and a reference picture. When using the motion information of temporally neighboring blocks in skip mode and merge mode, the highest picture in the reference picture list can be used as the reference picture.
[0077] In the case of the skip mode, unlike the merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.
[0078] In the case of MVP mode, the motion vector of the current block can be derived using the motion vector of the neighboring block as a motion vector predictor. Here, the neighboring block may include a spatial neighboring block and a temporal neighboring block.
[0079] When the merge mode is applied, for example, the motion vector of the reconstructed spatial neighboring block and / or the motion vector corresponding to the Col block as the temporal neighboring block can be used to generate a merge candidate list. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The above-mentioned information for prediction may include a merge index indicating the candidate block with the best motion vector selected from the candidate blocks included in the merge candidate list. Here, the predictor 230 can use the merge index to derive the motion vector of the current block.
[0080] As another example, when the MVP (motion vector prediction) mode is applied, a motion vector predictor candidate list can be generated using the reconstructed motion vectors of spatially neighboring blocks and / or the motion vectors corresponding to the Col block, which is a temporally neighboring block. That is, the reconstructed motion vectors of spatially neighboring blocks and / or the motion vectors corresponding to the Col block, which is a temporally neighboring block, can be used as motion vector candidates. The aforementioned information used for prediction may include a predicted motion vector index indicating the best motion vector selected from the motion vector candidates included in the list. Here, the predictor 230 may use the motion vector index to select a predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list. The predictor of the encoding device may obtain a motion vector difference (MVD) between the motion vector of the current block and a motion vector predictor, encode the MVD, and output the encoded MVD in the form of a bitstream. That is, the MVD may be obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the predictor 230 may obtain the motion vector included in the information used for prediction and derive the motion vector of the current block by adding the motion vector difference to the motion vector predictor. In addition, the predictor may obtain or derive a reference picture index indicating a reference picture from the above-mentioned information used for prediction.
[0081] Adder 240 can add residual samples to prediction samples to reconstruct the current block or current picture. Adder 240 can reconstruct the current picture by adding residual samples to prediction samples in units of blocks. When skip mode is applied, no residual is transmitted, so the prediction samples can become reconstructed samples. Although adder 240 is described as a separate component, adder 240 can be part of predictor 230. In addition, adder 240 can be referred to as a reconstructor or a reconstructed block generator.
[0082] The filter 250 may apply deblocking filtering, sample adaptive offset, and / or ALF to the reconstructed picture. Here, after deblocking filtering, sample adaptive offset may be applied in units of samples. ALF may be applied after deblocking filtering and / or applying sample adaptive offset.
[0083] The memory 260 can store reconstructed pictures (decoded pictures) or information required for decoding. Here, the reconstructed pictures can be reconstructed pictures filtered by the filter 250. For example, the memory 260 can store pictures used for inter-frame prediction. Here, the pictures used for inter-frame prediction can be specified based on a reference picture set or a reference picture list. The reconstructed pictures can be used as reference pictures for other pictures. The memory 260 can output the reconstructed pictures in the output order.
[0084] Furthermore, a lower-frequency transform coefficient of the residual block of the current block may be derived through the above-described transform, and a zero tail may be derived at the end of the residual block.
[0085] Specifically, the transformation may consist of two main processes, and these main processes may include a kernel transformation and a secondary transformation. The transformation including the kernel transformation and the secondary transformation may be denoted as a multiple transformation technique.
[0086] Figure 3 Schematic representation of the multiple transformation technique according to the present disclosure.
[0087] Reference Figure 3 , the converter can correspond to the above Figure 1 The converter in the encoding device, and the inverse converter may correspond to the above Figure 1 The inverse transformer in the encoding device or Figure 2 An inverse transformer in a decoding device.
[0088] The transformer may derive (primary) transform coefficients by performing a primary transform based on the residual samples (residual sample array) in the residual block (S310). In this regard, the primary transform may include an adaptive multi-core transform (AMT). The adaptive multi-core transform may be represented as a multi-transform set (MTS).
[0089] Adaptive multi-core transform may refer to a method for performing transform using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1. That is, adaptive multi-core transform may refer to a transform method that transforms a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this regard, from the perspective of the transformer, the primary transform coefficients may be referred to as temporary transform coefficients.
[0090] In other words, if an existing transform method is applied, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, if an adaptive multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DCT type 7, DCT type 8, and / or DST type 1. Here, DCT type 2, DCT type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores.
[0091] For reference, DCT / DST transform types may be defined based on basis functions, and the basis functions may be represented as in the following table.
[0092] [Table 1]
[0093]
[0094] If adaptive multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for the target block may be selected from the transform kernels, and a vertical transform for the target block may be performed based on the vertical transform kernel, and a horizontal transform for the target block may be performed based on the horizontal transform kernel. Here, the horizontal transform may refer to a transform for the horizontal component of the target block, and the vertical transform may refer to a transform for the vertical component of the target block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on a transform index indicating a prediction mode and / or a transform subset of the target block (CU or subblock) containing the residual block.
[0095] For example, when both the width and height of the target block are less than or equal to 64, an adaptive multi-core transform can be applied, and the determination can be made based on a CU-level flag related to whether to apply the adaptive multi-core transform of the target block. Specifically, if the CU-level flag is 0, the above-mentioned existing transform method can be applied. That is, if the CU-level flag is 0, a transform coefficient can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, and the transform coefficient can be encoded. In addition, here, the target block can be a CU. If the CU-level flag is 0, an adaptive multi-core transform can be applied to the target block.
[0096] In addition, if the target block to which the adaptive multi-core transform is applied is a luminance block, two additional flags may be signaled, and the vertical transform kernel and the horizontal transform kernel may be selected based on the flags. The flag of the vertical transform kernel may be represented as an AMT vertical flag, and AMT_TU_vertical_flag (or EMT_TU_vertical_flag) may represent a syntax element of the AMT vertical flag. The flag of the horizontal transform kernel may be represented as an AMT horizontal flag, and AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) may represent a syntax element of the AMT horizontal flag. The AMT vertical flag may indicate one of the transform kernel candidates included in the transform subset of the vertical transform kernel, and the transform kernel candidate indicated by the AMT vertical flag may be derived as the vertical transform kernel of the target block. In addition, the AMT horizontal flag may indicate one of the transform kernel candidates included in the transform subset of the horizontal transform kernel, and the transform kernel candidate indicated by the AMT horizontal flag may be derived as the horizontal transform kernel of the target block. Furthermore, the AMT vertical flag may be expressed as an MTS vertical flag, and the AMT horizontal flag may be expressed as an MTS horizontal flag.
[0097] Furthermore, the three transform subsets may be predetermined, and based on the intra prediction mode applied to the target block, one of the transform subsets may be derived as a transform subset for a vertical transform kernel. Furthermore, based on the intra prediction mode applied to the target block, one of the transform subsets may be derived as a transform subset for a horizontal transform kernel. For example, the predetermined transform subsets may be derived as shown in the following table.
[0098] [Table 2]
[0099] Transformation Set Transformation Candidates 0 DST-VII, DCT-VIII 1 DST-VII, DST-I 2 DST-VII, DCT-VIII
[0100] Reference Figure 2 , a transform subset with an index value of 0 may represent a transform subset including DST type 7 and DCT type 8 as transform core candidates, a transform subset with an index value of 1 may represent a transform subset including DST type 7 and DST type 1 as transform core candidates, and a transform subset with an index value of 2 may represent a transform subset including DST type 7 and DCT type 8 as transform core candidates.
[0101] A transform subset of a vertical transform kernel and a transform subset of a horizontal transform kernel derived based on an intra prediction mode applied to a target block may be derived as shown in the following table.
[0102] [Table 3]
[0103]
[0104] Here, V represents a transform subset of a vertical transform kernel, and H represents a transform subset of a horizontal transform kernel.
[0105] If the value of the AMT flag (or EMT_CU_flag) of the target block is 1, as shown in Table 3, the transform subset of the vertical transform kernel and the transform subset of the horizontal transform kernel can be derived based on the intra prediction mode of the target block. Thereafter, the transform kernel candidates included in the transform subset of the vertical transform kernel indicated by the AMT vertical flag of the target block can be derived as the vertical transform kernel of the target block, and the transform kernel candidates included in the transform subset of the horizontal transform kernel indicated by the AMT horizontal flag of the target block can be derived as the horizontal transform kernel of the target block. In addition, the AMT flag can be expressed as an MTS flag.
[0106] For reference, in the example, the intra-frame prediction mode may include two non-directional (or non-angle) intra-frame prediction modes and 65 directional (or angle) intra-frame prediction modes. The non-directional intra-frame prediction mode may include a plane intra-frame prediction mode No. 0 and a DC intra-frame prediction mode No. 1, and the directional intra-frame prediction mode may include 65 intra-frame prediction modes between an intra-frame prediction mode No. 2 and an intra-frame prediction mode No. 66. However, this is an example, and the present disclosure may be applied to cases where there are different numbers of intra-frame prediction modes. In addition, depending on the situation, an intra-frame prediction mode No. 67 may also be used, and the intra-frame prediction mode No. 67 may represent a linear model (LM) mode.
[0107] Figure 4 The intra directional mode exemplarily shows 65 prediction directions.
[0108] Reference Figure 4 , based on the intra prediction mode No. 34 having the upper left diagonal prediction direction, the intra prediction mode having the horizontal directionality and the intra prediction mode having the vertical directionality can be classified. Figure 4 H and V mean horizontal and vertical directivity, respectively, and numbers -32 to 32 indicate a shift in units of 1 / 32 at the sample grid position. Intra-frame prediction modes 2 to 33 have horizontal directivity, and intra-frame prediction modes 34 to 66 have vertical directivity. Intra-frame prediction mode 18 and intra-frame prediction mode 50 represent horizontal intra-frame prediction mode and vertical intra-frame prediction mode, respectively, and intra-frame prediction mode 2 can be called a lower-left diagonal intra-frame prediction mode; intra-frame prediction mode 34 can be called an upper-left diagonal intra-frame prediction mode; and intra-frame prediction mode 66 can be called an upper-right diagonal intra-frame prediction mode.
[0109] The converter can derive (secondary) transform coefficients (S320) by performing a secondary transform based on the (primary) transform coefficients. If the primary transform is a transform from the spatial domain to the frequency domain, the secondary transform can be regarded as a transform from the frequency domain to the frequency domain. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may represent a transform that generates transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix. At this time, the vertical transform and the horizontal transform are not applied separately to the (primary) transform coefficients (or the horizontal and vertical transforms are not applied independently), but the transform can be applied at one time based on the non-separable transform matrix. In other words, the non-separable secondary transform may represent a transform method that generates transform coefficients (or secondary transform coefficients) by transforming them together based on the non-separable transform matrix instead of separating the vertical component and the horizontal component of the (primary) transform coefficient. A non-separable secondary transform may be applied to the upper left region of a block configured with (primary) transform coefficients (hereinafter referred to as a transform coefficient block or a target block). For example, if both the width (W) and height (H) of the transform coefficient block are equal to or greater than 8, an 8×8 non-separable secondary transform may be applied to the upper left 8×8 region of the transform coefficient block (hereinafter referred to as the upper left target region). In addition, if both the width (W) and height (H) of the transform coefficient block are equal to or greater than 4 and the width (W) and height (H) of the transform coefficient block are less than 8, a 4×4 non-separable secondary transform may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0110] Specifically, for example, if a 4×4 input block is used, a non-separable secondary transform may be performed as follows.
[0111] A 4×4 input block X can be represented as follows.
[0112] [Formula 1]
[0113]
[0114] If X is represented in vector form, the vector can be represented as follows
[0115] [Formula 2]
[0116]
[0117] In this case, the quadratic non-separable transform can be calculated as follows.
[0118] [Formula 3]
[0119]
[0120] in, denotes a transform coefficient vector, and T denotes a 16x16 (non-separable) transform matrix.
[0121] By using the above formula 3, the 16×1 transform coefficient vector can be derived And can be scanned in order (horizontally, vertically, diagonally, etc.) Reorganize into 4×4 blocks. However, the above calculation is an example, and Hypercube-Givens Transform (HyGT) or the like may be used to calculate the non-separable quadratic transform in order to reduce the computational complexity of the non-separable quadratic transform.
[0122] Furthermore, in the non-separable secondary transform, the transform kernel (or transform core, transform type) may be selected so that it is mode-dependent. In this case, the mode may include an intra prediction mode and / or an inter prediction mode.
[0123] As described above, a non-separable secondary transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of a transform coefficient block. That is, a non-separable secondary transform can be performed based on an 8×8 sub-block size or a 4×4 sub-block size. For example, to select a mode-dependent transform kernel, 35 sets of three non-separable secondary transform kernels for the non-separable secondary transform can be configured for both the 8×8 sub-block size and the 4×4 sub-block size. That is, 35 transform sets can be configured for the 8×8 sub-block size, and 35 transform sets can be configured for the 4×4 sub-block size. In this case, each of the 35 transform sets for the 8×8 sub-block size can include three 8×8 transform kernels, and in this case, each of the 35 transform sets for the 4×4 sub-block size can include three 4×4 transform kernels. However, the transform sub-block size, the number of sets, and the number of transform kernels in the set are examples, and any size other than 8×8 or 4×4 can be used, or n sets can be configured, each set including k kernels.
[0124] The transform set may be referred to as an NSST set, and the transform core in the NSST set may be referred to as an NSST core.For example, selection of a specific set among the transform sets may be performed based on the intra prediction mode of the target block (CU or subblock).
[0125] In this case, for example, mapping between 35 transform sets and intra prediction modes may be represented as in the following table. For reference, if the LM mode is applied to the target block, the secondary transform may not be applied to the target block.
[0126] [Table 4]
[0127]
[0128] In addition, if it is determined that a specific set is to be used, one of the k transform cores in the specific set can be selected by the inseparable secondary transform index. The encoding device can derive an inseparable secondary transform index indicating a specific transform core based on a rate-distortion (RD) check, and can signal the inseparable secondary transform index to the decoding device. The decoding device can select one of the k transform cores in the specific set based on the inseparable secondary transform index. For example, an NSST index value of 0 can indicate a first inseparable secondary transform core, an NSST index value of 1 can indicate a second inseparable secondary transform core, and an NSST index value of 2 can indicate a third inseparable secondary transform core. Alternatively, an NSST index value of 0 can indicate that the first inseparable secondary transform is not applied to the target block, and NSST index values 1 to 3 can indicate three transform cores.
[0129] Return to reference Figure 3 The transformer can perform a non-separable secondary transform based on the selected transform kernel and obtain (secondary) transform coefficients. As described above, the transform coefficients can be derived as transform coefficients quantized by a quantizer and can be encoded and signaled to a decoding device and transmitted to an inverse quantizer / inverse transformer in an encoding device.
[0130] In addition, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device and transmitted to the inverse quantizer / inverse transformer in the encoding device.
[0131] The inverse transformer may perform a series of processes in the reverse order of the order in which they have been performed in the above-mentioned transformer. The inverse transformer may receive the (inverse quantized) transform coefficients and derive the (primary) transform coefficients by performing a secondary (inverse) transform (step S350), and may obtain a residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S360). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.
[0132] In addition, as described above, if the secondary (inverse) transform is omitted, the (inverse quantized) transform coefficients can be received, the primary (separable) transform can be performed, and the residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0133] Figures 5a to 5c is a diagram for explaining selective transformation according to an example of the present disclosure.
[0134] In this specification, the term "target block" may mean a current block or a residual block on which encoding is performed.
[0135] Figure 5a An example of deriving transform coefficients through transformation is shown.
[0136] Transforms in video coding can be represented by Figure 5a The process of transforming an input vector R by the transformation matrix shown in to generate a transform coefficient vector C for the input vector R. The input vector R can represent primary transform coefficients. Alternatively, the input vector R can represent a residual vector, i.e., residual samples. Furthermore, the transform coefficient vector C can be represented as an output vector C.
[0137] Figure 5b Represents a specific example of deriving transform coefficients through transformation. Figure 5b Specific expression Figure 5a The transformation process shown in FIG. Figure 3 As described in
[15] , in a non-separable secondary transform (hereinafter, referred to as "NSST"), after dividing an M×M block into block data of transform coefficients obtained by applying a primary transform, M can be performed on each M×M block. 2 ×M 2 NSST. For example, M can be 4 or 8, but is not limited thereto. M 2 It can be N. In this case, if Figure 5b As shown in FIG, the input vector R may be a vector comprising (primary) transform coefficients r1 to r N A (1×N)-dimensional vector, and the transform coefficient vector C may be a vector including transform coefficients c1 to c N That is, the input vector R may include N (primary) transform coefficients r1 to r N , and the size of the input vector R may be 1×N. In addition, the transform coefficient vector C may include N transform coefficients c1 to c N , and the size of the transform coefficient vector C can be N×1.
[0138] In order to derive the transformation coefficient vector C, the input vector R may be passed through a transformation matrix. That is, the input vector R may be derived into a transformation coefficient vector C based on the transformation matrix.
[0139] In addition, the transformation matrix may include N basis vectors B1 to B N .like Figure 5b As shown in N It can be a (1×N)-dimensional vector. That is, the basic vectors B1 to B NThe size of may be (1×N). A transform coefficient vector C may be generated based on the (primary) transform coefficient of the input vector R and each of the basis vectors of the transform matrix. For example, the inner product between the input vector and each of the basis vectors may be derived as the transform coefficient vector C.
[0140] Furthermore, in the above transformation, two main problems arise: in particular, the high computational complexity associated with the number of multiplications and additions required to generate the output vector and the memory requirement for storing the generated coefficients may appear as main problems.
[0141] For example, the computational complexity and storage requirements required for separable and non-separable transforms can be derived as shown in the following table.
[0142] [Table 5]
[0143] Separable Transform (NxN) Non-separable transform (NxN) storage <![CDATA[N 2 ]]> <![CDATA[N 4 ]]> multiplication <![CDATA[2N 3 ]]> <![CDATA[N 4 ]]>
[0144] Referring to Table 5, the memory required to store the coefficients generated by the separable transform may be N 2 , and the number of calculations can be 2N 3 The number of calculations indicates the computational complexity. In addition, the memory required to store the coefficients generated by the non-separable transform may be N 4 , and the number of calculations can be N 4 The number of calculations indicates the computational complexity. That is, the greater the number of calculations, the higher the computational complexity, and the fewer the number of calculations, the lower the computational complexity.
[0145] As shown in Table 5, the memory requirement and number of computations for a non-separable transform increase significantly compared to a separable transform. Furthermore, as the size of the target block on which the non-separable transform is performed increases, that is, as N becomes larger, the difference between the memory requirement and number of computations for a separable transform and the memory requirement and number of computations for a non-separable transform increases.
[0146] The non-separable transform provides better coding gain than the separable transform, but as shown in Table 5, due to the computational complexity of the non-separable transform, the non-separable transform is not used in the existing video coding standards, and because the computational complexity of the separable transform also increases with the increase of the size of the target block, in the existing HEVC standard, it is proposed to use the separable transform only in the target block size of 32×32 or smaller.
[0147] Therefore, the present disclosure proposes selective transformation. Selective transformation can greatly reduce computational complexity and storage requirements, thereby achieving effects such as increased efficiency of computationally intensive transformation blocks and improved coding efficiency. That is, selective transformation can be used to solve the computational complexity problem that arises when transforming inseparable transformations or large-sized blocks. Selective transformation can be used for any type of transformation such as primary transformation (or can be called kernel transformation), secondary transformation, etc. For example, selective transformation can be applied as a kernel transformation of an encoding device / decoding device, and can have the effect of greatly reducing encoding time / decoding time.
[0148] Figure 5c An example of deriving a transform coefficient through selective transform may be referred to as a transform performed on a target block based on a transform matrix including a basis vector having a selective number of elements.
[0149] Simplified transformation is a method that is proposed with the motivation of reducing computational complexity and storage requirements by excluding redundant or unimportant elements among the N elements of the basis vectors of the transformation matrix. Figure 5c , among the N elements of the basis vector B1, the Z0 elements may not be important elements. In this case, a truncated basis vector B1 including only N1 elements may be derived. Here, N1 may be N-Z0. The truncated basis vector B1 may be expressed as a modified basis vector B1.
[0150] Reference Figure 5c If the modified basis vector B1 is applied to the input vector R as part of the transformation matrix, the transformation coefficient C1 can be derived. Experimental results show that the transformation coefficient C1 is the same value as the transformation coefficient C1 derived by applying the existing basis vector B1 as part of the transformation matrix to the input vector R. That is, by deriving the result assuming that the unimportant elements of each basis vector are 0, the number of unnecessary multiplications can be significantly reduced without much difference in the result. Furthermore, the number of elements that must be stored for the calculation (i.e., elements of the transformation matrix) can be reduced.
[0151] In order to define the position of elements among the elements of the unimportant (or significant) basis vectors, an association vector is proposed. To derive the modified N1-dimensional basis vector B1, a (1×N)-dimensional association vector A1 can be considered. That is, to derive a 1×N1-sized modified basis vector B1 (i.e., a modified basis vector B1 including N1 elements), a 1×N-sized association vector A1 can be considered.
[0152] Reference Figure 5c , by applying theN The value derived from the associated vector of each of the input vectors R can be transferred to the base vector. Thus, the elements of the base vector can be calculated using only some elements of the input vector R. Specifically, the associated vector may include 0 and 1, and an operation is performed such that the elements selected from the elements of the input vector R are multiplied by 1 and the elements not selected therefrom are multiplied by 0, so that only the selected elements can pass through it and be transferred to the base vector.
[0153] For example, the association vector A1 may be applied to the input vector R, and the inner product with the base vector B1 may be calculated using only N1 elements of the input vector R that have been specified by the association vector A1. The inner product may represent C1 of the transform coefficient vector C. In this regard, the base vector B1 may include N1 elements, and N1 may be N or less. N and basis vectors B2 to B N The above-mentioned operation of the associated vector A1 and the base vector B1 with the input vector R is performed.
[0154] Since the association vector only includes binary values of 0 and / or 1, it is advantageous to store the association vector. A 0 in the association vector may indicate that the elements of the input vector R for 0 are not transferred to the transformation matrix for inner product calculation, and a 1 in the association vector may indicate that the elements of the input vector R for 1 are transferred to the transformation matrix for inner product calculation. For example, an association vector A of size 1×N k Can include A k1 To A kn If the associated vector A k A kn If it is 0, the input vector R is not allowed. n That is, the r of the input vector R n May not be transferred to the transformation vector B n In addition, if A kn If it is 1, then the input vector R is allowed to n That is, the r of the input vector R n Can be transferred to the transformation vector B n and can be used to derive the transform coefficient vector C c n used in the calculation of .
[0155] Figure 6 Schematic representation of a multiple transform technique where selective transforms are applied as secondary transforms.
[0156] Reference Figure 6 , the converter can correspond to the above Figure 1 The converter in the encoding device, and the inverse converter may correspond to the above Figure 1 Inverse transformer in the encoding device or above Figure 2 An inverse transformer in a decoding device.
[0157] The transformer may derive (primary) transform coefficients by performing a primary transform based on residual samples (residual sample array) in the residual block (S610). Here, the first transform may include the aforementioned AMT.
[0158] If an adaptive multi-core transform is applied, a transform from the spatial domain to the frequency domain can be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8 and / or DST type 1 to generate transform coefficients (or primary transform coefficients). In this regard, from the perspective of the transformer, the primary transform coefficients can be referred to as temporary transform coefficients. In addition, DCT type 2, DST type 7, DCT type 8 and DST type 1 can be referred to as transform types, transform cores or transform cores. For reference, the DCT / DST transform type can be defined based on a basis function, and the basis function can be represented as in Table 1 above. Specifically, the process of deriving the primary transform coefficients by applying an adaptive multi-core transform is as described above.
[0159] The transformer may derive (secondary) transform coefficients by performing selective transform based on the (primary) transform coefficients (S620). Selective transform may mean a transform performed on the (primary) transform coefficients of the target block based on a transform matrix including a modified base vector and an association matrix including an association vector of the base vector. The modified base vector may represent a base vector including N or fewer elements. That is, the modified base vector may represent a base vector including a specific number of elements selected from the N elements. For example, the modified base vector B n It can be (1×N n )-dimensional vector, and N n Can be less than or equal to N. That is, the modified basis vector B n It can be (1×N n ) size, and N n It may be less than or equal to N. Here, N may be the product of the height and width of the upper left target area of the target block to which the selective transformation is applied. Alternatively, N may be the total number of transform coefficients of the upper left target area of the target block to which the selective transformation is applied. In addition, the transform matrix including the modified base vectors may be expressed as a modified transform matrix. In addition, the transform matrix may be expressed as a transform basis block (TBB), and the association matrix may be expressed as an association vector block (AVB).
[0160] The transformer can perform a selective transform based on the modified transform matrix and the correlation matrix and obtain (secondary) transform coefficients. As described above, the transform coefficients can be derived as transform coefficients quantized by a quantizer, and can be encoded and signaled to a decoding device and transmitted to an inverse quantizer / inverse transformer in the encoding device.
[0161] The inverse transformer may perform a series of processes in the reverse order of the order in which they have been performed in the above-mentioned transformer. The inverse transformer may receive the (inverse quantized) transform coefficients and derive the (primary) transform coefficients by performing a selective (inverse) transform (step S650), and may obtain a residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S660). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.
[0162] Furthermore, the present disclosure proposes selective transformation combined with simplified transformation as an example of selective transformation.
[0163] In the present specification, the term 'simplification transform' may mean a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a simplification factor.
[0164] In the simplified transformation according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space so that a simplified transformation matrix can be determined, wherein R is less than N. That is, the simplified transformation may mean a transformation performed on the residual samples of the target block based on a simplified transformation matrix including R basis vectors. Here, N may mean the square of the length of the side of the block (or target area) to which the transformation is applied, or the total number of transformation coefficients corresponding to the block (or target area) to which the transformation is applied, and the simplification factor may mean an R / N value. The simplification factor may be referred to as various terms such as a simplified factor, a reduction factor, a simplified factor, a simple factor, and others. In addition, R may be referred to as a simplified coefficient, but depending on the situation, the simplification factor may mean R. In addition, depending on the situation, the simplification factor may mean an N / R value.
[0165] The size of the simplified transformation matrix according to an example may be R×N, which is smaller than the size of the conventional transformation matrix N×N, and may be defined as in Equation 4 below.
[0166] [Formula 4]
[0167]
[0168] If the simplified transformation matrix T RxNBy multiplying the transform coefficients of the primary transform of the target block to which the primary transform has been applied, the (secondary) transform coefficients of the target block can be derived.
[0169] If RST is applied, since a simplified transform matrix of size R×N is applied to the secondary transform, transform coefficients R+1 to N may implicitly become 0. In other words, if the transform coefficients of the target block are derived by applying RST, the values of transform coefficients R+1 to N may be 0. Here, transform coefficients R+1 to N may represent the R+1th to Nth transform coefficients among the transform coefficients. Specifically, the arrangement of the transform coefficients of the target block can be explained as follows.
[0170] Figure 7 is a diagram for explaining the arrangement of transform coefficients based on a target block according to an example of the present disclosure. Figure 7 The explanation of the transformation described in can be similarly applied to the inverse transformation. NSST based on the primary transformation and the simplified transformation can be performed on the target block (or residual block) 700. In the example, Figure 7 The 16×16 block shown in FIG7 may represent the target block 700, and the 4×4 blocks labeled A to P may represent subblocks of the target block 700. A primary transform may be performed on the entire range of the target block 700, and after the primary transform has been performed, NSST may be applied to an 8×8 block consisting of subgroups A, B, E, and F (hereinafter, the upper left target region). At this time, if NSST based on a simplified transform is performed, since only R (herein, R means a simplified coefficient, and R is less than N) NSST transform coefficients are derived, the R+1th NSST transform coefficient to the Nth NSST transform coefficient may be determined to be 0. For example, if R is 16, the 16 transform coefficients derived by performing NSST based on the simplified transform may be assigned to each block included in subgroup A, which is the upper left 4×4 block included in the upper left target region of the target block 700, and a transform coefficient of 0 may be assigned to each of NR (i.e., 64-16=48) blocks included in subgroups B, E, and F. The primary transform coefficients for which simplified transform-based NSST has not been performed may be allocated to the respective blocks included in subgroups C, D, G, H, I, J, K, L, M, N, O, and P.
[0171] Figure 8 This shows an example of deriving transform coefficients by combining a simplified transform with a selective transform. Figure 8 , the transformation matrix may include R basis vectors, and the incidence matrix may include R incidence vectors. Here, the transformation matrix including R basis vectors may be expressed as a simplified transformation matrix, and the incidence matrix including R incidence vectors may be expressed as a simplified incidence matrix.
[0172] In addition, each of the basis vectors may include only elements selected from N elements. Figure 8 , the base vector B1 may be a 1×N1-dimensional vector including N1 elements, the base vector B2 may be a 1×N2-dimensional vector including N2 elements, and the base vector B R Can include N R 1×N elements R dimensional vector. N1, N2 and N R It may be a value equal to or smaller than N. Transformation in which simplified transformation and selective transformation are combined with each other may be used for any type of transformation such as secondary transformation, primary transformation, and the like.
[0173] Reference Figure 8 , the encoding device and the decoding device may apply a transformation that combines a simplified transformation and a selective transformation as a secondary transformation. For example, a selective transformation may be performed based on a simplified transformation matrix including modified basis vectors and a simplified incidence matrix, and (secondary) transformation coefficients may be obtained. In addition, in order to reduce the computational complexity of the selective transformation, a hypercube-Givens transform (HyGT) or the like may be used for the calculation of the selective transformation.
[0174] Figure 9 In the example of selective transformation, there may be no correlation vectors A1, A2, ... A for the correlation matrix. N Alternatively, in another example of selective transformation, the correlation vectors A1, A2, ... A of the correlation matrix N can be derived in the same form.
[0175] Specifically, for example, each associated vector may have the same number of 1s. For example, if the number of 1s is M, the associated vector may have M 1s and (NM) 0s. In this case, M transform coefficients among the (primary) transform coefficients of the input vector R may be transferred to the base vector. Therefore, the length of the base vector may also be M. That is, Figure 9 As shown in , each basis vector may include M elements, may be a (1×M)-dimensional vector, and may be derived as N1=N2=...=N N =M. Figure 9 The incidence matrix and modified transformation matrix architecture shown in can be expressed as a symmetric architecture using selective transformations.
[0176] Additionally, in another example, there may be a specific pattern of elements with a value of 1, and this pattern may be repeated, rotated, and / or translated in an arbitrary manner such that an associated vector may be derived.
[0177] Furthermore, the above-described selective transformations may be applied together with other transformation techniques as well as simplified transformations and / or HyGT.
[0178] Furthermore, the present disclosure proposes a method for simplifying the correlation vectors used in the aforementioned selective transformation. By simplifying the correlation vectors, the storage of information required to perform the selective transformation and the processing of the selective transformation can be further improved. Specifically, the storage burden of performing the selective transformation can be reduced, and the processing capability of the selective transformation can be further improved.
[0179] If the non-zero elements in the correlation vector have a continuous distribution, the effect of simplifying the correlation vector can be more clearly shown. k It can include a continuous string of 1s. In this case, it can be obtained by two factors A ks and A kL To represent A k That is, based on factor A ks and A kL To derive the correlation vector A k Here, A ks can be a factor representing the starting point of a non-zero element (e.g., 1), and A kL It can be a factor representing the length of the non-zero element. The associated vector A represented by these factors can be derived as shown in the following table: k .
[0180] [Table 6]
[0181]
[0182] Referring to Table 6, the correlation vector A k It can contain 16 elements. That is, the associated vector A k It can be a 1×16 dimensional vector. It represents the associated vector A k The factor of the starting point of the non-zero elements in A ks The value of can be deduced to be 0, and in this case, the factor A ks The associated vector A can be indicated k The first element of is the starting point of non-zero elements. In addition, it represents the associated vector A k Factors of the lengths of the non-zero elements in A kL The value of can be derived as 8, and in this case, the factor A kL It can be indicated that the length of the non-zero elements is 8. Therefore, as shown in Table 6, A can be divided based on this factor. k It is derived as a vector whose first to eighth elements are 1 and the rest are 0.
[0183] Figure 10An example of performing a selective transformation by deriving an association vector based on two factors of the association vector is shown. Figure 10 , an association matrix may be derived based on the factors of each of the association vectors, and a transform coefficient of the target block may be derived based on the association matrix and the modified transform matrix.
[0184] Furthermore, the start point of non-zero elements in each of the associated vectors and the number of non-zero elements may be derived as fixed values, or the start point of non-zero elements in each of the associated vectors and the number of non-zero elements may be derived in various ways.
[0185] In addition, for example, the starting point of the non-zero elements in the association vector and the number of non-zero elements can be derived based on the size of the upper left target region on which the transformation is performed. Here, the size of the upper left target region can represent the number of transform coefficients of the upper left target region, or can represent the product of the height and width of the upper left target region. In another example, the starting point of the non-zero elements in the association vector and the number of non-zero elements can be derived based on the intra-frame prediction mode of the target block. Specifically, for example, the starting point of the non-zero elements in the association vector and the number of non-zero elements can be derived based on whether the intra-frame prediction mode of the target block is a non-directional intra-frame prediction mode.
[0186] Alternatively, for example, the starting point of the non-zero elements in the association vector and the number of non-zero elements may be predetermined. Alternatively, for example, information indicating the starting point of the non-zero elements in the association vector and information indicating the number of non-zero elements may be signaled, and the association vector may be derived based on the information indicating the starting point of the non-zero elements and the information indicating the number of non-zero elements. Alternatively, other information may be used in place of the information indicating the starting point of the non-zero elements. For example, instead of the information indicating the starting point of the non-zero elements, information indicating the last position of the non-zero elements may be used, and the association vector may be derived based on this information indicating the last position of the non-zero elements.
[0187] Furthermore, the method of deriving the correlation vector based on these factors can also be applied to separable transforms and inseparable transforms such as the simplified transform HyGT.
[0188] Figure 11 The following schematically illustrates an image encoding method performed by an encoding device according to the present disclosure. Figure 11 The method disclosed in Figure 1 Specifically, for example, Figure 11 S1100 of the encoding device may be performed by a subtractor; S1110 may be performed by a transformer of the encoding device; and S1120 may be performed by an entropy encoder of the encoding device. In addition, although not shown, the process of deriving the prediction sample may be performed by a predictor of the encoding device.
[0189] The encoding device derives residual samples of the target block (step S1100). For example, the encoding device may determine whether to perform inter-frame prediction or intra-frame prediction on the target block, and may determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. Based on the determined mode, the encoding device may derive predicted samples of the target block, and may derive residual samples by adding the original samples of the target block to the predicted samples.
[0190] The encoding device derives the transform coefficients of the target block based on the selective transform of the residual samples (step S1110). The selective transform can be performed based on the modified transform matrix, which is a matrix including the modified basis vectors, and the modified basis vectors can include a specific number of elements selected from N elements. In addition, the selective transform can be performed on the upper left target area of the target block, and N can be the number of residual samples located in the upper left target area. In addition, N can be a value obtained by multiplying the width and height of the upper left target area. For example, N can be 16 or 64.
[0191] The encoding device can derive modified transform coefficients by performing a kernel transform on the residual samples, and can derive transform coefficients of the target block by performing a selective transform on the modified transform coefficients located in the upper left target area of the target block based on an association matrix including an association vector of the modified base vector and a modified transform matrix.
[0192] Specifically, the kernel transform of the residual samples can be performed as follows. The encoding device can determine whether to apply an adaptive multi-kernel transform (AMT) to the target block. In this case, an AMT flag indicating whether to apply the adaptive multi-kernel transform of the target block can be generated. If the AMT is not applied to the target block, the encoding device can derive DCT type 2 as the transform kernel for the target block and can derive modified transform coefficients by performing a transform on the residual samples based on DCT type 2.
[0193] If AMT is applied to the target block, the encoding device can configure a transform subset of a horizontal transform kernel and a transform subset of a vertical transform kernel, can derive the horizontal transform kernel and the vertical transform kernel based on the transform subset, and can derive the modified transform coefficient by performing a transform on the residual sample based on the horizontal transform kernel and the vertical transform kernel. At this point, the transform subset of the horizontal transform kernel and the transform subset of the vertical transform kernel can include DCT type 2, DST type 7, DCT type 8 and / or DST type 1 as candidates. In addition, transform index information can be generated, and the transform index information can include an AMT horizontal flag indicating the horizontal transform kernel and an AMT vertical flag indicating the vertical transform kernel. In addition, the transform kernel can be referred to as a transform type or a transform core.
[0194] If the modified transform coefficient is derived, the encoding apparatus may derive the transform coefficient of the target block by performing selective transform on the modified transform coefficient located in the upper left target area of the target block based on the correlation matrix including the correlation vector of the modified basis vector and the modified transform matrix. Other modified transform coefficients except the modified transform coefficient located in the upper left area of the target block may be derived as the transform coefficient of the target block without any change.
[0195] Specifically, the modified transform coefficient of the element whose associated vector has a value of 1 among the modified transform coefficients located in the upper left target area can be derived, and the transform coefficient of the target block can be derived based on the modified base vector and the derived modified transform coefficient. In this regard, the associated vector of the modified base vector may include N elements, the N elements may include elements with a value of 1 and / or elements with a value of 0, and the number of elements with a value of 1 may be A. In addition, the modified base vector may include A elements.
[0196] Furthermore, in one example, the modified transformation matrix may include N modified basis vectors, and the association matrix may include N association vectors. The association vectors may include the same number of elements having a value of 1, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of elements having a value of 1, and the modified basis vectors may not all include the same number of elements.
[0197] Alternatively, in another example, the modified transformation matrix may include R modified basis vectors, and the association matrix may include R association vectors. R may be a simplified coefficient, and R may be less than N. The association vectors may include the same number of elements with a value of 1, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of elements with a value of 1, and the modified basis vectors may not all include the same number of elements.
[0198] Furthermore, the association vector may be configured so that elements having a value of 1 are arranged continuously. In this case, in one example, information about the association vector may be entropy-encoded. For example, the information about the association vector may include information indicating the starting point of the element having a value of 1 and information indicating the number of elements having a value of 1. Alternatively, for example, the information about the association vector may include information indicating the last position of the element having a value of 1 and information indicating the number of elements having a value of 1.
[0199] In another example, the association vector may be derived based on the size of the upper left target area. For example, the starting point and the number of elements with a value of 1 in the association vector may be derived based on the size of the upper left target area.
[0200] Alternatively, in another example, the association vector may be derived based on the intra prediction mode of the target block. For example, the starting point of the element with a value of 1 and the number of elements with a value of 1 in the association vector may be derived based on the intra prediction mode. Furthermore, for example, the starting point of the element with a value of 1 and the number of elements with a value of 1 in the association vector may be derived based on whether the intra prediction mode is a non-directional intra prediction mode.
[0201] The encoding device encodes information about the transform coefficients (S1330). The information about the transform coefficients may include information about the size and position of the transform coefficients. In addition, as described above, the information about the association vector may be entropy encoded. For example, the information about the association vector may include information indicating the starting point of an element with a value of 1 and information indicating the number of elements with a value of 1. Alternatively, for example, the information about the association vector may include information indicating the last position of an element with a value of 1 and information indicating the number of elements with a value of 1.
[0202] Image information including information about transform coefficients and / or information about correlation vectors may be output in the form of a bitstream. Furthermore, the image information may further include prediction information. The prediction information may include information about motion information (for example, when inter-frame prediction is applied) and prediction mode information as a plurality of information related to the prediction process.
[0203] The output bit stream may be transmitted to a decoding device via a storage medium or a network.
[0204] Figure 12 The figure schematically shows an encoding device for performing an image encoding method according to the present disclosure. Figure 11 The method disclosed in Figure 12 Specifically, for example, Figure 12 The adder of the encoding device can perform Figure 11 In step S1100, the transformer of the encoding device may perform S1110, and the entropy encoder of the encoding device may perform S1120 to S1130. In addition, although not shown, the process of deriving the prediction sample may be performed by the predictor of the encoding device.
[0205] Figure 13 The figure schematically shows an image decoding method performed by a decoding device according to the present disclosure. Figure 13 The method disclosed in Figure 2Specifically, for example, the entropy decoder of the decoding device may be used to perform Figure 13 S1300 to S1310 in the decoding apparatus; S1320 may be performed by the inverse transformer of the decoding apparatus; and S1330 may be performed by the adder of the decoding apparatus. In addition, although not shown, the process of deriving the prediction sample may be performed by the predictor of the decoding apparatus.
[0206] The decoding apparatus derives the transform coefficients of the target block from the bitstream (S1300). The decoding apparatus may derive the transform coefficients of the target block by decoding information about the transform coefficients of the target block received through the bitstream. The received information about the transform coefficients of the target block may be represented as residual information.
[0207] The decoding device derives residual samples of the target block based on the selective transformation of the transform coefficients (S1310). The selective transformation can be performed based on the modified transform matrix, which is a matrix including the modified basis vectors, and the modified basis vectors can include a specific number of elements selected from N elements. In addition, the selective transformation can be performed on the transform coefficients located in the upper left target area of the target block, and N can be the number of transform coefficients located in the upper left target area. In addition, N can be a value obtained by multiplying the width and height of the upper left target area. For example, N can be 16 or 64.
[0208] The decoding apparatus may derive the modified transform coefficient by performing selective transform on the transform coefficient located in the upper left target region of the target block based on the modified transform matrix and an association matrix including the association vectors of the modified basis vectors.
[0209] Specifically, the transform coefficients of the elements of the associated vector with a value of 1 among the transform coefficients located in the upper left target area can be derived, and the modified transform coefficients can be derived based on the derived transform coefficients and the modified base vector. In this regard, the associated vector of the modified base vector may include N elements, the N elements may include elements with a value of 1 and / or elements with a value of 0, and the number of elements with a value of 1 may be A. In addition, the modified base vector may include A elements.
[0210] Furthermore, in one example, the modified transformation matrix may include N modified basis vectors, and the association matrix may include N association vectors. The association vectors may include the same number of elements having a value of 1, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of elements having a value of 1, and the modified basis vectors may not all include the same number of elements.
[0211] Alternatively, in another example, the modified transformation matrix may include R modified basis vectors, and the association matrix may include R association vectors. R may be a simplified coefficient, and R may be less than N. The association vectors may include the same number of elements with a value of 1, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of elements with a value of 1, and the modified basis vectors may not all include the same number of elements.
[0212] Furthermore, the association vector may be configured so that elements having a value of 1 are arranged continuously. In this case, in one example, information about the association vector may be obtained from the bitstream, and the association vector may be derived based on the information about the association vector. For example, the information about the association vector may include information indicating the starting point of the element having a value of 1 and information indicating the number of elements having a value of 1. Alternatively, for example, the information about the association vector may include information indicating the last position of the element having a value of 1 and information indicating the number of elements having a value of 1.
[0213] Alternatively, in another example, the association vector may be derived based on the size of the upper left target area. For example, the starting point of the element with a value of 1 and the number of elements with a value of 1 in the association vector may be derived based on the size of the upper left target area.
[0214] Alternatively, in another example, the association vector may be derived based on the intra prediction mode of the target block. For example, the starting point of the element with a value of 1 and the number of elements with a value of 1 in the association vector may be derived based on the intra prediction mode. Furthermore, for example, the starting point of the element with a value of 1 and the number of elements with a value of 1 in the association vector may be derived based on whether the intra prediction mode is a non-directional intra prediction mode.
[0215] If the modified transform coefficient is derived, the decoding apparatus may induce residual samples by performing kernel transform on the target block including the modified transform coefficient.
[0216] The kernel transform of the target block may be performed as follows: the decoding apparatus may obtain an adaptive multi-kernel transform (AMT) flag indicating whether an AMT is applied from a bitstream, and if the value of the AMT flag is 0, the decoding apparatus may derive DCT type 2 as a transform kernel of the target block, and may derive residual samples by performing an inverse transform on the target block including the modified transform coefficient based on DCT type 2.
[0217] If the value of the AMT flag is 1, the decoding device can configure a transform subset of a horizontal transform kernel and a transform subset of a vertical transform kernel, can derive the horizontal transform kernel and the vertical transform kernel based on the transform subset and the transform index information obtained from the bitstream, and can derive residual samples by performing an inverse transform on a target block including modified transform coefficients based on the horizontal transform kernel and the vertical transform kernel. At this point, the transform subset of the horizontal transform kernel and the transform subset of the vertical transform kernel may include DCT type 2, DST type 7, DCT type 8 and / or DST type 1 as candidates. In addition, the transform index information may include an AMT horizontal flag indicating one of the candidates included in the transform subset of the horizontal transform kernel and an AMT vertical flag indicating one of the candidates included in the transform subset of the vertical transform kernel. In addition, the transform kernel may be referred to as a transform type or a transform core.
[0218] The decoding device generates a reconstructed picture based on the residual samples (S1320). The decoding device may generate the reconstructed picture based on the residual samples. For example, the decoding device may perform inter-frame prediction or intra-frame prediction on the target block based on the prediction information received through the bitstream, may derive prediction samples, and may generate a reconstructed picture by adding the prediction samples to the residual samples. Thereafter, as described above, a loop filtering process such as an ALF process, SAO and / or deblocking filtering may be applied to the reconstructed picture as needed to improve subjective / objective video quality.
[0219] Figure 14 A decoding device for performing an image decoding method according to the present disclosure is schematically shown. Figure 13 The method disclosed in Figure 14 Specifically, for example, Figure 14 The entropy decoder of the decoding device can perform Figure 13 S1300 in; Figure 14 The inverse transformer of the decoding device can perform Figure 13 S1310 in; and Figure 14 The adder of the decoding device can perform Figure 13 In addition, although not shown, it is possible to Figure 14 The predictor of the decoding device performs a process of deriving prediction samples.
[0220] According to the present disclosure described above, through efficient transformation, the amount of data that must be transmitted for residual processing can be reduced, and residual coding efficiency can be increased.
[0221] In addition, according to the present disclosure, an inseparable transform can be performed based on a transform matrix composed of basis vectors including a specific number of selected elements, thereby reducing the storage burden and computational complexity of the inseparable transform and increasing the residual coding efficiency.
[0222] In addition, according to the present disclosure, a non-separable transform can be performed based on a transform matrix of a simplified architecture, thereby reducing the amount of data that must be transmitted for residual processing and increasing residual coding efficiency.
[0223] In the above embodiments, the method is explained based on a flowchart with the aid of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may occur in an order or step different from the above order or step, or may occur simultaneously with another step. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive, and another step or one or more steps of the flowchart that may be incorporated may be removed without affecting the scope of the present disclosure.
[0224] The above-mentioned method according to the present disclosure may be implemented in a software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0225] When the embodiments in the present disclosure are implemented by software, the above methods can be implemented as modules (processes, functions, etc.) to perform the above functions. The modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller or a chip. In addition, the functional units shown in each figure can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip.
[0226] In addition, the decoding device and the encoding device to which the present disclosure is applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), and the like.
[0227] In addition, the processing method of the present invention can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. The multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include all kinds of storage devices and distributed storage devices in which computer-readable data are stored. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium, or can be transmitted through a wired or wireless communication network. In addition, the embodiments of the present invention can be implemented as a computer program product through program code, and the program code can be executed on a computer through the embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0228] In addition, the content streaming system to which the present disclosure is applied may mainly include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0229] The encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and transmit it to the streaming server. As another example, if the multimedia input device such as a smartphone, camera, and camcorder directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. Furthermore, the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0230] The streaming server transmits multimedia data to a user device via a network server based on a user's request. The network server serves as a tool for informing the user of available services. When the user requests a desired service, the network server transmits it to the streaming server, which then transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, in which case the control server is used to control commands and responses between the various devices in the content streaming system.
[0231] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to provide a streaming service smoothly.
[0232] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touch-screen PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each of the servers in the content streaming system may operate as a distributed server, and in this case, data received by each server may be processed in a distributed manner.
Claims
1. A decoding device for image decoding, the decoding device comprising: Memory; as well as at least one processor connected to the memory, the at least one processor configured to: Derives inverse quantized transform coefficients of the target block from the bitstream; deriving primary transform coefficients of the target block based on a non-separable transform of the inverse quantized transform coefficients; deriving residual samples of the target block based on a separable transform of the primary transform coefficients; and generating a reconstructed picture based on the residual samples of the target block and the predicted samples of the target block, Wherein, the inseparable transformation is performed based on the transformation matrix, Wherein, the transformation matrix includes the modified basis vectors, wherein each of the modified basis vectors comprises fewer than N elements, wherein N is equal to the number of inversely quantized transform coefficients located in a region of the target block to which the inseparable transform is applied, wherein the region to which the non-separable transform is applied is an 8×8 upper left target region in the target block, and N is equal to 64, and The number of the modified basis vectors is less than N.
2. The decoding device according to claim 1, wherein performing the non-separable transform on the inversely quantized transform coefficients located in the 8×8 upper left target area based on the transform matrix and an association matrix including an association vector of the modified basis vector, The association vector includes M elements with a value of 1 and NM elements with a value of 0 in sequence. wherein the M elements among the N transform coefficients located in the 8×8 upper left target region are derived by multiplying the elements by a value of 1 in an associated vector, and The primary transform coefficients are derived by performing the inseparable transform on the M elements among the N transform coefficients located in the 8×8 upper left target area.
3. The decoding device according to claim 2, wherein The modified basis vector includes M elements.
4. The decoding device according to claim 2, wherein: The transformation matrix includes N modified basis vectors, and the incidence matrix includes N incidence vectors.
5. The decoding device according to claim 4, wherein The association vectors include the same number of elements having a value of 1, and the modified basis vectors all include the same number of elements.
6. The decoding device according to claim 2, wherein The transformation matrix includes R modified basis vectors, and the incidence matrix includes R incidence vectors, and R is less than N.
7. The decoding device according to claim 2, wherein: obtaining information about the association vector from the bitstream and deriving the association vector based on the information about the association vector, and The information about the association vector includes information indicating the starting point of the element whose value is 1 and information indicating the number of elements whose value is 1.
8. A coding device for image coding, the coding device comprising: Memory; as well as at least one processor connected to the memory, the at least one processor configured to: Derived residual samples of the target block; deriving primary transform coefficients of the target block based on a separable transform of the residual samples; deriving secondary transform coefficients of the target block based on the non-separable transform of the primary transform coefficients; deriving quantized transform coefficients of the target block based on the secondary transform; as well as encoding information about the quantized transform coefficients, Wherein, the inseparable transformation is performed based on the transformation matrix, Wherein, the transformation matrix includes the modified basis vectors, wherein each of the modified basis vectors comprises fewer than N elements, wherein N is equal to the number of primary transform coefficients located in a region of the target block to which the non-separable transform is applied, wherein the region to which the non-separable transform is applied is an 8×8 upper left target region in the target block, and N is equal to 64, and The number of the modified basis vectors is less than N.
9. The encoding device according to claim 8, wherein performing the non-separable transform on the primary transform coefficients located in the 8×8 upper left target area of the target block based on the transform matrix and an association matrix including the association vectors of the modified basis vectors, The association vector includes M elements with a value of 1 and NM elements with a value of 0 in sequence. wherein the M elements among the N transform coefficients located in the 8×8 upper left target region are derived by multiplying the elements by a value of 1 in an associated vector, and The secondary transform coefficients are derived by performing the inseparable transform on the M elements among the N transform coefficients located in the 8×8 upper left target area.
10. The encoding device according to claim 9, wherein The transformation matrix includes R modified basis vectors, and the incidence matrix includes R incidence vectors, and R is less than N.
11. The encoding device according to claim 9, wherein The information about the association vector is entropy-coded, and includes information indicating a start point of an element having a value of 1 and information indicating the number of elements having a value of 1.
12. A device for transmitting data for an image, the device comprising: at least one processor configured to obtain a bitstream for an image, wherein the bitstream is generated based on the following steps: deriving residual samples of a target block, deriving primary transform coefficients of the target block by performing a separable transform on the residual samples, deriving secondary transform coefficients of the target block based on a non-separable transform on the primary transform coefficients, deriving quantized transform coefficients of the target block based on the secondary transform, and encoding information about the quantized transform coefficients; and a transmitter configured to transmit data comprising the bit stream, Wherein, the inseparable transformation is performed based on the transformation matrix, Wherein, the transformation matrix includes the modified basis vectors, wherein each of the modified basis vectors comprises fewer than N elements, wherein N is equal to the number of primary transform coefficients located in a region of the target block to which the non-separable transform is applied, wherein the region to which the non-separable transform is applied is an 8×8 upper left target region in the target block, and N is equal to 64, and The number of the modified basis vectors is less than N.
Citation Information
Patent Citations
Method for coding and reconstructing a pixel block and corresponding devices
CN103002279A
Vidio signal encoding method
CN104378637A