Method for picture coding based on selective conversion and device therefor
The selective transform method in video coding reduces data transmission and computational complexity by using a modified transform matrix, enhancing residual coding efficiency and addressing the high-cost issue of high-resolution video storage and transmission.
Patent Information
- Application Number
- JP2025120336
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-21
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-11
- Estimated Expiration
- 2038-12-21
AI Technical Summary
The increasing demand for high-resolution, high-quality video leads to higher transmission and storage costs due to increased video data, necessitating more efficient video coding techniques.
A video decoding method and apparatus that employs a selective transform using a modified transform matrix with specific basis vectors to reduce data transmission and computational complexity, improving residual coding efficiency.
The method reduces the amount of data required for residual processing and lowers memory and computational load by utilizing a non-separable transform with a simplified structure.
Smart Images

Figure 2025134049000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to video coding techniques, and more particularly to a method and apparatus for video decoding according to selective transformation in a video coding system. [Background technology]
[0002] Recently, demand for high-resolution, high-quality video, such as HD (High Definition) video and UHD (Ultra High Definition) video, is increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relative to existing video data. Therefore, when video data is transmitted using a medium such as an existing wired or wireless broadband line or when video data is stored using an existing storage medium, transmission costs and storage costs increase.
[0003] This requires highly efficient video compression techniques to effectively transmit, store, and play back high-resolution, high-quality video information. Summary of the Invention [Problem to be solved by the invention]
[0004] SUMMARY OF THE INVENTION A technical object of the present invention is to provide a method and apparatus for improving video coding efficiency.
[0005] Another technical object of the present invention is to provide a method and apparatus for increasing the conversion efficiency.
[0006] It is still another technical object of the present invention to provide a method and apparatus for improving the efficiency of residual coding through transforms.
[0007] It is still another technical object of the present invention to provide a video coding method and apparatus based on selective transform. [Means for solving the problem]
[0008] According to an embodiment of the present invention, there is provided a video decoding method performed by a decoding device, the method including the steps of: deriving transform coefficients of a current block from a bitstream; deriving residual samples for the current block based on a selective transform of the transform coefficients; and generating a reconstructed picture based on the residual samples for the current block and the prediction samples for the current block, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix being a matrix having modified basis vectors, and the modified basis vectors having a selected specific number of elements from N elements.
[0009] According to another embodiment of the present invention, there is provided a decoding device for decoding video, the decoding device including: an entropy decoding unit for deriving transform coefficients of a current block from a bitstream; an inverse transform unit for deriving residual samples for the current block based on a selective transform of the transform coefficients; and an adder unit for generating a reconstructed picture based on the residual samples for the current block and prediction samples for the current block, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix being a matrix having modified basis vectors, and the modified basis vectors having a specific number of elements selected from N elements.
[0010] According to yet another embodiment of the present invention, there is provided a video encoding method performed by an encoding device, the method comprising: deriving residual samples of a current block; deriving transform coefficients of the current block based on a selective transform of the residual samples; and encoding information about the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix being a matrix having modified basis vectors, and the modified basis vectors having a specific number of elements selected from N elements.
[0011] According to yet another embodiment of the present invention, there is provided a video encoding device, comprising: an adder for deriving residual samples of a current block; a transformer for deriving transform coefficients of the current block based on a selective transform of the residual samples; and an entropy encoding unit for encoding information related to the transform coefficients, wherein the selective transform is performed based on a modified transform matrix, the modified transform matrix being a matrix having modified basis vectors, and the modified basis vectors having a specific number of elements selected from N elements. [Effects of the Invention]
[0012] According to the present invention, the amount of data that must be transmitted for residual processing can be reduced through efficient conversion, thereby improving residual coding efficiency.
[0013] According to the present invention, a non-separable transform can be performed based on a transform matrix composed of basis vectors including a selected specific number of elements, thereby reducing the memory load and computational complexity for the non-separable transform and improving residual coding efficiency.
[0014] According to the present invention, a non-separable transform can be performed based on a transform matrix with a simplified structure, thereby reducing the amount of data that must be transmitted for residual processing and improving residual coding efficiency. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram illustrating the configuration of a video encoding device to which the present invention can be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video decoding device to which the present invention can be applied; [Figure 3] FIG. 2 is a diagram illustrating a multiple transform technique according to the present invention; [Figure 4] FIG. 10 is a diagram illustrating an example of intra-directional modes of 65 prediction directions. [Figure 5A] FIG. 2 illustrates a selective transform according to one embodiment of the present invention. [Figure 5B] FIG. 2 illustrates a selective transform according to one embodiment of the present invention. [Figure 5C] FIG. 2 illustrates a selective transform according to one embodiment of the present invention. [Figure 6] 1 is a diagram illustrating a multiple transform technique in which the selective transform is applied to a secondary transform. [Figure 7] FIG. 2 is a diagram illustrating an arrangement of transform coefficients based on a current block according to an embodiment of the present invention. [Figure 8] 10 is a diagram illustrating an example of deriving transform coefficients through a transform in which the simplified transform and the selective transform are combined. [Figure 9] 10 is a diagram illustrating an example of deriving transform coefficients through the selective transform. FIG. [Figure 10] FIG. 10 is a diagram illustrating an example of deriving and selectively transforming related vectors based on two factors for the related vectors. [Figure 11] 1 is a diagram illustrating a video encoding method using an encoding device according to the present invention; [Figure 12] 1 is a diagram illustrating an encoding device for performing a video encoding method according to the present invention; [Figure 13] 1 is a diagram illustrating a video decoding method using a decoding device according to the present invention; [Figure 14] 1 is a diagram illustrating a decoding device for performing a video decoding method according to the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0016] The present invention may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments. The terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical spirit of the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0017] Meanwhile, each component in the drawings described in the present invention is illustrated independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0018] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. In the following, the same reference numerals are used to designate the same components in the drawings, and redundant descriptions of the same components will be omitted.
[0019] Meanwhile, the present invention relates to video / image coding. For example, the methods / embodiments disclosed in the present invention can be applied to methods disclosed in the Versatile Video Coding (VVC) standard or next generation video / image coding.
[0020] In this specification, a picture generally refers to a unit representing one image at a specific time, and a slice is a unit constituting a part of a picture in coding. One picture may be composed of multiple slices, and pictures and slices may be used interchangeably as needed.
[0021] A pixel or a pel may refer to the smallest unit constituting a picture (or an image). A term corresponding to a pixel may also be used: "sample." A sample generally refers to a pixel or a pixel value, and may refer to only the value of a pixel / pixel of a luminance (luma) component, or may refer to only the value of a pixel / pixel of a chroma component.
[0022] A unit refers to a basic unit of video processing. A unit may include at least one of a specific region of a picture and information about the region. The term unit may be used interchangeably with terms such as block or area. In general, an MxN block may refer to a set of samples or transform coefficients consisting of M columns and N rows.
[0023] FIG. 1 is a diagram for explaining the configuration of a video encoding device to which the present invention can be applied.
[0024] 1, a video encoding device 100 may include a picture division unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a transformation unit 122, a quantization unit 123, a realignment unit 124, an inverse quantization unit 125, and an inverse transformation unit 126.
[0025] The picture division unit 105 can divide an input picture into at least one processing unit.
[0026] For example, the processing unit is called a coding unit (CU). In this case, the coding units may be recursively divided from the largest coding unit (LCU) using a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into a plurality of deeper coding units based on a quad-tree structure and / or a binary tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure. Alternatively, the binary tree structure may be applied first. A coding procedure according to the present invention may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include prediction, transformation, restoration, and other procedures, which will be described later.
[0027] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding units may be split into deeper coding units from the largest coding unit (LCU) using a quadtree structure. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively split into lower depth coding units as needed, and a coding unit of an optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that serves as a basis for partitioning or dividing into prediction units or transform units. A prediction unit is a unit that is partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into subblocks. The transform unit may be divided from the coding unit according to a quadtree structure and is a unit that derives transform coefficients and / or a unit that derives a residual signal from the transform coefficients. Hereinafter, the coding unit is also referred to as a coding block (CB), the prediction unit is also referred to as a prediction block (PB), and the transform unit is also referred to as a transform block (TB). A prediction block or a prediction unit refers to a specific region in the form of a block within a picture and may include an array of prediction samples.Also, a transform block or transform unit refers to a specific region in block form within a picture, and can include an array of transform coefficients or residual samples.
[0028] The prediction unit 110 performs prediction on a current block to be processed (hereinafter, referred to as a current block) and generates a predicted block including prediction samples for the current block. The prediction unit 110 performs prediction on a coding block, a transform block, or a prediction block.
[0029] The prediction unit 110 may determine whether intra prediction or inter prediction is applied to the current block. For example, the prediction unit 110 may determine whether intra prediction or inter prediction is applied on a CU basis.
[0030] In intra prediction, the predictor 110 may derive a prediction sample for a current block based on a reference sample outside the current block within a picture to which the current block belongs (hereinafter, the current picture). In this case, the predictor 110 may (i) derive a prediction sample based on an average or interpolation of neighboring reference samples of the current block, or (ii) derive a prediction sample based on a reference sample located in a specific (prediction) direction relative to the prediction sample among the neighboring reference samples of the current block. (i) is referred to as a non-directional mode or a non-angular mode, and (ii) is referred to as a directional mode or an angular mode. Prediction modes in intra prediction may include, for example, 33 directional prediction modes and at least two or more non-directional modes. Non-directional modes may include a DC prediction mode and a planar mode. The predictor 110 may also determine a prediction mode to be applied to the current block using a prediction mode applied to a neighboring block.
[0031] In the case of inter prediction, the predictor 110 may derive a predicted sample for the current block based on a sample identified by a motion vector on a reference picture. The predictor 110 may derive a predicted sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the skip mode and the merge mode, the predictor 110 may use motion information of a neighboring block as motion information of the current block. In the skip mode, unlike the merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted. In the MVP mode, the motion vector of the neighboring block may be used as a motion vector predictor to derive the motion vector of the current block.
[0032] In the case of inter prediction, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. A reference picture including the temporal neighboring blocks is also called a collocated picture (colPic). Motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and output in the form of a bitstream.
[0033] When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on a reference picture list can be used as a reference picture. Reference pictures included in a reference picture list can be sorted based on the difference in POC (Picture Order Count) between the current picture and the corresponding reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0034] The subtractor 121 generates residual samples, which are the differences between the original samples and the predicted samples, or, if skip mode is applied, does not generate residual samples, as described above.
[0035] The transform unit 122 transforms residual samples in units of transform blocks to generate transform coefficients. The transform unit 122 may perform the transform according to the size of the corresponding transform block and the prediction mode applied to a coding block or a prediction block spatially overlapping with the corresponding transform block. For example, if intra prediction is applied to the coding block or the prediction block overlapping with the transform block and the transform block is a 4x4 residual array, the residual samples may be transformed using a Discrete Sine Transform (DST) transform kernel; otherwise, the residual samples may be transformed using a Discrete Cosine Transform (DCT) transform kernel.
[0036] The quantization unit 123 can quantize the transform coefficients to generate quantized transform coefficients.
[0037] The rearrangement unit 124 rearranges the quantized transform coefficients. The rearrangement unit 124 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form through a coefficient scanning method. Here, the rearrangement unit 124 has been described as a separate component, but it may also be a part of the quantization unit 123.
[0038] The entropy encoding unit 130 may perform entropy encoding on the quantized transform coefficients. Entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also encode information required for video reconstruction (e.g., syntax element values) in addition to the quantized transform coefficients, either together with or separately from the quantized transform coefficients, using entropy encoding or a preset method. The encoded information may be transmitted or stored in Network Abstraction Layer (NAL) unit units in the form of a bitstream.
[0039] The inverse quantization unit 125 inversely quantizes the values (quantized transform coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inversely transforms the values inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0040] The adder 140 reconstructs a picture by adding residual samples and prediction samples. The residual samples and prediction samples may be added in block units to generate reconstructed blocks. Although the adder 140 has been described as a separate component, it may be part of the prediction unit 110. Meanwhile, the adder 140 may also be referred to as a reconstruction module or a reconstructed block generator.
[0041] The filter unit 150 may apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through the deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process may be corrected. The sample adaptive offset may be applied on a sample-by-sample basis and may be applied after the deblocking filtering process is completed. The filter unit 150 may also apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filtering and / or the sample adaptive offset have been applied.
[0042] The memory 160 may store a reconstructed picture (decoded picture) or information necessary for encoding / decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 150. The stored reconstructed picture may be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 may store (reference) pictures used for inter prediction. In this case, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list.
[0043] FIG. 2 is a diagram illustrating the configuration of a video decoding device to which the present invention can be applied.
[0044] 2, the video decoding device 200 may include an entropy decoding unit 210, a residual processing unit 220, a prediction unit 230, an addition unit 240, a filter unit 250, and a memory 260. Here, the residual processing unit 220 may include a realignment unit 221, an inverse quantization unit 222, and an inverse transform unit 223.
[0045] When a bitstream containing video information is input, the video decoding device 200 can restore the video in accordance with the process by which the video information was processed in the video encoding device.
[0046] For example, the video decoding device 200 may perform video decoding using a processing unit applied in a video encoding device. Accordingly, a processing unit block for video decoding may be a coding unit, for example, or a coding unit, a prediction unit, or a transform unit, for example. The coding unit may be divided from the largest coding unit using a quadtree structure and / or a binary tree structure.
[0047] A prediction unit and a transform unit may also be used in some cases, where a prediction block is a block derived or partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into sub-blocks. A transform unit may be divided from a coding unit according to a quadtree structure and is a unit that derives transform coefficients or a unit that derives a residual signal from the transform coefficients.
[0048] The entropy decoding unit 210 may parse the bitstream and output information necessary for video reconstruction or picture reconstruction. For example, the entropy decoding unit 210 may decode information in the bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for video reconstruction and quantized values of transform coefficients for residuals.
[0049] More specifically, the CABAC entropy decoding method receives BINs corresponding to each syntax element in a bitstream, determines a context model using information on the syntax element to be decoded and decoding information on adjacent and current blocks or information on symbols / BINs decoded in previous steps, predicts the occurrence probability of the BINs according to the determined context model, and performs arithmetic decoding of the BINs to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method can update the context model using information on the decoded symbols / BINs for the context model of the next symbol / BIN.
[0050] Prediction information among the information decoded by the entropy decoding unit 210 is provided to the prediction unit 230, and residual values, i.e., quantized transform coefficients, on which entropy decoding is performed by the entropy decoding unit 210 can be input to the reordering unit 221.
[0051] The rearrangement unit 221 may rearrange the quantized transform coefficients in a two-dimensional block format. The rearrangement unit 221 may perform rearrangement in response to a coefficient scan performed by the encoding device. Here, although the rearrangement unit 221 has been described as a separate component, it may also be a part of the inverse quantization unit 222.
[0052] The inverse quantization unit 222 may inversely quantize the quantized transform coefficients based on the (inverse) quantization parameter and output the transform coefficients. In this case, information for deriving the quantization parameter may be signaled from the encoding device.
[0053] The inverse transform unit 223 can inverse transform the transform coefficients to derive residual samples.
[0054] The prediction unit 230 may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The unit of prediction performed by the prediction unit 230 is a coding block, a transform block, or a prediction block.
[0055] The prediction unit 230 may determine whether to apply intra prediction or inter prediction based on the information related to prediction. In this case, the unit for determining whether to apply intra prediction or inter prediction differs from the unit for generating predicted samples. Furthermore, the unit for generating predicted samples differs between inter prediction and intra prediction. For example, whether to apply inter prediction or intra prediction may be determined on a CU basis. Furthermore, for example, in inter prediction, a prediction mode may be determined on a PU basis to generate predicted samples, and in intra prediction, a prediction mode may be determined on a PU basis to generate predicted samples on a TU basis.
[0056] In the case of intra prediction, the prediction unit 230 may derive prediction samples for the current block based on neighboring reference samples in the current picture. The prediction unit 230 may derive prediction samples for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined using the intra prediction mode of the neighboring block.
[0057] In the case of inter prediction, the prediction unit 230 may derive a prediction sample for the current block based on a sample identified on the reference picture by a motion vector on the reference picture. The prediction unit 230 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP mode. In this case, motion information required for inter prediction of the current block provided by the video encoding device, such as information on a motion vector, a reference picture index, etc., may be acquired or derived based on the information on the prediction.
[0058] In the skip mode and merge mode, the motion information of neighboring blocks can be used as the motion information of the current block, and the neighboring blocks can include spatial and temporal neighboring blocks.
[0059] The predictor 230 constructs a merge candidate list using the motion information of available neighboring blocks, and can use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. The motion information can include a motion vector and a reference picture. When the motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture.
[0060] In skip mode, unlike merge mode, the difference between the predicted samples and the original samples (residuals) is not transmitted.
[0061] In the MVP mode, the motion vector of the current block can be derived using the motion vector of a neighboring block as a motion vector predictor, where the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0062] For example, when a merge mode is applied, a merge candidate list may be generated using the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vector corresponding to the Col block, which is a temporally neighboring block. In the merge mode, the motion vector of a candidate block selected from the merge candidate list is used as the motion vector of the current block. The prediction information may include a merge index indicating a candidate block having an optimal motion vector selected from the candidate blocks included in the merge candidate list. In this case, the prediction unit 230 may derive the motion vector of the current block using the merge index.
[0063] As another example, when the Motion Vector Prediction (MVP) mode is applied, a motion vector predictor candidate list may be generated using the motion vector of a reconstructed spatial neighboring block and / or the motion vector corresponding to a Col block, which is a temporal neighboring block. That is, the motion vector of a reconstructed spatial neighboring block and / or the motion vector corresponding to a Col block, which is a temporal neighboring block, may be used as a motion vector candidate. The prediction information may include a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list. In this case, the prediction unit 230 may select a predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. A prediction unit of the encoding device may obtain a motion vector difference (MVD) between the motion vector of the current block and a motion vector predictor, encode the MVD, and output it in the form of a bitstream. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction unit 230 may obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit can also obtain or derive a reference picture index indicating a reference picture from the information related to the prediction.
[0064] The adder 240 may reconstruct a current block or a current picture by adding residual samples and prediction samples. The adder 240 may also reconstruct a current picture by adding residual samples and prediction samples in block units. When a skip mode is applied, the residual is not transmitted, and therefore the prediction samples may become reconstructed samples. Here, although the adder 240 has been described as a separate component, it may also be part of the prediction unit 230. Meanwhile, the adder 240 may also be referred to as a reconstruction module or a reconstructed block generation unit.
[0065] The filter unit 250 may apply deblock filtering, sample adaptive offset, and / or ALF to the reconstructed picture. In this case, the sample adaptive offset may be applied in sample units or may be applied after deblock filtering. The ALF may be applied after deblock filtering and / or sample adaptive offset.
[0066] The memory 260 may store a reconstructed picture (decoded picture) or information required for decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 250. For example, the memory 260 may store a picture used for inter prediction. In this case, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for another picture. The memory 260 may also output the reconstructed picture in an output order.
[0067] Meanwhile, through the above-described transformation, low frequency transform coefficients for the residual block of the current block can be derived, and a zero tail can be derived at the end of the residual block.
[0068] Specifically, the transformation can be composed of two main processes, which can include a core transform and a secondary transform. The transformation including the core transform and the secondary transform can be referred to as a multiple transformation technique.
[0069] FIG. 3 illustrates a schematic diagram of a multiple conversion technique according to the present invention.
[0070] Referring to Figure 3, the transform unit may correspond to the transform unit in the encoding device of Figure 1 described above, and the inverse transform unit may correspond to the inverse transform unit in the encoding device of Figure 1 described above or the inverse transform unit in the decoding device of Figure 2.
[0071] The transform unit may perform a first-order transform based on the residual samples (residual sample array) in the residual block to derive (first-order) transform coefficients (S310). Here, the first-order transform may include an adaptive multiple core transform (AMT). The adaptive multiple core transform may be referred to as a multiple transform set (MTS).
[0072] The adaptive multi-core transform may refer to a transform method additionally using a Discrete Cosine Transform (DCT) type 2 and a Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1. That is, the adaptive multi-core transform may refer to a transform method of transforming a spatial domain residual signal (or a residual block) into a frequency domain transform coefficient (or a primary transform coefficient) based on a plurality of transform kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary transform coefficient may also be referred to as a temporary transform coefficient from the perspective of a transform unit.
[0073] In other words, when an existing transform method is applied, a spatial-domain to frequency-domain transform may be applied to the residual signal (or residual block) based on DCT type 2 to generate transform coefficients. In contrast, when the adaptive multi-core transform is applied, a spatial-domain to frequency-domain transform may be applied to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc. to generate transform coefficients (or primary transform coefficients). Here, DCT type 2, DST type 7, DCT type 8, DST type 1, etc. may be referred to as transform types, transform kernels, or transform cores.
[0074] For reference, the DCT / DST transformation types can be defined based on basis functions, which can be shown as follows:
[0075] [Table 1]
[0076] When the adaptive multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for a current block may be selected from the transform kernels, and a vertical transform for the current block may be performed based on the vertical transform kernel, and a horizontal transform for the current block may be performed based on the horizontal transform kernel. Here, the horizontal transform may indicate a transform for a horizontal component of the current block, and the vertical transform may indicate a transform for a vertical component of the current block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on a transform index indicating a prediction mode and / or a transform subset of a current block (CU or sub-block) that surrounds a residual block.
[0077] For example, the adaptive multi-core transform may be applied when the width and height of the target block are both less than or equal to 64, and whether the adaptive multi-core transform of the target block is applied may be determined based on a CU level flag. Specifically, when the CU level flag is 0, the existing transform method described above may be applied. That is, when the CU level flag is 0, a spatial domain to frequency domain transform may be applied to a residual signal (or a residual block) based on the DCT type 2 to generate transform coefficients, and the transform coefficients may be encoded. Meanwhile, the target block may be a CU. When the CU level flag is 0, the adaptive multi-core transform may be applied to the target block.
[0078] Furthermore, for the luma block of the current block to which the adaptive multi-core transform is applied, two additional flags may be signaled, and a vertical transform kernel and a horizontal transform kernel may be selected based on the flags. The flag for the vertical transform kernel may be indicated as an AMT vertical flag, and AMT_TU_vertical_flag (or EMT_TU_vertical_flag) may indicate a syntax element of the AMT vertical flag. The flag for the horizontal transform kernel may be indicated as an AMT horizontal flag, and AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) may indicate a syntax element of the AMT horizontal flag. The AMT vertical flag may indicate one transform kernel candidate from among the transform kernel candidates included in a transform subset for the vertical transform kernel, and the transform kernel candidate indicated by the AMT vertical flag may be derived as the vertical transform kernel for the current block. In addition, the AMT horizontal flag may indicate one of the transform kernel candidates included in the transform subset for the horizontal transform kernel, and the transform kernel candidate indicated by the AMT horizontal flag may be derived as the horizontal transform kernel for the current block. Meanwhile, the AMT vertical flag may also be indicated as an MTS vertical flag, and the AMT horizontal flag may also be indicated as an MTS horizontal flag.
[0079] Meanwhile, three transform subsets may be preset, and one of the transform subsets may be derived as a transform subset for the vertical transform kernel based on an intra prediction mode applied to the current block. Also, one of the transform subsets may be derived as a transform subset for the horizontal transform kernel based on an intra prediction mode applied to the current block. For example, the preset transform subsets may be derived as shown in the following table.
[0080] [Table 2]
[0081] Referring to Table 2, a transform subset with an index value of 0 may indicate a transform subset that includes DST type 7 and DCT type 8 as transform kernel candidates, a transform subset with an index value of 1 may indicate a transform subset that includes DST type 7 and DST type 1 as transform kernel candidate candidates, and a transform subset with an index value of 2 may indicate a transform subset that includes DST type 7 and DCT type 8 as transform kernel candidate candidates.
[0082] The transform subset for the vertical transform kernel and the transform subset for the horizontal transform kernel, which are derived based on the intra prediction mode applied to the current block, can be derived as shown in the following table.
[0083] [Table 3]
[0084] where V denotes the transform subset for the vertical transform kernel and H denotes the transform subset for the horizontal transform kernel.
[0085] When the value of the AMT flag (or EMT_CU_flag) for the target block is 1, a transform subset for the vertical transform kernel and a transform subset for the horizontal transform kernel may be derived based on the intra prediction mode of the target block, as shown in Table 3. Then, among the transform kernel candidates included in the transform subset for the vertical transform kernel, a transform kernel candidate indicated by the AMT vertical flag of the target block may be derived as the vertical transform kernel of the target block, and among the transform kernel candidates included in the transform subset for the horizontal transform kernel, a transform kernel candidate indicated by the AMT horizontal flag of the target block may be derived as the horizontal transform kernel of the target block. Meanwhile, the AMT flag may also be represented as an MTS flag.
[0086] For reference, for example, the intra prediction modes may include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes may include a planar intra prediction mode numbered 0 and a DC intra prediction mode numbered 1, and the directional intra prediction modes may include 65 intra prediction modes numbered 2 to 66. However, this is merely an example, and the present invention may be applied to cases where the number of intra prediction modes is different. Meanwhile, a 67th intra prediction mode may also be used in some cases, and the 67th intra prediction mode may indicate a Linear Model (LM) mode.
[0087] FIG. 4 exemplarily shows the intra-directional modes of 65 prediction directions.
[0088] Referring to FIG. 4, intra prediction modes can be divided into those with horizontal directionality and those with vertical directionality, with the 34th intra prediction mode having a left-up diagonal prediction direction as the center. H and V in FIG. 4 represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate displacements in 1 / 32 units on the sample grid position. The 2nd to 33rd intra prediction modes have horizontal directionality, while the 34th to 66th intra prediction modes have vertical directionality. The 18th and 50th intra prediction modes represent horizontal and vertical intra prediction modes, respectively. The 2nd intra prediction mode can be referred to as a left-down diagonal intra prediction mode, the 34th intra prediction mode as a left-up diagonal intra prediction mode, and the 66th intra prediction mode as a right-up diagonal intra prediction mode.
[0089] The transform unit may derive (secondary) transform coefficients by performing a secondary transform based on the (first) transform coefficients (S320). If the first transform is a transform from the spatial domain to the frequency domain, the secondary transform can be considered a transform from the frequency domain to the frequency domain. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may refer to a transform that generates transform coefficients (or secondary transform coefficients) for a residual signal by performing a secondary transform on the (first) transform coefficients derived through the first transform based on a non-separable transform matrix. Here, the transform may be applied simultaneously to the (first) transform coefficients based on the non-separable transform matrix, rather than separately applying a vertical transform and a horizontal transform (or independently applying a horizontal-vertical transform). In other words, the non-separable quadratic transform may refer to a transform method in which vertical and horizontal components of the (first-order) transform coefficients are not separated based on the non-separable transform matrix, but are transformed together to generate transform coefficients (or second-order transform coefficients). The non-separable quadratic transform may be applied to a top-left region of a block (hereinafter referred to as a transform coefficient block or a target block) composed of the (first-order) transform coefficients. For example, if the width (W) and height (H) of the transform coefficient block are both 8 or greater, an 8x8 non-separable quadratic transform may be applied to the top-left 8x8 region of the transform coefficient block (hereinafter referred to as the top-left target region). Also, if the width (W) and height (H) of the transform coefficient block are both 4 or greater and the width (W) or height (H) of the transform coefficient block is less than 8, a 4x4 non-separable quadratic transform may be applied to the top-left min(8,W)xmin(8,H) region of the transform coefficient block.
[0090] Specifically, for example, if a 4x4 input block is used, a non-separable quadratic transform can be performed as follows:
[0091] The above 4×4 input block X can be expressed as follows:
[0092] <Formula 1>
number
[0093] If we express the above X in vector form, the vector TIFF2025134049000006.tif84 can be displayed as follows:
[0094] <Formula 2>
number
[0095] In this case, the above second-order non-separable transform can be calculated as follows:
[0096] <Formula 3>
number
[0097] where: TIFF2025134049000009.tif75 denotes the transform coefficient vector, and T denotes the 16x16 (non-separable) transform matrix.
[0098] 16×1 transform coefficient vector through Equation 3 TIFF2025134049000010.tif75 can be derived from the above TIFF2025134049000011.tif75 can be re-organized into 4x4 blocks through the scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is an example, and in order to reduce the computational complexity of non-separable quadratic transforms, HyGT (Hypercube-Givens Transform) etc. can also be used to calculate non-separable quadratic transforms.
[0099] Meanwhile, in the non-separable quadratic transform, the transform kernel (or transform core, transform type) may be selected in a mode-dependent manner, where the mode may include an intra-prediction mode and / or an inter-prediction mode.
[0100] As described above, the non-separable quadratic transform may be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. That is, the non-separable quadratic transform may be performed based on an 8×8 sub-block size or a 4×4 sub-block size. For example, for the mode-dependent transform kernel selection, 35 sets of three non-separable quadratic transform kernels for the non-separable quadratic transform may be configured for both the 8×8 sub-block size and the 4×4 sub-block size. That is, 35 transform sets may be configured for the 8×8 sub-block size, and 35 transform sets may be configured for the 4×4 sub-block size. In this case, each of the 35 transform sets for the 8×8 sub-block size may include three 8×8 transform kernels, and each of the 35 transform sets for the 4×4 sub-block size may include three 4×4 transform kernels. However, the transform sub-block size, the number of sets, and the number of transform kernels in a set may be, for example, sizes other than 8x8 or 4x4, or n sets may be constructed, with each set containing k transform kernels.
[0101] The transform set may be referred to as an NSST set, and a transform kernel in the NSST set may be referred to as an NSST kernel. Selection of a particular set from the transform set may be performed based on, for example, an intra prediction mode of a current block (CU or sub-block).
[0102] In this case, the mapping between the 35 transform sets and the intra prediction modes may be shown, for example, as follows: For reference, when the LM mode is applied to the current block, the secondary transform may not be applied to the current block.
[0103] [Table 4]
[0104] On the other hand, if it is determined that a specific set is to be used, one of the k transform kernels in the specific set can be selected through a non-separable secondary transform index. The encoding device can derive a non-separable secondary transform index that points to a specific transform kernel based on a rate-distortion (RD) check and signal the non-separable secondary transform index to a decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, an NSST index value of 0 can point to the first non-separable secondary transform kernel, an NSST index value of 1 can point to the second non-separable secondary transform kernel, and an NSST index value of 2 can point to the third non-separable secondary transform kernel. Alternatively, an NSST index value of 0 can indicate that the first non-separable secondary transform is not applied to the current block, and NSST index values of 1 to 3 can point to the three transform kernels.
[0105] 3, the transform unit may perform the non-separable quadratic transform based on the selected transform kernel to obtain (quadratic) transform coefficients. As described above, the transform coefficients may be derived as quantized transform coefficients through a quantizer, encoded, and signaled to a decoding device and transmitted to an inverse quantization / inverse transform unit in an encoding device.
[0106] On the other hand, if the secondary transform is omitted, the (primary) transform coefficients, which are the output of the primary (separate) transform, can be derived as quantized transform coefficients through the quantization unit as described above, encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse transform unit in the encoding device.
[0107] The inverse transform unit may perform a series of procedures in the reverse order of the procedures performed by the transform unit described above. The inverse transform unit may receive (dequantized) transform coefficients, perform a secondary (inverse) transform on them to derive (primary) transform coefficients (S350), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain residual blocks (residual samples) (S360). Here, the primary transform coefficients may be referred to as modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding and decoding devices may generate reconstructed blocks based on the residual blocks and predicted blocks, and generate reconstructed pictures based on the reconstructed blocks.
[0108] On the other hand, as described above, if the secondary (inverse) transform is omitted, the residual block (residual sample) can be obtained by receiving (dequantized) transform coefficients and performing the primary (separate) transform. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the predicted block, and generate a reconstructed picture based on the reconstructed block.
[0109] 5A-5C are diagrams illustrating selective transforms according to one embodiment of the present invention.
[0110] In this specification, the term "current block" may refer to a current block or a residual block to be coded.
[0111] FIG. 5A shows an example of deriving transform coefficients through a transform.
[0112] Transform in video coding may refer to a process of transforming an input vector R based on a transform matrix to generate a transform coefficient vector C for the input vector R, as shown in FIG. 5A. The input vector R may refer to a primary transform coefficient. Alternatively, the input vector R may refer to a residual vector, i.e., a residual sample. Meanwhile, the transform coefficient vector C may be referred to as an output vector C.
[0113] FIG. 5B shows a specific example of deriving transform coefficients through a transform. FIG. 5B shows the transform process shown in FIG. 5A in detail. As described above with reference to FIG. 3, in the non-separable quadratic transform (hereinafter referred to as 'NSST'), block data of transform coefficients obtained by applying a linear transform is divided into MxM blocks, and then MxM is calculated for each MxM block. 2 xM 2 NSST can be performed. M can be, for example, 4 or 8, but is not limited thereto. 2 can be N. In this case, as shown in FIG. 5B, the input vector R is divided into (first-order) transform coefficients r1 to r N The transform coefficient vector C can be a (1xN) dimensional vector including transform coefficients c1 to c N That is, the input vector R can be an (Nx1) dimensional vector including N (first-order) transform coefficients r1 to r N The input vector R may have a size of 1×N. The transform coefficient vector C may include N transform coefficients c1 to c N and the size of the transform coefficient vector C can be Nx1.
[0114] The input vector R can go through the transformation matrix to derive the transformation coefficient vector C. That is, the input vector R can be derived as the transformation coefficient vector C based on the transformation matrix.
[0115] On the other hand, the transformation matrix is made up of N basis vectors B1 to B N As shown in FIG. 5B, the basis vectors B1 to B N can be a (1xN) dimensional vector. That is, the above basis vectors B1 to B N The size of may be 1xN. A transform coefficient vector C may be generated based on the (first-order) transform coefficients of the input vector R and each of the basis vectors of the transform matrix. For example, the inner product of the input vector and each basis vector may be derived as the transform coefficient vector C.
[0116] However, the above transformations have two main issues: high computational complexity associated with the number of multiplications and additions required to generate the output vector, and memory requirements for storing the generated coefficients.
[0117] For example, the computational complexity and memory requirements for a separable transform and a non-separable transform can be derived as follows:
[0118] [Table 5]
[0119] Referring to Table 5, the memory for storing the coefficients generated through the separation transformation is N2 The number of calculations is 2N 3 The number of calculations indicates the computational complexity. In addition, N4 memories may be required to store the coefficients generated through the non-separable transform, and the number of calculations may be N4. The number of calculations indicates the computational complexity. That is, the more the number of calculations, the higher the computational complexity may be, and the fewer the number of calculations, the lower the computational complexity may be.
[0120] As shown in Table 5, the memory requirements and calculation times for the non-separable transform can be significantly greater than those for the separable transform. Also, as the size of the block to which the non-separable transform is performed increases, i.e., as N increases, the discrepancy between the memory requirements and calculation times for the separable transform and the non-separable transform can increase.
[0121] The non-separable transform provides better coding gain than the separable transform. However, as shown in Table 5 above, the non-separable transform is not used in existing video coding standards due to the computational complexity of the non-separable transform. Furthermore, since the computational complexity of the separable transform increases as the size of the target block increases, the existing HEVC standard restricts the separable transform to be used only for target blocks whose size is 32x32 or less.
[0122] The present invention proposes a selective transform. The selective transform can significantly reduce computational complexity and memory requirements, thereby increasing the efficiency of computationally intensive transform blocks and improving coding efficiency. That is, the selective transform can be used to resolve the computational complexity issues that arise when transforming large blocks or performing non-separable transforms. The selective transform can be used for any type of transform, such as a primary transform (also referred to as a core transform) or a secondary transform. For example, the selective transform can be applied to the core transform for an encoding device / decoding device, significantly reducing encoding / decoding times.
[0123] 5C illustrates an example of deriving transform coefficients through the selective transform. The selective transform may refer to a transform performed on a current block based on a transform matrix including basis vectors having a selective number of elements.
[0124] The simplified transformation is a method proposed based on the motivation that, among the N elements of the basis vectors of a transformation matrix, redundant or insignificant elements may be included, and excluding these elements can reduce computational complexity and memory requirements. For example, referring to FIG. 5C, among the N elements of a basis vector B1, Z0 elements may be insignificant elements. In this case, a truncated basis vector B1 including only N1 elements may be derived. Here, N1 may be N-Z0. The truncated basis vector B1 may be referred to as a modified basis vector B1.
[0125] Referring to Figure 5C, when the modified basis vector B1 is applied to the input vector R as part of a transformation matrix, a transformation coefficient C1 can be derived. Experimental results show that the transformation coefficient C1 is the same as the transformation coefficient C1 derived by applying the existing basis vector B1 to the input vector R as part of a transformation matrix. That is, deriving the result assuming that the smallest elements of each basis vector are 0 means that the number of required multiplications can be significantly reduced without a significant difference in the result. This also means that the number of elements (i.e., elements of the transformation matrix) that need to be stored for this operation can be reduced.
[0126] An association vector is proposed to define the position of insignificant (or meaningful) elements among the elements of the basis vectors. The (1xN)-dimensional association vector A1 can be considered to derive the N1-dimensional modified basis vector B1. That is, the 1xN-dimensional association vector A1 can be considered to derive the 1xN1-sized modified basis vector B1 (i.e., the modified basis vector B1 including N1 elements).
[0127] Referring to FIG. 5C, the input vector R is assigned to the basis vectors B1 to B2. N The associated vectors for each of the input vectors R may be applied, and the derived values may be transferred to the basis vectors. Through this, only some elements of the input vector R may be calculated as elements of the basis vectors. Specifically, the associated vectors may include 0 and 1, and may be calculated such that selected elements of the input vector R are multiplied by 1 and unselected elements are multiplied by 0, and only the selected elements may be passed through and transferred to the basis vectors.
[0128] For example, the related vector A1 may be applied to the input vector R, and only N1 elements of the input vector R designated by the related vector A1 may be used to calculate the dot product with the basis vector B1. The dot product may indicate C1 of the transform coefficient vector C. Here, the basis vector B1 may include N1 elements, and N1 may be less than or equal to N. The above-described operation between the input vector R, the related vector A1, and the basis vector B1 may be performed using related vectors A2 to A3. N and basis vectors B2 to B N It can also be performed on
[0129] Since the related vector includes only binary values, 0 and / or 1, it may be advantageous to store the related vector. A 0 in the related vector may indicate that the element of the input vector R corresponding to the 0 is not transferred to the transformation matrix for the dot product calculation, and a 1 in the related vector may indicate that the element of the input vector R corresponding to the 1 is transferred to the transformation matrix for the dot product calculation. For example, a 1xN size related vector A may be stored as follows: k is A k1 or A kN The above related vector A k A kn If is 0, then r of the above input vector R n In other words, the r of the above input vector R may not be passed. n is the above transformation vector B n In addition, the above A kn If is 1, then r of the above input vector R n can be passed. That is, r of the above input vector R n is the above transformation vector B n and the c of the above transformation coefficient vector C can be n can be used in calculations to derive
[0130] FIG. 6 shows a schematic diagram of a multiple transform technique in which the selective transform is applied to a secondary transform.
[0131] Referring to Figure 6, the transform unit may correspond to the transform unit in the encoding device of Figure 1 described above, and the inverse transform unit may correspond to the inverse transform unit in the encoding device of Figure 1 described above or the inverse transform unit in the decoding device of Figure 2.
[0132] The transform unit may perform a linear transform based on the residual samples (residual sample array) in the residual block to derive (first-order) transform coefficients (S610). Here, the linear transform may include the AMT described above.
[0133] When the adaptive multi-core transform is applied, a spatial-domain to frequency-domain transform may be applied to a residual signal (or a residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, to generate transform coefficients (or primary transform coefficients). Here, the primary transform coefficients may be referred to as temporary transform coefficients from the perspective of a transform unit. Also, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. For reference, the DCT / DST transform types may be defined based on basis functions, which may be shown in Table 1 above. Specifically, the process of applying the adaptive multi-core transform to derive the primary transform coefficients is as described above.
[0134] The transform unit may perform a selective transform based on the (primary) transform coefficients to derive (secondary) transform coefficients (S620). The selective transform may refer to a transform performed on the (primary) transform coefficients for the current block based on a transform matrix including modified basis vectors and an association matrix including association vectors for the basis vectors. The modified basis vector may indicate a basis vector including N or less elements. That is, the modified basis vector may indicate a basis vector including a selected specific number of elements among the N elements. For example, the modified basis vector B n is (1xN n )-dimensional vector, and the above N n can be smaller or equal to the above N. That is, the modified basis vector B n The size of (1xN n ) size, above N n may be less than or equal to N. Here, N may be the product of the height and width of the upper left target region of the target block to which the selected transform is applied. Alternatively, N may be the total number of transform coefficients of the upper left target region of the target block to which the selected transform is applied. Meanwhile, the transform matrix including the modified basis vectors may be referred to as a modified transform matrix. The transform matrix may also be referred to as a transform bases block (TBB), and the association matrix may be referred to as an association vectors block (AVB).
[0135] The transform unit may perform the selective transform based on the modified transform matrix and the associated matrix to obtain (secondary) transform coefficients. The transform coefficients may be derived as quantized transform coefficients through a quantizer, as described above, and may be encoded and signaled to a decoding device and transmitted to an inverse quantizer / inverse transform unit in an encoding device.
[0136] The inverse transform unit may perform a series of procedures in the reverse order of the procedures performed by the transform unit described above. The inverse transform unit may receive (dequantized) transform coefficients and perform a selective (inverse) transform to derive (first-order) transform coefficients (S650), and may perform a first-order (inverse) transform on the (first-order) transform coefficients to obtain residual blocks (residual samples) (S660). Here, the first-order transform coefficients may be referred to as modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the predicted block, and generate a reconstructed picture based on the reconstructed block.
[0137] Meanwhile, in one embodiment of the selective transformation, the present invention proposes a selective transformation combined with a simplifying transformation.
[0138] As used herein, a "simplification transform" may refer to a transform performed on residual samples for a current block based on a transform matrix whose size is reduced by a simplification factor.
[0139] In a simplified transform according to one embodiment, an N-dimensional vector may be mapped to an R-dimensional vector located in another space to determine a simplified transform matrix, where R is less than N. That is, the simplified transform may refer to a transform performed on residual samples for a target block based on a reduced transform matrix including R basis vectors. Here, N may refer to the square of the length of one side of the block (or target region) to which the transform is applied or the total number of transform coefficients corresponding to the block (or target region) to which the transform is applied, and the simplification factor may refer to an R / N value. The simplification factor may be referred to by various terms such as a reduced factor, a reduction factor, a simplified factor, or a simple factor. Meanwhile, R may be referred to as a reduced coefficient, and in some cases, the simplification factor may refer to R. In other cases, the simplification factor may refer to an N / R value.
[0140] The size of the simplified transformation matrix according to one embodiment is RxN, which is smaller than the size of a normal transformation matrix NxN, and can be defined as Equation 4 below.
[0141] <Formula 4>
number
[0142] The simplified transformation matrix T is used for the transformation coefficients to which the linear transformation of the target block has been applied. RxN When multiplied by , the (second order) transform coefficients for the current block can be derived.
[0143] When the RST is applied, a simplified transform matrix of size RxN is applied to the secondary transform, so that the transform coefficients R+1 to N can be implicitly 0. In other words, when the transform coefficients of the target block are derived by applying the RST, the values of the transform coefficients R+1 to N can be 0. Here, the transform coefficients R+1 to N can refer to the R+1th to Nth transform coefficients among the transform coefficients. Specifically, the arrangement of the transform coefficients of the target block can be described as follows:
[0144] FIG. 7 is a diagram illustrating an arrangement of transform coefficients based on a target block according to an embodiment of the present invention. The description of the transform described below with reference to FIG. 7 also applies to the inverse transform. NSST based on a linear transform and a simplified transform may be performed on a target block (or residual block) 700. In one example, the 16x16 block illustrated in FIG. 7 represents the target block 700, and the 4x4 blocks labeled A through P may represent subgroups of the target block 700. The linear transform may be performed on the entire target block 700. After the linear transform is performed, NSST may be applied to an 8x8 block (hereinafter, referred to as the upper left target region) consisting of subgroups A, B, E, and F. In this case, if NSST based on the simplified transform is performed, only R NSST coefficients (where R represents a simplified coefficient and is less than N) are derived, and therefore, the R+1th to Nth NSST coefficients may each be set to 0. For example, when R is 16, 16 transform coefficients derived by performing NSST based on the simplified transform can be assigned to each block included in subgroup A, which is the upper left 4x4 block included in the upper left target area of target block 700, and a transform coefficient of 0 can be assigned to each of the NR blocks, i.e., 64-16=48 blocks, included in subgroups B, E, and F. Primary transform coefficients not subjected to NSST based on the simplified transform can be assigned to each block included in subgroups C, D, G, H, I, J, K, L, M, N, O, and P.
[0145] 8 shows an example of deriving transform coefficients through a transform in which the simplified transform and the selective transform are combined. Referring to FIG. 8, the transform matrix may include R basis vectors, and the association matrix may include R association vectors. Here, the transform matrix including the R basis vectors may be referred to as a reduced transform matrix, and the association matrix including the R association vectors may be referred to as a reduced association matrix.
[0146] Also, each of the basis vectors may include only selected elements from the N elements. For example, referring to FIG. 8, basis vector B1 may be a 1xN1-dimensional vector including N1 elements, basis vector B2 may be a 1xN2-dimensional vector including N2 elements, and basis vector B R is N R 1xN containing elements R can be dimensional vectors. R can be a value equal to or less than N. The combined transformation of the simplifying transformation and the selective transformation can be used for any type of transformation, such as a quadratic transformation or a linear transformation.
[0147] 8, the encoding device and the decoding device may apply a transform obtained by combining the simplified transform and the selective transform to a secondary transform. For example, the selective transform may be performed based on a simplified transform matrix and a simplified relation matrix including the modified basis vectors, and (secondary) transform coefficients may be obtained. In addition, in order to reduce the computational complexity of the selective transform, a Hypercube-Givens Transform (HyGT) or the like may be used to calculate the selective transform.
[0148] 9 shows an example of deriving transform coefficients through the selective transform. In one embodiment of the selective transform, the associated vectors A1, A2, ... A NThere may be no pattern for the association vectors A1, A2, ... A1 of the association matrix, and the association vectors A1, A2, ... A1 may be derived in different forms. N can be derived in the same form.
[0149] Specifically, for example, the related vectors may contain the same number of 1's. For example, if the number of 1's is M, the related vector may contain M 1's and NM 0's. In this case, M transform coefficients among the (primary) transform coefficients of the input vector R can be transferred to the basis vectors. Therefore, the length of the basis vectors may also be M. That is, as shown in FIG. 9, the basis vectors may contain M elements and may be (1xM)-dimensional vectors, where N1=N2=...=N N =M. The associated matrix and modified transformation matrix structure illustrated in Figure 9 can be shown as a symmetric architecture for the above selective transformation.
[0150] In another example, there may be a particular pattern of elements with a value of 1, and the pattern may be repeated, rotated, and / or translated in any manner to derive the associated vector.
[0151] Meanwhile, the selective conversion described above can be applied not only to the simplified conversion and / or the HyGT, but also to other conversion techniques.
[0152] The present invention also proposes a method for simplifying the association vector in the selective transformation. By simplifying the association vector, the storage of information for the selective transformation and the handling of the selective transformation can be improved. That is, the memory load for the selective transformation can be reduced, and the handling capability of the selective transformation can be improved.
[0153] The effect of simplifying the related vector can be more clearly seen when the non-zero elements of the related vector have a continuous distribution. For example, the related vector A k can contain a continuous string of 1's. In this case, k is the sum of two factors A ks and A kL That is, the above factor A ks and A kL Based on the above related vector A k can be derived. Here, the above A ks can be a factor indicating the start point of a non-zero element (e.g., 1), and kL can be a factor indicating the length of a non-zero element. k can be derived as shown in the following table:
[0154] [Table 6]
[0155] Referring to Table 6, the above related vector A k can contain 16 elements. That is, the above related vector A k can be a 1x16 dimensional vector. k A factor A indicating the starting point of the non-zero elements of ks The value of can be derived as 0, in which case the above factor A ks The starting point of the non-zero element is the related vector A k Also, the above related vector A k A factor A indicating the length of the non-zero elements of kL The value of can be derived as 8, in which case the above factor A kLcan indicate that the length of the non-zero elements is 8. k Based on the above factors, Θ can be derived as a vector in which the first to eighth elements are 1 and the remaining elements are 0, as shown in Table 6.
[0156] 10 shows an example of deriving related vectors based on two factors for the related vectors to perform selective transformation. Referring to FIG. 10, a related matrix can be derived based on the factors for each related vector, and transform coefficients for the current block can be derived based on the related matrix and the modified transform matrix.
[0157] On the other hand, the starting point of each non-zero element of the associated vector and the number of non-zero elements can be derived as fixed values, or the starting point of each non-zero element of the associated vector and the number of non-zero elements can be derived in various ways.
[0158] For example, the starting points of the non-zero elements of the associated vector and the number of non-zero elements may be derived based on the size of the upper-left target region to be transformed. Here, the size of the upper-left target region may indicate the number of transform coefficients of the upper-left target region or the product of the height and width of the upper-left target region. In another example, the starting points of the non-zero elements of the associated vector and the number of non-zero elements may be derived based on the intra-prediction mode of the current block. Specifically, for example, the starting points of the non-zero elements of the associated vector and the number of non-zero elements may be derived based on whether the intra-prediction mode of the current block is a non-directional intra-prediction mode.
[0159] Alternatively, for example, the starting points of the non-zero elements of the associated vector and the number of non-zero elements may be preset. Alternatively, for example, information indicating the starting points of the non-zero elements of the associated vector and information indicating the number of non-zero elements may be signaled, and the associated vector may be derived based on the information indicating the starting points of the non-zero elements and the information indicating the number of non-zero elements. Alternatively, other information may be used instead of the information indicating the starting points of the non-zero elements. For example, information indicating the last position of the non-zero elements may be used instead of the information indicating the starting points of the non-zero elements, and the associated vector may be derived based on the information indicating the last position of the non-zero elements.
[0160] Meanwhile, the method of deriving the related vectors based on the factors can be applied to separable transformations and non-separable transformations such as the simplified transformation HyGT.
[0161] Figure 11 schematically illustrates a video encoding method using an encoding device according to the present invention. The method disclosed in Figure 11 can be performed by the encoding device disclosed in Figure 1. Specifically, for example, S1100 in Figure 11 can be performed by a subtraction unit of the encoding device, S1110 can be performed by a transformation unit of the encoding device, and S1120 can be performed by an entropy encoding unit of the encoding device. Also, although not shown, a process of deriving a prediction sample can be performed by a prediction unit of the encoding device.
[0162] The encoding device derives residual samples of a current block (S1100). For example, the encoding device may determine whether to perform inter prediction or intra prediction on the current block, and may determine a specific inter prediction mode or a specific intra prediction mode based on an RD cost. Depending on the determined mode, the encoding device may derive predicted samples for the current block, and may derive the residual samples by adding original samples for the current block and the predicted samples.
[0163] The encoding apparatus derives transform coefficients of the current block based on a selective transform of the residual samples (S1110). The selective transform may be performed based on a modified transform matrix, where the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors may include a selected specific number of elements from N elements. The selective transform may also be performed on an upper-left target region of the current block, where N may be the number of residual samples located in the upper-left target region. Alternatively, N may be a value obtained by multiplying the width and height of the upper-left target region. For example, N may be 16 or 64.
[0164] The encoding device can perform a core transform on the residual samples to derive modified transform coefficients, and can perform the selective transform on the modified transform coefficients located in the upper left target area of the target block based on an association matrix including association vectors for the modified transform matrix and the modified basis vectors to derive transform coefficients of the target block.
[0165] Specifically, the core transform for the residual samples may be performed as follows: The encoding apparatus may determine whether to apply an adaptive multiple core transform (AMT) to the current block. In this case, an AMT flag indicating whether the adaptive multiple core transform for the current block is applied may be generated. If the AMT is not applied to the current block, the encoding apparatus may derive a DCT type 2 as a transform kernel for the current block, and may perform a transform on the residual samples based on the DCT type 2 to derive the modified transform coefficients.
[0166] When the AMT is applied to the current block, the encoding apparatus may configure a transform subset for a horizontal transform kernel and a transform subset for a vertical transform kernel, derive a horizontal transform kernel and a vertical transform kernel based on the transform subset, and perform a transform on the residual sample based on the horizontal transform kernel and the vertical transform kernel to derive modified transform coefficients. Here, the transform subset for the horizontal transform kernel and the transform subset for the vertical transform kernel may include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Transform index information may also be generated, and the transform index information may include an AMT horizontal flag indicating the horizontal transform kernel and an AMT vertical flag indicating the vertical transform kernel. Meanwhile, the transform kernels may be referred to as transform types or transform cores.
[0167] When the modified transform coefficients are derived, the encoding device may derive the transform coefficients of the target block by performing the selective transform on the modified transform coefficients located in the upper left target area of the target block based on the modified transform matrix and an association matrix including association vectors for the modified basis vectors. Modified transform coefficients other than the modified transform coefficient located in the upper left target area of the target block may be derived as the transform coefficients of the target block as they are.
[0168] Specifically, modified transform coefficients for elements of the associated vector that are 1 among the modified transform coefficients located in the upper left target area may be derived, and transform coefficients of the target block may be derived based on the derived modified transform coefficients and the modified basis vector. Here, the associated vector for the modified basis vector may include N elements, and the N elements may include elements of 1 and / or elements of 0, and the number of elements of 1 may be A. Also, the modified basis vector may include the A elements.
[0169] Meanwhile, in one example, the modified transformation matrix may include N modified basis vectors, and the association matrix may include N association vectors. The association vectors may include the same number of 1 elements, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of 1 elements, and the modified basis vectors may not all include the same number of elements.
[0170] Alternatively, in another example, the modified transformation matrix may include R modified basis vectors, and the association matrix may include R association vectors. R may be a reduced coefficient, and R may be less than N. The association vectors may include the same number of 1 elements, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of 1 elements, and the modified basis vectors may not all include the same number of elements.
[0171] Meanwhile, the related vector may be configured to have consecutive elements of 1. In this case, in one example, information about the related vector may be entropy encoded. For example, the information about the related vector may include information indicating the start point of the elements of 1 and information indicating the number of elements of 1. Alternatively, for example, the information about the related vector may include information indicating the last position of the elements of 1 and information indicating the number of elements of 1.
[0172] In another example, the associated vector may be derived based on the size of the upper left target region, for example, the starting point of one element of the associated vector and the number of one elements may be derived based on the size of the upper left target region.
[0173] Alternatively, in another example, the associated vector may be derived based on the intra prediction mode of the current block. For example, the starting point of one element of the associated vector and the number of elements may be derived based on the intra prediction mode. Furthermore, for example, the starting point of one element of the associated vector and the number of elements may be derived based on whether the intra prediction mode is a non-directional intra prediction mode.
[0174] The encoding device encodes information about transform coefficients (S1330). The information about the transform coefficients may include information about the size, position, etc. of the transform coefficients. As described above, the information about the associated vectors may be entropy encoded. For example, the information about the associated vectors may include information indicating the start point of an element of 1 and information indicating the number of elements of 1. Alternatively, for example, the information about the associated vectors may include information indicating the last position of an element of 1 and information indicating the number of elements of 1.
[0175] The video information including the information about the transform coefficients and / or the information about the associated vectors may be output in the form of a bitstream. The video information may further include prediction information. The prediction information may include, as information related to the prediction procedure, prediction mode information and information about motion information (e.g., when inter-prediction is applied).
[0176] The output bitstream can be transmitted to a decoding device via a storage medium or a network.
[0177] Figure 12 schematically illustrates an encoding device that performs a video encoding method according to the present invention. The method disclosed in Figure 11 can be performed by the encoding device disclosed in Figure 12. Specifically, for example, the adder unit of the encoding device of Figure 12 can perform S1100 of Figure 11, the transform unit of the encoding device can perform S1110, and the entropy encoding unit of the encoding device can perform S1120 to S1130. Also, although not shown, the process of deriving predicted samples can be performed by a prediction unit of the encoding device.
[0178] Figure 13 schematically illustrates a video decoding method by a decoding device according to the present invention. The method disclosed in Figure 13 can be performed by the decoding device disclosed in Figure 2. Specifically, for example, steps S1300 to S1310 in Figure 13 can be performed by an entropy decoding unit of the decoding device, step S1320 can be performed by an inverse transform unit of the decoding device, and step S1330 can be performed by an adder unit of the decoding device. Also, although not shown, the process of deriving predicted samples can be performed by a prediction unit of the decoding device.
[0179] The decoding device derives transform coefficients of the target block from the bitstream (S1300). The decoding device may derive the transform coefficients of the target block by decoding information about the transform coefficients of the target block received through the bitstream. The received information about the transform coefficients of the target block may be represented as residual information.
[0180] The decoding apparatus derives residual samples for the current block based on a selective transform of the transform coefficients (S1310). The selective transform may be performed based on a modified transform matrix, where the modified transform matrix is a matrix including modified basis vectors, and the modified basis vectors may include a selected specific number of elements from among N elements. The selective transform may also be performed on transform coefficients located in an upper-left target region of the current block, where N may be the number of transform coefficients located in the upper-left target region. Alternatively, N may be a value obtained by multiplying the width and height of the upper-left target region. For example, N may be 16 or 64.
[0181] The decoding device can perform the selective transformation on the transformation coefficients located in the upper left target area of the target block based on the modified transformation matrix and an association matrix including an association vector for the modified basis vector, thereby deriving modified transformation coefficients.
[0182] Specifically, transform coefficients for elements of the associated vector set to 1 among the transform coefficients located in the upper left target area may be derived, and modified transform coefficients may be derived based on the derived transform coefficients and the modified basis vector. Here, the associated vector for the modified basis vector may include N elements, and the N elements may include elements set to 1 and / or elements set to 0, and the number of elements set to 1 may be A. Also, the modified basis vector may include the A elements.
[0183] Meanwhile, in one example, the modified transformation matrix may include N modified basis vectors, and the association matrix may include N association vectors. The association vectors may include the same number of 1 elements, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of 1 elements, and the modified basis vectors may not all include the same number of elements.
[0184] Alternatively, in another example, the modified transformation matrix may include R modified basis vectors, and the association matrix may include R association vectors. R may be a reduced coefficient, and R may be less than N. The association vectors may include the same number of 1 elements, and the modified basis vectors may all include the same number of elements. Alternatively, the association vectors may not include the same number of 1 elements, and the modified basis vectors may not all include the same number of elements.
[0185] On the other hand, the related vector may be configured to have consecutive 1 elements. In this case, in one example, information about the related vector may be acquired from a bitstream, and the related vector may be derived based on the information about the related vector. For example, the information about the related vector may include information indicating the start point of the 1 elements and information indicating the number of 1 elements. Alternatively, for example, the information about the related vector may include information indicating the last position of the 1 elements and information indicating the number of 1 elements.
[0186] Alternatively, in another example, the associated vector may be derived based on the size of the upper left target region, for example, the starting point of an element of the associated vector and the number of elements of the element may be derived based on the size of the upper left target region.
[0187] Alternatively, in another example, the associated vector may be derived based on the intra prediction mode of the current block. For example, the starting point of one element of the associated vector and the number of elements may be derived based on the intra prediction mode. Furthermore, for example, the starting point of one element of the associated vector and the number of elements may be derived based on whether the intra prediction mode is a non-directional intra prediction mode.
[0188] If the modified transform coefficients are derived, the decoding device may perform a core transform on the current block including the modified transform coefficients to derive the residual samples.
[0189] The core transform for the current block may be performed as follows: A decoding device may acquire an Adaptive Multiple Transform (AMT) flag indicating whether an AMT is applied from a bitstream, and if the value of the AMT flag is 0, the decoding device may derive a DCT type 2 as a transform kernel for the current block, and may perform an inverse transform on the current block including the modified transform coefficients based on the DCT type 2 to derive the residual samples.
[0190] When the AMT flag has a value of 1, the decoding device may configure a transform subset for a horizontal transform kernel and a transform subset for a vertical transform kernel, derive a horizontal transform kernel and a vertical transform kernel based on transform index information acquired from the bitstream and the transform subsets, and derive the residual sample by performing an inverse transform on the current block including the modified transform coefficients based on the horizontal transform kernel and the vertical transform kernel. Here, the transform subset for the horizontal transform kernel and the transform subset for the vertical transform kernel may include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. In addition, the transform index information may include an AMT horizontal flag indicating one of the candidates included in the transform subset for the horizontal transform kernel and an AMT vertical flag indicating one of the candidates included in the transform subset for the vertical transform kernel. Meanwhile, the transform kernels may be referred to as transform types or transform cores.
[0191] The decoding device generates a reconstructed picture based on the residual samples (S1320). The decoding device may generate the reconstructed picture based on the residual samples. For example, the decoding device may perform inter-prediction or intra-prediction on a current block based on prediction information received through a bitstream to derive prediction samples, and may generate the reconstructed picture by adding the prediction samples and the residual samples. As described above, in-loop filtering procedures such as deblock filtering, SAO, and / or ALF procedures may be applied to the reconstructed picture as needed to improve subjective / objective image quality.
[0192] Figure 14 schematically illustrates a decoding device that performs a video decoding method according to the present invention. The method disclosed in Figure 13 can be performed by the decoding device disclosed in Figure 14. Specifically, for example, the entropy decoding unit of the decoding device of Figure 14 can perform S1300 of Figure 13, the inverse transform unit of the decoding device of Figure 14 can perform S1310 of Figure 13, and the adder unit of the decoding device of Figure 16 can perform S1320 of Figure 15. Also, although not shown, the process of deriving predicted samples can be performed by the prediction unit of the decoding device of Figure 14.
[0193] According to the present invention, the amount of data that must be transmitted for residual processing can be reduced through efficient conversion, thereby improving residual coding efficiency.
[0194] In addition, according to the present invention, a non-separable transform can be performed based on a transform matrix composed of basis vectors including a selected specific number of elements, thereby reducing the memory load and computational complexity for the non-separable transform and improving residual coding efficiency.
[0195] In addition, according to the present invention, a non-separable transform can be performed based on a transform matrix with a simplified structure, thereby reducing the amount of data that must be transmitted for residual processing and improving residual coding efficiency.
[0196] In the above-described embodiments, the method is described based on a sequence (flowchart) of steps or blocks, but the present invention is not limited to the sequence of steps, and some steps may occur in a different sequence or simultaneously than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps in the flowcharts may be deleted without affecting the scope of the present invention.
[0197] The above-described method according to the present invention can be implemented in software form, and the encoding device and / or decoding device according to the present invention can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0198] When an embodiment of the present invention is implemented in software, the above-described methods may be implemented as modules (processes, functions, etc.) that perform the above-described functions. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the drawings may be implemented and performed on a computer, processor, microprocessor, controller, or chip.
[0199] In addition, the decoding device and encoding device to which the present invention is applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, a medical video device, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0200] Furthermore, a processing method according to the present invention can be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices on which computer-readable data is stored. Examples of the computer-readable recording medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmitted via the Internet). A bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. Furthermore, embodiments of the present invention can be embodied in a computer program product using program code, which can be executed by a computer according to embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0201] Furthermore, the content streaming system to which the present invention is applied can be broadly divided into an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0202] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. In another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server may be omitted. The bitstream may be generated by an encoding method or a bitstream generation method to which the present invention is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0203] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary that informs the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0204] The streaming server may receive content from a media storage and / or an encoding server. For example, if the content is received from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0205] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ULTRABOOK (registered trademark), a wearable device (e.g., a smartwatch, smart glasses, or a head mounted display (HMD)), a digital TV, a desktop computer, and a digital signage. Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
Claims
1. A video decoding method performed by a decoding device, comprising: deriving dequantized coefficients of the current block from the bitstream; deriving modified transform coefficients by performing a secondary inverse transform according to a non-separable transform on the dequantized coefficients; deriving residual samples for the current block by performing an inverse linear transform on the modified transform coefficients; generating a reconstructed picture by adding the residual samples for the current block and prediction samples for the current block; the quadratic inverse transformation according to the non-separable transformation is performed by using a transformation matrix; the transformation matrix includes basis vectors, each of the basis vectors includes M elements, where M is less than N; the number of the basis vectors is R, and R is less than N; N is the number of transform coefficients located in the upper left target region to which the non-separable transform is applied in the target block, For the upper left target area of size 8x8, N is equal to 64, A method wherein the number of the dequantized coefficients to which the quadratic inverse transform is applied is the same as the number R of the basis vectors.
2. A video encoding method performed by an encoding device, comprising: deriving residual samples for a current block; deriving modified transform coefficients of the current block by performing a linear transform on the residual samples; deriving transform coefficients of the current block by performing a quadratic transform according to a non-separable transform on the modified transform coefficients; encoding information about the transform coefficients; the quadratic transformation according to the non-separable transformation is performed by using a transformation matrix; the transformation matrix includes basis vectors, each of the basis vectors includes M elements, where M is less than N; the number of the basis vectors is R, and R is less than N; N is the number of modified transform coefficients located in the upper-left target region to which the non-separable transform is applied in the target block; For the top left target region of size 8x8, N is equal to 64, A method wherein the number of transform coefficients derived by performing the quadratic transform is the same as the number R of basis vectors in the transform matrix.
3. 1. A transmission method for data including a video related bitstream, comprising: generating the bitstream for the video, the bitstream comprising: deriving residual samples for a current block; deriving modified transform coefficients of the current block by performing a linear transform on the residual samples; deriving transform coefficients of the current block by performing a quadratic transform according to a non-separable transform on the modified transform coefficients; encoding information about the transform coefficients; transmitting the data including the bitstream; the quadratic transformation according to the non-separable transformation is performed by using a transformation matrix; the transformation matrix includes basis vectors, each of the basis vectors includes M elements, where M is less than N; the number of the basis vectors is R, and R is less than N; N is the number of modified transform coefficients located in the upper-left target region to which the non-separable transform is applied in the target block; For the top left target region of size 8x8, N is equal to 64, A method wherein the number of transform coefficients derived by performing the quadratic transform is the same as the number R of basis vectors in the transform matrix.
Citation Information
Patent Citations
Method and device for the transformation and method and device for the reverse transformation of images
US20130195177A1
Reduced size inverse transform for decoding and encoding
US20170034530A1
Non-separable secondary transform for video coding
US20170094313A1
Primary transform and secondary transform in video coding
US20180103252A1