Video decoding method and device
The video decoding method addresses the challenge of determining optimal primary transforms by using specific parameters to determine conversion types and zero-out regions, resulting in improved compression performance.
Patent Information
- Application Number
- JP2025046894
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-11-25
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-12-16
AI Technical Summary
Existing video decoding methods lack the ability to efficiently determine and apply the optimal primary transform type for blocks during the decoding process, leading to suboptimal compression performance.
A video decoding method that acquires specific parameters indicating the applicability of Multiple Transform Sets (MTS) and the size of the block to be decoded, determines the appropriate conversion type, sets a zero-out region, and performs inverse conversion based on these determinations.
This approach enables efficient inverse conversion under specific conditions and improves compression performance by applying an optimized conversion type to each block during decoding.
Smart Images

Figure 2025089399000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video coding technology, and more particularly, to a method for determining the type of primary transform of a block to be decoded during the video decoding process.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video data. Therefore, when storing video data using a medium such as an existing wired or wireless broadband line, the transmission cost and storage cost increase.
[0003] Since the establishment of the HEVC (High Efficiency Video Coding) video codec in 2013, immersive videos and virtual reality services using 4K and 8K video images have spread, leading to the initiation of the standardization work for the VVC (Versatile Video Coding), a next-generation video codec aiming for a performance improvement of more than twice that of HEVC. Currently, the standardization work is in full swing. VVC is being developed by the JVET (Joint Video Exploration Team), jointly formed by the ISO / IEC MPEG (Moving Picture Experts Group), a video coding standardization group, and the ITU-T VCEG (Video Coding Experts Group), with the goal of improving the coding compression performance by more than twice that of HEVC. The VVC standardization started in earnest in January 2018 at the 121st Gwangju MPEG and 9th JVET meetings with the release of a Call for proposal, and in the 122nd San Diego MPEG and 10th JVET meetings, a total of 23 organizations proposed video codec technologies, marking the beginning of full-fledged video standardization. At the 122nd MPEG and 10th JVET meetings, technical reviews, objective compression performance, and subjective image quality evaluations were conducted on the video codec technologies proposed by each organization, and some of the many technologies were adopted to release the Working Draft (WD) 1.0 and the VTM (VVC Test Mode) 1.0, a video reference software. After the conclusion of the 127th MPEG and 15th JVET meetings in July 2019, the Committee Draft (CD) of the VVC standard was completed, and standardization is progressing towards the goal of formulating the Final Draft International Standard (FDIS) in October 2020.
[0004] Conventionally, in the encoding structure hierarchically divided by a quadtree in HEVC, VVC adopts a partitioning block structure that combines a QTBT (QuadTree Binary Tree) and a TT (Ternary Tree). This enables more flexible generation or processing of the prediction residual signal compared to HEVC, resulting in further improved compression performance compared to HEVC. In addition to such a basic block structure, new technologies not used in existing codecs, such as Adaptive Loop Filter (ALF) technology, Affine Motion Prediction (AMP) technology as a motion prediction technology, and Decoder-side Motion Vector Refinement (DMVR) technology, have been adopted as standard technologies. As for the conversion and quantization technology, the DCT-II, which is a conversion kernel widely used in existing video codecs, continues to be used, and it has been changed to be applicable to even larger block sizes of the applicable blocks. Also, in existing HEVC, the DST-7 kernel, which has been applied to small conversion blocks such as 4×4, has been extended to larger conversion blocks, and a new conversion kernel, DCT-8, has been added as a conversion kernel.
[0005] On the other hand, in the HEVC standard, when encoding or decoding video, since conversion is performed using one type of conversion, it was not necessary to transmit information regarding the type of conversion for the video. However, in the new technology, since Multiple Transform Selection using DCT-II, DCT-8, and DCT-7 can be applied, a technology for defining whether to apply MTS and which primary conversion type is applied during decoding is required.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The technical problem of the present invention is to perform inverse conversion in a method pre-defined under specific conditions.
[0007] Another technical problem of the present invention is to perform decoding by applying a conversion type optimized for a block to be decoded.
Means for Solving the Problems
[0008] According to one aspect of the present invention, there is provided a video decoding method performed by a video decoder. The video decoding method includes steps of acquiring information on a parameter indicating whether a Multiple Transform Set (MTS) for a block to be decoded is applicable, the width of the block to be decoded, and the height of the block to be decoded; determining a conversion type of the block to be decoded based on at least one of the parameter indicating whether the Multiple Transform Set for the block to be decoded is applicable, the width of the block to be decoded, and the height of the block to be decoded; setting a zero-out region of the block to be decoded based on at least one of the parameter indicating whether the Multiple Transform Set for the block to be decoded is applicable, the width of the block to be decoded, and the height of the block to be decoded; and performing inverse conversion on the block to be decoded based on the determined zero-out region and conversion type of the block to be decoded.
[0009] According to another aspect of the present invention, in the step of determining the conversion type of the block to be decoded, when at least one of the width or height of the block to be decoded has a value greater than 32, it is determined that the block to be decoded is converted using a default conversion.
[0010] According to still another aspect of the present invention, in the step of setting the zero-out region of the block to be decoded, when one of the width or height of the block to be decoded has a value greater than 32, a region where the width or height of the block to be decoded is greater than 32 is set as the zero-out region.
[0011] According to still another aspect of the present invention, the parameter indicating whether a multiple transformation set is applicable to the block to be decoded is sps_mts_enabled_flag.
[0012] According to still another aspect of the present invention, a video decoding method performed by a video decoder is provided. The video decoding method includes steps of obtaining at least one of information regarding applicability of a multiple transformation set to a block to be decoded, information regarding a prediction mode, information regarding applicability of a secondary transformation, information regarding applicability of prediction using a matrix, and information regarding the size of the block to be decoded; determining whether an implicit multiple transformation set is applied to the block to be decoded based on at least one of information regarding applicability of a multiple transformation set to the block to be decoded, information regarding a prediction mode, information regarding applicability of a secondary transformation, and information regarding applicability of prediction using a matrix; obtaining information regarding a transformation type based on information regarding whether the implicit multiple transformation set is applied to the block to be decoded and information regarding the size of the block to be decoded; and performing an inverse transformation based on the information regarding the transformation type.
[0013] According to still another aspect of the present invention, the step of determining whether the implicit multiple transformation set is applied includes determining whether the implicit multiple transformation set is applied by using information regarding applicability of a multiple transformation set to the block to be decoded, information regarding a prediction mode, information regarding applicability of a secondary transformation, and information regarding applicability of prediction using a matrix.
[0014] According to another aspect of the present invention, the implicit multiple transformation set includes one default transform and at least one extra transform.
[0015] According to another aspect of the present invention, based on the information regarding the size of the block to be decoded, in the step of obtaining information regarding the transform type, when the horizontal axis length of the block to be decoded is all 4 or more and 16 or less, at least one of the extra transform types is applied to the block to be decoded in the horizontal axis direction.
[0016] According to another aspect of the present invention, based on the information regarding the size of the block to be decoded, in the step of obtaining information regarding the transform type, when the vertical axis length of the block to be decoded is all 4 or more and 16 or less, at least one of the extra transform types is applied to the block to be decoded in the vertical axis direction.
[0017] According to another aspect of the present invention, the information regarding the application of the multiple transformation set of the block to be decoded includes at least one of sps_mts_enabled_flag and sps_explicit_mts_intra_enabled_flag.
[0018] According to another aspect of the present invention, the information regarding the prediction mode includes CuPredMode.
[0019] According to another aspect of the present invention, the information regarding the application of the secondary transformation includes lfnst_idx.
[0020] According to another aspect of the present invention, the information regarding the application of the prediction using the matrix includes intra_mip_flag.
[0021] According to another aspect of the present invention, the information regarding the conversion type of the block to be decoded includes information regarding the horizontal axis conversion type and information regarding the vertical axis conversion type, respectively.
[0022] According to another aspect of the present invention, the step of determining whether an implicit multiple conversion set is applied to the block to be decoded is obtained by further confirming whether the block to be decoded is a luma block.
[0023] According to another aspect of the present invention, a video decoding apparatus including a memory and at least one processor is provided. The video decoding apparatus obtains at least one of information regarding whether a multiple conversion set is applied to a block to be decoded, information regarding a prediction mode, information regarding whether a secondary conversion is applied, information regarding whether prediction using a matrix is applied, and information regarding the size of the block to be decoded, determines whether an implicit multiple conversion set is applied to the block to be decoded based on at least one of information regarding whether a multiple conversion set is applied to the block to be decoded, information regarding a prediction mode, information regarding whether a secondary conversion is applied, and information regarding whether prediction using a matrix is applied, obtains information regarding the conversion type based on information regarding whether an implicit multiple conversion set is applied to the block to be decoded and information regarding the size of the block to be decoded, and includes at least one processor including an inverse conversion unit that performs an inverse conversion based on the information regarding the conversion type.
Advantages of the Invention
[0024] According to the present invention, an inverse conversion can be performed in a method pre-defined under specific conditions.
[0025] In addition, by performing decoding by applying an optimized conversion type to the block to be decoded, an improvement effect in compression performance can be expected.
Brief Description of the Drawings
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Best Mode for Carrying Out the Invention
[0027] The present invention may be subject to various modifications and may have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail below. However, this is not intended to limit the present invention to specific embodiments. The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the technical idea of the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0028] On the other hand, each configuration in the drawings described in the present invention is shown independently for the convenience of explaining separate characteristic functions, and does not mean that each configuration is realized by separate hardware or separate software. For example, two or more of the configurations may be combined to form one configuration, or one configuration may be divided into a plurality of configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0029] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components will be omitted.
[0030] On the one hand, the present invention relates to video / video coding. For example, the method / embodiment disclosed in the present invention can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0031] In this specification, a "picture" generally means a unit indicating one video in a specific time period, and a "slice" is a unit constituting a part of a picture in coding. One picture may be composed of a plurality of slices, and if necessary, pictures and slices may be used interchangeably.
[0032] A "pixel" or "pel" may mean the smallest unit constituting one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value, or only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0033] A "unit" indicates a basic unit of video processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block may indicate a set of samples or transform coefficients consisting of M columns and N rows.
[0034] FIG. 1 is a diagram schematically showing the configuration of a video encoding apparatus to which the present invention is applied.
[0035] Referring to FIG. 1, the video encoding apparatus 100 may include a picture splitting unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a conversion unit 122, a quantization unit 123, a reordering unit 124, an inverse quantization unit 125, and an inverse transformation unit 126.
[0036] The picture splitting unit 105 splits the input picture into at least one processing unit.
[0037] As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a Coding Tree Unit (CTU) by a Quad-tree binary-tree (QTBT) structure. For example, one coding unit may be split into a plurality of nodes at a deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure may be applied later. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that cannot be further split, the coding procedure according to the present invention may be performed. In this case, based on the coding efficiency according to video characteristics, etc., the largest coding unit may be immediately used as the final coding unit, or if necessary, the coding unit may be recursively split into coding units at a further deeper depth, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, conversion, and restoration described later.
[0038] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit may be split from the largest coding unit (LCU) into coding units of lower depths according to a quad-tree structure. In this case, based on coding efficiency according to video characteristics, etc., the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of even lower depths so that a coding unit of an optimal size is used as the final coding unit. When the smallest coding unit (SCU) is set, the coding unit is not split into coding units smaller than the smallest coding unit. Here, the final coding unit means a coding unit that serves as a basis for partitioning or splitting into a prediction unit or a transform unit. The prediction unit is a unit that is partitioned from the coding unit and may be a unit for sample prediction. At this time, the prediction unit may be divided into subblocks. The transform unit may be split from the coding unit according to a quad-tree structure and may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transform unit may be referred to as a transform block (TB). The prediction block or prediction unit means a specific region in block form within a picture and may include an array of prediction samples.Also, a transform block or a transform unit means a specific area in block form within a picture and may include an array of transform coefficients or residual samples.
[0039] The prediction unit 110 performs prediction on a block to be processed (hereinafter referred to as the current block) and generates a predicted block including prediction samples for the current block. The unit of prediction performed by the prediction unit 110 may be a coding block, a transform block, or a prediction block.
[0040] The prediction unit 110 determines whether intra prediction or inter prediction is applied to the current block. As an example, the prediction unit 110 determines whether intra prediction or inter prediction is applied in CU units.
[0041] In the case of intra prediction, the prediction unit 110 can derive prediction samples for the current block based on reference samples outside the current block within the picture (hereinafter referred to as the current picture) to which the current block belongs. At this time, the prediction unit 110 may (i) derive prediction samples based on the average or interpolation of neighboring reference samples of the current block, or (ii) derive the prediction samples based on reference samples existing in a specific (prediction) direction with respect to the prediction samples among the neighboring reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angle mode, and in the case of (ii), it is called a directional mode or an angular mode. The prediction mode in intra prediction may have, for example, 33 directional prediction modes and at least 2 non-directional modes. The non-directional modes may include a DC prediction mode and a planar mode. The prediction unit 110 may determine the prediction mode applied to the current block using the prediction mode applied to the adjacent block.
[0042] In the case of inter prediction, the prediction unit 110 can derive a prediction sample for the current block based on samples specified by motion vectors on the reference picture. The prediction unit 110 can derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP (motion vector prediction) mode. In the case of the skip mode and the merge mode, the prediction unit 110 may use the motion information of an adjacent block as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, the motion vector of an adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be derived by using it as the motion vector predictor of the current block.
[0043] In the case of inter prediction, the adjacent blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in a reference picture. The reference picture including the temporal neighboring blocks may be referred to as a collocated picture (colPic). The motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and output in the form of a bit stream.
[0044] When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top picture on the reference picture list may be used as a reference picture. The reference pictures included in the reference picture list may be sorted based on the difference in POC (Picture order count) between the current picture and the reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0045] The subtraction unit 121 generates a residual sample, which is the difference between the original sample and the predicted sample. When the skip mode is applied, it may not be necessary to generate the residual sample as described above.
[0046] The conversion unit 122 converts the residual samples in units of conversion blocks to generate transform coefficients. The conversion unit 122 can perform conversion according to the size of the conversion block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the conversion block. For example, if intra prediction is applied to the coding block or the prediction block that overlaps with the conversion block, and the conversion block is a 4×4 residual array, the residual samples are converted using a DST (Discrete Sine Transform) conversion kernel, and in other cases, the residual samples are converted using a DCT (Discrete Cosine Transform) conversion kernel.
[0047] The quantization unit 123 quantizes the transform coefficients to generate quantized transform coefficients.
[0048] The reordering unit 124 reorders the quantized transform coefficients. The reordering unit 124 can reorder the quantized transform coefficients in block form into a one-dimensional vector form by a coefficient scanning method. Here, although the reordering unit 124 is described separately, it may be a part of the quantization unit 123.
[0049] The entropy encoding unit 130 performs entropy encoding on the quantized transform coefficients. The entropy encoding may include encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 130 may encode, together or separately, information necessary for video restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The entropy-encoded information may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units.
[0050] The inverse quantization unit 125 inverse-quantizes the values (quantized transform coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inverse-transforms the values inverse-quantized by the inverse quantization unit 125 to generate residual samples.
[0051] The addition unit 140 combines the residual samples and the prediction samples to restore the picture. The residual samples and the prediction samples are added in block units to generate a restored block. Here, although the addition unit 140 has been described in a separate configuration, it may be a part of the prediction unit 110. On the other hand, the addition unit 140 is also called a restoration unit or a restored block generation unit.
[0052] For the reconstructed picture, the filter unit 150 can apply a deblocking filter and / or a sample adaptive offset. Through deblocking filtering and / or sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process can be corrected. The sample adaptive offset may be applied on a sample-by-sample basis or may be applied after the deblocking filtering process is completed. The filter unit 150 can also apply an ALF (Adaptive Loop Filter) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset have been applied.
[0053] The memory 160 stores the reconstructed picture (decoded picture) or information necessary for encoding / decoding. Here, the reconstructed picture may be the reconstructed picture for which the filtering procedure has been completed by the filter unit 150. The stored reconstructed picture may be utilized as a reference picture for (inter) prediction of other pictures. For example, the memory 160 can store the (reference) picture used for inter prediction. At this time, the picture used for inter prediction may be specified by a reference picture set or a reference picture list.
[0054] Figure 2 shows an example of a video encoding method performed by a video encoding device. Referring to Figure 2, the video encoding method may include a block partitioning, intra / inter prediction, transform, quantization, and entropy encoding process. For example, a current picture may be divided into a plurality of blocks, and a predicted block of the current block may be generated by intra / inter prediction. A residual block of the current block may be generated by subtracting the input block of the current block from the predicted block. Thereafter, a coefficient block, that is, a transform coefficient of the current block, may be generated by a transform on the residual block. The transform coefficient may be quantized and entropy encoded and stored in a bitstream.
[0055] Figure 3 is a diagram schematically explaining a configuration of a video decoding device to which the present invention is applied.
[0056] Referring to Figure 3, the video decoding device 300 may include an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 may include a reordering unit 321, an inverse quantization unit 322, and an inverse transform unit 323.
[0057] When a bitstream including video information is input, the video decoding device 300 can restore the video corresponding to the process in which the video information is processed by the video encoding device.
[0058] For example, the video decoding device 300 can perform video decoding using the processing units applied in the video encoding device. Therefore, the processing unit block for video decoding may be, as an example, a coding unit, or may be, as another example, a coding unit, a prediction unit, or a transformation unit. The coding unit may be divided from the largest coding unit by a quad tree structure and / or a binary tree structure.
[0059] The prediction unit and the transformation unit may be further used as appropriate. In this case, the prediction block may be a block derived from or partitioned from the coding unit and may be a unit for sample prediction. At this time, the prediction unit may be divided into sub-blocks. The transformation unit may be divided from the coding unit by a quad tree structure and may be a unit for deriving transformation coefficients or a unit for deriving a residual signal from the transformation coefficients.
[0060] The entropy decoding unit 310 parses the bit stream and outputs information necessary for video restoration or picture restoration. For example, the entropy decoding unit 310 can decode the information in the bit stream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transformation coefficient for the residual.
[0061] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information adjacent to the decoding target block, and the decoding information of the decoding target block or the information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.
[0062] Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient, is input to the reordering unit 421.
[0063] The reordering unit 321 reorders the quantized transform coefficients in a two-dimensional block form. The reordering unit 321 can perform reordering corresponding to the coefficient scanning performed in the encoding device. Here, although the reordering unit 321 is described separately, it may be a part of the inverse quantization unit 322.
[0064] The inverse quantization unit 322 can inverse-quantize the quantized transform coefficients based on the (inverse) quantization parameters to output the transform coefficients. At this time, the information for deriving the quantization parameters may be signaled from the encoding device.
[0065] The inverse transform unit 323 can inverse-transform the transform coefficients to derive residual samples.
[0066] The prediction unit 330 can perform a prediction on the current block and generate a predicted block including a prediction sample for the current block. The unit of prediction performed by the prediction unit 330 may be a coding block, a transform block, or a prediction block.
[0067] The prediction unit 330 determines whether to apply intra prediction or inter prediction based on the information related to the prediction. At this time, the unit for determining which of intra prediction and inter prediction to apply is different from the unit for generating a prediction sample. Also, in inter prediction and intra prediction, the units for generating a prediction sample are different. For example, which of inter prediction and intra prediction to apply can be determined in CU units. Also, for example, in inter prediction, the prediction mode may be determined in PU units to generate a prediction sample, and in intra prediction, the prediction mode may be determined in PU units and the prediction sample may be generated in TU units.
[0068] In the case of intra prediction, the prediction unit 330 can derive a prediction sample for the current block based on adjacent reference samples within the current picture. The prediction unit 330 can apply a directional mode or a non-directional mode based on the adjacent reference samples of the current block to derive a prediction sample for the current block. At this time, the prediction mode to be applied to the current block may be determined using the intra prediction mode of an adjacent block. On the other hand, MIP (Matrix-based Intra Prediction) for performing a prediction based on a pre-trained matrix may be used. In this case, the number of MIP modes and the size of the matrix are defined for each block size. After downsampling the reference samples to match the size of the matrix, the matrix determined by the mode number is multiplied, and interpolation is performed to match the size of the prediction block to generate a predicted value.
[0069] In the case of inter prediction, the prediction unit 330 can derive a prediction sample for the current block based on samples specified on the reference picture by the motion vector on the reference picture. The prediction unit 330 can apply any one of the skip mode, merge mode, and MVP mode to derive a prediction sample for the current block. At this time, motion information necessary for inter prediction of the current block provided from the video encoding device, for example, information regarding a motion vector, a reference picture index, etc. may be acquired or derived based on the information regarding the prediction.
[0070] In the case of the skip mode and the merge mode, the motion information of an adjacent block may be used as the motion information of the current block. At this time, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0071] The prediction unit 330 may configure a merge candidate list with the motion information of available adjacent blocks, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled from the encoding device. The motion information may include a motion vector and a reference picture. In the skip mode and the merge mode, when the motion information of a temporally adjacent block is used, the top picture on the reference picture list may be used as the reference picture.
[0072] In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted.
[0073] In the case of the MVP mode, the motion vector of an adjacent block may be used as a motion vector predictor to derive the motion vector of the current block. At this time, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0074] As an example, when the merge mode is applied, a merge candidate list may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks that are temporally adjacent blocks. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The information related to the prediction may include a merge index indicating a candidate block having an optimal motion vector selected from among the candidate blocks included in the merge candidate list. At this time, the prediction unit 330 may derive the motion vector of the current block using the merge index.
[0075] As another example, when the MVP (Motion Vector Prediction) mode is applied, a motion vector predictor candidate list may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks that are temporally adjacent blocks. That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks that are temporally adjacent blocks may be used as motion vector candidates. The information related to the prediction may include a predicted motion vector index that indicates the optimal motion vector selected from among the motion vector candidates included in the list. At this time, the prediction unit 330 can select the predicted motion vector of the current block from among the motion vector candidates included in the motion vector candidate list using the motion vector index. The prediction unit of the encoding device obtains the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encodes this, and outputs it in the form of a bit stream. That is, the MVD is obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit 330 can obtain the motion vector difference included in the information related to the prediction, and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. Also, the prediction unit can obtain or derive a reference picture index that indicates a reference picture, etc. from the information related to the prediction.
[0076] The addition unit 340 adds the residual sample and the predicted sample to restore the current block or the current picture. The addition unit 340 may add the residual sample and the predicted sample in block units to restore the current picture. When the skip mode is applied, since the residual is not transmitted, the predicted sample becomes the restored sample. Here, the addition unit 340 has been described as a separate configuration, but it may also be a part of the prediction unit 330. On the other hand, the addition unit 340 is also called a restoration unit or a restored block generation unit.
[0077] The filter unit 350 may apply a deblocking filter sample adaptive offset and / or an ALF or the like to the restored picture. At this time, the sample adaptive offset may be applied in units of samples or may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or the sample adaptive offset.
[0078] The memory 360 stores the restored picture (decoded picture) or information necessary for decoding. Here, the restored picture may be a restored picture for which the filtering procedure has been completed by the filter unit 350. For example, the memory 360 may store a picture used for inter prediction. At this time, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The restored picture may be used as a reference picture for other pictures. Also, the memory 360 may output the restored pictures in output order.
[0079] FIG. 4 is a diagram showing an example of a video decoding method performed by a decoding apparatus. Referring to FIG. 4, the video decoding method may include entropy decoding, inverse quantization, inverse transform, and intra / inter prediction processes. For example, in the decoding apparatus, the reverse process of the encoding method may be performed. Specifically, the quantization coefficients quantized by entropy decoding of the bitstream may be obtained, and the coefficient block of the current block, that is, the quantization coefficients, may be obtained by an inverse quantization process for the quantized coefficients. The residual block of the current block may be derived by inverse transformation of the quantization coefficients, and the restored block of the current block may be derived by addition of the prediction block of the current block derived by intra / inter prediction and the residual block.
[0080] On the one hand, the operator in the embodiments described below may be defined as follows in the following table.
[0081]
Table 1
[0082] Referring to Table 1, Floor(x) represents the largest integer value less than or equal to x, Log2(u) represents the logarithm value of u with base 2, and Ceil(x) represents the smallest integer value greater than or equal to x. For example, in the case of Floor(5.93), the largest integer value less than or equal to 5.93 is 5, so it represents 5.
[0083] Also referring to Table 1, x>>y represents an operator that right-shifts x by y bits, and x<<y represents an operator that left-shifts x by y bits.
[0084] <Introduction> The HEVC standard generally uses DCT, which is a single transform type. Therefore, there is no need to transmit a separate determination process for the transform type and information regarding the determined transform type. However, currently, when the size of the luma block is 4x4 and intra prediction is performed, the DST transform type is used exceptionally.
[0085] Among the quantized coefficients that have undergone the transformation and quantization processes, the information representing the positions of non-zero coefficients can be broadly classified into three types.
[0086] 1. Position (x, y) of the last significant coefficient: The position of the lowest-ranked non-zero coefficient (hereinafter defined as the last position) in the scan order within the block to be coded
[0087] 2. Coded sub - block flag: A flag that divides a block to be coded into a number of sub - blocks and indicates whether each sub - block contains one or more non - zero coefficients (alternatively, a flag that indicates whether all coefficients are zero).
[0088] 3. Significant coefficient flag: A flag that indicates whether each coefficient within a sub - block is non - zero or zero.
[0089] Here, the position of the last significant coefficient is displayed separately for the x - axis component and the y - axis component, and each component is expressed separately as a prefix and a suffix. That is, the syntax for indicating the non - zero position of the quantized coefficient includes the following six syntaxes in total.
[0090] 1. last_sig_coeff_x_prefix 2. last_sig_coeff_y_prefix 3. last_sig_coeff_x_suffix 4. last_sig_coeff_y_suffix 5. coded_sub_block_flag 6. sig_coeff_flag
[0091] The last_sig_coeff_x_prefix indicates the prefix of the x - axis component showing the position of the last significant coefficient, and the last_sig_coeff_y_prefix indicates the prefix of the y - axis component showing the position of the last significant coefficient. Also, the last_sig_coeff_x_suffix indicates the suffix of the x - component showing the position of the last significant coefficient, and the last_sig_coeff_y_suffix indicates the suffix of the y - component showing the position of the last significant coefficient.
[0092] On the one hand, the coded_sub_block_flag indicates "0" if all coefficients within the sub-block are all zero, and indicates "1" if there is one or more non-zero coefficients. The sig_coeff_flag indicates "0" for zero coefficients and indicates "1" for non-zero coefficients. Considering the position of the last valid coefficient within the block to be coded, in the scan order, the coded_sub_block_flag syntax is transmitted only for the previously existing sub-blocks. When the coded_sub_block_flag is "1", that is, when there is one or more non-zero coefficients, the sig_coeff_flag syntax for each of all the coefficients within the sub-block is transmitted.
[0093] The HEVC standard supports the following three types of scans for coefficients.
[0094] 1) Upright diagonal 2) Horizontal 3) Vertical
[0095] When the block to be coded is coded using the inter-picture prediction method, the coefficients of the block are scanned in the upright diagonal method. When the block is coded using the intra-picture prediction method, one of the above three types is selected according to the intra-picture prediction mode to scan the coefficients of the block.
[0096] That is, when the block to be coded is coded by the video coding device and the inter-picture prediction method is used, the coefficients of the block are scanned in the upright diagonal method. When the intra-picture prediction method is used in the coding of the block to be coded, the video coding device selects one of the above three types according to the intra-picture prediction mode to scan the coefficients of the block. The above scan may be performed by the reordering unit 124 in the video coding device shown in FIG. 1, and by the scan, the two-dimensional block-form coefficients may be changed into one-dimensional vector form.
[0097] FIG. 5 is a diagram showing the scan order of sub-blocks and coefficients for the diagonal scan method.
[0098] Referring to FIG. 5, when the block of FIG. 5 is scanned in the diagonal scan method by the rearrangement unit 124 of the video encoding apparatus, the scan is performed from the sub-block 1, which is the sub-block at the uppermost left end, downward and in the upper end direction of the diagonal line, and the scan is finally performed on the 16th sub-block at the lower right end. That is, the rearrangement unit 124 performs the scan in the order of sub-blocks 1, 2, 3, …, 14, 15, 16, and rearranges the quantized transform coefficients in the two-dimensional block form into a one-dimensional vector form. Similarly, the rearrangement unit 124 of the video encoding apparatus also performs the scan on the coefficients in each sub-block in the same diagonal scan method as the scan method of the sub-block. For example, in the first sub-block, the scan is performed in the order of coefficients 0, 1, 2, …, 13, 14, 15.
[0099] However, when the scanned coefficients are stored in the bit stream, the storage order is the reverse of the scan order. That is, when the block of FIG. 10 is scanned by the rearrangement unit 124 of the video encoding apparatus, the scan is performed in the order of coefficients from 0 to 255, but the order in which each pixel is stored in the bit stream is stored in the bit stream in the order of the pixel at the 255th position to the pixel at the 0th position.
[0100] FIG. 6 is a diagram showing an example of a 32×32 block to be encoded after quantization by a video encoding device. Here, when the 32×32 block shown in FIG. 6 is scanned by the video encoding device, the diagonal method is arbitrarily used. In FIG. 6, the pixels indicated by diagonal hatching indicate non-zero coefficients, and the pixels indicated by x indicate the last significant coefficients. All other white coefficients have zero values. Here, if the coded_sub_block_flag syntax is substituted into the block of FIG. 6, among the total 64 sub-blocks, the coded_sub_block_flag information for 24 sub-blocks existing before the last position in the scan order, that is, the sub-blocks shown by thick lines in FIG. 6, is required. Among the 24 sub-blocks, the coded_sub_block_flag values for the first sub-block including the DC value and the 24th sub-block including the last position coefficient are induced to be "1", and the coded_sub_block_flag values for the remaining 22 sub-blocks are transmitted to the video decoding device by the bit stream. At this time, in the case of a sub-block including one or more non-zero coefficients among the 22 sub-blocks, the coded_sub_block_flag value is set to "1" by the video encoding device. In FIG. 6, among the 22 sub-blocks excluding the first sub-block and the 24th sub-block, the coded_sub_block_flag values of the 4th, 5th, 11th, and 18th sub-blocks, which are sub-blocks including the pixels marked in gray, are set to "1".
[0101] 1. Method for determining the primary conversion type of the block to be decoded This specification discloses a method for determining the type of primary transformation of a block to be decoded during the video decoding process. That is, when a block to be decoded is decoded by a video decoding device, in the transformation process by the video encoding device, a process of determining what transformation type was used for primary transformation and encoding is required. The primary transformation type consists of one default transformation and a plurality of additional transformations. The block to be decoded uses either the default transformation or a set of multiple transformations including the default transformation and additional transformations depending on conditions. That is, the block to be decoded may be transformed using only the default transformation in the transformation process, or may be transformed using a set of multiple transformations including the default transformation and additional transformations. From the perspective of the video decoding device, conversely, it may perform decoding by grasping whether the block to be decoded uses only the default transformation or a set of multiple transformations including the default transformation and additional transformations. When the block to be decoded uses MTS, information regarding the actually used transformation among the multiple transformations is transmitted or induced. Here, the information regarding the actually used transformation may separately exist for the horizontal axis transformation type and the vertical axis transformation type. That is, when the block to be decoded is transformed using MTS, the video decoding device may receive or determine what transformation type among the multiple transformation types was used for transformation and perform decoding.
[0102] According to one embodiment, DCT-II may be set as the default conversion, and DST-7 and DCT-8 may be set as additional conversions. At this time, the maximum size supported by DCT-II, which is the default conversion, is up to 64×64, and the maximum sizes supported by DST-7 and DCT-8, which are additional conversions, are up to 32×32. For example, when the size of the block to be decoded is 64×64, one 64×64 DCT-II is applied in the conversion process. That is, when one or more of the width and height of the block to be decoded are greater than 32 (exceeding 32), the default conversion (*) is applied immediately without applying MTS. That is, from the perspective of the video decoding device, it is only necessary to determine whether the MTS has been used for conversion when both the horizontal and vertical sizes of the block to be decoded are 32 or less. On the other hand, when one of the horizontal or vertical sizes of the block to be decoded is greater than 32, it may be determined that the default conversion has been applied for conversion. Thus, when the block to be decoded is converted by the default conversion, there is no syntax information transmitted in relation to MTS. In the present invention, for convenience, the conversion type value of DCT-II is set to "0", the conversion type value of DST-7 is set to "1", and the conversion type value of DCT-8 is set to "2", but it is not limited thereto. Table 2 below defines the conversion types assigned to each value of the trType syntax.
[0103]
Table 2
[0104] Tables 3 and 4 show an example of the conversion kernels of DST-7 and DCT-8 when the size of the block to be multiplexed is 4×4.
[0105] Table 3 shows the coefficient values of the conversion kernel when tyType is "1" (DST-7) and the size of the block to be decoded is 4×4, and Table 4 shows the coefficient values of the conversion kernel when tyType is "2" (DCT-8) and the size of the block to be decoded is 4×4.
[0106]
Table 3
[0107]
Table 4
[0108] The entire conversion area of the block to be decoded may include a zero-out area. The conversion converts the values in the pixel domain to frequency domain values. At this time, the upper left frequency area is referred to as the low frequency area, and the lower right frequency area is referred to as the high frequency area. The low frequency components reflect the general (average) characteristics of the block, and the high frequency components reflect the sharp (peculiar) characteristics of the block. Therefore, there are a large number of large values in the low frequency components, and there are a small number of small values in the high frequency components. Through the quantization process after conversion, the small number of small values in the high frequency area will mostly have 0 values. Here, in addition to the low frequency area belonging to the upper left side, the remaining area with mostly 0 values is referred to as the zero-out area, and the zero-out area may be excluded in the signaling process. The area excluding the zero-out area in the block to be decoded is referred to as the valid area.
[0109] FIG. 7 is a diagram showing the remaining zero-out area excluding m×n in the area of the M×N block to be decoded.
[0110] Referring to FIG. 7, the gray area on the upper left side indicates the low frequency area, and the white area indicates the high frequency zero-out.
[0111] As another example, when the block to be decoded is 64×64, the upper left 32×32 area becomes the valid area, and the remaining area excluding this becomes the zero-out area and is not signaled.
[0112] Also, when the size of the block to be decoded is one of 64×64, 64×32, or 32×64, the upper left 32×32 region becomes the valid region, and the remaining part excluding this becomes the zero-out region. When the decoder parses the syntax for the quantized coefficients, since it knows the size of the block, the zero-out region is not signaled. That is, a region where the width or height of the block to be decoded is greater than 32 is set as the zero-out region. At this time, since it corresponds to the case where the horizontal or vertical size of the block to be decoded is greater than 32, the transformation used is DCT-II, which is the default transformation. Since the maximum sizes of the additional transformations DST-7 and DCT-8 are supported up to 32×32, MTS is not applied to blocks of such sizes.
[0113] Also, when the size of the block to be decoded is one of 32×32, 32×16, or 16×32 and MTS is applied to the block to be decoded (for example, when using DST-7 or DCT-8), the upper left 16×16 region becomes the valid region, and the remaining part excluding this is set as the zero-out region. Here, the zero-out region may be signaled by the position of the last valid coefficient and the scan method. This is because after the encoder finishes signaling the syntax related to the quantized coefficients, it signals the MTS index (mts_idx) value. So, when the decoder parses the syntax for the quantized coefficients, it does not know the information about the transformation type. When the zero-out region is signaled in this way, the decoder can perform the transformation only on the valid region after ignoring or removing the quantized coefficients corresponding to the zero-out region. Here, the information about the actually used transformation type may exist separately for the horizontal axis transformation type and the vertical axis transformation type. For example, when the horizontal axis transformation type of the block to be decoded is DST-7 or DCT-8, the horizontal axis (width) valid region of the block to be decoded is 16, and when the vertical axis transformation type of the block to be decoded is DST-7 or DCT-8, the vertical axis (height) valid region of the block to be decoded is 16.
[0114] On the other hand, when the size of the block to be decoded is one of 32×32, 32×16, and 16×32 and MTS is not applied to the block to be decoded (for example, when DCT-II which is the default transform is used), all regions are valid regions, there is no zero-out region, and it is transformed using DCT-II which is the default transform. Here, information regarding the actually used transform may separately exist for the horizontal axis transform type and the vertical axis transform type. For example, when the horizontal axis transform type of the block to be decoded is DCT-II, the horizontal axis (width) valid region of the block to be decoded is the width of the block, and when the vertical axis transform type of the block to be decoded is DCT-II, the vertical axis (height) valid region of the block to be decoded is the height of the block. That is, all regions (width×height) of the block to be decoded become valid regions.
[0115] On the other hand, when the block to be decoded has a size smaller than 16 that is not defined as above, all regions become valid regions and there is no zero-out region. Whether MTS is applied to the block to be decoded and the transform type value are determined by implicit and / or explicit embodiments.
[0116] In the present invention, regarding the content of the syntax that represents the positions of non-zero coefficients among the quantized coefficients that have undergone the transform and quantization processes, it is the same as that of the HEVC method. However, the syntax name of coded_sub_block_flag is changed to sb_coded_flag and used. Also, the method of scanning the quantized coefficients uses the upright diagonal method.
[0117] In the present invention, MTS may be applied to a luma block (not applied to a chroma block). Also, the MTS function can be turned on / off using a flag indicating whether MTS is used, i.e., sps_mts_enabled_flag. When using the MTS function, sps_mts_enabled_flag is set to on, and the use of the explicit MTS function for each of intra-picture prediction and inter-picture prediction can be set. That is, the sps_explicit_mts_intra_enabled_flag, which is a flag indicating whether MTS is used during intra-picture prediction, and the sps_explicit_mts_inter_enabled_flag, which is a flag indicating whether MTS is used during inter-picture prediction, may be set separately. In this specification, for convenience, it is described that the values of the sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, and sps_explicit_mts_inter_enabled_flag, which are flags indicating the use of three types of MTS, are located within the SPS (sequence parameter set), but it is not limited thereto. That is, the three flags may be set at one or more positions of DCI (decoding capability information), VPS (video parameter set), SPS (sequence parameter set), PPS (picture parameter set), PH (picture header), and SH (slice header), respectively. Also, the three flags indicating the use of MTS may be defined as HLS (high level syntax).
[0118] The use of MTS can be divided into explicit and implicit methods. The explicit use of MTS means that in the SPS, the flag value indicating the use of MTS within and / or between screens is set to on, and when specific conditions are met, MTS-related information (e.g., the conversion information actually used) is transmitted. That is, the video decoder can receive the MTS-related information, based on which it can confirm the conversion type used for the block to be decoded and perform decoding accordingly. For example, in an environment where explicit MTS is used, the three flags may be set as follows.
[0119] 1. sps_mts_enabled_flag = on 2. sps_explicit_mts_intra_enabled_flag = on 3. sps_explicit_mts_inter_enabled_flag = on
[0120] The implicit use of MTS means that in the SPS, when the value of sps_mts_enabled_flag among the three flags is set to on and specific conditions are met, MTS-related information (e.g., the conversion information actually used) is induced. For example, in an environment where implicit MTS is used, the three flags may be set as follows.
[0121] 1. sps_mts_enabled_flag = on 2. sps_explicit_mts_intra_enabled_flag = off 3. sps_explicit_mts_inter_enabled_flag = off (regardless of on or off)
[0122] Hereinafter, several embodiments will be used to describe the implicit MTS method and the explicit MTS method.
[0123] 2. First Embodiment (Implicit MTS) The implicit MTS described in this embodiment may be used when the block to be decoded is encoded by the intra prediction method. That is, when the block to be decoded is encoded by the video encoding apparatus, if it is encoded by the intra prediction method, encoding and / or decoding may be performed using the implicit MTS by the video encoding apparatus and / or the video decoding apparatus. On the other hand, when decoding the block to be decoded, whether to use the implicit MTS may be indicated by the implicitMtsEnabled parameter. The video decoding apparatus may check the value of the implicitMtsEnabled parameter and determine whether to perform decoding using the implicit MTS. For example, when the implicit MTS is used for decoding, the implicitMtsEnable parameter may have a value of 1, and otherwise, the implicitMtsEnable parameter may have a value of 0. On the other hand, in this specification, the implicitMtsEnabled may be displayed as "implicit_MTS_enabled" depending on the case.
[0124] When examining the conditions of the HLS (high level syntax) for applying such implicit MTS, since the sps_mts_enabled_flag is a flag indicating whether MTS is applied regardless of whether it is implicit or explicit, it must be set to "on" for implicit MTS to be applied. On the other hand, implicit MTS is used when the block to be decoded is encoded by the video encoder using the intra prediction method. Therefore, the video decoder can determine whether to use implicit MTS by checking the value of sps_explicit_mts_intra_enabled_flag. However, the sps_explicit_mts_intra_enabled_flag is set to "on" when the block to be decoded is encoded by the video encoder using the intra prediction method and explicit MTS is applied. Therefore, when the block to be decoded is encoded by the video encoder using implicit MTS, the sps_explicit_mts_intra_enabled_flag is set to "off". On the other hand, as described above, implicit MTS is used when the block to be decoded is encoded by the video encoder using the intra prediction method. Therefore, the value of sps_explicit_mts_inter_enabled_flag indicating explicit MTS when the block to be decoded is encoded by the video encoder using the intra prediction method is not important. On the other hand, since implicit MTS may be used when the block to be decoded is encoded by the video encoder using the intra prediction method, it may be applied when CuPredMode has the value MODE_INTRA.
[0125] In summary, the conditions for the block to be decoded to be decoded using implicit MTS by the video decoder can be listed as follows.
[0126] 1) sps_mts_enabled_flag is equal to 1 2) The sps_explicit_mts_intra_enabled_flag is equal to 0 3) The CuPredMode is equal to MODE_INTRA (intra prediction method)
[0127] On the other hand, CuPredMode[0][xTbY][yTbY] indicating the prediction mode of the current position in the luma block may have a MODE_INTRA value.
[0128] The additional conditions for enabling implicit MTS are as follows.
[0129] 4) The lfnst_idx is equal to 0 5) The intra_mip_flag is equal to 0
[0130] Here, the lfnst_idx value indicates a secondary transform. When lfnst_idx = 0, it means that no secondary transform is used. The Intra_mip_flag value indicates whether a prediction method using a matrix (matrix-based intra prediction: mip), which is one of the intra prediction methods, is used. When intra_mip_flag = 0, it means that no prediction using a matrix is used, and when intra_mip_flag = 1, it means that a prediction using a matrix is used.
[0131] That is, this embodiment describes a method for setting a primary transform type (or MTS) for a decoding target block that does not use a secondary transform while predicting using a general intra prediction method. When all of the above five conditions are satisfied, the implicit MTS function is activated (see FIG. 13).
[0132] FIG. 8 is a diagram showing a method for determining whether to apply the implicit MTS function according to an embodiment of the present invention. Each step in FIG. 8 may be performed in a video decoding apparatus.
[0133] Referring to FIG. 8, the video decoding apparatus determines whether the sps_mts_enable_flag has a value of 1, the sps_explicit_mts_intra_enable_flag has a value of 0, and the CuPredMode has a MODE_INTRA value (S810). As a result of the determination, if all the conditions of S810 are satisfied, the video decoding apparatus determines whether the lfnst_idx has a value of 0 and the intra_mip_flag has a value of 0 (S820). If all the conditions of S810 and S820 are satisfied, the implicit_MTS_enabled value is set to 1 (S830). On the other hand, if the video decoding apparatus does not satisfy any of the conditions of S810 or S820, the implicit_MTS_enabled value is set to 0 (S840).
[0134] When the implicit MTS function for the block to be decoded is activated (implicit_MTS_enabled = on), the MTS value (the conversion information actually used) is derived based on the width and height of the block (see FIG. 9). At this time, the conversion must not be a sub-block transform (sbt) in which only a part of the target block undergoes the conversion process. That is, the cu_sbt_flag value of the target block is "0".
[0135] FIG. 9 is a diagram showing a method of deriving conversion information based on the width and height of the block of implicit MTS according to an embodiment of the present invention. Each step in FIG. 9 may be performed in the video decoding apparatus.
[0136] Referring to FIG. 9, the video decoding apparatus determines whether the implicit_MTS_enabled value is "1" (S910). At this time, although not shown, the video decoding apparatus may further check whether the cu_sbt_flag value has a value of "0". At this time, when the cu_sbt_flag has a value of "1", it indicates that the decoding target block has been converted by sub-block conversion in which only a part of the target block undergoes the conversion process. On the other hand, when the cu_sbt_flag has a value of "0", it indicates that the decoding target block has not been converted by sub-block conversion in which only a part of the target block undergoes the conversion process. Therefore, the operation according to FIG. 14 may be set to operate only when the cu_sbt_flag has a value of "0".
[0137] When the implicit_MTS_enabled value is 1, it is determined whether the value of nTbW is 4 or more and 16 or less (S920). When the implicit_MTS_enabled value is not "1", the operation ends. nTbW indicates the width of the conversion block and is used to determine whether DST-7, which is an additional conversion type, is used in the horizontal axis direction.
[0138] As a result of the determination in step S920, when the value of nTbW is 4 or more and 16 or less, trTypeHor is set to "1" (S930). When the value of nTbW is not 4 or more and 16 or less, trTypeHor is set to "0" (S940). At this time, the nTbW indicates the width of the conversion block and is used to determine whether DST-7, which is an additional conversion type, is used in the horizontal axis direction. At this time, when tyTypeHor is set to "0", it can be determined that the conversion block has been converted using DCT-II conversion, which is the default type conversion, in the horizontal axis direction. On the other hand, when the trTypeHor is set to "1", it can be determined that the conversion block has been converted using DST-7 conversion, which is one of the additional conversion types, in the horizontal axis direction.
[0139] Further, the video decoding device determines whether the value of nTbH has a value of 4 or more and 16 or less (S950). When the value of nTbH has a value of 4 or more and 16 or less, trTypeVer is set to "1" (S960). When the value of nTbW does not have a value of 4 or more and 16 or less, trTypeVer is set to "0" (S970). The nTbH indicates the height of the conversion block and is used to determine whether DST-7, which is an additional conversion type in the vertical axis direction, is used. At this time, when trTypeVer is set to "0", it can be determined that the conversion block is converted using DCT-II conversion, which is the default type conversion in the vertical axis direction. On the other hand, when trTypeVer is set to "1", it can be determined that the conversion block is converted using DST-7 conversion, which is one of the additional conversion types in the vertical axis direction.
[0140] FIG. 10 is a diagram showing a method of executing inverse conversion based on conversion-related parameters according to an embodiment of the present invention. Each step in FIG. 10 may be performed by a video decoding device, for example, by an inverse conversion unit of the decoding device.
[0141] Referring to FIG. 10, the video decoding device acquires sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, Y0], NTbW, and nTbH (S1010). At this time, regarding what each of sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, and IntraMipFlag[x0, Y0] indicates, it is described in detail in the related description of FIG. 8. The parameters are used to determine whether the decoding target block can apply implicit MTS. Also, NTbW and nTbH respectively indicate the width and height of the conversion block and are used to determine whether DST-7, which is an additional conversion type, is used.
[0142] Next, the video decoding device sets implicit_MTS_enabled based on the values of sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], Lfnst_idx, and IntraMipFlag[x0,Y0] (S1020). At this time, the implicit_MTS_enable may be set by performing the process of FIG. 13.
[0143] Next, the video decoding device sets trTypeHor and trTypeVer based on the values of implicit_MTS_enabled, nTbW, and nTbH (S1030). At this time, the trTypeHor and trTypeVer may be set by performing the process of FIG. 9.
[0144] Next, the video decoding device performs inverse transformation based on trTypeHor and trTypeVer (S1040). The inverse transformation applied by trTypeHor and trTypeVer may be configured according to Table 2. For example, when trTypeHor is "1" and trTypeVer is "0", DST-7 may be applied in the horizontal axis direction and DST-II may be applied in the vertical axis direction.
[0145] On the other hand, although not shown, from the perspective of the video encoding device, in order to set whether to use implicit MTS, sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, Y0], NTbW, and nTbH may be set.
[0146] 3. Second Embodiment (Explicit MTS) In this embodiment, when the MTS function is explicitly activated in the HLS (high level syntax), the conversion method applied to the decoding target block will be described. When examining the HLS (high level syntax) conditions for the application of explicit MTS, since the sps_mts_enabled_flag is a flag indicating whether MTS is applied or not regardless of whether it is implicit or explicit, it must be set to "on" for implicit MTS to be applied. On the other hand, since explicit MTS may be applied in all cases where the decoding target block is encoded by the intra-prediction method or the inter-prediction method, when explicit MTS is applied, both sps_explicit_mts_intra_enabled_flag and / or sps_explicit_mts_intra_enabled_flag must be set to "on". Summing up, it can be listed as the following conditions.
[0147] 1) sps_mts_enabled_flag = on 2) sps_explicit_mts_intra_enabled_flag = on 3) sps_explicit_mts_inter_enabled_flag = on
[0148] Here, when the decoding target block is encoded by the intra-prediction method, check the condition that sps_explicit_mts_intra_enabled_flag = "on", and when the decoding target block is encoded by the inter-prediction method, check the condition that sps_explicit_mts_inter_enabled_flag = "on".
[0149] The additional conditions for the use of explicit MTS are as follows.
[0150] 4) lfnst_idx is equal to 0 (implicit MTS reference) 5) transform_skip_flag is equal to 0 6) The intra_subpartitions_mode_flag is equal to 0 7) The cu_sbt_flag is equal to 0 (implicit MTS reference) 8) Valid MTS area 9) The width and height of the target block are 32 or less
[0151] Here, the lfnst_idx value indicates the secondary transform. When lfnst_idx = 0, it means that the secondary transform is not used.
[0152] The transform_skip_flag value indicates to omit the conversion process. When transform_skip_flag = 0, it indicates that the conversion proceeds normally without omitting the conversion process. The intra_subpartitions_mode_flag value, as one of the in-picture prediction methods, indicates that the target block is divided into multiple sub-blocks and undergoes prediction, conversion, and quantization processes. That is, when the flag value (intra_subpartitions_mode_flag) is "0", it means that in-picture prediction is performed without dividing the target block into sub-blocks. On the other hand, the use of MTS may be restricted depending on the size supported by additional transforms (DST-7 and DCT-8) (supporting up to 32×32 as described above). That is, if the width and height of the target block are not 32 or less, MTS cannot be used. That is, if either the width or height exceeds 32, DCT-II, which is the default transform (*) (where MTS cannot be used), is performed.
[0153] The cu_sbt_flag indicates whether it is a sub-block transform (sbt) in which only a part of the target block undergoes the conversion process. That is, when the cu_sbt_flag value is "0", it means that it is not a sub-block transform in which only a part of the target block undergoes the conversion process.
[0154] Hereinafter, the effective area (hereinafter referred to as the effective MTS area) will be described in detail.
[0155] FIG. 11 is a diagram showing the effective MTS area marked in bold within a 32×32 block to be decoded.
[0156] Referring to FIG. 11, the upper left 16×16 area excluding the DC coefficient becomes the effective MTS area. That is, the upper left 16×16 area excluding the 1×1 (DC) area is the effective MTS area. For example, if the positions of all non-zero coefficients within the target block belong to the effective MTS area, MTS is applicable. If one or more non-zero coefficient values are outside the effective MTS area, MTS application is not possible, and the default transform (*), DCT-II, is performed. This is the same concept as the zero-out area described above. That is, when a 32×32 target block uses MTS (i.e., DST-7 or DCT-8), the upper left 16×16 becomes the effective area, and the remaining part becomes the zero-out area. Similarly, if all non-zero coefficients within the 32×32 target block are located in the upper left 16×16 area, MTS (i.e., DST-7 or DCT-8) can be applied. However, exceptionally, if there is only one non-zero coefficient within the block and its position is DC (1x1), MTS application is not possible, and the default transform (*), DCT-II, is performed.
[0157] As a result, in this embodiment, in order to determine whether MTS is applicable, the effective MTS area must be checked. To check the effective MTS area, the following two conditions must be checked.
[0158] (a) When there is only one non-zero coefficient within the block, whether its position is DC (1x1) (b) Whether all non-zero coefficients within the block are located in the upper left 16×16 area
[0159] To confirm the condition (a), the information of the last position may be utilized. Here, the last position means the position of the last non-zero coefficient, i.e., the position of the last valid coefficient, in the scan order within the target block. As an example, the information of the last position, i.e., the last sub-block including the last non-zero coefficient, may be utilized. For example, if the position of the last sub-block is not (0, 0), the condition (a) can be satisfied. In other words, if the position of the last sub-block is not "0" (greater than 0) in the scan order of the sub-blocks within the target block, the condition (a) can be satisfied. Or, if the position of the last sub-block is "0", the information of the last scan position indicating the relative position of the last position within the sub-block may be utilized. For example, if the last scan position is not "0" (greater than 0) in the scan order of the coefficients within the sub-block, the condition (a) can be satisfied (refer to FIG. 12). Also, as described above, the MTS of the present invention is applied to the luma block.
[0160] To verify the condition (b), sub-block information including one or more non-zero coefficients may be utilized. Here, the sub-block information including one or more non-zero coefficients can be verified by the sb_coded_flag value of the sub-block. When the flag value is "1" (sb_coded_flag = 1), it means that one or more non-zero coefficients are located within the sub-block. When sb_coded_flag = 0, it means that all the coefficients within the sub-block are zero. That is, if the positions of all sub-blocks with sb_coded_flag value "1" within the target block are within (0, 0) to (3, 3), the condition (b) can be satisfied. On the contrary, if even one of the sub-blocks with sb_coded_flag value "1" within the target block is outside the position from (0, 0) to (3, 3), the condition (b) cannot be satisfied. In other words, if even one of the sub-blocks with sb_coded_flag value "1" within the target block has a value greater than 3 for either the x-coordinate or y-coordinate of the sub-block, the condition (b) cannot be satisfied (refer to Fig. 18). In other embodiments, in the scanning order of the sub-blocks within the target block, if the first sub-block with sb_coded_flag value "1" and having a value greater than 3 for either the x-coordinate or y-coordinate of the sub-block is found, the condition (b) may be set to false, and the confirmation process for subsequent sub-blocks with sb_coded_flag value "1" in the scanning order may be omitted (refer to Fig. 13). Also, as described above, the MTS of the present invention is applied to the luma block.
[0161] Fig. 12 is a diagram showing a method for determining a valid MTS according to an embodiment of the present invention. The embodiment of Fig. 12 relates to a method for verifying the condition (a) among the two conditions for verifying the valid MTS region described above. Each step in Fig. 12 may be performed within a video decoding device.
[0162] Referring to FIG. 12, the video decoding apparatus sets MtsDcOnlyFlag to "1" (S1210). The MtsDcOnlyFlag can indicate whether there is one non-zero coefficient in the block and its position is the DC. For example, when there is one non-zero coefficient in the block and its position is the DC, the MtsDcOnlyFlag has a value of "1", and in other cases, the MtsDcOnlyFlag has a value of "0". At this time, when the value of the MtsDcOnlyFlag is "0", the MTS may be applied. The reason for setting the MtsDcOnlyFlag to "1" in step S1210 is that when the block satisfies the condition of not being at the DC position when there is one non-zero coefficient in the following block, the MtsDcOnlyFlag is reset to "0", and otherwise, the MTS is not applied.
[0163] Next, the video decoding apparatus determines whether the target block is a luma block (S1220). The purpose of determining whether the target block is a luma block is that, as described above, the MTS is only applied to luma blocks.
[0164] Next, the video decoding apparatus determines whether the last sub-block is greater than 0 (S1230). If the last sub-block is greater than 0, the MtsDcOnlyFlag is set to "0" (S1240), and the process ends.
[0165] As a result of the determination in step S1230, if the last sub-block is not greater than 0, it is determined whether the last scan position is greater than 0 (S1250).
[0166] As a result of the determination in step S1250, if the last scan position is greater than 0, the MtsDcOnlyFlag is set to "0" (S1240), and the process ends.
[0167] As a result of the determination in step S1250, if the last scan position is not greater than 0, the process ends.
[0168] According to this embodiment, if the last sub-block is greater than 0 or the last scan position is greater than 0, set MtsDcOnlyFlag to "0"; otherwise, set MtsDcOnlyFlag to "1".
[0169] Thereafter, when determining whether to apply MTS, if MtsDcOnlyFlag is checked and has a value of "1", DCT-II, which is the default conversion, may be applied without applying MTS.
[0170] FIG. 13 is a diagram showing a method for determining a valid MTS region according to another embodiment of the present invention. The embodiment of FIG. 13 specifically shows a method for checking condition (b) among the two conditions for checking the valid MTS region described above. Each step in FIG. 13 may be performed in a video decoding apparatus.
[0171] Referring to FIG. 13, the video decoding apparatus sets MtsZerooutFlag to "1" (S1305). The MtsZerooutFlag indicates whether non-zero coefficients in the block exist in the zero-out region. For example, when at least one of the non-zero coefficients in the block exists in the zero-out region, MtsZerooutFlag may have a value of "0", and when all non-zero coefficients in the block do not exist in the zero-out region, MtsZerooutFlag may have a value of "1". In this embodiment, assuming that all non-zero coefficients in the block do not exist in the zero-out region, the initial value of MtsZerooutFlag is set to "1", and when the conditions of the zero-out region and the non-zero coefficients are satisfied simultaneously, MtsZerooutFlag may be set to "0". At this time, if there is a MtsZerooutFlag having a value of "0", it may not be necessary to apply explicit MTS.
[0172] Next, the video decoding device sets the initial value of variable i to the value of the last sub-block, and subtracts 1 from the value of variable i one by one, and repeats the processes of the following steps S1325 to S1350 until the value of variable i becomes 0 (S1320). The purpose of repeatedly performing the routine of step S1820 is to check the sb_coded_flag values of all sub-blocks from the last sub-block to the first sub-block. As described above, when the flag value is "1", it means that there is one or more non-zero coefficients in the sub-block, and when the flag value is "0", it means that there are no non-zero coefficients in the sub-block. Therefore, referring to FIG. 11, if the positions of all sub-blocks in the target block with sb_coded_flag value of "1" only exist within (0, 0) to (3, 3), that is, only exist within 0 to 8 based on variable i, it may be determined that the condition (b) for applying explicit MTS is satisfied.
[0173] Next, the video decoding device determines whether variable i simultaneously satisfies the conditions that variable i is smaller than the last sub-block (i < last sub-block) and variable i is larger than 0 (i > 0) (S1325). For example, when the routine of step S1320 is executed for the first time, since the initial value of variable i is set to the same value as the last sub-block, the condition of step S1325 is not satisfied.
[0174] As a result of the determination in step S1325, when variable i simultaneously satisfies the conditions that variable i is smaller than the last sub-block (i < last sub-block) and variable i is larger than 0 (i > 0), parse sb_coded_flag (S1830). When the two conditions are not simultaneously satisfied, set sb_coded_flag to "1" (S1835).
[0175] At this time, the parsed sb_coded_flag indicates whether there is one or more non-zero coefficients in the sub-block. If there is one or more non-zero coefficients in the sub-block, the sb_coded_flag has a value of "1", and if there is no non-zero coefficient in the sub-block, the sb_coded_flag has a value of "0".
[0176] On the other hand, step S1835 is performed only when i indicates the last sub-block and the first sub-block. That is, since the last position coefficient is included in the last sub-block, the sb_coded_flag value is parsed as the value "1", and since the DC coefficient exists in the first sub-block, the sb_coded_flag value is parsed as the value "1".
[0177] Next, the video decoding device determines whether the block is a luma block (S1340). The purpose of determining whether the target block is a luma block is that, as described above, MTS is only applied to luma blocks.
[0178] As a result of the determination in step S1340, if the block is a luma block, it is determined whether the condition "sb_coded_flag && (xSb>3||ySb>3)" is satisfied (S1845). If the condition in step S1845 is satisfied, MtsZerooutFlag is set to "0" (S1350).
[0179] According to this embodiment, if one or more non-zero coefficients are found in a sub-block other than the sub-block (3, 3) in the target block, that is, in the zero-out region, MtsZerooutFlag may be set to "0" and it may be determined that explicit MTS cannot be applied.
[0180] FIG. 14 is a diagram showing a method for determining an effective MTS according to another embodiment of the present invention. The embodiment of FIG. 14 specifically shows a method for checking the condition (b) among the two conditions for checking the effective MTS region described above. However, in the embodiment of FIG. 13, the effective MTS region is checked by checking the sb_coded_flag of all sub-blocks, while in the embodiment of FIG. 14, when the first invalid MTS is found, it is different that the subsequent sb_coded_flag does not need to be checked. Each step of FIG. 14 may be performed in a video decoding device.
[0181] Referring to FIG. 14, the video decoding device sets the MtsZerooutFlag to "1" (S1405). The MtsZerooutFlag indicates whether non-zero coefficients in the block exist in the zero-out region. For example, when at least one of the non-zero coefficients in the block exists in the zero-out region, the MtsZerooutFlag may have a value of "0", and when all non-zero coefficients in the block do not exist in the zero-out region, the MtsZerooutFlag may have a value of "1". In this embodiment, assuming that all non-zero coefficients in the block do not exist in the zero-out region, the initial value of the MtsZerooutFlag is set to "1", and when the conditions of the zero-out region and the non-zero coefficients are simultaneously satisfied, the MtsZerooutFlag may be set to "0". At this time, if there is a MtsZerooutFlag having a value of "0", it may not be necessary to apply the explicit MTS.
[0182] Next, the video decoding apparatus sets the initial value of variable i to the value of the last sub-block, subtracts 1 from the value of variable i one by one, and repeats the processes of the following step S1425 to step S1450 until the value of variable i becomes 0 (S1420). The purpose of repeatedly performing the routine of step S1420 is to check the sb_coded_flag values of all sub-blocks from the last sub-block to the first sub-block. As described above, when the sb_coded_flag value is "1", it means that there is one or more non-zero coefficients in the sub-block, and when the sb_coded_flag value is "0", it means that there are no non-zero coefficients in the sub-block. Therefore, referring to FIG. 16, if the positions of all sub-blocks in the target block with the sb_coded_flag value of "1" only exist within (0, 0) to (3, 3), that is, only within 0 to 8 based on variable i, it may be determined that the condition (b) for applying explicit MTS is satisfied.
[0183] Next, the video decoding apparatus determines whether the variable i simultaneously satisfies the conditions that the variable i is smaller than the last sub-block (i < the last sub-block) and the variable i is larger than 0 (i > 0) (S1425). For example, when the routine of step S1920 is executed for the first time, since the initial value of variable i is set to the same value as the last sub-block, the condition of step S1425 is not satisfied.
[0184] As a result of the determination in step S1425, when the variable i simultaneously satisfies the conditions that the variable i is smaller than the last sub-block (i < the last sub-block) and the variable i is larger than 0 (i > 0), parse sb_coded_flag (S1430), and when the two conditions are not simultaneously satisfied, set sb_coded_flag to "1" (S1435).
[0185] At this time, the sb_coded_flag to be parsed indicates whether there is one or more non-zero coefficients in the sub-block. If there is one or more non-zero coefficients in the sub-block, the sb_coded_flag has a value of "1". If there is no non-zero coefficient in the sub-block, the sb_coded_flag has a value of "0".
[0186] On the other hand, step S1435 is performed only when i indicates the last sub-block and the first sub-block. That is, since the last position coefficient is included in the last sub-block, the sb_coded_flag value is parsed as the value of "1". Since the DC coefficient exists in the first sub-block, the sb_coded_flag value is parsed as the value of "1".
[0187] Next, the video decoding device determines whether the condition of "MtsZerooutFlag && luma block" is satisfied (S1440).
[0188] As a result of the determination in step S1440, if the condition of "MtsZerooutFlag && luma block" is satisfied, it is further determined whether the condition of "sb_coded_flag && (xSb>3 ||ySb>3)" is satisfied (S1445). If the condition of "sb_coded_flag && (xSb>3||ySb>3)" is satisfied, MtsZerooutFlag is set to "0" (S1450).
[0189] As a result of the determination in step S1440, if the condition of "MtsZerooutFlag && luma block" is not satisfied, the process in the sub-block is terminated.
[0190] According to this embodiment, when the variable i, that is, the MtsZerooutFlag value is set to "0" even once in the sub-block, in the next routine of the variable i-1, a false value is derived in step S1940, and there is no need to further check the sb_coded_flag value.
[0191] On the other hand, when the block to be decrypted satisfies all the conditions of (a) and (b) above, the use of explicit MTS is determined, and the conversion information actually used for the block is transmitted in index form (mts_idx). In contrast, when not all the conditions are satisfied, DCT-II, which is the default conversion (*), is used (see Figure 15). Table 5 shows the conversion types of the horizontal axis and the vertical axis according to the mts_idx value.
[0192] [Table 5]
[0193] In Table 5, trTypeHor means the horizontal axis conversion type, and trTypeVer means the vertical axis conversion type. The values of the conversion types in Table 5 mean the trType values in Table 2. For example, when the mts_idx value is "2", DCT-8(2) may be used for the horizontal axis conversion and DST-7(1) may be used for the vertical axis conversion.
[0194] In the present invention, in all cases where the above-mentioned default conversion (*) DCT-II is used / executed / applied, it may be replaced with the expression "induce the mts_idx value to be "0"". That is, when the mts_idx value is "0", DCT-II(0) is set for both the horizontal axis and the vertical axis conversions.
[0195] In the present invention, the binary method of mts_idx uses the TR (truncated rice) method, and the cMax value, which is the parameter value for TR, is "4", and the cRiceParam value is "0". Table 6 shows the codewords of the MTS index.
[0196] [Table 6]
[0197] Referring to Table 6, when the mts_idx value is "0", the corresponding codeword is "0"; when the mts_idx value is "1", the corresponding codeword is "10"; when the mts_idx value is "2", the corresponding codeword is "110"; when the mts_idx value is "3", the corresponding codeword is "1110"; when the mts_idx value is "4", the corresponding codeword is "1111", which is confirmed.
[0198] FIG. 15 is a diagram showing a method for determining whether to apply an explicit MTS function according to an embodiment of the present invention. Each step in FIG. 15 may be performed in a video decoding apparatus.
[0199] Referring to FIG. 15, the video decoding apparatus determines whether the condition of "(sps_explicit_mts_intra_enabled_flag && CuPredMode = MODE_INTRA) || (sps_explicit_mts_inter_enabled_flag && CuPredMode = MODE_INTER)" is satisfied (S1510).
[0200] sps_explicit_mts_intra_enabled_flag is a flag indicating whether to use explicit MTS during intra prediction, and sps_explicit_mts_inter_enabled_flag is a flag indicating whether to use explicit MTS during inter prediction. sps_explicit_mts_intra_enabled_flag has a value of "1" when using explicit MTS during intra prediction, and a value of "0" otherwise. sps_explicit_mts_inter_enabled_flag has a value of "1" when using explicit MTS during inter prediction, and a value of "0" otherwise.
[0201] CuPredMode indicates whether the block to be decoded was encoded using which prediction method. When the block to be decoded was encoded using the intra prediction method, CuPredMode has the value MODE_INTRA, and when the block to be decoded was encoded using the inter prediction method, CuPredMode has the value MODE_INTER.
[0202] Therefore, when the block to be decoded uses intra prediction and explicit MTS, "sps_explicit_mts_intra_enabled_flag && CuPredMode = MODE_INTRA" has the value "1", and when the block to be decoded uses inter prediction and explicit MTS, "sps_explicit_mts_inter_enabled_flag && CuPredMode = MODE_INTER" has the value "1". Therefore, in step S2010, by checking the values of sps_explicit_mts_intra_enabled_flag, sps_explicit_mts_inter_enabled_flag, and CuPredMode, it is possible to determine whether the block to be decoded uses explicit MTS.
[0203] When the condition of step S1510 is satisfied, the video decoder determines whether the condition "lfnst_idx = 0 && transform_skip_flag = 0 && cbW < 32 && cbH < 32 && intra_subpartitions_mode_flag = 0 && cu_sbt_flag = 0" is satisfied (S1520).
[0204] Here, the lfnst_idx value indicates the secondary transform. When lfnst_idx = 0, it means that the secondary transform is not being used.
[0205] The transform_skip_flag value indicates whether transform skip is applied to the current block. That is, it indicates whether to omit the transformation process for the current block. When transform_skip_flag = 0, it indicates that transform skip is not applied to the current block.
[0206] cbW and cbH indicate the width and height of the current block respectively. As described above, the maximum size supported for DCT-II, which is the default transformation, is up to 64×64, and the maximum size supported for DST-7 and DCT-8, which are additional transformations, is up to 32×32. For example, when the size of the block to be decoded is 64×64, one 64×64 DCT-II is applied to the transformation process. That is, when one or more of the width and height of the block to be decoded are greater than 32 (exceed 32), the default transformation (*) is applied immediately without applying MTS. Therefore, for MTS to be applied, all the values of cbW and cbH must have values of 32 or less.
[0207] The intra_subpartitions_mode_flag indicates whether the intra subpartition mode is applied. The intra subpartition mode is one of the in-picture prediction methods, and it indicates that the target block is divided into a number of sub-blocks and undergoes prediction, transformation, and quantization processes. That is, when the flag value (intra_subpartitions_mode_flag) is "0", it means that general in-picture prediction is performed without dividing the target block into sub-blocks.
[0208] The cu_sbt_flag indicates whether sub-block transformation, where only a part of the target block undergoes the transformation process, is applied. That is, when the cu_sbt_flag value is "0", it means that sub-block transformation, where only a part of the target block undergoes the transformation process, is not applied.
[0209] Therefore, it may be determined whether the block to be decoded can apply explicit MTS according to whether the condition of step S1520 is satisfied.
[0210] When the condition of step S1510 is not satisfied, the video decoding device sets the value of mts_idx to "0" (S1530) and ends the process.
[0211] When the condition of step S1520 is satisfied, the video decoding device determines whether the condition of "MtsZeroOutFlag = 1 && MtsDcOnlyFlag = 0" is satisfied (S1540).
[0212] The MtsZerooutFlag indicates whether non-zero coefficients in the block exist in the zero-out region. When at least one of the non-zero coefficients in the block exists in the zero-out region, the MtsZerooutFlag may have a value of "0", and when all non-zero coefficients in the block do not exist in the zero-out region, the MtsZerooutFlag may have a value of "1". At this time, the value of the MtsZerooutFlag may be determined by performing the process of FIG. 13 or FIG. 14.
[0213] MtsDcOnlyFlag indicates whether the non-zero coefficient in the block is one and its position is DC. When the non-zero coefficient in the block is one and its position is DC, the MtsDcOnlyFlag has a value of "1", and in other cases, the MtsDcOnlyFlag has a value of "0". At this time, the value of the MtsDcOnlyFlag may be determined by performing the process of FIG. 17.
[0214] On the other hand, when the condition of step S1520 is not satisfied, the video decoding device sets the value of mts_idx to "0" (S1530) and ends the process.
[0215] When the condition of step S1540 is satisfied, the video decoding apparatus parses mts_idx (S1550) and ends the process. At this time, the conversion types for the horizontal and vertical axes according to the value of mts_idx may be assigned according to Table 5. At this time, the value of the conversion type in Table 5 means the trType value in Table 2. For example, when the mts_idx value is "2", DCT-8 may be applied for the horizontal axis conversion and DST-7 may be applied for the vertical axis conversion.
[0216] Also, even when the condition of step S1540 is not satisfied, the video decoding apparatus sets the value of mts_idx to "0" (S1530) and ends the process.
[0217] FIG. 16 is a diagram showing a method of performing inverse conversion based on conversion-related parameters according to another embodiment of the present invention. Each step in FIG. 16 may be performed by a video decoding apparatus, for example, by an inverse conversion unit of the decoding apparatus.
[0218] Referring to FIG. 16, the video decoding apparatus acquires the values of sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, and cu_sbt_flag (S1610). At this time, what each of sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, and cu_sbt_flag indicates is described in detail in the related description of FIG. 15, and the parameters are used to determine whether the block to be decoded can apply explicit MTS.
[0219] Next, the video decoding device acquires the MtsZerooutFlag and MtsDcOnlyFlag values (S1620). At this time, the MtsZerooutFlag can be acquired by performing the process of FIG. 13 or FIG. 14, and the MtsDcOnlyFlag can be acquired by performing the process of FIG. 12.
[0220] Next, the video decoding device acquires the mts_idx value based on the parameters acquired in step S1610 and step S1620 (S1630). That is, the video decoding device acquires the mts_idx value based on sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, cu_sbt_flag, MtsZerooutFlag, and MtsDcOnlyFlag. At this time, the mts_idx may be acquired by performing the process of FIG. 15.
[0221] Next, the video decoding device performs inverse transformation based on mts_idx (S1640). The inverse transformation applied by the mts_idx value may be configured according to Table 5 and Table 2. For example, when the mts_idx value is "2", DCT-8 may be applied in the horizontal axis direction and DST-7 may be applied in the vertical axis direction.
[0222] On the other hand, although not shown in the drawings, from the perspective of a video encoding device, in order to set whether to use explicit MTS, sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, cu_sbt_flag, MtsZerooutFlag, and MtsDcOnlyFlag can be set. In the above-described embodiments, the method has been described based on a sequence diagram as a series of steps or blocks, but the present invention is not limited to the order of the steps, and a certain step may occur in a different order from or simultaneously with the steps different from the above. Also, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more steps of the sequence diagram may be deleted without affecting the scope of the present invention.
[0223] The embodiments described in this document may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., instruction information) or an algorithm may be stored in a digital storage medium.
[0224] In addition, the decoding device and encoding device to which the present invention is applied may be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video intercom devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, picture phone video devices, transportation means terminals (e.g., vehicle terminals, airplane terminals, ship terminals, etc.), and medical video devices, etc., and may be used to process video signals or data signals. For example, as over-the-top (OTT) video devices, game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc. may be mentioned.
[0225] In addition, the processing method to which the present invention is applied may be produced in the form of a program executed by a computer, or may be stored in a computer-readable storage medium. Multimedia data having a data structure according to the present invention may also be stored in a computer-readable storage medium. The computer-readable storage medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored and memorized. The above computer-readable storage medium may include, for example, Blu-ray discs (BDs), universal serial bus (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. Further, the computer-readable storage medium includes a medium realized in the form of a carrier wave (e.g., transmission via the Internet). Also, the bitstream generated by the encoding method may be stored in a computer-readable storage medium or transmitted via a wired or wireless communication network.
[0226] Further, an embodiment of the present invention may be realized by a computer program product with program code, and the program code may be executed by a computer according to an embodiment of the present invention. The program code may be stored on a computer-readable carrier.
Claims
1. A video decoding method performed by a video decoding device, comprising: obtaining a parameter related to whether a Multiple Transform Set (MTS) is applicable to a block to be decoded from a bitstream; determining a type of transform to be applied to the block to be decoded based on at least one of a parameter relating to whether a set of multiple transforms is applicable to the block to be decoded and a size of the block to be decoded; setting a validity region including non-zero coefficients of the block to be decoded based on at least one of a parameter relating to whether a multiple transformation set for the block to be decoded is applicable or not and a size of the block to be decoded; and Reconstructing the block to be decoded based on a transform type applied to the block to be decoded and a valid region including non-zero coefficients of the block to be decoded; When the size of the block to be decoded is any one of 64 width by 64 height, 64 width by 32 height, and 32 width by 64 height, a valid area including non-zero coefficients of the block to be decoded is set to a 32 width by 32 height area including an upper left sample in the block to be decoded, regardless of a value of a parameter related to whether a multiple transform set is applicable to the block to be decoded; When the size of the block to be decoded is any one of 32 width × 32 height, 32 width × 16 height, and 16 width × 32 height, a valid area including a non-zero coefficient of the block to be decoded is set to an area having a different size according to a value of the parameter related to whether a multiple transformation set for the block to be decoded is applicable or not; The image decoding method according to claim 1, wherein the transform type is one of a first transform type and a second transform type, the first transform type including a type indicating a same transform kernel for a vertical direction and a horizontal direction, and the second transform type including a type indicating different transform kernels for a vertical direction and a horizontal direction.
2. A video encoding method performed by a video encoding device, comprising: determining a type of transform to be applied to the block to be coded according to whether a Multiple Transform Set (MTS) for the block to be coded is applicable; generating transform coefficients by performing a transform corresponding to a transform type on the current block to be coded based on a valid area including non-zero coefficients for the current block to be coded, the valid area being determined taking into account a transform type to be applied to the current block to be coded; generating a bitstream including at least one of a parameter related to whether a multiple transform set is applicable to the current block, a parameter indicating a transform type to be applied to the current block, and the transform coefficients; when the size of the encoding target block is any one of width 64×height 64, width 64×height 32, and width 32×height 64, a valid area including nonzero coefficients of the encoding target block is a width 32×height 32 area including an upper left sample in the encoding target block, When the size of the block to be coded is any one of 32 width × 32 height, 32 width × 16 height, and 16 width × 32 height, a valid area including non-zero coefficients of the block to be coded is determined by a value of a parameter related to whether a multiple transformation set is applicable or not; The video encoding method, wherein the transform type is either a first transform type or a second transform type, the first transform type including a type indicating a same transform kernel for a vertical direction and a horizontal direction, and the second transform type including a type indicating different transform kernels for a vertical direction and a horizontal direction.
3. A method for transmitting a bitstream encoded by the video encoding method of claim 2, comprising the step of transmitting the bitstream.
Citation Information
Patent Citations
Method for encoding / decoding video signals and apparatus therefor
WO2020060364A1
Transform-based image coding method, and device therefor
WO2021066616A1
Methods and apparatus on transform and coefficient signaling
WO2021102424A1
Coding concepts for a transformed representation of a sample block
WO2021105255A1