Video decoding method and device
The method optimizes video decoding by determining conversion types for blocks using MTS parameters, enhancing compression performance and reducing costs in high-resolution video data transmission and storage.
Patent Information
- Application Number
- JP2025046891
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-11-25
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-12-16
AI Technical Summary
The increasing demand for high-resolution and high-quality videos has led to a surge in video data, resulting in higher transmission and storage costs, and the need for improved video coding technologies like VVC, which requires determining the appropriate conversion type for blocks during decoding.
A method for video decoding that involves obtaining parameters indicating the applicability of Multiple Transform Sets (MTS) for blocks, determining the conversion type based on block size and other factors, and performing inverse conversion accordingly, optimizing the decoding process.
This approach allows for efficient decoding by applying optimized conversion types, improving compression performance and reducing costs associated with high-resolution video data transmission and storage.
Smart Images

Figure 2025094156000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video coding technology, and more particularly, to a method for determining the type of primary transform of a block to be decoded during the video decoding process.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video data. Therefore, when storing video data using a medium such as an existing wired or wireless broadband line, the transmission cost and storage cost increase.
[0003] Since the establishment of the HEVC (High Efficiency Video Coding) video codec in 2013, with the spread of immersive videos and virtual reality services using 4K and 8K video images, work has begun on the standardization of VVC (Versatile Video Coding), a next-generation video codec aiming for more than twice the performance improvement of HEVC. Currently, the standardization work is in full swing. VVC is being developed by the JVET (Joint Video Exploration Team), jointly formed by the ISO / IEC MPEG (Moving Picture Experts Group), a video coding standardization group, and the ITU-T VCEG (Video Coding Experts Group), with the goal of improving the coding compression performance by more than twice that of HEVC. VVC standardization began in earnest in January 2018, when a Call for proposal was announced at the 121st Gwangju MPEG and 9th JVET meetings, and at the 122nd San Diego MPEG and 10th JVET meetings, a total of 23 organizations proposed video codec technologies, starting the formal video standardization. At the 122nd MPEG and 10th JVET meetings, technical reviews, objective compression performance, and subjective picture quality evaluations were conducted on the video codec technologies proposed by each organization, and some of the many technologies were adopted to release Working Draft (WD) 1.0 and VTM (VVC Test Mode) 1.0, the video reference software. After the conclusion of the 127th MPEG and 15th JVET meetings in July 2019, the Committee Draft (CD) of the VVC standard was completed, and standardization is progressing towards the goal of formulating the Final Draft International Standard (FDIS) in October 2020.
[0004] Conventionally, in the encoding structure hierarchically divided by a quadtree in HEVC, VVC adopted a split block structure combining QTBT (QuadTree Binary Tree) and TT (Ternary Tree). This enables more flexible generation or processing of prediction residual signals compared to HEVC, resulting in further improved compression performance compared to HEVC. In addition to such a basic block structure, new technologies not used in existing codecs, such as Adaptive Loop Filter (ALF) technology, Affine Motion Prediction (AMP) technology as a motion prediction technology, and Decoder-side Motion Vector Refinement (DMVR) technology, were adopted as standard technologies. As for the conversion and quantization technology, DCT-II, a conversion kernel widely used in existing video codecs, continues to be used, and it has been changed to be applicable to even larger block sizes of the applicable blocks. Also, in existing HEVC, the DST-7 kernel, which has been applied to small conversion blocks such as 4×4, has been extended to larger conversion blocks, and DCT-8, a new conversion kernel, has been added as a conversion kernel.
[0005] On the other hand, in the HEVC standard, when encoding or decoding video, since conversion is performed using one type of conversion, there was no need to transmit information regarding the conversion type for the video. However, in the new technology, Multiple Transform Selection using DCT-II, DCT-8, and DCT-7 can be applied. Therefore, in reality, a technique for defining whether to apply MTS and which primary conversion type is applied during decoding is required.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The technical problem of the present invention is to perform inverse conversion in a method pre-justified under specific conditions.
[0007] Another technical problem of the present invention is to perform decoding by applying a conversion type optimized for a block to be decoded.
Means for Solving the Problem
[0008] According to one aspect of the present invention, there is provided a video decoding method performed by a video decoder. The video decoding method includes steps of: obtaining information regarding a parameter indicating whether a Multiple Transform Set (MTS) for a block to be decoded is applicable, the width of the block to be decoded, and the height of the block to be decoded; determining a conversion type of the block to be decoded based on at least one of the parameter indicating whether the Multiple Transform Set for the block to be decoded is applicable, the width of the block to be decoded, and the height of the block to be decoded; setting a zero-out region of the block to be decoded based on at least one of the parameter indicating whether the Multiple Transform Set for the block to be decoded is applicable, the width of the block to be decoded, and the height of the block to be decoded; and performing inverse conversion on the block to be decoded based on a result of determination regarding the zero-out region and the conversion type of the block to be decoded.
[0009] According to another aspect of the present invention, in the step of determining the conversion type of the block to be decoded, when at least one of the width or height of the block to be decoded has a value greater than 32, it is determined that the block to be decoded is converted using a default conversion.
[0010] According to still another aspect of the present invention, in the step of setting the zero-out region of the block to be decoded, when one of the width or height of the block to be decoded has a value greater than 32, a region where the width or height of the block to be decoded is greater than 32 is set as the zero-out region.
[0011] According to still another aspect of the present invention, the parameter indicating whether a multiple transform set is applicable to the block to be decoded is sps_mts_enabled_flag.
[0012] According to still another aspect of the present invention, there is provided a video decoding method performed by a video decoding apparatus. The video decoding method includes steps of obtaining at least one of information regarding applicability of a multiple transform set to a block to be decoded, information regarding a prediction mode, information regarding applicability of secondary transformation, information regarding applicability of prediction using a matrix, and information regarding the size of the block to be decoded; determining whether an implicit multiple transform set is applicable to the block to be decoded based on at least one of information regarding applicability of a multiple transform set to the block to be decoded, information regarding a prediction mode, information regarding applicability of secondary transformation, and information regarding applicability of prediction using a matrix; obtaining information regarding a transform type based on information regarding whether the implicit multiple transform set is applicable to the block to be decoded and information regarding the size of the block to be decoded; and performing an inverse transform based on the information regarding the transform type.
[0013] According to still another aspect of the present invention, the step of determining whether the implicit multiple transform set is applicable includes determining whether the implicit multiple transform set is applicable by using information regarding applicability of a multiple transform set to the block to be decoded, information regarding a prediction mode, information regarding applicability of secondary transformation, and information regarding applicability of prediction using a matrix.
[0014] According to another aspect of the present invention, the implicit multiple transformation set includes one default transform and at least one extra transform.
[0015] According to another aspect of the present invention, based on the information regarding the size of the block to be decoded, the step of obtaining information regarding the transform type is such that when the horizontal axis length of the block to be decoded is all 4 or more and 16 or less, at least one of the extra transform types is applied to the block to be decoded in the horizontal axis direction.
[0016] According to another aspect of the present invention, based on the information regarding the size of the block to be decoded, the step of obtaining information regarding the transform type is such that when the vertical axis length of the block to be decoded is all 4 or more and 16 or less, at least one of the extra transform types is applied to the block to be decoded in the vertical axis direction.
[0017] According to another aspect of the present invention, the information regarding the application of the multiple transformation set of the block to be decoded includes at least one of sps_mts_enabled_flag and sps_explicit_mts_intra_enabled_flag.
[0018] According to another aspect of the present invention, the information regarding the prediction mode includes CuPredMode.
[0019] According to another aspect of the present invention, the information regarding the application of the secondary transformation includes lfnst_idx.
[0020] According to another aspect of the present invention, the information regarding the application of the prediction using the matrix includes intra_mip_flag.
[0021] According to another aspect of the present invention, the information regarding the conversion type of the block to be decoded includes information regarding the horizontal-axis conversion type and information regarding the vertical-axis conversion type, respectively.
[0022] According to another aspect of the present invention, the step of determining whether an implicit multiple conversion set is applied to the block to be decoded is obtained by further confirming whether the block to be decoded is a luma block.
[0023] According to another aspect of the present invention, a video decoding apparatus including a memory and at least one processor is provided. The video decoding apparatus obtains at least one of information regarding the application of a multiple conversion set to a block to be decoded, information regarding a prediction mode, information regarding the application of secondary conversion, information regarding the application of prediction using a matrix, and information regarding the size of the block to be decoded, determines whether an implicit multiple conversion set is applied to the block to be decoded based on at least one of the information regarding the application of a multiple conversion set to the block to be decoded, information regarding a prediction mode, information regarding the application of secondary conversion, and information regarding the application of prediction using a matrix, obtains information regarding the conversion type based on the information regarding whether an implicit multiple conversion set is applied to the block to be decoded and the information regarding the size of the block to be decoded, and includes at least one processor including an inverse conversion unit that performs inverse conversion based on the information regarding the conversion type.
Advantages of the Invention
[0024] According to the present invention, inverse conversion can be performed in a method pre-defined according to specific conditions.
[0025] In addition, by performing decoding by applying an optimized conversion type to the block to be decoded, an improvement effect in compression performance can be expected.
Brief Description of the Drawings
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Embodiments for Carrying Out the Invention
[0027] The present invention may be subject to various modifications and may have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail below. However, this is not intended to limit the present invention to specific embodiments. The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the technical idea of the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0028] On the other hand, each configuration in the drawings described in the present invention is shown independently for the convenience of explaining separate characteristic functions, and does not mean that each configuration is realized by separate hardware or separate software. For example, two or more of the configurations may be combined to form one configuration, or one configuration may be divided into a plurality of configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0029] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components will be omitted.
[0030] On the one hand, the present invention relates to video / video coding. For example, the methods / embodiments disclosed in the present invention can be applied to the methods disclosed in the VVC (versatile video coding) standard, the EVC (Essential Video Coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or the next-generation video / image coding standard (e.g., H.267, H.268, etc.).
[0031] In this specification, a "picture" generally means a unit indicating one video in a specific time period, and a "slice" is a unit constituting a part of a picture in coding. One picture may be composed of a plurality of slices, and if necessary, pictures and slices may be used interchangeably.
[0032] A "pixel" or "pel" may mean the smallest unit constituting one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value, or only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0033] A "unit" indicates a basic unit of video processing. The unit may include at least one of a specific region of a picture and information related to the region. The unit may, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block may indicate a set of samples or transform coefficients consisting of M columns and N rows.
[0034] FIG. 1 is a diagram schematically showing the configuration of a video encoding apparatus to which the present invention is applied.
[0035] Referring to FIG. 1, the video encoding apparatus 100 may include a picture splitting unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a conversion unit 122, a quantization unit 123, a reordering unit 124, an inverse quantization unit 125, and an inverse transformation unit 126.
[0036] The picture splitting unit 105 splits the input picture into at least one processing unit.
[0037] As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a Coding Tree Unit (CTU) by a Quad-tree binary-tree (QTBT) structure. For example, one coding unit may be split into a plurality of nodes at a deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure may be applied later. Alternatively, the binary-tree structure may be applied first. Based on the final coding unit that is no longer split, the coding procedure according to the present invention may be performed. In this case, based on the coding efficiency and the like according to the video characteristics, the largest coding unit may be immediately used as the final coding unit, or if necessary, the coding unit may be recursively split into coding units at a deeper depth, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, conversion, and restoration described later.
[0038] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit may be split from the largest coding unit (LCU) into coding units of lower depths according to a quad-tree structure. In this case, based on coding efficiency according to video characteristics, etc., the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of even lower depths so that a coding unit of an optimal size is used as the final coding unit. When the smallest coding unit (SCU) is set, the coding unit is not split into coding units smaller than the smallest coding unit. Here, the final coding unit means a coding unit that serves as a basis for partitioning or splitting into a prediction unit or a transform unit. The prediction unit is a unit partitioned from the coding unit and may be a unit for sample prediction. At this time, the prediction unit may be divided into subblocks. The transform unit may be split from the coding unit according to a quad-tree structure and may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transform unit may be referred to as a transform block (TB). The prediction block or prediction unit means a specific region in block form within a picture and may include an array of prediction samples.Further, the transform block or transform unit means a specific area in block form within a picture, and may include an array of transform coefficients or residual samples.
[0039] The prediction unit 110 performs prediction on a block to be processed (hereinafter referred to as the current block), and generates a predicted block including prediction samples for the current block. The unit of prediction performed by the prediction unit 110 may be a coding block, a transform block, or a prediction block.
[0040] The prediction unit 110 determines whether intra prediction or inter prediction is applied to the current block. As an example, the prediction unit 110 determines whether intra prediction or inter prediction is applied in CU units.
[0041] In the case of intra prediction, the prediction unit 110 can derive prediction samples for the current block based on reference samples outside the current block within the picture (hereinafter referred to as the current picture) to which the current block belongs. At this time, the prediction unit 110 may derive prediction samples based on (i) the average or interpolation of neighboring reference samples of the current block, or (ii) reference samples existing in a specific (prediction) direction with respect to the prediction samples among the neighboring reference samples of the current block. In the case of (i), it is called a non - directional mode or a non - angular mode, and in the case of (ii), it is called a directional mode or an angular mode. The prediction modes in intra prediction may have, for example, 33 directional prediction modes and at least 2 non - directional modes. The non - directional modes may include a DC prediction mode and a Planar mode. The prediction unit 110 may determine the prediction mode applied to the current block using the prediction mode applied to the adjacent block.
[0042] In the case of inter prediction, the prediction unit 110 can derive a prediction sample for the current block based on samples specified by motion vectors on the reference picture. The prediction unit 110 can derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP (motion vector prediction) mode. In the case of the skip mode and the merge mode, the prediction unit 110 may use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, the motion vector of an adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be derived by using it as the motion vector predictor of the current block.
[0043] In the case of inter prediction, the adjacent blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in a reference picture. The reference picture including the temporal neighboring blocks may be referred to as a collocated picture (colPic). The motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and output in the form of a bit stream.
[0044] When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top picture on the reference picture list may be used as a reference picture. The reference pictures included in the reference picture list may be sorted based on the difference in POC (Picture Order Count) between the current picture and the reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0045] The subtraction unit 121 generates a residual sample, which is the difference between the original sample and the predicted sample. When the skip mode is applied, it may not be necessary to generate the residual sample as described above.
[0046] The conversion unit 122 converts the residual samples in units of conversion blocks to generate transform coefficients. The conversion unit 122 can perform conversion according to the size of the conversion block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the conversion block. For example, if intra prediction is applied to the coding block or the prediction block that overlaps with the conversion block, and the conversion block is a 4×4 residual array, the residual samples are converted using a DST (Discrete Sine Transform) conversion kernel, and in other cases, the residual samples are converted using a DCT (Discrete Cosine Transform) conversion kernel.
[0047] The quantization unit 123 quantizes the transform coefficients to generate quantized transform coefficients.
[0048] The reordering unit 124 reorders the quantized transform coefficients. The reordering unit 124 can reorder the quantized transform coefficients in block form into a one-dimensional vector form by a coefficient scanning method. Here, although the reordering unit 124 is described separately, it may be a part of the quantization unit 123.
[0049] The entropy encoding unit 130 performs entropy encoding on the quantized transform coefficients. The entropy encoding may include encoding methods such as, for example, exponential Golomb, CAVLC (context - adaptive variable length coding), CABAC (context - adaptive binary arithmetic coding), etc. The entropy encoding unit 130 may encode, together or separately, information necessary for video restoration (for example, values of syntax elements, etc.) in addition to the quantized transform coefficients. The entropy - encoded information may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units.
[0050] The inverse quantization unit 125 inverse - quantizes the values (quantized transform coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inverse - transforms the values inverse - quantized by the inverse quantization unit 125 to generate residual samples.
[0051] The addition unit 140 combines the residual samples and the prediction samples to restore the picture. The residual samples and the prediction samples are added in block units to generate a restored block. Here, although the addition unit 140 has been described with a separate configuration, it may be a part of the prediction unit 110. On the other hand, the addition unit 140 is also called a restoration unit or a restored block generation unit.
[0052] For the reconstructed picture, the filter unit 150 can apply a deblocking filter and / or a sample adaptive offset. Through deblocking filtering and / or sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process can be corrected. The sample adaptive offset may be applied on a sample-by-sample basis or after the deblocking filtering process is completed. The filter unit 150 can also apply an ALF (Adaptive Loop Filter) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset have been applied.
[0053] The memory 160 stores the reconstructed picture (decoded picture) or information necessary for encoding / decoding. Here, the reconstructed picture may be the one on which the filtering procedure by the filter unit 150 has been completed. The stored reconstructed picture may be utilized as a reference picture for (inter) prediction of other pictures. For example, the memory 160 can store the (reference) picture used for inter prediction. At this time, the picture used for inter prediction may be specified by a reference picture set or a reference picture list.
[0054] Figure 2 shows an example of a video encoding method performed by a video encoding device. Referring to Figure 2, the video encoding method may include a block partitioning, intra / inter prediction, transform, quantization, and entropy encoding process. For example, a current picture may be divided into a plurality of blocks, and a predicted block of the current block may be generated by intra / inter prediction. A residual block of the current block may be generated by subtracting the input block of the current block from the predicted block. Thereafter, a coefficient block, i.e., a transform coefficient of the current block, may be generated by performing a transform on the residual block. The transform coefficient may be quantized and entropy encoded and stored in a bitstream.
[0055] Figure 3 is a diagram schematically explaining the configuration of a video decoding device to which the present invention is applied.
[0056] Referring to Figure 3, the video decoding device 300 may include an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 may include a reordering unit 321, an inverse quantization unit 322, and an inverse transform unit 323.
[0057] When a bitstream including video information is input, the video decoding device 300 can restore the video corresponding to the process in which the video information is processed by the video encoding device.
[0058] For example, the video decoding device 300 can perform video decoding using the processing units applied in the video encoding device. Therefore, the processing unit block for video decoding may be, as an example, a coding unit, or may be, as another example, a coding unit, a prediction unit, or a conversion unit. The coding unit may be divided from the maximum coding unit by a quad-tree structure and / or a binary-tree structure.
[0059] The prediction unit and the conversion unit may be further used as appropriate. In this case, the prediction block is a block derived from or partitioned from the coding unit, and may be a unit for sample prediction. At this time, the prediction unit may be divided into sub-blocks. The conversion unit may be divided from the coding unit by a quad-tree structure, and may be a unit for deriving conversion coefficients or a unit for deriving a residual signal from the conversion coefficients.
[0060] The entropy decoding unit 310 parses the bitstream and outputs information necessary for video restoration or picture restoration. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the conversion coefficient for the residual.
[0061] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information adjacent to the block to be decoded, and the decoding information of the block to be decoded or the information of the symbol / bin decoded in the previous step, and can predict the occurrence probability of the bin based on the determined context model and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.
[0062] Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient, is input to the reordering unit 421.
[0063] The reordering unit 321 reorders the quantized transform coefficients in a two-dimensional block form. The reordering unit 321 can perform reordering corresponding to the coefficient scanning performed in the encoding device. Here, although the reordering unit 321 has been described separately, it may be a part of the inverse quantization unit 322.
[0064] The inverse quantization unit 322 can inverse-quantize the quantized transform coefficients based on the (inverse) quantization parameters and output the transform coefficients. At this time, the information for deriving the quantization parameters may be signaled from the encoding device.
[0065] The inverse transform unit 323 can inverse-transform the transform coefficients to derive residual samples.
[0066] The prediction unit 330 can perform a prediction on the current block and generate a predicted block including a prediction sample for the current block. The unit of prediction performed by the prediction unit 330 may be a coding block, a transform block, or a prediction block.
[0067] The prediction unit 330 determines whether to apply intra prediction or inter prediction based on the information regarding the prediction. At this time, the unit for determining which of intra prediction and inter prediction to apply is different from the unit for generating a prediction sample. Also, in inter prediction and intra prediction, the units for generating a prediction sample are different. For example, which of inter prediction and intra prediction to apply can be determined in CU units. Also, for example, in inter prediction, the prediction mode may be determined in PU units to generate a prediction sample, and in intra prediction, the prediction mode may be determined in PU units and the prediction sample may be generated in TU units.
[0068] In the case of intra prediction, the prediction unit 330 can derive a prediction sample for the current block based on adjacent reference samples within the current picture. The prediction unit 330 can apply a directional mode or a non-directional mode based on the adjacent reference samples of the current block to derive a prediction sample for the current block. At this time, the prediction mode to be applied to the current block may be determined using the intra prediction mode of an adjacent block. On the other hand, MIP (Matrix-based Intra Prediction) for performing a prediction based on a pre-trained matrix may be used. In this case, the number of MIP modes and the size of the matrix are defined for each block size. After downsampling the reference samples to match the size of the matrix, the matrix determined by the mode number is multiplied, and interpolation is performed to match the size of the prediction block to generate a prediction value.
[0069] In the case of inter prediction, the prediction unit 330 can derive a prediction sample for the current block based on samples specified on the reference picture by the motion vector on the reference picture. The prediction unit 330 can apply any one of the skip mode, merge mode, and MVP mode to derive a prediction sample for the current block. At this time, the motion information necessary for the inter prediction of the current block provided from the video encoding device, for example, information regarding a motion vector, a reference picture index, etc., may be obtained or derived based on the information regarding the prediction.
[0070] In the case of the skip mode and the merge mode, the motion information of adjacent blocks may be used as the motion information of the current block. At this time, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0071] The prediction unit 330 may configure a merge candidate list with the motion information of available adjacent blocks, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled from the encoding device. The motion information may include a motion vector and a reference picture. In the skip mode and the merge mode, when the motion information of temporally adjacent blocks is used, the top picture on the reference picture list may be used as the reference picture.
[0072] In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted.
[0073] In the case of the MVP mode, the motion vector of the current block may be derived by using the motion vector of an adjacent block as a motion vector predictor. At this time, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0074] As an example, when the merge mode is applied, a merge candidate list may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks that are temporally adjacent blocks. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The information regarding the prediction may include a merge index indicating a candidate block having an optimal motion vector selected from among the candidate blocks included in the merge candidate list. At this time, the prediction unit 330 may derive the motion vector of the current block using the merge index.
[0075] As another example, when the MVP (Motion Vector Prediction) mode is applied, a motion vector predictor candidate list may be generated using the motion vectors of restored spatially adjacent blocks and / or the motion vectors corresponding to Col blocks that are temporally adjacent blocks. That is, the motion vectors of restored spatially adjacent blocks and / or the motion vectors corresponding to Col blocks that are temporally adjacent blocks may be used as motion vector candidates. The information related to the prediction may include a predicted motion vector index indicating an optimal motion vector selected from among the motion vector candidates included in the list. At this time, the prediction unit 330 can select a predicted motion vector of the current block from among the motion vector candidates included in the motion vector candidate list using the motion vector index. The prediction unit of the encoding device obtains a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encodes this, and outputs it in the form of a bitstream. That is, the MVD is obtained as a value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit 330 can obtain the motion vector difference included in the information related to the prediction, and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. Also, the prediction unit can obtain or derive a reference picture index indicating a reference picture, etc. from the information related to the prediction.
[0076] The addition unit 340 adds the residual sample and the prediction sample to restore the current block or the current picture. The addition unit 340 may add the residual sample and the prediction sample in block units to restore the current picture. When the skip mode is applied, since the residual is not transmitted, the prediction sample becomes the restored sample. Here, the addition unit 340 has been described as a separate configuration, but it may also be a part of the prediction unit 330. On the other hand, the addition unit 340 is also called a restoration unit or a restored block generation unit.
[0077] The filter unit 350 may apply a deblocking filter sample adaptive offset and / or ALF or the like to the restored picture. At this time, the sample adaptive offset may be applied in units of samples or may be applied after deblocking filtering. ALF may be applied after deblocking filtering and / or the sample adaptive offset.
[0078] The memory 360 stores the restored picture (decoded picture) or information necessary for decoding. Here, the restored picture may be a restored picture for which the filtering procedure has been completed by the filter unit 350. For example, the memory 360 may store a picture used for inter prediction. At this time, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The restored picture may be used as a reference picture for other pictures. Also, the memory 360 may output the restored pictures in output order.
[0079] FIG. 4 is a diagram showing an example of a video decoding method performed by a decoding apparatus. Referring to FIG. 4, the video decoding method may include entropy decoding, inverse quantization, inverse transform, and intra / inter prediction processes. For example, in the decoding apparatus, the reverse process of the encoding method may be performed. Specifically, the quantization coefficients quantized by entropy decoding for the bitstream may be obtained, and the coefficient block of the current block, that is, the quantization coefficients, may be obtained by an inverse quantization process for the quantized coefficients. The residual block of the current block may be derived by inverse transformation of the quantization coefficients, and the reconstructed block of the current block may be derived by adding the prediction block of the current block derived by intra / inter prediction and the residual block.
[0080] On the one hand, the operator in the embodiments described below may be defined as follows in the following table.
[0081]
Table 1
[0082] Referring to Table 1, Floor(x) represents the largest integer value less than or equal to x, Log2(u) represents the logarithm value of u with base 2, and Ceil(x) represents the smallest integer value greater than or equal to x. For example, in the case of Floor(5.93), since the largest integer value less than or equal to 5.93 is 5, it represents 5.
[0083] Also, referring to Table 1, x>>y represents an operator that right-shifts x by y bits, and x<<y represents an operator that left-shifts x by y bits.
[0084] <Introduction> The HEVC standard generally uses DCT, which is one type of transform type. Therefore, there is no need to transmit a separate determination process for the transform type and information regarding the determined transform type. However, currently, when the size of the luma block is 4x4 and intra prediction is performed, the DST transform type is exceptionally used.
[0085] Among the quantized coefficients that have undergone the transformation and quantization processes, the information representing the positions of non-zero coefficients can be broadly classified into three types.
[0086] 1. Position (x, y) of the last significant coefficient: The position of the lowest-rank non-zero coefficient (hereinafter defined as the last position) in the scan order within the block to be encoded
[0087] 2. Coded sub-block flag: A flag that divides a block to be coded into a number of sub-blocks and indicates whether each sub-block contains one or more non-zero coefficients (or a flag that indicates whether all coefficients are zero).
[0088] 3. Significant coefficient flag: A flag that indicates whether each coefficient within a sub-block is non-zero or zero.
[0089] Here, the position of the last significant coefficient is displayed separately for the x-axis component and the y-axis component, and each component is expressed separately as a prefix and a suffix. That is, the syntax that indicates the non-zero position of the quantized coefficient includes the following six syntaxes in total.
[0090] 1. last_sig_coeff_x_prefix 2. last_sig_coeff_y_prefix 3. last_sig_coeff_x_suffix 4. last_sig_coeff_y_suffix 5. coded_sub_block_flag 6. sig_coeff_flag
[0091] The last_sig_coeff_x_prefix indicates the prefix of the x-axis component indicating the position of the last significant coefficient, and the last_sig_coeff_y_prefix indicates the prefix of the y-axis component indicating the position of the last significant coefficient. Also, the last_sig_coeff_x_suffix indicates the suffix of the x-component indicating the position of the last significant coefficient, and the last_sig_coeff_y_suffix indicates the suffix of the y-component indicating the position of the last significant coefficient.
[0092] On the one hand, the coded_sub_block_flag indicates "0" if all coefficients within the sub-block are all zero, and indicates "1" if there is one or more non-zero coefficients. The sig_coeff_flag indicates "0" for zero coefficients and indicates "1" for non-zero coefficients. Considering the position of the last valid coefficient within the block to be coded, in the scan order, the coded_sub_block_flag syntax is transmitted only for the previously existing sub-blocks. When the coded_sub_block_flag is "1", that is, when there is one or more non-zero coefficients, the sig_coeff_flag syntax for each of all the coefficients within the sub-block is transmitted.
[0093] The HEVC standard supports the following three types of scans for coefficients.
[0094] 1) Up-right diagonal 2) Horizontal 3) Vertical
[0095] When the block to be coded is coded using the inter-picture prediction method, the coefficients of the block are scanned in the up-right diagonal method. When the block is coded using the intra-picture prediction method, one of the above three types is selected according to the intra-picture prediction mode, and the coefficients of the block are scanned.
[0096] That is, when the block to be coded is coded by the video coding device and the inter-picture prediction method is used, the coefficients of the block are scanned in the up-right diagonal method. When the intra-picture prediction method is used in the coding of the block to be coded, the video coding device selects one of the above three types according to the intra-picture prediction mode and scans the coefficients of the block. The above scan may be performed by the reordering unit 124 in the video coding device shown in FIG. 1, and the scan may change the two-dimensional block-form coefficients into one-dimensional vector form.
[0097] FIG. 5 is a diagram showing the scan order of sub-blocks and coefficients for the diagonal scan method.
[0098] Referring to FIG. 5, when the block of FIG. 5 is scanned in the diagonal scan method by the rearrangement unit 124 of the video encoding apparatus, scanning is performed from the sub-block 1, which is the sub-block at the uppermost left end, downward and in the upper end direction of the diagonal line, and finally scanning is performed on the 16th sub-block at the lower right end. That is, the rearrangement unit 124 performs scanning in the order of sub-blocks 1, 2, 3, …, 14, 15, 16, and rearranges the quantized transform coefficients in the two-dimensional block form into a one-dimensional vector form. Similarly, the rearrangement unit 124 of the video encoding apparatus also performs scanning on the coefficients in each sub-block in the same diagonal scan method as the sub-block scan method. For example, in the 1st sub-block, scanning is performed in the order of coefficients 0, 1, 2, …, 13, 14, 15.
[0099] However, when the scanned coefficients are stored in the bit stream, the storage order is the reverse of the scan order. That is, when the block of FIG. 10 is scanned by the rearrangement unit 124 of the video encoding apparatus, scanning is performed in the order of coefficients from 0 to 255, but the order in which each pixel is stored in the bit stream is stored in the bit stream in the order of pixels from the pixel at the 255th position to the pixel at the 0th position.
[0100] FIG. 6 is a diagram showing an example of a 32×32 block to be encoded after quantization by a video encoding device. Here, when the 32×32 block shown in FIG. 6 is scanned by the video encoding device, a diagonal method is arbitrarily used. In FIG. 6, the pixels indicated by diagonal hatching represent non-zero coefficients, and the pixels indicated by "x" represent the last significant coefficient. All the other white coefficients have zero values. Here, if the coded_sub_block_flag syntax is substituted into the block of FIG. 6, coded_sub_block_flag information for 24 sub-blocks existing before the last position in the scan order among a total of 64 sub-blocks, that is, the sub-blocks shown by thick lines in FIG. 6, is required. Among the 24 sub-blocks, the coded_sub_block_flag values for the first sub-block including the DC value and the 24th sub-block including the last position coefficient are induced to be "1", and the coded_sub_block_flag values for the remaining 22 sub-blocks are transmitted to the video decoding device by the bit stream. At this time, in the case of a sub-block including one or more non-zero coefficients among the 22 sub-blocks, the coded_sub_block_flag value is set to "1" by the video encoding device. In FIG. 6, among the 22 sub-blocks excluding the first sub-block and the 24th sub-block, the coded_sub_block_flag values of the 4th, 5th, 11th, and 18th sub-blocks, which are sub-blocks including the pixels marked in gray, are set to "1".
[0101] 1. Method for determining the primary conversion type of the block to be decoded This specification discloses a method for determining the type of primary transformation of a block to be decoded during the video decoding process. That is, when a block to be decoded is decoded by a video decoding device, in the transformation process by the video encoding device, a process of determining what transformation type was used for the primary transformation and encoding is required. The primary transformation type is composed of one default transformation and a plurality of additional transformations. The block to be decoded, depending on the conditions, uses the default transformation or a set of multiple transformations including the default transformation and additional transformations. That is, the block to be decoded may be transformed using only the default transformation in the transformation process, or may be transformed using a set of multiple transformations including the default transformation and additional transformations. From the perspective of the video decoding device, conversely, it may perform decoding by grasping whether the block to be decoded uses only the default transformation or a set of multiple transformations including the default transformation and additional transformations. When the block to be decoded uses MTS, information regarding the actually used transformation among the multiple transformations is transmitted or induced. Here, the information regarding the actually used transformation may separately exist for the horizontal axis transformation type and the vertical axis transformation type. That is, when the block to be decoded is transformed using MTS, the video decoding device may receive or determine what transformation type was used for the transformation among the multiple transformation types and then perform decoding.
[0102] According to one embodiment, DCT-II may be set as the default conversion, and DST-7 and DCT-8 may be set as additional conversions. At this time, the maximum size of DCT-II, which is the default conversion, is supported up to 64×64, and the maximum sizes of DST-7 and DCT-8, which are additional conversions, are supported up to 32×32. For example, when the size of the block to be decoded is 64×64, one 64×64 DCT-II is applied to the conversion process. That is, when one or more of the width and height of the block to be decoded are greater than 32 (exceeding 32), the default conversion (*) is applied immediately without applying MTS. That is, from the perspective of the video decoding device, it is only necessary to determine whether MTS is used for conversion when the horizontal and vertical sizes of the block to be decoded are both 32 or less. On the other hand, when one of the horizontal or vertical sizes of the block to be decoded is greater than 32, it may be determined that the default conversion is applied and the conversion is performed. In this way, when the block to be decoded is converted by the default conversion, there is no syntax information transmitted in relation to MTS. In the present invention, for convenience, the conversion type value of DCT-II is set to "0", the conversion type value of DST-7 is set to "1", and the conversion type value of DCT-8 is set to "2", but it is not limited thereto. Table 2 below defines the conversion types assigned to each value of the trType syntax.
[0103]
Table 2
[0104] Tables 3 and 4 show an example of the conversion kernels of DST-7 and DCT-8 when the size of the block to be multiplexed is 4×4.
[0105] Table 3 shows the coefficient values of the conversion kernel when tyType is "1" (DST-7) and the size of the block to be decoded is 4×4, and Table 4 shows the coefficient values of the conversion kernel when tyType is "2" (DCT-8) and the size of the block to be decoded is 4×4.
[0106]
Table 3
[0107]
Table 4
[0108] The entire conversion area of the block to be decoded may include a zero-out area. The conversion converts the values in the pixel domain to frequency domain values. At this time, the upper left frequency area is referred to as the low-frequency area, and the lower right frequency area is referred to as the high-frequency area. The low-frequency components reflect the general (average) characteristics of the block, and the high-frequency components reflect the sharp (peculiar) characteristics of the block. Therefore, there are a large number of large values in the low-frequency components and a small number of small values in the high-frequency components. Through the quantization process after conversion, the small number of small values in the high-frequency area will mostly have 0 values. Here, in addition to the low-frequency area belonging to the upper left side, the remaining area with mostly 0 values is referred to as the zero-out area, and the zero-out area may be excluded in the signaling process. The area excluding the zero-out area in the block to be decoded is referred to as the valid area.
[0109] FIG. 7 is a diagram showing the remaining zero-out area excluding m×n in the area of the M×N block to be decoded.
[0110] Referring to FIG. 7, the gray area on the upper left side indicates the low-frequency area, and the white area indicates the high-frequency zero-out.
[0111] As another example, when the block to be decoded is 64×64, the upper left 32×32 area becomes the valid area, and the remaining area excluding this becomes the zero-out area and is not signaled.
[0112] Also, when the size of the block to be decoded is one of 64×64, 64×32, or 32×64, the upper left 32×32 region becomes the valid region, and the remaining part excluding this becomes the zero-out region. When the decoder parses the syntax for the quantized coefficients, since it knows the size of the block, the zero-out region is not signaled. That is, a region where the width or height of the block to be decoded is greater than 32 is set as the zero-out region. At this time, since it corresponds to the case where the horizontal or vertical size of the block to be decoded is greater than 32, the transform used is the default transform, DCT-II. Since the maximum sizes of the additional transforms DST-7 and DCT-8 are supported up to 32×32, MTS is not applied to blocks of such sizes.
[0113] Also, when the size of the block to be decoded is one of 32×32, 32×16, or 16×32 and MTS is applied to the block to be decoded (for example, when using DST-7 or DCT-8), the upper left 16×16 region becomes the valid region, and the remaining part excluding this is set as the zero-out region. Here, the zero-out region may be signaled by the position of the last valid coefficient and the scanning method. This is because after the encoder finishes signaling the syntax related to the quantized coefficients, it signals the MTS index (mts_idx) value. So, when the decoder parses the syntax for the quantized coefficients, it does not know the information about the transform type. When the zero-out region is signaled in this way, the decoder can perform the transform only on the valid region after ignoring or removing the quantized coefficients corresponding to the zero-out region. Here, the information about the actually used transform type may exist separately for the horizontal axis transform type and the vertical axis transform type. For example, when the horizontal axis transform type of the block to be decoded is DST-7 or DCT-8, the horizontal axis (width) valid region of the block to be decoded is 16, and when the vertical axis transform type of the block to be decoded is DST-7 or DCT-8, the vertical axis (height) valid region of the block to be decoded is 16.
[0114] On the one hand, when the size of the block to be decoded is one of 32×32, 32×16, or 16×32 and MTS is not applied to the block to be decoded (for example, when using DCT-II which is the default transform), all regions are valid regions, there is no zero-out region, and it is transformed using DCT-II which is the default transform. Here, information regarding the actually used transform may separately exist for the horizontal axis transform type and the vertical axis transform type. For example, when the horizontal axis transform type of the block to be decoded is DCT-II, the horizontal axis (width) valid region of the block to be decoded is the width of the block, and when the vertical axis transform type of the block to be decoded is DCT-II, the vertical axis (height) valid region of the block to be decoded is the height of the block. That is, all regions (width × height) of the block to be decoded become valid regions.
[0115] On the other hand, when the block to be decoded has a size smaller than 16 which is not defined as above, all regions become valid regions and there is no zero-out region. Whether MTS is applied to the block to be decoded and the transform type value are determined by implicit and / or explicit embodiments.
[0116] In the present invention, among the quantized coefficients that have undergone the transform and quantization processes, the content of the syntax representing the positions of non-zero coefficients is the same as that of the HEVC method. However, the syntax name of coded_sub_block_flag is changed to sb_coded_flag and used. Also, the method of scanning the quantized coefficients uses the upright diagonal method.
[0117] In the present invention, MTS may be applied to a luma block (not applied to a chroma block). Also, the MTS function can be turned on / off using a flag indicating whether MTS is used, i.e., sps_mts_enabled_flag. When using the MTS function, sps_mts_enabled_flag is set to on, and it is possible to set whether to use the explicit MTS function for each of intra-picture prediction and inter-picture prediction. That is, sps_explicit_mts_intra_enabled_flag, a flag indicating whether to use MTS during intra-picture prediction, and sps_explicit_mts_inter_enabled_flag, a flag indicating whether to use MTS during inter-picture prediction, may be separately set. In this specification, for convenience, it is described that the values of the three flags sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, and sps_explicit_mts_inter_enabled_flag indicating whether to use MTS are located within the SPS (sequence parameter set), but it is not limited thereto. That is, the three flags may be set at one or more positions of DCI (decoding capability information), VPS (video parameter set), SPS (sequence parameter set), PPS (picture parameter set), PH (picture header), and SH (slice header), respectively. Also, the three flags indicating whether to use MTS may be defined as HLS (high level syntax).
[0118] The use of MTS can be divided into explicit and implicit methods. For the explicit use of MTS, a flag value indicating how MTS is used within the SPS on the screen and / or between screens is set to on, and when specific conditions are met, MTS-related information (e.g., the conversion information actually used) is transmitted. That is, the video decoder can receive the MTS-related information, based on which it can confirm what conversion type the block to be decoded was converted with, and based on this, perform decoding. For example, in an environment where explicit MTS is used, the three flags may be set as follows.
[0119] 1. sps_mts_enabled_flag = on 2. sps_explicit_mts_intra_enabled_flag = on 3. sps_explicit_mts_inter_enabled_flag = on
[0120] For the implicit use of MTS, among the three flags within the SPS, when the value of sps_mts_enabled_flag is set to on and specific conditions are met, MTS-related information (e.g., the conversion information actually used) is induced. For example, in an environment where implicit MTS is used, the three flags may be set as follows.
[0121] 1. sps_mts_enabled_flag = on 2. sps_explicit_mts_intra_enabled_flag = off 3. sps_explicit_mts_inter_enabled_flag = off (regardless of on or off)
[0122] Hereinafter, several embodiments will be used to describe the implicit MTS method and the explicit MTS method.
[0123] 2. First Embodiment (Implicit MTS) The implicit MTS described in this embodiment may be used when the block to be decoded is encoded by the intra prediction method. That is, when the block to be decoded is encoded by the video encoding apparatus, if it is encoded by the intra prediction method, encoding and / or decoding may be performed using the implicit MTS by the video encoding apparatus and / or the video decoding apparatus. On the other hand, when decoding the block to be decoded, whether to use the implicit MTS may be indicated by the implicitMtsEnabled parameter. The video decoding apparatus may check the value of the implicitMtsEnabled parameter and determine whether to perform decoding using the implicit MTS. For example, when the implicit MTS is used for decoding, the implicitMtsEnable parameter may have a value of 1, and otherwise, the implicitMtsEnable parameter may have a value of 0. On the other hand, in this specification, the implicitMtsEnabled may be displayed as "implicit_MTS_enabled" depending on the case.
[0124] When examining the HLS (high level syntax) conditions for applying such implicit MTS, since the sps_mts_enabled_flag is a flag indicating whether MTS is applied regardless of whether it is implicit or explicit, it must be set to "on" for implicit MTS to be applied. On the other hand, implicit MTS is used when the block to be decoded is encoded by the video encoding device using the intra prediction method. Therefore, the video decoding device can determine whether to use implicit MTS by checking the value of the sps_explicit_mts_intra_enabled_flag. However, the sps_explicit_mts_intra_enabled_flag is set to "on" when the block to be decoded is encoded by the video encoding device using the intra prediction method and explicit MTS is applied. Therefore, when the block to be decoded is encoded by the video encoding device using implicit MTS, the sps_explicit_mts_intra_enabled_flag is set to "off". On the other hand, as explained above, implicit MTS is used when the block to be decoded is encoded by the video encoding device using the intra prediction method. Therefore, the value of the sps_explicit_mts_inter_enabled_flag indicating explicit MTS when the block to be decoded is encoded by the video encoding device using the intra prediction method is not important. On the other hand, since implicit MTS may be used when the block to be decoded is encoded by the video encoding device using the intra prediction method, it may be applied when the CuPredMode has the value MODE_INTRA.
[0125] To summarize, the conditions for the block to be decoded to be decoded using implicit MTS by the video decoding device can be listed as follows.
[0126] 1) sps_mts_enabled_flag is equal to 1 2) The sps_explicit_mts_intra_enabled_flag is equal to 0 3) The CuPredMode is equal to MODE_INTRA (intra-prediction method)
[0127] On the other hand, CuPredMode[0][xTbY][yTbY] indicating the prediction mode of the current position in the luma block may have a MODE_INTRA value.
[0128] The additional conditions for enabling implicit MTS are as follows.
[0129] 4) The lfnst_idx is equal to 0 5) The intra_mip_flag is equal to 0
[0130] Here, the lfnst_idx value indicates a secondary transform. When lfnst_idx = 0, it means that no secondary transform is used. The Intra_mip_flag value indicates whether a prediction method using a matrix (matrix-based intra prediction: mip), which is one of the intra-prediction methods, is used. When intra_mip_flag = 0, it means that no prediction using a matrix is used, and when intra_mip_flag = 1, it means that a prediction using a matrix is used.
[0131] That is, this embodiment describes a method of setting a primary transform type (or MTS) for a decoding target block that does not use a secondary transform while predicting using a general intra-prediction method. When all of the above five conditions are satisfied, the implicit MTS function is activated (see FIG. 13).
[0132] FIG. 8 is a diagram showing a method for determining whether to apply the implicit MTS function according to an embodiment of the present invention. Each step in FIG. 8 may be performed in a video decoding apparatus.
[0133] Referring to FIG. 8, the video decoding apparatus determines whether the sps_mts_enable_flag has a value of 1, the sps_explicit_mts_intra_enable_flag has a value of 0, and the CuPredMode has a MODE_INTRA value (S810). As a result of the determination, if all the conditions of S810 are satisfied, the video decoding apparatus determines whether the lfnst_idx has a value of 0 and the intra_mip_flag has a value of 0 (S820). If all the conditions of S810 and S820 are satisfied, the implicit_MTS_enabled value is set to 1 (S830). On the other hand, if the video decoding apparatus does not satisfy any of the conditions of S810 or S820, the implicit_MTS_enabled value is set to 0 (S840).
[0134] When the implicit MTS function for the block to be decoded is activated (implicit_MTS_enabled = on), the MTS value (the conversion information actually used) is derived based on the width and height of the block (see FIG. 9). At this time, the conversion must not be a sub-block transform (sbt) in which only a part of the target block undergoes the conversion process. That is, the cu_sbt_flag value of the target block is "0".
[0135] FIG. 9 is a diagram showing a method of deriving conversion information based on the width and height of the block of implicit MTS according to an embodiment of the present invention. Each step in FIG. 9 may be performed in the video decoding apparatus.
[0136] Referring to FIG. 9, the video decoding apparatus determines whether the implicit_MTS_enabled value is "1" (S910). At this time, although not shown in the figure, the video decoding apparatus may further confirm whether the cu_sbt_flag value has a value of "0". At this time, when the cu_sbt_flag has a value of "1", it indicates that the decoding target block has been converted by sub-block conversion in which only a part of the target block undergoes a conversion process. On the contrary, when the cu_sbt_flag has a value of "0", it indicates that the decoding target block has not been converted by sub-block conversion in which only a part of the target block undergoes a conversion process. Therefore, the operation according to FIG. 14 may be set to operate only when the cu_sbt_flag has a value of "0".
[0137] When the implicit_MTS_enabled value is 1, it is determined whether the value of nTbW is 4 or more and 16 or less (S920). When the implicit_MTS_enabled value is not "1", the operation ends. nTbW indicates the width of the conversion block and is used to determine whether DST-7, which is an additional conversion type, is used in the horizontal axis direction.
[0138] As a result of the determination in step S920, when the value of nTbW is 4 or more and 16 or less, trTypeHor is set to "1" (S930). When the value of nTbW is not 4 or more and 16 or less, trTypeHor is set to "0" (S940). At this time, the nTbW indicates the width of the conversion block and is used to determine whether DST-7, which is an additional conversion type, is used in the horizontal axis direction. At this time, when tyTypeHor is set to "0", it can be determined that the conversion block has been converted using DCT-II conversion, which is the default type conversion in the horizontal axis direction. On the other hand, when the trTypeHor is set to "1", it can be determined that the conversion block has been converted using DST-7 conversion, which is one of the additional conversion types, in the horizontal axis direction.
[0139] Further, the video decoding device determines whether the value of nTbH has a value of 4 or more and 16 or less (S950). When the value of nTbH has a value of 4 or more and 16 or less, trTypeVer is set to "1" (S960). When the value of nTbW does not have a value of 4 or more and 16 or less, trTypeVer is set to "0" (S970). The nTbH indicates the height of the conversion block and is used to determine whether DST-7, which is an additional conversion type in the vertical axis direction, is used. At this time, when trTypeVer is set to "0", it can be determined that the conversion block is converted using DCT-II conversion, which is the default type conversion in the vertical axis direction. On the other hand, when trTypeVer is set to "1", it can be determined that the conversion block is converted using DST-7 conversion, which is one of the additional conversion types in the vertical axis direction.
[0140] FIG. 10 is a diagram showing a method of executing inverse conversion based on conversion-related parameters according to an embodiment of the present invention. Each step in FIG. 10 may be performed by a video decoding device, for example, by an inverse conversion unit of the decoding device.
[0141] Referring to FIG. 10, the video decoding device acquires sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, Y0], NTbW, and nTbH (S1010). At this time, what each of sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, and IntraMipFlag[x0, Y0] indicates is described in detail in the related description of FIG. 8. The parameters are used to determine whether the decoding target block can apply implicit MTS. Also, NTbW and nTbH respectively indicate the width and height of the conversion block and are used to determine whether DST-7, which is an additional conversion type, is used.
[0142] Next, the video decoding device sets implicit_MTS_enabled based on the values of sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], Lfnst_idx, and IntraMipFlag[x0,Y0] (S1020). At this time, the implicit_MTS_enable may be set by performing the process of FIG. 13.
[0143] Next, the video decoding device sets trTypeHor and trTypeVer based on the values of implicit_MTS_enabled, nTbW, and nTbH (S1030). At this time, the trTypeHor and trTypeVer may be set by performing the process of FIG. 9.
[0144] Next, the video decoding device performs inverse transformation based on trTypeHor and trTypeVer (S1040). The inverse transformation applied by trTypeHor and trTypeVer may be configured according to Table 2. For example, when trTypeHor is "1" and trTypeVer is "0", DST-7 may be applied in the horizontal axis direction and DST-II may be applied in the vertical axis direction.
[0145] On the other hand, although not shown, from the perspective of the video encoding device, in order to set whether to use implicit MTS, sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, Y0], NTbW, and nTbH may be set.
[0146] 3. Second Embodiment (Explicit MTS) In this embodiment, when the MTS function is explicitly activated in the HLS (high level syntax), the conversion method applied to the block to be decoded will be described. When examining the HLS (high level syntax) conditions for the application of explicit MTS, since the sps_mts_enabled_flag is a flag indicating whether MTS is applied regardless of whether it is implicit or explicit, it must be set to "on" for implicit MTS to be applied. On the other hand, explicit MTS may be applied in all cases where the block to be decoded is encoded by the intra-prediction method or the inter-prediction method. Therefore, when explicit MTS is applied, both sps_explicit_mts_intra_enabled_flag and / or sps_explicit_mts_intra_enabled_flag must be set to "on". Summing up, it can be listed as the following conditions.
[0147] 1) sps_mts_enabled_flag = on 2) sps_explicit_mts_intra_enabled_flag = on 3) sps_explicit_mts_inter_enabled_flag = on
[0148] Here, when the block to be decoded is encoded by the intra-prediction method, check the condition that sps_explicit_mts_intra_enabled_flag = "on", and when the block to be decoded is encoded by the inter-prediction method, check the condition that sps_explicit_mts_inter_enabled_flag = "on".
[0149] The additional conditions for the use of explicit MTS are as follows.
[0150] 4) lfnst_idx is equal to 0 (for implicit MTS reference) 5) transform_skip_flag is equal to 0 6) The intra_subpartitions_mode_flag is equal to 0 7) The cu_sbt_flag is equal to 0 (implicit MTS reference) 8) Valid MTS area 9) The width and height of the target block are 32 or less
[0151] Here, the lfnst_idx value indicates a secondary transform. When lfnst_idx = 0, it means that the secondary transform is not used.
[0152] The transform_skip_flag value indicates whether to skip the conversion process. When transform_skip_flag = 0, it indicates that the conversion proceeds normally without skipping the conversion process. The intra_subpartitions_mode_flag value, as one of the in-picture prediction methods, instructs to divide the target block into multiple sub-blocks and go through the prediction, conversion, and quantization processes. That is, when the flag value (intra_subpartitions_mode_flag) is "0", it means that the general in-picture prediction is performed without dividing the target block into sub-blocks. On the other hand, the use of MTS may be restricted depending on the size supported by the additional transforms (DST-7 and DCT-8) (supporting up to 32×32 as described above). That is, if the width and height of the target block are not 32 or less, the use of MTS is not possible. That is, if either the width or the height exceeds 32, the DCT-II, which is the default transform (*) (where the use of MTS is not possible), is performed.
[0153] The cu_sbt_flag indicates whether it is a sub-block transform (sbt) in which only a part of the target block goes through the conversion process. That is, when the cu_sbt_flag value is "0", it means that it is not a sub-block transform in which only a part of the target block goes through the conversion process.
[0154] Hereinafter, the effective area (hereinafter referred to as the effective MTS area) will be described in detail.
[0155] FIG. 11 is a diagram showing the effective MTS area marked with a thick line within a 32×32 block to be decoded.
[0156] Referring to FIG. 11, the upper left 16×16 area excluding the DC coefficient becomes the effective MTS area. That is, the upper left 16×16 area excluding the 1×1 (DC) area is the effective MTS area. For example, if the positions of all non-zero coefficients within the target block belong to the effective MTS area, MTS is applicable. If one or more non-zero coefficient values are outside the effective MTS area, MTS application is not possible, and DCT-II, which is the default transformation (*), is performed. This is the same concept as the zero-out area described above. That is, when a 32×32 target block uses MTS (i.e., DST-7 or DCT-8), the upper left 16×16 becomes the effective area, and the remaining part becomes the zero-out area. Similarly, if all non-zero coefficients within the 32×32 target block are located in the upper left 16×16 area, MTS (i.e., DST-7 or DCT-8) can be applied. However, exceptionally, if there is only one non-zero coefficient within the block and its position is DC (1x1), MTS application is not possible, and DCT-II, which is the default transformation (*), is performed.
[0157] As a result, in this embodiment, in order to determine whether MTS is applicable, the effective MTS area must be checked. To check the effective MTS area, the following two conditions must be checked.
[0158] (a) When there is only one non-zero coefficient within the block, whether its position is DC (1x1) (b) Whether all non-zero coefficients within the block are located in the upper left 16×16 area
[0159] To confirm the condition (a), the information of the last position may be utilized. Here, the last position means the position of the last non-zero coefficient, i.e., the position of the last valid coefficient, in the scanning order within the target block. As an example, the information of the last position, i.e., the last sub-block including the last non-zero coefficient, may be utilized. For example, if the position of the last sub-block is not (0, 0), the condition (a) can be satisfied. In other words, if the position of the last sub-block in the scanning order of the sub-blocks within the target block is not "0" (greater than 0), the condition (a) can be satisfied. Or, if the position of the last sub-block is "0", the information of the last scan position indicating the relative position of the last position within the sub-block may be utilized. For example, if the last scan position is not "0" (greater than 0) in the scanning order of the coefficients within the sub-block, the condition (a) can be satisfied (refer to FIG. 12). Also, as described above, the MTS of the present invention is applied to the luma block.
[0160] To confirm the condition (b), sub-block information including one or more non-zero coefficients may be utilized. Here, the sub-block information including one or more non-zero coefficients can be confirmed by the sb_coded_flag value of the sub-block. When the flag value is "1" (sb_coded_flag = 1), it means that one or more non-zero coefficients are located within the sub-block, and when sb_coded_flag = 0, it means that all coefficients within the sub-block are zero. That is, if the positions of all sub-blocks with sb_coded_flag value "1" within the target block are within (0,0) to (3,3), the condition (b) can be satisfied. On the contrary, if even one of the sub-blocks with sb_coded_flag value "1" within the target block is outside the position from (0,0) to (3,3), the condition (b) cannot be satisfied. In other words, if even one of the sub-blocks with sb_coded_flag value "1" within the target block has a value greater than 3 for either the x-coordinate or y-coordinate of the sub-block, the condition (b) cannot be satisfied (refer to FIG. 18). In other embodiments, in the scanning order of the sub-blocks within the target block, if the first sub-block with sb_coded_flag value "1" but having a value greater than 3 for either the x-coordinate or y-coordinate of the sub-block is found, the condition (b) may be set to false, and the confirmation process for subsequent sub-blocks with sb_coded_flag value "1" in the scanning order may be omitted (refer to FIG. 13). Also, as described above, the MTS of the present invention is applied to the luma block.
[0161] FIG. 12 is a diagram showing a method for determining a valid MTS according to an embodiment of the present invention. The embodiment of FIG. 12 relates to a method for confirming the condition (a) among the two conditions for confirming the valid MTS region described above. Each step in FIG. 12 may be performed within a video decoding device.
[0162] Referring to FIG. 12, the video decoding apparatus sets MtsDcOnlyFlag to "1" (S1210). The MtsDcOnlyFlag can indicate whether the non-zero coefficient in the block is one and its position is DC. For example, if there is one non-zero coefficient in the block and its position is DC, the MtsDcOnlyFlag has a value of "1", and in other cases, the MtsDcOnlyFlag has a value of "0". At this time, when the value of the MtsDcOnlyFlag is "0", MTS may be applied. The reason for setting MtsDcOnlyFlag to "1" in step S1210 is that when the block satisfies the condition of not being in the DC position when there is one non-zero coefficient in the following block, the MtsDcOnlyFlag is reset to "0", and when it is not the case, MTS is not applied.
[0163] Next, the video decoding apparatus determines whether the target block is a luma block (S1220). The purpose of determining whether the target block is a luma block is that, as described above, MTS is only applied to luma blocks.
[0164] Next, the video decoding apparatus determines whether the last sub-block is greater than 0 (S1230). If the last sub-block is greater than 0, the MtsDcOnlyFlag is set to "0" (S1240), and the process ends.
[0165] As a result of the determination in step S1230, if the last sub-block is not greater than 0, it is determined whether the last scan position is greater than 0 (S1250).
[0166] As a result of the determination in step S1250, if the last scan position is greater than 0, the MtsDcOnlyFlag is set to "0" (S1240), and the process ends.
[0167] As a result of the determination in step S1250, if the last scan position is not greater than 0, the process ends.
[0168] According to this embodiment, if the last sub-block is greater than 0 or the last scan position is greater than 0, set MtsDcOnlyFlag to "0"; otherwise, set MtsDcOnlyFlag to "1".
[0169] Subsequently, when determining whether to apply MTS, if MtsDcOnlyFlag is checked and has a value of "1", DCT-II, which is the default conversion, may be applied without applying MTS.
[0170] FIG. 13 is a diagram showing a method for determining a valid MTS region according to another embodiment of the present invention. The embodiment of FIG. 13 specifically shows a method for checking condition (b) among the two conditions for checking the valid MTS region described above. Each step in FIG. 13 may be performed in a video decoding apparatus.
[0171] Referring to FIG. 13, the video decoding apparatus sets MtsZerooutFlag to "1" (S1305). The MtsZerooutFlag indicates whether non-zero coefficients in the block exist in the zero-out region. For example, when at least one of the non-zero coefficients in the block exists in the zero-out region, MtsZerooutFlag may have a value of "0", and when all non-zero coefficients in the block do not exist in the zero-out region, MtsZerooutFlag may have a value of "1". In this embodiment, assuming that all non-zero coefficients in the block do not exist in the zero-out region, the initial value of MtsZerooutFlag is set to "1", and when the conditions of the zero-out region and the non-zero coefficient are satisfied simultaneously, MtsZerooutFlag may be set to "0". At this time, if there is a MtsZerooutFlag having a value of "0", it may not be necessary to apply explicit MTS.
[0172] Next, the video decoding apparatus sets the initial value of variable i to the value of the last sub-block, and subtracts 1 from the value of variable i one by one, and repeats the processes of the following steps S1325 to S1350 until the value of variable i becomes 0 (S1320). The purpose of repeatedly executing the routine of step S1820 is to check the sb_coded_flag values of all sub-blocks from the last sub-block to the first sub-block. As described above, when the flag value is "1", it means that there is one or more non-zero coefficients in the sub-block, and when the flag value is "0", it means that there are no non-zero coefficients in the sub-block. Therefore, referring to FIG. 11, if the positions of all sub-blocks in the target block with sb_coded_flag value of "1" only exist within (0, 0) to (3, 3), that is, only exist within 0 to 8 based on variable i, it may be determined that the condition (b) for applying explicit MTS is satisfied.
[0173] Next, the video decoding apparatus determines whether the variable i simultaneously satisfies the conditions that the variable i is smaller than the last sub-block (i < the last sub-block) and the variable i is larger than 0 (i > 0) (S1325). For example, when the routine of step S1320 is executed for the first time, since the initial value of variable i is set to the same value as the last sub-block, the condition of step S1325 is not satisfied.
[0174] As a result of the determination in step S1325, when the variable i simultaneously satisfies the conditions that the variable i is smaller than the last sub-block (i < the last sub-block) and the variable i is larger than 0 (i > 0), parse sb_coded_flag (S1830). When the two conditions are not simultaneously satisfied, set sb_coded_flag to "1" (S1835).
[0175] At this time, the sb_coded_flag to be parsed indicates whether one or more non-zero coefficients exist within the sub-block. If one or more non-zero coefficients exist within the sub-block, the sb_coded_flag has a value of "1", and if no non-zero coefficients exist within the sub-block, the sb_coded_flag has a value of "0".
[0176] On the other hand, step S1835 is performed only when i indicates the last sub-block and the first sub-block. That is, since the last position coefficient is included in the last sub-block, the sb_coded_flag value is parsed as a value of "1", and since the DC coefficient exists in the first sub-block, the sb_coded_flag value is parsed as a value of "1".
[0177] Next, the video decoding device determines whether the block is a luma block (S1340). The purpose of determining whether the target block is a luma block is that, as explained above, MTS is applied only to luma blocks.
[0178] As a result of the determination in step S1340, if the block is a luma block, it is determined whether the condition "sb_coded_flag && (xSb>3||ySb>3)" is satisfied (S1845). If the condition in step S1845 is satisfied, MtsZerooutFlag is set to "0" (S1350).
[0179] According to this embodiment, if even one non-zero coefficient is found in a sub-block other than the sub-block (3,3) within the target block, that is, in the zero-out region, MtsZerooutFlag may be set to "0" and it may be determined that explicit MTS cannot be applied.
[0180] FIG. 14 is a diagram showing a method for determining an effective MTS according to another embodiment of the present invention. The embodiment of FIG. 14 specifically shows a method for checking the condition (b) among the two conditions for checking the effective MTS region described above. However, in the embodiment of FIG. 13, the effective MTS region is checked by checking the sb_coded_flag of all sub-blocks, while in the embodiment of FIG. 14, when the first invalid MTS is found, it is different that the subsequent sb_coded_flag does not need to be checked. Each step in FIG. 14 may be performed in a video decoding device.
[0181] Referring to FIG. 14, the video decoding device sets the MtsZerooutFlag to "1" (S1405). The MtsZerooutFlag indicates whether non-zero coefficients in the block exist in the zero-out region. For example, when at least one of the non-zero coefficients in the block exists in the zero-out region, the MtsZerooutFlag may have a value of "0", and when all non-zero coefficients in the block do not exist in the zero-out region, the MtsZerooutFlag may have a value of "1". In this embodiment, assuming that all non-zero coefficients in the block do not exist in the zero-out region, the initial value of the MtsZerooutFlag is set to "1", and when the conditions of the zero-out region and the non-zero coefficient are simultaneously satisfied, the MtsZerooutFlag may be set to "0". At this time, if there is a MtsZerooutFlag having a value of "0", an explicit MTS may not be applied.
[0182] Next, the video decoding device sets the initial value of variable i to the value of the last sub-block, subtracts 1 from the value of variable i one by one, and repeats the processes of the following step S1425 to step S1450 until the value of variable i becomes 0 (S1420). The purpose of repeatedly performing the routine of step S1420 is to check the sb_coded_flag values of all sub-blocks from the last sub-block to the first sub-block. As described above, when the sb_coded_flag value is "1", it means that there is one or more non-zero coefficients in the sub-block, and when the sb_coded_flag value is "0", it means that there are no non-zero coefficients in the sub-block. Therefore, referring to FIG. 16, if the positions of all sub-blocks with sb_coded_flag value "1" in the target block only exist within (0, 0) to (3, 3), that is, only within 0 to 8 based on variable i, it may be determined that the condition (b) for applying explicit MTS is satisfied.
[0183] Next, the video decoding device determines whether the variable i simultaneously satisfies the conditions that the variable i is smaller than the last sub-block (i < last sub-block) and the variable i is larger than 0 (i > 0) (S1425). For example, when the routine of step S1920 is executed for the first time, since the initial value of variable i is set to the same value as the last sub-block, the condition of step S1425 is not satisfied.
[0184] As a result of the determination in step S1425, when the variable i simultaneously satisfies the conditions that the variable i is smaller than the last sub-block (i < last sub-block) and the variable i is larger than 0 (i > 0), parse sb_coded_flag (S1430), and when the two conditions are not simultaneously satisfied, set sb_coded_flag to "1" (S1435).
[0185] At this time, the parsed sb_coded_flag indicates whether one or more non-zero coefficients exist within the sub-block. If one or more non-zero coefficients exist within the sub-block, the sb_coded_flag has a value of "1", and if no non-zero coefficients exist within the sub-block, the sb_coded_flag has a value of "0".
[0186] On the other hand, step S1435 is performed only when i indicates the last sub-block and the first sub-block. That is, since the last position coefficient is included in the last sub-block, the sb_coded_flag value parsed is a value of "1", and since the DC coefficient exists in the first sub-block, the sb_coded_flag value parsed is a value of "1".
[0187] Next, the video decoding device determines whether the condition "MtsZerooutFlag && luma block" is satisfied (S1440).
[0188] As a result of the determination in step S1440, if the condition "MtsZerooutFlag && luma block" is satisfied, it is further determined whether the condition "sb_coded_flag && (xSb>3 || ySb>3)" is satisfied (S1445). If the condition "sb_coded_flag && (xSb>3 || ySb>3)" is satisfied, MtsZerooutFlag is set to "0" (S1450).
[0189] As a result of the determination in step S1440, if the condition "MtsZerooutFlag && luma block" is not satisfied, the process for the sub-block is terminated.
[0190] According to this embodiment, when the variable i, that is, the MtsZerooutFlag value is set to "0" even once in the sub-block, in the next routine of the variable i - 1, a false value is derived in step S1940, and there is no need to further check the sb_coded_flag value.
[0191] On the one hand, when the block to be decoded satisfies all the conditions of (a) and (b) above, the use of explicit MTS is determined, and the conversion information actually used for the block is transmitted in index form (mts_idx). On the contrary, when not all conditions are satisfied, DCT-II which is the default conversion (*) is used (refer to Fig. 15). Table 5 shows the conversion types of the horizontal axis and the vertical axis according to the mts_idx value.
[0192]
Table 5
[0193] In Table 5, trTypeHor means the horizontal axis conversion type, and trTypeVer means the vertical axis conversion type. The values of the conversion types in Table 5 mean the trType values in Table 2. For example, when the mts_idx value is "2", DCT-8(2) may be used for the horizontal axis conversion and DST-7(1) may be used for the vertical axis conversion.
[0194] In the present invention, in all cases where the above-mentioned default conversion (*) DCT-II is used / executed / applied, it may be replaced by the expression "inducing the mts_idx value to be "0"". That is, when the mts_idx value is "0", DCT-II(0) is set for both the horizontal axis and the vertical axis conversions.
[0195] In the present invention, the binary method of mts_idx uses the TR (truncated rice) method, and the cMax value which is the parameter value for TR is "4", and the cRiceParam value is "0". Table 6 shows the codewords of the MTS index.
[0196]
Table 6
[0197] Referring to Table 6, when the mts_idx value is "0", the corresponding codeword is "0"; when the mts_idx value is "1", the corresponding codeword is "10"; when the mts_idx value is "2", the corresponding codeword is "110"; when the mts_idx value is "3", the corresponding codeword is "1110"; and when the mts_idx value is "4", the corresponding codeword is "1111".
[0198] FIG. 15 is a diagram showing a method for determining whether to apply an explicit MTS function according to an embodiment of the present invention. Each step in FIG. 15 may be performed in a video decoding device.
[0199] Referring to FIG. 15, the video decoding device determines whether the condition of "(sps_explicit_mts_intra_enabled_flag && CuPredMode = MODE_INTRA) || (sps_explicit_mts_inter_enabled_flag && CuPredMode = MODE_INTER)" is satisfied (S1510).
[0200] sps_explicit_mts_intra_enabled_flag is a flag indicating whether to use explicit MTS during intra prediction, and sps_explicit_mts_inter_enabled_flag is a flag indicating whether to use explicit MTS during inter prediction. sps_explicit_mts_intra_enabled_flag has a value of "1" when using explicit MTS during intra prediction, and a value of "0" otherwise. sps_explicit_mts_inter_enabled_flag has a value of "1" when using explicit MTS during inter prediction, and a value of "0" otherwise.
[0201] The CuPredMode indicates how the block to be decoded was encoded. If the block to be decoded was encoded using the intra prediction method, the CuPredMode has the value MODE_INTRA, and if the block to be decoded was encoded using the inter prediction method, the CuPredMode has the value MODE_INTER.
[0202] Therefore, when the block to be decoded uses intra prediction and explicit MTS, "sps_explicit_mts_intra_enabled_flag && CuPredMode = MODE_INTRA" has the value "1", and when the block to be decoded uses inter prediction and explicit MTS, "sps_explicit_mts_inter_enabled_flag && CuPredMode = MODE_INTER" has the value "1". Therefore, in step S2010, it is possible to determine whether the block to be decoded uses explicit MTS by checking the values of sps_explicit_mts_intra_enabled_flag, sps_explicit_mts_inter_enabled_flag, and CuPredMode.
[0203] When the condition of step S1510 is satisfied, the video decoder determines whether the condition "lfnst_idx = 0 && transform_skip_flag = 0 && cbW < 32 && cbH < 32 && intra_subpartitions_mode_flag = 0 && cu_sbt_flag = 0" is satisfied (S1520).
[0204] Here, the lfnst_idx value indicates the secondary transform, and when lfnst_idx = 0, it means that the secondary transform is not being used.
[0205] The transform_skip_flag value indicates whether transform skip is applied to the current block, i.e., whether to omit the transformation process for the current block. When transform_skip_flag = 0, it indicates that transform skip is not applied to the current block.
[0206] cbW and cbH indicate the width and height of the current block respectively. As described above, the maximum size supported by the default transformation DCT-II is up to 64×64, and the maximum sizes supported by the additional transformations DST-7 and DCT-8 are up to 32×32. For example, when the size of the block to be decoded is 64×64, one 64×64 DCT-II is applied to the transformation process. That is, if one or more of the width and height of the block to be decoded are greater than 32 (exceed 32), the default transformation (*) is applied immediately without applying MTS. Therefore, for MTS to be applied, all the values of cbW and cbH must have values of 32 or less.
[0207] The intra_subpartitions_mode_flag indicates whether the intra subpartition mode is applied. The intra subpartition mode is one of the in-picture prediction methods, which indicates dividing the target block into multiple sub-blocks and going through the prediction, transformation, and quantization processes. That is, when the flag value (intra_subpartitions_mode_flag) is "0", it means performing general in-picture prediction without dividing the target block into sub-blocks.
[0208] The cu_sbt_flag indicates whether sub-block transformation, where only a part of the target block goes through the transformation process, is applied. That is, when the cu_sbt_flag value is "0", it means that sub-block transformation where only a part of the target block goes through the transformation process is not applied.
[0209] Therefore, it may be determined whether an explicit MTS can be applied to the block to be decoded according to whether the condition of step S1520 is satisfied.
[0210] If the condition of step S1510 is not satisfied, the video decoding apparatus sets the value of mts_idx to "0" (S1530) and ends the process.
[0211] If the condition of step S1520 is satisfied, the video decoding apparatus determines whether the condition of "MtsZeroOutFlag = 1 && MtsDcOnlyFlag = 0" is satisfied (S1540).
[0212] The MtsZerooutFlag indicates whether non-zero coefficients in the block exist in the zero-out region. When at least one of the non-zero coefficients in the block exists in the zero-out region, the MtsZerooutFlag may have a value of "0", and when all of the non-zero coefficients in the block do not exist in the zero-out region, the MtsZerooutFlag may have a value of "1". At this time, the value of the MtsZerooutFlag may be determined by performing the process of FIG. 13 or FIG. 14.
[0213] The MtsDcOnlyFlag indicates whether the number of non-zero coefficients in the block is one and its position is DC. When the number of non-zero coefficients in the block is one and its position is DC, the MtsDcOnlyFlag has a value of "1", and in other cases, the MtsDcOnlyFlag has a value of "0". At this time, the value of the MtsDcOnlyFlag may be determined by performing the process of FIG. 17.
[0214] On the other hand, if the condition of step S1520 is not satisfied, the video decoding apparatus sets the value of mts_idx to "0" (S1530) and ends the process.
[0215] When the condition of step S1540 is satisfied, the video decoding device parses mts_idx (S1550) and ends the process. At this time, the conversion types for the horizontal axis and the vertical axis according to the value of mts_idx may be assigned according to Table 5. At this time, the value of the conversion type in Table 5 means the trType value in Table 2. For example, when the mts_idx value is "2", DCT-8 may be applied for the horizontal axis conversion and DST-7 may be applied for the vertical axis conversion.
[0216] Also, even when the condition of step S1540 is not satisfied, the video decoding device sets the value of mts_idx to "0" (S1530) and ends the process.
[0217] FIG. 16 is a diagram showing a method of performing inverse conversion based on conversion-related parameters according to another embodiment of the present invention. Each step in FIG. 16 may be performed by a video decoding device, for example, by an inverse conversion unit of the decoding device.
[0218] Referring to FIG. 16, the video decoding device acquires the values of sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, and cu_sbt_flag (S1610). At this time, what each of sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, and cu_sbt_flag indicates is described in detail in the related description of FIG. 15, and the parameters are used to determine whether the decoding target block can apply explicit MTS.
[0219] Next, the video decoding device acquires the MtsZerooutFlag and MtsDcOnlyFlag values (S1620). At this time, the MtsZerooutFlag can be acquired by performing the process of FIG. 13 or FIG. 14, and the MtsDcOnlyFlag can be acquired by performing the process of FIG. 12.
[0220] Next, the video decoding device acquires the mts_idx value based on the parameters acquired in step S1610 and step S1620 (S1630). That is, the video decoding device acquires the mts_idx value based on sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, cu_sbt_flag, MtsZerooutFlag, and MtsDcOnlyFlag. At this time, the mts_idx may be acquired by performing the process of FIG. 15.
[0221] Next, the video decoding device performs inverse transformation based on mts_idx (S1640). The inverse transformation applied by the mts_idx value may be configured according to Table 5 and Table 2. For example, when the mts_idx value is "2", DCT-8 may be applied in the horizontal axis direction and DST-7 may be applied in the vertical axis direction.
[0222] On the one hand, although not shown, from the perspective of a video encoding device, in order to set whether to use explicit MTS, sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, cu_sbt_flag, MtsZerooutFlag, and MtsDcOnlyFlag can be set. In the above-described embodiments, the method has been described based on a sequence diagram as a series of steps or blocks, but the present invention is not limited to the order of the steps, and a certain step may occur in a different order from or simultaneously with the steps different from the above. Also, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more of the steps in the sequence diagram may be deleted without affecting the scope of the present invention.
[0223] The embodiments described in this document may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., instruction information) or an algorithm may be stored in a digital storage medium.
[0224] In addition, the decoding device and encoding device to which the present invention is applied may be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video intercom device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a videophone video device, a transportation means terminal (e.g., a vehicle terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and may be used to process a video signal or a data signal. For example, as an over-the-top (OTT) video device, a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc. may be mentioned.
[0225] In addition, the processing method to which the present invention is applied may be produced in the form of a program executed by a computer, or may be stored in a computer-readable storage medium. Multimedia data having a data structure according to the present invention may also be stored in a computer-readable storage medium. The computer-readable storage medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored and memorized. The above computer-readable storage medium may include, for example, a Blu-ray Disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Further, the computer-readable storage medium includes a medium realized in the form of a carrier wave (e.g., transmission via the Internet). In addition, the bitstream generated by the encoding method may be stored in a computer-readable storage medium or transmitted via a wired or wireless communication network.
[0226] Further, an embodiment of the present invention may be realized by a computer program product with program code, and the program code may be executed by a computer according to an embodiment of the present invention. The program code may be stored on a computer-readable carrier.
Claims
1. A video decoding method performed by a video decoding device, comprising: obtaining a parameter related to whether a Multiple Transform Set (MTS) is applicable to a block to be decoded from a bitstream; determining a type of transform to be applied to the block to be decoded based on at least one of a parameter relating to whether a set of multiple transforms is applicable to the block to be decoded and a size of the block to be decoded; setting a validity region including non-zero coefficients of the block to be decoded based on at least one of a parameter relating to whether a multiple transformation set for the block to be decoded is applicable or not and a size of the block to be decoded; and Reconstructing the block to be decoded based on a transform type applied to the block to be decoded and a valid region including non-zero coefficients of the block to be decoded; When the size of the block to be decoded is any one of 64 width by 64 height, 64 width by 32 height, and 32 width by 64 height, a valid area including non-zero coefficients of the block to be decoded is set to a 32 width by 32 height area including an upper left sample in the block to be decoded, regardless of a value of a parameter related to whether a multiple transform set is applicable to the block to be decoded; When the size of the block to be decoded is any one of 32 width × 32 height, 32 width × 16 height, and 16 width × 32 height, a valid area including a non-zero coefficient of the block to be decoded is set to an area having a different size according to a value of the parameter related to whether a multiple transformation set for the block to be decoded is applicable or not; A video decoding method, characterized in that, when a parameter indicating a transform type to be applied to a block to be decoded is acquired from the bitstream, the transform type to be applied to the block to be decoded is determined by the parameter indicating the transform type to be applied to the block to be decoded.
2. A video encoding method performed by a video encoding device, comprising: determining a type of transform to be applied to the block to be coded according to whether a Multiple Transform Set (MTS) for the block to be coded is applicable; generating parameters indicating a type of transformation to be applied to the block to be coded; generating transform coefficients by performing a transform corresponding to a transform type on the current block to be coded based on a valid area including non-zero coefficients for the current block to be coded, the valid area being determined taking into account a transform type to be applied to the current block to be coded; generating a bitstream including at least one of a parameter related to whether a multiple transform set is applicable to the current block, a parameter indicating a transform type to be applied to the current block, and the transform coefficients; when the size of the encoding target block is any one of width 64×height 64, width 64×height 32, and width 32×height 64, a valid area including nonzero coefficients of the encoding target block is a width 32×height 32 area including an upper left sample in the encoding target block, A video coding method characterized in that, when the size of the block to be coded is any one of 32 width x 32 height, 32 width x 16 height, or 16 width x 32 height, a valid area including non-zero coefficients of the block to be coded is determined by the value of a parameter related to whether a multiple transformation set is applicable.
3. A method for transmitting a bitstream encoded by the video encoding method of claim 2, comprising the step of transmitting the bitstream.
Citation Information
Patent Citations
Method for encoding / decoding video signals and apparatus therefor
WO2020060364A1
Transform-based image coding method, and device therefor
WO2021066616A1
Methods and apparatus on transform and coefficient signaling
WO2021102424A1
Coding concepts for a transformed representation of a sample block
WO2021105255A1