Image decoding method and device
By obtaining and determining the conversion type parameters of the decoding object block during the image decoding process, the problem of difficulty in determining the optimal conversion type in the prior art is solved, and more efficient image encoding is achieved.
Patent Information
- Application Number
- CN202510346539.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-25
- Filing Date
- 2020-12-16
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively determine and use the optimal conversion type during image decoding, resulting in low encoding efficiency.
By obtaining the parameters, width and height information of the block that are applicable to the multiple conversion set of the decoding object block, judge and set the conversion type, and perform inverse conversion to improve encoding efficiency.
The inverse conversion method defined in advance according to specific conditions is implemented, which improves the compression performance of the decoding object block and enhances the encoding efficiency.
Smart Images

Figure CN119996710A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with application date of December 16, 2020, application number 202080050502.9, and title “Image decoding method and device”. Technical Field
[0002] The present invention relates to video coding technology, and more particularly to a method for determining a type of a primary transform of a decoding target block during an image decoding process. Background Art
[0003] Recently, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has increased in various fields. The higher the resolution and the higher the quality of image data, the greater the amount of information or bit volume to be transmitted compared to existing image data, and the higher the transmission and storage costs when storing image data using existing media such as wired / wireless broadband lines.
[0004] After HEVC video compression was formulated in 2013, the standardization of the next-generation video compression, Versatile Video Coding (VVC), which aims to improve the performance by more than 2 times compared to HEVC, has been actively carried out as 4K and 8K are expanded to real images and virtual reality services that use video images. VVC is being developed in the Joint Video Exploration Team (JVET), which is a joint video coding standardization group, the Moving Picture Experts Group (ISO / ICEMPEG) and the Video Coding Experts Group (ITU-TVCEG), with the goal of improving the coding compression performance by more than 2 times compared to HEVC. In January 2018, the Call for Proposal was published at the 121st Gwangju MPEG and the 9th JVET meeting. A total of 23 organizations proposed video compression technologies at the 122nd San Diego MPEG and the 10th JVET meeting, and formal video standardization began. At 122 MPEG and 10 JVET meetings, technical research on video compression technologies proposed by various organizations and evaluation of objective compression performance and subjective image quality were carried out, and Working Draft (WD) 1.0 and VTM (VVC Test Mode) 1.0, a video reference software, were published by adopting a part of each technology. After the 127th MPEG and 15 JVET meetings ended in July 2019, the VVC standard was completed as a committee draft (CD), and standardization is being carried out with the goal of designating the final draft international standard (FDIS: Final Draft International Standard) in October 2020.
[0005] In the existing HEVC, in the coding structure divided by quadtree layers, VVC adopts a block structure that combines quadtree + binary tree (QTBT: Quad Tree Binary Tree) and ternary tree (TT: Ternary Tree). Compared with HEVC, it generates or processes prediction residual signals more flexibly and has higher compression performance than HEVC. In addition to this basic block structure, new technologies that are not used in existing compression, such as adaptive loop filtering technology (ALF: Adaptive Loop Filter), motion prediction technology, AMP (Affine Motion Prediction), and decoder-side motion vector correction (DMVR: Decoder-side Motion Vector Refinement) are adopted as standard technologies. As a conversion and quantization technology, the conversion kernel DCT-II, which is mainly used in existing video compression, continues to be used, and the block size is changed to a larger block size than the applicable block size. Furthermore, in the conventional HEVC, the DCT-7 kernel applicable to small transform blocks such as 4×4 is extended to large transform blocks until a new transform kernel, namely, DCT-8, is added as a transform kernel.
[0006] In addition, in the HEVC standard, when encoding or decoding an image, a transformation is performed using one transformation type, and there is no need to transmit information about the transformation type of the image. However, in the new technology, multiple transformation selection (Multiple Transform Selection) using DCT-II, DCT-8, and DCT-7 is applied. During decoding, it is actually required to define whether MTS is applicable and whether any transformation type is applicable. Summary of the invention
[0007] Technical problem to be solved by the invention The technical problem of the present invention is to perform the inverse conversion by a method defined in advance according to specific conditions.
[0008] Another technical problem of the present invention is to apply an optimal transformation type to a decoding target block and perform decoding.
[0009] Technical solutions to solve problems According to one aspect of the present invention, a video decoding method performed by a video decoding device is provided. The video decoding method includes the following steps: obtaining a parameter indicating whether a multiple transform set (MTS) is applicable to a decoding object block, the width of the decoding object block, and information about the height of the decoding object block; determining a transformation type of the decoding object block based on at least one of the information indicating whether the multiple transform set (MTS) is applicable to the decoding object block, the width of the decoding object block, and the height of the decoding object block; setting a zeroing area of the decoding object block based on at least one of the information indicating whether the multiple transform set (MTS) is applicable to the decoding object block, the width of the decoding object block, and the height of the decoding object block; and performing an inverse transformation of the decoding object block based on the determination result of the zeroing area and the transformation type of the decoding object block.
[0010] According to another aspect of the present invention, in the step of determining the conversion type of the decoding object block, in the case where at least one of the width or height of the decoding object block is greater than 32, the decoding object block is determined to be converted by using a default conversion.
[0011] According to another aspect of the present invention, in the step of setting the zeroing area of the decoding object block, in the case where one of the width or height of the decoding object block has a value greater than 32, the area where the width or height of the decoding object block is greater than 32 is set as the zeroing area.
[0012] According to still another aspect of the present invention, the parameter indicating whether the multiple transform set of the decoding target block is applicable is sps_mts_enabled_flag.
[0013] According to another aspect of the present invention, there is provided an image decoding method performed by an image decoding device. The image decoding method comprises the following steps: obtaining at least one of information on whether a multiple transform set (MTS) of a decoding object block is applicable, information on a prediction mode, information on whether a secondary transform is applicable, information on whether a prediction using rows and columns is applicable, and information on the size of the decoding object block; judging whether the multiple transform set is applicable to the decoding object block by default based on at least one of information on whether the multiple transform set is applicable, information on the prediction mode, information on whether a secondary transform is applicable, and information on whether a prediction using rows and columns is applicable to the decoding object block; obtaining information on a transform type based on information on whether the default multiple transform set is applicable to the decoding object block and information on the size of the decoding object block; and performing inverse transform based on the information on the transform type.
[0014] According to another aspect of the present invention, the step of determining whether the default multiple transform set is applicable utilizes information on whether the multiple transform set (MTS) of the decoding object block is applicable, information on the prediction mode, information on whether the secondary transform is applicable, and information on whether the prediction using rows and columns is applicable to determine whether the default multiple transform set is applicable.
[0015] According to yet another aspect of the present invention, the default multiple transformation set includes a default transformation (defaulttransform) and at least one extra transformation (extra transform).
[0016] According to another aspect of the present invention, in the step of obtaining the transformation type information based on the size information of the decoding object block, when the horizontal axis lengths of the decoding object block are all 4 to 16, the decoding object block applies at least one of the extra transform types for the horizontal axis direction.
[0017] According to another aspect of the present invention, in the step of obtaining the transformation type information based on the size information of the decoding object block, when the vertical axis lengths of the decoding object block are all 4 to 16, the decoding object block applies at least one of the extra transformation types for the vertical axis direction.
[0018] According to still another aspect of the present invention, the information on whether a multiple transform set (MTS) of the decoding target block is applicable includes at least one of sps_mts_enabled_flag and sps_explicit_mts_intra_enabled_flag.
[0019] According to still another aspect of the present invention, the prediction mode information includes CuPredMode.
[0020] According to yet another aspect of the present invention, the information on whether the secondary conversion is applicable includes lfnst_idx.
[0021] According to still another aspect of the present invention, the information on whether prediction using the row or column is applicable includes intra_mip_flag.
[0022] According to still another aspect of the present invention, the information on the transformation type of the decoding target block includes information on each horizontal-axis transformation type and information on a vertical-axis transformation type.
[0023] According to still another aspect of the present invention, the step of determining whether a default multiple transform set is applicable to the decoding target block is achieved by additionally confirming whether the decoding target block is a luma block.
[0024] According to another aspect of the present invention, there is provided an image decoding device including a memory and at least one processor. The image decoding device includes at least one processor including an inverse transform unit, which obtains at least one of information on whether a multiple transform set (MTS) of a decoding object block is applicable, information on a prediction mode, information on whether a secondary transform is applicable, information on whether a prediction using rows and columns is applicable, and information related to the size of the decoding object block, and determines whether a default multiple transform set is applicable to the decoding object block based on at least one of the information on whether the multiple transform set of the decoding object block is applicable, information on the prediction mode, information on whether a secondary transform is applicable, and information on whether a prediction using rows and columns is applicable, and obtains information related to a transform type based on the information on whether the default multiple transform set is applicable to the decoding object block and information related to the size of the decoding object block, and performs inverse transform based on the information on the transform type.
[0025] Effects of the Invention According to the present invention, the inverse conversion can be performed in a method defined in advance through specific conditions.
[0026] Furthermore, an optimum transform type is applied to the decoding target block and decoding is performed, thereby increasing the effect of improving the compression performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A diagram schematically showing the structure of a video encoding device to which the present invention is applied.
[0028] Figure 2 A diagram showing an example of a video encoding method performed by a video encoding device.
[0029] Figure 3 A diagram schematically showing the structure of a video decoding device to which the present invention is applied.
[0030] Figure 4 A diagram showing an example of a video decoding method performed by a decoding device.
[0031] Figure 5 A diagram showing the scanning order of sub-blocks and parameters in a diagonal scanning method.
[0032] Figure 6 This is a diagram showing an example of a quantized 32×32 encoding target block.
[0033] Figure 7 This is a diagram showing a zero-out area other than the mxn area in the MxN decoding target block.
[0034] Figure 8 The figure shows a method for determining whether the default MTS function is applicable or not according to an embodiment of the present invention.
[0035] Fig. 9 The diagram shows a method of converting information of width and height of a corresponding block of a default MTS according to an embodiment of the present invention.
[0036] Fig.10 FIG. 4 is a diagram showing a method for performing an inverse transformation based on transformation-related parameters according to an embodiment of the present invention.
[0037] Fig.11 This figure shows a valid MTS area marked by a bold line in a 32×32 decoding target block.
[0038] Fig.12 This is a diagram showing a method of determining a valid MTS according to an embodiment of the present invention.
[0039] Fig.13 This is a diagram showing a method of determining a valid MTS according to another embodiment of the present invention.
[0040] Fig.14 This is a diagram showing a method of determining a valid MTS according to still another embodiment of the present invention.
[0041] Fig.15 This is a diagram showing a method for determining whether the explicit MTS function is applicable or not according to an embodiment of the present invention.
[0042] Fig.16FIG. 1 is a diagram showing a method for performing inverse conversion of conversion-related parameters according to another embodiment of the present invention. DETAILED DESCRIPTION
[0043] The present invention is subject to various changes and has various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, the present invention is not limited to specific embodiments. The terms used in this specification are used only to illustrate specific embodiments and are not intended to limit the technical ideas of the present invention. Expressions in the singular that are not clearly defined differently in the context include plural expressions. In this specification, the terms "including" or "having" specify the existence of features, numbers, steps, actions, constituent elements, parts or features, numbers, steps, actions, constituent elements, and combinations of parts recorded in the specification, and should be understood as not excluding in advance the existence or additional possibility of one or more other features or numbers, steps, actions, constituent elements, parts or features, numbers, steps, actions, constituent elements, and combinations of parts.
[0044] In addition, the various structures on the drawings described in the present invention are independently displayed for the convenience of explaining the functions of different features, and the various structures are not implemented by separate hardware or other software. For example, two or more structures in each structure can be combined to form a structure, and a structure can also be divided into multiple structures. The embodiments of each structure in combination and / or separation do not deviate from the essence of the present invention, but are included in the scope of the claims of the present invention.
[0045] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. The same reference numerals are used for the same components in the drawings, and repeated description of the same components is omitted.
[0046] In addition, the present invention relates to video / image coding. For example, the method / embodiment disclosed in the present invention is applicable to methods disclosed in VVC (versatile video coding) standard, EVC (Essential Video Coding) standard, AV1 (AOMediaVideo 1) standard, AVS2 (2nd generation of audio video coding standard) or new generation video / image coding standards (for example, H.267, H.268, etc.).
[0047] In this specification, a picture generally refers to a unit that displays an image in a specific period of time, and a slice refers to a unit that constitutes a part of a picture during encoding. A picture is composed of multiple slices, and pictures and slices can be mixed and used as needed.
[0048] A pixel or a picture element (pel) is the smallest unit that makes up an image (or video). In addition, a "sample" is used as a term corresponding to a pixel. A sample generally displays a pixel or a pixel value, but it can also display only the pixel / pixel value of the brightness (luma) component, or it can display only the pixel / pixel value of the chromaticity (chroma) component.
[0049] Unit represents the basic unit of image processing. A unit includes at least one of a specific area of an image and information related to the corresponding area. Unit is used interchangeably with terms such as block or area depending on the situation. In general, an MxN block represents a set of samples or transform coefficients consisting of M columns and N rows.
[0050] Figure 1 A diagram briefly illustrating the structure of a video encoding device to which the present invention is applied.
[0051] Reference Figure 1 The video encoding device 100 includes: an image segmentation unit 105, a prediction unit 110, a residual processing unit 120, an entropy coding unit 130, an addition unit 140, a filter unit 150 and a memory 160. The residual processing unit 120 includes: a subtraction unit 121, a conversion unit 122, a quantization unit 123, a rearrangement unit 124, an inverse quantization unit 125 and an inverse conversion unit 126.
[0052] The image segmentation unit 105 segments the input image into at least one processing unit.
[0053] For example, the processing unit is called a coding unit (CU). In this case, the coding unit is recursively split from the coding tree unit (Coding Tree Unit) according to the QTBT (Quad-treebinary-tree) structure. For example, a coding tree unit is split into multiple coding units of a lower (deeper) depth based on a quadtree structure and / or a binary tree structure. In this case, for example, the quadtree structure is first applied, and the binary tree structure is finally applied. Or the binary tree structure can also be applied first. Based on the final coding unit that cannot be split any further, the encoding procedure of the present invention is executed. In this case, based on the encoding efficiency of the image characteristics, etc., the maximum coding unit is directly used as the final coding unit, or as needed, the coding unit is recursively split into coding units of a lower depth, and the coding unit of the most suitable size is used as the final coding unit. Here, the so-called encoding procedure includes the following prediction, conversion and recovery procedures.
[0054] In another example, the processing unit also includes: a coding unit (CU), a prediction unit (PU) or a transform unit (TU). The coding unit is split into deeper coding units from the largest coding unit (LCU) according to the quadtree structure. In this case, based on the coding efficiency according to the image characteristics, the largest coding unit is directly used as the final coding unit, or the coding unit is recursively split into deeper coding units as needed, and the coding unit of the optimal size is used as the final coding unit. In the case of setting the smallest coding unit (SCU), the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the so-called final coding unit refers to a coding unit based on partitioning or partitioning into a prediction unit or a transform unit. The prediction unit is a unit of sample prediction as a unit partitioning from the coding unit. At this time, the prediction unit can also be divided into sub-blocks. The conversion unit is divided from the coding unit according to the quadtree structure, and is a unit that guides the conversion parameter and / or a unit that guides the residual signal from the conversion parameter. Hereinafter, the coding unit is referred to as a coding block (CB), the prediction unit is referred to as a prediction block (PB), and the conversion unit is referred to as a transform block (TB). A prediction block or a prediction unit refers to a specific block-shaped area within an image, including an array of prediction samples. Furthermore, a transform block or a transform unit refers to a specific block-shaped area within an image, including an array of conversion parameters or residual samples.
[0055] The prediction unit 110 performs prediction on a processing target block (hereinafter referred to as a current block) to generate a predicted block (predicted block) including prediction samples of the current block. The unit of prediction performed in the prediction unit 110 is a coding block, which may also be a conversion block or a prediction block.
[0056] The prediction unit 110 determines whether to apply intra prediction or inter prediction in the current block. For example, the prediction unit 110 determines whether to apply intra prediction or inter prediction in units of CUs.
[0057] In the case of intra-frame prediction, the prediction unit 110 guides the prediction samples of the current block based on the reference samples outside the current block in the image to which the current block belongs (hereinafter referred to as the current image). At this time, the prediction unit 110 (i) guides the prediction samples based on the average or interpolation of the neighboring reference samples of the current block, and (ii) for the prediction samples in the neighboring reference samples of the current block, it is also possible to guide the prediction samples based on the reference samples existing in a specific (prediction) direction. The case of (i) is called a non-directional mode or a non-angular mode, and the case of (ii) is called a directional mode or an angular mode. In intra-frame prediction, the prediction mode has, for example, 33 directional prediction modes and at least two non-directional modes. The non-directional mode includes a DC prediction mode and a planar mode. The prediction unit 110 can also determine the prediction mode applicable to the current block by using the prediction mode applicable to the neighboring blocks.
[0058] In the case of inter-frame prediction, the prediction unit 110 can guide the prediction sample of the current block based on the sample specified by the motion vector on the reference image. The prediction unit 110 applies any one of the skip mode, merge mode and motion vector prediction (MVP: motion vector prediction) mode to guide the prediction sample of the current block. In the case of skip mode and merge mode, the prediction unit 110 uses the motion information of the surrounding blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, the residual between the prediction sample and the original sample is not transmitted. In the case of MVP mode, the motion vector of the surrounding block is used as a motion vector predictor (Motion Vector Predictor), and the motion vector predictor of the current block is used to guide the motion vector of the current block.
[0059] In the case of inter-frame prediction, the neighboring blocks include spatial neighboring blocks in the current image and temporal neighboring blocks in the reference picture. The reference picture containing the temporal neighboring blocks is also called a collocated picture (colPic). Motion information includes motion vectors and reference picture indexes. Information such as prediction mode information and motion information is (entropy) encoded and output in the form of a bitstream.
[0060] In skip mode and merge mode, the top picture on the reference picture list can also be used as a reference picture when using the temporal motion information of surrounding blocks. The reference pictures included in the reference picture list are arranged based on the POC (Picture Order Count) difference between the current picture and the corresponding reference picture. POC corresponds to the display order of the pictures and is distinguished from the encoding order.
[0061] The subtraction unit 121 generates residual samples which are differences between original samples and predicted samples. When the skip mode is applied, the residual samples are not generated as described above.
[0062] The conversion unit 122 converts the residual samples in units of conversion blocks and generates a conversion parameter (transform coefficient). The conversion unit 122 performs the conversion according to the size of the corresponding conversion block, the coding block spatially overlapping with the corresponding conversion block, or the prediction mode applicable to the prediction block. For example, in the case where intra prediction is applied in the coding block or the prediction block overlapping with the conversion block, and the conversion block is a 4×4 residual array, the residual samples are converted using a DST (Discrete Sine Transform) conversion kernel. For other cases, the residual samples are converted using a DCT (Discrete Cosine Transform) conversion kernel.
[0063] The quantization unit 123 quantizes the conversion parameter to generate a quantized conversion parameter.
[0064] The rearrangement unit 124 rearranges the quantized conversion parameters. The rearrangement unit 124 rearranges the block-shaped quantized conversion parameters in a one-dimensional vector form by a scanning parameter method. Here, the rearrangement unit 124 is described with a separate structure, but the rearrangement unit 124 is a part of the quantization unit 123.
[0065] The entropy coding unit 130 performs entropy coding on the quantized transformation parameters. Entropy coding includes, for example, coding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition to the quantized transformation parameters, the entropy coding unit 130 can also encode information required for video recovery (for example, the value of a syntax element) together or separately. The entropy coded information is in the form of a bit stream and is transmitted or stored in units of NAL (network abstraction layer).
[0066] The inverse quantization unit 125 inversely quantizes the value quantized by the quantization unit 123 (quantized conversion parameter), and the inverse conversion unit 126 inversely converts the value inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0067] The adding unit 140 combines the residual samples and the prediction samples to restore the image. The residual samples and the prediction samples are added in units of blocks to generate a restored block. Here, the adding unit 140 is described as a separate structure, but the adding unit 140 is a part of the prediction unit 110. In addition, the adding unit 140 can also be called a restoration unit or a restoration block generation unit.
[0068] For the reconstructed picture, the filter unit 150 applies deblocking filtering and / or sample adaptive offset. The artifacts of block boundaries in the reconstructed picture or the distortion of the quantization process are corrected by deblocking filtering and / or sample adaptive offset. Sample adaptive offset is applied in units of samples and is applied after the deblocking filtering process is completed. The filter unit 150 can also apply to the reconstructed ALF (Adaptive Loop Filter) image. ALF is applied to the reconstructed picture after the deblocking filter and / or sample adaptive offset are applied.
[0069] The memory 160 stores the restored image (decoded image) or information required for encoding / decoding. Here, the restored image is a restored image that has been filtered through the filter unit 150. The stored restored image is used as a reference image (inter-frame) for predicting other images. For example, the memory 160 stores a (reference) image used in inter-frame prediction. At this time, the image used in the inter-frame prediction is specified by a reference picture set or a reference picture list.
[0070] Figure 2 An example of a video encoding method performed by a video encoding device is shown. Figure 2 , the image encoding method includes: block partitioning, intra / inter prediction, transform, quantization and entropy encoding processes. For example, the current image is divided into multiple blocks, and a prediction block of the current block is generated by intra / inter prediction, and a residual block of the current block is generated by subtracting the input block of the current block and the prediction block. Afterwards, a parameter block, that is, a transformation parameter of the current block, is generated by transforming the residual block. The transformation parameters are quantized and entropy encoded and stored in the bit stream.
[0071] Figure 3 A diagram schematically illustrating the structure of a video decoding device to which the present invention is applied.
[0072] Reference Figure 3 The video decoding device 300 includes an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 includes a rearrangement unit 321, an inverse quantization unit 322, and an inverse conversion unit 323.
[0073] When a bit stream including video information is input, the video decoding device 300 restores the video in accordance with a program for processing the video information in the video encoding device.
[0074] For example, the video decoding device 300 performs video decoding using a processing unit applicable to a video encoding device. Therefore, the processing unit block of the video decoding is, for example, a coding unit, another example is a coding unit, a prediction unit, or a conversion unit. The coding unit is divided from the maximum coding unit according to a quadtree structure and / or a binary tree structure.
[0075] The prediction unit and the conversion unit are further applicable according to this case, for which the prediction block is a block derived or partitioned from the coding unit, and is a unit of sample prediction. In this case, the prediction unit can also be divided into sub-blocks. The conversion unit is a unit that is split from the coding unit according to the quadtree structure and leads to a unit of a residual signal from a unit that guides a conversion parameter or a conversion parameter.
[0076] The entropy decoding unit 310 parses the bitstream and outputs information required for video restoration or image restoration. For example, the entropy decoding unit 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of syntax elements required for video restoration and the quantized values of conversion parameters related to the residual.
[0077] More specifically, the CABAC entropy decoding method receives a binary corresponding to each sentence element from a bitstream, determines a context model using information about the decoding object sentence element and decoding information of surrounding and decoding object blocks or information about symbols / bins decoded in the previous step, and predicts the probability of occurrence of a binary (bin) based on the determined context model, performs arithmetic decoding of the binary, and generates a symbol corresponding to the value of each sentence element. At this time, after determining the context model, the CABAC entropy decoding method updates the context model using information about symbols / bins decoded using the context model for symbols / bins.
[0078] Information related to prediction among the information decoded in the entropy decoding unit 310 is provided to the prediction unit 330 , and residual values entropy-decoded in the entropy decoding unit 310 , that is, quantized conversion parameters, are input to the rearrangement unit 321 .
[0079] The rearrangement unit 321 rearranges the quantized conversion parameters in a two-dimensional block. The rearrangement unit 321 performs the rearrangement in accordance with the parameter scan performed in the encoding device. Here, the rearrangement unit 321 is described as a separate structure, but the rearrangement unit 321 is a part of the inverse quantization unit 322.
[0080] The inverse quantization unit 322 inversely quantizes the quantized conversion parameter based on the (inverse) quantization parameter and outputs the conversion parameter. At this time, information for guiding the quantization parameter is signaled by the encoding device.
[0081] The inverse transformation unit 323 performs inverse transformation on the transformation parameters to induce residual samples.
[0082] The prediction unit 330 performs prediction on the current block and generates a predicted block including prediction samples of the current block. The unit of the prediction performed in the prediction unit 330 can be a coding block, a transform block, or a prediction block.
[0083] The prediction unit 330 determines whether to apply intra-frame prediction or inter-frame prediction based on the predicted information. At this time, the unit for determining whether intra-frame prediction or inter-frame prediction is applicable is different from the unit for generating prediction samples. Moreover, the unit for generating prediction samples is also different for inter-frame prediction and intra-frame prediction. For example, it is determined by the CU unit to which either inter-frame prediction or intra-frame prediction is applied. Moreover, for example, for inter-frame prediction, the prediction mode is determined by the PU unit and the prediction sample is generated, and for intra-frame prediction, the prediction mode is determined by the PU unit, and the prediction sample can also be generated by the TU unit.
[0084] In the case of intra-frame prediction, the prediction unit 330 guides the prediction samples of the current block based on the surrounding reference samples in the current image. The prediction unit 330 applies a directional mode or a non-directional mode based on the surrounding reference samples of the current block, thereby guiding the prediction samples of the current block. At this time, the prediction mode applicable to the current block can also be determined by using the intra-frame prediction mode of the surrounding blocks. In addition, matrix-based intra-frame prediction (MIP: Matrix-based Intra Prediction) is used to perform prediction based on pre-trained rows and columns. In this case, the number of MIP modes and the size of rows and columns are defined according to the size of each block. After the reference samples are downsampled to match the size of the rows and columns, the rows and columns determined are multiplied by the mode number to meet the prediction block size and interpolate to generate prediction values.
[0085] In the case of inter-frame prediction, the prediction unit 330 guides the prediction sample of the current block based on the sample specified on the reference image through the motion vector on the reference image. The prediction unit 330 applies any one of the skip mode, merge mode and MVP mode to guide the prediction sample of the current block. At this time, the motion information required for the inter-frame prediction of the current block provided by the video encoding device, such as the motion vector, the reference image index and other related information, is obtained or guided based on the prediction-related information.
[0086] In the case of skip mode and merge mode, the motion information of the neighboring blocks is used as the motion information of the current block. At this time, the neighboring blocks include spatial neighboring blocks and temporal neighboring blocks.
[0087] The prediction unit 330 forms a merge alternative list from the motion information of the available surrounding blocks, and uses the information indicated by the merge index on the merge alternative list as the motion vector of the current block. The merge index is signaled by the encoding device. The motion information includes a motion vector and a reference image. When the motion information of the temporal surrounding blocks is used in the skip mode and the merge mode, the highest image on the reference image list is used as the reference image.
[0088] In the case of skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.
[0089] In the case of the MVP mode, the motion vector of the surrounding blocks is used as a motion vector predictor to guide the motion vector of the current block. In this case, the surrounding blocks include spatial surrounding blocks and temporal surrounding blocks.
[0090] For example, in the case where the merge mode is applied, a merge alternative list is generated using the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks, i.e., Col blocks. In the merge mode, the motion vector of the substitute block selected from the merge alternative list is used as the motion vector of the current block. The prediction-related information includes a merge index indicating a substitute block having the best motion vector selected from the substitute blocks included in the merge alternative list. At this time, the prediction unit 330 uses the merge index to obtain the motion vector of the current block.
[0091] In another example, in the case of applying the MVP (Motion vector prediction) mode, a motion vector prediction substitute list is generated using the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks, i.e., Col blocks. That is, the motion vectors corresponding to the restored spatial neighboring blocks and / or the temporal neighboring blocks, i.e., Col blocks, are used as motion vector substitutes. The information related to the prediction includes a predicted motion vector index indicating the best motion vector selected from the motion vector substitutes contained in the list. At this time, the prediction unit 330 uses the motion vector index to select the predicted motion vector of the current block from the motion vector substitutes contained in the motion vector substitute list. The prediction unit of the encoding device seeks the motion vector difference (MVD) between the motion vector of the current block and the motion vector prediction amount, encodes it, and outputs it in the form of a bit stream. That is, the MVD is calculated by subtracting the motion vector prediction amount from the motion vector of the current block. At this time, the prediction unit 330 obtains the motion vector difference contained in the information related to the prediction, and derives the motion vector of the current block by adding the motion vector difference and the motion vector prediction amount. The prediction unit acquires or guides a reference picture index or the like of the reference picture from the information related to the prediction.
[0092] The adding unit 340 adds the residual sample and the prediction sample to restore the current block or the current image. The adding unit 340 can also restore the current image by adding the residual sample and the prediction sample in units of blocks. When the skip mode is applied, the residual is not transmitted, so the prediction sample is a restored sample. Here, the adding unit 340 is described as a separate structure, but the adding unit 340 can also be a part of the prediction unit 330. In addition, the adding unit 340 can also be called a restoration unit or a restoration block generation unit.
[0093] The filter unit 350 applies deblocking filtering, sample adaptive offset, and / or ALF to the restored image. In this case, sample adaptive offset is also applied in units of samples, and can also be applied after deblocking filtering. ALF can also be applied after deblocking filtering and / or sample adaptive offset.
[0094] The memory 360 stores the restored image (decoded image) or information required for decoding. Here, the restored image is a restored image after the filtering process is completed by the filter unit 350. For example, the memory 360 stores an image used for inter-frame prediction. At this time, the image used for inter-frame prediction can also be specified by a reference image set or a reference image list. The restored image is used as a reference image of other images. In addition, the memory 360 can also output the restored image according to the output order.
[0095] Figure 4 An example of a video decoding method performed by a decoding device is shown. Figure 4 , the image decoding method includes: entropy decoding, inverse quantization, inverse transform and intra / inter prediction process. For example, in a decoding device, the inverse process of the encoding method is performed. Specifically, the quantized transformation parameters are obtained by entropy decoding of the bit stream, and the parameter block of the current block, that is, the transformation parameters, is obtained by the inverse quantization process of the quantized transformation parameters. The residual block of the current block is derived by the inverse transformation of the transformation parameters, and the recovery block of the current block is derived by adding the prediction block of the current block derived by intra / inter prediction and the residual block.
[0096] In addition, keywords in the embodiments described below are defined as described in the following table.
[0097]
Table 1
[0098] Referring to Table 1, Floor(x) shows the largest integer value less than or equal to x, Log2(u) shows the logarithm of u to the base 2, and Ceil(x) shows the smallest integer value greater than or equal to x. For example, in the case of Floor(5.93), since the largest integer value less than 5.93 is 5, it shows 5.
[0099] Also, referring to Table 1, x>>y shows the operator that right-shifts x along the y-axis, and x<<y shows the operator that left-shifts x along the y-axis.
[0100] <Import> The HEVC standard generally uses one transform type, namely the discrete cosine transform (DCT). Therefore, there is no need to transmit the information of another determination process for the transform type and the determined transform type. However, when the size of the current luma block is 4x4 and for the case of performing intra prediction, the DST (discrete sine transform) transform type is used exceptionally.
[0101] The information indicating the positions of non-zero parameters in the quantized coefficients after the transformation and quantization processes is roughly classified into three types.
[0102] 1. The position (x, y) of the last significant coefficient: The position of the last non-zero parameter (coefficient) in the scan order within the coding target block (hereinafter defined as the last position) 2. Coded sub-block flag: A flag that divides the coding target block into multiple sub-blocks and indicates whether each sub-block contains one or more non-zero parameters (or a flag for all-zero coefficients) 3. Significant coefficient flag: A flag that indicates whether each parameter within a sub-block is non-zero or zero Here, the position of the last significant coefficient is displayed by separating the x-axis component and the y-axis component, and each component is separated by a prefix and a suffix. That is, the syntax for notifying the non-zero position of the quantized parameter includes the following 6 syntaxes.
[0103] 1.last_sig_coeff_x_prefix, 2.last_sig_coeff_y_prefix 3.last_sig_coeff_x_suffix, 4.last_sig_coeff_y_suffix 5.coded_sub_block_flag 6.sig_coeff_flag The last_sig_coeff_x_prefix indicates the prefix of the x-axis component indicating the position of the last valid parameter, and the last_sig_coeff_y_prefix indicates the prefix of the y-axis component indicating the position of the last valid parameter. In addition, last_sig_coeff_x_suffix indicates the suffix of the x-component indicating the position of the last valid parameter, and last_sig_coeff_y_suffix indicates the suffix of the y-component indicating the position of the last valid parameter.
[0104] In addition, coded_sub_block_flag is displayed as "0" when all parameters in the corresponding sub-block are all zero (allzero), and is displayed as "1" when there are more than one non-zero (non-zero) parameters. For the case where sig_coeff_flag is a zero (zero) parameter, it is displayed as "0", and is displayed as "1" when it is a non-zero (non-zero) parameter. The coded_sub_block_flag syntax is transmitted for the sub-block that exists before in the scan order based on the last significant parameter position in the encoding target block. When coded_sub_block_flag is "1", that is, when there are more than one non-zero (non-zero) parameters, the sig_coeff_flag syntax of all parameters in the corresponding sub-block is transmitted.
[0105] The HEVC standard supports the following three forms of scanning for parameters.
[0106] 1) Up-right diagonal 2) Horizontal 3) Vertical When the encoding target block is encoded using the inter-picture prediction method, the parameters of the corresponding block are scanned in an up-right diagonal manner. When the block is encoded using the intra-picture prediction method, one of the three forms is selected according to the intra-picture prediction mode to scan the parameters of the corresponding block.
[0107] That is, when the image encoding device uses the inter-picture prediction method to encode the encoding target block, the parameters of the corresponding block are scanned in an up-right diagonal manner. When the encoding of the encoding target block uses the intra-picture prediction method, the image encoding device selects one of the three forms according to the intra-picture prediction mode and scans the parameters of the corresponding block. The scanning is performed in Figure 1 In the image encoding device, it is executed by the rearrangement unit 124, and the two-dimensional block form parameters are changed into a one-dimensional vector form by scanning.
[0108] Figure 5 Displays the scanning order of sub-blocks and parameters by diagonal scanning.
[0109] Reference Figure 5 , Figure 5 When the blocks are scanned by the rearrangement unit 124 of the image encoding device in a diagonal scanning manner, the scanning is performed from the upper leftmost sub-block, that is, sub-block No. 1, in the downward direction and the diagonal upper direction, and the scanning is performed last for sub-block No. 16 at the lower right end. That is, the rearrangement unit 124 performs scanning in the order of sub-blocks No. 1, 2, 3, ..., 14, 15, and 16 to rearrange the quantized conversion parameters in the two-dimensional block form in the form of a one-dimensional vector. Similarly, the rearrangement unit 124 of the image encoding device scans the parameters in each sub-block in the same diagonal scanning method as the scanning method of the sub-block. For example, in sub-block No. 1, the scanning is performed in the order of parameters No. 0, 1, 2, ..., 13, 14, and 15.
[0110] However, when the parameters for performing the scan are stored in the bit stream, the order in which they are stored is in the reverse order of the scan order. That is, Fig.10When the block is scanned by the rearrangement unit 124 of the image encoding device, the scanning is performed in the order from parameter No. 0 to parameter No. 255, but the order in which each pixel is stored in the bit stream is stored in the bit stream in the order from the pixel at position 255 to the pixel at position 0.
[0111] Figure 6 An example of a 32×32 encoding target block quantized by a video encoding device is shown. Figure 6 The 32×32 blocks shown are arbitrarily scanned diagonally by the video encoding device. Figure 6 In the example, pixels indicated by diagonal lines indicate non-zero parameters, and pixels indicated by x indicate the last significant parameter. All other white parameters have a value of zero. Figure 6 When scanning a block, the 24 sub-blocks that exist before the last position in the scanning order in the total 64 sub-blocks are scanned, that is, Figure 6 The coded_sub_block_flag information of the sub-block needs to be displayed in bold. The coded_sub_block_flag values of the first sub-block containing the DC value and the 24th sub-block containing the last position parameter among the 24 sub-blocks are led by "1", and the coded_sub_block_flag values of the remaining 22 sub-blocks are transmitted to the image decoding device through the bit stream. At this time, in the case where the 22 sub-blocks contain a sub-block with more than one non-zero coefficient, the coded_sub_block_flag value is set to "1" by the image encoding device. Figure 6 , among the 22 sub-blocks excluding the first sub-block and the 24th sub-block, the coded_sub_block_flag values of the 4th, 5th, 11th, and 18th sub-blocks including pixels marked in gray are set to "1".
[0112] 1. Method for determining the primary transform type of the decoding target block In this specification, a method for determining the type (type) of a primary transform (primary transform) of a decoding target block in the decoding process of an image is disclosed. That is, when the decoding target block is decoded by an image decoding device, in the conversion process of the image encoding device, it is necessary to determine whether to perform primary transform according to any transform type (type) for encoding. The primary transform type (primary transform type) is composed of a default transform (default transform) and a plurality of extra transforms (extra transforms). The decoding target block uses the default transform (default transform) or a multiple transform set (multiple transform set; MTS) including the default transform (default transform) and extra transforms (extra transforms) according to conditions. That is, the decoding target block is transformed using only the default transform or using the multiple transform set (MTS) including the default transform and the extra transform in the transformation process. From the perspective of the image decoding device, it is understood whether the decoding target block is decoded using only the default transform or using the multiple transform set (MTS) including the default transform and the extra transform. When the decoding target block uses MTS, information about the transform actually used among multiple transforms is transmitted or guided. Here, the information about the transform actually used includes transform type on the horizontal axis and transform type on the vertical axis. That is, when the image decoding device transforms the decoding target block using MTS, it receives whether to use any transform type among multiple transform types and performs decoding.
[0113] According to one embodiment, DCT-II is set as the default transform, and DST-7 and DCT-8 are set as extra transforms. At this time, the maximum size of the default transform, i.e., DCT-II, is supported up to 64×64, and the maximum size of the extra transforms, i.e., DST-7 and DCT-8, is supported up to 32×32. For example, when the size of the decoding object block is 64×64, a 64×64 DCT-II is applied to the transform process. That is, when one or more of the width and height of the decoding object block is greater than 32 (exceeds 32), the default transform is directly applied when MTS is not applied. ). That is, from the perspective of the image decoding device, only when the horizontal and vertical dimensions of the decoding object block are all less than 32, it is determined whether to use MTS for conversion. On the contrary, when one of the horizontal or vertical dimensions of the decoding object block is greater than 32, it is determined to apply the default conversion and convert. Therefore, for the case where the decoding object block is converted by the default conversion, the syntax information transmitted related to MTS does not exist. In the present invention, for convenience, the transform type value of DCT-II is set to "0", the transform type value of DST-7 is set to "1", and the transform type value of DCT-8 is set to "2", but it is not limited to this. Table 2 below defines the transform type assigned to the value of the trType syntax.
[0114]
Table 2
[0115] Table 3 and Table 4 show examples of transform kernels of DST-7 and DCT-8 when the size of the decoding target block is 4×4.
[0116] Table 3 shows the parameter values of the corresponding transform kernel when tyType is "1" (DST-7) and the size of the decoding object block is 4x4. Table 4 shows the parameter values of the corresponding transform kernel when tyType is "2" (DCT-8) and the size of the decoding object block is 4x4.
[0117]
Table 3
[0118]
Table 4
[0119] The entire transform area of the decoding object block includes a zero-out area. Transform converts the value of the pixel domain to the frequency domain value. At this time, the upper left frequency area is called the low frequency area, and the lower right frequency area is called the high frequency area. The low frequency component reflects the general (average) characteristics of the corresponding block, and the high frequency component reflects the sharp (specific) characteristics of the corresponding block. Therefore, there are many large values in the low frequency component and a few small values in the high frequency component. The few small values in the high frequency area have mostly zero values through the quantization process after the transformation. Here, the remaining area with mostly zero values outside the low frequency area belonging to the upper left side is called the zero-out area, and the zero-out area is excluded during the signaling process. The area in the decoding object block except the zero-out area is called the valid area.
[0120] Figure 7 It is shown that a zero-out area is left except for the mxn area in the MxN decoding target block.
[0121] Referring to 7, the gray area on the upper left side is the low frequency area, and the white area shows the zeroing of the high frequency.
[0122] For another example, when the decoding target block is 64×64, the upper left 32×32 area is a valid area, and the remaining area is a zero-out area, and no signaling is sent.
[0123] Furthermore, when the size of the decoding object block is one of 64x64, 64x32, and 32x64, the upper left 32x32 area is a valid area, and the rest is a zero-out area. When issuing the syntax of the decoder quantization parameter, the zero-out area is not signaled because the block size is notified. That is, the area where the width or height of the decoding object block is larger than 32 is set as a zero-out area. At this time, since it corresponds to the case where the horizontal or vertical size of the decoding object block is larger than 32, the transform used is the default transform, namely DCT-II. The maximum size supported by the extra transforms, namely DST-7 and DCT-8, is up to 32×32, and MTS is not applicable to the object block of this size.
[0124] Furthermore, when the size of the decoding object block is one of 32x32, 32x16, and 16x32, in the case where MTS is applied to the decoding object block (for example, when DST-7 or DCT-8 is used), the upper left 16x16 area is a valid area, and the remaining part is set as a zero-out area. Here, the zero-out area can also be signaled according to the position and scanning method of the last valid parameter. After completing the signaling of the syntax related to the quantized parameters, the encoder signals the MTS index (mts_idx) value because the information of the conversion type is not informed when the syntax of the quantized parameters of the decoder is issued. As described above, in the case of signaling the zero-out area, the decoder performs conversion only on the valid area after ignoring or removing the quantized parameters corresponding to the zero-out area. Here, the information of the conversion type actually used also includes a horizontal axis conversion type and a vertical axis conversion type. For example, when the horizontal axis transformation type of the decoding object block is DST-7 or DCT-8, the horizontal axis (width) valid area of the decoding object block is 16; when the vertical axis transformation type of the decoding object block is DST-7 or DCT-8, the vertical axis (height) valid area of the decoding object block is 16.
[0125] However, when the size of the decoding object block is one of 32x32, 32x16, and 16x32 and MTS is not applied to the decoding object block (for example, when the default transformation, i.e., DCT-II, is used), the default transformation, i.e., DCT-II, is used for transformation as all regions are valid regions and there is no zero-out region. Here, the information related to the transformation (transform) actually used also includes the horizontal axis transformation type (transform type) and the vertical axis transformation type (transform type). For example, when the horizontal axis transformation type of the decoding object block is DCT-II, the horizontal axis (width) valid region of the decoding object block is the width of the corresponding block, and when the vertical axis transformation type of the decoding object block is DCT-II, the vertical axis (height) valid region of the decoding object block is the height of the corresponding block. That is, all regions (width x height) of the decoding object block are valid regions.
[0126] In addition, when the decoding target block has a size smaller than 16 not defined above, all areas are valid areas and there is no zero area. Whether the MTS of the decoding target block is applicable and the conversion type value are determined by default and / or explicit embodiments.
[0127] In the present invention, the syntax content showing the position of non-zero parameters in the quantized coefficient after the conversion and quantization process is the same as the HEVC method. However, the syntax name of coded_sub_block_flag is changed to sb_coded_flag for use. In addition, the method of scanning the quantized parameters uses the up-right diagonal method.
[0128] In the present invention, MTS is applied to luma blocks (not applied to chroma blocks). In addition, a flag indicating whether MTS is used or not is used, that is, the MTS function is turned on / off using sps_mts_enabled_flag. In the case of using the MTS function, sps_mts_enabled_flag=on is set to set whether the explicit MTS function is used for intra-picture prediction and inter-picture prediction. That is, a flag sps_explicit_mts_intra_enabled_flag indicating whether MTS is used when intra-picture prediction is used and a flag sps_explicit_mts_inter_enabled_flag indicating whether MTS is used when inter-picture prediction are used are also set. In this specification, for convenience, the flags sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, and sps_explicit_mts_inter_enabled_flag indicating whether the three MTSs are used are described in SPS (sequence parameter set), but the present invention is not limited thereto. That is, the three flags are respectively set in one or more positions in DCI (decoding capability information), video parameter set (VPS), sequence parameter set SPS (SPS), picture parameter set (PPS), picture layer header (PH), and slice header (SH). In addition, the flag indicating whether the three MTSs are used is defined as high level syntax (HLS).
[0129] In the use of MTS, there are two methods: explicit method and default method. Explicit use of MTS is to transmit MTS related information (e.g., conversion information actually used) when the flag value indicating whether MTS is used is set to on within the screen and / or between the screens in the SPS and when specific conditions are met. That is, the image decoding device receives MTS related information, based on which the decoding target block uses the conversion type to confirm whether the conversion is performed, and based on this, performs decoding. For example, in an environment where explicit MTS is used, the three flags are set as follows.
[0130] 1.sps_mts_enabled_flag=on 2.sps_explicit_mts_intra_enabled_flag=on 3.sps_explicit_mts_inter_enabled_flag=on The use of the default MTS is to guide MTS related information (eg, conversion information actually used) when the sps_mts_enabled_flag value is set to on among the three flags in the SPS and specific conditions are met. For example, in an environment where the default MTS is used, the three flags are set as follows.
[0131] 1.sps_mts_enabled_flag=on 2.sps_explicit_mts_intra_enabled_flag=off 3.sps_explicit_mts_inter_enabled_flag=off (on and off are irrelevant) The default MTS method and the explicit MTS method are described below through several embodiments.
[0132] 2. First embodiment (default MTS) The default MTS described in the present embodiment is used when the decoding object block is encoded by the intra-screen prediction method. That is, when the decoding object block is encoded by the image encoding device, in the case of encoding by the intra-screen prediction method, the default MTS is used by the image encoding device and / or the image decoding device to perform encoding and / or decoding. In addition, when decoding the decoding object block, whether to use the default MTS is indicated by the implicitMtsEnabled parameter. The image decoding device confirms the value of the implicitMtsEnabled parameter and uses the default MTS to determine whether to perform decoding. For example, in the case of using the default MTS in decoding, the implicitMtsEnable parameter has a value of 1, otherwise the implicitMtsEnable parameter has a value of 0. In addition, in this specification, the implicitMtsEnabled can also be displayed as "implicit_MTS_enabled" according to the situation.
[0133] In the case of explaining the conditions of HLS (high level syntax) for applying the default MTS, sps_mts_enabled_flag has no relation to whether it is default or explicit, and is a flag indicating whether the MTS is applied or not. Therefore, in order to apply the default MTS, it is set to "on". In addition, the default MTS is used when the decoding object block is encoded by the image encoding device, and is encoded by a prediction method. Therefore, the image decoding device determines whether the default MTS is applicable by confirming the value of sps_explicit_mts_intra_enabled_flag. However, when the decoding object block is encoded by the intra-screen prediction method when the decoding object block is encoded by the image encoding device, and the explicit MTS is applied, sps_explicit_mts_intra_enabled_flag is set to "on". Therefore, when the decoding object block is encoded by the default MTS by the image encoding device, sps_explicit_mts_intra_enabled_flag is set to "off". In addition, as described above, the default MTS is used when the decoding target block is encoded by the intra-picture prediction method by the image encoding device. Therefore, it is not important whether the sps_explicit_mts_inter_enabled_flag value of the explicit MTS indicating the case where the decoding target block is encoded by the intra-picture prediction method by the image encoding device has the value. In addition, the default MTS is used when the decoding target block is encoded by the intra-picture prediction method by the image encoding device, and the case where CuPredMode has the MODE_INTRA value is applied.
[0134] When these are arranged, the conditions for decoding the decoding target block by the video decoding device using the default MTS are listed as follows.
[0135] 1) sps_mts_enabled_flag is equal to 1 2) sps_explicit_mts_intra_enabled_flag is equal to 0 3) CuPredMode is equal to MODE_INTRA (intra-picture prediction method) In addition, CuPredMode[0][xTbY][yTbY] indicating the prediction mode of the current position in the luma block has a MODE_INTRA value.
[0136] Additional conditions under which the default MTS can be used are as follows.
[0137] 4) lfnst_idx is equal to 0 5) intra_mip_flag is equal to 0 Here, the lfnst_idx value indicates secondary transform, and when lfnst_idx=0, secondary transform is not used. The intra_mip_flag value indicates whether a matrix-based intra prediction method (matrix-based intra prediction: mip) is used or not. When intra_mip_flag=0, matrix-based prediction is not used, and when intra_mip_flag=1, matrix-based prediction is used.
[0138] That is, this embodiment describes a method of predicting by a general intra-frame prediction method and setting a primary transform type (or MTS) for a decoding target block that does not use a secondary transform. When all the five conditions are met, the default MTS function is activated (see 13).
[0139] Figure 8 A method for determining whether the default MTS function is applicable or not according to an embodiment of the present invention is shown. Figure 8 The various steps are executed in the image decoding device.
[0140] Reference Figure 8 , the image decoding device determines whether sps_mts_enable_flag has a value of 1, sps_explicit_mts_intra_enable_flag has a value of 0, and CuPredMode has a value of MODE_INTRA (S810). As a result of the judgment, for the case where all the conditions of S810 are satisfied, the image decoding device determines whether lfnst_idx has a value of 0 and whether intra_mip_flag has a value of 0 (S820). For the case where all the conditions of S810 and S820 are satisfied, the implicit_MTS_enabled value is set to 1 (S830). In addition, for the case where one of the conditions of S810 or S820 cannot be satisfied, the image decoding device sets the implicit_MTS_enabled value to 0 (S840).
[0141] When the default MTS for the decoding target block is activated (implicit_MTS_enabled=on), the MTS value (transform information actually used) is guided according to the width and height of the corresponding block (refer to Fig. 9 ). At this time, only a part of the object block is transformed (transform) without undergoing the sub-block transform (sbt) of the transform process. That is, the cu_sbt_flag value of the object block is "0".
[0142] Fig. 9 A method for converting the width and height of a corresponding block of a guide default MTS according to an embodiment of the present invention is shown. Fig. 9 Each step can be executed in the image decoding device.
[0143] In reference Fig. 9 In this case, the image decoding device determines whether the implicit_MTS_enabled value is "1" (S910). At this time, although not shown in the accompanying drawings, the image decoding device further confirms whether the cu_sbt_flag value has a value of "0". At this time, when cu_sbt_flag has a value of "1", it is displayed that only a portion of the decoding object block is converted through the sub-block conversion through the conversion process. Moreover, when cu_sbt_flag has a value of "0", it is displayed that only a portion of the decoding object block is not converted through the sub-block conversion through the conversion process. Therefore, Fig.14 The action is set only when the cu_sbt_flag has a value of "0".
[0144] When the implicit_MTS_enabled value is 1, determine whether the value of nTbW is greater than 4 and less than 16 (S920). When the implicit_MTS_enabled value is not "1", complete the operation. nTbW indicates the width of the corresponding conversion block and is used to determine whether to use the additional conversion type, i.e., DST-7, in the horizontal direction.
[0145] As a result of the judgment in step S920, if the value of nTbW has a value greater than or equal to 4 and less than or equal to 16, trTypeHor is set to "1" (S930), and if the value of nTbW does not have a value greater than or equal to 16, trTypeHor is set to "0" (S940). At this time, the nTbW indicates the width of the corresponding conversion block, in order to determine whether to use the additional conversion type, i.e., DST-7, in the horizontal axis direction. At this time, if tyTypeHor is set to "0", it is determined that the corresponding conversion block is converted using the default type conversion, i.e., DCT-II conversion, in the horizontal axis direction. In addition, if trTypeHor is set to "1", it is determined that the corresponding conversion block is converted using one of the additional conversion types, i.e., DST-7 conversion, in the horizontal axis direction.
[0146] Furthermore, the image decoding device determines whether the value of nTbH has a value greater than or equal to 4 and less than or equal to 16 (S950), and sets trTypeVer to "1" (S960) if the value of nTbH has a value greater than or equal to 4 and less than or equal to 16, and sets trTypeVer to "0" (S970) if the value of nTbW does not have a value greater than or equal to 4 and less than or equal to 16. The nTbH indicates the height of the corresponding conversion block, and is used to determine whether to use the additional conversion type, i.e., DST-7, in the vertical axis direction. At this time, when trTypeVer is set to "0", it is determined that the corresponding conversion block is converted using the default type conversion, i.e., DCT-II conversion, in the vertical axis direction. In addition, when trTypeVer is set to "1", it is determined that the corresponding conversion block is converted using one of the additional conversion types, i.e., DST-7 conversion, in the vertical axis direction.
[0147] Fig.10 A method for performing an inverse conversion of conversion-related parameters according to an embodiment of the present invention is shown. Fig.10 The various steps are performed in an image decoding device, for example, in an inverse conversion unit of the decoding device.
[0148] Reference Fig.10, the video decoding device obtains sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, y0], NTbW, nTbH (S1010). At this time, sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, y0] respectively indicate what is in Figure 8 The parameters are used to determine whether the decoding target block is suitable for the default MTS. In addition, NTbW and nTbH respectively show the width and height of the corresponding conversion block, and are used to determine whether to use the additional conversion type, namely DST-7.
[0149] Next, the video decoding apparatus sets implicit_MTS_enabled based on sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], Lfnst_idx, and IntraMipFlag[x0, y0] values (S1020). At this time, the implicit_MTS_enabled flag is set by executing Fig.13 is set according to the process.
[0150] Next, the video decoding device sets trTypeHor and trTypeVer based on the implicit_MTS_enabled, nTbW, and nTbH values (S1030). At this time, the methods of trTypeHor and trTypeVer are executed by Fig. 9 is set according to the process.
[0151] Next, the video decoding device performs inverse conversion based on trTypeHor and trTypeVer (S1040). The inverse conversion applied by trTypeHor and trTypeVer is constructed according to Table 2. For example, when trTypeHor is "1" and trTypeVer is "0", DST-7 is applied in the horizontal direction and DST-II is applied in the vertical direction.
[0152] In addition, although not shown in the drawings, from the perspective of the image encoding device, in order to set whether the default MTS is used, sps_mts_enabled_flag, sps_explicit_mts_intra_enabled_flag, CuPredMode[0][xTbY][yTbY], lfnst_idx, IntraMipFlag[x0, y0], NTbW, and nTbH are set.
[0153] 3. Second embodiment (explicit MTS) In this embodiment, for the case where the MTS function is explicitly activated in HLS (high level syntax), a transformation method applicable to the decoding object block is described. The conditions of HLS (high level syntax) for applying the explicit MTS are described. sps_mts_enabled_flag is a flag that shows whether MTS is applied, regardless of whether it is default or explicit. In order to apply the default MTS, it is necessary to set it to "on". In addition, the explicit MTS is applicable when the decoding object block is encoded by the intra-picture prediction method and when it is encoded by the inter-picture prediction method. For the case where the explicit MTS is applied, sps_explicit_mts_intra_enabled_flag and / or sps_explicit_mts_intra_enabled_flag are all set to "on". When it is organized, it is listed as shown in the following conditions.
[0154] 1) sps_mts_enabled_flag = on 2) sps_explicit_mts_intra_enabled_flag=on 3) sps_explicit_mts_inter_enabled_flag=on Here, when the decoding target block is encoded by the intra-picture prediction method, the condition of sps_explicit_mts_intra_enabled_flag = "on" is confirmed, and when the decoding target block is encoded by the inter-picture prediction method, the condition of sps_explicit_mts_inter_enabled_flag = "on" is confirmed.
[0155] Other conditions under which explicit MTS can be used are as follows.
[0156] 4) lfnst_idx is equal to 0 (refer to the default MTS) 5) transform_skip_flag is equal to 0 6) intra_subpartitions_mode_flag is equal to 0 7) cu_sbt_flag is equal to 0 (refer to default MTS) 8) Valid MTS area 9) The width and height of the object block are less than 32 Here, the lfnst_idx value indicates the secondary transform, and in the case of lfnst_idx=0, it means that the secondary transform is not used.
[0157] The transform_skip_flag value indicates that the transform process is omitted. For the case where transform_skip_flag=0, the transform is displayed as normal without omitting the transform process. The intra_subpartitions_mode_flag value indicates that the object block is divided into multiple sub-blocks through one of the intra-screen prediction methods, and undergoes the prediction, conversion, and quantization processes. That is, for the case where the corresponding flag value (intra_subpartitions_mode_flag) is "0", the object block is not divided into sub-blocks, which means that general intra-screen prediction is performed. In addition, the use of MTS is limited by the supported sizes (supported up to 32x32, as described above) of the extra transforms (DST-7 and DCT-8). That is, the width and height of the object block must be less than 32 to use MTS. That is, even if one of the width and height exceeds 32, the default transform (MTS cannot be used) is performed ( ) is DCT-II.
[0158] cu_sbt_flag indicates whether only a portion of the object block undergoes sub-block transform (sbt) of the transform process. That is, if the cu_sbt_flag value is "0", it means that only a portion of the object block does not undergo sub-block transform (sbt) of the transform process.
[0159] Next, the valid area (hereinafter, valid MTS area) will be described in detail.
[0160] Fig.11 An example of a valid MTS area marked by a bold line in a 32x32 decoding target block is shown.
[0161] Reference Fig.11 , the 16x16 area on the upper left side except the DC coefficient is the valid MTS area. That is, the 16x16 area on the upper left side except the 1x1 (DC) area is the valid MTS area. For example, when all non-zero coefficients in the object block belong to the valid MTS area, MTS can be applied. When one or more non-zero coefficients are out of the valid MTS area, MTS cannot be applied and the default transform is performed ( ) is DCT-II. It is the same concept as the zero-out area described above. That is, when the 32x32 object block uses MTS (i.e., DST-7 or DCT-8), the upper left 16x16 is the valid area, and the remaining part is as shown in the zero-out area. When all non-zero coefficients in the 32x32 object block are in the upper left 16x16 area, MTS (i.e., DST-7 or DCT-8) can be applied. However, in the exceptional case where there is only one non-zero coefficient in the block and the position is DC (1x1), MTS cannot be applied and the default transform (default transform) is performed ( ) is DCT-II.
[0162] As a result, in this embodiment, in order to determine whether MTS is applicable or not, it is necessary to confirm the valid MTS area. In order to confirm the valid MTS area, the following two conditions should be confirmed.
[0163] (a) When the non-zero coefficient in a block is one, the position is identified as DC (1x1) (b) Are all non-zero coefficients in the block in the upper left 16x16 region? In order to confirm the above-mentioned condition (a), the last position information is used. Here, the last position refers to the position of the last non-zero coefficient, that is, the last significant coefficient, in the scanning order within the object block. For example, the last sub-block information containing the last position, that is, the last non-zero coefficient, is used. For example, when the position of the last sub-block is not (0, 0), the condition (a) is satisfied. In other words, when the position of the last sub-block in the scanning order of the sub-blocks within the object block is not "0" (when it is greater than 0), the condition (a) is satisfied. Or when the position of the last sub-block is "0", the last scan position information showing the relative position of the last position in the corresponding sub-block can also be used. For example, when the last scan position in the coefficients scanning order in the corresponding sub-block is not "0" (is greater than 0), condition (a) is satisfied (refer to Fig.12 ). And, as mentioned above, the MTS of the present invention is applicable to luma blocks.
[0164] In order to confirm the above-mentioned condition (b), sub-block information containing one or more non-zero coefficients is used. Here, the sub-block information containing one or more non-zero coefficients can be confirmed by the sb_coded_flag value of the corresponding sub-block. For the case where the corresponding flag value is "1" (sb_coded_flag=1), it means that one or more non-zero coefficients are set in the corresponding sub-block, and for the case where sb_coded_flag=0, it means that all coefficients in the corresponding sub-block are all zero. That is, when the positions of all sub-blocks whose sb_coded_flag values are all "1" in the object block are within (0, 0) to (3, 3), condition (b) is satisfied. On the contrary, even if the sb_coded_flag value in the target block is "1", that is, one of the sub blocks, if it is out of the position between (0, 0) and (3, 3), condition (b) cannot be satisfied. In other words, even if the sb_coded_flag value in the target block is "1", that is, one of the sub blocks, if the x-coordinate or y-coordinate of the sub block has a value greater than 3, condition (b) cannot be satisfied (see Fig.13 ). For another example, when the sb_coded_flag value is "1" in the sub-block scanning order within the object block, and the first sub-block having a value greater than 3 in the x-coordinate or y-coordinate of the sub-block is found, the confirmation process of the sub-blocks whose sb_coded_flag value is "1" in the subsequent scanning order can be omitted by setting condition (b) to false (refer to Fig.13 ). And, as mentioned above, the MTS of the present invention is applicable to luma blocks.
[0165] Fig.12 A method for determining a valid MTS according to an embodiment of the present invention is shown. Fig.12 The embodiment relates to a method of confirming the condition (a) of the two conditions for confirming a valid MTS region as described above. Fig.12 Each step can be executed in the image decoding device.
[0166] Reference Fig.12, the image decoding device sets MtsDcOnlyFlag to "1" (S1210). The MtsDcOnlyFlag indicates whether there is one non-zero coefficient in the block and whether the position is DC. For example, in the case where there is one non-zero coefficient in the block and the position is DC, the MtsDcOnlyFlag has a value of "1", and in other cases, the MtsDcOnlyFlag has a value of "0". At this time, the image decoding device applies MTS when the MtsDcOnlyFlag value has a value of "0". In step S1210, the reason for setting MtsDcOnlyFlag to "1" is that when the corresponding block has one non-zero coefficient in the following block, if the condition of not being in the DC position is met, the MtsDcOnlyFlag is reset to "0", otherwise MTS is not applied.
[0167] Next, the video decoding apparatus determines whether the target block is a luma block (S1220). The purpose of determining whether the target block is a luma block is as described above, because MTS is only applicable to luma blocks.
[0168] Next, the video decoding device determines whether the last sub block is greater than 0 ( S1230 ). If the last sub block is greater than 0, MtsDcOnlyFlag is set to “0” ( S1240 ), and the process ends.
[0169] If the result of the determination in step S1230 is that the last sub-block is not greater than 0, it is determined whether the last scan position is greater than 0 ( S1250 ).
[0170] As a result of the determination in step S1250 , if the last scan position is greater than 0, MtsDcOnlyFlag is set to “0” ( S1240 ), and the process ends.
[0171] The judgment result of step S1250 is that the process ends when the last scanning position is not greater than 0.
[0172] According to this embodiment, when the last sub-block is greater than 0 or the last scanning position is greater than 0, MtsDcOnlyFlag is set to "0", otherwise MtsDcOnlyFlag is set to "1".
[0173] When determining whether to apply MTS later, if MtsDcOnlyFlag is confirmed and has a value of "1", MTS is not applied and DCT-II, which is a default transform, can be applied.
[0174] Fig.13 A method for determining a valid MTS area according to another embodiment of the present invention is shown. Fig.13 The embodiment specifically shows a method for confirming the condition (b) of the two conditions for the effective MTS region described above. Fig.13 Each step can be executed in the image decoding device.
[0175] Reference Fig.13 , the image decoding device sets MtsZerooutFlag to "1" (S1305). The MtsZerooutFlag indicates whether the non-zero parameters in the block exist in the zero-out area. For example, in the case where at least one of the non-zero parameters in the block exists in the zero-out area, MtsZerooutFlag has a value of "0", and in the case where all non-zero parameters in the block do not exist in the zero-out area, MtsZerooutFlag has a value of "1". In this embodiment, in the case where the initial value of MtsZerooutFlag is set to "1" assuming that all non-zero parameters in the block do not exist in the zero-out area, and the condition of the zero-out area and the condition of the non-zero parameter are satisfied at the same time, MtsZerooutFlag is set to "0". At this time, in the case where there is an MtsZerooutFlag with a value of "0", the explicit MTS is not applied.
[0176] Next, the image decoding device sets the initial value of the variable i to the value of the last sub block, and repeatedly executes the following process from step S1325 to step S1350 (S1320) until the value of the variable i is offset once and the value of the variable i is 0. The purpose of repeatedly executing the routine of step S1820 is to confirm the sb_coded_flag values of all sub blocks from the last sub block to the first sub block. As described above, for the case of the corresponding flag value "1", there is more than one non-zero parameter in the corresponding sub block, and for the case of the corresponding flag value "0", there is no non-zero parameter in the corresponding sub block. Therefore, refer to Fig.11 For the case where the positions of all sub-blocks with sb_coded_flag value "1" in the object block exist only within (0, 0) to (3, 3), that is, based on the variable i, they exist only within 0 to 8, it is judged that the (b) condition for applying explicit MTS is satisfied.
[0177] Next, the video decoding device determines whether the variable i simultaneously satisfies the conditions of being smaller than the last sub-block (i < last sub block) and being greater than 0 (i > 0) (S1325). For example, when the routine of step S1320 is first executed, since the initial value of the variable i is set to the same value as the last sub-block (last sub block), the condition of step S1325 is not satisfied.
[0178] For the case where the determination result of step S1325 is that the variable i simultaneously satisfies the conditions of being smaller than the last sub-block (i < last sub block) and being greater than 0 (i > 0), sb_coded_flag is issued (S1830). For the case where the two conditions are not simultaneously satisfied, sb_coded_flag is set to "1" (S1835).
[0179] At this time, the issued sb_coded_flag indicates whether there is more than one non-zero parameter in the corresponding sub-block. In the case where there is more than one non-zero parameter in the corresponding sub-block, sb_coded_flag has a value of "1", and in the case where there is no non-zero parameter in the corresponding sub-block, sb_coded_flag has a value of "0".
[0180] In addition, when i only indicates the last sub-block and the first sub-block, step S1835 is executed. That is, in the last sub-block, the last position parameter is included, and the sb_coded_flag value is issued as "1", and in the first sub-block, the DC parameter exists, and the sb_coded_flag value is issued as "1".
[0181] Next, the video decoding device determines whether the corresponding block is a luma block (S1340). The purpose of determining whether the block to be determined is a luma block is as described above, and MTS is only applicable to luma blocks.
[0182] For the case where the determination result of step S1340 is that the corresponding block is a luma block, it is determined whether the condition "sb_coded_flag && (xSb > 3 || ySb > 3)" is satisfied (S1845). For the case where the condition of step S1845 is satisfied, MtsZerooutFlag is set to "0" (S1350).
[0183] According to this embodiment, for sub-blocks other than sub-block (3, 3) in the object block, that is, for the case where a non-zero parameter is found in the zeroing region, MtsZerooutFlag is set to "0" and it is determined that explicit MTS cannot be applied.
[0184] Fig.14 A method for determining a valid MTS according to another embodiment of the present invention is shown. Fig.14 The embodiment specifically shows a method for confirming the condition (b) of the two conditions for confirming the effective MTS area described above. Fig.13 In the embodiment of the present invention, the valid MTS area is confirmed by confirming the sb_coded_flag of all sub-blocks, but in Fig.14 In the embodiment, for the case where an invalid MTS is found for the first time, there is a difference in the sb_coded_flag even if it is not confirmed afterwards. Fig.14 Each step can be executed in the image decoding device.
[0185] Reference Fig.14 , the image decoding device sets MtsZerooutFlag to "1" (S1405). The MtsZerooutFlag indicates whether the non-zero parameters in the block exist in the zero-out area. For example, in the case where at least one of the non-zero parameters in the block exists in the zero-out area, MtsZerooutFlag has a value of "0", and in the case where all non-zero parameters in the block do not exist in the zero-out area, MtsZerooutFlag has a value of "1". In this embodiment, it is assumed that all non-zero parameters in the block do not exist in the zero-out area and the initial value of MtsZerooutFlag is set to "1". In the case where both the condition of the zero-out area and the condition of the non-zero parameter are satisfied, MtsZerooutFlag is set to "0". At this time, in the case where there is an MtsZerooutFlag with a value of "0", the explicit MTS is not applied.
[0186] Next, the image decoding device sets the initial value of the variable i to the value of the last sub block, offsets the value of the variable i once, and repeatedly executes the following process from step S1425 to step S1450 (S1420) until the value of the variable i is 0. The purpose of the routine of repeatedly executing step S1420 is to confirm the sb_coded_flag values of all sub blocks from the last sub block to the first sub block. As described above, for the case where the sb_coded_flag value is "1", there is more than one non-zero parameter in the corresponding sub block, and for the case where the sb_coded_flag value is "0", there is no non-zero parameter in the corresponding sub block. Therefore, referring to Fig.16For the case where the positions of all sub - blocks with the sb_coded_flag value of "1" within the object block exist only within (0, 0) to (3, 3), that is, exist only within 0 to 8 based on the variable i, it is determined that the condition (b) for applying the explicit MTS is satisfied.
[0187] Next, the video decoding device determines whether the variable i simultaneously satisfies the conditions of being less than the last sub - block (i < last sub block) and greater than 0 (i > 0) (S1425). For example, when the routine of step S1920 is executed for the first time, since the initial value of the variable i is set to the same value as the last sub - block, the condition of step S1425 is not satisfied.
[0188] For the case where the determination result of step S1425 simultaneously satisfies the conditions of the variable i being less than the last sub - block (i < last sub block) and greater than 0 (i > 0), sb_coded_flag is issued (S1430). For the case where the two conditions are not simultaneously satisfied, sb_coded_flag is set to "1" (S1435).
[0189] At this time, the issued sb_coded_flag indicates whether there is more than one non - zero parameter within the corresponding sub - block. For the case where there is more than one non - zero parameter within the corresponding sub - block, sb_coded_flag has a value of "1", and for the case where there is no non - zero parameter within the corresponding sub - block, sb_coded_flag has a value of "0".
[0190] In addition, step S1435 is executed only when i indicates the last sub - block and the first sub - block. That is, since the last position parameter is included in the last sub - block, the value of sb_coded_flag is issued as "1", and since the DC parameter exists in the first sub - block, the value of sb_coded_flag is issued as "1".
[0191] Next, the video decoding device determines whether the condition of "MtsZerooutFlag && luma block" is satisfied (S1440).
[0192] The judgment result of step S1440 is that if the condition of "MtsZerooutFlag&&luma block" is met, it is further judged whether the condition of "sb_coded_flag&&(xSb>3||ySb>3)" is met (S1445). If the condition of "sb_coded_flag&&(xSb>3||ySb>3)" is met, MtsZerooutFlag is set to "0" (S1450).
[0193] If the judgment result of step S1440 does not satisfy the condition of "MtsZerooutFlag&&lumablock", the process in the corresponding sub-block is terminated.
[0194] According to the present embodiment, even when the corresponding variable i, ie, the corresponding subblock once sets the MtsZerooutFlag value to "0", in the following routine, ie, variable i-1, a false (False) value is derived in step S1940 without further confirming the sb_coded_flag value.
[0195] In addition, if the decoding target block satisfies all the conditions (a) and (b), it is determined to use the explicit MTS, and the transform information actually used for the corresponding block is transmitted in the index form (mts_idx). If all conditions cannot be met, the default transform (default transform) is used ( ) is DCT-II (refer to Fig.15 ). Table 5 shows the conversion types of the horizontal and vertical axes of the mts_idx value.
[0196]
Table 5
[0197] In Table 5, trTypeHor refers to the horizontal axis transform type, and trTypeVer refers to the vertical axis transform type. The transform type value in Table 5 refers to the trType value in Table 2. For example, when the mts_idx value is "2", DCT-8 (2) is used for horizontal axis transform, and DST-7 (1) is used for vertical axis transform.
[0198] In using / implementing / applying the present invention, the default transform (default transform) mentioned above ( ) In all cases of DCT-II, it is replaced by the expression "the mts_idx value is guided to "0"". That is, when the mts_idx value is "0", DCT-II (0) is set in all cases by transforming the horizontal axis and the vertical axis.
[0199] In the present invention, the binary method of mts_idx uses TR (truncated rice), and the parameter values for TR, namely cMax value is "4" and cRiceParam value is "0". Table 6 shows the codewords of the MTS index.
[0200]
Table 6
[0201] Referring to Table 6, it is confirmed as follows: when the mts_idx value is "0", the corresponding code word is "0", when the mts_idx value is "1", the corresponding code word is "10", when the mts_idx value is "2", the corresponding code word is "110", when the mts_idx value is "3", the corresponding code word is "1110", and when the mts_idx value is "4", the corresponding code word is "1111".
[0202] Fig.15 A method for determining whether the explicit MTS function according to an embodiment of the present invention is applicable is shown. Fig.15 Each step can be executed in the image decoding device.
[0203] Reference Fig.15 , the image decoding device determines whether the condition of “(sps_explicit_mts_intra_enabled_flag&&CuPredMode=MODE_INTRA)||(sps_explicit_mts_inter_enabled_flag&&CuPredMode=MODE_INTER)” is satisfied ( S1510 ).
[0204] sps_explicit_mts_intra_enabled_flag is a flag indicating whether or not explicit MTS is used during intra-picture prediction, and sps_explicit_mts_intra_enabled_flag is a flag indicating whether or not explicit MTS is used during inter-picture prediction. sps_explicit_mts_intra_enabled_flag has a value of "1" when explicit MTS is used during intra-picture prediction, and has a value of "0" in other cases. sps_explicit_mts_intra_enabled_flag has a value of "1" when explicit MTS is used during inter-picture prediction, and has a value of "0" in other cases.
[0205] CuPredMode indicates whether the decoding target block is encoded by any prediction method. If the decoding target block is encoded by the intra-picture prediction method, CuPredMode has a MODE_INTRA value, and if the decoding target block is encoded by the inter-picture prediction method, CuPredMode has a MODE_INTER value.
[0206] Therefore, when the decoding target block uses intra-picture prediction and explicit MTS, "sps_explicit_mts_intra_enabled_flag&&CuPredMode=MODE_INTRA" has a value of "1", and when the decoding target block uses inter-picture prediction and explicit MTS, "sps_explicit_mts_inter_enabled_flag&&CuPredMode=MODE_INTER" has a value of "1". Therefore, step S2010 determines whether the decoding target block uses explicit MTS by confirming the values of sps_explicit_mts_intra_enabled_flag, sps_explicit_mts_inter_enabled_flag, and CuPredMode.
[0207] When the condition of step S1510 is satisfied, the video decoding apparatus determines whether the condition of “lfnst_idx=0&&transform_skip_flag=0&&cbW<32&&cbH<32&&intra_subpartitions_mode_flag=0&&cu_sbt_flag=0” is satisfied ( S1520 ).
[0208] Here, the lfnst_idx value indicates the secondary transform, and in the case of lfnst_idx=0, it means that the secondary transform is not used.
[0209] The transform_skip_flag value indicates whether transform skip is applied to the current block. That is, it indicates whether the transform process is omitted in the current block. If transform_skip_flag=0, transform skip is not applied to the current block.
[0210] cbW and cbH respectively show the width and height of the current block. As shown in the above description, the maximum size supported by the default transform (defaulttransform), i.e. DCT-II, is up to 64×64, and the maximum size supported by the extra transforms (extra transforms), i.e. DST-7 and DCT-8, is up to 32×32. For example, when the size of the decoding object block is 64×64, a 64×64 DCT-II is applied to the transform process. That is, when one or more of the width and height of the decoding object block is greater than 32 (exceeds 32), the default transform (defaulttransform) is directly applied when MTS is not applicable ( ). Therefore, in order to apply MTS, the values of cbW and cbH are all less than 32.
[0211] intra_subpartitions_mode_flag indicates whether the intra subpartition mode (intrasubpartition mode) is applied or not. The intra subpartition mode indicates that the object block is divided into a plurality of sub-blocks in the intra-picture prediction method and then goes through the prediction, conversion, and quantization processes. That is, when the corresponding flag value (intra_subpartitions_mode_flag) is "0", it means that the object block is not divided into sub-blocks and ordinary intra-picture prediction is performed.
[0212] cu_sbt_flag indicates whether only a portion of the object block is subject to sub-block transform (sbt) through the transform process. That is, if the cu_sbt_flag value is "0", it means that only a portion of the object block is not subject to sub-block transform (sbt) through the transform process.
[0213] Therefore, whether the explicit MTS is applicable to the decoding target block is determined by whether the condition of step S1520 is satisfied.
[0214] If the condition of step S1510 is not satisfied, the video decoding apparatus sets the value of mts_idx to “0” ( S1530 ) and ends the process.
[0215] When the condition of step S1520 is satisfied, the video decoding apparatus determines whether the condition of “MtsZeroOutFlag=1&&MtsDcOnlyFlag=0” is satisfied ( S1540 ).
[0216] The MtsZerooutFlag indicates whether the non-zero parameters in the block exist in the zero-out region. If at least one of the non-zero parameters in the block exists in the zero-out region, MtsZerooutFlag has a value of "0", and if all non-zero parameters in the block do not exist in the zero-out region, MtsZerooutFlag has a value of "1". At this time, the value of MtsZerooutFlag is obtained by executing Fig.13 or Fig.14 determined by the process.
[0217] MtsDcOnlyFlag indicates whether there is one non-zero coefficient in the block and the position is DC. If there is one non-zero coefficient in the block and the position is DC, the MtsDcOnlyFlag has a value of "1". Otherwise, the MtsDcOnlyFlag has a value of "0". At this time, the value of MtsDcOnlyFlag is determined by executing Fig.12 determined by the process.
[0218] If the condition of step S1520 is not satisfied, the video decoding apparatus sets the value of mts_idx to “0” ( S1530 ) and ends the process.
[0219] If the condition of step S1540 is satisfied, the video decoding device issues mts_idx (S1550) and ends the process. At this time, the transform types of the horizontal axis and vertical axis of the mts_idx value are assigned by Table 5. At this time, the value of the transform type in Table 5 refers to the trType value in Table 2. For example, if the mts_idx value is "2", DCT-8 is applied to the horizontal axis transform, and DST-7 is applied to the vertical axis transform.
[0220] Furthermore, when the condition of step S1540 is not satisfied, the video decoding apparatus sets the value of mts_idx to “0” ( S1530 ), and ends the process.
[0221] Fig.16 A method for performing inverse conversion of conversion-related parameters according to another embodiment of the present invention is shown. Fig.16 The various steps are performed in the image decoding device, for example, in the inverse conversion unit of the decoding device.
[0222] Reference Fig.16 The video decoding device obtains the values of sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, and cu_sbt_flag (S1610). At this time, sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, and cu_sbt_flag are displayed. Fig.15 The parameter is used to determine whether the decoding object block is applicable to the explicit MTS.
[0223] Next, the video decoding device obtains the values of MtsZeroOutFlag and MtsDcOnlyFlag (S1620). Fig.13 or Fig.14 The process is obtained by executing Fig.12 obtained through the process.
[0224] Next, the video decoding apparatus obtains the mts_idx value based on the parameters obtained in step S1610 and step S1620 (S1630). That is, the video decoding apparatus obtains the mts_idx value based on sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, cu_sbt_flag, MtsZeroOutFlag, and MtsDcOnlyFlag. At this time, mts_idx is obtained by executing Fig.15 obtained through the process.
[0225] Next, the video decoding apparatus performs inverse transformation based on mts_idx (S1640). The inverse transformation applied according to the mts_idx value is configured according to Table 5 and Table 2. For example, when the mts_idx value is "2", DCT-8 is applied in the horizontal direction and DST-7 is applied in the vertical direction.
[0226] In addition, although not shown in the accompanying drawings, from the perspective of the image encoding device, in order to set whether to use explicit MTS, sps_explicit_mts_intra_enable_flag, sps_explicit_mts_inter_enable_flag, CuPredMode, lfnst_idx, transform_skip_flag, cbW, cbH, intra_subpartitions_mode_flag, cu_sbt_flag, MtsZeroOutFlag, and MtsDcOnlyFlag are set. In the above embodiment, the method is described based on the sequence diagram through a series of steps or blocks, but the present invention is not limited by the order of the steps, and any step occurs simultaneously through different steps and different orders from the above. In addition, those skilled in the art do not exclude the steps shown in the sequence diagram, and the inclusion of other steps or one or more steps in the sequence diagram does not affect the scope of the present invention and can be deleted.
[0227] The embodiments described in this specification are implemented and executed on a processor, microprocessor, controller or chip. For example, the functional units shown in each figure are implemented and executed on a computer, processor, microprocessor, controller or chip. Information (e.g. information on instructions) or algorithms for implementing this situation are stored in a digital storage medium.
[0228] Furthermore, the decoding device and encoding device to which the present invention is applicable are included in multimedia transmission devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video conversation devices, real-time communication devices for video communication, mobile streaming devices, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the top video) devices, network streaming service providers, three-dimensional (3D) video devices, videophone video devices, transportation device terminals (e.g. vehicle terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and are used for processing video signals or data signals. For example, game consoles, Blu-ray players, network-connected televisions, home theater systems, smartphones, tablet computers, DVRs (Digital Video Recoders), etc. are included as OTT video (Over the top video) devices.
[0229] Furthermore, the processing method applicable to the present invention is produced in the form of a program executed by a computer and stored in a computer-readable recording medium. The multimedia data with a data structure of the present invention is also stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices that store data read by a computer. The computer-readable recording medium includes such as a Blu-ray disc (BD), a serial bus interface (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Furthermore, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (for example, transmission through a network). Furthermore, the bit stream generated by the encoding method is stored in a computer-recordable recording medium or transmitted through a wired / wireless communication network.
[0230] Furthermore, the embodiments of the present invention are implemented by a computer program product according to a program code, wherein the program code is executed by a computer according to the embodiments of the present invention. The program code is stored in a carrier readable by a computer.
Claims
1. An image decoding method performed by an image decoding device comprises the following steps: Acquire, from the bit stream, a parameter related to whether a multiple transform set is applicable to a decoding target block; determining a transformation type of the decoding target block based on at least one of a parameter related to whether a multiple transformation set is applicable to the decoding target block and a size of the decoding target block; The valid area of the decoding target block is set based on at least one of a parameter related to whether a multiple transform set is applicable to the decoding target block and a size of the decoding target block, wherein: The valid region includes at least one non-zero coefficient; as well as restoring the decoding object block based on the conversion type of the decoding object block and the valid area of the decoding object block, When the size of the decoding target block is width 64×height 64, width 64×height 32, or width 32×height 64, the valid area of the decoding target block is set to the upper left side width 32×height 32 area of the decoding target block, and coefficient information of the remaining area of the decoding target block except the valid area is not sent to the image decoding device. When the size of the decoding target block is width 64×height 64, width 64×height 32, or width 32×height 64, the effective area is set regardless of the parameter. When the size of the decoding object block is width 32×height 32, width 32×height 16, or width 16×height 32 and a multiple transform set is applied to the decoding object block, the valid area of the decoding object block is set to the upper left 16×16 area of the decoding object block, When the size of the decoding object block is width 32×height 32, width 32×height 16, or width 16×height 32 and the multiple transform set is not applicable to the decoding object block, all areas of the decoding object block are the valid areas, The transformation type of the decoding object block is one of a first transformation type and a second transformation type, wherein the first transformation type includes a type indicating the same transformation kernel for a longitudinal direction and a transverse direction, and the second transformation type includes a type indicating different transformation kernels for a vertical direction and a horizontal direction.
2. A video encoding method performed by a video encoding device, comprising the following steps: Determining a transformation type of the encoding target block according to whether the multiple transformation sets are applicable to the encoding target block; The coefficient information is generated by performing a transformation corresponding to the transformation type on the encoding object block based on the effective area of the encoding object block, wherein The valid region of the encoding object block is determined by considering a transform type applicable to the encoding object block, wherein the valid region includes at least one non-zero coefficient; generating a bit stream including at least one of a parameter related to whether a multiple transform set is applicable to the encoding object block, a parameter indicating a transform type applicable to the encoding object block, and coefficient information, When the size of the encoding target block is width 64×height 64, width 64×height 32, or width 32×height 64, the valid area of the encoding target block is set to the upper left width 32×height 32 area of the encoding target block, and coefficient information of the remaining area of the encoding target block except the valid area is not sent to the image decoding device. When the size of the encoding target block is width 64×height 64, width 64×height 32, or width 32×height 64, the valid area is set regardless of the parameter related to whether the multiple transform set is applicable to the encoding target block. When the size of the encoding object block is width 32×height 32, width 32×height 16, or width 16×height 32 and a multiple transform set is applied to the encoding object block, the valid area of the encoding object block is set to the upper left 16×16 area of the encoding object block, When the size of the encoding target block is width 32×height 32, width 32×height 16, or width 16×height 32 and the multiple transform set is not applicable to the encoding target block, all areas of the encoding target block are the valid areas, The transformation type of the encoding object block is one of a first transformation type and a second transformation type, wherein the first transformation type includes a type indicating the same transformation kernel for a longitudinal direction and a transverse direction, and the second transformation type includes a type indicating different transformation kernels for a vertical direction and a horizontal direction.
3. A method for transmitting a bit stream encoded by the image encoding method according to claim 2, the method comprising the step of transmitting the bit stream.