Image coding method and apparatus based on non-separated secondary transformation.
The NSST-based image decoding method optimizes coding efficiency by deriving NSST indexes and coefficients, reducing bit rates and costs for high-resolution video transmission and storage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-08-07
- Publication Date
- 2026-05-26
Smart Images

Figure 0007866125000015 
Figure 0007866125000016 
Figure 0007866125000017
Abstract
Description
Technical Field
[0001] The present invention relates to image coding technology, and more particularly, to an image decoding method and apparatus by non-separable secondary conversion in an image coding system.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As the video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line or storing video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] As a result, a highly efficient video compression technology is required to effectively transmit, store, and reproduce high-resolution and high-quality video information.
Summary of the Invention
Problems to be Solved by the Invention
[0004] A technical problem of the present invention is to provide a method and apparatus for enhancing image coding efficiency.
[0005] Another technical problem of the present invention is to provide an image decoding method and apparatus for applying NSST to a target block.
[0006] Another technical problem of the present invention is to provide an image decoding method and apparatus for deriving the range of NSST indexes based on specific conditions of a target block.
[0007] Another technical problem of the present invention is to provide an image decoding method and apparatus for determining whether to code an NSST index based on the conversion coefficient of a target block. [Means for solving the problem]
[0008] According to one embodiment of the present invention, an image decoding method performed by a decoding device is provided. The method is characterized by comprising the steps of: deriving a transformation coefficient for a target block from a bitstream; deriving an NSST (Non-Separable Secondary Transform) index for the target block; deriving a residual sample of the target block by performing an inverse transform on the transformation coefficient of the target block based on the NSST index; and generating a restored picture based on the residual sample.
[0009] According to another embodiment of the present invention, a decoding device for image decoding is provided. The decoding device is characterized by including an entropy decoding unit that derives transformation coefficients for a target block from a bitstream and derives an NSST (Non-Separable Secondary Transform) index for the target block; an inverse transform unit that performs an inverse transform on the transformation coefficients of the target block based on the NSST index to derive a residual sample of the target block; and an additive unit that generates a restored picture based on the residual sample.
[0010] Another embodiment of the present invention provides a video encoding method performed by an encoding device. The method includes the steps of: deriving a residual sample of a target block; performing a transform on the residual sample to derive a transformation coefficient of the target block; determining whether an NSST index for the target block can be encoded; and encoding information regarding the transformation coefficient. The step of determining whether an NSST index can be encoded includes the steps of: scanning the R+1 to Nth transformation coefficients of the target block; and determining that the NSST index should not be encoded if a non-zero transformation coefficient is included among the R+1 to Nth transformation coefficients, wherein N is the number of samples in the upper left target region of the target block, R is a reduced coefficient, and R is smaller than N.
[0011] Another embodiment of the present invention provides a video encoding device. The encoding device includes an adder that derives the residual sample of a target block, a transformer that performs a transform on the residual sample to derive a transform coefficient of the target block, and an entropy encoding unit that determines whether or not the NSST index for the target block can be encoded and encodes information about the transform coefficient. The entropy encoding unit scans the R+1 to Nth transform coefficients of the target block, and if the R+1 to Nth transform coefficients include a non-zero transform coefficient, it is determined not to encode the NSST index, wherein N is the number of samples in the upper left target region of the target block, R is a reduced coefficient, and R is smaller than N. [Effects of the Invention]
[0012] According to the present invention, the range of the NSST index can be derived based on specific conditions of the target block, thereby reducing the amount of bits required for the NSST index and improving overall coding efficiency.
[0013] According to the present invention, the transmission of a syntax element to an NSST index is determined based on a conversion coefficient for the target block, thereby reducing the number of bits required for the NSST index and improving overall coding efficiency. [Brief explanation of the drawing]
[0014] [Figure 1] This diagram schematically illustrates the configuration of a video encoding device to which the present invention can be applied. [Figure 2] This shows an example of an image encoding method performed by a video encoding device. [Figure 3] This diagram schematically illustrates the configuration of a video decoding device to which the present invention can be applied. [Figure 4] An example of an image decoding method performed by a decoding device is shown. [Figure 5] A schematic representation of the multiple transformation technique according to the present invention is shown below. [Figure 6] Sixty-five intra-directional modes for predicting direction are shown as examples. [Figure 7a] This is a flowchart showing the coding process of conversion coefficients according to one embodiment. [Figure 7b] This is a flowchart showing the coding process of conversion coefficients according to one embodiment. [Figure 8] This figure illustrates the arrangement of conversion coefficients based on the target block according to an embodiment of the present invention. [Figure 9] Here is an example of scanning the conversion coefficients from R+1 to N. [Figure 10a] This is a flowchart showing the coding process for an NSST index according to one embodiment. [Figure 10b] It is a flowchart showing the coding process of the NSST index according to an embodiment. [Figure 11] An example of determining whether the NSST index is coded is shown. [Figure 12] An example of scanning the conversion coefficients from R + 1 to N for all components of the target block is shown. [Figure 13] The image encoding method by the encoding device according to the present invention is schematically shown. [Figure 14] The encoding device performing the image encoding method according to the present invention is schematically shown. [Figure 15] The image decoding method by the decoding device according to the present invention is schematically shown. [Figure 16] The decoding device performing the image decoding method according to the present invention is schematically shown.
Embodiments for Carrying Out the Invention
[0015] The present invention can be subjected to various modifications and can have various embodiments. Specific embodiments are illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments. The terms used in this specification are merely used to explain specific embodiments and are not intended to limit the technical idea of the present invention. Singular expressions include plural expressions unless the context clearly indicates a different meaning. In this specification, terms such as "including" or "having" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should not be understood that the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is precluded in advance.
[0016] On the other hand, each component shown in the drawings described in this invention is illustrated independently for the convenience of illustrating its distinct characteristic functions, and does not mean that each component is embodied in separate hardware or separate software. For example, two or more components may combine to form a single component, and a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the invention, as long as they do not deviate from the essence of the invention.
[0017] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. Hereafter, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.
[0018] On the other hand, the present invention relates to video / image coding, and for example, the methods / examples disclosed herein can be applied to methods disclosed in the VVC (versatile video coding) standard or next-generation video / image coding.
[0019] In this specification, "picture" generally refers to a unit representing a single image within a specific time period, and "slice" is a unit that constitutes a part of a picture in coding. A single picture may consist of multiple slices, and pictures and slices may be used interchangeably as needed.
[0020] A pixel or pel can refer to the smallest unit that makes up a picture (or image). The term "sample" can also be used as a counterpart to "pixel." A sample generally refers to a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance (luma) component, or only the pixel / pixel value of the chroma component.
[0021] A unit represents a basic unit of image processing. A unit can contain at least one of the following: a specific region of a picture and information about that region. A unit may be used interchangeably with terms such as block or area. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows.
[0022] Figure 1 is a schematic diagram illustrating the configuration of a video encoding device to which the present invention can be applied.
[0023] Referring to Figure 1, the video encoding device 100 may include a picture splitting unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a conversion unit 122, a quantization unit 123, a realignment unit 124, an inverse quantization unit 125, and an inverse conversion unit 126.
[0024] The picture splitting unit 105 can split the input picture into at least one processing unit.
[0025] For example, a processing unit is called a coding unit (CU). In this case, a coding unit can be recursively divided from the largest coding unit (LCU) by a quad-tree binary-tree (QTBT) structure. For example, a single coding unit can be divided into multiple coding units of deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure later, or the binary-tree structure may be applied first. Based on the final coding unit that is not further divided, the coding procedure according to the present invention can be executed. In this case, based on coding efficiency due to video characteristics, the largest coding unit may be used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of lower depth so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0026] As another example, a processing unit may include coding units (CU), prediction units (PU), or transform units (TU). A coding unit can be split from the largest coding unit (LCU) into deeper coding units using a quad-tree structure. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively split into even deeper coding units so that the optimally sized coding unit is used as the final coding unit. If a smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to the coding unit that forms the basis for partitioning or splitting into prediction units or transform units. A prediction unit is a unit partitioned from a coding unit and is a sample prediction unit. In this case, the prediction unit can also be divided into subblocks. A transformation unit can be separated from a coding unit by a quad-tree structure and is a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients. Hereinafter, coding units are also called coding blocks (CB), prediction units are called prediction blocks (PB), and transformation units are called transform blocks (TB). A prediction block or prediction unit refers to a specific region in block form within a picture and may contain an array of prediction samples.Furthermore, a conversion block or conversion unit refers to a specific region in block form within a picture and may include a conversion coefficient or an array of residual samples.
[0027] The prediction unit 110 performs a prediction on the block to be processed (hereinafter referred to as the current block) and can generate a predicted block that includes prediction samples for the current block. The unit of prediction performed by the prediction unit 110 is a coding block, or a transformation block, or a prediction block.
[0028] The prediction unit 110 can determine whether intra-prediction or inter-prediction is applied to the current block. For example, the prediction unit 110 can determine whether intra-prediction or inter-prediction is applied on a CU basis.
[0029] In the case of intra-prediction, the prediction unit 110 can guide a prediction sample for the current block based on a reference sample outside the current block within the picture to which the current block belongs (hereinafter referred to as the current picture). In this case, the prediction unit 110 can (i) guide a prediction sample based on the average or interpolation of neighboring reference samples of the current block, or (ii) guide a prediction sample based on a reference sample among the neighboring reference samples of the current block that is located in a specific (prediction) direction relative to the prediction sample. Case (i) is called a non-directional mode or non-angular mode, and case (ii) is called a directional mode or angular mode. The prediction modes in intra-prediction can, for example, have 33 directional prediction modes and at least 2 or more non-directional modes. Non-directional modes can include DC prediction modes and planar modes. The prediction unit 110 can also use the prediction modes applied to neighboring blocks to determine the prediction mode to be applied to the current block.
[0030] In interpretation, the prediction unit 110 can guide predicted samples for the current block based on samples identified by motion vectors on a reference picture. The prediction unit 110 can guide predicted samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP (motion vector prediction). In skip mode and merge mode, the prediction unit 110 can use motion information from adjacent blocks as motion information for the current block. In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted. In MVP mode, the motion vectors of adjacent blocks can be used as motion vector predictors for the current block to guide the motion vector of the current block.
[0031] In interpretation, adjacent blocks can include spatially adjacent blocks currently present in the picture and temporally adjacent blocks present in the reference picture. The reference picture containing the temporally adjacent blocks is also called a collocated picture (colPic). Motion information can include motion vectors and reference picture indices. Information such as prediction mode information and motion information can be (entropy) encoded and output in bitstream format.
[0032] In skip mode and merge mode, when motion information of temporally adjacent blocks is used, the top-level picture in the reference picture list can also be used as the reference picture. Reference pictures included in the picture order count can be sorted based on the difference in Picture Order Count (POC) between the current picture and the corresponding reference picture. The POC corresponds to the display order of the pictures and can be distinguished from the coding order.
[0033] The subtraction unit 121 generates a residual sample, which is the difference between the original sample and the predicted sample. When skip mode is applied, a residual sample is not generated, as described above.
[0034] The transformation unit 122 transforms residual samples in units of transformation blocks to generate transformation coefficients. The transformation unit 122 can perform the transformation depending on the size of the transformation block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the transformation block. For example, if intraprediction is applied to the coding block or prediction block that overlaps with the transformation block, and the transformation block is a 4x4 residual array, the residual samples are transformed using a Discrete Sine Transform (DST) transformation kernel. Otherwise, the residual samples can be transformed using a Discrete Cosine Transform (DCT) transformation kernel.
[0035] The quantization unit 123 can quantize the conversion coefficients and generate the quantized conversion coefficients.
[0036] The realignment unit 124 realigns the quantized transformation coefficients. The realignment unit 124 can realign the block-shaped quantized transformation coefficients into a one-dimensional vector form via a coefficient scanning method. Here, although the realignment unit 124 is described in a separate configuration, it may also be part of the quantization unit 123.
[0037] The entropy encoding unit 130 can perform entropy encoding on the quantized conversion coefficients. Entropy encoding can include encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). The entropy encoding unit 130 can also encode information necessary for video restoration (e.g., the values of syntax elements) in addition to the quantized conversion coefficients, either together or separately, by entropy encoding or a pre-configured method. The encoded information can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units.
[0038] The inverse quantization unit 125 inversely quantizes the values (quantized conversion coefficients) quantized by the quantization unit 123, and the inverse transformation unit 126 inversely transforms the values inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0039] The addition unit 140 reconstructs the picture by adding the residual sample and the predicted sample. The residual sample and the predicted sample can be added in block units to generate a reconstructed block. Here, although the addition unit 140 has been described in a separate configuration, it may also be part of the prediction unit 110. On the other hand, the addition unit 140 is also called the reconstruction module or the reconstructed block generation unit.
[0040] The filter unit 150 can apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through deblocking filtering and / or sample adaptive offset, artifacts at block boundaries and distortions during the quantization process within the reconstructed picture can be corrected. The sample adaptive offset can be applied on a sample-by-sample basis and can be applied after the deblocking filtering process is complete. The filter unit 150 can also apply an Adaptive Loop Filter (ALF) to the reconstructed picture. The ALF can be applied to the reconstructed picture after the deblocking filter and / or sample adaptive offset have been applied.
[0041] Memory 160 can store a restored picture (decoded picture) or information necessary for encoding / decoding. Here, the restored picture is a restored picture after the filtering procedure has been completed by the filter unit 150. The stored restored picture can be used as a reference picture for (inter)prediction of other pictures. For example, memory 160 can store a (reference) picture used for interprediction. In this case, the picture used for interprediction can be specified by a reference picture set or a reference picture list.
[0042] Figure 2 shows an example of an image encoding method performed by a video encoding device. As shown in Figure 2, the image encoding method includes the processes of intra / inter prediction, transformation, quantization, and entropy encoding. For example, intra / inter prediction generates a predicted block of the current block, and subtraction of the input block of the current block and the predicted block generates the residual block of the current block. Subsequently, a transformation of the residual block generates a coefficient block, i.e., the transformation coefficient of the current block. The transformation coefficient is quantized and entropy encoded and stored in a bitstream.
[0043] Figure 3 is a schematic diagram illustrating the configuration of a video decoding device to which the present invention can be applied.
[0044] As shown in Figure 3, the video decoding device 300 includes an entropy decoding unit 310, a resistive processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the resistive processing unit 320 may also include a realignment unit 321, an inverse quantization unit 322, and an inverse transformation unit 323.
[0045] When a bitstream containing video information is input, the video decoding device 300 can restore the video in accordance with the process by which the video information was processed in the video encoding device.
[0046] For example, the video decoding device 300 can perform video decoding using the processing units applied in the video encoding device. Therefore, a processing unit block for video decoding may, as an example, be a coding unit, and as an example, a coding unit, a prediction unit, or a conversion unit. A coding unit can be divided from a maximum coding unit according to a quad-tree structure and / or a binary tree structure.
[0047] Prediction units and transformation units may be used further, in which case the prediction block is a block derived or partitioned from the coding unit and may be a unit of sample prediction. Here, the prediction unit may also be divided into subblocks. The transformation unit is partitioned from the coding unit according to a quad-tree structure and may be a unit that derives transformation coefficients or a unit that derives a residual signal from transformation coefficients.
[0048] The entropy decoding unit 310 parses the bitstream and outputs information necessary for video or picture restoration. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for video restoration and the quantized values of conversion coefficients related to resistivity.
[0049] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntactic element in the bitstream, determines a context model using the information of the syntactic element to be decoded, the decoded information of the surrounding and decoded blocks, or symbol / bin information decoded in a previous stage, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntactic element. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.
[0050] The information decoded in the entropy decoding unit 310 that relates to predictions is provided to the prediction unit 330, and the residual values that have been entropy decoded in the entropy decoding unit 310, i.e., the quantized conversion coefficients, are input to the realignment unit 321.
[0051] The realignment unit 321 rearranges the quantized conversion coefficients into a two-dimensional block form. The realignment unit 321 can perform the realignment in response to the coefficient scanning performed in the encoding device. Here, the realignment unit 321 has been described as a separate configuration, but the realignment unit 321 may be part of the inverse quantization unit 322.
[0052] The inverse quantization unit 322 inversely quantizes the quantized conversion coefficients based on (inverse) quantization parameters and outputs the conversion coefficients. Here, information for inducing the quantization parameters is signaled from the encoding device.
[0053] The inverse transform unit 323 inversely transforms the transformation coefficients to induce a residual sample.
[0054] The prediction unit 330 performs a prediction on the current block and generates a predicted block that includes prediction samples for the current block. The unit of prediction performed by the prediction unit 330 can be a coding block, a transformation block, or a prediction block.
[0055] The prediction unit 330 determines whether to apply intra-prediction or inter-prediction based on the prediction information. Here, the unit for determining whether to apply intra-prediction or inter-prediction and the unit for generating prediction samples may be different. Similarly, the units for generating prediction samples in inter-prediction and intra-prediction may also be different. For example, the decision on whether to apply inter-prediction or intra-prediction can be made in CU units. Alternatively, for example, in inter-prediction, the prediction mode can be determined in PU units and prediction samples can be generated, while in intra-prediction, the prediction mode can be determined in PU units and prediction samples can be generated in TU units.
[0056] In the case of intra-prediction, the prediction unit 330 can guide prediction samples for the current block based on surrounding reference samples in the current picture. The prediction unit 330 can guide prediction samples for the current block by applying a directional mode or a non-directional mode based on surrounding reference samples of the current block. Here, the prediction mode to be applied to the current block can also be determined by using the intra-prediction mode of the surrounding block.
[0057] In interpretation, the prediction unit 330 can guide prediction samples for the current block based on samples identified on the reference picture by motion vectors on the reference picture. The prediction unit 330 guides prediction samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP mode. Here, motion information necessary for interpretation of the current block, such as motion vectors and reference picture indices, provided by the video encoding device, is acquired or guided based on the prediction information.
[0058] In skip mode and merge mode, the movement information of surrounding blocks may be used as the movement information of the current block. Here, surrounding blocks include spatially surrounding blocks and temporally surrounding blocks.
[0059] The prediction unit 330 constructs a merge candidate list as motion information of available surrounding blocks, and the merge index can use the information indicated on the merge candidate list as the motion vector of the current block. The merge index is signaled by the encoding device. The motion information may include motion vectors and reference pictures. When motion information of temporally surrounding blocks is used in skip mode and merge mode, the top-level picture on the reference picture list can be used as the reference picture.
[0060] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not sent.
[0061] In MVP mode, the motion vector of the current block is derived using the motion vectors of surrounding blocks as a motion vector predictor. Here, surrounding blocks include spatially surrounding blocks and temporally surrounding blocks.
[0062] For example, when merge mode is applied, a merge candidate list is generated using the motion vectors of the restored spatially surrounding blocks and / or the motion vectors corresponding to the temporally surrounding block, Col. In merge mode, the motion vectors of the candidate blocks selected in the merge candidate list are used as the motion vectors of the current block. The prediction information includes a merge index that indicates the candidate block with the selected optimal motion vector from among the candidate blocks included in the merge candidate list. Here, the prediction unit 330 uses the merge index to derive the motion vector of the current block.
[0063] As another example, when the MVP (Motion Vector Prediction) mode is applied, a list of motion vector predictor candidates is generated using the motion vectors of the restored spatially surrounding blocks and / or the motion vectors corresponding to the temporally surrounding block, Col block. That is, the motion vectors of the restored spatially surrounding blocks and / or the motion vectors corresponding to the temporally surrounding block, Col block, can be used as motion vector candidates. The prediction information includes a predicted motion vector index that indicates the optimal motion vector selected from the motion vector candidates included in the list. Here, the prediction unit 330 can use the motion vector index to select the predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list. The prediction unit of the encoding device calculates the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encodes it, and outputs it in bitstream format. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the prediction unit 330 obtains the motion vector difference included in the prediction information and derives the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit can also obtain or derive a reference picture index that indicates a reference picture from the information related to the prediction.
[0064] The adder 340 restores the current block or current picture by adding the residual sample and the predicted sample. The adder 340 can also restore the current picture by adding the residual sample and the predicted sample in block units. If skip mode is applied, the residual is not transmitted, so the predicted sample may become the restored sample. Here, the adder 340 is described as a separate configuration, but the adder 340 may be part of the prediction unit 330. On the other hand, the adder 340 may also be called the restoration unit or the restored block generation unit.
[0065] The filter unit 350 can apply deblocking filtering, sample-adaptive offset, and / or ALF to the restored picture. Here, the sample-adaptive offset is applied on a sample-by-sample basis and may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or sample-adaptive offset.
[0066] Memory 360 stores the restored picture (decoded picture) or information necessary for decoding. Here, the restored picture may be a restored picture that has undergone the filtering procedure by the filter unit 350. For example, memory 360 stores the picture used for inter prediction. Here, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The restored picture may also be used as a reference picture for other pictures. Memory 360 may also output the restored pictures in the order of output.
[0067] Figure 4 shows an example of an image decoding method performed by a decoding device. As shown in Figure 4, the image decoding method includes the processes of entropy decoding, inverse quantization, inverse transform, and intra / inter prediction. For example, the decoding device can perform the reverse process of the encoding method. Specifically, quantized transformation coefficients are obtained by entropy decoding of the bitstream, and the coefficient block of the current block, i.e., the transformation coefficients, is obtained by the inverse quantization process of the quantized transformation coefficients. The residual block of the current block is derived by the inverse transform of the transformation coefficients, and the reconstructed block of the current block is derived by adding the predicted block of the current block, which is derived by intra / inter prediction, and the residual block.
[0068] On the other hand, the conversion described above derives a lower frequency conversion coefficient for the current block relative to the resistive block, and a zero tail is derived at the end of the resistive block.
[0069] Specifically, the transformation consists of two main processes, the main processes including a core transform and a secondary transform. A transformation including the core transform and the secondary transform can be called a multiple transformation technique.
[0070] Figure 5 schematically shows a multiplexing technique according to the present invention.
[0071] As shown in Figure 5, the conversion unit corresponds to the conversion unit in the encoding device shown in Figure 1, and the inverse conversion unit corresponds to the inverse conversion unit in the encoding device shown in Figure 1 or the inverse conversion unit in the decoding device shown in Figure 3.
[0072] The transformation unit performs a linear transformation based on the residual samples (residual sample array) in the residual block to derive (linear) transformation coefficients (S510). Here, the linear transformation includes an Adaptive Multiple Core Transform (AMT). The Adaptive Multiple Core Transform may also be expressed as an MTS (Multiple Transform Set).
[0073] The adaptive multicore transformation described above refers to a method of transformation that additionally uses a Discrete Cosine Transform (DCT) type 2, a Discrete Sine Transform (DST) type 7, a Discrete Sine Transform (DST) type 8, and / or a DST type 1. That is, the adaptive multicore transformation describes a transformation method that transforms a spatial domain residual signal (or residual block) into frequency domain transformation coefficients (or first-order transformation coefficients) based on a plurality of transformation kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the first-order transformation coefficients may also be called temporary transformation coefficients from the perspective of the transformation unit.
[0074] In other words, when an existing conversion method is applied, a spatial-domain to frequency-domain conversion of a residual signal (or residual block) is applied based on DCT type 2 to generate conversion coefficients. In contrast, when the adaptive multiple core conversion is applied, a spatial-domain to frequency-domain conversion of a residual signal (or residual block) is applied based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., to generate conversion coefficients (or first-order conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. may also be called conversion types, conversion kernels, or conversion cores.
[0075] For reference, the DCT / DST conversion type is defined based on basis functions, which can be expressed as shown in the table below.
[0076] [Table 1]
[0077] When the adaptive multicore conversion is performed, a vertical conversion kernel and a horizontal conversion kernel are selected from the conversion kernels for the target block, a vertical conversion is performed on the target block based on the vertical conversion kernel, and a horizontal conversion is performed on the target block based on the horizontal conversion kernel. Here, the horizontal conversion represents a conversion to the horizontal component of the target block, and the vertical conversion represents a conversion to the vertical component of the target block. The vertical conversion kernel / horizontal conversion kernel is adaptively determined based on a conversion index that indicates the prediction mode and / or conversion subset of the target block (CU or subblock) encompassing the residual block.
[0078] For example, the adaptive multicore conversion is applied when both the width and height of the target block are less than or equal to 64, and whether or not the adaptive multicore conversion is applied to the target block can be determined based on the CU level flag. Specifically, if the CU level flag is 0, the aforementioned existing conversion method may be applied. That is, if the CU level flag is 0, a spatial domain to frequency domain conversion is applied to the residual signal (or residual block) based on the DCT type 2 to generate conversion coefficients, and these conversion coefficients are encoded. Here, the target block may be a CU. If the CU level flag is 0, the adaptive multicore conversion can be applied to the target block.
[0079] Furthermore, in the case of a rumor block of the target block to which the adaptive multicore transformation is applied, two additional flags are signaled, and a vertical transformation kernel and a horizontal transformation kernel are selected based on these flags. The flag for the vertical transformation kernel may also be represented as the AMT vertical flag, where AMT_TU_vertical_flag (or EMT_TU_vertical_flag) represents the syntax element of the AMT vertical flag. The flag for the horizontal transformation kernel may also be represented as the AMT horizontal flag, where AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) represents the syntax element of the AMT horizontal flag. The AMT vertical flag indicates one transformation kernel candidate from among the transformation subset candidates for the vertical transformation kernel, and the transformation kernel candidate indicated by the AMT vertical flag is derived as the vertical transformation kernel for the target block. Furthermore, the AMT horizontal flag indicates one of the candidate conversion kernels included in the conversion subset for the horizontal conversion kernel, and the candidate conversion kernel indicated by the AMT horizontal flag is derived as the horizontal conversion kernel for the target block. On the other hand, the AMT vertical flag may be represented as the MTS vertical flag, and the AMT horizontal flag may be represented as the MTS horizontal flag.
[0080] On the other hand, three transformation subsets have already been set, and one of these transformation subsets is derived as the transformation subset for the vertical transformation kernel based on the intra-prediction mode applied to the target block. Also, one of these transformation subsets is derived as the transformation subset for the horizontal transformation kernel based on the intra-prediction mode applied to the target block. For example, the already set transformation subsets are derived as shown in the table below.
[0081] [Table 2]
[0082] Referring to Table 2, the transformation subsets with an index value of 0 represent transformation subsets that include DST type 7 and DCT type 8 as transformation kernel candidates, and the transformation subsets with an index value of 1 represent transformation subsets that include DST type 7 and DCT type 8 as transformation kernel candidates.
[0083] The transformation subsets for the vertical transformation kernel and the transformation subsets for the horizontal transformation kernel, derived based on the intra-prediction mode applied to the target block, are derived as shown in the table below.
[0084] [Table 3]
[0085] Here, V represents a subset of the transformations for the vertical transformation kernel, and H represents a subset of the transformations for the horizontal transformation kernel.
[0086] If the value of the AMT flag (or EMT_Cu_flag) for the target block is 1, a transformation subset for the vertical transformation kernel and a transformation subset for the horizontal transformation kernel are derived based on the intra-prediction mode of the target block, as shown in Table 3. Thereafter, among the transformation kernel candidates included in the transformation subset for the vertical transformation kernel, the transformation kernel candidate indicated by the AMT vertical flag of the target block is derived as the vertical transformation kernel of the target block, and the horizontal transformation kernel candidate is derived as the horizontal transformation kernel of the target block. On the other hand, the AMT flag may also be represented as the MTS flag.
[0087] For reference, for example, the intra-prediction mode includes two non-directional (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. The non-directional intra-prediction mode includes planar intra-prediction mode number 0 and DC intra-prediction mode number 1, and the directional intra-prediction mode includes 65 intra-prediction modes numbered 2 through 66. However, this is illustrative, and the present invention also applies when the number of intra-prediction modes differs. On the other hand, in some cases an additional intra-prediction mode number 67 may be used, and this 67th intra-prediction mode may represent a linear model (LM) mode.
[0088] Figure 6 illustrates 65 intradirectional modes for predicting directions.
[0089] As shown in Figure 6, intra-prediction modes can be divided into those with horizontal directionality and those with vertical directionality, centered around intra-prediction mode 34, which has a diagonal prediction direction pointing diagonally upward to the left. In Figure 6, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 units on the sample grid position. Intra-prediction modes 2 through 33 have horizontal directionality, while intra-prediction modes 34 through 66 have vertical directionality. Intra-prediction modes 18 and 50 represent the horizontal intra-prediction mode and the vertical intra-prediction mode, respectively. Intra-prediction mode 2 may also be called the diagonal intra-prediction mode pointing diagonally downward to the left, intra-prediction mode 34 the diagonal intra-prediction mode pointing diagonally upward to the left, and intra-prediction mode 66 the diagonal mode pointing upward to the right.
[0090] The conversion unit performs a quadratic transformation based on the (primary) transformation coefficients to derive (secondary) transformation coefficients (S520). If the primary transformation was a transformation from the spatial domain to the frequency domain, then the secondary transformation can be said to be a transformation from the frequency domain to the frequency domain. The secondary transformation includes a non-separable transform. In this case, the secondary transformation may be called a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The non-separable secondary transform represents a transformation that generates transformation coefficients (or secondary transformation coefficients) for a residual signal by performing a quadratic transformation on the (primary) transformation coefficients derived by the primary transformation based on a non-separable transform matrix. Here, the transformation can be applied at once to the (primary) transformation coefficients based on the non-separable transform matrix without separating the vertical and horizontal transformations (or applying the horizontal and vertical transformations independently). In other words, the non-separable quadratic transformation refers to a transformation method that generates transformation coefficients (or quadratic transformation coefficients) by transforming the vertical and horizontal components of the (linear) transformation coefficients together without separating them, based on the non-separable transformation matrix. The non-separable quadratic transformation is applied to the top-left region of a block composed of (linear) transformation coefficients (hereinafter, it may be called a transformation coefficient block or target block). For example, if both the width (W) and height (H) of the transformation coefficient block are 8 or greater, an 8x8 non-separable quadratic transformation is applied to the top-left 8x8 region of the transformation coefficient block (hereinafter, the top-left target region). Also, if both the width (W) and height (H) of the transformation coefficient block are 4 or greater, and either the width (W) or height (H) of the transformation coefficient block is less than 8, a 4x4 non-separable quadratic transformation is applied to the top-left min(8,W) x min(8,H) region of the transformation coefficient block.
[0091] Specifically, for example, when a 4x4 input block is used, the unseparated quadratic transform is performed as follows:
[0092] The aforementioned 4x4 input block X is represented as follows:
[0093]
number
[0094] When X is shown in vector form, JPEG0007866125000005.jpg84 is represented as follows:
[0095]
number
[0096] In this case, the aforementioned second-order inseparable transform is calculated as follows:
[0097]
number
[0098] Here, JPEG0007866125000008.jpg75 shows the transformation coefficient vector, and T shows the 16×16 (unseparable) transformation metric.
[0099] The above formula 3 gives a 16 × 1 transformation coefficient vector JPEG0007866125000009.jpg75 is derived, and the above JPEG0007866125000010.jpg75 is reorganized into 4x4 blocks according to the scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is illustrative, and to reduce the computational complexity of inseparable quadratic transforms, methods such as HyGT (Hypercube-Givens Transform) may also be used for calculating inseparable quadratic transforms.
[0100] On the other hand, the unseparated quadratic conversion can be mode-dependent in which the conversion kernel (or conversion core, conversion type) is selected. Here, the mode includes intra-predictive mode and / or inter-predictive mode.
[0101] As described above, the non-separable quadratic transformation is performed based on an 8x8 transformation or a 4x4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. That is, the non-separable quadratic transformation is performed based on an 8x8 subblock size or a 4x4 subblock size. For example, for the selection of the mode-based transformation kernel, 35 sets of non-separable quadratic transformation kernels are configured for both the 8x8 subblock size and the 4x4 subblock size, with three kernels for each. That is, 35 transformation sets are configured for the 8x8 subblock size, and 35 transformation sets are configured for the 4x4 subblock size. In this case, each of the 35 transformation sets for the 8x8 subblock size contains three 8x8 transformation kernels, and each of the 35 transformation sets for the 4x4 subblock size contains three 4x4 transformation kernels. However, the conversion subblock size, the number of sets, and the number of conversion kernels within a set are examples only, and sizes other than 8x8 or 4x4 may be used, or n sets may be configured, with each set containing k conversion kernels.
[0102] The aforementioned conversion set may also be called an NSST set, and the conversion kernel within the NSST set may also be called an NSST kernel. The selection of a particular set from among the aforementioned conversion sets is performed, for example, based on the intra-prediction mode of the target block (CU or subblock).
[0103] In this case, the mapping between the 35 transformation sets and the intra-prediction mode is shown, for example, in the table below. For reference, when the LM mode is applied to a target block, the quadratic transformation may not be applied to that target block.
[0104] [Table 4]
[0105] On the other hand, once it is determined that a specific set is to be used, one of the k transformation kernels within that specific set is selected using an unseparated quadratic transformation index. The encoding device derives an unseparated quadratic transformation index indicating the specific transformation kernel using an RD (rate-distortion) check board and signals the decoding device with the unseparated quadratic transformation index. The decoding device selects one of the k transformation kernels within the specific set based on the unseparated quadratic transformation index. For example, an NSST index value of 0 indicates the first unseparated quadratic transformation kernel, an NSST index value of 1 indicates the second unseparated quadratic transformation kernel, and an NSST index value of 2 indicates the third unseparated quadratic transformation kernel. Alternatively, an NSST index value of 0 indicates that the first unseparated quadratic transformation is not applied to the target block, and NSST index values of 1 to 3 indicate the three transformation kernels.
[0106] Referring again to Figure 5, the transformation unit can perform the non-separable quadratic transformation based on the selected transformation kernel and obtain (quadratic) transformation coefficients. These transformation coefficients are derived as quantized transformation coefficients by the quantization unit as described above, encoded, and transmitted to the decoding unit for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.
[0107] On the other hand, if the quadratic transformation is omitted, the (linear) transformation coefficients, which are the output of the linear (separated) transformation, are derived as quantized transformation coefficients by the quantization unit as described above, encoded, and transmitted to the decoder for signaling and to the inverse quantization / inverse transformation unit within the encoding unit.
[0108] The inverse transformer performs a series of steps in the reverse order of the steps performed in the transformer described above. The inverse transformer receives the (inversely quantized) transform coefficients, performs a quadratic (inverse) transform to derive (primary) transform coefficients (S550), and performs a primary (inverse) transform on the (primary) transform coefficients to obtain a residual block (residual sample). Here, the primary transform coefficients may also be called modified transform coefficients from the perspective of the inverse transformer. As described above, the encoding and decoding devices can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.
[0109] On the other hand, as mentioned above, if the quadratic (inverse) transformation is omitted, the (inversely quantized) transformation coefficients are received and the linear (separated) transformation is performed to obtain a residual block (residual sample). As mentioned above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the predicted block, and generate a reconstructed picture based on this.
[0110] On the other hand, the aforementioned non-separable quadratic transformation may not be applied to blocks coded in transformation-skip mode. For example, if an NSST index is signaled for a target CU and the value of the NSST index is not zero, the non-separable quadratic transformation may not be applied to blocks coded in transformation-skip mode within the target CU. Also, if the target CU, which includes blocks of all components (luma components, chroma components, etc.), is coded in transformation-skip mode, or if the number of non-zero transformation coefficients among the transformation coefficients for the target CU is less than 2, the NSST index may not be signaled. The specific process for coding the transformation coefficients is as follows.
[0111] Figures 7a and 7b are flowcharts showing the coding process of conversion coefficients according to one embodiment.
[0112] Each step disclosed in Figures 7a and 7b is performed by the encoding device 100 or decoding device 300 disclosed in Figures 1 and 3, and more specifically, by the entropy encoding unit 130 disclosed in Figure 1 and the entropy decoding unit 310 disclosed in Figure 3. Therefore, specific details that overlap with what has been described above in Figure 1 or Figure 3 are omitted or simplified in their explanation.
[0113] This specification uses terms or phrases to define specific information or concepts. For example, in this specification, "a flag indicating whether or not there is at least one non-zero conversion coefficient among the conversion coefficients for a target block" is expressed as cbf. However, since "cbf" can be replaced with various other terms such as coded_block_flag, when interpreting terms or phrases used to define specific information or concepts in this specification throughout the specification, the interpretation should not be limited to the names themselves, but rather should take into account the diverse actions, functions, and effects that result from the meaning of the terms.
[0114] Figure 7a shows the encoding process of the conversion coefficients.
[0115] In one embodiment, the encoding device 100 determines whether a flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is 1 (S700). If the flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is 1, then there is at least one non-zero conversion coefficient among the conversion coefficients for the target block. Conversely, if the flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is 0, then all conversion coefficients for the target block are 0.
[0116] A flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for a target block is, for example, expressed as the cbf flag. The cbf flag includes the cbf_luma[x0][y0][trafoDepth] flag for luma blocks and the cbf_cb[x0][y0][trafoDepth] and cbf_cr[x0][y0][trafoDepth] flags for chroma blocks. Here, array indices x0 and y0 represent the position of the top-left luma / chroma sample of the target block relative to the top-left luma / chroma sample of the picture, and array index trafoDepth may represent the level at which the coding block has been divided for the purpose of conversion coding. A block with trafoDepth of 0 corresponds to a coding block, and if the coding block and the conversion block are defined identically, trafoDepth is considered to be 0.
[0117] In one embodiment, the encoding device 100 encodes information regarding the conversion coefficients for the target block if, in S700, the flag indicating whether or not there is at least one non-zero conversion coefficient for the target block is set to 1 (S710).
[0118] Information regarding the conversion coefficients for the target block includes, for example, information regarding the position of the last non-zero conversion coefficient, group flag information indicating whether or not a subgroup of the target block contains a non-zero conversion coefficient, and information regarding the simplification coefficient, at least one of these. A detailed explanation of each piece of information will be provided later.
[0119] An encoding device 100 according to one embodiment determines whether the conditions for performing NSST are met (S720). More specifically, the encoding device 100 determines whether the conditions for encoding an NSST index are met. Here, the NSST index may be called, for example, a transform index.
[0120] In one embodiment, the encoding device 100 encodes the NSST index if it determines in S720 that the conditions for performing NSST are met (S730). More specifically, the encoding device 100 encodes the NSST index if it determines that the conditions for encoding the NSST index are met.
[0121] In one embodiment, the encoding device 100 can omit the operations according to S710, S720, and S730 if the flag indicating whether or not there is at least one non-zero conversion coefficient among the conversion coefficients for the target block in S700 is 0.
[0122] Furthermore, in one embodiment, if the encoding device 100 determines in S720 that the conditions for performing NSST are not met, the operation according to S730 can be omitted.
[0123] Figure 7b shows the decoding process of the conversion coefficients.
[0124] In one embodiment, the decoding device 300 determines whether a flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is 1 (S740). If the flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is 1, then there is at least one non-zero conversion coefficient among the conversion coefficients for the target block. Conversely, if the flag indicating whether there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is 0, then all conversion coefficients for the target block are 0.
[0125] In one embodiment, the decoding device 300 decodes information regarding the conversion coefficients for the target block (S750) if the flag indicating whether or not there is at least one non-zero conversion coefficient among the conversion coefficients for the target block is set to 1 in S740.
[0126] According to one embodiment, the decoding device 300 determines whether or not the conditions for performing NSST are met (S760). More specifically, the decoding device 300 determines whether or not the conditions for decoding the NSST index from the bitstream are met.
[0127] In one embodiment, the decoding device 300 determines in S760 that the conditions for performing NSST are met, and decodes the NSST index (S770).
[0128] In one embodiment, the decoding device 300 can omit the operations according to S750, S760, and S770 if the flag indicating whether or not there is at least one non-zero conversion coefficient among the conversion coefficients for the target block in S740 is 0.
[0129] Furthermore, in one embodiment, if the decoding device 300 determines in S760 that the conditions for performing NSST are not met, the operation according to S770 can be omitted.
[0130] As mentioned above, signaling the NSST index when NSST is not performed can reduce coding efficiency. Furthermore, different coding methods for the NSST index depending on specific conditions can improve overall image coding efficiency. Therefore, the present invention proposes various NSST index coding methods.
[0131] As an example, the NSST index range is determined based on specific conditions. In other words, the range of values for the NSST index can be determined based on specific conditions. Specifically, the maximum value of the NSST index is determined based on the specific conditions.
[0132] For example, the range of the NSST index value is determined based on the block size. Here, the block size is defined as minimum(W,H). W represents the width, and H represents the height. In this case, the range of the NSST index value is determined by comparing the width of the target block with W, and comparing the height of the target block with minimumH.
[0133] Alternatively, the block size is defined as the number of samples in the block (W*H). In this case, the range of the NSST index value is determined by comparing the number of samples in the target block (W*H) with a specific value.
[0134] Furthermore, for example, the range of the NSST index value can be determined based on the shape of the block, i.e., the block type. Here, the block type is defined as a square block or a non-square block. In this case, the range of the NSST index value is determined based on whether the target block is a square block or a non-square block.
[0135] Alternatively, the block type is defined as the ratio of the longest side (the longer side of the width and height) to the shortest side of the block. In this case, the range of the NSST index value is determined by comparing the ratio of the longest side to the shortest side of the target block with a previously set critical value (for example, 2 or 3). Here, the ratio represents the value obtained by dividing the longest side by the shortest side. For example, if the width of the target block is longer than its height, the range of the NSST index value is determined by comparing the value obtained by dividing the width by the height with the previously set critical value. Also, if the height of the target block is longer than its width, the range of the NSST index value is determined by comparing the value obtained by dividing the height by the width with the previously set critical value.
[0136] Furthermore, for example, the range of the NSST index value is determined based on the intra-prediction mode applied to the block. As an example, the range of the NSST index value is determined based on whether the intra-prediction mode applied to the target block is a non-directional intra-prediction mode or a directional intra-prediction mode.
[0137] Alternatively, as another example, the range of the NSST index value is determined based on whether the intra prediction mode applied to the target block is an intra prediction mode included in Category A or Category B. Here, as an example, Category A includes intra prediction modes 2, 10, 18, 26, 34, 42, 50, 58, and 66, and Category B includes intra prediction modes other than those included in Category A. The intra prediction modes included in Category A may already be set, and Category A and Category B may already be set to include intra prediction modes different from the examples described above.
[0138] Alternatively, as another example, the range of the NSST index value is determined based on the block's AMT factor. The AMT factor may also be denoted as the MTS factor.
[0139] For example, the AMT factor may be defined as the aforementioned AMT flag. In this case, the range of the NSST index value is determined based on the value of the AMT flag of the target block.
[0140] Alternatively, the AMT factor may be defined as the aforementioned AMT vertical flag and / or AMT horizontal flag. In this case, the range of the NSST index value is determined based on the values of the AMT vertical flag and / or AMT horizontal flag of the target block.
[0141] Alternatively, the AMT factor may be defined as the transformation kernel applied in the multicore transformation. In this case, the range of the NSST index value is determined based on the transformation kernel applied in the multicore transformation of the target block.
[0142] Alternatively, as another example, the range of the NSST index value is determined based on the components of the block. For example, the range of the NSST index value for the rumor block of the target block and the range of the NSST index value for the chroma block of the target block may be applied differently.
[0143] On the other hand, the range of the NSST index value may also be determined by a combination of the specific conditions mentioned above.
[0144] The range of values for the NSST index determined based on the aforementioned specific conditions, i.e., the maximum value of the NSST index, can be set in a variety of ways.
[0145] For example, based on the specified conditions, the maximum value of the NSST index is determined to be R1, R2, or R3. Specifically, if the specified conditions fall under category A, the maximum value of the NSST index is derived as R1; if the specified conditions fall under category B, the maximum value of the NSST index is derived as R2; and if the specified conditions fall under category C, the maximum value of the NSST index is derived as R3.
[0146] R1 for category A, R2 for category B, and R3 for category C are derived as shown in the table below.
[0147] [Table 5]
[0148] R1, R2, and R3 may already be set. For example, the relationship between R1, R2, and R3 can be derived as shown in the following formula.
[0149]
number
[0150] Referring to Equation 4, R1 is greater than or equal to 0, R2 is greater than R1, and R3 is greater than R2. On the other hand, if R1 is 0 and the maximum value of the NSST index for the target block is determined to be R1, the NSST index is not signaled, and the value of the NSST index is derived (inferred) as 0.
[0151] Furthermore, the present invention proposes an implicit NSST index coding method.
[0152] Generally, when NSST is applied, the distribution of non-zero transformation coefficients can be altered. In particular, when RST (reduced secondary transform) is used as a secondary transformation under certain conditions, NSST indices may not be coded.
[0153] Here, RST represents a quadratic transformation that uses a simplified transformation matrix as an inseparable transformation matrix, where the simplified transformation matrix is determined by mapping an N-dimensional vector to an R-dimensional vector located in another space, where R is less than N. N represents the square of the side length of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplified factor may represent the R / N value. The simplified factor is referred to by various terms such as reduced factor, reduction factor, simplified factor, and simple factor. R may also be called the reduced coefficient, although in some cases the simplified factor may mean R. Furthermore, in some cases the simplified factor may represent the N / R value.
[0154] The size of the simplified transformation matrix according to one embodiment is R×N, which is smaller than the size N×N of a normal transformation matrix, and is defined as shown in the following equation 5.
[0155]
number
[0156] Simplified transformation matrix T for the transformation coefficients to which the linear transformation of the target block has been applied. R×N When multiplied, a (quadratic) transformation coefficient for the target block is derived.
[0157] When the aforementioned RST is applied, a simplified transformation matrix of size R × N is applied to the quadratic transformation, so the transformation coefficients from R+1 to N can implicitly be 0. In other words, if the transformation coefficients of the target block are derived by applying the aforementioned RST, the values of the transformation coefficients from R+1 to N can be 0. Here, the transformation coefficients from R+1 to N refer to the transformation coefficients from the R+1th to the Nth among the transformation coefficients. Specifically, the array of transformation coefficients of the target block can be described as follows.
[0158] Figure 8 is a diagram illustrating the arrangement of transformation coefficients based on a target block according to an embodiment of the present invention. Hereinafter, the explanation of the transformation described in Figure 8 applies equally to the inverse transformation. An NSST (an example of a quadratic transformation) based on a linear transformation and a simplified transformation is performed on the target block (or residual block) 800. In one example, the 16×16 block shown in Figure 8 represents the target block 800, and the 4×4 blocks labeled A to P represent subgroups of the target block 800. The linear transformation is performed over the entire range of the target block 800, and after the linear transformation is performed, the NSST is applied to the 8×8 block (hereinafter, the upper left target region) composed of subgroups A, B, E, and F. Here, when an NSST based on a simplified transformation is performed, only R NSST transformation coefficients (where R means the simplified coefficient, and R is less than N) are derived, so the NSST transformation coefficients in the range from the R+1th to the Nth are each determined to be 0. If R is, for example, 16, the 16 transformation coefficients derived by NSST based on the simplified transformation are assigned to each block in subgroup A, which is the upper left 4x4 block included in the upper left target region of the target block 800, and a transformation coefficient of 0 is assigned to each of the NR blocks, i.e., 64-16=48 blocks, included in subgroups B, E, and F. The linear transformation coefficients for which NSST based on the simplified transformation is not performed are assigned to each block in subgroups C, D, G, H, I, J, K, L, M, N, O, and P.
[0159] Therefore, if scanning the conversion coefficients from R+1 to N yields at least one non-zero conversion coefficient, it is determined that the RST does not apply, and the value of the NSST index can implicitly become 0 without any further signaling. In other words, if scanning the conversion coefficients from R+1 to N yields at least one non-zero conversion coefficient, the RST does not apply, and the value of the NSST index is derived as 0 without any further signaling.
[0160] Figure 9 shows an example of scanning the conversion coefficients from R+1 to N.
[0161] As shown in Figure 9, the size of the target block to which the transformation is applied may be 64 × 64, and R = 16 (i.e., R / N = 16 / 64 = 1 / 4). That is, Figure 9 shows the upper left target region of the target block. A 16 × 64 size simplified transformation matrix may be applied to the quadratic transformation of the 64 samples in the upper left target region of the target block. In this case, when the RST is applied to the upper left target region, the values of the transformation coefficients from 17 to 64 (N) must be 0. In other words, if at least one non-zero transformation coefficient is derived among the transformation coefficients from 17 to 64 of the target block, the RST is not applied, and the value of the NSST index is derived as 0 without further signaling. Therefore, the decoding device decodes the conversion coefficients of the target block, scans the decoded conversion coefficients from 17 to 64, and if a non-zero conversion coefficient is derived, the value of the NSST index can be derived as 0 without separate syntax element signaling for the NSST index. On the other hand, if there are no non-zero conversion coefficients among the conversion coefficients from 17 to 64, the decoding device can receive and decode the NSST index.
[0162] Figures 10a and 10b are flowcharts showing the coding process of an NSST index according to one embodiment.
[0163] Figure 10a shows the encoding process of the NSST index.
[0164] The encoding device encodes the conversion coefficients for the target block (S1000). The encoding device performs entropy encoding on the quantized conversion coefficients. Entropy encoding includes, for example, encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding).
[0165] The encoding device determines whether or not an (explicit) NSST index for the target block is coded (S1010). Here, the explicit NSST index refers to the NSST index that is transmitted to the decoding device. That is, the encoding device can determine whether or not to generate an NSST index to be signaled. In other words, the encoding device can determine whether or not to allocate bits for syntax elements for the NSST index. If the decoding device can derive the value of the NSST index even if the NSST index is not signaled, as in the embodiment described above, the encoding device may not code the NSST index. The specific process for determining whether or not an NSST index is coded will be described later.
[0166] If it is determined that the (explicit) NSST index is to be coded, the encoding device encodes the NSST index (S1020).
[0167] Figure 10b shows the decoding process of the NSST index.
[0168] The decoding device decodes the conversion coefficient for the target block (S1030).
[0169] The decoding device determines whether or not an (explicit) NSST index is coded for the target block (S1040). Here, the explicit NSST index refers to the NSST index signaled by the encoding device. In the embodiments described above, if the decoding device can derive the value of the NSST index even if it is not signaled, the encoding device may not signal the NSST index. The specific process for determining whether or not an NSST index is coded will be described later.
[0170] If it is determined that the aforementioned (explicit) NSST index is to be coded, the encoding device decodes the NSST index (S1040).
[0171] Figure 11 shows an example of how to determine whether or not an NSST index is coded.
[0172] The encoding / decoding device determines whether the conditions for coding an NSST index for the target block are met (S1100). For example, if the cbf flag for the target block indicates 0, the encoding / decoding device determines not to code an NSST index for the target block. Alternatively, if the target block is coded in conversion skip mode, or if the number of non-zero conversion coefficients among the conversion coefficients for the target block is less than a previously set critical value, the encoding / decoding device determines not to code an NSST index for the target block. For example, the previously set critical value may be 2.
[0173] If the conditions for coding an NSST index for the target block are met, the encoding / decoding device scans for conversion coefficients from R+1 to N (S1110). The conversion coefficients from R+1 to N represent the conversion coefficients from the R+1th to the Nth in the scan order of the conversion coefficients.
[0174] The encoding / decoding device determines whether a non-zero conversion coefficient is derived from the conversion coefficients R+1 to N (S1120). If a non-zero conversion coefficient is derived from the conversion coefficients R+1 to N, the encoding / decoding device determines not to code an NSST index for the target block. In this case, the encoding / decoding device can derive an NSST index value of 0 for the target block. In other words, for example, if an NSST index with a value of 0 indicates that NSST is not applied, the encoding / decoding device may not perform NSST on the upper left target region of the target block.
[0175] On the other hand, if no non-zero conversion coefficients are derived among the conversion coefficients from R+1 to N, the encoding device encodes the NSST index for the target block, and the decoding device decodes the NSST index for the target block.
[0176] On the other hand, it is proposed to use the aforementioned NSST index, which has common components (luma component, chroma Cb component, chroma Cr component) in the target block.
[0177] For example, the same NSST index is used for the chroma Cb block and the chroma Cr block of the target block. Another example is that the same NSST index is used for the lumen block, the chroma Cb block, and the chroma Cr block of the target block.
[0178] If two or three components of the target block use the same NSST index, the encoding device scans the conversion coefficients from R+1 to N for all components (luma block, chroma Cb block, and chroma Cr block of the target block), and if at least one non-zero conversion coefficient is derived, the NSST index is not encoded, and the value of the NSST index is derived as 0. Similarly, the decoding device scans the conversion coefficients from R+1 to N for all components (luma block, chroma Cb block, and chroma Cr block of the target block), and if at least one non-zero conversion coefficient is derived, the NSST index is not decoded, and the value of the NSST index is derived as 0.
[0179] Figure 12 shows an example of scanning the conversion coefficients from R+1 to N for all components of the target block.
[0180] As shown in Figure 12, the size of the lumar block, chroma Cb block, and chroma Cr block to which the transformation is applied is 64 × 64, and R = 16 (i.e., R / N = 16 / 64 = 1 / 4). That is, Figure 12 shows the upper left corner target region of the lumar block, the upper left corner target region of the chroma Cb block, and the upper left corner target region of the chroma Cr block. Therefore, a 16 × 64 size simplified transformation matrix can be applied to the quadratic transformation of the 64 samples in each of the upper left corner target regions of the lumar block, the chroma Cb block, and the chroma Cr block. In this case, when the RST is applied to the upper left corner target regions of the lumar block, the chroma Cb block, and the chroma Cr block, the values of the transformation coefficients from 17 to 64(N) for each block must be 0. In other words, if at least one non-zero conversion coefficient is derived from the conversion coefficients 17 to 64 of each block, the RST is not applied, and the value of the NSST index is derived as 0 without any further signaling. Accordingly, the decoding device decodes the conversion coefficients for all components of the target block, scans the conversion coefficients 17 to 64 of the lumen block, chroma Cb block, and chroma Cr block from the decoded conversion coefficients, and if a non-zero conversion coefficient is derived, it derives the value of the NSST index as 0 without any further syntax element signaling for the NSST index. On the other hand, if there are no non-zero conversion coefficients among the conversion coefficients 17 to 64, the decoding device can receive and decode the NSST index. The NSST index is used as an index for the lumen block, chroma Cb block, and chroma Cr block.
[0181] Furthermore, the present invention proposes a method for signaling an NSST index indicator at a higher level. NSST_Idx_indicator represents the syntax element for the NSST index indicator. For example, the NSST index indicator is coded at the CTU (Coding Tree Unit) level, and the NSST index indicator indicates whether or not NSST is applicable to the target CTU. That is, the NSST index indicator indicates whether or not NSST is available to the target CTU. Specifically, when the NSST index indicator for the target CTU is activated (enabled) (when NSST is available to the target CTU), i.e., when the value of the NSST index indicator is 1, an NSST index for the CU or TU included in the target CTU is coded. When the NSST index indicator for the target CTU is not activated (when NSST is not available to the target CTU), i.e., when the value of the NSST index indicator is 0, an NSST index for the CU or TU included in the target CTU is not coded. On the other hand, the NSST index indicator can be coded at the CTU level, as described above, and can also be coded at any other sample group level of arbitrary size. For example, the NSST index indicator can also be coded at the CU (Coding Unit) level.
[0182] Figure 13 schematically shows an image encoding method using an encoding device according to the present invention. The method disclosed in Figure 13 can be performed using the encoding device disclosed in Figure 1. Specifically, for example, S1300 in Figure 13 can be performed by the subtraction unit of the encoding device, S1310 by the conversion unit of the encoding device, and S1320 to S1330 by the entropy encoding unit of the encoding device. Although not shown, the process of deriving the predicted sample can be performed by the prediction unit of the encoding device.
[0183] The encoding device derives the residual sample of the target block (S1300). For example, the encoding device decides whether to perform interpretation or intrapretation on the target block, and determines a specific interpretation mode or specific intrapretation mode on the RD cost basis. According to the determined mode, the encoding device derives a predicted sample for the target block, and derives the residual sample by adding the original sample and the predicted sample for the target block.
[0184] The encoding device performs a conversion on the residual sample to derive the conversion coefficient for the target block (S1310). The encoding device determines whether or not NSST is applicable to the target block.
[0185] When the NSST is applied to the target block, the encoding device performs a core transformation on the residual sample to derive a modified transformation coefficient, and then performs an NSST on the modified transformation coefficient located in the upper left target region of the target block based on the simplified transformation matrix to derive the transformation coefficient of the target block. Modified transformation coefficients other than those located in the upper left target region of the target block are derived as the transformation coefficient of the target block. The size of the simplified transformation matrix is R × N, where N is the number of samples in the upper left target region, and R is the reduced coefficient, where R is smaller than N.
[0186] Specifically, the core transformation to the residual sample is performed as follows: The encoding device can determine whether or not to apply Adaptive Multiple Core Transform (AMT) to the target block. In this case, an AMT flag is generated indicating whether or not the Adaptive Multiple Core Transform is applied to the target block. If the AMT is not applied to the target block, the encoding device derives DCT type 2 as the transformation kernel for the target block, and performs the transformation to the residual sample based on the DCT type 2 to derive the modified transformation coefficient.
[0187] When the AMT is applied to the target block, the encoding device configures a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, derives the horizontal conversion kernel and the vertical conversion kernel based on the conversion subset, and performs a conversion on the residual sample based on the horizontal conversion kernel and the vertical conversion kernel to derive the corrected conversion coefficients. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. In addition, conversion index information can be generated, and the conversion index information includes an AMT horizontal flag indicating the horizontal conversion kernel and an AMT vertical flag indicating the vertical conversion kernel. On the other hand, the conversion kernel may be called a conversion type or conversion core.
[0188] On the other hand, if the NSST is not applied to the target block, the encoding device can perform a core conversion on the residual sample to derive the conversion coefficient for the target block.
[0189] Specifically, the core transformation to the residual sample is performed as follows: The encoding device determines whether or not to apply Adaptive Multiple Core Transform (AMT) to the target block. In this case, an AMT flag is generated indicating whether or not the Adaptive Multiple Core Transform is applied to the target block. If the AMT is not applied to the target block, the encoding device derives DCT type 2 as the transformation kernel for the target block, and performs the transformation to the residual sample based on the DCT type 2 to derive the transformation coefficient.
[0190] When the AMT is applied to the target block, the encoding device configures a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, derives the horizontal conversion kernel and the vertical conversion kernel based on the conversion subset, and performs the conversion to the residual sample based on the horizontal conversion kernel and the vertical conversion kernel to derive the conversion coefficients. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. In addition, conversion index information can be generated, and the conversion index information includes an AMT horizontal flag indicating the horizontal conversion kernel and an AMT vertical flag indicating the vertical conversion kernel. On the other hand, the conversion kernel may be called a conversion type or conversion core.
[0191] The encoding device determines whether or not to encode the NSST index (S1320).
[0192] As an example, the encoding device scans the conversion coefficients of the target block from the R+1th to the Nth conversion coefficient, and if the R+1th to the Nth conversion coefficients include a non-zero conversion coefficient, it decides not to encode the NSST index. Here, N is the number of samples in the upper left target region, and R is the reduced coefficient, which is smaller than N. N is derived as the product of the width and height of the upper left target region.
[0193] Furthermore, if the R+1th to Nth conversion coefficients do not include any non-zero conversion coefficients, the encoding device decides to encode the NSST index. In this case, the information regarding the conversion coefficients includes the syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0194] On the other hand, the encoding device determines whether the conditions for performing the NSST are met, and if the NSST can be performed, it decides to encode the NSST index for the target block. For example, an NSST index indicator is generated from the bitstream for the target CTU containing the target block, and the NSST index indicator indicates whether or not NSST is applied to the target CTU. If the value of the NSST index indicator is 1, the encoding device decides to encode the NSST index for the target block, and if the value of the NSST index indicator is 0, the decoding device decides not to encode the NSST index for the target block. As in the example above, the NSST index indicator is signaled at the CTU level, and the NSST index indicator is also signaled at the CU level or other higher levels.
[0195] Furthermore, the NSST index is used for multiple components of the target block.
[0196] For example, the NSST index is used for the inverse transformation of the transformation coefficients of the lumar block, chroma Cb block, and chroma Cr block of the target block. In this case, the R+1 to Nth transformation coefficients of the lumar block, the R+1 to Nth transformation coefficients of the chroma Cb block, and the R+1 to Nth transformation coefficients of the chroma Cr block are scanned, and if the scanned transformation coefficients include non-zero transformation coefficients, it is determined that the NSST index is not encoded. If the scanned transformation coefficients do not include non-zero transformation coefficients, it is determined that the NSST index is encoded. In this case, the information regarding the transformation coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0197] As another example, the NSST index is used for the inverse transformation of the transformation coefficients of the rumor block and the chroma Cb block of the target block. In this case, the R+1 to Nth transformation coefficients of the rumor block and the R+1 to Nth transformation coefficients of the chroma Cb block are scanned, and if the scanned transformation coefficients include non-zero transformation coefficients, it is determined that the NSST index is not encoded. If the scanned transformation coefficients do not include non-zero transformation coefficients, it is determined that the NSST index is encoded. In this case, the information regarding the transformation coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0198] As another example, the NSST index is used for the inverse transformation of the transformation coefficients of the rumor block and the chroma Cr block of the target block. In this case, the R+1 to Nth transformation coefficients of the rumor block and the R+1 to Nth transformation coefficients of the chroma Cr block are scanned, and if the scanned transformation coefficients include non-zero transformation coefficients, it is determined that the NSST index is not encoded. If the scanned transformation coefficients do not include non-zero transformation coefficients, it is determined that the NSST index is encoded. In this case, the information regarding the transformation coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0199] On the other hand, the range of the NSST index can be derived based on specific conditions. For example, the maximum value of the NSST index can be derived based on the specific conditions, and the range can be derived as 0 to the derived maximum value. The derived value of the NSST index is included in the range.
[0200] For example, the range of the NSST index is derived based on the size of the target block. Specifically, if the minimum width and minimum height are already set, the range of the NSST index is derived based on the width and minimum width of the target block, the height and minimum height of the target block. Alternatively, the range of the NSST index is derived based on the number of samples of the target block and a specific value. The number of samples is the product of the width and height of the target block, and the specific value may already be set.
[0201] As another example, the range of the NSST index is derived based on the type of the target block. Specifically, the range of the NSST index is derived based on whether or not the target block is a non-square block. Furthermore, the range of the NSST index is derived based on the ratio of the width to the height of the target block and a specific value. The ratio of the width to the height of the target block is the value obtained by dividing the longer side of the width and height of the target block by the shorter side, and the specific value may already be set.
[0202] As another example, the range of the NSST index is derived based on the intra-prediction mode of the target block. Specifically, the range of the NSST index is derived based on whether the intra-prediction mode of the target block is a non-directional intra-prediction mode or a directional intra-prediction mode. Furthermore, the range of the NSST index is derived based on whether the intra-prediction mode of the target block is an intra-prediction mode included in Category A or Category B. Here, the intra-prediction modes included in Category A and the intra-prediction modes included in Category B may already be set. As an example, Category A includes intra-prediction modes 2, 10, 18, 26, 34, 42, 50, 58, and 66, and Category B includes intra-prediction modes other than those included in Category A.
[0203] As another example, the range of the NSST index is derived based on information regarding the core transform of the target block. For example, the range of the NSST index is derived based on the AMT flag, which indicates whether or not an Adaptive Multiple Core Transform (AMT) is applied. Furthermore, the range of the NSST index is derived based on the AMT horizontal flag, which indicates the horizontal transform kernel, and the AMT vertical flag, which indicates the vertical transform kernel.
[0204] On the other hand, if the value of the NSST index is 0, the NSST index indicates that NSST is not applied to the target block.
[0205] The encoding device encodes information regarding the conversion coefficients (S1330). The information regarding the conversion coefficients includes information regarding the size, position, etc. of the conversion coefficients. As mentioned above, the information regarding the conversion coefficients may further include the NSST index, the conversion index information, and / or the AMT flag. The image information including the information regarding the conversion coefficients is output in bitstream form. The image information may further include the NSST index indicator and / or prediction information. The prediction information is information regarding the prediction procedure and includes prediction mode information and motion information (for example, if interpretation is applied).
[0206] The output bitstream is transmitted to a decoding device via a storage medium or network.
[0207] Figure 14 schematically shows an encoding apparatus that performs an image encoding method according to the present invention. The method disclosed in Figure 13 can be performed by the encoding apparatus disclosed in Figure 14. Specifically, for example, the addition unit of the encoding apparatus in Figure 14 can perform S1300 in Figure 13, the conversion unit of the encoding apparatus can perform S1310, and the entropy encoding unit of the encoding apparatus can perform S1320 to S1330 in Figure 13. Although not shown, the process of deriving predicted samples can be performed by the prediction unit of the encoding apparatus.
[0208] Figure 15 schematically shows an image decoding method using a decoding device according to the present invention. The method disclosed in Figure 15 can be performed using the decoding device disclosed in Figure 3. Specifically, for example, steps S1500 to S1510 in Figure 15 can be performed by the entropy decoding unit of the decoding device, S1520 by the inverse transformation unit of the decoding device, and S1530 by the addition unit of the decoding device. Although not shown, the process of deriving the predicted sample is performed by the prediction unit of the decoding device.
[0209] The decoding device derives the conversion coefficient of the target block from the bitstream (S1500). The decoding device decodes the information regarding the conversion coefficient of the target block received via the bitstream to derive the conversion coefficient of the target block. The received information regarding the conversion coefficient of the target block is referred to as residual information.
[0210] On the other hand, the conversion coefficient of the target block includes the conversion coefficient of the lumens block of the target block, the conversion coefficient of the chroma Cb block of the target block, and the conversion coefficient of the chroma Cr block of the target block.
[0211] The decoding device derives an NSST (Non-Separable Secondary Transform) index for the target block (S1510).
[0212] As an example, the decoding device scans the R+1th to Nth conversion coefficients of the target block, and if a non-zero conversion coefficient is included among the R+1th to Nth conversion coefficients, it derives the value of the NSST index as 0. Here, N is the number of samples in the upper left corner target region of the target block, and R is the reduced coefficient, which is smaller than N. N is derived as the product of the width and height of the upper left corner target region.
[0213] Furthermore, if the R+1 to Nth conversion coefficients do not contain any non-zero conversion coefficients, the decoding device parses the syntax elements for the NSST index contained in the bitstream to derive the value of the NSST index. In other words, if the R+1 to Nth conversion coefficients do not contain any non-zero conversion coefficients, the bitstream contains syntax elements for the NSST index, and the decoding device parses the syntax elements for the NSST index received via the bitstream to derive the value of the NSST index.
[0214] On the other hand, the decoding device determines whether the conditions for NSST execution are met, and if NSST execution is possible, derives an NSST index for the target block. For example, an NSST index indicator is signaled from the bitstream to a target CTU containing the target block, and the NSST index indicator indicates whether NSST is enabled for the target CTU. If the value of the NSST index indicator is 1, the decoding device derives an NSST index for the target block; if the value of the NSST index indicator is 0, the decoding device may not derive an NSST index for the target block. As in the example above, the NSST index indicator is signaled at the CTU level, or the NSST index indicator is signaled at the CU level or other higher level.
[0215] Furthermore, the NSST index is used for multiple components of the target block.
[0216] For example, the NSST index is used for inverse transformations of the transformation coefficients of the lumar block, chroma Cb block, and chroma Cr block of the target block. In this case, the R+1 to Nth transformation coefficients of the lumar block, the R+1 to Nth transformation coefficients of the chroma Cb block, and the R+1 to Nth transformation coefficients of the chroma Cr block are scanned, and if the scanned transformation coefficients include non-zero transformation coefficients, the value of the NSST index is derived as 0. If the scanned transformation coefficients do not include non-zero transformation coefficients, the bitstream includes syntax elements for the NSST index, and the value of the NSST index is derived by parsing the syntax elements for the NSST index received via the bitstream.
[0217] As another example, the NSST index is used for the inverse transformation of the transformation coefficients of the rumor block and the chroma Cb block of the target block. In this case, the R+1 to Nth transformation coefficients of the rumor block and the R+1 to Nth transformation coefficients of the chroma Cb block are scanned, and if the scanned transformation coefficients include non-zero transformation coefficients, the value of the NSST index is derived as 0. If the scanned transformation coefficients do not include non-zero transformation coefficients, the bitstream includes syntax elements for the NSST index, and the value of the NSST index is derived by parsing the syntax elements for the NSST index received via the bitstream.
[0218] As another example, the NSST index is used for the inverse transformation of the transformation coefficients of the rumor block and the chroma Cr block of the target block. In this case, the R+1 to Nth transformation coefficients of the rumor block and the R+1 to Nth transformation coefficients of the chroma Cr block are scanned, and if the scanned transformation coefficients include non-zero transformation coefficients, the value of the NSST index is derived as 0. If the scanned transformation coefficients do not include non-zero transformation coefficients, the bitstream includes syntax elements for the NSST index, and the value of the NSST index is derived by parsing the syntax elements for the NSST index received via the bitstream.
[0219] On the other hand, the range of the NSST index can be derived based on specific conditions. For example, the maximum value of the NSST index can be derived based on the specific conditions, and the range can be derived as 0 to the derived maximum value. The derived value of the NSST index is included in the range.
[0220] For example, the range of the NSST index is derived based on the size of the target block. Specifically, if the minimum width and minimum height are already set, the range of the NSST index is derived based on the width and minimum width of the target block, the height and minimum height of the target block. Alternatively, the range of the NSST index is derived based on the number of samples of the target block and a specific value. The number of samples is the value obtained by multiplying the width and height of the target block, and the specific value may already be set.
[0221] As another example, the range of the NSST index can be derived based on the type of the target block. Specifically, the range of the NSST index can be derived based on whether or not the target block is a non-square block. Alternatively, the range of the NSST index can be derived based on the ratio of the width to the height of the target block and a specific value. The ratio of the width to the height of the target block is the value obtained by dividing the longer side of the width and height of the target block by the shorter side, and the specific value may already be set.
[0222] As another example, the range of the NSST index can be derived based on the intra-prediction mode of the target block. Specifically, the range of the NSST index is derived based on whether the intra-prediction mode of the target block is a non-directional intra-prediction mode or a directional intra-prediction mode. Alternatively, the range of the NSST index can be derived based on whether the intra-prediction mode of the target block is an intra-prediction mode included in Category A or Category B. Here, the intra-prediction modes included in Category A and Category B may already be set. For example, Category A includes intra-prediction modes 2, 10, 18, 26, 34, 42, 50, 58, and 66, and Category B includes intra-prediction modes other than those included in Category A.
[0223] Another example is that the range of the NSST index can be derived based on information regarding the core transform of the target block. For example, the range of the NSST index can be derived based on the AMT flag, which indicates whether or not an Adaptive Multiple core transform (AMT) is applied. Alternatively, the range of the NSST index can be derived based on the AMT horizontal flag, which indicates the horizontal transform kernel, and the AMT vertical flag, which indicates the vertical transform kernel.
[0224] On the other hand, if the value of the NSST index is 0, the NSST index indicates that NSST is not applied to the target block.
[0225] The decoding device performs an inverse transform on the transformation coefficient of the target block based on the NSST index to derive the residual sample of the target block (S1520).
[0226] For example, if the value of the NSST index is 0, the decoding device performs a core transform on the conversion coefficient of the target block to derive the residual sample.
[0227] Specifically, the decoding device obtains an AMT flag indicating whether or not Adaptive Multiple Core Transform (AMT) is applied to the bitstream.
[0228] If the value of the AMT flag is 0, the decoding device derives DCT type 2 as the conversion kernel for the target block and performs an inverse conversion on the conversion coefficients based on DCT type 2 to derive the residual sample.
[0229] If the value of the AMT flag is 1, the decoding device configures a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, derives the horizontal conversion kernel and the vertical conversion kernel based on the conversion index information obtained from the bitstream and the conversion subset, and derives the residual sample by performing an inverse conversion on the conversion coefficients based on the horizontal conversion kernel and the vertical conversion kernel. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. The conversion index information also includes an AMT horizontal flag indicating one of the candidates included in the conversion subset for the horizontal conversion kernel and an AMT vertical flag indicating one of the candidates included in the conversion subset for the vertical conversion kernel. On the other hand, the conversion kernel may be called a conversion type or a conversion core.
[0230] As another example, if the value of the NSST index is not 0, the decoder performs an NSST on the transformation coefficient located in the upper left corner target region of the target block based on the reduced transform matrix indicated by the NSST index to derive a modified transformation coefficient, and then performs a core transform on the target block containing the modified transformation coefficient to derive the residual sample. The size of the reduced transform matrix is R × N, where N is the number of samples in the upper left corner target region, and R is the reduced coefficient, where R is less than N.
[0231] The core transformation for the target block is performed as follows: The decoder obtains an AMT flag from the bitstream indicating whether or not an Adaptive Multiple Core Transform (AMT) is applied. If the value of the AMT flag is 0, the decoder derives a DCT type 2 as the transformation kernel for the target block and performs an inverse transformation on the target block, including the modified transformation coefficients, based on the DCT type 2 to derive the sample.
[0232] If the value of the AMT flag is 1, the decoder configures a conversion subset for the horizontal conversion kernel and a conversion subset for the vertical conversion kernel, derives the horizontal conversion kernel and the vertical conversion kernel based on the conversion index information obtained from the bitstream and the conversion subset, and performs an inverse conversion on the target block including the modified conversion coefficients based on the horizontal conversion kernel and the vertical conversion kernel to derive the residual sample. Here, the conversion subset for the horizontal conversion kernel and the conversion subset for the vertical conversion kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. The conversion index information also includes an AMT horizontal flag indicating one of the candidates included in the conversion subset for the horizontal conversion kernel and an AMT vertical flag indicating one of the candidates included in the conversion subset for the vertical conversion kernel. On the other hand, the conversion kernel may also be called a conversion type or conversion core.
[0233] The decoding device generates a reconstructed picture based on the residual samples (S1530). The decoding device generates a reconstructed picture based on the residual samples. For example, the decoding device can perform inter-prediction or intra-prediction for the target block based on prediction information received via the bitstream to derive predicted samples, and generate the reconstructed picture by adding the predicted samples and the residual samples. Thereafter, as described above, in-loop filtering procedures such as deblocking filtering, SAO and / or ALF procedures can be applied to the reconstructed picture as needed to improve subjective / objective image quality.
[0234] Figure 16 schematically shows a decoding device that performs an image decoding method according to the present invention. The method disclosed in Figure 15 can be performed by the decoding device disclosed in Figure 16. Specifically, for example, the entropy decoding unit of the decoding device in Figure 16 can perform S1500 to S1510 in Figure 15, the inverse transform unit of the decoding device in Figure 16 can perform S1520 in Figure 15, and the summing unit of the decoding device in Figure 16 can perform S1530 in Figure 15. In addition, although not shown, the process of deriving a predicted sample can be performed by the prediction unit of the decoding device in Figure 16.
[0235] According to the present invention described above, the range of the NSST index can be derived based on specific conditions of the target block, thereby reducing the amount of bits required for the NSST index and improving overall coding efficiency.
[0236] Furthermore, according to the present invention, the transmission of syntax elements to an NSST index is determined based on a conversion coefficient for the target block, thereby reducing the number of bits required for the NSST index and improving overall coding efficiency.
[0237] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks, but the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and that other steps may be included or one or more steps in the flowchart may be removed without affecting the scope of the present invention.
[0238] The method according to the present invention described above is implemented in software form, and the encoding and / or decoding device according to the present invention is included in, for example, an image processing device such as a television, computer, smartphone, set-top box, or display device.
[0239] When embodiments of the present invention are implemented by software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and executed by a processor. The memory is located inside or outside the processor and can be connected to the processor by a variety of well-known means. The processor includes an ASIC (Application Specific Integrated Circuit), other chipsets, logic circuits, and / or data processing devices. The memory includes ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in the present invention are implemented on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing are implemented on a computer, processor, microprocessor, controller, or chip.
[0240] Furthermore, the decoding and encoding devices to which the present invention applies can include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, OTT (Over the top video) equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals or data signals. For example, OTT video (Over the top video) equipment includes game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and DVRs (Digital Video Recorders).
[0241] Furthermore, the processing method to which the present invention is applied can be produced in the form of a program executed by a computer and can be stored on a computer-readable recording medium. Multimedia data having the data structure according to the present invention can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks (registered trademarks), and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of a carrier wave (for example, transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wireless communication network. Furthermore, embodiments of the present invention are implemented as computer program products using program code, and the program code is performed on a computer according to embodiments of the present invention. The program code can be stored on a computer-readable carrier.
[0242] Furthermore, the content streaming system to which the present invention applies includes an encoding server, a streaming server, a web server, a media storage facility, a user device, and a multimedia input device.
[0243] The encoding server is responsible for compressing content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream, and transmitting it to the streaming server. In another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted. The bitstream is generated by an encoding method or bitstream generation method to which the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0244] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. The content streaming system may also include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0245] The streaming server receives content from a media storage facility and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0246] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs (head-mounted displays)), digital TVs, desktop computers, and digital signage. Each server in the content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
Claims
1. In a video decoding method performed by a decoding device, Steps to obtain prediction mode information and residual information from a bitstream, The steps include: deriving a prediction sample for the target block based on the prediction mode information; The steps include: deriving the conversion coefficient of the target block based on the residual information; The steps include: deriving a residual sample of the target block based on the inverse transformation of the conversion coefficient of the target block; The steps include generating a reconstructed picture based on the predicted sample and the residual sample, The step of deriving the conversion coefficient for the target block is: The step of determining whether there is a non-zero conversion coefficient among the conversion coefficients of the target block, from the R+1th to the Nth conversion coefficient, The step of deriving an inseparable transformation index for the target block based on the determination of whether there is a non-zero transformation coefficient among the R+1th to Nth transformation coefficients, The inverse transformation is performed based on the non-separable transform index, If the value of the non-separable transformation index is not equal to 0, the inverse transformation for the coefficients included in the upper-left target region of the target block is performed based on the transformation matrix associated with the non-separable transformation index. The size of the transformation matrix is R × N. The above N is the number of coefficients included in the upper left symmetric region, The R is a video decoding method that is smaller than the N.
2. In a video encoding method performed by an encoding device, The steps include: deriving the prediction mode for the target block, The steps include: deriving a predicted sample for the target block based on the prediction mode; A step of generating prediction mode information based on the prediction mode, The steps include: deriving a residual sample of the target block based on the predicted sample; The steps include: deriving a conversion coefficient for the target block by performing a conversion on the residual sample; The step of determining whether there is a non-zero conversion coefficient among the conversion coefficients of the target block, from the R+1th to the Nth conversion coefficient, A step of determining whether to encode an unseparable transformation index for the transformation coefficients based on the determination of whether there are non-zero transformation coefficients among the R+1th to Nth transformation coefficients, The step includes encoding the prediction mode information and residual information including information related to the conversion coefficient, If the value of the non-separable transformation index is not equal to 0, the transformation based on the non-separable transformation for the coefficients included in the upper left target region of the target block is performed based on the transformation matrix associated with the non-separable transformation index. The size of the transformation matrix is R × N. The above N is the number of coefficients included in the upper left symmetric region, A video encoding method in which R is smaller than N.
3. In a transmission method for data including a bitstream of image information, A step of acquiring the bitstream of the image information, which includes prediction mode information and residual information, wherein the bitstream is The steps include: deriving the prediction mode for the target block, The steps include: deriving a predicted sample for the target block based on the prediction mode; A step of generating prediction mode information based on the prediction mode, The steps include: deriving a residual sample of the target block based on the predicted sample; The steps include: deriving a conversion coefficient for the target block by performing a conversion on the residual sample; The step of determining whether there is a non-zero conversion coefficient among the conversion coefficients of the target block, from the R+1th to the Nth conversion coefficient, A step of determining whether to encode an unseparable transformation index for the transformation coefficients based on the determination of whether there are non-zero transformation coefficients among the R+1th to Nth transformation coefficients, A step of encoding the prediction mode information and residual information including the conversion coefficients and outputting the bitstream, which is generated by the step, The step of transmitting the data, which includes the bitstream of the image information, which includes the prediction mode information and the residual information, If the value of the non-separable transformation index is not equal to 0, the transformation based on the non-separable transformation for the coefficients included in the upper left target region of the target block is performed based on the transformation matrix associated with the non-separable transformation index. The size of the transformation matrix is R × N. The above N is the number of coefficients included in the upper left symmetric region, The method wherein R is smaller than N.