Image coding method based on non-separable secondary transform and device therefor
The image decoding method using NSST indexes addresses the high data volume challenge of high-resolution video by optimizing coding efficiency through conditional encoding, reducing transmission costs.
Patent Information
- Application Number
- JP2025132010
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-15
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-06
- Estimated Expiration
- 2038-12-13
AI Technical Summary
The increasing demand for high-resolution, high-quality video has led to a surge in video data transmission and storage costs due to the higher amount of information required, necessitating highly efficient video compression techniques.
An image decoding method and apparatus that utilizes Non-Separable Secondary Transform (NSST) indexes, determining their range based on specific conditions of the target block, and decides whether to encode them based on transform coefficients, reducing the bit transmission for NSST indexes.
This approach enhances coding efficiency by reducing the amount of bits required for NSST indexes, thereby improving overall coding efficiency and transmission costs.
Smart Images

Figure 2025147205000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to image coding techniques, and more particularly to a method and apparatus for decoding images with non-separable quadratic transforms in an image coding system. [Background technology]
[0002] Recently, the demand for high-resolution, high-quality video such as HD (High Definition) video and UHD (Ultra High Definition) video has been increasing in various fields. As the video data has higher resolution and quality, the amount of information or bits to be transmitted increases relatively compared to existing video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line or storing video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] This requires highly efficient video compression techniques to effectively transmit, store, and play back high-resolution, high-quality video information. Summary of the Invention [Problem to be solved by the invention]
[0004] The technical problem of the present invention is to provide a method and apparatus for improving image coding efficiency.
[0005] Another technical object of the present invention is to provide an image decoding method and apparatus that applies NSST to a current block.
[0006] Another technical object of the present invention is to provide an image decoding method and apparatus that derives the range of NSST indexes based on specific conditions of the target block.
[0007] Another technical object of the present invention is to provide an image decoding method and apparatus for determining whether to code an NSST index based on the transform coefficients of a current block. [Means for solving the problem]
[0008] According to an embodiment of the present invention, there is provided an image decoding method performed by a decoding device, the method including the steps of: deriving transform coefficients of a current block from a bitstream, deriving a Non-Separable Secondary Transform (NSST) index for the current block, performing an inverse transform on the transform coefficients of the current block based on the NSST index to derive residual samples of the current block, and generating a reconstructed picture based on the residual samples.
[0009] According to another embodiment of the present invention, there is provided a decoding device for performing image decoding, including: an entropy decoding unit that derives transform coefficients of a current block from a bitstream and derives a Non-Separable Secondary Transform (NSST) index for the current block; an inverse transform unit that performs an inverse transform on the transform coefficients of the current block based on the NSST index to derive residual samples of the current block; and an adder that generates a reconstructed picture based on the residual samples.
[0010] According to another embodiment of the present invention, there is provided a video encoding method performed by an encoding device, the method including the steps of: deriving residual samples of a current block, transforming the residual samples to derive transform coefficients of the current block, determining whether to encode an NSST index for the current block, and encoding information about the transform coefficients, wherein the step of determining whether to encode the NSST index includes scanning (R+1)th to Nth transform coefficients of the current block, and determining not to encode the NSST index if a non-zero transform coefficient is included in the (R+1)th to Nth transform coefficients, where N is the number of samples in an upper left target region of the current block, R is a reduced coefficient, and R is less than N.
[0011] According to another embodiment of the present invention, there is provided a video encoding device, including an adder unit that derives residual samples of a current block, a transform unit that transforms the residual samples to derive transform coefficients of the current block, and an entropy encoding unit that determines whether to encode an NSST index for the current block and encodes information about the transform coefficients, wherein the entropy encoding unit scans (R+1)th to (N)th transform coefficients of the current block and determines not to encode the NSST index if a non-zero transform coefficient is included in the (R+1)th to (N)th transform coefficients, where N is the number of samples in an upper left target region of the current block, R is a reduced coefficient, and R is less than N. [Effects of the Invention]
[0012] According to the present invention, the range of the NSST index can be derived based on specific conditions of the target block, thereby reducing the amount of bits for the NSST index and improving overall coding efficiency.
[0013] According to the present invention, the transmission of syntax elements for NSST indexes is determined based on the transform coefficients for the target block, thereby reducing the amount of bits for the NSST indexes and improving overall coding efficiency. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating a configuration of a video encoding device to which the present invention can be applied. [Figure 2] 1 shows an example of an image encoding method performed by a video encoding device. [Figure 3] 1 is a diagram illustrating the configuration of a video decoding device to which the present invention can be applied. [Figure 4] 1 shows an example of an image decoding method performed by a decoding device. [Figure 5] 1 illustrates a schematic representation of a multiple conversion technique according to the present invention; [Figure 6] 65 intra-directional modes of prediction directions are shown exemplarily. [Figure 7a] 1 is a flowchart illustrating a process for coding transform coefficients according to an embodiment. [Figure 7b] 1 is a flowchart illustrating a process for coding transform coefficients according to an embodiment. [Figure 8] 10A and 10B are diagrams illustrating an arrangement of transform coefficients based on a current block according to an embodiment of the present invention; [Figure 9] An example of scanning transform coefficients from R+1 to N is shown below. [Figure 10a] 10 is a flowchart illustrating a coding process of an NSST index according to an embodiment. [Figure 10b] 10 is a flowchart illustrating a coding process of an NSST index according to an embodiment. [Figure 11] An example of determining whether an NSST index is coded will be described below. [Figure 12] An example of scanning transform coefficients R+1 to N for all components of a current block is shown below. [Figure 13] 1 illustrates a schematic diagram of an image encoding method using an encoding device according to the present invention. [Figure 14] 1 shows a schematic representation of an encoding device for carrying out an image encoding method according to the present invention; [Figure 15] 2 illustrates a schematic diagram of an image decoding method using a decoding device according to the present invention; [Figure 16] 1 shows a schematic representation of a decoding device for carrying out an image decoding method according to the invention; DETAILED DESCRIPTION OF THE INVENTION
[0015] The present invention may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this does not limit the present invention to the specific embodiments. The terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical spirit of the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0016] Meanwhile, each component in the drawings described in the present invention is illustrated independently for the convenience of explaining different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0017] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. In the following, the same reference numerals are used to designate the same components in the drawings, and redundant description of the same components will be omitted.
[0018] Meanwhile, the present invention relates to video / image coding, for example, the methods / embodiments disclosed in the present invention can be applied to methods disclosed in the VVC (versatile video coding) standard or next generation video / image coding.
[0019] In this specification, a picture generally refers to a unit representing one image in a specific time period, and a slice is a unit constituting a part of a picture in coding. One picture may be composed of multiple slices, and pictures and slices may be used interchangeably as needed.
[0020] A pixel or a pel may refer to the smallest unit constituting a picture (or an image). A term corresponding to a pixel may also be used: "sample." A sample generally refers to a pixel or a pixel value, and may refer to only the value of a pixel / pixel of a luminance (luma) component, or may refer to only the value of a pixel / pixel of a chroma component.
[0021] A unit refers to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information about the region. The term unit may be used interchangeably with terms such as block or area. In general, an MxN block may refer to a set of samples or transform coefficients consisting of M columns and N rows.
[0022] FIG. 1 is a diagram illustrating the configuration of a video encoding device to which the present invention can be applied.
[0023] 1, the video encoding apparatus 100 may include a picture division unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an adder 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a transform unit 122, a quantization unit 123, a realignment unit 124, an inverse quantization unit 125, and an inverse transform unit 126.
[0024] The picture division unit 105 can divide an input picture into at least one processing unit.
[0025] For example, the processing unit is called a coding unit (CU). In this case, the coding units may be recursively divided from the largest coding unit (LCU) using a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure and / or a binary tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure. Alternatively, the binary tree structure may be applied first. A coding procedure according to the present invention may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0026] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding units may be split into coding units of deeper depths using a quadtree structure, starting from the largest coding unit (LCU). In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to video characteristics, or the coding unit may be recursively split into coding units of lower depths as needed, and the coding unit of the optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that serves as the basis for partitioning or division into prediction units or transform units. A prediction unit is a unit that is partitioned from a coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into subblocks. A transform unit can be divided from a coding unit according to a quadtree structure and is a unit that derives transform coefficients and / or a unit that derives a residual signal from the transform coefficients. Hereinafter, a coding unit is also referred to as a coding block (CB), a prediction unit is also referred to as a prediction block (PB), and a transform unit is also referred to as a transform block (TB). A prediction block or a prediction unit refers to a specific region in a block form within a picture and can include an array of prediction samples.Also, a transform block or transform unit refers to a specific region in a picture in the form of a block, and can include an array of transform coefficients or residual samples.
[0027] The prediction unit 110 performs prediction on a current block to be processed (hereinafter, referred to as a current block) and generates a predicted block including prediction samples for the current block. The prediction unit 110 performs prediction on a coding block, a transform block, or a prediction block.
[0028] The predictor 110 may determine whether intra prediction or inter prediction is applied to the current block. For example, the predictor 110 may determine whether intra prediction or inter prediction is applied to each CU.
[0029] In intra prediction, the predictor 110 may derive a prediction sample for a current block based on a reference sample outside the current block within a picture to which the current block belongs (hereinafter, the current picture). In this case, the predictor 110 may (i) derive a prediction sample based on an average or interpolation of neighboring reference samples of the current block, or (ii) derive a prediction sample based on a reference sample present in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. (i) is referred to as a non-directional mode or a non-angular mode, and (ii) is referred to as a directional mode or an angular mode. Prediction modes in intra prediction may include, for example, 33 directional prediction modes and at least two or more non-directional modes. Non-directional modes may include a DC prediction mode and a planar mode. The predictor 110 may also determine a prediction mode to be applied to the current block using a prediction mode applied to a neighboring block.
[0030] In the case of inter prediction, the predictor 110 may derive a predicted sample for the current block based on a sample identified by a motion vector on a reference picture. The predictor 110 may derive a predicted sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the skip mode and the merge mode, the predictor 110 may use motion information of a neighboring block as motion information of the current block. In the skip mode, unlike the merge mode, a difference (residual) between a predicted sample and an original sample is not transmitted. In the MVP mode, the motion vector of the current block may be derived by using the motion vector of the neighboring block as a motion vector predictor.
[0031] In the case of inter-prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. The reference picture including the temporal neighboring blocks is also called a collocated picture (colPic). Motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and output in the form of a bitstream.
[0032] When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture. Reference pictures included in the reference picture list can be sorted based on the POC (Picture Order Count) difference between the current picture and the corresponding reference picture. POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0033] The subtractor 121 generates residual samples, which are the differences between the original samples and the predicted samples. When the skip mode is applied, the residual samples are not generated as described above.
[0034] The transform unit 122 transforms residual samples in units of transform blocks to generate transform coefficients. The transform unit 122 may perform the transform according to the size of the corresponding transform block and a prediction mode applied to a coding block or a prediction block spatially overlapping with the corresponding transform block. For example, if intra prediction is applied to the coding block or the prediction block overlapping with the transform block and the transform block is a 4x4 residual array, the residual samples may be transformed using a Discrete Sine Transform (DST) transform kernel; otherwise, the residual samples may be transformed using a Discrete Cosine Transform (DCT) transform kernel.
[0035] The quantization unit 123 can quantize the transform coefficients to generate quantized transform coefficients.
[0036] The rearrangement unit 124 rearranges the quantized transform coefficients. The rearrangement unit 124 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form through a coefficient scanning method. Here, the rearrangement unit 124 has been described as a separate component, but it may also be a part of the quantization unit 123.
[0037] The entropy encoding unit 130 may perform entropy encoding on the quantized transform coefficients. Entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also encode information required for video restoration (e.g., syntax element values) in addition to the quantized transform coefficients, either together with or separately from the quantized transform coefficients, using entropy encoding or a preset method. The encoded information may be transmitted or stored in network abstraction layer (NAL) unit units in the form of a bitstream.
[0038] The inverse quantization unit 125 inversely quantizes the values (quantized transformation coefficients) quantized by the quantization unit 123, and the inverse transform unit 126 inversely transforms the values inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0039] The adder 140 reconstructs a picture by adding residual samples and predicted samples. The residual samples and predicted samples may be added in block units to generate reconstructed blocks. Although the adder 140 has been described as a separate component, it may be part of the prediction unit 110. Meanwhile, the adder 140 may also be referred to as a reconstruction module or a reconstructed block generator.
[0040] The filter unit 150 may apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through the deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process may be corrected. The sample adaptive offset may be applied on a sample-by-sample basis and may be applied after the deblocking filtering process is completed. The filter unit 150 may also apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filtering and / or the sample adaptive offset have been applied.
[0041] The memory 160 may store a reconstructed picture (a decoded picture) or information necessary for encoding / decoding. Here, a reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter unit 150. The stored reconstructed picture may be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 may store (reference) pictures used for inter prediction. In this case, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list.
[0042] 2 illustrates an example of an image encoding method performed by a video encoding device. As shown in FIG. 2, the image encoding method includes intra / inter prediction, transform, quantization, and entropy encoding processes. For example, a predicted block of a current block is generated by intra / inter prediction, and a residual block of the current block is generated by subtracting an input block of the current block from the predicted block. Subsequently, a coefficient block, i.e., a transform coefficient of the current block, is generated by transforming the residual block. The transform coefficient is quantized and entropy encoded, and then stored in a bitstream.
[0043] FIG. 3 is a diagram for explaining the outline of the configuration of a video decoding device to which the present invention can be applied.
[0044] 3, the video decoding device 300 includes an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 may include a realignment unit 321, an inverse quantization unit 322, and an inverse transform unit 323.
[0045] When a bitstream containing video information is input, the video decoding device 300 can reconstruct the video in accordance with the process by which the video information was processed in the video encoding device.
[0046] For example, the video decoding device 300 may perform video decoding using a processing unit applied in a video encoding device. Accordingly, a processing unit block for video decoding may be a coding unit, for example, or a coding unit, a prediction unit, or a transform unit, for example. The coding unit may be divided from the largest coding unit according to a quad tree structure and / or a binary tree structure.
[0047] A prediction unit and a transform unit may also be used in some cases. In this case, a prediction block may be a block derived or partitioned from a coding unit and may be a unit of sample prediction. Here, the prediction unit may be divided into sub-blocks. A transform unit may be divided from a coding unit according to a quadtree structure and may be a unit that derives transform coefficients or a unit that derives a residual signal from the transform coefficients.
[0048] The entropy decoding unit 310 parses the bitstream and outputs information necessary for video or picture reconstruction. For example, the entropy decoding unit 310 may decode information in the bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for video reconstruction and quantized values of transform coefficients related to residuals.
[0049] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in a bitstream, determines a context model using information on the syntax element to be decoded and decode information on neighboring and target blocks to be decoded or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. Here, the CABAC entropy decoding method can update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin after determining the context model.
[0050] Among the information decoded in the entropy decoding unit 310, information related to prediction is provided to the prediction unit 330, and the residual values entropy decoded in the entropy decoding unit 310, i.e., the quantized transform coefficients, are input to the reordering unit 321.
[0051] The rearrangement unit 321 rearranges the quantized transform coefficients into a two-dimensional block format. The rearrangement unit 321 may perform the rearrangement in response to coefficient scanning performed in the encoding device. Here, the rearrangement unit 321 has been described as a separate component, but the rearrangement unit 321 may be a part of the inverse quantization unit 322.
[0052] The inverse quantization unit 322 inversely quantizes the quantized transform coefficients based on the (inverse) quantization parameter and outputs the transform coefficients, where information for deriving the quantization parameter is signaled from the encoding device.
[0053] The inverse transform unit 323 inversely transforms the transform coefficients to derive residual samples.
[0054] The prediction unit 330 performs prediction on the current block and generates a predicted block including prediction samples for the current block. The unit of prediction performed by the prediction unit 330 may be a coding block, a transform block, or a prediction block.
[0055] The prediction unit 330 determines whether to apply intra prediction or inter prediction based on the information related to the prediction. Here, the unit for determining whether to apply intra prediction or inter prediction may differ from the unit for generating prediction samples. In addition, the unit for generating prediction samples in inter prediction and intra prediction may also differ. For example, whether to apply inter prediction or intra prediction may be determined on a CU basis. Furthermore, for example, in inter prediction, a prediction mode may be determined on a PU basis and prediction samples may be generated, and in intra prediction, a prediction mode may be determined on a PU basis and prediction samples may be generated on a TU basis.
[0056] In the case of intra prediction, the prediction unit 330 may derive prediction samples for the current block based on neighboring reference samples in the current picture. The prediction unit 330 may derive prediction samples for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. Here, the prediction mode to be applied to the current block may be determined using the intra prediction mode of the neighboring blocks.
[0057] In the case of inter prediction, the predictor 330 may derive a prediction sample for the current block based on a sample identified on the reference picture by a motion vector on the reference picture. The predictor 330 derives a prediction sample for the current block by applying any one of a skip mode, a merge mode, and an MVP mode. Here, motion information required for inter prediction of the current block provided from the video encoding device, such as information on a motion vector and a reference picture index, is obtained or induced based on the information on the prediction.
[0058] In the skip mode and merge mode, motion information of neighboring blocks may be used as motion information of the current block, where neighboring blocks include spatial neighboring blocks and temporal neighboring blocks.
[0059] The prediction unit 330 constructs a merge candidate list as motion information of available neighboring blocks, and can use information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index is signaled from the encoding device. The motion information may include a motion vector and a reference picture. When motion information of temporal neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture.
[0060] In skip mode, unlike merge mode, the residual between the predicted sample and the original sample is not transmitted.
[0061] In the MVP mode, the motion vector of the current block is derived using the motion vectors of the surrounding blocks as motion vector predictors, where the surrounding blocks include spatial and temporal surrounding blocks.
[0062] For example, when a merge mode is applied, a merge candidate list is generated using the motion vectors of the reconstructed spatial neighboring blocks and / or the motion vector corresponding to the Col block, which is a temporal neighboring block. In the merge mode, the motion vector of a candidate block selected in the merge candidate list is used as the motion vector of the current block. The prediction information includes a merge index indicating a candidate block having an optimal motion vector selected from the candidate blocks included in the merge candidate list. Here, the prediction unit 330 derives the motion vector of the current block using the merge index.
[0063] As another example, when the Motion Vector Prediction (MVP) mode is applied, a motion vector predictor candidate list is generated using the motion vectors of the reconstructed spatially surrounding blocks and / or the motion vector corresponding to the Col block, which is a temporally surrounding block. That is, the motion vectors of the reconstructed spatially surrounding blocks and / or the motion vector corresponding to the Col block, which is a temporally surrounding block, may be used as motion vector candidates. The prediction information includes a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list. Here, the prediction unit 330 may select a predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. A prediction unit of the encoding device calculates a motion vector differential (MVD) between the motion vector of the current block and the motion vector predictor, encodes the MVD, and outputs it in the form of a bitstream. That is, the MVD is calculated by subtracting the motion vector predictor from the motion vector of the current block. Here, the prediction unit 330 obtains the motion vector differential included in the prediction information and derives the motion vector of the current block by adding the motion vector differential and the motion vector predictor. The predictor may also obtain or derive from information about the prediction a reference picture index or the like that indicates a reference picture.
[0064] The adder 340 reconstructs a current block or a current picture by adding residual samples and predicted samples. The adder 340 may also reconstruct a current picture by adding residual samples and predicted samples in block units. When skip mode is applied, residuals are not transmitted, so predicted samples may be reconstructed samples. Although the adder 340 is described as a separate component here, the adder 340 may be part of the prediction unit 330. Meanwhile, the adder 340 may also be referred to as a reconstruction unit or a reconstructed block generation unit.
[0065] The filter unit 350 may apply deblocking filtering, sample adaptive offset, and / or ALF to the reconstructed picture. Here, the sample adaptive offset is applied in sample units and may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or sample adaptive offset.
[0066] The memory 360 stores a reconstructed picture (decoded picture) or information required for decoding. Here, a reconstructed picture may be a reconstructed picture that has undergone a filtering procedure by the filter unit 350. For example, the memory 360 stores pictures used for inter prediction. Here, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for other pictures. The memory 360 may also output the reconstructed pictures in an output order.
[0067] 4 illustrates an example of an image decoding method performed by a decoding device. As shown in FIG. 4, the image decoding method includes processes of entropy decoding, inverse quantization, inverse transform, and intra / inter prediction. For example, the decoding device may perform the inverse process of the encoding method. Specifically, quantized transform coefficients are obtained by entropy decoding a bitstream, and a coefficient block of a current block, i.e., transform coefficients, is obtained by inverse quantization of the quantized transform coefficients. A residual block of the current block is derived by inverse transform of the transform coefficients, and a reconstructed block of the current block is derived by adding the residual block to a predicted block of the current block derived by intra / inter prediction.
[0068] Meanwhile, the above-described transformation derives low-frequency transform coefficients for the residual block of the current block, and a zero tail is derived at the end of the residual block.
[0069] Specifically, the transform is composed of two main processes, including a core transform and a secondary transform. The transform including the core transform and the secondary transform can be called a multi-transform technique.
[0070] FIG. 5 illustrates a schematic diagram of a multiple conversion technique according to the present invention.
[0071] As shown in Figure 5, the conversion unit corresponds to the conversion unit in the encoding device of Figure 1 described above, and the inverse conversion unit corresponds to the inverse conversion unit in the encoding device of Figure 1 described above or the inverse conversion unit in the decoding device of Figure 3.
[0072] The transform unit performs a linear transform based on the residual samples (residual sample array) in the residual block to derive (primary) transform coefficients (S510). Here, the linear transform includes an adaptive multiple core transform (AMT). The adaptive multiple core transform may also be referred to as a multiple transform set (MTS).
[0073] The adaptive multi-core transform refers to a method of performing transform using a Discrete Cosine Transform (DCT) type 2 and a Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1 additionally. That is, the adaptive multi-core transform refers to a transform method of transforming a spatial domain residual signal (or a residual block) into a frequency domain transform coefficient (or a primary transform coefficient) based on a plurality of transform kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary transform coefficient may be referred to as a temporary transform coefficient from the perspective of a transform unit.
[0074] In other words, when an existing transform method is applied, a spatial-domain to frequency-domain transform is applied to a residual signal (or a residual block) based on DCT type 2 to generate transform coefficients. In contrast, when the adaptive multi-core transform is applied, a spatial-domain to frequency-domain transform is applied to a residual signal (or a residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1 to generate transform coefficients (or primary transform coefficients). Here, DCT type 2, DST type 7, DCT type 8, DST type 1, etc. may be referred to as transform types, transform kernels, or transform cores.
[0075] For reference, the DCT / DST transformation types are defined based on basis functions, which can be expressed as shown in the table below.
[0076] [Table 1]
[0077] When the adaptive multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for a current block are selected from the transform kernels, and a vertical transform for the current block is performed based on the vertical transform kernel, and a horizontal transform for the current block is performed based on the horizontal transform kernel. Here, the horizontal transform indicates a transform for a horizontal component of the current block, and the vertical transform indicates a transform for a vertical component of the current block. The vertical transform kernel / horizontal transform kernel are adaptively determined based on a transform index indicating a prediction mode and / or a transform subset of a current block (CU or sub-block) that surrounds a residual block.
[0078] For example, the adaptive multi-core transform is applied when both the width and height of the current block are less than or equal to 64, and whether the adaptive multi-core transform of the current block is applied can be determined based on a CU level flag. Specifically, when the CU level flag is 0, the existing transform method described above can be applied. That is, when the CU level flag is 0, a spatial domain to frequency domain transform is applied to a residual signal (or a residual block) based on the DCT type 2 to generate transform coefficients, and the transform coefficients are encoded. Meanwhile, here, the current block may be a CU. When the CU level flag is 0, the adaptive multi-core transform can be applied to the current block.
[0079] Furthermore, for a luma block of a current block to which the adaptive multi-core transform is applied, two additional flags are signaled, and a vertical transform kernel and a horizontal transform kernel are selected based on the flags. The flag for the vertical transform kernel may be expressed as an AMT vertical flag, and AMT_TU_vertical_flag (or EMT_TU_vertical_flag) indicates a syntax element of the AMT vertical flag. The flag for the horizontal transform kernel may be expressed as an AMT horizontal flag, and AMT_TU_horizontal_flag (or EMT_TU_horizontal_flag) indicates a syntax element of the AMT horizontal flag. The AMT vertical flag indicates one transform kernel candidate from among the transform kernel candidates included in a transform subset for the vertical transform kernel, and the transform kernel candidate indicated by the AMT vertical flag is derived as the vertical transform kernel for the current block. In addition, the AMT horizontal flag indicates one of the transform kernel candidates included in the transform subset for the horizontal transform kernel, and the transform kernel candidate indicated by the AMT horizontal flag is derived as the horizontal transform kernel for the current block. Meanwhile, the AMT vertical flag may be expressed as an MTS vertical flag, and the AMT horizontal flag may be expressed as an MTS horizontal flag.
[0080] Meanwhile, three transform subsets are pre-defined, and one of the transform subsets is derived as the transform subset for the vertical transform kernel based on the intra prediction mode applied to the current block. Also, one of the transform subsets is derived as the transform subset for the horizontal transform kernel based on the intra prediction mode applied to the current block. For example, the pre-defined transform subsets are derived as shown in the following table.
[0081] [Table 2]
[0082] Referring to Table 2, a transform subset with an index value of 0 indicates a transform subset that includes DST type 7 and DCT type 8 as transform kernel candidates, and a transform subset with an index value of 1 indicates a transform subset that includes DST type 7 and DCT type 8 as transform kernel candidates.
[0083] The transform subset for the vertical transform kernel and the transform subset for the horizontal transform kernel, which are derived based on the intra prediction mode applied to the current block, are derived as shown in the following table.
[0084] [Table 3]
[0085] where V denotes a transform subset for the vertical transform kernel and H denotes a transform subset for the horizontal transform kernel.
[0086] When the value of the AMT flag (or EMT_Cu_flag) for the current block is 1, a transform subset for the vertical transform kernel and a transform subset for the horizontal transform kernel are derived based on the intra prediction mode of the current block, as shown in Table 3. Thereafter, among the transform kernel candidates included in the transform subset for the vertical transform kernel, the transform kernel candidate indicated by the AMT vertical flag for the current block is derived as the vertical transform kernel of the current block, and the horizontal transform kernel candidate is derived as the horizontal transform kernel of the current block. Meanwhile, the AMT flag may be expressed as an MTS flag.
[0087] For reference, for example, the intra prediction modes include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes include a planar intra prediction mode numbered 0 and a DC intra prediction mode numbered 1, and the directional intra prediction modes include 65 intra prediction modes numbered 2 through 66. However, this is merely an example, and the present invention is also applicable to cases where the number of intra prediction modes is different. Meanwhile, in some cases, an intra prediction mode numbered 67 may also be used, and the 67th intra prediction mode may indicate a linear model (LM) mode.
[0088] FIG. 6 exemplarily shows the intra-directional modes of 65 prediction directions.
[0089] As shown in FIG. 6, intra prediction modes with horizontal directionality and intra prediction modes with vertical directionality can be distinguished around intra prediction mode No. 34, which has a left-up diagonal prediction direction. In FIG. 6, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 represent displacements in 1 / 32 units on the sample grid position. Intra prediction modes No. 2 through No. 33 have horizontal directionality, while intra prediction modes No. 34 through No. 66 have vertical directionality. Intra prediction modes No. 18 and No. 50 represent horizontal and vertical intra prediction modes, respectively. Intra prediction mode No. 2 may be referred to as left-down diagonal intra prediction mode, intra prediction mode No. 34 as left-up diagonal intra prediction mode, and intra prediction mode No. 66 as right-up diagonal intra prediction mode.
[0090] The transform unit performs a secondary transform based on the (primary) transform coefficients to derive (secondary) transform coefficients (S520). If the primary transform is a transform from the spatial domain to the frequency domain, the secondary transform can be considered a transform from the frequency domain to the frequency domain. The secondary transform includes a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform refers to a transform that generates transform coefficients (or secondary transform coefficients) for a residual signal by performing a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix. Here, a transform can be applied to the (primary) transform coefficients at once based on the non-separable transform matrix without separately applying a vertical transform and a horizontal transform (or independently applying a horizontal-vertical transform). In other words, the non-separable quadratic transform refers to a transform method in which vertical and horizontal components of the (primary) transform coefficients are transformed together without separating them based on the non-separable transform matrix to generate transform coefficients (or secondary transform coefficients). The non-separable quadratic transform is applied to the top-left region of a block (hereinafter referred to as a transform coefficient block or a target block) composed of (primary) transform coefficients. For example, if the width (W) and height (H) of the transform coefficient block are both 8 or greater, an 8x8 non-separable quadratic transform is applied to the top-left 8x8 region of the transform coefficient block (hereinafter referred to as the top-left target region). Also, if the width (W) and height (H) of the transform coefficient block are both 4 or greater and the width (W) or height (H) of the transform coefficient block is less than 8, a 4x4 non-separable quadratic transform is applied to the top-left min(8,W) x min(8,H) region of the transform coefficient block.
[0091] Specifically, for example, if a 4x4 input block is used, the non-separable quadratic transform is performed as follows:
[0092] The 4×4 input block X is expressed as follows:
[0093]
number
[0094] When X is expressed in vector form, the vector JPEG2025147205000006.jpg84 is expressed as follows:
[0095]
number
[0096] In this case, the second-order non-separable transform is calculated as follows:
[0097]
number
[0098] where: JPEG2025147205000009.jpg75 denotes the transform coefficient vector, and T denotes the 16x16 (non-separable) transform metric.
[0099] The 16×1 transform coefficient vector JPEG2025147205000010.jpg75 is derived, JPEG2025147205000011.jpg75 is re-organized into 4x4 blocks according to the scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is an example, and in order to reduce the computational complexity of the non-separable quadratic transform, HyGT (Hypercube-Givens Transform) or the like may be used to calculate the non-separable quadratic transform.
[0100] Meanwhile, the non-separable secondary transform may select a transform kernel (or transform core, transform type) on a mode-dependent basis, where the mode includes an intra-prediction mode and / or an inter-prediction mode.
[0101] As described above, the non-separable quadratic transform is performed based on an 8x8 transform or a 4x4 transform determined based on the width (W) and height (H) of the transform coefficient block. That is, the non-separable quadratic transform is performed based on an 8x8 sub-block size or a 4x4 sub-block size. For example, for the mode-based transform kernel selection, 35 sets of three non-separable quadratic transform kernels for the non-separable quadratic transform are configured for both the 8x8 sub-block size and the 4x4 sub-block size. That is, 35 transform sets are configured for the 8x8 sub-block size, and 35 transform sets are configured for the 4x4 sub-block size. In this case, each of the 35 transform sets for the 8x8 sub-block size includes three 8x8 transform kernels, and each of the 35 transform sets for the 4x4 sub-block size includes three 4x4 transform kernels. However, the transform sub-block size, the number of sets, and the number of transform kernels in a set are merely examples, and sizes other than 8x8 or 4x4 may be used, or n sets may be constructed, with k transform kernels included in each set.
[0102] The transform set may be referred to as an NSST set, and the transform kernels in the NSST set may be referred to as NSST kernels. The selection of a particular set from the transform sets may be based on, for example, the intra prediction mode of the current block (CU or sub-block).
[0103] In this case, the mapping between the 35 transform sets and the intra prediction modes is shown in the following table, for example: For reference, when the LM mode is applied to the current block, the secondary transform may not be applied to the current block.
[0104] [Table 4]
[0105] On the other hand, if it is determined that a specific set is to be used, one of the k transform kernels in the specific set is selected based on a non-separable secondary transform index. The encoding device derives a non-separable secondary transform index indicating a specific transform kernel based on a rate-distortion (RD) check and signals the non-separable secondary transform index to a decoding device. The decoding device selects one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, an NSST index value of 0 indicates the first non-separable secondary transform kernel, an NSST index value of 1 indicates the second non-separable secondary transform kernel, and an NSST index value of 2 indicates the third non-separable secondary transform kernel. Alternatively, an NSST index value of 0 indicates that the first non-separable secondary transform is not applied to the current block, and NSST index values of 1 to 3 indicate the three transform kernels.
[0106] 5, the transform unit may perform the non-separable secondary transform based on the selected transform kernel to obtain (secondary) transform coefficients, which are derived as quantized transform coefficients by the quantizer unit as described above, encoded, and signaled to a decoding device and transmitted to an inverse quantization / inverse transform unit in an encoding device.
[0107] On the other hand, if the secondary transform is omitted, the (primary) transform coefficients, which are the output of the primary (separate) transform, are derived as quantized transform coefficients by the quantization unit as described above, encoded, and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.
[0108] The inverse transform unit performs a series of steps in reverse order to those performed by the transform unit. The inverse transform unit receives (dequantized) transform coefficients, performs a secondary (inverse) transform on them to derive (primary) transform coefficients (S550), and performs a primary (inverse) transform on the (primary) transform coefficients to obtain residual blocks (residual samples). Here, the primary transform coefficients may be called modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding device and the decoding device can generate reconstructed blocks based on the residual blocks and predicted blocks, and generate reconstructed pictures based on the reconstructed blocks.
[0109] On the other hand, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients are received and the primary (separate) transform is performed to obtain a residual block (residual sample). As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and a predicted block, and generate a reconstructed picture based on the reconstructed block.
[0110] Meanwhile, the non-separable secondary transform may not be applied to blocks coded in transform skip mode. For example, if an NSST index for a target CU is signaled and the value of the NSST index is not 0, the non-separable secondary transform may not be applied to blocks coded in transform skip mode in the target CU. Also, if the target CU including blocks of all components (e.g., luma component, chroma component) is coded in the transform skip mode, or if the number of non-zero transform coefficients among the transform coefficients for the target CU is less than 2, the NSST index may not be signaled. A specific transform coefficient coding process is as follows.
[0111] 7a and 7b are flowcharts illustrating a process for coding transform coefficients according to one embodiment.
[0112] The steps disclosed in Figures 7a and 7b are performed by the encoding device 100 or the decoding device 300 disclosed in Figures 1 and 3, and more specifically, by the entropy encoding unit 130 disclosed in Figure 1 and the entropy decoding unit 310 disclosed in Figure 3. Therefore, the description of the details that overlap with the details described above in Figures 1 and 3 will be omitted or simplified.
[0113] In this specification, terms or phrases are used to define specific information or concepts. For example, in this specification, "a flag indicating whether or not at least one non-zero transform coefficient exists among the transform coefficients of a target block" is expressed as "cbf." However, since "cbf" can be replaced with various terms such as "coded_block_flag," when interpreting terms or phrases used to define specific information or concepts in the specification as a whole, they should not be interpreted solely based on their names, but should be interpreted with a focus on various operations, functions, and effects according to the meanings of the terms.
[0114] FIG. 7a shows the encoding process of the transform coefficients.
[0115] An encoding device 100 according to an embodiment determines whether a flag indicating whether at least one non-zero transform coefficient exists among the transform coefficients of a current block indicates 1 (S700). If the flag indicating whether at least one non-zero transform coefficient exists among the transform coefficients of the current block indicates 1, at least one non-zero transform coefficient exists among the transform coefficients of the current block. Conversely, if the flag indicating whether at least one non-zero transform coefficient exists among the transform coefficients of the current block indicates 0, all of the transform coefficients of the current block indicate 0.
[0116] A flag indicating whether there is at least one non-zero transform coefficient among the transform coefficients of a current block is expressed as a cbf flag, for example. The cbf flag includes a cbf_luma[x0][y0][trafoDepth] flag for a luma block and a cbf_cb[x0][y0][trafoDepth] and cbf_cr[x0][y0][trafoDepth] flag for a chroma block. Here, the array indexes x0 and y0 refer to the positions of the top-left luma / chroma samples of the current block relative to the top-left luma / chroma samples of the current picture, and the array index trafoDepth may refer to the level at which a coding block is divided for transform coding. A block whose trafoDepth indicates 0 corresponds to a coding block, and if the coding block and the transform block are defined as the same, trafoDepth is considered to be 0.
[0117] If the flag indicating whether there is at least one non-zero transform coefficient among the transform coefficients of the current block indicates 1 in S700, the encoding apparatus 100 according to an embodiment encodes information about the transform coefficients of the current block (S710).
[0118] The information about the transform coefficients of the current block includes, for example, at least one of information about the position of the last non-zero transform coefficient, group flag information indicating whether a subgroup of the current block includes a non-zero transform coefficient, and information about simplification coefficients. Each piece of information will be described in detail later.
[0119] According to an embodiment, the encoding apparatus 100 determines whether a condition for performing NSST is met (S720). More specifically, the encoding apparatus 100 determines whether a condition for encoding an NSST index is met. Here, the NSST index may be referred to as a transform index, for example.
[0120] According to an embodiment, the encoding device 100 encodes an NSST index when it is determined in step S720 that a condition for performing NSST is met (S730). More specifically, when it is determined that a condition for encoding an NSST index is met, the encoding device 100 encodes the NSST index.
[0121] According to an embodiment, if the flag indicating whether there is at least one non-zero transform coefficient among the transform coefficients for the current block indicates 0 in S700, the encoding device 100 may omit operations of S710, S720, and S730.
[0122] Furthermore, if it is determined in S720 that the conditions for performing NSST are not met, the encoding apparatus 100 according to an embodiment may omit the operation in S730.
[0123] FIG. 7b shows the process of decoding the transform coefficients.
[0124] A decoding device 300 according to an embodiment determines whether a flag indicating whether at least one non-zero transform coefficient exists among the transform coefficients of a current block indicates 1 (S740). If the flag indicating whether at least one non-zero transform coefficient exists among the transform coefficients of the current block indicates 1, at least one non-zero transform coefficient exists among the transform coefficients of the current block. Conversely, if the flag indicating whether at least one non-zero transform coefficient exists among the transform coefficients of the current block indicates 0, all of the transform coefficients of the current block indicate 0.
[0125] If the flag indicating whether there is at least one non-zero transform coefficient among the transform coefficients of the current block indicates 1 in S740, the decoding device 300 according to an embodiment decodes information about the transform coefficients of the current block (S750).
[0126] The decoding apparatus 300 according to an embodiment determines whether a condition for performing NSST is met (S760). More specifically, the decoding apparatus 300 determines whether a condition for decoding an NSST index from a bitstream is met.
[0127] If it is determined in S760 that the condition for performing NSST is met, the decoding device 300 according to an embodiment decodes the NSST index (S770).
[0128] According to an embodiment, if the flag indicating whether there is at least one non-zero transform coefficient among the transform coefficients for the current block indicates 0 in S740, the decoding device 300 can omit the operations of S750, S760, and S770.
[0129] Furthermore, if it is determined in S760 that the conditions for performing NSST are not met, the decoding device 300 according to an embodiment may omit the operation in S770.
[0130] As described above, signaling the NSST index when NSST is not performed may reduce coding efficiency. Also, different coding methods for the NSST index depending on specific conditions can improve overall image coding efficiency. Therefore, the present invention proposes various NSST index coding methods.
[0131] For example, the NSST index range may be determined based on a specific condition. In other words, the range of values of the NSST index may be determined based on a specific condition. Specifically, the maximum value of the NSST index may be determined based on the specific condition.
[0132] For example, the range of values of the NSST index is determined based on the block size. Here, the block size is defined as a minimum (W, H), where W indicates width and H indicates height. In this case, the range of values of the NSST index is determined by comparing the width of the current block with the W and comparing the height of the current block with the minimum H.
[0133] Alternatively, the block size is defined as the number of samples in a block (W*H). In this case, the range of the NSST index value is determined by comparing W*H, the number of samples in the current block, with a specific value.
[0134] Also, the range of values of the NSST index can be determined based on the shape of the block, i.e., the block type. Here, the block type is defined as a square block or a non-square block. In this case, the range of values of the NSST index is determined based on whether the target block is a square block or a non-square block.
[0135] Alternatively, the block type is defined as the ratio of the long side (the longer side of the width and height) to the short side of the block. In this case, the range of the NSST index value is determined by comparing the ratio of the long side to the short side of the target block with a preset threshold value (e.g., 2 or 3). Here, the ratio indicates a value obtained by dividing the long side by the short side. For example, if the width of the target block is longer than its height, the range of the NSST index value is determined by comparing a value obtained by dividing the width by the height with the preset threshold value. Also, if the height of the target block is longer than its width, the range of the NSST index value is determined by comparing a value obtained by dividing the height by the width with the preset threshold value.
[0136] In addition, the range of values of the NSST index is determined based on, for example, an intra prediction mode applied to a block. For example, the range of values of the NSST index is determined based on whether the intra prediction mode applied to the current block is a non-directional intra prediction mode or a directional intra prediction mode.
[0137] Alternatively, the range of values of the NSST index may be determined based on whether the intra prediction mode applied to the current block is included in Category A or Category B. For example, Category A includes intra prediction modes 2, 10, 18, 26, 34, 42, 50, 58, and 66, and Category B includes intra prediction modes other than those included in Category A. The intra prediction modes included in Category A may be pre-defined, or Category A and Category B may be pre-defined to include intra prediction modes different from those in the above examples.
[0138] Alternatively, the range of values of the NSST index may be determined based on the AMT factor of the block, which may also be referred to as the MTS factor.
[0139] For example, the AMT factor may be defined as the AMT flag described above, in which case the range of values of the NSST index is determined based on the value of the AMT flag of the target block.
[0140] Alternatively, the AMT factor may be defined as the AMT vertical flag and / or the AMT horizontal flag described above. In this case, the range of the NSST index value is determined based on the value of the AMT vertical flag and / or the AMT horizontal flag of the target block.
[0141] Alternatively, the AMT factor may be defined as a transform kernel applied in a multi-core transform, in which case the range of values of the NSST index is determined based on the transform kernel applied in the multi-core transform of the current block.
[0142] Alternatively, the range of values of the NSST index may be determined based on the components of the block. For example, the range of values of the NSST index for the luma block of the target block and the range of values of the NSST index for the chroma block of the target block may be applied differently.
[0143] On the other hand, the range of the NSST index value may be determined by a combination of the above-mentioned specific conditions.
[0144] The range of the NSST index value determined based on the specific condition, i.e., the maximum value of the NSST index, can be set in various ways.
[0145] For example, the maximum value of the NSST index is determined as R1, R2, or R3 based on the specific condition. Specifically, if the specific condition corresponds to category A, the maximum value of the NSST index is derived as R1, if the specific condition corresponds to category B, the maximum value of the NSST index is derived as R2, and if the specific condition corresponds to category C, the maximum value of the NSST index is derived as R3.
[0146] R1 for the category A, R2 for the category B, and R3 for the category C are derived as shown in the table below.
[0147] [Table 5]
[0148] The R1, R2, and R3 may be already set. For example, the relationship between the R1, R2, and R3 is derived as shown in the following formula:
[0149]
number
[0150] Referring to Equation 4, R1 is greater than or equal to 0, R2 is greater than R1, and R3 is greater than R2. On the other hand, if R1 is 0 and the maximum value of the NSST index for the current block is determined to be R1, the NSST index is not signaled and the value of the NSST index is inferred as 0.
[0151] In addition, an implicit NSST index coding method is proposed in the present invention.
[0152] In general, when NSST is applied, the distribution of non-zero transform coefficients among transform coefficients may be changed. In particular, when a reduced secondary transform (RST) is used as a secondary transform under certain conditions, the NSST index may not be coded.
[0153] Here, the RST indicates a quadratic transformation using a simplified transformation matrix as a non-separable transformation matrix. The simplified transformation matrix is determined by mapping an N-dimensional vector to an R-dimensional vector located in another space, where R is smaller than N. N refers to the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor may refer to an R / N value. The simplification factor is also referred to by various terms such as a reduced factor, reduction factor, simplified factor, or simple factor. Meanwhile, R may be referred to as a simplification coefficient, but in some cases the simplification factor may refer to R. In other cases, the simplification factor may refer to an N / R value.
[0154] The size of the simplified transformation matrix according to one embodiment is R×N, which is smaller than the size N×N of a normal transformation matrix, and is defined as Equation 5 below.
[0155]
number
[0156] The simplified transformation matrix T is used for the transformation coefficients to which the linear transformation of the target block has been applied. R×N are multiplied to derive the (secondary) transform coefficients for the current block.
[0157] When the RST is applied, a simplified transform matrix of size R×N is applied to the secondary transform, so that the transform coefficients R+1 to N may implicitly be 0. In other words, when the transform coefficients of the target block are derived by applying the RST, the values of the transform coefficients R+1 to N may be 0. Here, the transform coefficients R+1 to N refer to the R+1th to Nth transform coefficients among the transform coefficients. Specifically, the arrangement of the transform coefficients of the target block can be described as follows.
[0158] FIG. 8 is a diagram illustrating an arrangement of transform coefficients based on a current block according to an embodiment of the present invention. The following description of the transform described below with reference to FIG. 8 also applies to the inverse transform. An NSST (an example of a secondary transform) based on a primary transform and a simplified transform is performed on a current block (or residual block) 800. In one example, the 16×16 block shown in FIG. 8 represents the current block 800, and the 4×4 blocks labeled A through P represent subgroups of the current block 800. The primary transform is performed on the entire current block 800, and after the primary transform, the NSST is applied to an 8×8 block (hereinafter referred to as the upper left target region) consisting of subgroups A, B, E, and F. Here, when the NSST based on the simplified transform is performed, only R NSST coefficients (where R represents a simplified coefficient and R is less than N) are derived, and therefore, the R+1th to Nth NSST coefficients are determined to be 0. For example, if R is 16, the 16 transform coefficients derived by performing NSST based on the simplified transform are assigned to each block included in subgroup A, which is the upper left 4x4 block included in the upper left target area of target block 800, and transform coefficient 0 is assigned to each of NR blocks, i.e., 64-16=48 blocks, included in subgroups B, E, and F. Primary transform coefficients not subjected to NSST based on the simplified transform are assigned to each block included in subgroups C, D, G, H, I, J, K, L, M, N, O, and P.
[0159] Therefore, if at least one non-zero transform coefficient is derived by scanning transform coefficients R+1 to N, it is determined that the RST is not applied, and the value of the NSST index may be implicitly set to 0 without any additional signaling. That is, if at least one non-zero transform coefficient is derived by scanning transform coefficients R+1 to N, the RST is not applied, and the value of the NSST index is derived as 0 without any additional signaling.
[0160] FIG. 9 shows an example of scanning transform coefficients from R+1 to N.
[0161] As shown in FIG. 9, the size of a target block to which a transform is applied may be 64×64, with R=16 (i.e., R / N=16 / 64=1 / 4). That is, FIG. 9 shows the upper left target region of the target block. A simplified transform matrix of 16×64 size may be applied to the secondary transform of 64 samples of the upper left target region of the target block. In this case, when the RST is applied to the upper left target region, the values of the transform coefficients 17 to 64(N) must be 0. In other words, if at least one non-zero transform coefficient is derived among the transform coefficients 17 to 64 of the target block, the RST is not applied, and the value of the NSST index is derived as 0 without any additional signaling. Therefore, the decoding device decodes the transform coefficients of the current block, scans transform coefficients 17 to 64 among the decoded transform coefficients, and if a non-zero transform coefficient is derived, the decoding device may derive the value of the NSST index as 0 without signaling a separate syntax element for the NSST index. On the other hand, if there is no non-zero transform coefficient among the 17 to 64 transform coefficients, the decoding device may receive and decode the NSST index.
[0162] 10a and 10b are flowcharts illustrating a coding process of an NSST index according to an embodiment.
[0163] Figure 10a shows the encoding process of the NSST index.
[0164] The encoding device encodes transform coefficients for a current block (S1000). The encoding device performs entropy encoding on the quantized transform coefficients. Entropy encoding includes encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).
[0165] The encoding device determines whether an (explicit) NSST index for a current block is to be coded (S1010). Here, the explicit NSST index indicates an NSST index to be transmitted to a decoding device. That is, the encoding device can determine whether to generate an NSST index to be signaled. In other words, the encoding device can determine whether to allocate bits for a syntax element for the NSST index. If the decoding device can derive the value of the NSST index even if the NSST index is not signaled, as in the above embodiment, the encoding device may not code the NSST index. A specific process of determining whether an NSST index is to be coded will be described later.
[0166] If it is determined that the (explicit) NSST index is to be coded, the encoding device encodes the NSST index (S1020).
[0167] Figure 10b shows the decoding process of the NSST index.
[0168] The decoding device decodes the transform coefficients for the current block (S1030).
[0169] The decoding apparatus determines whether an (explicit) NSST index for the current block is coded (S1040). Here, the explicit NSST index indicates an NSST index signaled from the encoding apparatus. If the decoding apparatus can derive the value of the NSST index even if the NSST index is not signaled, as in the above-described embodiment, the NSST index may not be signaled from the encoding apparatus. A detailed process of determining whether an NSST index is coded will be described later.
[0170] If it is determined that the (explicit) NSST index is to be coded, the encoding device decodes the NSST index (S1040).
[0171] FIG. 11 shows an example of determining whether an NSST index is coded.
[0172] The encoding / decoding device determines whether a condition for coding an NSST index for a current block is met (S1100). For example, if the cbf flag for the current block indicates 0, the encoding / decoding device determines not to code an NSST index for the current block. Alternatively, if the current block is coded in a transform skip mode or if the number of non-zero transform coefficients for the current block is less than a predetermined threshold, the encoding / decoding device determines not to code an NSST index for the current block. For example, the predetermined threshold may be 2.
[0173] If a condition for coding an NSST index for the current block is met, the encoding / decoding device scans transform coefficients R+1 to N (S1110), which indicate the R+1 to N transform coefficients in the scan order.
[0174] The encoding / decoding device determines whether a non-zero transform coefficient is derived from the transform coefficients R+1 to N (S1120). If a non-zero transform coefficient is derived from the transform coefficients R+1 to N, the encoding / decoding device determines not to code an NSST index for the current block. In this case, the encoding / decoding device may derive an NSST index value for the current block as 0. In other words, for example, if an NSST index value of 0 indicates that NSST is not applied, the encoding / decoding device may not perform NSST on the top-left target region of the current block.
[0175] On the other hand, if no non-zero transform coefficients are derived from the transform coefficients R+1 to N, the encoding device encodes an NSST index for the current block, and the decoding device decodes the NSST index for the current block.
[0176] Meanwhile, it is proposed to use the NSST index common to the components (luma component, chroma Cb component, chroma Cr component) of the target block.
[0177] For example, the same NSST index is used for the chroma Cb block of the target block and the chroma Cr block of the target block.As another example, the same NSST index is used for the luma block of the target block, the chroma Cb block of the target block, and the chroma Cr block of the target block.
[0178] When two or three components of the current block use the same NSST index, the encoding device scans transform coefficients R+1 to N of all components (the luma block, chroma Cb block, and chroma Cr block of the current block), and if at least one non-zero transform coefficient is derived, the encoding device does not encode the NSST index but derives the value of the NSST index as 0. Also, the decoding device scans transform coefficients R+1 to N of all components (the luma block, chroma Cb block, and chroma Cr block of the current block), and if at least one non-zero transform coefficient is derived, the decoding device does not decode the NSST index but derives the value of the NSST index as 0.
[0179] FIG. 12 shows an example of scanning transform coefficients R+1 to N for all components of a current block.
[0180] As shown in FIG. 12, the luma block, chroma Cb block, and chroma Cr block of a target block to which a transform is applied may have a size of 64×64, where R=16 (i.e., R / N=16 / 64=1 / 4). That is, FIG. 12 shows the upper left target region of the luma block, the upper left target region of the chroma Cb block, and the upper left target region of the chroma Cr block. Therefore, a simplified transform matrix of 16×64 size may be applied to the secondary transform of 64 samples in each of the upper left target region of the luma block, the upper left target region of the chroma Cb block, and the upper left target region of the chroma Cr block. In this case, when the RST is applied to the upper left target region of the luma block, the upper left target region of the chroma Cb block, and the upper left target region of the chroma Cr block, the values of transform coefficients 17 to 64(N) of each block must be 0. In other words, if at least one non-zero transform coefficient is derived among the 17 to 64 transform coefficients of each block, the RST is not applied and the value of the NSST index is derived as 0 without any additional signaling. Therefore, a decoding device decodes transform coefficients for all components of a current block, scans the 17 to 64 transform coefficients of the luma block, the chroma Cb block, and the chroma Cr block among the decoded transform coefficients, and if a non-zero transform coefficient is derived, derives the value of the NSST index as 0 without signaling a separate syntax element for the NSST index. On the other hand, if there is no non-zero transform coefficient among the 17 to 64 transform coefficients, the decoding device may receive and decode the NSST index. The NSST index is used as an index for the luma block, the chroma Cb block, and the chroma Cr block.
[0181] Furthermore, the present invention proposes a scheme for signaling an NSST index indicator at an upper level. NSST_Idx_indicator indicates a syntax element for the NSST index indicator. For example, the NSST index indicator is coded at a CTU (Coding Tree Unit) level, and the NSST index indicator indicates whether NSST is applied to a target CTU. That is, the NSST index indicator indicates whether NSST is available for a target CTU. Specifically, if the NSST index indicator for the target CTU is enabled (if NSST is available for the target CTU), i.e., if the value of the NSST index indicator is 1, an NSST index for a CU or TU included in the target CTU is coded. If the NSST index indicator for the target CTU is not activated (if NSST is not available for the target CTU), i.e., if the value of the NSST index indicator is 0, an NSST index for a CU or TU included in the target CTU is not coded. Meanwhile, the NSST index indicator can be coded at the CTU level as described above, or can be coded at the sample group level of any other size, for example, the NSST index indicator can be coded at the CU (Coding Unit) level.
[0182] Figure 13 schematically illustrates an image encoding method using an encoding device according to the present invention. The method disclosed in Figure 13 can be performed by the encoding device disclosed in Figure 1. Specifically, for example, S1300 in Figure 13 can be performed by a subtraction unit of the encoding device, S1310 can be performed by a transformation unit of the encoding device, and S1320 to S1330 can be performed by an entropy encoding unit of the encoding device. Also, although not shown, the process of deriving predicted samples can be performed by a prediction unit of the encoding device.
[0183] The encoding device derives residual samples of a current block (S1300). For example, the encoding device determines whether to perform inter prediction or intra prediction on the current block, and determines a specific inter prediction mode or a specific intra prediction mode based on an RD cost. The encoding device derives predicted samples for the current block according to the determined mode, and derives the residual samples by adding original samples for the current block and the predicted samples.
[0184] The encoding apparatus performs a transform on the residual samples to derive transform coefficients of the current block (S1310). The encoding apparatus determines whether to apply NSST to the current block.
[0185] When the NSST is applied to the target block, the encoding device derives modified transform coefficients by performing a core transform on the residual samples, and then derives the transform coefficients of the target block by performing an NSST on the modified transform coefficients located in the upper left target region of the target block based on a simplified transform matrix. Modified transform coefficients other than the modified transform coefficient located in the upper left target region of the target block are directly derived as the transform coefficients of the target block. The size of the simplified transform matrix is R×N, where N is the number of samples in the upper left target region, R is a reduced coefficient, and R is less than N.
[0186] Specifically, the core transform for the residual samples is performed as follows: The encoding device may determine whether to apply an adaptive multiple core transform (AMT) to the current block. In this case, an AMT flag indicating whether the adaptive multiple core transform for the current block is applied is generated. If the AMT is not applied to the current block, the encoding device derives a DCT type 2 as a transform kernel for the current block, and performs a transform on the residual samples based on the DCT type 2 to derive the modified transform coefficients.
[0187] When the AMT is applied to the current block, the encoding apparatus forms a transform subset for a horizontal transform kernel and a transform subset for a vertical transform kernel, derives a horizontal transform kernel and a vertical transform kernel based on the transform subset, and performs a transform on the residual samples based on the horizontal transform kernel and the vertical transform kernel to derive modified transform coefficients. Here, the transform subset for the horizontal transform kernel and the transform subset for the vertical transform kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Transform index information may also be generated, and the transform index information includes an AMT horizontal flag indicating the horizontal transform kernel and an AMT vertical flag indicating the vertical transform kernel. Meanwhile, the transform kernels may also be referred to as transform types or transform cores.
[0188] On the other hand, if the NSST is not applied to the current block, the encoding apparatus may perform a core transform on the residual samples to derive the transform coefficients of the current block.
[0189] Specifically, the core transform for the residual samples is performed as follows: The encoding device determines whether to apply an adaptive multiple core transform (AMT) to the current block. In this case, an AMT flag indicating whether the adaptive multiple core transform for the current block is applied is generated. If the AMT is not applied to the current block, the encoding device derives a DCT type 2 as a transform kernel for the current block, and performs a transform on the residual samples based on the DCT type 2 to derive the transform coefficients.
[0190] When the AMT is applied to the current block, the encoding apparatus forms a transform subset for a horizontal transform kernel and a transform subset for a vertical transform kernel, derives a horizontal transform kernel and a vertical transform kernel based on the transform subset, and performs a transform on the residual samples based on the horizontal transform kernel and the vertical transform kernel to derive transform coefficients. Here, the transform subset for the horizontal transform kernel and the transform subset for the vertical transform kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Transform index information may also be generated, and the transform index information includes an AMT horizontal flag indicating the horizontal transform kernel and an AMT vertical flag indicating the vertical transform kernel. Meanwhile, the transform kernels may also be referred to as transform types or transform cores.
[0191] The encoding device determines whether to encode the NSST index (S1320).
[0192] For example, the encoding device scans the (R+1)th to (N)th transform coefficients of the target block, and if the (R+1)th to (N)th transform coefficients include a non-zero transform coefficient, determines not to encode the NSST index, where N is the number of samples in the upper left target region, R is a reduced coefficient, and R is less than N. N is derived as the product of the width and height of the upper left target region.
[0193] Furthermore, if the (R+1)th to Nth transform coefficients do not include a non-zero transform coefficient, the encoding apparatus determines to encode the NSST index. In this case, the information about the transform coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits for the syntax element for the NSST index are allocated.
[0194] Meanwhile, the encoding device determines whether the conditions for performing NSST are met, and if the NSST is feasible, determines to encode an NSST index for the current block. For example, an NSST index indicator for a current CTU including the current block is generated from a bitstream, and the NSST index indicator indicates whether NSST is applied to the current CTU. If the value of the NSST index indicator is 1, the encoding device determines to encode an NSST index for the current block, and if the value of the NSST index indicator is 0, the decoding device determines not to encode an NSST index for the current block. As in the above example, the NSST index indicator is signaled at a CTU level, and the NSST index indicator is signaled at a CU level or another higher level.
[0195] The NSST index is also used for multiple components of the current block.
[0196] For example, the NSST index is used for inverse transform of the transform coefficients of the luma block, the transform coefficients of the chroma Cb block, and the transform coefficients of the chroma Cr block of the current block. In this case, the R+1th to Nth transform coefficients of the luma block, the R+1th to Nth transform coefficients of the chroma Cb block, and the R+1th to Nth transform coefficients of the chroma Cr block are scanned. If the scanned transform coefficients include a non-zero transform coefficient, it is determined that the NSST index is not to be encoded. If the scanned transform coefficients do not include a non-zero transform coefficient, it is determined that the NSST index is to be encoded. In this case, information about the transform coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0197] As another example, the NSST index is used for inverse transform of the transform coefficients of the luma block and the transform coefficients of the chroma Cb block of the current block. In this case, the R+1th to Nth transform coefficients of the luma block and the R+1th to Nth transform coefficients of the chroma Cb block are scanned, and if the scanned transform coefficients include a non-zero transform coefficient, it is determined that the NSST index is not to be encoded. If the scanned transform coefficients do not include a non-zero transform coefficient, it is determined that the NSST index is to be encoded. In this case, information about the transform coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0198] As another example, the NSST index is used for inverse transform of the transform coefficients of the luma block and the transform coefficients of the chroma Cr block of the current block. In this case, the R+1th to Nth transform coefficients of the luma block and the R+1th to Nth transform coefficients of the chroma Cr block are scanned, and if the scanned transform coefficients include a non-zero transform coefficient, it is determined that the NSST index is not to be encoded. If the scanned transform coefficients do not include a non-zero transform coefficient, it is determined that the NSST index is to be encoded. In this case, information about the transform coefficients includes a syntax element for the NSST index. That is, the syntax element for the NSST index is encoded. In other words, bits are allocated for the syntax element for the NSST index.
[0199] Meanwhile, a range of the NSST index may be derived based on a specific condition. For example, a maximum value of the NSST index may be derived based on the specific condition, and the range may be derived as 0 to the derived maximum value. The derived NSST index value may be included in the range.
[0200] For example, the range of the NSST index is derived based on the size of the target block. Specifically, a minimum width and a minimum height are pre-set, and the range of the NSST index is derived based on the width and the minimum width of the target block, and the height and the minimum height of the target block. Also, the range of the NSST index is derived based on the number of samples of the target block and a specific value. The number of samples is a value obtained by multiplying the width and height of the target block, and the specific value may be pre-set.
[0201] As another example, the range of the NSST index is derived based on the type of the target block. Specifically, the range of the NSST index is derived based on whether the target block is a non-square block. The range of the NSST index is derived based on a ratio between the width and height of the target block and a specific value. The ratio between the width and height of the target block is a value obtained by dividing the longer side of the width and height of the target block by the shorter side, and the specific value may be preset.
[0202] As another example, the range of the NSST index may be derived based on the intra prediction mode of the current block. Specifically, the range of the NSST index may be derived based on whether the intra prediction mode of the current block is a non-directional intra prediction mode or a directional intra prediction mode. The range of the NSST index may be derived based on whether the intra prediction mode of the current block is an intra prediction mode included in Category A or Category B. Here, the intra prediction modes included in Category A and the intra prediction modes included in Category B may be pre-set. As an example, Category A includes intra prediction modes 2, 10, 18, 26, 34, 42, 50, 58, and 66, and Category B includes intra prediction modes other than those included in Category A.
[0203] As another example, the range of NSST indexes may be derived based on information about a core transform of the current block. For example, the range of NSST indexes may be derived based on an Adaptive Multiple Core Transform (AMT) flag indicating whether an AMT is applied. Furthermore, the range of NSST indexes may be derived based on an AMT horizontal flag indicating a horizontal transform kernel and an AMT vertical flag indicating a vertical transform kernel.
[0204] On the other hand, if the value of the NSST index is 0, the NSST index indicates that NSST is not applied to the current block.
[0205] The encoding device encodes information about transform coefficients (S1330). The information about the transform coefficients includes information about the size, position, etc. of the transform coefficients. As described above, the information about the transform coefficients may further include the NSST index, the transform index information, and / or the AMT flag. Image information including the information about the transform coefficients is output in the form of a bitstream. The image information may further include the NSST index indicator and / or prediction information. The prediction information is information about the prediction procedure, and includes prediction mode information and information about motion information (e.g., when inter prediction is applied).
[0206] The output bitstream is transmitted to a decoding device via a storage medium or a network.
[0207] Figure 14 schematically illustrates an encoding device that performs an image encoding method according to the present invention. The method disclosed in Figure 13 can be performed by the encoding device disclosed in Figure 14. Specifically, for example, an adder unit of the encoding device of Figure 14 can perform S1300 of Figure 13, a transform unit of the encoding device can perform S1310, and an entropy encoder unit of the encoding device can perform S1320 to S1330 of Figure 13. Also, although not shown, a process of deriving predicted samples can be performed by a prediction unit of the encoding device.
[0208] Figure 15 schematically illustrates an image decoding method by a decoding device according to the present invention. The method disclosed in Figure 15 can be performed by the decoding device disclosed in Figure 3. Specifically, for example, steps S1500 to S1510 in Figure 15 can be performed by an entropy decoding unit of the decoding device, step S1520 can be performed by an inverse transform unit of the decoding device, and step S1530 can be performed by an adder unit of the decoding device. Also, although not shown, the process of deriving predicted samples is performed by a prediction unit of the decoding device.
[0209] The decoding device derives transform coefficients of the current block from the bitstream (S1500). The decoding device derives the transform coefficients of the current block by decoding information about the transform coefficients of the current block received through the bitstream. The received information about the transform coefficients of the current block is referred to as residual information.
[0210] Meanwhile, the transform coefficients of the current block include transform coefficients of the luma block of the current block, transform coefficients of the chroma Cb block of the current block, and transform coefficients of the chroma Cr block of the current block.
[0211] The decoding device derives a non-separable secondary transform (NSST) index for the current block (S1510).
[0212] For example, the decoding device scans the R+1th to Nth transform coefficients of the target block, and if the R+1th to Nth transform coefficients include a non-zero transform coefficient, derives the value of the NSST index as 0. Here, N is the number of samples in the upper left target region of the target block, and R is a reduced coefficient, which is smaller than N. N is derived as the product of the width and height of the upper left target region.
[0213] Furthermore, if the (R+1)th to (N)th transform coefficients do not include a non-zero transform coefficient, the decoding device derives the value of the NSST index by parsing a syntax element for the NSST index included in the bitstream. That is, if the (R+1)th to (N)th transform coefficients do not include a non-zero transform coefficient, the bitstream includes a syntax element for the NSST index, and the decoding device derives the value of the NSST index by parsing the syntax element for the NSST index received via the bitstream.
[0214] Meanwhile, the decoding device determines whether the condition that NSST is executable is met, and if the condition that NSST is executable is met, derives an NSST index for the current block. For example, an NSST index indicator for a current CTU including the current block is signaled from a bitstream, and the NSST index indicator indicates whether NSST is enabled for the current CTU. If the value of the NSST index indicator is 1, the decoding device derives an NSST index for the current block, and if the value of the NSST index indicator is 0, the decoding device may not derive an NSST index for the current block. As in the above example, the NSST index indicator may be signaled at a CTU level, or the NSST index indicator may be signaled at a CU level or another higher level.
[0215] Also, the NSST index is used for multiple components of the current block.
[0216] For example, the NSST index is used for inverse transform of the transform coefficients of the luma block, the transform coefficients of the chroma Cb block, and the transform coefficients of the chroma Cr block of the current block. In this case, the R+1th to Nth transform coefficients of the luma block, the R+1th to Nth transform coefficients of the chroma Cb block, and the R+1th to Nth transform coefficients of the chroma Cr block are scanned, and if a non-zero transform coefficient is included in the scanned transform coefficients, the value of the NSST index is derived as 0. If a non-zero transform coefficient is not included in the scanned transform coefficients, the bitstream includes a syntax element for the NSST index, and the value of the NSST index is derived by parsing the syntax element for the NSST index received via the bitstream.
[0217] As another example, the NSST index is used for inverse transform of the transform coefficients of the luma block and the transform coefficients of the chroma Cb block of the current block. In this case, the R+1th to Nth transform coefficients of the luma block and the R+1th to Nth transform coefficients of the chroma Cb block are scanned, and if the scanned transform coefficients include a non-zero transform coefficient, the value of the NSST index is derived as 0. If the scanned transform coefficients do not include a non-zero transform coefficient, the bitstream includes a syntax element for the NSST index, and the value of the NSST index is derived by parsing the syntax element for the NSST index received via the bitstream.
[0218] As another example, the NSST index is used for inverse transform of the transform coefficients of the luma block and the transform coefficients of the chroma Cr block of the current block. In this case, the R+1-th to N-th transform coefficients of the luma block and the R+1-th to N-th transform coefficients of the chroma Cr block are scanned, and if the scanned transform coefficients include a non-zero transform coefficient, the value of the NSST index is derived as 0. If the scanned transform coefficients do not include a non-zero transform coefficient, the bitstream includes a syntax element for the NSST index, and the value of the NSST index is derived by parsing the syntax element for the NSST index received via the bitstream.
[0219] Meanwhile, a range of the NSST index may be derived based on a specific condition. For example, a maximum value of the NSST index may be derived based on the specific condition, and the range may be derived as 0 to the derived maximum value. The derived NSST index value may be included in the range.
[0220] For example, the range of the NSST index is derived based on the size of the target block. Specifically, a minimum width and a minimum height are pre-set, and the range of the NSST index is derived based on the width and the minimum width of the target block, and the height and the minimum height of the target block. Also, the range of the NSST index is derived based on the number of samples of the target block and a specific value. The number of samples is a value obtained by multiplying the width and height of the target block, and the specific value may be pre-set.
[0221] As another example, the range of the NSST index may be derived based on the type of the target block. Specifically, the range of the NSST index may be derived based on whether the target block is a non-square block. The range of the NSST index may also be derived based on a ratio between the width and height of the target block and a specific value. The ratio between the width and height of the target block is a value obtained by dividing the longer side of the width and height of the target block by the shorter side, and the specific value may be preset.
[0222] As another example, the range of the NSST index may be derived based on the intra prediction mode of the current block. Specifically, the range of the NSST index may be derived based on whether the intra prediction mode of the current block is a non-directional intra prediction mode or a directional intra prediction mode. The range of the NSST index may also be derived based on whether the intra prediction mode of the current block is an intra prediction mode included in Category A or Category B. Here, the intra prediction modes included in Category A and the intra prediction modes included in Category B may be pre-set. For example, Category A includes intra prediction modes 2, 10, 18, 26, 34, 42, 50, 58, and 66, and Category B includes intra prediction modes other than those included in Category A.
[0223] As another example, the range of NSST indexes may be derived based on information about a core transform of the current block. For example, the range of NSST indexes may be derived based on an Adaptive Multiple Core Transform (AMT) flag indicating whether an AMT is applied. Furthermore, the range of NSST indexes may be derived based on an AMT horizontal flag indicating a horizontal transform kernel and an AMT vertical flag indicating a vertical transform kernel.
[0224] On the other hand, if the value of the NSST index is 0, the NSST index indicates that NSST is not applied to the current block.
[0225] The decoding apparatus performs an inversed transform on the transform coefficients of the current block based on the NSST index to derive residual samples of the current block (S1520).
[0226] For example, if the value of the NSST index is 0, the decoding device performs a core transform on the transform coefficients of the current block to derive the residual samples.
[0227] Specifically, the decoding device obtains an Adaptive Multiple Core Transform (AMT) flag from the bitstream, which indicates whether or not AMT is applied.
[0228] If the value of the AMT flag is 0, the decoding apparatus derives a DCT type 2 as a transform kernel for the current block, and performs an inverse transform on the transform coefficients based on the DCT type 2 to derive the residual samples.
[0229] If the AMT flag has a value of 1, the decoding device configures a transform subset for a horizontal transform kernel and a transform subset for a vertical transform kernel, derives a horizontal transform kernel and a vertical transform kernel based on transform index information acquired from the bitstream and the transform subsets, and performs an inverse transform on the transform coefficients based on the horizontal transform kernel and the vertical transform kernel to derive the residual samples. Here, the transform subset for the horizontal transform kernel and the transform subset for the vertical transform kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. Also, the transform index information includes an AMT horizontal flag indicating one of the candidates included in the transform subset for the horizontal transform kernel and an AMT vertical flag indicating one of the candidates included in the transform subset for the vertical transform kernel. Meanwhile, the transform kernels may be referred to as transform types or transform cores.
[0230] As another example, if the value of the NSST index is not 0, the decoding device derives modified transform coefficients by performing NSST on transform coefficients located in the upper left target region of the target block based on a reduced transform matrix indicated by the NSST index, and derives the residual samples by performing core transform on the target block including the modified transform coefficients. The size of the reduced transform matrix is R×N, where N is the number of samples in the upper left target region, R is a reduced coefficient, and R is less than N.
[0231] The core transform for the current block is performed as follows: The decoding device obtains an Adaptive Multiple Core Transform (AMT) flag from a bitstream, which indicates whether an AMT is applied, and if the value of the AMT flag is 0, the decoding device derives a DCT type 2 as a transform kernel for the current block, and performs an inverse transform for the current block including the modified transform coefficients based on the DCT type 2 to derive the samples.
[0232] If the AMT flag has a value of 1, the decoding apparatus configures a transform subset for a horizontal transform kernel and a transform subset for a vertical transform kernel, derives a horizontal transform kernel and a vertical transform kernel based on transform index information acquired from the bitstream and the transform subsets, and performs an inverse transform on the current block including the modified transform coefficients based on the horizontal transform kernel and the vertical transform kernel to derive the residual sample. Here, the transform subset for the horizontal transform kernel and the transform subset for the vertical transform kernel include DCT type 2, DST type 7, DCT type 8, and / or DST type 1 as candidates. In addition, the transform index information includes an AMT horizontal flag indicating one of the candidates included in the transform subset for the horizontal transform kernel and an AMT vertical flag indicating one of the candidates included in the transform subset for the vertical transform kernel. Meanwhile, the transform kernels may also be referred to as transform types or transform cores.
[0233] The decoding device generates a reconstructed picture based on the residual samples (S1530). The decoding device generates the reconstructed picture based on the residual samples. For example, the decoding device may derive predicted samples by performing inter-prediction or intra-prediction on a current block based on prediction information received via a bitstream, and generate the reconstructed picture by adding the predicted samples and the residual samples. As described above, in-loop filtering procedures such as deblocking filtering, SAO, and / or ALF procedures may be applied to the reconstructed picture as needed to improve subjective / objective image quality.
[0234] Figure 16 schematically illustrates a decoding device that performs an image decoding method according to the present invention. The method disclosed in Figure 15 can be performed by the decoding device disclosed in Figure 16. Specifically, for example, an entropy decoding unit of the decoding device of Figure 16 can perform steps S1500 to S1510 of Figure 15, an inverse transform unit of the decoding device of Figure 16 can perform step S1520 of Figure 15, and an adder unit of the decoding device of Figure 16 can perform step S1530 of Figure 15. Also, although not shown, a process of deriving predicted samples can be performed by a prediction unit of the decoding device of Figure 16.
[0235] According to the present invention described above, the range of the NSST index is derived based on specific conditions of the target block, thereby reducing the amount of bits for the NSST index and improving the overall coding efficiency.
[0236] In addition, according to the present invention, the transmission of syntax elements for the NSST index is determined based on the transform coefficients for the current block, thereby reducing the amount of bits for the NSST index and improving overall coding efficiency.
[0237] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but the present invention is not limited to the order of steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included or one or more steps of the flowcharts may be omitted without affecting the scope of the present invention.
[0238] The above-described method according to the present invention is implemented in software form, and the encoding device and / or decoding device according to the present invention is included in an image processing device such as a television, a computer, a smartphone, a set-top box, or a display device.
[0239] When an embodiment of the present invention is implemented in software, the above-described method may be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (Application Specific Integrated Circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (Read-Only Memory), RAM (Random Access Memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described herein may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be implemented on a computer, processor, microprocessor, controller, or chip.
[0240] In addition, the decoding device and encoding device to which the present invention is applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video on demand (VoD) service providing device, an over-the-top (OTT) video device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, a medical video device, etc., and may be used to process a video signal or a data signal. For example, over-the-top (OTT) video devices include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0241] Furthermore, a processing method according to the present invention can be produced in the form of a computer-executable program and stored on a computer-readable recording medium. Multimedia data having a data structure according to the present invention can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices on which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. The computer-readable recording medium also includes media embodied in the form of a carrier wave (e.g., transmission via the Internet). The bitstream generated by the encoding method can be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network. Furthermore, an embodiment of the present invention can be realized as a computer program product using program code, which is executed on a computer according to an embodiment of the present invention. The program code can be stored on a computer-readable carrier.
[0242] The content streaming system to which the present invention is applied includes an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0243] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server may be omitted. The bitstream is generated by an encoding method or a bitstream generation method to which the present invention is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0244] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. Here, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0245] The streaming server receives content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0246] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
Claims
1. An image decoding method performed by a decoding device, obtaining residual information for a current block from the bitstream; deriving transform coefficients of the current block based on the residual information; scanning the R+1th to Nth transform coefficients of the transform coefficients of the current block; deriving a non-separable transform index for the current block based on scanning the R+1 through N transform coefficients; deriving residual samples of the current block based on an inverse transform of the transform coefficients of the current block, the inverse transform being based on the non-separable transform index; generating a reconstructed picture based on the residual samples; applying deblocking filtering to the reconstructed picture; The value of the non-separable transform index is derived as 0 based on a case where the R+1th to Nth transform coefficients include a non-zero transform coefficient; When the R+1th to Nth transform coefficients do not include the non-zero transform coefficient, the inverse transform based on a non-separable transform is performed on coefficients in an upper-left object region of the object block based on a transform matrix associated with the non-separable transform index; The size of the transformation matrix is R×N, N is the number of samples in the top left region of interest, The image decoding method, wherein R is smaller than N.
2. 1. A method of image encoding performed by an encoding device, comprising: deriving a residual sample of the current block; deriving transform coefficients for the current block by performing a transform based on the residual samples; scanning the R+1th to Nth transform coefficients of the transform coefficients of the current block; determining whether to encode a non-separable transform index for the transform coefficient based on the step of scanning the R+1 th to N th transform coefficients; encoding residual information including information related to the transform coefficients and information related to the non-separable transform indexes; determining that the non-separable transform index is not to be encoded based on a case where the R+1th to Nth transform coefficients include a non-zero transform coefficient; determining that the non-separable transform index is to be encoded based on a case where the R+1th to Nth transform coefficients do not include the non-zero transform coefficient; based on the value of the non-separable transform index not equal to 0, the transform based on a non-separable transform is performed on coefficients in a top-left object region of the object block based on a transform matrix associated with the non-separable transform index; The size of the transformation matrix is R×N, N is the number of samples in the top left region of interest, The image encoding method, wherein R is smaller than N.
3. 1. A method for transmitting data including a bitstream of image information, comprising: obtaining the bitstream of the image information including residual information, the residual information including information related to transform coefficients and information related to non-separable transform indexes, the bitstream being generated by deriving residual samples of a current block, deriving transform coefficients of the current block by performing a transform based on the residual samples, scanning R+1th to Nth transform coefficients of the current block, determining whether to encode non-separable transform indexes for the transform coefficients based on scanning the R+1th to Nth transform coefficients, and encoding the residual information including the information related to the transform coefficients and the information related to the non-separable transform indexes; transmitting the data including the bitstream of the image information including the residual information; determining that the non-separable transform index is not to be encoded based on a case where the R+1th to Nth transform coefficients include a non-zero transform coefficient; determining that the non-separable transform index is to be encoded based on a case where the R+1th to Nth transform coefficients do not include the non-zero transform coefficient; based on the value of the non-separable transform index not equal to 0, the transform based on a non-separable transform is performed on coefficients in a top-left object region of the object block based on a transform matrix associated with the non-separable transform index; The size of the transformation matrix is R×N, N is the number of samples in the top left region of interest, A data transmission method, wherein R is smaller than N.
Citation Information
Patent Citations
Method and device for the transformation and method and device for the reverse transformation of images
US20130195177A1
Non-separable secondary transform for video coding with reorganizing
US20170094314A1
Enhanced multiple transforms for prediction residual
WO2016123091A1
Non-separable secondary transform for video coding
WO2017058614A1
Binarizing secondary transform index
WO2017192705A1