Image encoding / decoding method and apparatus, and recording medium having bitstream stored therein

The video encoding/decoding method addresses the challenge of high-resolution image data by adaptively determining transform sets and kernels based on prediction modes, enhancing encoding/decoding efficiency and reducing costs.

WO2025135643A1PCT designated stage expired Publication Date: 2025-06-26HYUNDAI MOTOR CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019984
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-06
Filing Date
2024-12-06
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images, such as UHD images, leads to a significant increase in image data, resulting in higher transmission and storage costs. Conventional non-separable transforms are limited in application due to difficulties in designing the transform kernel and optimized transform structure.

Method used

A video encoding/decoding method that determines a transform set and transform kernel based on the prediction mode of a current block, allowing for non-separable transformations to be applied not only to intra prediction blocks but also to inter prediction blocks and special intra prediction modes.

Benefits of technology

This approach improves encoding/decoding efficiency by adaptively selecting the appropriate transform set and kernel, reducing data redundancy and thereby decreasing transmission and storage costs while maintaining high image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019984_26062025_PF_FP_ABST
    Figure KR2024019984_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image encoding / decoding method and apparatus, a recording medium having a bitstream stored therein, and a transmission method. The image decoding method comprises the steps of: deriving a prediction mode of the current block; determining a transform set of the current block on the basis of the prediction mode; and performing inverse transformation on the current block on the basis of the transform set, wherein the transform set may be determined as any one of a first transform set for first intra prediction, a second transform set for second intra prediction, and a third transform set for inter prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method, device, and recording medium storing bitstream

[0001] The present invention relates to a video encoding / decoding method, a device, and a recording medium storing a bitstream. Specifically, the present invention relates to a video encoding / decoding method, a device, and a recording medium storing a bitstream based on a method for determining a transform set.

[0002] Recently, the demand for high-resolution, high-quality images, such as UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired or wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image encoding / decoding technologies for higher-resolution and higher-quality images are required.

[0003] Conventional non-separable transformations are applied only to intra prediction blocks, and their application is limited due to the difficulty in designing transformation kernels and optimized transformation structures.

[0004] The purpose of the present invention is to provide a video encoding / decoding method and device with improved encoding / decoding efficiency.

[0005] In addition, the present invention aims to provide a recording medium storing a bitstream generated by an image decoding method or device according to the present invention.

[0006] In addition, the present invention aims to provide a transformation set and a transformation kernel determination method to solve the above problems.

[0007] A video decoding method according to one embodiment of the present invention includes the steps of deriving a prediction mode of a current block, determining a transform set of the current block based on the prediction mode, and performing an inverse transform on the current block based on the transform set, wherein the transform set may be determined as any one of a first transform set for a first intra prediction, a second transform set for a second intra prediction, and a third transform set for inter prediction.

[0008] In the above image decoding method, when the prediction mode is an intra template matching mode, the transformation set can be determined as the third transformation set.

[0009] In the above image decoding method, the transformation set may be a transformation set of non-separable surroundings.

[0010] In the above image decoding method, when the prediction mode is an intra block copy mode, the transformation set can be determined as the third transformation set.

[0011] In the above image decoding method, the first intra prediction may include non-directional intra prediction and directional intra prediction, and the second intra prediction may include SGPM (Spatial geometric prediction mode) intra prediction.

[0012] In the above image decoding method, the step of performing the inverse transformation may include the step of determining a transformation kernel of the current block among transformation kernels included in the transformation set, and the step of performing the inverse transformation based on the transformation kernel.

[0013] In the above image decoding method, the step of determining the transformation kernel may include the step of deriving a transformation kernel determined intra prediction mode and the step of determining the transformation kernel based on the transformation kernel determined intra prediction mode.

[0014] In the above image decoding method, the transform kernel determination intra prediction mode can be derived based on direction information of some pixels included in the prediction block of the current block.

[0015] In the above image decoding method, the transformation kernel may be a transformation kernel with the smallest distortion among the transformation kernels included in the transformation set.

[0016] In the above image decoding method, if the prediction mode of the current block is one of an intra template matching mode and an intra block copy mode, the method further includes a step of obtaining intra prediction mode usage information indicating whether an intra prediction mode is used when determining the transformation kernel, and if the intra prediction mode usage information indicates that an intra prediction mode is used, the step of determining the transformation kernel may include a step of deriving a transformation kernel determined intra prediction mode and a step of determining the transformation kernel based on the transformation kernel determined intra prediction mode.

[0017] In the above image decoding method, when the current block is divided into a first sub-block and a second sub-block according to a division boundary, the transformation set can be determined as the second transformation set.

[0018] In the above image decoding method, when the current block is divided into a first sub-block and a second sub-block according to a division boundary, the transformation kernel can be determined based on a prediction mode of the first sub-block and a prediction mode of the second sub-block.

[0019] In the above image decoding method, when the current block is divided into a first sub-block and a second sub-block according to a division boundary, the method may further include a step of obtaining sub-block prediction mode usage information indicating whether a prediction mode of the first sub-block and a prediction mode of the second sub-block are used when determining the transformation kernel, and when the sub-block prediction mode usage information indicates that the prediction mode of the first sub-block and the prediction mode of the second sub-block are used, the transformation kernel may be determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.

[0020] In the above image decoding method, when the prediction mode of the current block is a SCIIP (Spatial combined intra-intra prediction) mode, the transformation kernel of the current block can be determined based on a first intra prediction mode of the SCIIP mode and a second intra prediction mode of the SCIIP mode.

[0021] A video encoding method according to one embodiment of the present invention includes the steps of determining a prediction mode of a current block, determining a transform set of the current block based on the prediction mode, and performing a transform on the current block based on the transform set, wherein the transform set may be determined as any one of a first transform set for a first intra prediction, a second transform set for a second intra prediction, and a third transform set for inter prediction.

[0022] A non-transitory computer-readable recording medium storing a bitstream generated by an image encoding method according to one embodiment of the present invention can store a bitstream generated by the image encoding method.

[0023] A method for transmitting a bitstream generated by a video encoding method according to one embodiment of the present invention includes a step of transmitting the bitstream, and the video encoding method may be:

[0024] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0025] According to the present invention, a video encoding / decoding method and device with improved encoding / decoding efficiency can be provided.

[0026] Additionally, according to the present invention, a method for determining a transformation set for inverse transformation of a current block can be provided.

[0027] Additionally, according to the present invention, a method for determining a transformation kernel for inverse transformation of a current block can be provided.

[0028] In addition, according to the present invention, it is possible to determine a suitable transformation set and transformation kernel, thereby improving transformation efficiency.

[0029] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0030] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0031] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.

[0032] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.

[0033] FIG. 4 is a diagram for explaining a method for determining a transformation set when the prediction mode of the current block is an intra template matching mode according to one embodiment of the present invention.

[0034] FIG. 5 is a diagram for explaining a method for determining a transformation set when the prediction mode of the current block is an intra block copy mode according to one embodiment of the present invention.

[0035] FIG. 6 is a diagram for explaining a method for determining a transformation set and a transformation kernel in a spatial geometric partitioning mode (SGPM) according to one embodiment of the present invention.

[0036] FIG. 7 is a flowchart illustrating a method for determining a transformation set according to one embodiment of the present invention.

[0037] FIG. 8 is a drawing exemplarily showing a content streaming system to which an embodiment according to the present invention can be applied.

[0038] A video decoding method according to one embodiment of the present invention includes the steps of deriving a prediction mode of a current block, determining a transform set of the current block based on the prediction mode, and performing an inverse transform on the current block based on the transform set, wherein the transform set may be determined as any one of a first transform set for a first intra prediction, a second transform set for a second intra prediction, and a third transform set for inter prediction.

[0039] The present invention is susceptible to various modifications and embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and substitutes falling within the spirit and scope of the present invention. In the drawings, similar reference numerals designate the same or similar functions throughout. The shape and size of elements in the drawings may be provided by way of example only for clarity. The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments, while different, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the embodiment. Accordingly, the detailed description set forth below is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments, if properly described, is defined only by the appended claims, along with the full scope equivalents to which such claims are entitled.

[0040] In the present invention, terms such as first, second, etc. may be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes a combination of multiple related described items or any of multiple related described items.

[0041] The components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and do not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0042] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In addition, some components of the present invention are not essential components that perform essential functions in the present invention and may be optional components merely for performance enhancement. The present invention may be implemented by including only components essential to realizing the essence of the present invention, excluding components used only for performance enhancement, and a structure including only essential components, excluding optional components used only for performance enhancement, is also within the scope of the present invention.

[0043] In an embodiment, the term "at least one" may mean one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In an embodiment, the term "a plurality of" may mean one of a number greater than or equal to 2, such as 2, 3, and 4.

[0044] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function may obscure the gist of this specification, the detailed description will be omitted. The same reference numerals will be used for identical components in the drawings, and duplicate descriptions of identical components will be omitted.

[0045] Glossary of Terms

[0046] Hereinafter, “video” may mean a single picture constituting a video, or may refer to the video itself. For example, “encoding and / or decoding of a video” may mean “encoding and / or decoding of a video,” or may mean “encoding and / or decoding of one of the videos constituting the video.”

[0047] Hereinafter, the terms "video" and "movie" may be used interchangeably and have the same meaning. Furthermore, the target image may be an encoding target image, which is the target of encoding, and / or a decoding target image, which is the target of decoding. Furthermore, the target image may be an input image input to an encoding device, or an input image input to a decoding device. Here, the target image may have the same meaning as the current image.

[0048] Hereinafter, the terms encoder and image encoding device may be used interchangeably and have the same meaning.

[0049] Hereinafter, the terms decoder and image decoding device may be used interchangeably and have the same meaning.

[0050] Hereinafter, “image”, “picture”, “frame” and “screen” may be used with the same meaning and may be used interchangeably.

[0051] Hereinafter, the term "target block" may refer to an encoding target block, which is the target of encoding, and / or a decoding target block, which is the target of decoding. Furthermore, the target block may refer to a current block, which is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably and have the same meaning.

[0052] Hereinafter, "block" and "unit" may be used with the same meaning and may be used interchangeably. In addition, "unit" may mean including a luminance component block and a corresponding chroma component block to distinguish it from a block. For example, a coding tree unit (CTU) may be composed of one luma component (Y) coding tree block (CTB) and two chroma component (Cb, Cr) coding tree blocks associated with it.

[0053] Hereinafter, “sample,” “pixel,” and “pixel” may be used interchangeably and have the same meaning. Here, a sample may represent a basic unit that constitutes a block.

[0054] Hereinafter, “inter” and “between screens” may be used interchangeably and have the same meaning.

[0055] Hereinafter, “intra” and “within screen” may be used interchangeably and have the same meaning.

[0056]

[0057] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0058] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device (100) may sequentially encode one or more images.

[0059] Referring to FIG. 1, the encoding device (100) may include an image segmentation unit (110), an intra prediction unit (120), a motion prediction unit (121), a motion compensation unit (122), a switch (115), a subtractor (113), a transformation unit (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse transformation unit (170), an adder (117), a filter unit (180), and a reference picture buffer (190).

[0060] Additionally, the encoding device (100) can generate a bitstream including encoded information through encoding an input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or can be streamed via a wired / wireless transmission medium.

[0061] The video segmentation unit (110) can segment the input video into various forms to increase the efficiency of video encoding / decoding. That is, the input video is composed of multiple pictures, and one picture can be hierarchically segmented and processed for compression efficiency, parallel processing, etc. For example, one picture can be segmented into one or more tiles or slices, which can then be segmented into multiple Coding Tree Units (CTUs). Alternatively, one picture can first be segmented into multiple sub-pictures defined as groups of rectangular slices, and each sub-picture can then be segmented into the tiles / slices. Here, the sub-pictures can be utilized to support the function of partially independently encoding / decoding and transmitting the picture. Since multiple sub-pictures can each be individually restored, there is an advantage of easy editing in applications that configure multi-channel input into a single picture. In addition, tiles can be segmented horizontally to generate bricks. Here, a brick can be utilized as the basic unit of intra-picture parallel processing. In addition, one CTU can be recursively split into a quadtree (QT), and the terminal node of the split can be defined as a coding unit (CU). The CU can be split into a prediction unit (PU) and a transformation unit (TU), and prediction and splitting can be performed. Meanwhile, the CU can be utilized as a prediction unit and / or a transformation unit itself. Here, for flexible splitting, each CTU can be recursively split into a multi-type tree (MTT) as well as a quadtree (QT). Splitting of a CTU into a multi-type tree can start from the terminal node of a QT, and the MTT can be composed of a binary tree (BT) and a triple tree (TT).For example, the MTT structure can be divided into vertical binary split mode (SPLIT_BT_VER), horizontal binary split mode (SPLIT_BT_HOR), vertical ternary split mode (SPLIT_TT_VER), and horizontal ternary split mode (SPLIT_TT_HOR). In addition, the minimum block size (MinQTSize) of the quad tree of the luminance block during splitting can be set to 16x16, the maximum block size (MaxBtSize) of the binary tree can be set to 128x128, and the maximum block size (MaxTtSize) of the triple tree can be set to 64x64. In addition, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the triple tree can be set to 4x4, and the maximum depth (MaxMttDepth) of the multi-type tree can be set to 4. Additionally, to improve the encoding efficiency of the I slice, a dual tree can be applied that uses different CTU partition structures for luminance and chrominance components. On the other hand, in the P and B slices, the luminance and chrominance CTBs (Coding Tree Blocks) within the CTU can be partitioned into a single tree that shares the coding tree structure.

[0062] The encoding device (100) may perform encoding on the input image in intra mode and / or inter mode. Alternatively, the encoding device (100) may perform encoding on the input image in a third mode (e.g., IBC mode, Palette mode, etc.) other than the intra mode and inter mode. However, if the third mode has functional characteristics similar to the intra mode or inter mode, it may be classified as intra mode or inter mode for convenience of explanation. In the present invention, the third mode will be classified and described separately only when a specific description is required.

[0063] When the intra mode is used as the prediction mode, the switch (115) can be switched to intra, and when the inter mode is used as the prediction mode, the switch (115) can be switched to inter. Here, the intra mode can mean an intra-screen prediction mode, and the inter mode can mean an inter-screen prediction mode. The encoding device (100) can generate a prediction block for an input block of an input image. In addition, after the prediction block is generated, the encoding device (100) can encode a residual block using a residual of the input block and the prediction block. The input image can be referred to as a current image that is currently a target of encoding. The input block can be referred to as a current block that is currently a target of encoding or an encoding target block.

[0064] When the prediction mode is intra mode, the intra prediction unit (120) can use samples of blocks already encoded / decoded around the current block as reference samples. The intra prediction unit (120) can perform spatial prediction on the current block using the reference samples, and can generate prediction samples for the input block through spatial prediction. Here, intra prediction can mean prediction within the screen.

[0065] As an intra prediction method, non-directional prediction modes such as DC mode and Planar mode, as well as directional prediction modes (e.g., 65 directions) can be applied. Here, the intra prediction method can be expressed as an intra prediction mode or an intra-screen prediction mode.

[0066] When the prediction mode is inter mode, the motion prediction unit (121) can search for an area that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched area. At this time, the area can be used as a search area. The reference image can be stored in the reference picture buffer (190). Here, when encoding / decoding for the reference image is processed, it can be stored in the reference picture buffer (190).

[0067] The motion compensation unit (122) can generate a prediction block for the current block by performing motion compensation using a motion vector. Here, inter prediction may mean inter-screen prediction or motion compensation.

[0068] The above motion prediction unit (121) and motion compensation unit (122) can generate a prediction block by applying an interpolation filter to a portion of an area within a reference image when the value of the motion vector does not have an integer value. In order to perform inter-screen prediction or motion compensation, it is possible to determine whether the motion prediction and motion compensation method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, Advanced Motion Vector Prediction (AMVP) mode, and Intra Block Copy (IBC) mode based on the encoding unit, and perform inter-screen prediction or motion compensation according to each mode.

[0069] In addition, based on the above inter-screen prediction method, the AFFINE mode of sub-PU based prediction, the SbTMVP (Subblock-based Temporal Motion Vector Prediction) mode, and the MMVD (Merge with MVD) mode and the GPM (Geometric Partitioning Mode) mode of PU based prediction can be applied. In addition, in order to improve the performance of each mode, the HMVP (History based MVP), the PAMVP (Pairwise Average MVP), the CIIP (Combined Intra / Inter Prediction), the AMVR (Adaptive Motion Vector Resolution), the BDOF (Bi-Directional Optical-Flow), the BCW (Bi-predictive with CU Weights), the LIC (Local Illumination Compensation), the TM (Template Matching), and the OBMC (Overlapped Block Motion Compensation) can be applied.

[0070] Among these, AFFINE mode is a technology that is used in both AMVP and MERGE modes and also has high encoding efficiency. In the existing video coding standard, since MC (Motion Compensation) is performed by considering only the parallel translation of the block, there was a disadvantage in that it could not properly compensate for motions that occur in reality, such as zoom in / out and rotation. To supplement this, a 4-parameter affine motion model using two control point motion vectors (CPMV) and a 6-parameter affine motion model using three control point motion vectors can be applied to inter prediction. Here, CPMV is a vector representing the affine motion model of one of the upper left, upper right, and lower left of the current block.

[0071] The subtractor (113) can generate a residual block using the difference between the input block and the predicted block. The residual block may also be referred to as a residual signal. The residual signal may refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a residual signal in block units.

[0072] The transform unit (130) can perform a transform on the residual block to generate a transform coefficient and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit (130) may also skip the transform on the residual block.

[0073] Quantized levels can be generated by applying quantization to transform coefficients or residual signals. In the following embodiments, quantized levels may also be referred to as transform coefficients.

[0074] For example, a 4x4 luminance residual block generated through intra prediction can be transformed using a basis vector based on DST (Discrete Sine Transform), and the remaining residual blocks can be transformed using a basis vector based on DCT (Discrete Cosine Transform). In addition, through RQT (Residual Quad Tree) technology, the transform block is divided into a quad tree shape for one block, and after performing transformation and quantization on each transform block divided through RQT, a coded block flag (cbf) can be transmitted to increase encoding efficiency when all coefficients become 0.

[0075] Another alternative is to apply Multiple Transform Selection (MTS) technology, which selectively performs transformation using multiple transformation bases. That is, instead of dividing CUs into TUs via RQT, a Sub-block Transform (SBT) technology can perform a function similar to TU division. Specifically, SBT is applied only to inter-screen prediction blocks, and unlike RQT, it can divide the current block into ½ or ¼ blocks vertically or horizontally, and then perform transformation on only one of the blocks. For example, in a vertically divided block, the transformation can be performed on the leftmost or rightmost block, and in a horizontally divided block, the transformation can be performed on the topmost or bottommost block.

[0076] Additionally, LFNST (Low Frequency Non-Separable Transform), a secondary transform technique that further transforms the residual signal converted to the frequency domain through DCT or DST, can be applied. LFNST additionally performs a transform on the low-frequency region of 4x4 or 8x8 in the upper left, which allows the residual coefficients to be concentrated in the upper left.

[0077] The quantization unit (140) can generate a quantized level by quantizing a transform coefficient or residual signal according to a quantization parameter (QP), and can output the generated quantized level. At this time, the quantization unit (140) can quantize the transform coefficient using a quantization matrix.

[0078] For example, a quantizer with QP values ​​of 0 to 51 can be used. Alternatively, if the image size is larger and high encoding efficiency is required, a QP of 0 to 63 can be used. In addition, a Dependent Quantization (DQ) method that uses two quantizers instead of a single quantizer can be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a specific quantizer, the quantizer to be used for the next transform coefficient can be selected based on the current state through a state transition model.

[0079] The entropy encoding unit (150) can generate a bitstream by performing entropy encoding according to a probability distribution on values ​​produced by the quantization unit (140) or coding parameter values ​​produced during the encoding process, and can output the bitstream. The entropy encoding unit (150) can perform entropy encoding on information about image samples and information for decoding the image. For example, the information for decoding the image can include syntax elements, etc.

[0080] When entropy encoding is applied, a small number of bits are allocated to symbols with a high occurrence probability, and a large number of bits are allocated to symbols with a low occurrence probability, thereby representing the symbols, whereby the size of the bit string for the symbols to be encoded can be reduced. The entropy encoding unit (150) can use an encoding method such as exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), or Context-Adaptive Binary Arithmetic Coding (CABAC) for entropy encoding. For example, the entropy encoding unit (150) can perform entropy encoding using a Variable Length Coding / Code (VLC) table. In addition, the entropy encoding unit (150) may perform arithmetic encoding using the binarization method, probability model, and context model derived from the binarization method of the target symbol and the probability model of the target symbol / bin.

[0081] In this regard, when applying CABAC, the table probability update method can be changed to a simple formula-based table update method to reduce the size of the probability table stored in the decryption device. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.

[0082] The entropy encoding unit (150) can change a two-dimensional block form coefficient into a one-dimensional vector form through a transform coefficient scanning method to encode a transform coefficient level (quantized level).

[0083] Coding parameters may include not only information (flags, indexes, etc.) encoded in an encoding device (100) and signaled to a decoding device (200), such as syntax elements, but also information derived during an encoding process or a decoding process, and may mean information necessary when encoding or decoding an image.

[0084] Here, signaling a flag or index may mean that the encoder entropy encodes the flag or index and includes it in the bitstream, and that the decoder entropy decodes the flag or index from the bitstream.

[0085] The encoded current image can be used as a reference image for other images to be processed later. Accordingly, the encoding device (100) can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference picture buffer (190).

[0086] The quantized level can be dequantized in the dequantization unit (160) and inversely transformed in the inverse transformation unit (170). The dequantized and / or inversely transformed coefficients can be combined with a prediction block through an adder (117), and a reconstructed block can be generated by combining the dequantized and / or inversely transformed coefficients and the prediction block. Here, the dequantized and / or inversely transformed coefficients refer to coefficients on which at least one of dequantization and inverse transformation has been performed, and may refer to a reconstructed residual block. The dequantization unit (160) and the inverse transformation unit (170) can be performed in the reverse process of the quantization unit (140) and the transformation unit (130).

[0087] The restoration block may pass through a filter unit (180). The filter unit (180) may apply a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter (BIF), a Luma Mapping with Chroma Scaling (LMCS), etc. as a filtering technique, in whole or in part, to the restoration sample, restoration block, or restoration image. The filter unit (180) may also be referred to as an in-loop filter. In this case, the in-loop filter is also used as a name excluding LMCS.

[0088] A deblocking filter can remove block distortion that occurs at the boundaries between blocks. Whether to apply a deblocking filter to the current block can be determined based on the samples contained in several columns or rows within the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering strength.

[0089] Sample adaptive offset can be used to compensate for encoding errors by adding an appropriate offset value to sample values. Sample adaptive offset can compensate for the offset from the original image on a sample-by-sample basis for deblocked images. This can be done by dividing the samples contained in the image into a fixed number of regions, determining the regions to be offset, and applying the offset to those regions. Alternatively, the offset can be applied by considering the edge information of each sample.

[0090] Bilateral filter (BIF) can also compensate for the offset from the original image on a sample-by-sample basis for the deblocked image.

[0091] An adaptive loop filter can perform filtering based on a comparison between a reconstructed image and the original image. By dividing the samples contained in the image into predetermined groups and determining the filter to be applied to each group, filtering can be performed differentially for each group. Information regarding whether to apply an adaptive loop filter can be signaled for each coding unit (CU), and the shape and filter coefficients of the adaptive loop filter applied to each block can vary.

[0092] In LMCS (Luma Mapping with Chroma Scaling), luma mapping (LM) refers to remapping luminance values ​​through a piece-wise linear model, and chroma scaling (CS) refers to a technique that scales the residual values ​​of chrominance components according to the average luminance value of the prediction signal. In particular, LMCS can be utilized as an HDR correction technique that reflects the characteristics of HDR (High Dynamic Range) images.

[0093] The restored block or restored image that has passed through the filter unit (180) may be stored in the reference picture buffer (190). The restored block that has passed through the filter unit (180) may be a part of the reference image. In other words, the reference image may be a restored image composed of restored blocks that have passed through the filter unit (180). The stored reference image may be used for inter-screen prediction or motion compensation thereafter.

[0094] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.

[0095] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.

[0096] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), a motion compensation unit (250), an adder (201), a switch (203), a filter unit (260), and a reference picture buffer (270).

[0097] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium, or a bitstream streamed through a wired / wireless transmission medium. The decoding device (200) can perform decoding on the bitstream in intra mode or inter mode. In addition, the decoding device (200) can generate a restored image or a decoded image through decoding, and can output the restored image or the decoded image.

[0098] If the prediction mode used for decryption is intra mode, the switch (203) can be switched to intra. If the prediction mode used for decryption is inter mode, the switch (203) can be switched to inter.

[0099] The decoding device (200) can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded may be referred to as a current block.

[0100] The entropy decoding unit (210) can generate symbols by performing entropy decoding according to a probability distribution for the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the reverse process of the entropy encoding method described above.

[0101] The entropy decoding unit (210) can change a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a transform coefficient scanning method to decode a transform coefficient level (quantized level).

[0102] The quantized level can be inversely quantized in the inverse quantization unit (220) and inversely transformed in the inverse transformation unit (230). The quantized level can be generated as a restored residual block as a result of performing inverse quantization and / or inverse transformation. At this time, the inverse quantization unit (220) can apply a quantization matrix to the quantized level. The inverse quantization unit (220) and inverse transformation unit (230) applied to the decoding device can apply the same technology as the inverse quantization unit (160) and inverse transformation unit (170) applied to the encoding device described above.

[0103] When intra mode is used, the intra prediction unit (240) can generate a predicted block by performing spatial prediction on the current block using sample values ​​of already decoded blocks surrounding the block to be decoded. The intra prediction unit (240) applied to the decoding device can apply the same technology as the intra prediction unit (120) applied to the encoding device described above.

[0104] When the inter mode is used, the motion compensation unit (250) can generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer (270) on the current block. The motion compensation unit (250) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. In order to perform motion compensation, it is possible to determine whether the motion compensation method of the prediction unit included in the corresponding encoding unit is skip mode, merge mode, AMVP mode, or current picture reference mode based on the encoding unit, and motion compensation can be performed according to each mode. The motion compensation unit (250) applied to the decoding device can apply the same technology as the motion compensation unit (122) applied to the encoding device described above.

[0105] The adder (201) can add the restored residual block and the predicted block to generate a restored block. The filter unit (260) can apply at least one of an Inverse-LMCS, a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the restored block or restored image. The filter unit (260) applied to the decoding device can apply the same filtering technology as that applied to the filter unit (180) applied to the encoding device described above.

[0106] The filter unit (260) can output a restored image. The restored block or restored image can be stored in the reference picture buffer (270) and used for inter prediction. The restored block that has passed through the filter unit (260) can be a part of the reference image. In other words, the reference image can be a restored image composed of restored blocks that have passed through the filter unit (260). The stored reference image can be used for inter-screen prediction or motion compensation thereafter.

[0107] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.

[0108] A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) in the form of a file or streaming through a digital storage medium or a network.

[0109] An encoding device (10) according to one embodiment may include a video source generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a reception unit (21), a decoding unit (22), and a rendering unit (23). The encoding unit (12) may be referred to as a video / image encoding unit, and the decoding unit (22) may be referred to as a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The reception unit (21) may be included in the decoding unit (22). The rendering unit (23) may include a display unit, and the display unit may be configured as a separate device or an external component.

[0110] The video source generation unit (11) can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit (11) can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data.

[0111] The encoding unit (12) can encode the input video / image. The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit (12) can output encoded data (encoded video / image information) in the form of a bitstream. The detailed configuration of the encoding unit (12) can also be configured in the same manner as the encoding device (100) of FIG. 1 described above.

[0112] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (21) of the decoding device (20) via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22).

[0113] The decoding unit (22) can decode video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12). The detailed configuration of the decoding unit (22) can also be configured in the same manner as the decoding device (200) of FIG. 2 described above.

[0114] The rendering unit (23) can render the decrypted video / image. The rendered video / image can be displayed through the display unit.

[0115]

[0116] In this specification, a transform set may refer to a group including a transform kernel, and a transform kernel may refer to a transform core used when applying a transformation from a spatial domain to a frequency domain. In addition, a transform set may refer to a kernel cluster.

[0117] A secondary transform may refer to a transform performed based on the correlation of the primary transform coefficients generated by the primary transform after the primary transform from the spatial domain to the frequency domain has been performed. Here, the primary transform and the primary transform may have the same meaning. By performing the secondary transform, higher compression efficiency can be achieved compared to when performing only the primary transform.

[0118] The first and second transformations can be performed as separable or non-separable transformations. Here, a separable transformation can mean that a transformation in the horizontal direction is performed first, followed by a transformation in the vertical direction. In addition, a non-separable transformation can mean that a transformation is performed simultaneously, rather than in the horizontal and vertical directions, using a non-separable transformation kernel in the form of a matrix.

[0119] In non-separable transformations, basis vectors are used. A basis vector is a vector whose size corresponds to the total number of pixels in the two-dimensional space where the transformation is performed. It is used to determine the overall characteristics and similarity of pixels in the two-dimensional space where the transformation is performed. Therefore, while the complexity of non-separable transformations is higher than that of separable transformations, they can achieve high encoding efficiency.

[0120] For secondary transforms such as LFNST (Low frequency non-separable transform), non-separable secondary transform kernels of various sizes can be used to improve coding efficiency and transform complexity for specific input block sizes.

[0121] In the case of NSPT (Non-separable primary transform), although it is a primary transform, similar to LFNST, non-separable primary transform kernels of various sizes can be used to improve coding efficiency and transform complexity for specific input block sizes.

[0122] A selected transformation kernel among multiple transformation kernels can be used to transform the current block. This can be called MTS (Multiple Transform Selection).

[0123] In addition, a transformation set can be selected from among multiple transformation sets, and transformation of the current block can be performed using at least one of multiple transformation kernels included in the selected transformation set. Here, selecting a transformation set from among multiple transformation sets can be referred to as MTSS (Multiple Transform Set Selection).

[0124]

[0125] Both conventional NSPT and LFNST could only be used for intra-prediction blocks. Hereinafter, this specification provides a transform set and a method for determining a transform kernel for an inter-prediction block. In addition, a transform set and a method for determining a transform kernel for a block that performs special intra-prediction, such as intra-template matching prediction (IntraTMP), intra-block copy (IBC), spatial geometric partition mode (SGPM), and spatial combined intra-intra prediction (SCIIP), are provided.

[0126]

[0127] Hereinafter, a method for determining a transformation set according to the present invention will be described. The method for determining a transformation set refers to a method for determining a transformation set for transformation / inverse transformation of a current block among a plurality of transformation sets.

[0128] For example, a method for determining a transformation set may refer to a method for determining a transformation set for transformation / inverse transformation of a current block among three transformation sets. At this time, the first transformation set (hereinafter, the first transformation set) and the second transformation set (hereinafter, the second transformation set) may be transformation sets used for transformation / inverse transformation of a block on which intra prediction is performed, and the third transformation set (hereinafter, the third transformation set) may be a transformation set used for transformation / inverse transformation of a block on which inter prediction is performed. In addition, a transformation set for transformation / inverse transformation of the current block among the three transformation sets described above may be determined.

[0129] Meanwhile, the first transformation set may be a transformation set used for transformation / inverse transformation of a conventional intra prediction block. Here, the conventional intra prediction block may mean a prediction block whose intra prediction mode is either a non-directional prediction mode such as a DC mode or a Planar mode, or a directional prediction mode (e.g., 65 directions).

[0130] Meanwhile, the second transformation set may be a transformation set used for transformation / inverse transformation of a block on which intra prediction is performed using any one of SGPM (spatial geometric prediction mode) technology, TIMD (Template-based Intra Mode Derivation) technology, DIMD (Decoder Side Intra Mode Derivation) technology, EIP (Extrapolation filter-based Intra Prediction) technology, and MIP (Matrix-based intra prediction) technology. Here, SGPM may refer to a technology in which intra prediction is performed by spatially dividing the current block into at least two or more regions. TIMD may refer to a technology in which the prediction mode of the current block is derived using a reference region of a template. In addition, DIMD may refer to a technology in which the prediction mode of the current block is derived in the decoder without parsing, EIP may refer to a technology in which prediction is performed using an extrapolation filter, and MIP may refer to a technology in which prediction is performed based on a matrix.

[0131] Meanwhile, the aforementioned transformation set may be a transformation set of a nonseparable peripheral transformation (NSPT). Furthermore, the aforementioned transformation set may be a transformation set of a nonseparable second-order transformation (LFNST).

[0132] Meanwhile, the method for determining a transformation set according to the present invention can be applied even when the prediction mode of the current block is not an intra prediction mode.

[0133] Meanwhile, according to the above-described embodiment, there are three transformation sets, but this is just one example, and the number of transformation sets may be S. Here, S is an integer greater than or equal to 2.

[0134] Meanwhile, according to the above-described embodiment, the first transformation set and the second transformation set are transformation sets used for inverse transformation of the current block on which intra prediction is performed, and the third transformation set is a transformation set used for inverse transformation of the current block on which inter prediction is performed. However, this is only one example, and any transformation set may correspond to inverse transformation of the current block on which prediction is performed in any way.

[0135]

[0136]

[0137] According to one embodiment of the present invention, the transformation set may be determined based on the prediction mode of the current block. For example, if the prediction mode of the current block is an inter prediction mode, the transformation set of the current block may be determined as the third transformation set among the first transformation set, the second transformation set, and the third transformation set described above. In this way, the present invention allows non-separable transformation to be performed on blocks other than intra-prediction blocks.

[0138] Figures 4 and 5 are diagrams for explaining how a transformation set is determined based on the prediction mode of the current block.

[0139] FIG. 4 is a diagram for explaining a method for determining a transformation set when the prediction mode of the current block is an intra template matching mode according to one embodiment of the present invention.

[0140] The intra-template matching prediction method uses template matching to find the optimal prediction block in the reconstructed area of ​​the current picture and copy it to generate the prediction block of the current block. If the prediction mode of the current block is intra-template matching mode, the prediction block of the current block is generated based on the intra-template matching prediction method.

[0141] Referring to FIG. 4, the neighboring L-shaped areas (i.e., the left, top, and upper left areas) of the current block (Current block, 410) can be defined as the current template (Current template, 420). Then, a reference template (Reference template, 440) most similar to the current template (420) can be searched within a predefined search range (R1, R2, R3, R4) of the reconstructed area (Reconstructed area, 430) of the current picture (Current picture, 400). Then, the prediction block of the current block (410) can be derived based on the corresponding matching block (Matching block, 450) of the determined reference template (440).

[0142] In Fig. 4, the predefined search ranges R1, R2, R3, and R4 can be defined as the current CTU (Coding Tree Unit) including the current block, the upper left CTU, the upper CTU, and the left CTU, respectively. Here, the upper left CTU may refer to a CTU located at the upper left of the current CTU within a pre-set area from the current CTU, the upper CTU may refer to a CTU located above the current CTU within a pre-set area from the current CTU, and the left CTU may refer to a CTU located to the left of the current CTU within a pre-set area from the current CTU. The pre-set area may be any one of the current tile, the current slice, and the current picture including the current block.

[0143] Additionally, the predefined search range may be searched for reference templates based on a predefined search order. For example, the reference templates may be searched in a zigzag order of R1, R4, R3, R2.

[0144] Meanwhile, information regarding the search range and the size and shape of the current template can be determined by the encoder and transmitted to the decoder. Furthermore, the search range and the size and shape of the current template can be set to preset values ​​in the encoder / decoder.

[0145] For example, instead of an L-shaped template that uses both the left restoration area and the top restoration area, you can use only the top restoration area or the left restoration area as a template.

[0146] For example, the size of the current template can be determined as (w x L2) + (L1 xh) + (L1 x L2) as shown in Fig. 4. Here, w and h represent the width and height of the current block, and the values ​​of L1 and L2 can be determined as any positive integer.

[0147] With respect to the template shape, the example of FIG. 4 uses one L-shaped template to search for a reference template most similar to the current template, but this can also be done using, for example, an L-shaped template, a template using only the top region, a template using only the left region, or a template of any shape to search for a reference template most similar to the current template.

[0148] Meanwhile, if the prediction mode of the current block is the intra template matching mode, the transformation / inverse transformation can be performed in the same manner as the transformation / inverse transformation of the intra prediction block. That is, the transformation set of the intra block can be used even when the prediction mode of the current block is the intra template matching mode. For example, if the prediction mode of the current block is the intra template matching mode, non-separable transformation / inverse transformation can be performed using NSPT and LFNST. As another example, if the prediction mode of the current block is the intra template matching mode, the transformation set can be determined as the first transformation set described above.

[0149] Alternatively, for prediction blocks in intra-template matching mode, the characteristics of the residuals and the characteristics of the primary transform coefficients generated by the surrounding transformation may be similar to those of prediction blocks in inter-prediction mode due to the high prediction accuracy. Therefore, the transformation set of the inter-block may also be used for blocks in intra-template matching mode. For example, if the prediction mode of the current block is intra-template matching mode, the transformation set of the current block may be determined as the third transformation set described above.

[0150] Alternatively, if the prediction mode of the current block is the intra template matching mode, the transformation set may be determined as either the transformation set of the intra block (hereinafter, the intra transformation set) or the transformation set of the inter block (hereinafter, the inter transformation set). Specifically, the transformation set of the current block may be determined as the transformation set with a smaller cost value among the intra transformation set and the inter transformation set.

[0151] For example, in an encoder, a rate-distortion cost (RD cost) is calculated for each transformation set, so that a transformation set with small distortion can be selected as the optimal transformation set.

[0152] In addition to the information typically signaled in a transformation, information about the optimal transformation set can be signaled to the decoder. Here, the information about the optimal transformation set can be either a flag indicating whether to use the intra transformation set or a flag indicating whether to use the inter transformation set.

[0153] And, in the decoder, the transformation set of the current block can be determined based on the information about the optimal transformation set. For example, if the information about the optimal transformation set is a flag regarding whether an intra transformation set is used and the information indicates that an intra transformation set is used, the transformation set of the current block can be determined as an intra transformation set. On the other hand, if the information does not indicate that an intra transformation set is used, the transformation set of the current block can be determined as an inter transformation set. Similarly, if the information about the optimal transformation set is a flag regarding whether an inter transformation set is used and the information indicates that an inter transformation set is used, the transformation set of the current block can be determined as an inter transformation set. However, if the information does not indicate that an inter transformation set is used, the transformation set of the current block can be determined as an intra transformation set.

[0154] Meanwhile, information about the optimal transformation set can be signaled as a one-bit flag at any one of the block level, slice level, picture level, and sequence level.

[0155]

[0156] FIG. 5 is a diagram for explaining a method for determining a transformation set when the prediction mode of the current block is an intra block copy mode according to one embodiment of the present invention.

[0157] The intra-block copy method refers to a method of searching for an optimal prediction block in a reconstructed area of ​​the current picture using a block vector and copying it to generate a prediction block for the current block. If the prediction mode of the current block is intra-block copy mode, the prediction block for the current block is generated based on the intra-block copy method.

[0158] Referring to FIG. 5, a matching block (Matching Block, 540) can be derived within a predefined search range (R1, R2, R3, R4) of a reconstructed area (530) of a current picture (500) based on a block vector (Block vector, 520) of a current block (510). In addition, a prediction block of the current block (510) can be derived based on the matching block (540). Meanwhile, the block vector (520) is a vector indicating the matching block (540) of the current block (510) in an intra block copy.

[0159] In Fig. 5, the predefined search ranges R1, R2, R3, and R4 can be defined as the current CTU (Coding Tree Unit) including the current block, the upper left CTU, the upper CTU, and the left CTU, respectively. Here, the upper left CTU may refer to a CTU located at the upper left of the current CTU within a pre-set area from the current CTU, the upper CTU may refer to a CTU located above the current CTU within a pre-set area from the current CTU, and the left CTU may refer to a CTU located to the left of the current CTU within a pre-set area from the current CTU. The pre-set area may be any one of the current tile, the current slice, and the current picture including the current block.

[0160] Additionally, matching blocks may be searched based on a predefined search order within a predefined search range. For example, matching blocks may be searched in a zigzag order of R1, R4, R3, R2.

[0161] Meanwhile, information about the search range can be determined by the encoder and transmitted to the decoder. Furthermore, the search range can be set to a preset value in the encoder / decoder.

[0162] Meanwhile, if the prediction mode of the current block is an intra block copy mode, the transformation / inverse transformation can be performed in the same manner as the transformation / inverse transformation of the intra prediction block. That is, the transformation set of the intra block can be used even if the prediction mode of the current block is an intra block copy mode. For example, if the prediction mode of the current block is an intra block copy mode, non-separable transformation / inverse transformation can be performed using NSPT and LFNST. As another example, if the prediction mode of the current block is an intra block copy mode, the transformation set can be determined as the first transformation set described above.

[0163] Alternatively, for a prediction block in intra-block copy mode, the characteristics of the residual and the characteristics of the primary transform coefficients generated by the surrounding transformation may be similar to those of a block in inter-prediction mode due to the high prediction accuracy. Therefore, the transformation set of the inter-block can also be used for a block in intra-block copy mode. For example, if the prediction mode of the current block is intra-block copy mode, the transformation set of the current block can be determined as the third transformation set described above.

[0164] Alternatively, if the prediction mode of the current block is an intra block copy mode, the transformation set may be determined as either a transformation set of the intra block (hereinafter, referred to as an intra transformation set) or a transformation set of the inter block (hereinafter, referred to as an inter transformation set). Specifically, the transformation set of the current block may be determined as a transformation set with a smaller cost value among the intra transformation set and the inter transformation set.

[0165] For example, the cost value can be calculated in the encoder. In the encoder, a rate-distortion cost (RD cost) is calculated for each transform set, so that a transform set with small distortion can be selected as the optimal transform set.

[0166] In addition to the information typically signaled in a transformation, information about the optimal transformation set can be signaled to the decoder. Here, the information about the optimal transformation set can be either a flag indicating whether to use the intra transformation set or a flag indicating whether to use the inter transformation set.

[0167] And, in the decoder, the transformation set of the current block can be determined based on the information about the optimal transformation set. For example, if the information about the optimal transformation set is a flag regarding whether an intra transformation set is used and the information indicates that an intra transformation set is used, the transformation set of the current block can be determined as an intra transformation set. On the other hand, if the information does not indicate that an intra transformation set is used, the transformation set of the current block can be determined as an inter transformation set. Similarly, if the information about the optimal transformation set is a flag regarding whether an inter transformation set is used and the information indicates that an inter transformation set is used, the transformation set of the current block can be determined as an inter transformation set. However, if the information does not indicate that an inter transformation set is used, the transformation set of the current block can be determined as an intra transformation set.

[0168] Meanwhile, information about the optimal transformation set can be signaled as a one-bit flag at any one of the block level, slice level, picture level, and sequence level.

[0169] Meanwhile, in the aforementioned embodiment, the transformation set of the current block can be determined based on the prediction mode of the current block without separate signaling. For example, if the prediction mode of the current block is any one of inter prediction mode, intra template matching mode, and intra block copy mode, the transformation set of the current block in the decoder can be determined as the third transformation set described above.

[0170]

[0171] According to one embodiment of the present invention, the transform set of the current block may be determined in a manner other than being determined based on the prediction mode of the current block.

[0172] For example, the transformation set may be determined based on directional information of pixels contained in the entire region or a portion of the prediction block of the current block. Here, the directional information may refer to information obtained by calculating the gradient of the corresponding pixels.

[0173] As another example, the transformation set can be determined based on the directional information of the pixels contained in the entire region or a portion of the prediction block of the current block and the directional information of the subpixels obtained by applying an interpolation filter to the region.

[0174] As another example, the transformation set can be determined based on the directional information of pixels contained in the entire region or a portion of one of the templates neighboring the current block or the template corresponding to the current block.

[0175] As another example, the transformation set can be determined based on the directional information of the pixels contained in the entire region or a portion of the region, either of the templates neighboring the current block or the templates corresponding to the current block, and the directional information of the subpixels obtained by applying an interpolation filter to the region.

[0176] Meanwhile, the aforementioned directional information may refer to information obtained by calculating the slope of the corresponding pixels. Specifically, the directional information may refer to information obtained by accumulating or processing slope information calculated by applying a filter to the corresponding pixels. The filter applied to the corresponding pixels may be any one of a Sobel filter, a Roberts cross filter, a Prewitt filter, a Scharr filter, and a Laplacian filter.

[0177] Meanwhile, the template neighboring the current block mentioned above may mean an arbitrary area neighboring the current block, and the template corresponding to the current block may mean an arbitrary area having the same size as the template neighboring the current block.

[0178] As another example, if the prediction mode of the current block is an inter prediction mode, the transformation set may be determined based on the magnitude of at least one of the x-direction motion vector and the y-direction motion vector of the current block. Alternatively, if the prediction mode of the current block is an inter prediction mode, the transformation set may be determined based on a processed value of the magnitude of at least one of the x-direction motion vector and the y-direction motion vector of the current block.

[0179] As another example, if the prediction mode of the current block is the inter prediction mode, the transformation set can be determined based on the precision of at least one of the x-direction motion vector and the y-direction motion vector of the current block. Specifically, the transformation set can be determined based on whether the precision of the corresponding motion vector is integer-pel precision or sub-pel precision.

[0180] It can be determined based on whether the (fractional-pel precision) is correct.

[0181] As another example, the transformation set may be determined based on at least one of the width and height of the current block.

[0182] As another example, the transformation set can be determined based on the number of pixels contained in the current block.

[0183] As another example, the transformation set can be determined based on information about the partitioning of the current block.

[0184] As another example, a set of transformations can be determined based on explicitly signaled syntax elements. Specifically, once a set of transformations is determined at the encoder, the syntax for the determined set of transformations can be signaled to the decoder. The decoder can then determine a set of transformations for inverse transformation based on the corresponding syntax elements.

[0185] As another example, the transformation set may be determined based on the values ​​of surrounding pixels of the current block. Here, the surrounding pixels may refer to all or some of the pixels contained in a block adjacent to the current block, or all or some of the pixels contained in a block corresponding to the current block.

[0186] As another example, the transform set can be determined based on the quantization coefficients.

[0187] Meanwhile, using the aforementioned methods, a transformation set for the transformation of the current block can be determined among S transformation sets. Here, S is an integer greater than or equal to 2.

[0188] Meanwhile, the above-described transformation set determination method can be performed even when the prediction mode of the current block is an inter prediction mode.

[0189]

[0190] According to one embodiment of the present invention, the transformation set of the current block can be determined by deriving an intra-prediction mode for determining the transformation set. Specifically, the intra-prediction mode for determining the transformation set is derived based on directionality information of pixels included in a prediction block of the current block, and the transformation set can be determined using the intra-prediction mode. Meanwhile, the directionality information may refer to information obtained by calculating the gradient of pixels.

[0191] First, the gradient of pixels included in the prediction block of the current block can be calculated. At this time, the gradient can be calculated by applying at least one of a Sobel filter, a Roberts cross filter, a Prewitt filter, a Schar filter, and a Laplacian filter to the pixels included in the prediction block.

[0192] Meanwhile, the gradients of only some of the pixels included in the prediction block of the current block may be calculated. For example, if a sub-block transform (SBT) is applied to the current block to split the current block and the transformation is performed on only one of the blocks, the gradients of only the pixels included in the region of the prediction block corresponding to the block on which the transformation is performed may be calculated. As another example, after downsampling is performed on the pixels included in the prediction block of the current block, only the gradients of the downsampled pixels may be calculated.

[0193] And, a histogram of gradient (HoG) can be generated based on the calculated gradient.

[0194] And, an intra prediction mode for determining a transformation set can be derived using the gradient histogram.

[0195] For example, the intra prediction mode for determining the transformation set may be the intra prediction mode corresponding to the most accumulated direction in the gradient histogram.

[0196] As another example, the intra prediction mode for determining the transformation set may be the prediction mode with the smallest assigned prediction mode number among the N prediction modes corresponding to the most accumulated direction in the gradient histogram, where N is any positive integer.

[0197] As another example, N intra prediction modes corresponding to the most accumulated directionality in the gradient histogram can be derived, and an index for determining a transformation set can be derived using these indices. Then, the transformation set for the current block can be determined based on the indices. Here, N is an arbitrary positive integer.

[0198] As another example, if the accumulated directionality in the gradient histogram is lower than a predetermined threshold, the intra prediction mode for determining the transformation set can be derived as a predetermined intra prediction mode. Here, the predetermined intra prediction mode can be any one of a DC mode, a vertical DC mode, a horizontal DC mode, a planar mode, a vertical planar mode, and a horizontal planar mode.

[0199] Meanwhile, a predetermined threshold value may be defined in each encoder and decoder. Alternatively, the predetermined threshold value may be defined in the encoder and signaled to the decoder. Alternatively, the predetermined threshold value may be defined by an agreement between the encoder and decoder.

[0200] And, a transformation set can be determined based on the derived intra prediction mode. Specifically, the transformation set is determined as a transformation set that maps to the derived intra prediction mode among multiple transformation sets. Here, the mapping between the intra prediction mode and the transformation set can be performed in the same manner as the transformation of the intra block. For example, it can be performed in a manner similar to LFNST and NSPT.

[0201] Table 1 shows examples of how the 95 intra prediction modes in LFNST and NPST are mapped to the transform set.

[0202]

[0203]

[0204] Meanwhile, in the aforementioned embodiment, when accumulating the directionality in the slope histogram, a method of accumulating in a single bin or a method of accumulating in multiple bins may be used. In this case, the method of accumulating in multiple bins may mean a method of accumulating in a single bin corresponding to the directionality of the slope calculated through the filter and M bins surrounding it. M is an integer greater than or equal to 1.

[0205] Meanwhile, in the above-described embodiment, the directional information is obtained by calculating the gradient of pixels included in the prediction block of the current block, but this is only one example, and the directional information can be obtained by calculating the gradient of the subpixels obtained by applying an interpolation filter to the entire area or a part of the prediction block of the current block and the corresponding area.

[0206] Meanwhile, in the above-described embodiment, the directional information is obtained by calculating the gradient of pixels included in the prediction block of the current block, but this is just one example, and the directional information can be obtained by calculating the gradient of pixels included in either the template neighboring the current block or the template corresponding to the current block. Alternatively, the directional information can be obtained by calculating the gradient of an integer of the entire region or a part of the region of either the template neighboring the current block or the template corresponding to the current block, and a subpixel obtained by applying an interpolation filter to the corresponding region. Here, the template neighboring the current block can mean any region neighboring the current block, and the template corresponding to the current block can mean any region with the same size as the template neighboring the current block.

[0207]

[0208] Hereinafter, a method for determining a transformation kernel for transformation / inverse transformation of a current block among the transformation kernels included in a transformation set according to the present invention will be described.

[0209] Once a transformation set is determined, transformation / inverse transformation of the current block can be performed based on the transformation set. Specifically, a transformation kernel for transformation / inverse transformation of the current block can be determined from among K transformation kernels included in the transformation set, and transformation / inverse transformation of the current block can be performed based on the transformation kernel. Here, K is an integer greater than or equal to 2.

[0210] Meanwhile, if the transformation set includes one transformation kernel, the transformation kernel of the current block can be determined by that transformation kernel.

[0211]

[0212] According to one embodiment of the present invention, if the prediction mode of the current block is an inter prediction mode, the transformation kernel of the current block may be determined as a special transformation kernel predefined in the encoder / decoder. In this case, the encoder / decoder may check whether the prediction mode of the current block is an inter prediction mode, and if it is an inter prediction mode, perform transformation / inverse transformation based on the predefined transformation kernel.

[0213]

[0214] According to one embodiment of the present invention, a transformation kernel may be determined based on the directionality of pixels included in a prediction block of a current block. Specifically, an intra-prediction mode for determining a transformation kernel may be derived based on directionality information of pixels included in a prediction block of the current block, and the transformation kernel may be determined using the intra-prediction mode. Meanwhile, the directionality information may refer to information obtained by calculating the gradient of pixels.

[0215] First, the gradient of pixels included in the prediction block of the current block can be calculated. At this time, the gradient can be calculated by applying at least one of a Sobel filter, a Roberts cross filter, a Prewitt filter, a Schar filter, and a Laplacian filter to the pixels included in the prediction block.

[0216] Meanwhile, the gradients of only some of the pixels included in the prediction block of the current block may be calculated. For example, if a sub-block transform (SBT) is applied to the current block to split the current block, and the transformation is performed on only one of the blocks, the gradients of only the pixels included in the region of the prediction block corresponding to the block on which the transformation is performed may be calculated. As another example, after downsampling is performed on the pixels included in the prediction block of the current block, only the gradients of the downsampled pixels may be calculated.

[0217] And, a histogram of gradient (HoG) can be generated based on the calculated gradient.

[0218] And, an intra prediction mode for determining a transformation kernel can be derived using a gradient histogram. Here, the intra prediction mode for determining a transformation kernel can have the same meaning as the transformation kernel determination intra prediction mode.

[0219]

[0220] According to one embodiment of the present invention, the transformation kernel of the current block can be determined in the following manner.

[0221] For example, the transformation kernel may be determined based on directional information of pixels contained in the entire region or a portion of the prediction block of the current block. Here, the directional information may refer to information obtained by calculating the gradient of the corresponding pixels.

[0222] As another example, the transformation kernel can be determined based on the directional information of the pixels contained in the entire region or a portion of the prediction block of the current block and the directional information of the subpixels obtained by applying an interpolation filter to the corresponding region.

[0223] As another example, the transformation kernel may be determined based on directional information of pixels contained in the entire region or a portion of one of the templates neighboring the current block or the template corresponding to the current block.

[0224] As another example, the transformation kernel can be determined based on the directional information of the integers contained in the entire region or a portion of one of the templates neighboring the current block or the template corresponding to the current block, and the directional information of the subpixels obtained by applying an interpolation filter to the corresponding region.

[0225] Meanwhile, the aforementioned directional information may refer to information obtained by calculating the slope of the corresponding pixels. Specifically, the directional information may refer to information obtained by accumulating or processing slope information calculated by applying a filter to the corresponding pixels. The filter applied to the corresponding pixels may be any one of a Sobel filter, a Roberts cross filter, a Prewitt filter, a Scharr filter, and a Laplacian filter.

[0226] Meanwhile, the template neighboring the current block mentioned above may mean an arbitrary area neighboring the current block, and the template corresponding to the current block may mean an arbitrary area having the same size as the template neighboring the current block.

[0227] As another example, if the prediction mode of the current block is the inter prediction mode, the transformation kernel may be determined based on the magnitude of at least one of the x-direction motion vector and the y-direction motion vector of the current block. Alternatively, if the prediction mode of the current block is the inter prediction mode, the transformation kernel may be determined based on a processed value of the magnitude of at least one of the x-direction motion vector and the y-direction motion vector of the current block.

[0228] As another example, if the prediction mode of the current block is the inter prediction mode, the transformation kernel can be determined based on the precision of at least one of the x-direction motion vector and the y-direction motion vector of the current block. Specifically, the transformation kernel determines whether the precision of the corresponding motion vector is integer-pel precision or sub-pel precision.

[0229] It can be determined based on whether the (fractional-pel precision) is correct.

[0230] As another example, the transformation kernel may be determined based on at least one of the width and height of the current block.

[0231] As another example, the transformation kernel may be determined based on the number of pixels contained in the current block.

[0232] As another example, the transformation kernel can be determined based on information about the partitioning of the current block.

[0233] As another example, the transformation kernel may be determined based on the values ​​of surrounding pixels of the current block. Here, the surrounding pixels may refer to all or some of the pixels contained in a block adjacent to the current block, or all or some of the pixels contained in a block corresponding to the current block.

[0234] As another example, the transform kernel can be determined based on the quantization coefficients.

[0235] Meanwhile, using the aforementioned methods, a transformation kernel for transformation of the current block can be determined among K transformation kernels. Here, K is an integer greater than or equal to 2.

[0236] Meanwhile, the above-described transformation kernel determination method can be performed even when the prediction mode of the current block is an inter prediction mode.

[0237] Meanwhile, the aforementioned transform kernel determination method can be performed in the encoder / decoder. After performing the transformation based on the transform kernel determined in the encoder, signaling for the transform kernel is not required. The decoder also performs the inverse transformation based on the determined transform kernel.

[0238] Alternatively, the transform kernel can be determined based on an explicitly signaled syntax element. Specifically, once the transform kernel is determined at the encoder, the syntax for the determined set of transforms can be signaled to the decoder. The decoder can then determine the transform kernel for the inverse transform based on the corresponding syntax element.

[0239] Meanwhile, the above-described transformation kernel determination method can be performed even when the prediction mode of the current block is an inter prediction mode.

[0240]

[0241] According to one embodiment of the present invention, the transformation kernel may be determined as the transformation kernel having the smallest cost value among K transformation kernels included in the transformation set, where K is an integer greater than or equal to 2.

[0242] Specifically, a rate-distortion cost (RD cost) is calculated for each transformation kernel in the encoder, and the kernel with the smallest distortion can be selected as the optimal transformation kernel.

[0243] And, information about the optimal transformation kernel can be signaled to the decoder. Here, the information about the optimal transformation kernel can be a flag when K is 2, or a kernel index when K is 3 or greater. The decoder determines the transformation kernel based on the information, and inverse transformation can be performed based on the determined transformation kernel.

[0244] Meanwhile, according to the above-described embodiments, the transformation kernel is determined after the transformation set is determined. However, this is only an example. If a separate transformation set does not exist, the transformation kernel of the current block may be determined without determining the transformation set. In this case, the method for determining the transformation kernel may be performed in the same manner as the method for determining the transformation kernel according to the above-described embodiments.

[0245] Meanwhile, according to the above-described embodiments, the transformation kernel is determined after the transformation set is determined, but the determination of the transformation set and the transformation kernel can be performed simultaneously.

[0246]

[0247] FIG. 6 is a diagram for explaining a method for determining a transformation set and a transformation kernel in a spatial geometric partitioning mode (SGPM) according to an embodiment of the present invention. Here, the spatial geometric partitioning mode may refer to a mode in which a block is divided into two regions by a straight-line partitioning boundary and intra prediction is performed. Specifically, according to the spatial geometric partitioning mode, each divided region is independently intra-predicted to generate an intra-prediction block of each region, and these can be weighted and added to generate an intra-prediction block of the current block.

[0248] Referring to Fig. 6, when the prediction mode of the current block (Current block, 600) is the spatial geometric partitioning mode, it is divided into two regions based on the geometric partitioning boundary (Geometric partitioning line, 601). Here, each of the divided regions may be referred to as sub-region 0 (Sub partition 0, 602) and sub-region 1 (Sub partition 1, 603). Prediction can be performed independently for each sub-region based on each intra prediction mode. Specifically, a prediction block can be generated based on the intra prediction mode of sub-region 0, and a prediction block can be generated based on the intra prediction mode of sub-region 1. Then, each prediction block can be weighted and combined to generate a final intra prediction block of the current block.

[0249] According to one embodiment of the present invention, when the prediction mode of the current block is a spatial geometry segmentation mode, a transformation set and kernel for transformation of the current block can be determined based on the intra prediction mode of each sub-region.

[0250] For example, among the intra prediction mode of sub-region 0 and the intra prediction mode of sub-region 1, an intra prediction mode with a smaller assigned prediction mode number may be used to determine at least one of the transformation set and transformation kernel of the current block.

[0251] As another example, an intra prediction mode with a lower cost value among the intra prediction mode of sub-region 0 and the intra prediction mode of sub-region 1 may be used to determine at least one of the transform set and transform kernel of the current block. Specifically, a bit-rate-distortion cost may be calculated for the performance result of each intra prediction mode, so that an intra prediction mode with a lower distortion may be used.

[0252] As another example, a cost value may be calculated by performing prediction using the intra prediction mode of sub-region 0 and the intra prediction mode of sub-region 1 for either a template neighboring the current block or a template corresponding to the current block. Specifically, a bit-rate-distortion cost may be calculated for the performance result of each intra prediction mode. In addition, an intra prediction mode with small distortion may be used to determine at least one of the transform set and the transform kernel of the current block. Here, the neighboring template may mean an arbitrary region neighboring the current block, and the template corresponding to the current block may mean an arbitrary region having the same size as the template neighboring the current block.

[0253] Meanwhile, in the aforementioned example, the method of determining the transformation set using the intra prediction mode can be performed in the same manner as the mapping in the transformation of the intra block. For example, it can be performed in the same manner as LFNST and NSPT.

[0254] Meanwhile, when the intra prediction mode is used in the above-described example to determine the transformation kernel, the transformation kernel may be mapped to the intra prediction mode through any method. Alternatively, an arbitrary lookup table for the transformation kernel mapped to the intra prediction mode may be defined.

[0255] Meanwhile, in spatial geometry partitioning mode, sub-regions and sub-blocks can have the same meaning.

[0256]

[0257] According to one embodiment of the present invention, when the prediction mode of the current block is a spatial geometry division mode, one of the above-described transformation set determination method and transformation kernel determination method may be performed instead of determining the transformation set and transformation kernel based on the intra prediction mode of each sub-region.

[0258] For example, at least one of the transformation set and the transformation kernel may be determined based on directionality information of pixels included in the entire region or a portion of one of the templates neighboring the current block or the template corresponding to the current block.

[0259] As another example, for a prediction block in the spatial geometry partitioning mode, the characteristics of the residuals and the characteristics of the primary transform coefficients generated by the main transformation may be similar to those of a block in the inter prediction mode due to the high prediction accuracy. Therefore, the transformation set for the transformation / inverse transformation of the inter block may also be used for a block in the spatial geometry partitioning mode. For example, if the prediction mode of the current block is the spatial geometry mode, the transformation set of the current block may be determined as the third transformation set described above.

[0260] As another example, if the prediction mode of the current block is the spatial geometric partitioning mode, the transformation / inverse transformation can be performed in the same manner as the transformation / inverse transformation of the intra block. That is, the transformation set of the intra block can also be used for a block with the spatial geometric partitioning mode. If the prediction mode of the current block is the spatial geometric partitioning mode, the transformation set of the current block can be determined as the second transformation set described above. Alternatively, it can be determined as the first transformation set described above.

[0261] As another example, if the prediction mode of the current block is geometric partitioning mode, the transformation set of the current block may be determined as either the transformation set of the intra block (intra transformation set) or the transformation set of the inter block (inter transformation set). Specifically, the transformation set of the current block may be determined as the transformation set with a smaller cost value among the intra transformation set and the inter transformation set. Even in this case, in addition to the information typically signaled for the transformation, information about the optimal transformation set is signaled to the decoder.

[0262]

[0263] Hereinafter, a method for determining a transformation set and a transformation kernel in a spatial combined intra-intra prediction (SCIIP) mode according to one embodiment of the present invention will be described. Here, the spatial combined intra-intra prediction mode may refer to a mode in which intra prediction is performed based on two different intra prediction modes.

[0264] Specifically, according to the spatial intra-picture combined prediction mode, two prediction blocks can be generated based on different intra prediction modes. Then, the generated prediction blocks can be weighted and combined to generate a final intra prediction block of the current block. Here, the different intra prediction modes can be referred to as the first intra prediction mode and the second intra prediction mode, respectively.

[0265] According to one embodiment, when the prediction mode of the current block is a spatial intra-intra-picture combined prediction mode, a transform set and kernel for transforming the current block can be determined based on the first intra prediction mode and the second intra prediction mode.

[0266] For example, among the first intra prediction mode and the second intra prediction mode, an intra prediction mode with a smaller assigned prediction mode number may be used to determine at least one of the transformation set and transformation kernel of the current block.

[0267] As another example, an intra prediction mode having a lower cost value among the first intra prediction mode and the second intra prediction mode may be used to determine at least one of the transform set and transform kernel of the current block.

[0268] Meanwhile, the method for determining the transformation set using the intra prediction mode can be performed in the same manner as the mapping in the transformation of the intra block. For example, the mapping can be performed in the same manner as the mapping of the intra prediction mode and the transformation set in LFNST and NSPT.

[0269] Meanwhile, the intra prediction mode used when determining the transformation kernel based on the first intra prediction mode and the second intra prediction mode may be mapped to the transformation kernel through any method. Alternatively, an arbitrary lookup table for the transformation kernel mapped to the intra prediction mode used may be defined.

[0270]

[0271] According to one embodiment of the present invention, even if the prediction mode of the current block is a spatial intra-picture combined prediction mode, one of the above-described transformation set determination method and transformation kernel determination method may be performed instead of determining the transformation set and transformation kernel based on the intra prediction mode of each sub-region.

[0272] For example, at least one of the transformation set and the transformation kernel may be determined based on directionality information of pixels included in the entire region or a portion of one of the templates neighboring the current block or the template corresponding to the current block.

[0273] As another example, for a prediction block in a spatial intra-picture joint prediction mode, the characteristics of the residual and the characteristics of the primary transform coefficients generated by the main transformation may be similar to those of a prediction block in an inter-prediction mode due to the high prediction accuracy. Therefore, the transformation set of the inter-block may also be used for a block in the current spatial intra-picture joint prediction mode. For example, if the prediction mode of the current block is a spatial intra-picture joint prediction mode, the transformation set of the current block may be determined as the third transformation set described above.

[0274] As another example, if the prediction mode of the current block is a spatial intra-picture combined prediction mode, the transformation / inverse transformation can be performed in the same manner as the transformation / inverse transformation of the intra block. That is, the transformation set of the intra block can also be used for a block that is a spatial intra-picture combined prediction mode. If the prediction mode of the current block is a spatial intra-picture combined prediction mode, the transformation set of the current block can be determined as the second transformation set described above. Alternatively, it can be determined as the first transformation set described above.

[0275] As another example, if the prediction mode of the current block is a spatial intra-internal combined prediction mode, the transformation set of the current block may be determined as either a transformation set of an intra block (hereinafter, referred to as an intra transformation set) or a transformation set for transformation / inverse transformation of an inter block (hereinafter, referred to as an inter transformation set). Specifically, the transformation set of the current block may be determined as a transformation set with a smaller cost value among the intra transformation set and the inter transformation set. In this case, in addition to the information generally signaled for the transformation, information about the corresponding optimal transformation set is signaled to the decoder.

[0276]

[0277] Fig. 7 is a flowchart illustrating a method for determining a transformation set according to one embodiment of the present invention. The method for determining a transformation set of Fig. 7 can be performed by an image decoding device.

[0278] The video decoding device can derive the prediction mode of the current block (S700).

[0279] And, the image decoding device can determine the transformation set of the current block based on the prediction mode of the current block (S710).

[0280] Meanwhile, the transformation set may be determined as one of a first transformation set for the first intra prediction, a second transformation set for the second intra prediction, and a third transformation set for the inter prediction.

[0281] Meanwhile, if the prediction mode is an intra template matching mode, the transformation set can be determined as the third transformation set.

[0282] Meanwhile, the above transformation set may be a transformation set of non-separable surroundings.

[0283] Meanwhile, if the prediction mode is an intra block copy mode, the transformation set can be determined as the third transformation set.

[0284] Meanwhile, the first intra prediction may include non-directional intra prediction and directional intra prediction, and the second intra prediction may include SGPM (Spatial geometric prediction mode) intra prediction.

[0285] Meanwhile, if the current block is divided into a first sub-block and a second sub-block according to a division boundary, the transformation set can be determined as the second transformation set.

[0286] And the image decoding device can perform reverse transformation on the current block based on the transformation set (S720).

[0287] Meanwhile, the step of performing the inverse transformation may include the step of determining a transformation kernel of the current block among the transformation kernels included in the transformation set and the step of performing the inverse transformation based on the transformation kernel.

[0288] Meanwhile, the step of determining the transformation kernel may include the step of deriving a transformation kernel determination intra prediction mode and the step of determining the transformation kernel based on the transformation kernel determination intra prediction mode.

[0289] Meanwhile, the intra prediction mode of the transformation kernel decision can be determined using direction information of some pixels included in the prediction block of the current block.

[0290] Meanwhile, the above transformation kernel may be a transformation kernel with the smallest distortion among the transformation kernels included in the above transformation set.

[0291] Meanwhile, if the prediction mode of the current block is one of the intra template matching mode and the intra block copy mode, the method further includes a step of obtaining intra prediction mode usage information indicating whether the intra prediction mode is used when determining the transformation kernel, and if the intra prediction mode usage information indicates that the intra prediction mode is used, the step of determining the transformation kernel may include a step of deriving a transformation kernel determined intra prediction mode and a step of determining the transformation kernel based on the transformation kernel determined intra prediction mode.

[0292] Meanwhile, when the current block is divided into a first sub-block and a second sub-block according to a division boundary, the transformation kernel can be determined based on a prediction mode of the first sub-block and a prediction mode of the second sub-block.

[0293] Meanwhile, in a case where the current block is divided into a first sub-block and a second sub-block according to a division boundary, the method further includes a step of obtaining sub-block prediction mode usage information indicating whether the prediction mode of the first sub-block and the prediction mode of the second sub-block are used when determining the transformation kernel, and in a case where the sub-block prediction mode usage information indicates that the prediction mode of the first sub-block and the prediction mode of the second sub-block are used, the transformation kernel may be determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.

[0294] Meanwhile, when the prediction mode of the current block is the SCIIP (Spatial combined intra-intra prediction) mode, the transformation kernel of the current block can be determined based on the first intra prediction mode of the SCIIP mode and the second intra prediction mode of the SCIIP mode.

[0295] Meanwhile, the steps described in FIG. 7 can be performed in the same manner in an image encoding method. Furthermore, a bitstream can be generated by an image encoding method including the steps described in FIG. 7. The bitstream can be stored on a non-transitory computer-readable recording medium and can also be transmitted (or streamed).

[0296]

[0297] FIG. 8 is a drawing exemplarily showing a content streaming system to which an embodiment according to the present invention can be applied.

[0298] As illustrated in FIG. 8, a content streaming system to which an embodiment of the present invention is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0299] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and CCTVs into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and CCTVs directly generate bitstreams, the encoding server may be omitted.

[0300] The above bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0301] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system.

[0302] The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0303] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0304] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0305]

[0306] The above embodiments can be performed in the same or corresponding manner in an encoding device and a decoding device. In addition, an image can be encoded / decoded using at least one or a combination of at least one of the above embodiments.

[0307] The order in which the above embodiments are applied may be different in the encoding device and the decoding device. Alternatively, the order in which the above embodiments are applied may be the same in the encoding device and the decoding device.

[0308] The above embodiments can be performed for each of the luminance and chrominance signals. Alternatively, the above embodiments can be performed identically for the luminance and chrominance signals.

[0309] In the above embodiments, the methods are described based on a flowchart as a series of steps or units. However, the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and that other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the present invention.

[0310] The above embodiments may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be those specifically designed and constructed for the present invention, or may be known and usable by those skilled in the art of computer software.

[0311] The bitstream generated by the encoding method according to the above embodiment can be stored in a non-transitory computer-readable recording medium. In addition, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiment.

[0312] Here, examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.

[0313] Although the present invention has been described above with specific details such as specific components and limited examples and drawings, these are provided only to help a more general understanding of the present invention, and the present invention is not limited to the above examples, and those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from this description.

[0314] Therefore, the idea of ​​the present invention should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the following claims as well as the claims are considered to fall within the scope of the idea of ​​the present invention.

[0315] The present invention can be used in a device for encoding / decoding an image and a recording medium storing a bitstream.

Claims

1. In the video decryption method, A step of deriving a prediction mode of the current block; a step of determining a transformation set of the current block based on the above prediction mode; and Comprising a step of performing inverse transformation on the current block based on the above transformation set, A method for decoding an image, characterized in that the above transformation set is determined as one of a first transformation set for a first intra prediction, a second transformation set for a second intra prediction, and a third transformation set for inter prediction.

2. In paragraph 1, An image decoding method, characterized in that when the above prediction mode is an intra template matching mode, the transformation set is determined as the third transformation set.

3. In paragraph 1, An image decoding method, characterized in that the above transformation set is a transformation set of non-separable surroundings.

4. In paragraph 1, A video decoding method, characterized in that when the above prediction mode is an intra block copy mode, the transformation set is determined as the third transformation set.

5. In paragraph 1, The above first intra prediction includes non-directional intra prediction and directional intra prediction, An image decoding method, characterized in that the second intra prediction includes SGPM (Spatial geometric prediction mode) intra prediction.

6. In paragraph 1, The steps for performing the above inverse transformation are: A step of determining a transformation kernel of the current block among the transformation kernels included in the transformation set; and An image decoding method, characterized by comprising a step of performing inverse transformation based on the above transformation kernel.

7. In paragraph 6, The step of determining the above transformation kernel is: A step of deriving an intra prediction mode for determining a transformation kernel; and An image decoding method, characterized by comprising a step of determining the transform kernel based on the intra prediction mode determined by the transform kernel.

8. In paragraph 7, An image decoding method, characterized in that the intra prediction mode determined by the above transformation kernel is derived based on directionality information of some pixels included in a prediction block of the current block.

9. In paragraph 6, An image decoding method, characterized in that the above transformation kernel is a transformation kernel having the smallest distortion among the transformation kernels included in the above transformation set.

10. In paragraph 6, If the prediction mode of the current block is one of the intra template matching mode and the intra block copy mode, the method further includes a step of obtaining intra prediction mode usage information indicating whether the intra prediction mode is used when determining the transformation kernel. If the intra prediction mode usage information indicates that the intra prediction mode is used, the step of determining the transformation kernel is: A step of deriving an intra prediction mode for determining a transformation kernel; and An image decoding method, characterized by comprising a step of determining the transform kernel based on the intra prediction mode determined by the transform kernel.

11. In paragraph 1, An image decoding method, characterized in that when the current block is divided into a first sub-block and a second sub-block according to a division boundary, the transformation set is determined as the second transformation set.

12. In paragraph 6, An image decoding method, characterized in that when the current block is divided into a first sub-block and a second sub-block according to a division boundary, the transformation kernel is determined based on a prediction mode of the first sub-block and a prediction mode of the second sub-block.

13. In paragraph 6, In a case where the current block is divided into a first sub-block and a second sub-block according to a division boundary, the method further comprises a step of obtaining sub-block prediction mode usage information indicating whether the prediction mode of the first sub-block and the prediction mode of the second sub-block are used when determining the transformation kernel. An image decoding method, characterized in that when the sub-block prediction mode usage information indicates that the prediction mode of the first sub-block and the prediction mode of the second sub-block are used, the transform kernel is determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.

14. In paragraph 6, An image decoding method, characterized in that when the prediction mode of the current block is a SCIIP (Spatial combined intra-intra prediction) mode, the transform kernel of the current block is determined based on a first intra prediction mode of the SCIIP mode and a second intra prediction mode of the SCIIP mode.

15. In a video encoding method, A step for determining the prediction mode of the current block; a step of determining a transformation set of the current block based on the above prediction mode; and Comprising a step of performing a transformation on the current block based on the transformation set, A method for decoding an image, characterized in that the above transformation set is determined as one of a first transformation set for a first intra prediction, a second transformation set for a second intra prediction, and a third transformation set for inter prediction.

16. A non-transitory computer-readable recording medium storing a bitstream generated by a video encoding method, The above image encoding method is, A step for determining the prediction mode of the current block; a step of determining a transformation set of the current block based on the above prediction mode; and Comprising a step of performing a transformation on the current block based on the transformation set, A non-transitory computer-readable recording medium, characterized in that the transform set is determined as one of a first transform set for the first intra prediction, a second transform set for the second intra prediction, and a third transform set for the inter prediction.

17. In a method for transmitting a bitstream generated by a video encoding method, The above transmission method comprises a step of transmitting the bitstream, The above image encoding method is, A step for determining the prediction mode of the current block; a step of determining a transformation set of the current block based on the above prediction mode; and Comprising a step of performing a transformation on the current block based on the transformation set, A transmission method, characterized in that the above transformation set is determined as one of a first transformation set for a first intra prediction, a second transformation set for a second intra prediction, and a third transformation set for inter prediction.

Citation Information

Patent Citations

  • Methods and apparatus for transform selection in video encoding and decoding

    KR1020170118956A

  • 30m class cell-guide steel structure production automation system

    KR102606349B1

  • Military Important Facilities Vigilance Operation System

    KR102633616B1