Image encoding / decoding method and apparatus and recording medium storing bitstream
Patent Information
- Application Number
- CN202480087651.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-06
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]由于传统的不可分离变换仅应用于帧内预测块,并且难以设计变换核和优化变换结构,因此传统的不可分离变换的应用受到限制
根据本发明,可以提供用于对具有提高的编码/解码效率的图像进行编码/解码的方法和装置。
Smart Images

Figure CN122804403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image encoding / decoding methods and apparatus, and recording media for storing bitstreams. Specifically, this invention relates to image encoding / decoding methods and apparatus based on a method for determining a transform set, and recording media for storing bitstreams. Background Technology
[0002] Recently, there has been an increasing demand for high-resolution and high-quality images, such as Ultra High Definition (UHD) images, across various application areas. As image data becomes higher in resolution and quality, the relative data volume increases compared to existing image data. Therefore, transmission and storage costs increase when using existing media such as wired or wireless broadband channels to transmit this image data, or when using existing storage media to store it. To address these issues arising from the increasing resolution and quality of image data, efficient image encoding / decoding technologies are needed for images with higher resolution and image quality.
[0003] Because traditional non-separable transforms are only applied to intra-frame prediction blocks, and it is difficult to design transform kernels and optimize transform structures, the application of traditional non-separable transforms is limited. Summary of the Invention
[0004] Technical issues The object of this invention is to provide a method and apparatus for encoding / decoding images with improved encoding / decoding efficiency.
[0005] Another object of the present invention is to provide a recording medium for storing a bitstream generated by a method or apparatus for decoding images according to the present invention.
[0006] Another object of the present invention is to provide a method for determining the transform set and the transform kernel, thereby solving the above-mentioned problems.
[0007] Technical solution The image decoding method according to an embodiment of the present invention includes: deriving a prediction mode of the current block; determining a transform set of the current block based on the prediction mode; and performing an inverse transform on the current block based on the transform set, wherein the transform set can be determined as any one of a first transform set for intra-frame prediction, a second transform set for intra-frame prediction, and a third transform set for inter-frame prediction.
[0008] In image decoding methods, when the prediction mode is intra-frame template matching mode, the transform set can be determined as the third transform set.
[0009] In image decoding methods, the transform set can be a transform set that is indivisible and can be a single transform.
[0010] In image decoding methods, when the prediction mode is intra-block copy mode, the transform set can be determined as the third transform set.
[0011] In the image decoding method, the first intra-frame prediction may include non-directional intra-frame prediction and directional intra-frame prediction, and the second intra-frame prediction may include spatial geometric prediction mode (SGPM) intra-frame prediction.
[0012] In an image decoding method, performing an inverse transform may include: determining the transform kernel of the current block from transform kernels included in a transform set, and performing an inverse transform based on the transform kernel.
[0013] In image decoding methods, determining the transform kernel can include: deriving the transform kernel to determine the intra-frame prediction mode, and determining the transform kernel based on the intra-frame prediction mode determined by the transform kernel.
[0014] In image decoding methods, the transform kernel can be derived to determine the intra-frame prediction mode based on the directional information of some pixels included in the prediction block of the current block.
[0015] In image decoding methods, the transform kernel can be the transform kernel with the least distortion among the transform kernels included in the transform set.
[0016] The image decoding method further includes: when the prediction mode of the current block is either intra template matching mode or intra block copy mode, obtaining intra prediction mode utilization information indicating whether to use intra prediction mode to determine the transform kernel, and when the intra prediction mode utilization information indicates that the intra prediction mode is used, determining the transform kernel may include: deriving the transform kernel to determine the intra prediction mode, and determining the transform kernel based on the intra prediction mode determined by the transform kernel.
[0017] In image decoding methods, when the current block is partitioned into a first sub-block and a second sub-block according to the partition boundaries, the transform set can be determined as the second transform set.
[0018] In image decoding methods, when the current block is divided into a first sub-block and a second sub-block according to the partition boundary, the transform kernel can be determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.
[0019] The image decoding method further includes: when the current block is partitioned into a first sub-block and a second sub-block according to the partition boundary, obtaining sub-block prediction mode utilization information that indicates whether the prediction mode of the first sub-block and the prediction mode of the second sub-block are used to determine the transform kernel; and when the sub-block prediction mode utilization information indicates that the prediction mode of the first sub-block and the prediction mode of the second sub-block are used, the transform kernel can be determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.
[0020] In image decoding methods, when the prediction mode of the current block is Spatial Combination Intra-Intra Prediction (SCIIP) mode, the transform kernel of the current block can be determined based on the first intra-prediction mode and the second intra-prediction mode of SCIIP mode.
[0021] The image coding method according to an embodiment of the present invention includes: determining a prediction mode for a current block; determining a transform set for the current block based on the prediction mode; and performing a transform on the current block based on the transform set, wherein the transform set can be determined as any one of a first transform set for intra-frame prediction, a second transform set for intra-frame prediction, and a third transform set for inter-frame prediction.
[0022] According to embodiments of the present invention, a non-volatile computer-readable recording medium for storing bit streams generated by an image encoding method can store bit streams generated by an image encoding method.
[0023] A method for transmitting a bitstream generated by an image coding method according to an embodiment of the present invention includes transmitting the bitstream, and the image coding method may include: determining a prediction mode of a current block; determining a transform set of the current block based on the prediction mode; and performing a transform on the current block based on the transform set, wherein the transform set is determined to be any one of a first transform set for a first intra-frame prediction, a second transform set for a second intra-frame prediction, and a third transform set for inter-frame prediction.
[0024] The features briefly outlined above are provided as examples to illustrate the detailed description and should not be construed as limiting the scope of the invention.
[0025] Beneficial effects According to the present invention, a method and apparatus for encoding / decoding images with improved encoding / decoding efficiency can be provided.
[0026] In addition, according to the present invention, a method for determining the transformation set for the inverse transformation of the current block can be provided.
[0027] In addition, according to the present invention, a method for determining a transform kernel for the inverse transform of the current block can be provided.
[0028] Furthermore, according to the present invention, a suitable transform set and transform kernel can be determined, thereby improving the transform efficiency.
[0029] The effects that can be obtained from the present invention are not limited to those described above, and other effects not mentioned will be clearly understood by those skilled in the art based on the following description. Attached Figure Description
[0030] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0031] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.
[0032] Figure 3 This is a schematic diagram illustrating a video encoding / decoding system to which the present invention can be applied.
[0033] Figure 4 This is a schematic diagram illustrating a method for determining a transform set when the prediction mode of the current block is an intra-frame template matching mode, according to an embodiment of the present invention.
[0034] Figure 5 This is a schematic diagram illustrating a method for determining a transform set when the prediction mode of the current block is an intra-block copy mode, according to an embodiment of the present invention.
[0035] Figure 6 This is a schematic diagram illustrating a method for determining the transform set and transform kernel under the Spatial Geometric Partitioning Mode (SGPM) according to an embodiment of the present invention.
[0036] Figure 7 This is a flowchart describing a method for determining a transformation set according to an embodiment of the present invention.
[0037] Figure 8 This is a schematic diagram illustrating a content streaming system applicable to embodiments of the present invention. Detailed Implementation
[0038] Best mode The image decoding method according to an embodiment of the present invention includes: deriving a prediction mode of the current block; determining a transform set of the current block based on the prediction mode; and performing an inverse transform on the current block based on the transform set, wherein the transform set can be determined as any one of a first transform set for intra-frame prediction, a second transform set for intra-frame prediction, and a third transform set for inter-frame prediction.
[0039] The mode of the present invention This invention can have various modifications and embodiments, and specific embodiments are shown in the accompanying drawings and described in detail in the specification. However, this is not intended to limit the invention to the specific embodiments, but rather to include all modifications, equivalents, or alternatives contained within the spirit and scope of the invention. The same reference numerals in the drawings indicate the same or similar functions in various aspects. For clarity, the shapes and dimensions of the elements in the drawings are provided by way of example. The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice them. It should be understood that the various embodiments differ from one another but are not necessarily mutually exclusive. For example, the specific shapes, structures, and features described herein may be implemented in other embodiments without departing from the spirit and scope of the invention with reference to one embodiment. It should also be understood that the position or arrangement of the various components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiments. Accordingly, the detailed description set forth below is not intended to be restrictive, and the scope of the exemplary embodiments is defined only by the appended claims and the full scope of their equivalents (if appropriately described).
[0040] In this invention, the terms first, second, etc., may be used to describe various components, but the components should not be limited by the terms. The terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the invention, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. The term "and / or" includes a combination of multiple related descriptive terms or any item from multiple related descriptive terms.
[0041] The components shown in the embodiments of the invention are depicted independently to indicate different functional characteristics, and do not represent each component as a separate hardware or software configuration unit. That is, for ease of interpretation, each component is listed and included as a separate component, and at least two components may be combined to form a single component, or a component may be divided into multiple components to perform functions, as long as it does not depart from the spirit of the invention. Embodiments in which components are integrated and embodiments in which each component is divided are also included within the scope of the invention.
[0042] The terminology used in this invention is for describing specific embodiments only and is not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, some components of this invention are not essential for performing the necessary functions and may be optional components used only to improve performance. This invention can be implemented by including only the essential components necessary for achieving the spirit of the invention, excluding components used only to improve performance, and structures that include only the essential components and excluding optional components used only to improve performance are also included within the scope of this invention.
[0043] In the implementation, the term "at least one" can mean one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In the implementation, the term "a plurality of" can mean one of a number greater than or equal to 2, such as 2, 3, and 4.
[0044] In the following, embodiments of the invention will be described in detail with reference to the accompanying drawings. When describing embodiments of this specification, detailed descriptions will be omitted if determining that a detailed description of a related known configuration or function would obscure the subject matter of this specification; the same reference numerals will be used for the same components in the drawings, and repeated descriptions of the same components will be omitted.
[0045] Description of terms In the following text, "image" can refer to a picture that constitutes a video, or it can refer to the video itself. For example, "encoding and / or decoding of an image" can mean "encoding and / or decoding of a video," or it can mean "encoding and / or decoding of one of the images that constitute a video."
[0046] In the following text, "moving image" and "video" can be used with the same meaning and can be used interchangeably. Additionally, the target image can be an encoded target image that is the target of encoding and / or a decoded target image that is the target of decoding. Furthermore, the target image can be an input image to an encoding device and can also be an input image to a decoding device. Here, the target image can have the same meaning as the current image.
[0047] In the following text, encoder and image encoding device can be used with the same meaning and can be used interchangeably.
[0048] In the following text, "decoder" and "image decoding device" can be used with the same meaning and can be used interchangeably.
[0049] In the following text, “image,” “picture,” “frame,” and “screen” can be used with the same meaning and can be used interchangeably.
[0050] In the following text, "target block" can refer to an encoding target block that is the target of encoding and / or a decoding target block that is the target of decoding. Additionally, a target block can refer to the current block that is the target of the current encoding and / or decoding. For example, "target block" and "current block" can be used with the same meaning and can be used interchangeably.
[0051] In the following text, "block" and "unit" can be used with the same meaning and can be used interchangeably. Additionally, "unit" can refer to a block comprising a luma component block and its corresponding chroma component block, in order to distinguish it from a block. For example, a Coding Tree Unit (CTU) can consist of a luma component (Y) Coding Tree Block (CTB) and two associated chroma component (Cb, Cr) Coding Tree Blocks.
[0052] In the following text, "sample," "image element," and "pixel" can be used with the same meaning and interchangeably. In this article, a sample can represent the basic unit that makes up a block.
[0053] In the following text, "between frames" and "between images" can be used with the same meaning and can be used interchangeably.
[0054] In the following text, "within the frame" and "within the screen" can be used with the same meaning and can be used interchangeably.
[0055] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0056] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images. The encoding device 100 may encode one or more images sequentially.
[0057] refer to Figure 1 The encoding device 100 may include: an image partitioning unit 110, an intra-frame prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switcher 115, a subtractor 113, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 117, a filtering unit 180, and a reference image buffer 190.
[0058] Additionally, the encoding device 100 can generate a bitstream including information encoded by encoding the input image, and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium.
[0059] Image partitioning unit 110 can partition the input image into various forms to improve the efficiency of video encoding / decoding. That is, the input video consists of multiple images, and for compression efficiency, parallel processing, etc., an image can be partitioned and processed hierarchically. For example, an image can be partitioned into one or more tiles or slices, and then further partitioned into multiple codec unit units (CTUs). Alternatively, an image can first be partitioned into multiple sub-images defined as groups of rectangular slices, and each sub-image can be partitioned into tiles / slices. Here, sub-images can be used to support the function of partially independent encoding / decoding and transmission of images. Since multiple sub-images can be reconstructed individually, this has the advantage of ease of editing in applications where multi-channel input is configured as a single image. Additionally, tiles can be horizontally partitioned to generate bricks. Here, bricks can be used as the basic unit for parallel processing within an image. Furthermore, a CTU can be recursively partitioned into a quadtree (QT), and the terminal node of the partition can be defined as a codec unit (CU). A CTU can be partitioned into a Prediction Unit (PU) and a Transform Unit (TU) to perform prediction and partitioning. Alternatively, a CTU can be used as both a Prediction Unit and / or a Transform Unit itself. For flexible partitioning, each CTU can be recursively partitioned into multi-type trees (MTTs) and quadtrees (QTs). Partitioning a CTU into multi-type trees can begin from the terminal node of a QT, and an MTT can consist of binary trees (BTs) and triple trees (TTs). For example, an MTT structure can be categorized into vertical binary partitioning patterns (SPLIT_BT_VER), horizontal binary partitioning patterns (SPLIT_BT_HOR), vertical ternary partitioning patterns (SPLIT_TT_VER), and horizontal ternary partitioning patterns (SPLIT_TT_HOR). Additionally, during partitioning, the minimum block size (MinQTSize) of the quadtree for the luma block can be set to 16×16, the maximum block size (MaxBtSize) of the binary tree can be set to 128×128, and the maximum block size (MaxTtSize) of the ternary tree can be set to 64×64. Furthermore, the minimum block sizes (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the ternary tree can be specified as 4×4, and the maximum depth (MaxMttDepth) of the multi-type tree can be specified as 4. Moreover, to improve the coding efficiency of I-slices, a dual-tree CTU partitioning structure using different luma and chroma components can be applied.On the other hand, in P and B slices, the luminance and chrominance codec tree blocks (CTBs) within the CTU can be partitioned into single trees sharing a codec tree structure.
[0060] The encoding device 100 can encode the input image in intra-frame mode and / or inter-frame mode. Alternatively, the encoding device 100 can encode the input image in a third mode other than intra-frame mode and inter-frame mode (e.g., IBC mode, palette mode, etc.). However, if the third mode has similar functional characteristics to the intra-frame mode or inter-frame mode, it can be classified as an intra-frame mode or inter-frame mode for ease of explanation. In this invention, the third mode is classified and described separately only when a specific explanation of the third mode is required.
[0061] When using intra-frame mode as the prediction mode, switcher 115 can switch to intra-frame mode, and when using inter-frame mode as the prediction mode, switcher 115 can switch to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. Encoding device 100 can generate prediction blocks for input blocks of the input image. In addition, encoding device 100 can encode residual blocks using the residuals of the input blocks and prediction blocks after generating prediction blocks. The input image can be referred to as the current image as the current coding target. The input block can be referred to as the current block as the current coding target or the coding target block.
[0062] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples of already encoded / decoded blocks surrounding the current block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction of the current block using the reference samples, or generate prediction samples for the input block through spatial prediction. In this paper, intra-frame prediction can refer to intra-frame prediction.
[0063] As an intra-frame prediction method, non-directional prediction modes such as DC mode and planar mode, as well as directional prediction modes (e.g., 65 directions) can be applied. Here, the intra-frame prediction method can be expressed as an intra-frame prediction mode or an intra-frame prediction mode.
[0064] When the prediction mode is inter-frame mode, the motion prediction unit 121 can retrieve the region that best matches the input block from the reference image during motion prediction processing, and derive the motion vector by utilizing the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference image buffer 190. Here, it can be stored in the reference image buffer 190 when encoding / decoding the reference image is performed.
[0065] The motion compensation unit 122 can generate a predicted block for the current block by performing motion compensation using motion vectors. In this paper, inter-frame prediction can refer to inter-picture prediction or motion compensation.
[0066] When the value of the motion vector is not an integer, the motion prediction unit 121 and the motion compensation unit 122 can generate prediction blocks by applying an interpolation filter to a portion of the reference image. To perform inter-frame prediction or motion compensation, the motion prediction and motion compensation modes of the prediction units included in the codec unit can be determined based on the codec unit. These modes include skip mode, merge mode, Advanced Motion Vector Prediction (AMVP) mode, and Intra Block Copy (IBC) mode, and inter-frame prediction or motion compensation can be performed according to each mode.
[0067] In addition, based on the above inter-frame prediction methods, the following modes can be applied: Affine mode based on sub-PU prediction, Subblock-based Temporal Motion Vector Prediction (SbTMVP) mode, Merge with MVD (MMVD) mode based on PU prediction, and Geometric Partitioning Mode (GPM) mode. In addition, to improve the performance of each mode, the following can be applied: History-based MVP (HMVP), Pairwise Average MVP (PAMVP), Combined Intra / Inter Prediction (CIIP), Adaptive Motion Vector Resolution (AMVR), Bi-Directional Optical-Flow (BDOF), Bi-predictive with CU Weights (BCW), Local Illumination Compensation (LIC), Template Matching (TM), and Overlapped Block Motion Compensation (OBMC).
[0068] Affine mode is a technique used in both AMVP and MERGE modes, and it also boasts high coding efficiency. In existing video codec standards, motion compensation (MC) is performed by considering only the parallel movement of blocks, thus it has the drawback of not being able to adequately compensate for real-world motion (e.g., zooming in / out and rotation). To compensate for this, a four-parameter affine motion model utilizing two control point motion vectors (CPMV) and a six-parameter affine motion model utilizing three control point motion vectors can be used and applied to inter-frame prediction. Here, CPMV is a vector representing one of the affine motion models of the upper left, upper right, and lower left sides of the current block.
[0069] Subtractor 113 can generate residual blocks by utilizing the difference between the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Alternatively, the residual signal can be a signal generated by transforming or quantizing, or transforming and quantizing the difference between the original signal and the prediction signal. The residual block can be a residual signal on a block-by-block basis.
[0070] Transform unit 130 can generate transform coefficients by performing a transform on the residual block and output the generated transform coefficients. In this document, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, transform unit 130 can skip the transform of the residual block.
[0071] Quantization levels can be generated by applying quantization to the transform coefficients or residual signal. In the following text, quantization levels may also be referred to as transform coefficients in the implementation scheme.
[0072] For example, the 4×4 lumen residual block generated by intra-frame prediction can be transformed using basis vectors based on Discrete Sine Transform (DST), and the remaining residual block can be transformed using basis vectors based on Discrete Cosine Transform (DCT). Furthermore, the transformed block can be partitioned into a quadtree shape using Residual Quad Tree (RQT) technology, and after performing transformation and quantization on each transformed block partitioned by RQT, the coded block flag (CBF) can be transmitted when all coefficients become 0 to improve coding efficiency.
[0073] As an alternative, Multiple Transform Selection (MTS) can be applied, selectively using multiple transform bases to perform the transform. That is, instead of partitioning the CU into TUs via RQT, a sub-block transform (SBT) technique can perform a function similar to TU partitioning. Specifically, SBT is applied only to inter-frame prediction blocks, and unlike RQT, it can partition the current block into 1 / 2 or 1 / 4 sizes vertically or horizontally, and then perform the transform on only one block. For example, if the current block is vertically partitioned, the transform can be performed on the leftmost or rightmost block, and if the current block is horizontally partitioned, the transform can be performed on the topmost or bottommost block.
[0074] In addition, the Low Frequency Non-Separable Transform (LFNST) can be applied. This is a secondary transform technique that additionally transforms the residual signal into the frequency domain using either DCT or DST. LFNST performs an additional transform on the upper left 4×4 or 8×8 low-frequency region, allowing the residual coefficients to be concentrated on the upper left.
[0075] The quantization unit 140 can generate quantization levels by quantizing the transform coefficients or residual signals according to the quantization parameters (QP), and output the generated quantization levels. In this paper, the quantization unit 140 can quantize the transform coefficients using a quantization matrix.
[0076] For example, quantizers with QP values from 0 to 51 can be used. Alternatively, if the image size is large and high coding efficiency is required, QP values from 0 to 63 can be used. Furthermore, dependent quantization (DQ) methods that utilize two quantizers instead of one can be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a particular quantizer, a state transition model can be used to select the quantizer for the next transform coefficient based on the current state.
[0077] Entropy coding unit 150 can generate and output a bitstream by performing entropy coding on values calculated by quantization unit 140 or on coding parameter values calculated during encoding, according to a probability distribution. Entropy coding unit 150 can perform entropy coding on information about samples of the image and information used to decode the image. For example, information used to decode the image may include syntax elements.
[0078] When entropy coding is applied, symbols are represented such that fewer bits are allocated to symbols with high occurrence probabilities and more bits are allocated to symbols with low occurrence probabilities, thus reducing the size of the bitstream used for the symbols to be encoded. The entropy coding unit 150 can perform entropy coding using coding methods such as exponential Golomb, context-adaptive Variable Length Coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by utilizing a variable-length coding / code (VLC) table. Alternatively, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / binary representation, and perform arithmetic coding by utilizing the derived binarization method and context model.
[0079] Relatedly, when applying CABAC, to reduce the size of the probability table stored in the decoding device, the table probability update method can be modified into a table update method using a simple equation and then applied. Alternatively, two different probability models can be used to obtain more accurate symbol probability values.
[0080] In order to encode the transform coefficient level (quantization level), the entropy coding unit 150 can change the coefficients in two-dimensional block form into one-dimensional vector form by the transform coefficient scanning method.
[0081] The encoding parameters may include information (flags, indexes, etc.) encoded in the encoding device 100 and signaled to the decoding device 200, such as syntax elements, as well as information derived during the encoding or decoding process, and may represent the information required when encoding or decoding an image.
[0082] In this paper, emitting a flag or index with a signal can indicate that the corresponding flag or index is entropy encoded in the encoder and included in the bitstream, and can also indicate that the corresponding flag or index is entropy decoded from the bitstream in the decoder.
[0083] The encoded current image can be used as a reference image for another image to be processed later. Therefore, the encoding device 100 can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference image buffer 190.
[0084] The quantization level can be dequantized in dequantization unit 160, or inverse transformed in inverse transform unit 170. The coefficients of dequantization and / or inverse transform can be added to the prediction block via adder 117. In this document, the coefficients of dequantization and / or inverse transform can represent coefficients from which at least one dequantization and inverse transform have been performed, and can represent the reconstructed residual block. Dequantization unit 160 and inverse transform unit 170 can be performed as the inverse of quantization unit 140 and transform unit 130.
[0085] The reconstructed block can be processed by filtering unit 180. Filtering unit 180 can apply all or some of the following filtering techniques to the reconstructed sample, reconstructed block, or reconstructed image: deblocking filter, Sample Adaptive Offset (SAO) filter, Adaptive Loop Filter (ALF), Bilateral Filter (BIF), Luma Mapping with Chroma Scaling (LMCS) filter, etc. Filtering unit 180 can be referred to as an in-loop filter. In this case, in-loop filter can also be used as a name excluding LMCS.
[0086] Deblocking filters remove block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, the application of the filter to the current block can be determined based on samples included in several rows or columns contained within the block. When applying a deblocking filter to a block, different filters can be applied depending on the desired deblocking filtering intensity.
[0087] To compensate for coding errors using sample-adaptive offsets, appropriate offset values can be added to the sample values. Sample-adaptive offsets can correct the offset between the deblocked image and the original image on a sample-by-sample basis. Methods include dividing the image into a predetermined number of regions, determining the regions to which the offset will be applied, and then applying the offset to those regions; or considering edge information about each sample when applying the offset.
[0088] Bilateral filters (BIF) can also correct the offset from the original image on a sample-by-sample basis for images that have already undergone deblocking.
[0089] Adaptive loop filters can perform filtering based on a comparison between the reconstructed image and the original image. Samples included in the image can be partitioned into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information regarding whether to apply the ALF can be emitted by the codec unit (CU) using a signal, and the form and coefficients of the adaptive loop filter to be applied to each block can vary.
[0090] In Luma Mapping with Chroma Scaling (LMCS), luminance mapping (LM) represents remapping luminance values using a piecewise linear model, and chroma scaling (CS) represents scaling the residual values of the chroma components based on the average luminance values of the predicted signal. Specifically, LMCS can be used as an HDR correction technique that reflects the characteristics of High Dynamic Range (HDR) images.
[0091] The reconstructed block or reconstructed image that has passed through filtering unit 180 can be stored in reference image buffer 190. The reconstructed block that has passed through filtering unit 180 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks that have passed through filtering unit 180. The stored reference image can be used later for inter-frame prediction or motion compensation.
[0092] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.
[0093] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.
[0094] refer to Figure 2 The decoding device 200 may include: an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 201, a switcher 203, a filtering unit 260, and a reference image buffer 270.
[0095] The decoding device 200 can receive a bitstream output from the encoding device 100. The decoding device 200 can receive a bitstream stored in a computer-readable recording medium, or it can receive a bitstream streamed via a wired / wireless transmission medium. The decoding device 200 can decode the bitstream in intra-frame mode or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image through decoding, and output the reconstructed image or the decoded image.
[0096] When the prediction mode used for decoding is intra-frame mode, switcher 203 can switch to intra-frame mode. Alternatively, when the prediction mode used for decoding is inter-frame mode, switcher 203 can switch to inter-frame mode.
[0097] The decoding device 200 can obtain a reconstructed residual block and generate a prediction block by decoding the input bitstream. When the reconstructed residual block and prediction block are obtained, the decoding device 200 can add them together to generate a reconstructed block that becomes the decoding target. The decoding target block can be referred to as the current block.
[0098] The entropy decoding unit 210 can generate symbols by entropy decoding the bitstream according to the probability distribution. The generated symbols can include symbols in quantization level form. In this paper, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0099] The entropy decoding unit 210 can change the coefficients of a one-dimensional vector shape into coefficients of a two-dimensional block shape through the transformation coefficient scanning method to decode the transformation coefficient level (quantization level).
[0100] The quantization level can be dequantized in the dequantization unit 220 or inverse transformed in the inverse transform unit 230. The quantization level can be the result of dequantization and / or inverse transform, and can be generated as a reconstructed residual block. In this document, the dequantization unit 220 can apply a quantization matrix to the quantization level. The dequantization unit 220 and inverse transform unit 230 applied to the decoding device can employ the same techniques as those applied to the dequantization unit 160 and inverse transform unit 170 of the aforementioned encoding device.
[0101] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, which uses sample values from blocks that have already been decoded around the target block. The intra-frame prediction unit 240 applied to the decoding apparatus can employ the same techniques as the intra-frame prediction unit 120 applied to the aforementioned encoding apparatus.
[0102] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block using motion vectors and a reference image stored in reference image buffer 270. When the value of the motion vector is not an integer value, motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion compensation mode of the prediction unit included in the corresponding encoding / decoding unit can be determined based on the encoding / decoding unit, whether it is a skip mode, merge mode, AMVP mode, or current image reference mode, and motion compensation can be performed according to each mode. The motion compensation unit 250 applied to the decoding device can apply the same techniques as the motion compensation unit 122 applied to the encoding device described above.
[0103] Adder 201 generates a reconstructed block by adding the reconstructed residual block and the prediction block. Filtering unit 260 can apply at least one of inverse LMCS, deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or reconstructed image. Filtering unit 260 applied to the decoding apparatus can apply the same filtering techniques as filtering unit 180 applied to the aforementioned encoding apparatus.
[0104] Filtering unit 260 can output a reconstructed image. The reconstructed block or reconstructed image can be stored in reference image buffer 270 and used for inter-frame prediction. The reconstructed block that has passed through filtering unit 260 can be a portion of the reference image. That is, the reference image can be a reconstructed image composed of reconstructed blocks that have passed through filtering unit 260. The stored reference image can be used later for inter-frame prediction or motion compensation.
[0105] Figure 3 This is a schematic diagram illustrating a video encoding / decoding system to which the present invention can be applied.
[0106] The video encoding / decoding system according to the implementation scheme may include an encoding device 10 and a decoding device 20. The encoding device 10 may send encoded video and / or image information or data to the decoding device 20 in the form of files or streaming via a digital storage medium or network.
[0107] The encoding apparatus 10 according to the embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding apparatus 20 according to the embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be configured as a separate device or an external component.
[0108] The video source generation unit 11 can obtain video / images through processes of capturing, compositing, or generating video / images. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., in which case the process of generating related data can replace the video / image capture process.
[0109] Encoding unit 12 can encode the input video / image. For compression and encoding efficiency, encoding unit 12 can perform a series of processes such as prediction, transformation, and quantization. Encoding unit 12 can output the encoded data (encoded video / image information) as a bitstream. Detailed configuration of encoding unit 12 can also be found in the above description. Figure 1 The encoding device 100 is configured in the same way.
[0110] The transmitting unit 13 can send encoded video / image information or data, output in bitstream form, to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in the form of a file or stream. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 can include elements for generating media files according to a predetermined file format, and may include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and send it to the decoding unit 22.
[0111] Decoding unit 22 can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, corresponding to the operations of encoding unit 12. Detailed configuration of decoding unit 22 can also be found in... Figure 2 The decoding device 200 described above is configured in the same manner.
[0112] Rendering unit 23 can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0113] In this specification, a transform set can refer to a group that includes transform kernels, and a transform kernel can refer to a transform kernel used when applying a transform from the spatial domain to the frequency domain. Additionally, a transform set can refer to a kernel cluster.
[0114] A quadratic transform can refer to a transform performed after a primary transform from the spatial domain to the frequency domain, based on the correlation of the primary transform coefficients generated by the primary transform. Here, the primary transform and the main transform can have the same meaning. By performing a quadratic transform, higher compression efficiency can be achieved than by performing only a primary transform.
[0115] First-order and second-order transformations can be performed as separable or non-separable transformations. Here, a separable transformation can refer to performing a transformation in the horizontal direction first, and then a transformation in the vertical direction. Conversely, a non-separable transformation can refer to performing the transformation in one step using a matrix-form non-separable transformation kernel, rather than performing the transformation in the horizontal and vertical directions separately.
[0116] In non-separable transformations, basis vectors are used. A basis vector is a vector with a size corresponding to the total number of pixels in the two-dimensional space where the transformation is performed, and it is used to identify the overall features and similarity of pixels in the two-dimensional space where the transformation is performed. Therefore, non-separable transformations are more complex than separable transformations, but they can achieve high coding efficiency.
[0117] In a quadratic transform similar to the Low Frequency Inseparable Transform (LFNST), considering the enhanced encoding / decoding efficiency and the transform complexity for a specific input block size, inseparable kernels of various sizes can be used.
[0118] In the Non-separable primary transform (NSPT), similar to LFNST, non-separable primary transform kernels of various sizes can be used to enhance encoding / decoding efficiency and reduce the transform complexity for specific input block sizes.
[0119] A transform kernel selected from multiple transform kernels can be used for the transform of the current block. This can be called multiple transform selection (MTS).
[0120] Alternatively, a transform set can be selected from multiple transform sets, and the transformation of the current block can be performed by utilizing at least one of the multiple transform kernels included in the selected transform set. Here, selecting a transform set from multiple transform sets can be referred to as Multiple Transform Set Selection (MTSS).
[0121] Traditional NSPT and LFNST can both be used only for intra-prediction blocks. Below, this specification provides methods for determining the transform set and transform kernel for inter-prediction blocks. Additionally, this specification provides methods for determining the transform set and transform kernel for blocks performing specific intra-prediction operations (e.g., Intra-template matching prediction (IntraTMP), Intra-block copy (IBC), Spatial geometric partition mode (SGPM), and Spatial combined intra-intra prediction (SCIIP)).
[0122] The method for determining a transform set according to the present invention will be described below. The method for determining a transform set refers to a method for determining, from multiple transform sets, a transform set for the transform / inverse transform of the current block.
[0123] As an example, the method for determining the transform set refers to a method for determining the transform / inverse transform set for the current block from three transform sets. Here, the first transform set (hereinafter referred to as the first transform set) and the second transform set (hereinafter referred to as the second transform set) can be transform sets for performing transform / inverse transforms of blocks for intra-frame prediction, and the third transform set (hereinafter referred to as the third transform set) can be transform sets for performing transform / inverse transforms of blocks for inter-frame prediction. Additionally, the transform / inverse transform set for the current block can be determined from the aforementioned three transform sets.
[0124] On the other hand, the first transform set can be a transform / inverse transform set used for conventional intra-prediction blocks. Here, conventional intra-prediction blocks can refer to prediction blocks whose intra-prediction mode is either a non-directional prediction mode (e.g., DC mode and planar mode) or a directional prediction mode (e.g., 65 directions).
[0125] On the other hand, the second transform set can be a transform / inverse transform set for blocks that perform intra-prediction by utilizing any of the following techniques: spatial geometric prediction mode (SGPM), template-based intra-mode derivation (TIMD), decoder-side intra-mode derivation (DIMD), extrapolation filter-based intra-prediction (EIP), and matrix-based intra-prediction (MIP). Here, SGPM can refer to a technique that spatially partitions the current block into at least two or more regions and performs intra-prediction, and TIMD can refer to a technique that derives the prediction mode of the current block by utilizing a reference region of a template. Additionally, DIMD can refer to a technique that derives the prediction mode of the current block in the decoder without resolution, EIP can refer to a technique that performs prediction by utilizing an extrapolation filter, and MIP can refer to a matrix-based prediction technique.
[0126] On the other hand, the above-mentioned transformation set can be a transformation set of the Inseparable First-Order Transformation (NSPT). Additionally, the above-mentioned transformation set can be a transformation set of the Inseparable Second-Order Transformation (LFNST).
[0127] On the other hand, the transform set determination method according to the present invention can be applied even when the prediction mode of the current block is not the intra-frame prediction mode.
[0128] On the other hand, according to the above implementation scheme, the number of transformation sets is three, but this is only an example, and the number of transformation sets can be S. Here, S is an integer of 2 or greater.
[0129] On the other hand, according to the above implementation scheme, the first transform set and the second transform set are transform sets for performing the inverse transform of the current block for intra-frame prediction, and the third transform set is a transform set for performing the inverse transform of the current block for inter-frame prediction. However, this is only an example, and any transform set can correspond to performing the inverse transform of the current block for prediction in any way.
[0130] According to an embodiment of the present invention, the transform set can be determined based on the prediction mode of the current block. As an example, when the prediction mode of the current block is an inter-frame prediction mode, the transform set of the current block can be determined as the third transform set among the first, second, and third transform sets described above. As stated above, in the present invention, non-separable transforms can be performed on blocks other than intra-frame prediction blocks.
[0131] Figure 4 and Figure 5 It is a schematic diagram used to describe the determination of the transform set based on the prediction mode of the current block.
[0132] Figure 4 This is a schematic diagram illustrating a method for determining a transform set when the prediction mode of the current block is an intra-frame template matching mode, according to an embodiment of the present invention.
[0133] The IntraTMP method refers to a method that uses template matching to search for the best prediction block in the reconstructed region of the current image and generates the prediction block for the current block by copying the best prediction block. When the prediction mode of the current block is intra-template matching mode, the prediction block for the current block is generated based on the intra-template matching prediction method.
[0134] refer to Figure 4 The adjacent L-shaped regions of the current block 410 (i.e., the left region, the upper region, and the upper left region) can be defined as the current template 420. Additionally, a reference template 440 most similar to the current template 420 can be searched within the predefined search ranges R1, R2, R3, and R4 of the reconstructed region 430 of the current image 400. Furthermore, the predicted block of the current block 410 can be derived based on the matching block 450 corresponding to the determined reference template 440.
[0135] Can Figure 4The predefined search ranges R1, R2, R3, and R4 are defined as including the current codec tree unit (CTU), the top-left CTU, the top CTU, and the left-side CTU of the current block, respectively. Here, the top-left CTU can refer to the CTU to the top left of the current CTU within the preset area of the current CTU, the top CTU can refer to the CTU above the current CTU within the preset area, and the left-side CTU can refer to the CTU to the left of the current CTU within the preset area. The preset area can be any region including the current tile, the current slice, and the current image of the current block.
[0136] Additionally, reference templates can be searched within a predefined search range based on a predefined search order. For example, reference templates can be searched in a zigzag order of R1, R4, R3, and R2.
[0137] On the other hand, information about the search range, as well as the size and shape of the current template, can be determined in the encoder and sent to the decoder. Additionally, the search range and the size and shape of the current template can be set to predetermined values in the encoder / decoder.
[0138] As an example, in addition to using both the left and top reconstruction regions as L-shaped templates, either the top reconstruction region or the left reconstruction region can also be used as a template.
[0139] As an example, the size of the current template can be determined by (w×L2) + (L1×h) + (L1×L2), such as Figure 4 As shown. Here, w and h can represent the width and height of the current block, and the values of L1 and L2 can be any positive integers.
[0140] Regarding template shape Figure 4 The example uses an L-shaped template to search for the most similar reference template to the current template, but this is just an example, and you can search for the most similar reference template to the current template by using an L-shaped template, a template that only uses the upper region, a template that only uses the left region, or a template with any shape.
[0141] On the other hand, when the prediction mode of the current block is intra-template matching mode, the transform / inverse transform can be performed in the same way as the transform / inverse transform method for intra-predictive blocks. That is, even when the prediction mode of the current block is intra-template matching mode, the transform set of the intra-block can be used. As an example, when the prediction mode of the current block is intra-template matching mode, an inseparable transform / inverse transform can be performed using NSPT and LFNST. As another example, when the prediction mode of the current block is intra-template matching mode, the transform set can be determined as the first transform set described above.
[0142] Alternatively, in the case of a prediction block in intra-template matching mode, due to high prediction accuracy, the characteristics of the residuals and the characteristics of the first-order transform coefficients generated by the first-order transform can be similar to those of the prediction block in inter-frame prediction mode. Therefore, the transform set of the inter-frame block can be used for the block in intra-template matching mode. As an example, when the prediction mode of the current block is intra-template matching mode, the transform set of the current block can be determined as the third transform set mentioned above.
[0143] Alternatively, when the prediction mode for the current block is intra-template matching mode, the transform set can be determined as either the transform set of intra-blocks (hereinafter referred to as the intra-transform set) or the transform set of inter-blocks (hereinafter referred to as the inter-transform set). Specifically, the transform set for the current block can be determined as the transform set with the smaller cost value between the intra-transform set and the inter-transform set.
[0144] For example, the rate-distortion cost (RD cost) can be calculated for each transform set in the encoder, and the transform set with the smaller distortion can be selected as the optimal transform set.
[0145] In addition to the information normally emitted as signals during the transform, information about the optimal transform set can be emitted as signals to the decoder. Here, the information about the optimal transform set can be either a flag indicating whether to use the intra-frame transform set or a flag indicating whether to use the inter-frame transform set.
[0146] Additionally, the transform set of the current block can be determined in the decoder based on information about the optimal transform set. For example, if the information about the optimal transform set is a flag indicating whether an intra-frame transform set is used, and this information indicates that an intra-frame transform set is used, the transform set of the current block can be determined as an intra-frame transform set. On the other hand, if this information does not indicate that an intra-frame transform set is used, the transform set of the current block can be determined as an inter-frame transform set. Similarly, if the information about the optimal transform set is a flag indicating whether an inter-frame transform set is used, and this information indicates that an inter-frame transform set is used, the transform set of the current block can be determined as an inter-frame transform set; however, if this information does not indicate that an inter-frame transform set is used, the transform set of the current block can be determined as an intra-frame transform set.
[0147] On the other hand, information about the optimal transform set can be a bit flag and can be signaled at any level, such as block level, slice level, picture level, or sequence level.
[0148] Figure 5 This is a schematic diagram illustrating a method for determining a transform set when the prediction mode of the current block is an intra-block copy mode, according to an embodiment of the present invention.
[0149] The intra-block copying method searches for the best prediction block in the reconstruction region of the current image using block vectors, and then generates the prediction block for the current block by copying the best prediction block. When the prediction mode for the current block is intra-block copying mode, the prediction block for the current block is generated based on the intra-block copying method.
[0150] refer to Figure 5 Based on the block vector 520 of the current block 510, a matching block 540 can be derived within a predefined search range R1, R2, R3, and R4 of the reconstruction region 530 of the current image 500. Furthermore, a predicted block for the current block 510 can be derived based on the matching block 540. On the other hand, the block vector 520 is a vector indicating the matching block 540 of the current block 510 in intra-frame block copying.
[0151] Can Figure 5 The predefined search ranges R1, R2, R3, and R4 are defined as including the current codec tree unit (CTU), the top-left CTU, the top CTU, and the left-side CTU of the current block, respectively. Here, the top-left CTU can refer to the CTU to the top left of the current CTU within the preset area of the current CTU, the top CTU can refer to the CTU above the current CTU within the preset area, and the left-side CTU can refer to the CTU to the left of the current CTU within the preset area. The preset area can be any region including the current tile, the current slice, and the current image of the current block.
[0152] Additionally, matching blocks can be searched within a predefined search range based on a predefined search order. For example, matching blocks can be searched in a zigzag order of R1, R4, R3, and R2.
[0153] On the other hand, information about the search range can be determined in the encoder and sent to the decoder. Additionally, the search range can be set to a predetermined value in the encoder / decoder.
[0154] On the other hand, when the prediction mode of the current block is intra-block copy mode, the transform / inverse transform can be performed in the same way as the transform / inverse transform method for intra-predicted blocks. That is, even when the prediction mode of the current block is intra-block copy mode, the transform set of the intra-block can be used. As an example, when the prediction mode of the current block is intra-block copy mode, an inseparable transform / inverse transform can be performed using NSPT and LFNST. As another example, when the prediction mode of the current block is intra-block copy mode, the transform set can be determined as the first transform set described above.
[0155] Alternatively, in the case of a prediction block in intra-block copy mode, due to high prediction accuracy, the characteristics of the residuals and the characteristics of the first-order transform coefficients generated by the first-order transform can be similar to those of the prediction block in inter-frame prediction mode. Therefore, the transform set of the inter-frame block can be used for the block in intra-block copy mode. As an example, when the prediction mode of the current block is intra-block copy mode, the transform set of the current block can be determined as the third transform set mentioned above.
[0156] Alternatively, when the prediction mode for the current block is intra-block copy mode, the transform set can be determined as either the transform set for intra-blocks (hereinafter referred to as the intra-transform set) or the transform set for inter-blocks (hereinafter referred to as the inter-transform set). Specifically, the transform set for the current block can be determined as the transform set with the smaller cost value between the intra-transform set and the inter-transform set.
[0157] For example, cost values can be calculated within the encoder. Rate-distortion cost (RD cost) can be calculated for each transform set within the encoder, and the transform set with the lowest distortion can be selected as the optimal transform set.
[0158] In addition to the information normally emitted as signals during the transform, information about the optimal transform set can be emitted as signals to the decoder. Here, the information about the optimal transform set can be either a flag indicating whether to use the intra-frame transform set or a flag indicating whether to use the inter-frame transform set.
[0159] Additionally, the transform set of the current block can be determined in the decoder based on information about the optimal transform set. For example, if the information about the optimal transform set is a flag indicating whether an intra-frame transform set is used, and this information indicates that an intra-frame transform set is used, the transform set of the current block can be determined as an intra-frame transform set. On the other hand, if this information does not indicate that an intra-frame transform set is used, the transform set of the current block can be determined as an inter-frame transform set. Similarly, if the information about the optimal transform set is a flag indicating whether an inter-frame transform set is used, and this information indicates that an inter-frame transform set is used, the transform set of the current block can be determined as an inter-frame transform set; however, if this information does not indicate that an inter-frame transform set is used, the transform set of the current block can be determined as an intra-frame transform set.
[0160] On the other hand, information about the optimal transform set can be a bit flag and can be signaled at any level, such as block level, slice level, picture level, or sequence level.
[0161] On the other hand, in the above implementation scheme, the transform set of the current block can be determined based on the prediction mode of the current block without separate signaling. For example, when the prediction mode of the current block is any one of inter-frame prediction mode, intra-frame template matching mode, and intra-frame block copy mode, the transform set of the current block can be determined as the aforementioned third transform set in the decoder.
[0162] According to an embodiment of the present invention, in addition to determining the transformation set of the current block based on the prediction mode of the current block, the transformation set of the current block can also be determined in other ways.
[0163] As an example, the transform set can be determined based on the directional information of pixels in the entire or partial region of the prediction block included in the current block. Here, directional information can refer to information obtained by calculating the gradient of the pixels.
[0164] As another example, the transform set can be determined based on the directional information of integer pixels in the entire or partial region of the prediction block included in the current block, as well as the directional information of subpixels obtained by applying an interpolation filter to that region.
[0165] As another example, the transform set can be determined based on the directional information of pixels in the entire or partial region of either the template adjacent to the current block or the template corresponding to the current block.
[0166] As another example, the transform set can be determined based on the directional information of integer pixels in a whole or part of a region included in either the template adjacent to the current block or the template corresponding to the current block, and the directional information of subpixels obtained by applying an interpolation filter to that region.
[0167] On the other hand, the aforementioned directional information can refer to information obtained by calculating the gradient of a pixel. Specifically, directional information can refer to information obtained by accumulating or processing the gradient information calculated by applying a filter to a pixel. The filter applied to the pixel can be any of the following: Sobel filter, Roberts cross filter, Prewitt filter, Scharr filter, and Laplacian filter.
[0168] On the other hand, the templates adjacent to the current block can refer to any region adjacent to the current block, and the templates corresponding to the current block can refer to any region with the same size as the templates adjacent to the current block.
[0169] As another example, when the prediction mode of the current block is inter-frame prediction mode, the transform set can be determined based on the magnitude of at least one of the motion vectors in the x-direction and y-direction of the current block. Alternatively, when the prediction mode of the current block is inter-frame prediction mode, the transform set can be determined based on values obtained by processing the magnitude of at least one of the motion vectors in the x-direction and y-direction of the current block.
[0170] As another example, when the prediction mode for the current block is inter-frame prediction mode, the transform set can be determined based on the precision of at least one of the motion vectors in the x-direction and y-direction of the current block. Specifically, the transform set can be determined based on whether the precision of the motion vectors is integer pixel precision or fractional pixel precision.
[0171] As another example, the transformation set can be determined based on at least one of the horizontal and vertical lengths of the current block.
[0172] As yet another example, the transform set can be determined based on the number of pixels included in the current block.
[0173] As yet another example, the transform set can be determined based on information about the partitions of the current block.
[0174] As another example, the transform set can be determined based on syntax elements explicitly signaled. Specifically, when determining the transform set in the encoder, the syntax for determining the transform set can be signaled to the decoder. The transform set for the inverse transform can then be determined in the decoder based on the syntax elements.
[0175] As another example, the transform set can be determined based on the values of the neighboring pixels of the current block. Here, neighboring pixels can refer to all or some pixels included in the block adjacent to the current block, or it can refer to all or some pixels included in the block corresponding to the current block.
[0176] As yet another example, the transform set can be determined based on the quantization coefficients.
[0177] On the other hand, the transformation set for the current block can be determined from the S transformation sets using the method described above. Here, S is an integer of 2 or greater.
[0178] On the other hand, the above-described method for determining the transform set can be performed even when the prediction mode of the current block is the inter-frame prediction mode.
[0179] According to an embodiment of the present invention, the transform set of the current block can be determined by deriving an intra-prediction mode for transform set determination. Specifically, the intra-prediction mode for transform set determination can be derived based on the directional information of pixels included in the prediction block of the current block, and the transform set can be determined by utilizing the intra-prediction mode. On the other hand, directional information can refer to information obtained by calculating the gradient of pixels.
[0180] First, the gradients of the pixels included in the prediction block of the current block can be calculated. Here, the gradients can be calculated by applying at least one of the Sobel filter, Roberts cross filter, Prewitt filter, Scharr filter, and Laplacian filter to the pixels included in the prediction block.
[0181] On the other hand, gradients can be computed only for some pixels included in the prediction block of the current block. For example, when a sub-block transform (SBT) is applied to the current block and the current block is partitioned, and the transform is performed on only one block, gradients can be computed only for pixels included in the region of the prediction block corresponding to the block on which the transform was performed. As another example, after downsampling pixels included in the prediction block of the current block, gradients can be computed only for the downsampled pixels.
[0182] Additionally, a gradient histogram (HoG) can be generated based on the calculated gradient.
[0183] Additionally, HoG can be used to derive intra-prediction modes for transform set determination.
[0184] As an example, the intra-prediction mode used for transform set determination can be an intra-prediction mode that corresponds to the directionality with the highest accumulation in HoG.
[0185] As another example, the intra-prediction mode used for transform set determination can be the prediction mode with the fewest allocated number of prediction modes among the N prediction modes corresponding to the highest cumulative directionality in HoG. Here, N is any positive integer.
[0186] As another example, N intra-prediction modes corresponding to the highest cumulative directionality in HoG can be derived, and indices for transform set determination can be derived by utilizing these N intra-prediction modes. Furthermore, the transform set for the current block can be determined based on these indices. Here, N is any positive integer.
[0187] As another example, when the directionality accumulated in the HoG is below a predetermined threshold, the intra-prediction mode used for transform set determination can be derived as a predetermined intra-prediction mode. Here, the predetermined intra-prediction mode can be any one of DC mode, vertical DC mode, horizontal DC mode, planar mode, vertical planar mode, and horizontal planar mode.
[0188] On the other hand, predetermined thresholds can be defined separately in the encoder and decoder. Alternatively, predetermined thresholds can be defined in the encoder and transmitted to the decoder as a signal. Alternatively, predetermined thresholds can be defined based on a protocol between the encoder and decoder.
[0189] Alternatively, the transform set can be determined based on the derived intra-prediction mode. Specifically, the transform set is determined as a set of transforms that map from multiple transform sets to the derived intra-prediction mode. Here, the mapping between the intra-prediction mode and the transform set can be performed in the same way as the transforms of intra-blocks. For example, the mapping can be performed in the same way as LFNST and NSPT.
[0190] Table 1 is a graph showing an example of mapping 95 intra-frame prediction modes to transform sets in LFNST and NSPT.
[0191] [Table 1] On the other hand, in the above implementation, when accumulating the directionality in HoG, a method of accumulation using one interval or a method of accumulation using multiple intervals can be used. Here, the method of accumulation using multiple intervals can refer to a method of accumulation using one interval corresponding to the directionality of the gradient calculated by the filter and M intervals surrounding said one interval. M is an integer of 1 or greater.
[0192] On the other hand, in the above implementation, directional information is obtained by calculating the gradient of the pixels included in the prediction block of the current block. However, this is just an example, and directional information can also be obtained by calculating the gradient of integer pixels in the entire or part of the prediction block of the current block and the gradient of sub-pixels obtained by applying an interpolation filter to that region.
[0193] On the other hand, in the above implementation, directional information is obtained by calculating the gradients of pixels included in the prediction block of the current block. However, this is just an example, and directional information can also be obtained by calculating the gradients of pixels included in either the template adjacent to the current block or the template corresponding to the current block. Alternatively, directional information can be obtained by calculating the gradients of integer pixels in the entire region or a portion of the template adjacent to the current block and the sub-pixels obtained by applying an interpolation filter to that region. Here, the template adjacent to the current block can refer to any region adjacent to the current block, and the template corresponding to the current block can refer to any region of the same size as the template adjacent to the current block.
[0194] Hereinafter, a method according to the invention for determining a transform kernel from a transform kernel included in a transform set for a transform / inverse transform of a current block will be described.
[0195] Once the transform set is determined, the transform / inverse transform of the current block can be performed based on the transform set. Specifically, a transform kernel for the transform / inverse transform of the current block can be determined from K transform kernels included in the transform set, and the transform / inverse transform of the current block can be performed based on that transform kernel. Here, K is an integer of 2 or greater.
[0196] On the other hand, when the transform set includes a transform kernel, the transform kernel of the current block can be determined as the transform kernel.
[0197] According to an embodiment of the present invention, when the prediction mode of the current block is inter-frame prediction mode, the transform kernel of the current block can be determined as a predefined special transform kernel in the encoder / decoder. In this case, the encoder / decoder can check whether the prediction mode of the current block is inter-frame prediction mode, and can perform transform / inverse transform based on the predefined transform kernel when the prediction mode is inter-frame prediction mode.
[0198] According to embodiments of the present invention, the transform kernel can be determined based on the orientation of pixels included in the prediction block of the current block. Specifically, an intra-prediction mode for transform kernel determination can be derived based on the orientation information of pixels included in the prediction block of the current block, and the transform kernel can be determined by utilizing the intra-prediction mode. On the other hand, the orientation information can refer to information obtained by calculating the gradient of pixels.
[0199] First, the gradients of the pixels included in the prediction block of the current block can be calculated. Here, the gradients can be calculated by applying at least one of the Sobel filter, Roberts cross filter, Prewitt filter, Scharr filter, and Laplacian filter to the pixels included in the prediction block.
[0200] On the other hand, gradients can be computed only for some pixels included in the prediction block of the current block. For example, when a sub-block transform (SBT) is applied to the current block and the current block is partitioned, and then the transform is performed on only one block, gradients can be computed only for pixels included in the region of the prediction block corresponding to the block on which the transform was performed. As another example, after downsampling pixels included in the prediction block of the current block, gradients can be computed only for the downsampled pixels.
[0201] Additionally, a gradient histogram (HoG) can be generated based on the calculated gradient.
[0202] Additionally, the intra-prediction mode used for transform kernel determination can be derived using HoG. Here, the intra-prediction mode used for transform kernel determination can have the same meaning as the intra-prediction mode determined by transform kernel determination.
[0203] According to an embodiment of the present invention, the transformation kernel of the current block can be determined as follows.
[0204] As an example, the transform kernel can be determined based on the directional information of pixels in the entire or partial region of the prediction block included in the current block. Here, directional information can refer to information obtained by calculating the gradient of the pixels.
[0205] As another example, the transform kernel can be determined based on the directional information of integer pixels in the entire or partial region of the prediction block included in the current block, as well as the directional information of subpixels obtained by applying an interpolation filter to that region.
[0206] As another example, the transform kernel can be determined based on the directional information of pixels in the entire region or a portion of the region of either the template adjacent to the current block or the template corresponding to the current block.
[0207] As another example, the transform kernel can be determined based on the directional information of integer pixels in a whole or part of a region of either a template adjacent to the current block or a template corresponding to the current block, and the directional information of subpixels obtained by applying an interpolation filter to that region.
[0208] On the other hand, the aforementioned directional information can refer to information obtained by calculating the gradient of a pixel. Specifically, directional information can refer to information obtained by accumulating or processing the gradient information calculated by applying a filter to a pixel. The filter applied to the pixel can be any of the following: Sobel filter, Roberts cross filter, Prewitt filter, Scharr filter, and Laplacian filter.
[0209] On the other hand, the templates adjacent to the current block can refer to any region adjacent to the current block, and the templates corresponding to the current block can refer to any region with the same size as the templates adjacent to the current block.
[0210] As another example, when the prediction mode of the current block is inter-frame prediction mode, the transform kernel can be determined based on the magnitude of at least one of the motion vectors in the x-direction and y-direction of the current block. Alternatively, when the prediction mode of the current block is inter-frame prediction mode, the transform kernel can be determined based on a value obtained by processing the magnitude of at least one of the motion vectors in the x-direction and y-direction of the current block.
[0211] As another example, when the prediction mode for the current block is inter-frame prediction mode, the transform kernel can be determined based on the precision of at least one of the motion vectors in the x-direction and y-direction of the current block. Specifically, the transform kernel can be determined based on whether the precision of the motion vectors is integer pixel precision or fractional pixel precision.
[0212] As another example, the transformation kernel can be determined based on at least one of the horizontal and vertical lengths of the current block.
[0213] As yet another example, the transform kernel can be determined based on the number of pixels included in the current block.
[0214] As another example, the transform kernel can be determined based on information about the partitions of the current block.
[0215] As another example, the transform kernel can be determined based on the values of the neighboring pixels of the current block. Here, neighboring pixels can refer to all or some pixels included in the block adjacent to the current block, or it can refer to all or some pixels included in the block corresponding to the current block.
[0216] As yet another example, the transform kernel can be determined based on the quantization coefficients.
[0217] On the other hand, the transformation kernel for the current block can be determined from K transformation kernels using the method described above. Here, K is an integer of 2 or greater.
[0218] On the other hand, the above-described method for determining the transform kernel can be performed even when the prediction mode of the current block is the inter-frame prediction mode.
[0219] On the other hand, the transform kernel determination method described above can be performed in the encoder / decoder. After the transform is performed based on the transform kernel determined in the encoder, no signaling for the transform kernel is required. In the decoder, the inverse transform is also performed based on the determined transform kernel.
[0220] Alternatively, the transform kernel can be determined based on syntax elements explicitly signaled. Specifically, when determining the transform kernel in the encoder, the syntax for the determined transform set can be signaled to the decoder. The transform kernel for the inverse transform can then be determined in the decoder based on the syntax elements.
[0221] On the other hand, the above-described method for determining the transform kernel can be performed even when the prediction mode of the current block is the inter-frame prediction mode.
[0222] According to an embodiment of the invention, the transform kernel can be determined as the transform kernel with the minimum cost value among K transform kernels included in the transform set. Here, K is an integer of 2 or greater.
[0223] Specifically, the RD cost can be calculated for each transform core in the encoder, and the core with the least distortion can be selected as the optimal transform core.
[0224] Additionally, information about the optimal transform kernel can be signaled to the decoder. Here, this information can be a flag when K is 2, and a kernel index when K is 3 or higher. In the decoder, the transform kernel can be determined based on this information, and the inverse transform can be performed based on the determined transform kernel.
[0225] On the other hand, according to the above implementation scheme, the transform kernel is determined after the transform set is determined. However, this is just an example, and when a separate transform set does not exist, the transform kernel of the current block can be determined without knowing the transform set. In this case, the transform kernel determination method can be performed in the same manner as the transform kernel determination method according to the above implementation scheme.
[0226] On the other hand, according to the above implementation scheme, the transformation kernel is determined after the transformation set is determined, but the determination of the transformation set and the determination of the transformation kernel can be performed simultaneously.
[0227] Figure 6 This is a schematic diagram illustrating a method for determining transform sets and transform kernels under a Spatial Geometric Partitioning Mode (SGPM) according to an embodiment of the present invention. Here, SGPM can refer to a mode that partitions a block into two regions by a straight partition boundary and performs intra-prediction. Specifically, according to SGPM, intra-prediction is performed separately for each partitioned region to generate an intra-prediction block for each region, and the intra-prediction block of the current block can be generated by calculating a weighted sum of the intra-prediction blocks.
[0228] refer to Figure 6When the prediction mode of the current block 600 is SGPM, the current block is separated into two regions based on the geometric partition line 601. Here, each of the separated regions can be referred to as sub-partition 0 (602) and sub-partition 1 (603). Prediction can be performed separately for each sub-partition based on each intra-prediction mode. Specifically, a prediction block can be generated based on the intra-prediction mode of sub-partition 0, and a prediction block can be generated based on the intra-prediction mode of sub-partition 1. In addition, a weighted sum can be performed on each prediction block to generate the final intra-prediction block of the current block.
[0229] According to an embodiment of the present invention, when the prediction mode of the current block is SGPM, the transform set and transform kernel for the transform of the current block can be determined based on the intra-prediction mode of each sub-partition.
[0230] For example, at least one of the transform set and transform kernel of the current block can be determined by utilizing intra-prediction modes that have a smaller number of allocated prediction modes between the intra-prediction modes of sub-partition 0 and the intra-prediction modes of sub-partition 1.
[0231] As another example, at least one of the transform set and transform kernel of the current block can be determined by utilizing the intra-prediction mode with the smaller cost value between the intra-prediction mode of sub-partition 0 and the intra-prediction mode of sub-partition 1. Specifically, the RD cost can be calculated based on the result of performing each intra-prediction mode, and an intra-prediction mode with smaller distortion can be used.
[0232] As another example, prediction can be performed on either the template adjacent to the current block or the template corresponding to the current block using the intra prediction modes of sub-partition 0 and sub-partition 1, and a cost value can be calculated. Specifically, the RD cost can be calculated based on the result of performing each intra prediction mode. Additionally, at least one transform set and transform kernel of the current block can be determined using intra prediction modes with lower distortion. Here, adjacent templates can refer to any region adjacent to the current block, and the template corresponding to the current block can refer to any region of the same size as the template adjacent to the current block.
[0233] On the other hand, in the above example, the method of determining the transform set by utilizing intra-prediction modes can be performed in the same manner as mapping in the transform of intra-blocks. For example, mapping can be performed in the same way as LFNST and NSPT.
[0234] On the other hand, in the example above, when determining the transform kernel by utilizing the intra-prediction mode, the transform kernel can be mapped to the intra-prediction mode using any method. Alternatively, an arbitrary lookup table can be defined to map the transform kernel to the intra-prediction mode.
[0235] On the other hand, under SGPM, sub-partitions and sub-blocks can have the same meaning.
[0236] According to an embodiment of the present invention, even when the prediction mode of the current block is SGPM, either of the above-described transform set determination method and transform kernel determination method can be performed, instead of determining the transform set and transform kernel based on the intra-prediction mode of each sub-partition.
[0237] As an example, at least one transform set and transform kernel can be determined based on the directional information of pixels in a whole or part of a region of either a template adjacent to the current block or a template corresponding to the current block.
[0238] As another example, in the case of prediction blocks under SGPM, due to the high prediction accuracy, the characteristics of the residuals and the characteristics of the first-order transform coefficients generated by the first-order transform can be similar to those of prediction blocks in inter-frame prediction mode. Therefore, the transform set used for the transform / inverse transform of inter-frame blocks can be used for blocks under SGPM. As an example, when the prediction mode of the current block is spatial geometry mode, the transform set of the current block can be determined as the third transform set mentioned above.
[0239] As another example, when the prediction mode of the current block is SGPM, the transform / inverse transform can be performed in the same way as the transform / inverse transform method for intra-blocks. That is, the transform set of an intra-block can be used for blocks under SGPM. When the prediction mode of the current block is SGPM, the transform set of the current block can be determined as the second transform set described above. Alternatively, the transform set can be determined as the first transform set described above.
[0240] As another example, when the prediction mode for the current block is geometric partitioning, the transform set for the current block can be determined as either the transform set for intra-frame blocks (intra-transform set) or the transform set for inter-frame blocks (inter-transform set). Specifically, the transform set for the current block can be determined as the transform set with the smaller cost value between the intra-frame and inter-frame transform sets. In this case, in addition to the information normally signaled during the transform, information about the optimal transform set is also signaled to the decoder.
[0241] The following describes a method for determining a transform set and transform kernel in a Spatial Combining Intra-Intra Prediction (SCIIP) mode according to an embodiment of the present invention. Here, SCIIP can refer to a mode that performs intra prediction based on two different intra prediction modes.
[0242] Specifically, according to SCIIP, two prediction blocks can be generated based on different intra-prediction modes. Furthermore, the generated prediction blocks can be weighted and summed to generate the final intra-prediction block for the current block. Here, the different intra-prediction modes can be referred to as the first intra-prediction mode and the second intra-prediction mode, respectively.
[0243] According to the implementation plan, when the prediction mode of the current block is SCIIP, the transform set and transform kernel for the current block can be determined based on the first intra-frame prediction mode and the second intra-frame prediction mode.
[0244] For example, at least one of the transform set and transform kernel of the current block can be determined by utilizing intra-prediction modes that have a smaller number of allocated prediction modes between the first intra-prediction mode and the second intra-prediction mode.
[0245] As another example, at least one of the transform set and transform kernel of the current block can be determined by utilizing an intra-prediction mode that has a smaller cost value between the first intra-prediction mode and the second intra-prediction mode.
[0246] On the other hand, the method of determining the transform set by utilizing intra-prediction modes can be performed in the same manner as the mapping in the transform of intra-blocks. For example, the method can be performed in the same manner as the mapping between intra-prediction modes and transform sets in LFNST and NSPT.
[0247] On the other hand, the intra-prediction modes used to determine the transform kernel based on the first and second intra-prediction modes can be mapped to the transform kernel using any method. Alternatively, an arbitrary lookup table can be defined to map the transform kernel to the intra-prediction modes.
[0248] According to an embodiment of the present invention, even when the prediction mode of the current block is SCIIP, either of the above-described transform set determination method and transform kernel determination method can be performed, instead of determining the transform set and transform kernel based on the intra-prediction mode of each sub-partition.
[0249] As an example, at least one transform set and transform kernel can be determined based on the directional information of pixels in a whole or part of a region of either a template adjacent to the current block or a template corresponding to the current block.
[0250] As another example, in the case of a prediction block under SCIIP, due to the high prediction accuracy, the characteristics of the residuals and the characteristics of the first-order transform coefficients generated by the first-order transform can be similar to those of the prediction block in the inter-frame prediction mode. Therefore, the transform set of the inter-frame block can be used for the block under SCIIP. As an example, when the prediction mode of the current block is SCIIP, the transform set of the current block can be determined as the third transform set mentioned above.
[0251] As another example, when the prediction mode of the current block is SCIIP, the transform / inverse transform can be performed in the same manner as the transform / inverse transform method for intra-blocks. That is, the transform set of an intra-block can be used for blocks under SCIIP. When the prediction mode of the current block is SCIIP, the transform set of the current block can be determined as the second transform set described above. Alternatively, the transform set can be determined as the first transform set described above.
[0252] As another example, when the prediction mode for the current block is SCIIP, the transform set for the current block can be determined as either the transform set for intra-blocks (hereinafter referred to as the intra-transform set) or the transform set for the transform / inverse transform of inter-blocks (hereinafter referred to as the inter-transform set). Specifically, the transform set for the current block can be determined as the transform set with the smaller cost value between the intra-transform set and the inter-transform set. In this case, in addition to the information normally signaled during the transform, information about the optimal transform set is also signaled to the decoder.
[0253] Figure 7 This is a flowchart describing a method for determining a transformation set according to an embodiment of the present invention. Figure 7 The method for determining the transform set can be performed by an image decoding device.
[0254] The image decoding device can deduce the prediction mode of the current block (S700).
[0255] In addition, the image decoding device can determine the transform set of the current block based on the prediction mode of the current block (S710).
[0256] On the other hand, the transform set can be determined as any one of a first transform set for intra-frame prediction, a second transform set for intra-frame prediction, and a third transform set for inter-frame prediction.
[0257] On the other hand, when the prediction mode is intra-frame template matching mode, the transform set can be determined as the third transform set.
[0258] On the other hand, a transformation set can be a transformation set that is indivisible by a single transformation.
[0259] On the other hand, when the prediction mode is intra-block copy mode, the transform set can be determined as the third transform set.
[0260] On the other hand, the first intra-frame prediction may include non-directional intra-frame prediction and directional intra-frame prediction, and the second intra-frame prediction may include SGPM intra-frame prediction.
[0261] On the other hand, when the current block is partitioned into a first sub-block and a second sub-block according to the partition boundary, the transformation set can be determined as the second transformation set.
[0262] In addition, the image decoding device can perform an inverse transform on the current block based on the transform set (S720).
[0263] On the other hand, performing an inverse transform may include: determining the transform kernel of the current block from among the transform kernels included in the transform set, and performing an inverse transform based on the transform kernel.
[0264] On the other hand, determining the transform kernel may include: deriving the transform kernel to determine the intra-prediction mode, and determining the transform kernel based on the intra-prediction mode determined by the transform kernel.
[0265] On the other hand, the intra-frame prediction mode can be determined by utilizing the directional information of some pixels included in the prediction block of the current block to determine the transform kernel.
[0266] On the other hand, the transform kernel can be the transform kernel with the least distortion among the transform kernels included in the transform set.
[0267] On the other hand, it further includes: when the prediction mode of the current block is either intra template matching mode or intra block copy mode, obtaining intra prediction mode utilization information indicating whether to use intra prediction mode to determine the transform kernel, and when the intra prediction mode utilization information indicates that intra prediction mode is used, determining the transform kernel may include: deriving the transform kernel to determine the intra prediction mode, and determining the transform kernel based on the intra prediction mode determined by the transform kernel.
[0268] On the other hand, when the current block is partitioned into a first sub-block and a second sub-block according to the partition boundary, the transform kernel can be determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.
[0269] On the other hand, it further includes: when the current block is partitioned into a first sub-block and a second sub-block according to the partition boundary, obtaining sub-block prediction mode utilization information that indicates whether to use the prediction mode of the first sub-block and the prediction mode of the second sub-block to determine the transform kernel, and when the sub-block prediction mode utilization information indicates that the prediction mode of the first sub-block and the prediction mode of the second sub-block are used, the transform kernel can be determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.
[0270] On the other hand, when the prediction mode of the current block is SCIIP mode, the transform kernel of the current block can be determined based on the first intra-frame prediction mode and the second intra-frame prediction mode of SCIIP mode.
[0271] On the other hand, a similar approach can be implemented in image encoding methods. Figure 7 The steps described herein. Additionally, it can be achieved by including... Figure 7The image encoding method described in the document generates a bitstream. The bitstream can be stored on a non-volatile computer-readable recording medium and can also be transmitted (or streamed).
[0272] Figure 8 An exemplary content streaming system applicable to embodiments of the present invention is shown.
[0273] like Figure 8 As shown, the content streaming system implementing the present invention can mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0274] The encoding server compresses content received from multimedia input devices such as smartphones, cameras, and CCTV into digital data to generate a bitstream, which is then sent to the streaming server. As another example, if multimedia input devices such as smartphones, cameras, and CCTV directly generate the bitstream, the encoding server can be omitted.
[0275] The bitstream can be generated by the image encoding method and / or image encoding apparatus of the present invention, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0276] A streaming server sends multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to notify the user of any available services. When a user requests a desired service from the web server, the web server sends it to the streaming server, and the streaming server can send multimedia data to the user. In this case, the content streaming system may include a separate control server, which can control the commands / responses between devices within the content streaming system.
[0277] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a period of time.
[0278] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital televisions, desktop computers, digital signage, etc.
[0279] In the above-mentioned content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed and processed.
[0280] The above embodiments can be performed in the same or corresponding manner as the encoding and decoding devices. Furthermore, at least one or a combination of the above embodiments can be used to encode / decode images.
[0281] The order in which the above-described embodiments are applied can differ between the encoding and decoding devices. Alternatively, the order in which the above-described embodiments are applied can be the same between the encoding and decoding devices.
[0282] The above-described implementation scheme can be performed on each of the luminance and chrominance signals. Alternatively, the above-described implementation scheme can be performed on both the luminance and chrominance signals in the same way.
[0283] In the above embodiments, the method is described based on a flowchart having a series of steps or units. However, the present invention is not limited to the order of the steps; on the contrary, some steps may be performed simultaneously with other steps or in a different order. Furthermore, those skilled in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps may be added to the flowchart or some steps may be deleted from the flowchart without affecting the scope of the present invention.
[0284] The implementation scheme can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include individual program instructions, data files, data structures, or combinations thereof. The program instructions recorded in the computer-readable recording medium may be specifically designed and constructed for this invention or may be well known to those skilled in the art of computer software.
[0285] The bitstream generated by the encoding method according to the above embodiment can be stored in a non-volatile computer-readable recording medium. Furthermore, the bitstream stored in the non-volatile computer-readable recording medium can be decoded by the decoding method according to the above embodiment.
[0286] Examples of computer-readable recording media include: magnetic recording media such as hard disks, floppy disks, and magnetic tapes; optical data storage media such as CD-ROMs or DVD-ROMs; magneto-optical media such as magneto-optical disks; and hardware devices such as read-only memory (ROM), random access memory (RAM), flash memory, etc., specifically configured to store and execute program instructions. Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code that can be implemented by a computer using an interpreter. The hardware device may be configured to operate by one or more software modules or vice versa to perform the processes according to the invention.
[0287] While the invention has been described with respect to specific items such as detailed elements, as well as limited embodiments and drawings, these are provided only to aid in a more complete understanding of the invention, and the invention is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and changes can be made based on the above description.
[0288] Therefore, the spirit of the present invention should not be limited to the above embodiments, and the entire scope of the appended claims and their equivalents shall fall within the scope and spirit of the present invention.
[0289] Industrial applicability This invention can be used in apparatuses for encoding / decoding images and in recording media for storing bit streams.
Claims
1. An image decoding method, comprising: Derive the prediction pattern for the current block; The transformation set of the current block is determined based on the prediction pattern; as well as Perform an inverse transform on the current block based on the transform set. The transform set is determined to be any one of the first transform set for intra-frame prediction, the second transform set for intra-frame prediction, and the third transform set for inter-frame prediction.
2. The image decoding method according to claim 1, wherein, When the prediction mode is intra-frame template matching mode, the transform set is determined as the third transform set.
3. The image decoding method according to claim 1, wherein, The transformation set is a set of transformations that cannot be separated into a single transformation.
4. The image decoding method according to claim 1, wherein, When the prediction mode is intra-block copy mode, the transform set is determined to be the third transform set.
5. The image decoding method according to claim 1, wherein, The first intra-frame prediction includes non-directional intra-frame prediction and directional intra-frame prediction; and The second intra-frame prediction includes Spatial Geometric Prediction Mode (SGPM) intra-frame prediction.
6. The image decoding method according to claim 1, wherein, Performing the inverse transform includes: Determine the transformation kernel of the current block from the transformation kernels included in the transformation set; and Inverse transformation is performed based on the transform kernel.
7. The image decoding method according to claim 6, wherein, Determining the transform kernel includes: Derivation of the transform kernel to determine the intra-frame prediction mode; and The transform kernel is determined based on the intra-frame prediction mode.
8. The image decoding method according to claim 7, wherein, The transform kernel is derived to determine the intra-frame prediction mode based on the directional information of some pixels included in the prediction block of the current block.
9. The image decoding method according to claim 6, wherein, The transform kernel is the transform kernel with the least distortion among the transform kernels included in the transform set.
10. The image decoding method according to claim 6, further comprising: When the prediction mode of the current block is either intra template matching mode or intra block copy mode, obtain information indicating whether to use the intra prediction mode to determine the intra prediction mode utilization of the transform kernel. When the intra-prediction mode uses information to indicate the use of the intra-prediction mode, determining the transform kernel includes: Derivation of the transform kernel to determine the intra-frame prediction mode; and The transform kernel is determined based on the intra-frame prediction mode.
11. The image decoding method according to claim 1, wherein, When the current block is partitioned into a first sub-block and a second sub-block based on the partition boundaries, the transform set is determined as the second transform set.
12. The image decoding method according to claim 6, wherein, When the current block is partitioned into a first sub-block and a second sub-block according to the partition boundary, the transform kernel is determined based on the prediction patterns of the first sub-block and the second sub-block.
13. The image decoding method according to claim 6, further comprising: When the current block is partitioned into a first sub-block and a second sub-block based on the partition boundaries, information is obtained indicating whether to use the prediction modes of the first and second sub-blocks to determine the sub-block prediction modes of the transform kernel. Specifically, when the sub-block prediction mode uses information to indicate the use of the prediction mode of the first sub-block and the prediction mode of the second sub-block, the transformation kernel is determined based on the prediction mode of the first sub-block and the prediction mode of the second sub-block.
14. The image decoding method according to claim 6, wherein, When the prediction mode of the current block is Spatial Combination Intra-Intra Prediction (SCIIP) mode, the transform kernel of the current block is determined based on the first intra-prediction mode and the second intra-prediction mode of SCIIP mode.
15. An image encoding method, comprising: Determine the prediction mode for the current block; The transformation set of the current block is determined based on the prediction pattern; as well as Perform a transformation on the current block based on the transform set. The transform set is determined to be any one of the first transform set for intra-frame prediction, the second transform set for intra-frame prediction, and the third transform set for inter-frame prediction.
16. A non-volatile computer-readable recording medium for storing a bitstream, said bitstream being generated by an image encoding method, in, The image encoding method includes: Determine the prediction mode for the current block; Determine the transformation set of the current block based on the prediction pattern; and Perform a transformation on the current block based on the transform set. The transform set is determined to be any one of the first transform set for intra-frame prediction, the second transform set for intra-frame prediction, and the third transform set for inter-frame prediction.
17. A method for transmitting a bitstream generated by an image encoding method, the method comprising transmitting the bitstream, in, The image encoding method includes: Determine the prediction mode for the current block; Determine the transformation set of the current block based on the prediction pattern; and Perform a transformation on the current block based on the transform set. The transform set is determined to be any one of the first transform set for intra-frame prediction, the second transform set for intra-frame prediction, and the third transform set for inter-frame prediction.