Image encoding / decoding method based on inseparable transformation, method for transmitting bitstream, and recording medium for storing bitstream

By adopting an inseparable one-time transformation method in image encoding and decoding, combining intra prediction mode and mixed prediction mode, the problem of low high-resolution image encoding efficiency is solved, and more efficient image data transmission and storage is achieved.

CN120345253APending Publication Date: 2025-07-18LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085242.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-12
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the encoding and decoding process of high resolution and high quality images, the prior art has problems such as low encoding efficiency and high information volume, resulting in high transmission and storage costs.

Method used

The inseparable one-time transformation method is adopted, combining the intra prediction mode or mixed prediction mode of the matrix, the prediction mode of the image block is adaptively processed and the residual samples are transformed.

Benefits of technology

The efficiency of image encoding and decoding is improved, the amount of bits required for transmission and storage is reduced, and the storage and transmission process of image data is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345253A_ABST
    Figure CN120345253A_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, a method for transmitting a bitstream, and a computer-readable recording medium for storing the bitstream are provided. An image decoding method according to the present disclosure is an image decoding method performed by an image decoding apparatus, the method comprising the steps of: determining a prediction mode of a current block; determining whether the prediction mode of the current block is a predetermined prediction mode; and transforming a residual sample for the current block, in which the predetermined prediction mode includes at least one of a matrix-based intra prediction (MIP) mode or a hybrid prediction mode, and the hybrid prediction mode is a mode in which a prediction block is derived based on surrounding samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for encoding / decoding an image, a method for transmitting a bitstream, and a recording medium for storing a bitstream, and relates to a method for adaptively performing an inseparable primary transform. Background Art

[0002] In recent years, the need for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images has been increasing in various fields. As image data becomes high-resolution and high-quality, the amount of information or bitrate to be transmitted increases compared to conventional image data. The increase in the amount of information or bitrate to be transmitted leads to an increase in transmission costs and storage costs.

[0003] Therefore, there is a need for an efficient image compression technique to effectively transmit, store, and reproduce information of high-resolution and high-quality images. Summary of the Invention

[0004] Technical Problem

[0005] The present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.

[0006] In addition, the present disclosure is to provide a method for configuring a relationship between an inseparable primary transform and other encoding tools.

[0007] In addition, the present disclosure is to provide a method for not applying an inseparable primary transform for a predetermined prediction mode.

[0008] In addition, the present disclosure is to provide a method for adaptively applying an inseparable primary transform for a predetermined prediction mode.

[0009] In addition, the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method according to the present disclosure.

[0010] In addition, the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream received and decoded by an image decoding apparatus according to the present disclosure and for image reconstruction.

[0011] In addition, the present disclosure is to provide a method for transmitting a bitstream generated by an image encoding method according to the present disclosure.

[0012] The technical problems to be achieved by the present disclosure are not limited to the above technical problems, and other technical problems not mentioned can be clearly understood by those of ordinary skill in the art from the following description.

[0013] Technical Solution

[0014] An image decoding method according to an aspect of the present disclosure is an image decoding method performed by an image decoding apparatus, and may be an image decoding method including the following steps: determining a prediction mode of a current block; determining whether the prediction mode of the current block is a predetermined prediction mode; and transforming residual samples for the current block, wherein the predetermined prediction mode includes at least one of a matrix-based intra prediction (MIP) mode or a hybrid prediction mode, and the hybrid prediction mode is a mode of deriving a prediction block based on neighboring samples.

[0015] An image encoding method according to another aspect of the present disclosure is an image encoding method performed by an image encoding apparatus, and may be an image encoding method including the following steps: determining a prediction mode of a current block; determining whether the prediction mode of the current block is a predetermined prediction mode; and transforming residual samples for the current block, wherein the predetermined prediction mode includes at least one of a matrix-based intra prediction (MIP) mode or a hybrid prediction mode, and the hybrid prediction mode is a mode of deriving a prediction block based on neighboring samples.

[0016] A computer-readable recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0017] A transmission method according to another aspect of the present disclosure may transmit a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0018] The features briefly outlined above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure below and do not limit the scope of the present disclosure.

[0019] Advantageous Effects

[0020] According to the present disclosure, an image encoding / decoding method and apparatus having improved encoding / decoding efficiency can be provided.

[0021] In addition, according to the present disclosure, an inseparable transform can be utilized when performing a transform once.

[0022] In addition, according to the present disclosure, the amount of bits required to signal information about an inseparable single transform can be reduced, thereby improving bit efficiency.

[0023] In addition, according to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream generated by the image encoding method according to the present disclosure can be provided.

[0024] According to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction can be provided.

[0025] According to the present disclosure, a method for transmitting a bitstream generated by an image encoding method can be provided.

[0026] The effects obtainable from the present disclosure are not limited to the above effects, and other effects not described can be clearly understood by those of ordinary skill in the art from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 A schematic diagram showing a video coding system to which an embodiment according to the present disclosure can be applied is shown.

[0028] Figure 2 A schematic diagram showing an image encoding device to which an embodiment according to the present disclosure can be applied is shown.

[0029] Figure 3 A schematic diagram showing an image decoding device to which an embodiment according to the present disclosure can be applied is shown.

[0030] Figure 4 It is a diagram for describing an LFNST application method.

[0031] Figure 5 It is a flowchart showing a method for encoding an image based on intra prediction.

[0032] Figure 6 It is a diagram schematically showing an intra predictor of an image encoding device.

[0033] Figure 7 It is a flowchart showing a method for decoding an image based on intra prediction.

[0034] Figure 8 It is a diagram schematically showing an intra predictor of an image encoding device.

[0035] Figure 9 It is a diagram for describing template-based intra mode derivation (TIMD).

[0036] Figure 10 It is a diagram for describing a method for constructing a gradient histogram (HoG) in decoder-side intra mode derivation (DIMD).

[0037] Figure 11 It is a diagram for describing a method for constructing a prediction block when applying DIMD.

[0038] Figure 12 It is a diagram for describing a method for applying DIMD to a chrominance block.

[0039] Figure 13 and Figure 14 It is a diagram for describing combined inter and intra prediction (CIIP).

[0040] Figure 15 is a flowchart showing an image encoding / decoding method according to an embodiment of the present disclosure.

[0041] Figure 16 is a flowchart showing an image encoding / decoding method according to another embodiment of the present disclosure.

[0042] Figure 17 is a flowchart showing an image encoding / decoding method according to an embodiment of the present disclosure.

[0043] Figure 18 shows an example diagram of a content streaming system to which an embodiment of the present disclosure can be applied. Detailed Embodiments

[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those of ordinary skill in the art can easily implement them. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.

[0045] When describing embodiments of the present disclosure, when a detailed description of a well-known configuration or function is considered to obscure the gist of the present disclosure, its detailed description is omitted. Additionally, parts irrelevant to the description of the present disclosure are omitted from the drawings, and similar reference numerals are assigned to similar components.

[0046] In the present disclosure, when describing that a certain component is "connected", "coupled", or "linked" to another component, this includes not only direct connection but also indirect connection where another component may exist in between. Additionally, when describing that a certain component "includes" or "has" another component, unless otherwise clearly stated, this means that other components are not excluded but additional components may be further included.

[0047] In the present disclosure, terms such as first and second are only for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless otherwise clearly stated. Therefore, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0048] In the present disclosure, distinguishable components are described to clearly illustrate their corresponding characteristics and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, embodiments of such integration or distribution are also included in the scope of the present disclosure without being explicitly described.

[0049] In the present disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Additionally, embodiments including additional components in addition to the components described in each embodiment are also included within the scope of the present disclosure.

[0050] The present disclosure relates to encoding and decoding of images, and unless otherwise defined in the present disclosure, the terms used herein may have the ordinary meanings commonly used in the technical field to which the present disclosure pertains.

[0051] In the present disclosure, a "picture" generally refers to a unit representing a single image at a specific point in time. A slice / tile is an encoding unit that forms part of a picture, and a picture may be composed of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs).

[0052] In the present disclosure, a "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, the term "sample" may be used as a corresponding term for a pixel. A sample generally may represent a pixel or the value of a pixel, and may indicate only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0053] In the present disclosure, a "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific area of a picture or information related to the area. Depending on the context, the term "unit" may be used interchangeably with "sample array", "block", "region", etc. Generally, an M×N block may include a set (or array) of samples (or a sample array) or a set (or array) of transform coefficients composed of M columns and N rows.

[0054] In the present disclosure, the term "current block" may refer to one of "current encoding block", "current encoding unit", "encoding target block", "decoding target block", or "processing target block". When prediction is performed, the "current block" may refer to "current prediction block" or "prediction target block". When transform (inverse transform) / quantization (dequantization) is performed, the "current block" may refer to "current transform block" or "transform target block". When filtering is performed, the "current block" may refer to "filtering target block".

[0055] In the present disclosure, unless clearly stated as a chrominance block, the term "current block" may refer to a block including both a luminance component block and a chrominance component block, or may refer to the "luminance block of the current block". The luminance component block of the current block can be clearly denoted by terms such as "luminance block" or "current luminance block", which clearly indicate that it is a luminance component block. Additionally, the chrominance component block of the current block can be clearly denoted by terms such as "chrominance block" or "current chrominance block", which clearly indicate that it is a chrominance component block.

[0056] In the present disclosure, " / " and "," may refer to "and / or". For example, "A / B" and "A,B" may refer to "A and / or B". Additionally, "A / B / C" and "A,B,C" may refer to "at least one of A, B, and / or C".

[0057] In the present disclosure, "or" may refer to "and / or". For example, "A or B" may mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" may also mean "additionally or alternatively".

[0058] Overview of the video coding system

[0059] Figure 1 A schematic diagram showing a video coding system to which embodiments according to the present disclosure can be applied is shown.

[0060] The video coding system according to an embodiment may include an encoder device 10 and a decoder device 20. The encoder device 10 may send encoded video and / or image information or data to the decoder device 20 in the form of a file or a stream via a digital storage medium or a network.

[0061] The encoder device 10 according to an embodiment may include a video source generator 11, an encoder 12, and a transmitter 13. The decoder device 20 according to an embodiment may include a receiver 21, a decoder 22, and a renderer 23. The encoder 12 may be referred to as a video / image encoder, and the decoder 22 may be referred to as a video / image decoder. The transmitter 13 may be included in the encoder 12. The receiver 21 may be included in the decoder 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0062] The video source generator 11 can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source generator 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device includes, for example, a computer, a tablet, or a smart phone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer or the like, and in this case, the video / image capture process can be replaced by a process of generating relevant data.

[0063] The encoder 12 can encode the input video / images. The encoder 12 can perform a series of processes such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoder 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0064] The transmitter 13 can obtain the encoded video / image information or data output in the form of a bitstream and send it to the receiver 21 of the decoder device 20 or another external object in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include elements for generating a media file in a predetermined file format and elements for transmitting through a broadcast / communication network. The transmitter 13 can be provided as a transmission device separate from the encoding device 10. In this case, the transmission device can include at least one processor for obtaining the encoded video / image information or data in the form of a bitstream and a transmitter for delivering the data in the form of a file or a stream. The receiver 21 can extract / receive the bitstream from the storage medium or the network and send it to the decoder 22.

[0065] The decoder 22 can decode the video / images by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operations of the encoder 12.

[0066] The renderer 23 can render the decoded video / images. The rendered video / images can be displayed through a display unit.

[0067] Overview of the image coding device

[0068] Figure 2 A schematic diagram showing an image encoding device to which embodiments according to the present disclosure can be applied is shown.

[0069] As Figure 2As described above, the image encoding device 100 may include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a "predictor". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include the subtractor 115.

[0070] All or at least some of the plurality of components constituting the image encoding device 100 may be implemented as a single hardware component (i.e., an encoder or a processor) according to an embodiment. Additionally, the memory 170 may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium.

[0071] The image splitter 110 may split an input image (or picture, frame) input to the image encoding device 100 into at least one processing unit. As an example, the processing unit may be referred to as a coding unit (CU). The coding unit may be obtained by recursively splitting a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree, binary tree, or ternary tree (QT / BT / TT) structure. For example, the coding unit may be divided into coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. To split the coding unit, a quadtree structure may be first applied, and then a binary tree structure and / or a ternary tree structure may be applied. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer further split. The largest coding unit may be directly used as the final coding unit, or a coding unit of a deeper depth obtained by splitting the largest coding unit may be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and / or reconstruction described later. As another example, the processing unit for the encoding process may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or split from the final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or deriving a residual signal from the transformation coefficients.

[0072] The predictor (inter-frame predictor 180 or intra-frame predictor 185) may perform prediction on a target block (current block) and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction to the current block or a coding unit (CU). The predictor may generate various information related to the prediction of the current block and send it to the entropy encoder 190. The prediction-related information may be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0073] The intra predictor 185 can predict the current block by referring to samples within the current picture. The samples referred to can be located in the adjacent region of the current block, or can be located at a farther position according to the intra prediction mode and / or intra prediction method. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used according to the configuration. The intra predictor 185 can also determine the prediction mode applied to the current block by using the prediction mode applied to the adjacent block.

[0074] The inter predictor 180 can derive the prediction block of the current block based on the reference block (reference sample array) specified by the motion vector on the reference picture. To reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted at the block, sub-block, or sample level based on the correlation of the motion information between the adjacent block and the current block. The motion information can include the motion vector and the reference picture index. The motion information can also include information about the inter prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter prediction, the adjacent block can include the spatially adjacent block present in the current picture and the temporally adjacent block present in the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent block can be the same or different. The temporally adjacent block can be referred to as the collocated reference block or the collocated coding unit (colCU), and the reference picture including the temporally adjacent block can be referred to as the collocated picture (colPic). For example, the inter predictor 180 can construct a motion information candidate list based on the adjacent block and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. The inter prediction can be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter predictor 180 can use the motion information of the adjacent block as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be sent. In the motion vector prediction (MVP) mode, the motion vector of the adjacent block can be used as the motion vector predictor, and the motion vector of the current block can be signaled by encoding the motion vector difference and the indicator for the motion vector predictor. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.

[0075] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the predictor may apply intra prediction or inter prediction for the prediction of the current block, and may also apply intra prediction and inter prediction simultaneously. The prediction method that applies intra prediction and inter prediction simultaneously for the prediction of the current block may be referred to as combined intra-inter prediction (CIIP). Additionally, the predictor may perform intra-block copy (IBC) for the prediction of the current block. For example, intra-block copy can be used in applications such as game content image / video coding, such as screen content coding (SCC). IBC is a method of predicting the current block by using a pre-reconstructed reference block located within the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but since it derives the reference block within the current picture, it can operate similarly to inter prediction. In other words, IBC can use at least one of the inter prediction methods described in this disclosure.

[0076] The prediction signal generated by the predictor can be used to generate a reconstruction signal or generate a residual signal. The subtractor 115 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the predictor from the input image signal (original block, original sample array). The generated residual signal can be sent to the transformer 120.

[0077] The transformer 120 can generate transform coefficients by applying a transform method to the residual signal. For example, the transform method may include at least one of discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loeve transform (KLT), graph-based transform (GBT), or conditional non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. The transform process can be applied to pixel blocks of the same square size or non-square variable-size blocks.

[0078] The quantizer 130 can quantize the transform coefficients and send them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the block-shaped quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and can generate information about the quantized transform coefficients based on the one-dimensional vector of the quantized transform coefficients.

[0079] The entropy encoder 190 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can not only encode the quantized transform coefficients, but also encode the information required for video / image reconstruction (i.e., the values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (i.e., the encoded video / image information) can be sent or stored in the form of a bitstream in a network abstraction layer (NAL) unit. The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information can also include general constraint information. The signaling information, the transmitted information, and / or the syntax elements described in this disclosure can be encoded through the above encoding process and included in the bitstream.

[0080] The bitstream can be sent through a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for sending the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal can be provided as an internal / external element of the image encoding device 100, or the transmitter can be configured as a component of the entropy encoder 190.

[0081] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual samples) can be reconstructed by dequantizing and inverse-transforming the quantized transform coefficients via the dequantizer 140 and the inverse-transformer 150.

[0082] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-frame predictor 180 or the intra-frame predictor 185. When there is no residual for the target block (such as when the skip mode is applied), the predicted block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block within the current picture, and as will be described later, after filtering, it can also be used for inter-frame prediction of the next picture.

[0083] Filter 160 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, Filter 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in Memory 170, specifically in the DPB of Memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 160 may generate various filtering-related information as described later in the description of each filtering method and send it to Entropy Encoder 190. The filtering-related information may be encoded by Entropy Encoder 190 and output in the form of a bitstream.

[0084] The modified reconstructed picture sent to Memory 170 may be used as a reference picture in Inter-Frame Predictor 180. In this case, when inter-frame prediction is applied, Image Encoding Device 100 may avoid prediction mismatches between Image Encoding Device 100 and the Image Decoding Device, and may improve the encoding efficiency.

[0085] The DPB in Memory 170 may store the modified reconstructed picture to be used as a reference picture in Inter-Frame Predictor 180. Memory 170 may store the motion information of the blocks in the current picture for which motion information has been derived (or encoded) and / or the motion information of the blocks in the reconstructed image. The stored motion information may be sent to Inter-Frame Predictor 180 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and send them to Intra-Frame Predictor 185.

[0086] Overview of the image decoding device

[0087] Figure 3 A schematic diagram of an image decoding device to which embodiments according to the present disclosure may be applied is shown.

[0088] As Figure 3 shown, Image Decoding Device 200 may include Entropy Decoder 210, Dequantizer 220, Inverse Transformer 230, Adder 235, Filter 240, Memory 250, Inter-Frame Predictor 260, and Intra-Frame Predictor 265. Inter-Frame Predictor 260 and Intra-Frame Predictor 265 may be collectively referred to as "predictors". Dequantizer 220 and Inverse Transformer 230 may be included in the residual processor.

[0089] All or at least some of the multiple components constituting Image Decoding Device 200 may be implemented as a single hardware component (i.e., a decoder or a processor) according to embodiments. Additionally, Memory 170 may include a DPB and may be implemented by a digital storage medium.

[0090] Image Decoding Device 200 that receives a bitstream including video / image information may perform operations related to those Figure 2The processing corresponding to that performed by the image encoding device 100 in [ ] is performed to reconstruct the image. For example, the image decoding device 200 may perform decoding using the processing units applied in the image encoding device. Thus, the processing units for decoding may be, for example, encoding units. The encoding units may be coding tree units or may be obtained by splitting the largest coding unit. Additionally, the reconstructed image signal decoded and output by the image decoding device 200 may be played back by a playback device (not shown).

[0091] The image decoding device 200 may receive a signal output in the form of a bitstream from the Figure 2 image encoding device in [ ]. The received signal may be decoded by the entropy decoder 210. For example, the entropy decoder 210 may parse the bitstream to extract information required for image reconstruction (i.e., video / image information). The video / image information may also include information about various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may also include general constraint information. The image decoding device may additionally use the information about the parameter sets and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements described in the present disclosure may be obtained from the bitstream by being decoded through the decoding process. For example, the entropy decoder 210 may decode the information in the bitstream based on an encoding method such as Exponential Golomb, CAVLC, or CABAC, and may output the syntax element values required for image reconstruction and the quantization values of the transform coefficients related to the residuals. More specifically, the CABAC entropy decoding method may receive bins corresponding to the syntax elements in the bitstream, may determine a context model using the information of the decoding target syntax element, the decoding information of the decoding target block and adjacent blocks, or the information of the previously decoded symbols / bins, may predict the probability of the bin occurrence according to the determined context model, and may perform arithmetic decoding of the bin to generate symbols corresponding to each syntax element. In this case, the CABAC entropy decoding method may update the context model of the next symbol / bin using the decoded symbol / bin information after determining the context model. Among the decoding information from the entropy decoder 210, the prediction-related information may be provided to the predictors (the inter-frame predictor 260 and the intra-frame predictor 265), and the residual values entropy decoded by the entropy decoder 210 (in other words, the quantization transform coefficients and the related parameter information) may be input to the dequantizer 220. Additionally, among the decoding information from the entropy decoder 210, the filtering-related information may be provided to the filter 240. Furthermore, a receiver (not shown) that receives the signal output from the image encoding device may be additionally configured as an internal / external element of the image decoding device 200, or the receiver may be configured as a component of the entropy decoder 210.

[0092] In addition, the image decoding device according to the present disclosure may also be referred to as a video / image / picture decoding device. The image decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210, and the sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.

[0093] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients into a two-dimensional block. In this case, the rearrangement may be performed based on the coefficient scan order applied in the image encoding device. The dequantizer 220 may dequantize the quantized transform coefficients using a quantization parameter (i.e., quantization step information) and may obtain the transform coefficients.

[0094] The inverse transformer 230 may perform an inverse transform on the transform coefficients to obtain a residual signal (residual block or residual sample array).

[0095] The predictor may perform prediction on the current block and generate a prediction block including the prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction-related information output from the entropy decoder 210, and may determine a specific intra-frame / inter-frame prediction mode (prediction method).

[0096] The predictor may generate a prediction signal based on various prediction methods (techniques) described later, which is the same as the description of the predictor in the image encoding device 100.

[0097] The intra-frame predictor 265 may perform prediction on the current block by referring to the samples within the current picture. The description of the intra-frame predictor 185 may also be applied to the intra-frame predictor 265 in the same manner.

[0098] The inter - frame predictor 260 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, motion information can be predicted at the block, sub - block, or sample level based on the correlation of motion information between adjacent blocks and the current block. Motion information can include motion vectors and reference picture indices. Motion information can also include information about the inter - frame prediction direction (i.e., L0 prediction, L1 prediction, Bi - prediction, etc.). In inter - frame prediction, adjacent blocks can include spatially adjacent blocks in the current picture and temporally adjacent blocks in the reference picture. For example, the inter - frame predictor 260 can construct a candidate list of motion information based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter - frame prediction can be performed based on various prediction modes (methods), and prediction - related information can include information indicating the inter - frame prediction mode (method) applied to the current block.

[0099] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter - frame predictor 260 and / or the intra - frame predictor 265). When there is no residual for the target block (such as when the skip mode is applied), the prediction block can be used as the reconstructed block. The description of the adder 155 can also be applied to the adder 235 in the same way. The adder 235 can be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal can be used for intra - frame prediction of the next target block within the current picture and, as described later, can also be used for inter - frame prediction of the next picture after being filtered.

[0100] The filter 240 can apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 240 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture can be stored in the memory 250, specifically in the DPB of the memory 250. Various filtering methods can include, for example, de - blocking filtering, sample - adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0101] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter - frame predictor 260. The memory 250 can store the motion information of the blocks in the current picture for which motion information has been derived (or decoded) and / or the motion information of the blocks in the already - reconstructed images. The stored motion information can be sent to the inter - frame predictor 260 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 250 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra - frame predictor 265.

[0102] In this specification, the embodiments described for filter 160, inter - frame predictor 180, and intra - frame predictor 185 of image encoding device 100 may be applied to filter 240, inter - frame predictor 260, and intra - frame predictor 265 of image decoding device 200 in the same or corresponding manner.

[0103] Overview of the transform / inverse transform

[0104] As described above, the encoding apparatus may derive a residual block (residual samples) from a predicted block (predicted samples) predicted based on intra - frame / inter - frame / IBC prediction, etc., and may derive quantized transform coefficients by applying transform and quantization to the derived residual samples. Information about the quantized transform coefficients (residual information) may be included in the residual coding syntax and output in the form of a bitstream after encoding. The decoding apparatus may obtain information about (quantized) transform coefficients (residual information) from the bitstream and decode it to derive the quantized transform coefficients. The decoding apparatus may derive residual samples based on the quantized transform coefficients through de - quantization / inverse transform. As described above, at least one of the above quantization / de - quantization and / or transform / inverse transform may be omitted. When quantization / de - quantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for consistency of expression. Whether to omit transform / inverse transform may be signaled based on the transform_skip_flag.

[0105] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled through the residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by performing inverse transform (scaling) on the transform coefficients. The residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied to / stated in other parts of the present disclosure.

[0106] The transform / inverse transform may be performed based on a transform kernel. For example, according to the present disclosure, a multi - transform selection (MTS) scheme may be applied. In this case, some of a plurality of transform kernel sets may be selected and applied to the current block. The transform kernel may be referred to by various terms such as a transform matrix, a transform type, etc. For example, the transform kernel set may represent a combination of a vertical transform kernel (a vertical transform kernel) and a horizontal transform kernel (a horizontal transform kernel).

[0107] For example, in order to indicate one of the transform kernel sets, MTS index information (or mts_idx syntax element) may be generated / encoded in the encoding apparatus and signaled to the decoding apparatus. For example, according to the value of the MTS index information, the transform kernel set may be derived as shown in Table 1.

[0108] [Table 1]

[0109] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2

[0110] Table 1 shows the trTypeHor and trTypeVer values according to tu_mts_idx[x0][y0].

[0111] The transform kernel set can also be determined based on, for example, cu_sbt_horizontal_flag and cu_sbt_pos_flag as shown in Table 2.

[0112] [Table 2]

[0113] cu_sbt_horizontal_flag cu_sbt_pos_flag trTypeHor trTypeVer 0 0 2 1 0 1 1 1 1 0 1 2 1 1 1 1

[0114] Table 2 shows the trTypeHor and trTypeVer values according to cu_sbt_horizontal_flag and cu_sbt_pos_flag. Here, cu_sbt_horizontal_flag equal to 1 can indicate that the current coding unit is horizontally divided into two transform blocks. On the contrary, cu_sbt_horizontal_flag equal to 0 can indicate that the current coding unit is vertically divided into two transform blocks. In addition, cu_sbt_pos_flag equal to 1 can indicate that the syntax elements tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the first transform unit in the current coding unit do not exist in the bitstream. On the contrary, cu_sbt_pos_flag equal to 0 can indicate that the syntax elements tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the second transform unit in the current coding unit do not exist in the bitstream.

[0115] In addition, in Table 1 and Table 2, trTypeHor can represent the horizontal transform kernel, and trTypeVer can represent the vertical transform kernel. The value 0 of trTypeHor / trTypeVer can represent DCT2, the value 1 of trTypeHor / trTypeVer can represent DST7, and the value 2 of trTypeHor / trTypeVer can represent DCT8. However, this is an example, and other values can be mapped to other DCT / DST according to the convention.

[0116] Table 3 exemplarily represents the basis functions of the above-mentioned DCT2, DCT8, and DST7.

[0117] [Table 3]

[0118]

[0119] In the present disclosure, the MTS-based transform is applied as a primary transform, and a secondary transform may be further applied. The secondary transform may be applied only to the coefficients in the upper left w×h region of the coefficient block to which the primary transform has been applied, and may be referred to as a reduced secondary transform (RST). For example, w and / or h above may be 4 or 8. In the transform, the primary transform and the secondary transform may be sequentially applied to the residual block, and in the inverse transform, the inverse secondary transform and the inverse primary transform may be sequentially applied to the transform coefficients. The secondary transform (RST transform) may be referred to as a low-frequency coefficient transform (LFCT) or a low-frequency non-separable transform (LFNST). The inverse secondary transform may be referred to as the inverse LFCT or the inverse LFNST.

[0120] Figure 4 is a diagram for describing the method of applying LFNST.

[0121] Referring to Figure 4 , LFNST may be applied between the forward primary transform 411 and quantization 413 on the encoder side, and between the dequantization 421 and the inverse primary transform (or primary inverse transform) 423 on the decoder side.

[0122] In LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform may be (selectively) applied according to the block size. For example, a 4×4 LFNST may be applied to relatively small blocks (i.e., min(width, height) < 8), and an 8×8 LFNST may be applied to relatively large blocks (i.e., min(width, height) > 4). In Figure 4 , it is exemplarily shown that a 4×4 forward LFNST is applied to 16 input coefficients and an 8×8 forward LFNST is applied to 64 input coefficients. In addition, in Figure 4 , it is exemplarily shown that a 4×4 inverse LFNST may be applied to 8 input coefficients and an 8×8 inverse LFNST may be applied to 16 input coefficients.

[0123] In LFNST, a total of 4 transform sets and 2 non-separable transform matrices (kernels) for each transform set may be used. The mapping from the intra prediction mode to the transform set may be predefined as shown in Table 4.

[0124] [Table 4]

[0125] IntraPredMode Transform set index IntraPredMode < 0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode <= 80 1 81 <= IntraPredMode <= 83 0

[0126] Referring to Table 4, when three CCLM modes with prediction mode numbers from 81 to 83 (i.e., 81 ≤ IntraPredMode ≤ 83) are used for the current block, transform set 0 can be selected for the current chrominance block. For each transform set, the selected non-separable quadratic transform candidate can be additionally specified by an explicitly signaled LFNST index. The corresponding index can be signaled once in the bitstream for each intra CU after the transform coefficients.

[0127] In addition, the transform / inverse transform can be performed on a CU or TU basis. In other words, the transform / inverse transform can be applied to the residual samples within a CU or the residual samples within a TU. The CU size and TU size can be the same, or there can be multiple TUs within the CU region. Additionally, the CU size typically represents the luminance component (sample) CB size. The TU size typically represents the luminance component (sample) TB size. The chrominance component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio according to the color format (chrominance format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) of maxTbSize can be derived from the CU, and the transform / inverse transform can be performed on a TU (TB) basis. maxTbSize can be considered when determining whether to apply various intra prediction types such as ISP. Information about maxTbSize can be determined in advance, or can be generated and encoded in the encoding device and signaled to the decoding device.

[0128] As described above, the transform can be applied to the residual block. This is to decorrelate the residual block as much as possible, concentrate the coefficients on the low frequencies, and generate zero tails at the end of the block. In the JEM software, the transform part includes two main functions, namely the core transform and the quadratic transform. The core transform consists of the discrete cosine transform (DCT) and discrete sine transform (DST) transform families applied to all rows and columns of the residual block. Thereafter, the quadratic transform can be additionally applied to the upper left corner of the output of the core transform. Similarly, the inverse transform can be applied in the order of inverse quadratic transform and core inverse transform. Specifically, the inverse quadratic transform can be applied to the upper left corner of the coefficient block. Thereafter, the core inverse transform is applied to the rows and columns of the output of the inverse quadratic transform. The core transform / inverse transform can be referred to as the primary transform / inverse transform.

[0129] Overview of the intra prediction

[0130] Hereinafter, the intra prediction according to the present disclosure will be described.

[0131] Intra prediction may refer to a prediction method of generating prediction samples for a current block based on reference samples within a picture (hereinafter referred to as the current picture) to which the current block belongs. When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block may include: a total of 2×nH samples adjacent to the left boundary of the current block of size nW×nH and adjacent to the lower left of the current block, a total of 2×nW samples adjacent to the upper boundary of the current block and adjacent to the upper right of the current block, and one sample adjacent to the upper left of the current block. Alternatively, the adjacent reference samples of the current block may include multiple columns of upper adjacent samples and multiple rows of left adjacent samples. Additionally, the adjacent reference samples of the current block may include: a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right of the current block.

[0132] However, some adjacent reference samples of the current block may not have been decoded or may be unavailable. In this case, the image decoding device 200 may construct adjacent reference samples for prediction by replacing the unavailable samples with available samples. Alternatively, adjacent reference samples for prediction may be constructed by interpolating the available samples.

[0133] When the adjacent reference samples are derived, (i) prediction samples may be derived based on the average value or interpolation of the adjacent reference samples of the current block, and (ii) prediction samples may be derived based on the reference samples among the adjacent reference samples of the current block that are in a specific (prediction) direction of the prediction sample. Case (i) may be referred to as a non-directional mode or non-angle mode, and case (ii) may be referred to as a directional mode or angle mode.

[0134] Additionally, based on the prediction target sample of the current block among the adjacent reference samples, prediction samples may be generated by interpolation between a first adjacent sample in the prediction direction of the intra prediction mode of the current block and a second adjacent sample in the opposite direction. The above case may be referred to as linear interpolation intra prediction (LIP).

[0135] Additionally, chrominance prediction samples may be generated based on luminance samples using a linear model. This case may be referred to as the linear model (LM) mode.

[0136] Additionally, temporary prediction samples of the current block may be derived based on filtered adjacent reference samples, and prediction samples of the current block may be derived by calculating a weighted sum of at least one of the reference samples derived according to the intra prediction mode among the conventional adjacent reference samples (i.e., unfiltered adjacent reference samples) and the temporary prediction samples. This case is called position-dependent intra prediction (PDPC).

[0137] Additionally, a reference sample line with the highest prediction accuracy can be selected among multiple adjacent reference sample lines of the current block, and a prediction sample can be derived using the reference sample located in the prediction direction in the corresponding line. In this case, information about the used reference sample line (e.g., intra_luma_ref_idx) can be encoded and signaled in the bitstream. This case is referred to as multi-reference line intra prediction (MRL) or MRL-based intra prediction. When MRL is not applied, the reference sample can be derived from the reference sample line directly adjacent to the current block, and in this case, the information about the reference sample line may not be signaled.

[0138] Additionally, the current block can be divided into vertical sub-partitions or horizontal sub-partitions, and intra prediction can be performed for each sub-partition based on the same intra prediction mode. In this case, adjacent reference samples for intra prediction can be derived for each sub-partition unit. In other words, the reconstructed samples of the previous sub-partition in the coding / decoding order can be used as the adjacent reference samples of the current sub-partition. In this case, the intra prediction mode of the current block is applied identically to multiple sub-partitions, and adjacent reference samples are derived and used for each sub-partition unit, thereby improving the intra prediction performance in some cases. This prediction method is referred to as intra sub-partition (ISP) or ISP-based intra prediction.

[0139] The above intra prediction methods can be referred to by various terms, such as intra prediction type or additional intra prediction mode, to distinguish them from the directional or non-directional intra prediction modes. For example, the intra prediction method (i.e., intra prediction type or additional intra prediction mode, etc.) can include at least one of the above LIP, LM, PDPC, MRL, or ISP. Conventional intra prediction methods other than specific intra prediction types such as LIP, LM, PDPC, MRL, ISP, etc. can be referred to as normal intra prediction types. The normal intra prediction type is generally applied when the above specific intra prediction types are not used, and prediction can be performed based on the above intra prediction mode. Additionally, post-processing filtering can be performed on the derived prediction samples when necessary.

[0140] Specifically, the intra prediction process can include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. Additionally, a post-filtering step can be performed on the derived prediction samples when necessary.

[0141] In addition to the above intra prediction types, affine linear weighted intra prediction (ALWIP) can also be used. ALWIP can also be referred to as linear weighted intra prediction (LWIP), matrix weighted intra prediction, or matrix-based intra prediction (MIP). When applying MIP to the current block, the predicted samples of the current block can be derived through the following steps: i) using adjacent reference samples that have undergone an averaging process, ii) performing a matrix-vector multiplication process, and iii) further performing a horizontal / vertical interpolation process if necessary. The intra prediction mode for MIP can be different from the modes used in the above LIP, PDPC, MRL, ISP intra prediction, or normal intra prediction. The intra prediction mode for MIP can be referred to as the MIP intra prediction mode, the MIP prediction mode, or the MIP mode. For example, the matrix and offset used in the matrix-vector multiplication can be set differently according to the intra prediction mode of MIP. Here, the matrix can be referred to as the (MIP) weight matrix, and the offset can be referred to as the (MIP) offset vector or the (MIP) bias vector. The specific MIP method will be described later.

[0142] The intra predictor and the block reconstruction process based on intra prediction in the encoding device will be described with reference to Figure 5 and Figure 6 will be described.

[0143] Figure 5 is a flowchart showing a video / image encoding method based on intra prediction.

[0144] Figure 5 The encoding method of Figure 2 can be executed by the image encoding device 100 of

[0145] The image encoding device 100 may perform intra prediction S510 on a current block. The image encoding device 100 may determine an intra prediction mode / type of the current block, may derive adjacent reference samples of the current block, and may generate prediction samples within the current block based on the intra prediction mode / type and the adjacent reference samples. Here, the processes of determining the intra prediction mode / type, deriving the adjacent reference samples, and generating the prediction samples may be performed simultaneously, or one process may be performed before another process.

[0146] Figure 6 FIG. shows an example diagram of the configuration of the intra predictor 185 according to the present disclosure.

[0147] As Figure 6 shown, the intra predictor 185 of the image encoding device 100 may include an intra prediction mode / type determiner 186, a reference sample deriver 187, and / or a prediction sample deriver 188. The intra prediction mode / type determiner 186 may determine the intra prediction mode / type of the current block. The reference sample deriver 187 may derive the adjacent reference samples of the current block. The prediction sample deriver 188 may derive the prediction samples of the current block. Additionally, although not shown, when performing the prediction sample filtering process described later, the intra predictor 185 may further include a prediction sample filter (not shown).

[0148] The image encoding device 100 may determine the intra prediction mode / type to be applied to the current block among a plurality of intra prediction modes / types. The image encoding device 100 may compare the rate-distortion costs (RD costs) of the respective intra prediction modes / types and may determine the optimal intra prediction mode / type of the current block.

[0149] Furthermore, the image encoding device 100 may perform a prediction sample filtering process. Prediction sample filtering may be referred to as post-filtering. Through the prediction sample filtering process, some or all of the prediction samples may be filtered. In some cases, the prediction sample filtering process may be omitted.

[0150] Referring again to Figure 5 , the image encoding device 100 may generate residual samples S520 for the current block based on the prediction samples or the filtered prediction samples. The image encoding device 100 may derive the residual samples by subtracting the prediction samples from the original samples of the current block. In other words, the image encoding device 100 may derive the residual sample values by subtracting the corresponding prediction sample values from the original sample values.

[0151] The image encoding device 100 may encode image information including information on intra prediction (prediction information) and residual information on residual samples S530. The prediction information may include intra prediction mode information and / or intra prediction method information. The image encoding device 100 may output the encoded image information in the form of a bitstream. The output bitstream may be sent to the image decoding device 200 via a storage medium or a network.

[0152] The residual information may include the residual coding syntax described later. The image encoding device 100 may derive quantized transform coefficients by performing transform / quantization on the residual samples. The residual information may include information on the quantized transform coefficients.

[0153] In addition, as described above, the image encoding device 100 may generate a reconstructed picture (including reconstructed samples and reconstructed blocks). The image encoding device 100 may perform dequantization / inverse transform on the quantized transform coefficients to derive (modified) residual samples. The reason for applying dequantization / inverse transform after the transform / quantization of the residual samples is to derive the same residual samples as those derived in the image decoding device 200. The image encoding device 100 may generate a reconstructed block including the reconstructed samples of the current block based on the prediction samples and the (modified) residual samples. The reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, in-loop filtering processing may be further applied to the reconstructed picture.

[0154] Figure 7 is a flowchart showing a video / image decoding method based on intra prediction.

[0155] The image decoding device 200 may perform operations corresponding to those performed in the image encoding device 100.

[0156] Figure 7 The decoding method of Figure 3 may be performed by the image decoding device 200 of . Steps S710 to S730 may be performed by the intra predictor 265, and the prediction information in step S710 and the residual information in step S740 may be obtained by the entropy decoder 210 from the bitstream. The residual processor of the image decoding device 200 may derive the residual samples of the current block based on the residual information S740. Specifically, the dequantizer 220 of the residual processor may perform dequantization on the quantized transform coefficients derived based on the residual information to obtain the transform coefficients, and the inverse transformer 230 of the residual processor may perform inverse transform on the transform coefficients to derive the residual samples of the current block. Step S750 may be performed by the adder 235 or the reconstructor.

[0157] Specifically, the image decoding device 200 may derive the intra prediction mode / type S710 of the current block based on the received prediction information (i.e., intra prediction mode / type information). Additionally, the image decoding device 200 may derive the neighboring reference samples of the current block S720. The image decoding device 200 may generate prediction samples within the current block based on the intra prediction mode / type and the neighboring reference samples S730. In this case, the image decoding device 200 may perform a prediction sample filtering process. The prediction sample filtering may be referred to as post-filtering. Through this prediction sample filtering process, some or all of the prediction samples may be filtered. In some cases, the prediction sample filtering process may be omitted.

[0158] The image decoding device 200 may generate residual samples of the current block based on the received residual information S740. The image decoding device 200 may generate reconstructed samples of the current block based on the prediction samples and the residual samples, and derive a reconstructed block including the reconstructed samples S750. A reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, in-loop filtering processing may be further applied to the reconstructed picture.

[0159] Figure 8 An example diagram showing the configuration of the intra predictor 265 according to the present disclosure is shown.

[0160] As Figure 8 shown, the intra predictor 265 of the image decoding device 200 may include an intra prediction mode / type determiner 266, a reference sample deriver 267, and a prediction sample deriver 268. The intra prediction mode / type determiner 266 may determine the intra prediction mode / type of the current block based on the intra prediction mode / type information generated and signaled concurrently in the intra prediction mode / type determiner 186 of the image encoding device 100, and the reference sample deriver 267 may derive the neighboring reference samples of the current block from the reconstructed reference region of the current picture. The prediction sample deriver 268 may derive the prediction samples of the current block. Additionally, although not shown, when the above-described prediction sample filtering process is performed, the intra predictor 265 may further include a prediction sample filter (not shown).

[0161] The intra prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether the most probable mode (MPM) is applied to the current block or the remaining mode is applied, and when the MPM is applied to the current block, the intra prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) may be configured as an MPM candidate list or an MPM list. Additionally, when the MPM is not applied to the current block, the intra prediction mode information may further include remainder mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes other than the intra prediction mode candidates (MPM candidates). The image decoding device 200 may determine the intra prediction mode of the current block based on the intra prediction mode information.

[0162] Additionally, the intra prediction method information may be implemented in various forms. As an example, the intra prediction method information may include intra prediction method index information indicating one of the intra prediction methods. As another example, the intra prediction method information may include at least one of the following: reference sample row information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and which reference sample row is used when the MRL is applied, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartitions when the ISP is applied, flag information indicating whether the PDPC is applied, or flag information indicating whether the LIP is applied. Additionally, the intra prediction type information may include an MIP flag indicating whether the MIP is applied to the current block. In the present disclosure, the ISP flag information may be referred to as an ISP application indicator.

[0163] The intra prediction mode information and / or the intra prediction method information may be encoded / decoded by the encoding method described in the present disclosure. For example, the intra prediction mode information and / or the intra prediction method information may be encoded / decoded by entropy encoding (e.g., CABAC, CAVLC) based on a truncated (Rice) binary code.

[0164] In addition, in addition to the PLANAR mode, the DC mode, and the directional intra prediction mode, the intra prediction mode may further include a cross-component linear model (CCLM) mode for chrominance samples. The CCLM mode may be classified into L_CCLM, T_CCLM, and LT_CCLM according to whether the left sample, the top sample, or both are considered when deriving the CCLM parameters, and it may be applied only to the chrominance component.

[0165] For example, the intra prediction mode can be indexed as shown in Table 5 below.

[0166] [Table 5]

[0167] Intra prediction mode Associated name 0 INTRA_PLANAR 1 INTRA_DC 2..66 INTRA_ANGULAR2..INTRA_ANGULAR66 81..83 INTRA_LT_CCLM,INTRA_L_CCLM,INTRA_T_CCLM

[0168] In addition, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above-mentioned LIP, PDPC, MRL, ISP, or MIP. The intra prediction type may be indicated based on intra prediction type information, and the intra prediction type information may be implemented in various forms. As an example, the intra prediction type information may include an intra prediction type index indicating one of the intra prediction types. In another example, the intra prediction type information may include at least one of the following: reference sample row information (e.g., intra_luma_ref_idx) indicating whether MRL is applied to the current block and which reference sample row to use when MRL is applied, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition when ISP is applied, flag information indicating whether PDPC is applied, or flag information indicating whether LIP is applied. Additionally, the intra prediction type information may include an MIP flag (which may be referred to as intra_mip_flag) indicating whether MIP is applied to the current block.

[0169] Overview of the template-based intra mode derivation (TIMD)

[0170] Figure 9 FIG. is a diagram showing a template region and reference samples for TIMD according to the present disclosure. TIMD can be the following mode: The mode with the minimum SATD is selected as the intra mode of the current block by calculating the sum of absolute transform differences (SATD) between the predicted block predicted from the template region 910 and the actual reconstructed samples of the intra prediction mode (IPM) of adjacent intra blocks and inter blocks.

[0171] In the TIMD mode, for each intra prediction mode in the most probable mode (MPM), the SATD between the predicted samples and the reconstructed samples can be calculated. The two modes with the minimum SATD can be selected as the TIMD mode. These two TIMD modes can be fused by using weights, and this weighted intra prediction can be used to encode the current CU. Position-dependent intra prediction combination (PDPC) can be used to derive the TIMD mode.

[0172] The errors (costs) of the two selected prediction modes can be compared with a threshold, and an error factor 2 is applied in this comparison as shown in Equation 1.

[0173] [Formula 1]

[0174] costMode2 < 2 * costMode1

[0175] Here, costMode1 may refer to the mode with the minimum SATD. Additionally, costMode2 may refer to the mode with the second minimum SATD.

[0176] When the condition in Formula 1 is satisfied, the final prediction block can be generated by fusing the prediction blocks generated using the two modes. Otherwise, only the mode with the minimum SATD value can be used to generate the final prediction block.

[0177] The weighting ratio applied when fusing the prediction blocks generated using the two prediction modes with the minimum SATD can be as shown in Formula 2 below.

[0178] [Formula 2]

[0179] weight1 = costMode2 / (costMode1 + costMode2)

[0180] weight2 = 1 - weight1

[0181] Here, weight1 may refer to the weight applied to the prediction block generated based on the mode with the minimum SATD. Additionally, weight2 may refer to the weight applied to the prediction block generated based on the mode with the second minimum SATD.

[0182] Overview of the decoder-side intra mode derivation (DIMD)

[0183] Figure 10 is a diagram showing a method for constructing HoG in the DIMD mode.

[0184] According to the DIMD mode of the present disclosure, the intra prediction mode information can be derived from the image encoding device 100 and the image decoding device 200 and used without directly transmitting it. The DIMD mode can be performed by obtaining the horizontal gradient and the vertical gradient from the second adjacent reference columns and rows adjacent to the current block and constructing HoG based on them.

[0185] Referring to Figure 10 , the HoG can be obtained by applying the Sobel filter to the L-shaped columns and rows of three adjacent pixels 1010 around the current block. In this case, when the boundary of the block exists in different CTUs, the adjacent pixels of the current block may not be used for texture analysis.

[0186] Additionally, the Sobel filter can be referred to as the Sobel operator and can be an effective filter for edge detection. When using the Sobel filter, two types of Sobel filters can be used: the Sobel filter in the vertical direction and the Sobel filter in the horizontal direction.

[0187] Figure 11 is a diagram showing a method for constructing a prediction block when applying the DIMD mode. According to Figure 11 , the DIMD mode can be performed by selecting two intra modes 1110 with the highest histogram amplitudes and constructing the final prediction block 1150 by mixing the prediction blocks predicted by the two selected intra modes (1120, 1130) and the prediction block predicted by the planar mode (1040). In this case, the weights applied when mixing the prediction blocks can be derived from the histogram amplitudes 1160. Additionally, the DIMD flag can be sent on a per-block basis to determine whether to use DIMD.

[0188] DIMD chroma mode (DIMD in chroma)

[0189] As Figure 12 in, the DIMD chrominance mode can use the DIMD derivation method to derive the chrominance intra prediction mode of the current block based on the adjacent reconstructed Y samples of the second adjacent rows and columns ( Figure 12 the filled samples in (a)), Cb samples ( Figure 12 the filled samples in (b)), and Cr samples ( Figure 12 the filled samples in (c)).

[0190] Specifically, to construct the HoG, the horizontal gradient and the vertical gradient can be calculated for each reconstructed luma sample and the reconstructed Cb and Cr samples placed together with the current chrominance block. The chrominance intra prediction of the current chrominance block can be performed by using the intra prediction mode with the maximum histogram amplitude value.

[0191] When the intra prediction mode derived from the DIMD chrominance mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value can be used as the DIMD chrominance mode. A CU-level flag can be signaled to indicate whether the DIMD chrominance mode is applied.

[0192] CIIP (Combined inter and intra prediction)

[0193] The CIIP mode can be applied to the current block. An additional flag (e.g., ciip_flag) can be signaled to indicate whether the CIIP mode is applied to the current CU. For example, when the CU is encoded in the merge mode, if the CU includes at least 64 luma samples (i.e., the product of the CU width and the CU height is greater than or equal to 64) and both the width and height of the CU are less than 128 luma samples, the additional flag can be signaled. As its name implies, CIIP prediction can combine an inter-prediction signal and an intra-prediction signal. The inter-prediction signal P_inter of the CIIP mode can be derived by using the same processing as the inter-prediction processing applied to the general merge mode. The intra-prediction signal P_intra can be derived by using the conventional intra-prediction processing of the planar mode. Then, the intra- and inter-prediction signals can be combined by using a weighted average, where the weight value can be calculated as follows according to the coding modes of the upper adjacent block and the left adjacent block (refer to Figure 13 ).

[0194] When the upper adjacent block is available and intra-coded, isIntraTop is configured to 1, otherwise, isIntraTop is configured to 0;

[0195] When the left adjacent block is available and intra-coded, isIntraLeft is configured to 1, otherwise, isIntraLeft is configured to 0;

[0196] When (isIntraTop + isIntraLeft) is equal to 2, wt is configured to 2, otherwise, wt is configured to 1.

[0197] CIIP prediction can be constructed as in Equation 3.

[0198] [Equation 3]

[0199] P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2

[0200] P CIIP can represent the predicted block through CIIP prediction, P inter can represent the predicted block through inter-prediction, and P intra can represent the predicted block through intra-prediction.

[0201] In the CIIP mode, the predicted samples can be generated by applying weights to the inter-prediction signal obtained by using CIIP template matching (TM) to merge candidate predictions and the intra-prediction signal obtained by using TIMD prediction derived from the intra-prediction mode. This method can be applied only to coding blocks with an area less than or equal to 1024.

[0202] This TIMD derivation method can be used to derive the intra prediction mode in CIIP. Specifically, the intra prediction mode with the minimum SATD value in the TIMD mode list is selected, and this intra prediction mode can be mapped to one of the 67 conventional intra prediction modes.

[0203] In addition, when the derived intra prediction mode is an angular mode, the weights (wIntra, wInter) in the following two cases can be modified. For adjacent horizontal modes (2 <= angular mode index < 34), the current block can be split in the vertical direction as shown in Figure 14 (a) below. For adjacent vertical modes (34 <= angular mode index <= 66), the current block can be split in the horizontal direction as shown in Figure 14 (b) below.

[0204] The weights (wIntra, wInter) of different sub - blocks are shown in Table 6.

[0205] [Table 6]

[0206] Sub-block index (wIntra,wInter) 0 (6,2) 1 (5,3) 2 (3,5) 3 (2,6)

[0207] When using CIIP - TM, a CIIP - TM merge candidate list for the CIIP - TM mode can be constructed. The merge candidates can be refined through template matching. The CIIP - TM merge candidates can also be reordered into conventional merge candidates by the ARMC method. The maximum number of CIIP - TM merge candidates can be 2.

[0208] Implementation

[0209] The present invention relates to an inseparable forward / inverse transform method and its apparatus. The embodiments proposed in the present invention can be applied to existing video compression technologies including General Video Coding (VVC), and can also be used in combination with next - generation video coding technologies.

[0210] The "transform" mentioned in the present invention can refer to the transform performed by the image encoding device 100 or the inverse transform performed by the image decoding device 200. In other words, in the following description, "transform" can be used to include the meaning of inverse transform. In addition, the "Non - separable Single - Pass Transform (NSPT)" mentioned in the present invention can refer to the non - separable single - pass transform performed by the image encoding device 100 or the non - separable single - pass inverse transform performed by the image decoding device 200. In other words, in the following description, "non - separable single - pass transform" can be used to include the meaning of non - separable single - pass inverse transform. The embodiments mentioned in the following description can be performed by the image encoding device 100 and the image decoding device 200.

[0211] Figure 15 is a flowchart showing an image encoding / decoding method according to an embodiment of the present disclosure.

[0212] Reference Figure 15 ,the prediction mode of the current block can be determined in S1510. The prediction mode of the current block can be referred to as "the mode for prediction of the current block" or "the mode for current block prediction", etc.

[0213] It can be determined whether the prediction mode of the current block is a predetermined prediction mode in S1520. The predetermined prediction mode can include at least one of the MIP mode and the hybrid prediction mode. The hybrid prediction mode can be a mode of deriving a prediction block of the current block based on neighboring samples around the current block. For example, the hybrid prediction mode can include at least one of the TIMD mode, the DIMD mode, the DIMD chrominance mode, or the CIIP mode.

[0214] The residual samples of the current block can be transformed in S1530. The transformation of the residual samples can be performed based on whether the prediction mode of the current block is a predetermined prediction mode. In addition, the transformation of the residual samples can be performed based on the prediction mode corresponding to the current block among the prediction modes corresponding to the predetermined prediction mode.

[0215] Implementation 1

[0216] The method for transforming residual samples from the spatial domain to the frequency domain through a predefined transformation can be roughly divided into separable transformation and non-separable transformation.

[0217] Generally, since image pixels are values (pixel values or sample values) existing in a two-dimensional (position with horizontal and vertical coordinates) space, a transformation process of transforming them to another domain (e.g., the frequency domain) by using predefined basis vectors can be performed. For separable transformation, since each transformation in the horizontal and vertical directions is performed independently, the transformation coefficients are obtained through the inner product between each of the horizontal input value and the vertical input value and the basis vector. Therefore, compared with non-separable transformation, the required computational complexity can be relatively small. For non-separable transformation, in order to identify the overall characteristics of the pixel values existing in the two-dimensional space, the inner product between the size corresponding to the total number of pixels existing in the two-dimensional space and the basis vector (or kernel) must be applied. Therefore, compared with separable transformation, the required kernel size may be larger and the computational complexity may be higher. However, since non-separable transformation can well express the similarity existing between pixels in the two-dimensional space, a significantly improved coding efficiency can be provided.

[0218] Separable transforms and non-separable transforms are designed to efficiently transmit the energy of the residual signal through mathematical methods (transforming to the frequency domain) so as to ultimately well identify the characteristics of the residual signal and effectively express it. Generally, the residual signal (which can be called a residual block when encoded in blocks) refers to the difference between the original image block and the predicted image block, and for intra-mode prediction, the predicted image block is generated as an intra-prediction block by using the values of the already decoded adjacent pixels. Therefore, when using a direction-based intra-prediction mode such as a conventional video compression method, the characteristics of the predicted block are similarly formed according to the direction of the intra-prediction mode, which has the same effect on the characteristics of the residual block. Thus, the statistical characteristics of the residual image are affected according to the prediction mode of the given block. For example, for a block using the horizontal intra-prediction mode, since the prediction mode block replicates and uses pixel values in the horizontal direction, the residual image has a distribution characteristic of residual pixel values with a pattern similar to that of the predicted block. When similar intra-prediction modes are used even for different images, characteristics that are commonly shared with each other are retained in the residual image, and a single transform technique improves the coding efficiency by using basis vectors (kernels) that well reflect this characteristic.

[0219] In the present invention, in order to more efficiently use the single transform based on the characteristics of the residual signal that is the difference between the predicted pixels for the single transform and the pixel values existing in the original image, embodiments regarding the interaction between the single transform and coding tools using different prediction modes are proposed.

[0220] Embodiment 1 is an embodiment of a method for transforming residual samples based on whether the prediction mode of the current block is an MIP mode corresponding to a predetermined prediction mode.

[0221] As an example, when the prediction mode of the current block is the MIP mode, the non-separable single transform may not be applied. Due to the difference in the prediction method, the block encoded in the MIP mode (MIP mode block) may have different residual characteristics from those of the block not encoded in the MIP mode (non-MIP mode block). Therefore, in order to improve the coding efficiency or reduce the computational complexity, the non-separable single transform may be omitted for the MIP mode block. In this case, the syntax elements related to the non-separable single transform (non-separable single transform flag or / and non-separable single transform index, etc.) may not be signaled / parsed for the MIP mode block. This method may correspond to a method for specifying to encode the additional information about the non-separable single transform technique after the additional information about the MIP coding technique.

[0222] The flowchart of the method for not applying the non-separable single transform when the prediction mode of the current block is the MIP mode is as Figure 16 shown.

[0223] Referring toFigure 16 After determining the prediction mode of the current block in S1610, it is possible to determine whether the prediction mode of the current block is a predetermined prediction mode (MIP mode) in S1620. When the prediction mode of the current block is not the MIP mode, a non-separable first-order transform and / or a separable first-order transform can be applied to the residual block of the current block in S1640. In other words, when the prediction mode of the current block is not the MIP mode, both a transform other than the non-separable first-order transform and the non-separable first-order transform can be applied. On the contrary, when the prediction mode of the current block is the MIP mode, the non-separable first-order transform is not applied, and a transform other than the non-separable first-order transform can be applied in S1630.

[0224] As another example, even when the prediction mode of the current block is a predetermined prediction mode (MIP mode), the non-separable first-order transform can be applied. The flowchart of the method for applying the non-separable first-order transform when the prediction mode of the current block is the MIP mode is as Figure 17 shown.

[0225] Referring to Figure 17 After determining the prediction mode of the current block in S1710, it is possible to determine whether the prediction mode of the current block is a predetermined prediction mode (MIP mode) in S1720. When the prediction mode of the current block is not the MIP mode, a non-separable first-order transform and / or a separable first-order transform can be applied to the residual block of the current block in S1740. In other words, when the prediction mode of the current block is not the MIP mode, both a transform other than the non-separable first-order transform and the non-separable first-order transform can be applied. In contrast, when the prediction mode of the current block is the MIP mode, the non-separable first-order transform can be applied in S1730.

[0226] When applying the non-separable first-order transform to an MIP mode block, the transform set can be allocated according to a predetermined rule. In other words, the transform set for the non-separable first-order transform of the residual samples can be determined as the transform set candidate corresponding to the predetermined intra mode among the transform set candidates. Here, the predetermined intra mode can be the planar mode or the DC mode. For example, when deriving the transform set through the intra prediction mode in the non-separable first-order transform, the intra prediction mode of the MIP mode block can be assigned as the planar mode, and the transform set can be derived through the planar mode. When applying this method, some or all of the syntax elements such as the non-separable first-order transform flag and the non-separable first-order transform index may not be signaled / parsed for the MIP mode block, so the coding efficiency can be improved.

[0227] When applying an inseparable first transform to an MIP mode block, the transform kernel of the inseparable first transform for residual samples can be determined as the transform kernel corresponding to the MIP mode. Here, the transform kernel corresponding to the MIP mode is a kernel trained separately by MIP mode information and can be a kernel different from the kernels in the above transform kernel set. In other words, when applying an inseparable first transform to an MIP mode block, the kernel trained separately by MIP mode information can be assigned to the transform kernel for the inseparable first transform. Since the MIP mode block has residual characteristics different from those of non-MIP mode blocks due to different prediction methods, the separately trained kernel can be efficient. When using this method, it is not necessary to signal / parse some or all of the syntax elements such as the inseparable first transform flag, inseparable first transform index, etc. for the MIP mode block, so the coding efficiency can be improved.

[0228] When applying an inseparable first transform to an MIP mode block, the transform set of the inseparable first transform for residual samples can be determined based on the information of adjacent samples adjacent to the current block among the transform set candidates. The information of adjacent samples can be the information of all / some samples adjacent to the top / left side of the current block. For example, when deriving the transform set through the intra prediction mode in the inseparable first transform, the direction of the residual samples can be determined by analyzing the information of adjacent samples, an intra prediction mode suitable for it can be assigned, and the transform set corresponding to the direction of the residual samples can be derived. When using this method, it is not necessary to signal / parse some or all of the syntax elements such as the inseparable first transform flag, inseparable first transform index, etc. for the MIP mode block, so the coding efficiency can be improved.

[0229] When applying an inseparable first transform to an MIP mode block, the transform set for the inseparable first transform can be determined based on multiple MIP prediction mode information (multiple MIP prediction modes). The multiple MIP prediction mode information can be the above-mentioned "intra prediction mode for MIP". For example, different inseparable first transform sets can be derived for each of the multiple MIP prediction modes, or different inseparable first transform sets can be derived in units of groups in which multiple MIP prediction modes are grouped. When using this method, it is not necessary to signal / parse some or all of the syntax elements such as the inseparable first transform flag, inseparable first transform index, etc. for the MIP mode block, so the coding efficiency can be improved. The number of multiple MIP prediction modes can be variably determined according to the size of the block, and the inseparable first transform sets derived in units of multiple MIP prediction modes or in units of groups can also be variably determined according to the size of the block.

[0230] Implementation 2

[0231] As a method of one - time transformation from the spatial domain to the frequency domain, there may be a method for always applying only separable transformations, a method for always applying only non - separable transformations, a method for applying one of separable and non - separable transformations, and a method for applying both separable and non - separable transformations. Here, the one - time transformation method based on separable transformations may include DCT type 2, DST type 7, DCT type 8, DCT type 5, DST type 4, DST type 1, IDT (identity transformation), or other transformations not based on non - separable transformations (e.g., transformation skip), etc.

[0232] Embodiment 2 is an embodiment of a method for transforming residual samples based on whether the prediction mode of the current block is a hybrid prediction mode corresponding to a predetermined prediction mode.

[0233] As an example, when the prediction mode of the current block is a hybrid prediction mode, non - separable one - time transformation may not be applied. Due to different prediction methods, blocks encoded in various types of hybrid prediction modes (hybrid prediction mode blocks) other than the single intra - frame mode may have different residual characteristics from blocks not encoded in the hybrid prediction mode (non - hybrid prediction mode blocks). Therefore, to improve coding efficiency or reduce computational complexity, non - separable one - time transformation may be omitted for hybrid prediction mode blocks. In this case, syntax elements related to non - separable one - time transformation (non - separable one - time transformation flag or / and non - separable one - time transformation index, etc.) may not be signaled / parsed for hybrid prediction mode blocks. This method may correspond to a method in which additional information about non - separable one - time transformation technology is encoded after additional information about hybrid prediction mode.

[0234] The flowchart of the method for not applying non - separable one - time transformation when the prediction mode of the current block is a hybrid prediction mode is as Figure 16 shown.

[0235] Referring to Figure 16 , after determining the prediction mode of the current block S1610, it may be determined whether the prediction mode of the current block is a predetermined prediction mode (hybrid prediction mode) S1620. When the prediction mode of the current block is not a hybrid prediction mode, non - separable one - time transformation and / or separable one - time transformation may be applied to the residual block of the current block S1640. In other words, when the prediction mode of the current block is not a hybrid prediction mode, both transformations other than non - separable one - time transformation and non - separable one - time transformation may be applied. On the contrary, when the prediction mode of the current block is a hybrid prediction mode, non - separable one - time transformation is not applied and transformations other than non - separable one - time transformation may be applied S1630.

[0236] As another example, even when the prediction mode of the current block is a predetermined prediction mode (hybrid prediction mode), an inseparable first-order transform can be applied. The flowchart of the method for applying an inseparable first-order transform when the prediction mode of the current block is a hybrid prediction mode is as Figure 17 shown.

[0237] Referring to Figure 17 , after determining the prediction mode S1710 of the current block, it can be determined whether the prediction mode of the current block is a predetermined prediction mode (hybrid prediction mode) S1720. When the prediction mode of the current block is not a hybrid prediction mode, an inseparable first-order transform and / or a separable first-order transform can be applied to the residual block of the current block S1740. In other words, when the prediction mode of the current block is not a hybrid prediction mode, both a transform other than the inseparable first-order transform and the inseparable first-order transform can be applied. On the contrary, when the prediction mode of the current block is a hybrid prediction mode, an inseparable first-order transform can be applied S1730.

[0238] When applying an inseparable first-order transform to a hybrid prediction mode block, a transform set can be assigned according to a predetermined rule. In other words, the transform set for the inseparable first-order transform of the residual samples can be determined as the transform set candidate corresponding to the predetermined intra mode among the transform set candidates. Here, the predetermined intra mode can be a planar mode or a DC mode. For example, when deriving the transform set through the intra prediction mode in the inseparable first-order transform, the intra prediction mode of the hybrid prediction mode block can be assigned as a planar mode or a DC mode, and the transform set can be derived through this mode. When this method is applied, some or all of the syntax elements such as the inseparable first-order transform flag, the inseparable first-order transform index, etc. do not need to be signaled / parsed for the hybrid prediction mode block, so the coding efficiency can be improved.

[0239] When applying an inseparable first-order transform to a hybrid prediction mode block, the transform kernel for the inseparable first-order transform of the residual samples can be determined as the transform kernel corresponding to the hybrid prediction mode. Here, the transform kernel corresponding to the hybrid prediction mode is a kernel trained separately through the hybrid prediction mode information, and can be a kernel different from the kernels in the above transform kernel set. In other words, when applying an inseparable first-order transform to a hybrid prediction mode block, the kernel trained separately through the hybrid prediction mode information can be assigned to the transform kernel for the inseparable first-order transform. Since the hybrid prediction mode block has residual characteristics different from those of the non-hybrid prediction mode block due to different prediction methods, the separately trained kernel can be efficient. When using this method, some or all of the syntax elements such as the inseparable first-order transform flag, the inseparable first-order transform index, etc. do not need to be signaled / parsed for the hybrid prediction mode block, so the coding efficiency can be improved.

[0240] When applying a non-separable first-order transform to a hybrid prediction mode block, the transform set of the non-separable first-order transform for residual samples can be determined among the transform set candidates based on the information of neighboring samples adjacent to the current block. The information of neighboring samples can be the information of all / some samples adjacent to the top / left side of the current block. For example, when deriving the transform set through an intra prediction mode in a non-separable first-order transform, the direction of the residual samples can be determined by analyzing the information of neighboring samples, an intra prediction mode suitable for it can be assigned, and the transform set corresponding to the direction of the residual samples can be derived. When using this method, it is not necessary to signal / parse some or all of the syntax elements such as non-separable first-order transform flags, non-separable first-order transform indices, etc. for the hybrid prediction mode block, so the coding efficiency can be improved.

[0241] When applying a non-separable first-order transform to a hybrid prediction mode block, the number of non-separable first-order transform sets (the number of transform set candidates for the non-separable first-order transform) and the number of transform kernels in each set (the number of transform kernels included in each transform set candidate) can be configured differently according to the additional information derived from the hybrid prediction step (the information derived from the hybrid prediction mode). For example, when the hybrid prediction mode is the DIMD mode, the HoG can be constructed through the information of neighboring pixels, and two intra modes can be selected in the order of the magnitude of the large histogram. In this case, the number of transform set candidates and the number of transform kernels in each transform set can be determined based on the difference between the values of the two selected intra modes. Specifically, when the difference between the values of the selected intra modes is greater than or equal to a certain value, n1 transform sets and k1 transform kernels can be configured. If the difference between the values of the selected intra modes is less than or equal to a certain value, n2 transform sets and k2 transform kernels can be configured.

[0242] When an inseparable primary transform is applied to a hybrid prediction mode block, the transform set for the inseparable primary transform can be determined based on additional information derived from the hybrid prediction step (information derived from the hybrid prediction mode). For example, when the hybrid prediction mode is the DIMD mode, the HoG can be constructed with the information of adjacent pixels, and two intra modes can be selected in the order of the large histogram magnitude. In this case, the transform set can be determined by using the intra mode with the largest histogram magnitude among the two selected intra modes. In different embodiments, the transform set for the inseparable primary transform can also be determined based on the intra prediction mode information derived from the hybrid prediction step. For example, when the hybrid prediction mode is the TIMD mode, the intra prediction mode for the hybrid prediction mode is derived by using the pixel information of adjacent template regions, and in this process, when the difference between the adjacent reconstructed template region and the template region predicted from the adjacent pixels of the template region satisfies a specific condition, the intra prediction mode may not be mixed and a single intra prediction mode can be used. In this case, the transform set can be determined based on the single intra prediction mode derived from this process, and the determined transform set can be the same as the transform set used when the intra prediction mode of the block is not a hybrid prediction mode.

[0243] Figure 18 FIG. shows an exemplary schematic diagram of a content streaming system to which embodiments of the present disclosure can be applied.

[0244] As Figure 18 shown, the content streaming system to which embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0245] The encoding server compresses the content input from a multimedia input device such as a smart phone, a camera, or a video camera into digital data, generates a bitstream, and sends it to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0246] The bitstream can be generated by applying the video encoding method and / or the image encoding device of the embodiments of the present disclosure, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.

[0247] The streaming server can send multimedia data to a user device via a web server based on a user request, and the web server can be used as an intermediary for notifying the user of available services. When the user requests a required service from the web server, the web server can send the request to the streaming server, and the streaming server can send multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server can be used to control command / response exchanges between devices within the content streaming system.

[0248] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a seamless streaming service, the streaming server can store the bitstream for a period of time.

[0249] Examples of user devices can include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (i.e., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.

[0250] Each server within the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

[0251] The scope of the present disclosure includes software or machine-executable instructions (i.e., operating systems, applications, firmware, programs, etc.) that enable methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media in which such software or instructions are stored and that can be executed on a device or computer.

[0252] Industrial Applicability

[0253] Embodiments of the present disclosure can be used for encoding / decoding of images.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Determine a prediction mode of a current block; Determine whether the prediction mode of the current block is a predetermined prediction mode; And Perform a transform on residual samples for the current block, wherein the predetermined prediction mode includes at least one of a matrix-based intra prediction MIP mode or a hybrid prediction mode, and wherein the hybrid prediction mode is a mode of deriving a prediction block based on neighboring samples.

2. The method according to claim 1, wherein Based on the prediction mode of the current block being the predetermined prediction mode, a non-separable single transform is not applied to the residual samples.

3. The method according to claim 1, wherein Based on the prediction mode of the current block being the predetermined prediction mode, a transform set for the non-separable single transform for the residual samples is determined as a transform set candidate corresponding to a predetermined intra mode among transform set candidates.

4. The method according to claim 3, wherein The predetermined intra mode is a planar mode or a DC mode.

5. The method according to claim 1, wherein, Based on the prediction mode of the current block being the predetermined prediction mode, the transform set for the non-separable single transform for the residual samples is determined among the transform set candidates based on information of neighboring samples adjacent to the current block.

6. The method according to claim 5, wherein Derive a direction of the residual samples from the information of the neighboring samples, and wherein the transform set for the non-separable single transform is determined as a transform set candidate corresponding to the direction of the residual samples among the transform set candidates.

7. The method according to claim 1, wherein Based on the prediction mode of the current block being the MIP mode, a transform set for the non-separable single transform is determined for each of the MIP mode candidates for the current block, or for each of the groups into which the MIP mode candidates are grouped.

8. The method according to claim 1, wherein Based on the prediction mode of the current block being the hybrid prediction mode, the number of transform set candidates for the non-separable single transform and the number of transform kernels included in each transform set candidate are determined based on information derived from the hybrid prediction mode.

9. The method according to claim 8, wherein, Based on the prediction mode of the current block being the decoder-side intra mode derivation DIMD mode, the number of the transform set candidates and the number of the transform kernels are determined based on the difference between intra modes derived from the DIMD mode.

10. The method according to claim 1, wherein, Based on the prediction mode of the current block being the hybrid prediction mode, the transform set for the non-separable single transform is determined based on information derived from the hybrid prediction mode.

11. The method according to claim 10, wherein, Based on the prediction mode of the current block being the DIMD mode, the transform set is determined based on the intra mode derived from the DIMD mode.

12. The method according to claim 1, wherein, Based on the prediction mode of the current block being the predetermined prediction mode, the transform kernel for the non-separable single transform for the residual samples is determined as a transform kernel corresponding to the predetermined prediction mode.

13. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Determine a prediction mode of a current block; Determine whether the prediction mode of the current block is a predetermined prediction mode; And Perform a transform on residual samples for the current block, wherein, the predetermined prediction mode includes at least one of a matrix-based intra prediction MIP mode or a hybrid prediction mode, and wherein, the hybrid prediction mode is a mode of deriving a prediction block based on neighboring samples.

14. A method for transmitting a bitstream generated by an image coding method, the image coding method comprising the steps of: determining a prediction mode of a current block; determining whether the prediction mode of the current block is a predetermined prediction mode; and transforming residual samples for the current block, wherein, the predetermined prediction mode includes at least one of a matrix-based intra prediction MIP mode or a hybrid prediction mode, and wherein, the hybrid prediction mode is a mode of deriving a prediction block based on neighboring samples.

15. A computer-readable storage medium storing a bitstream generated by an image coding method according to claim 13.