Image encoding / decoding method and apparatus, and recording medium storing bit stream
By employing inseparable master transform and dimensionality reduction techniques in image decoding and encoding, the problem of insufficient compression efficiency for high-resolution and high-quality images is solved, achieving a more efficient image compression effect.
Patent Information
- Application Number
- CN202480037548.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-12
- Filing Date
- 2024-04-12
- Publication Date
- 2025-12-30
AI Technical Summary
Existing image compression techniques are inefficient when processing high-resolution and high-quality images, especially when using non-separable master transforms, and cannot effectively improve transform performance and coding efficiency.
The method employs the Non-Separable Principal Transform (NSPT) and the dimension-reduced Non-Separable Principal Transform kernel, combined with the coding parameter determination and signal notification of the Non-Separable Transform kernel, for the derivation of inverse transform and transform coefficients in the image decoding and encoding process, and distinguishes different transform methods for different block sizes and types.
By using an inseparable master transform and dimensionality reduction techniques, the transform performance and coding efficiency are improved, resulting in better compression of high-resolution and high-quality images.
Smart Images

Figure CN121241569A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to an image encoding / decoding method and apparatus and a recording medium storing a bitstream. BACKGROUND
[0002] Recently, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various application fields, and thus, high efficient image compression techniques are being discussed.
[0003] There are various techniques such as an inter prediction technique of predicting pixel values included in a current picture from a picture before or after the current picture using a video compression technique, an intra prediction technique of predicting pixel values included in a current picture by using pixel information in the current picture, an entropy encoding technique of assigning a short symbol to a value having a high frequency of occurrence and a long symbol to a value having a low frequency of occurrence, etc., which can be used to efficiently compress image data and transmit or store it. SUMMARY
[0004] TECHNICAL PROBLEM
[0005] The present disclosure provides a method and apparatus of performing a transform by using a non-separable primary transform.
[0006] The present disclosure provides a method and apparatus of performing a transform by using a non-separable primary transform kernel of reduced dimension.
[0007] The present disclosure provides a method and apparatus of determining / signaling a non-separable transform kernel based on an encoding parameter.
[0008] TECHNICAL SOLUTION
[0009] The image decoding method and apparatus according to the present disclosure can obtain residual information from a bitstream, derive transform coefficients of a current block based on the residual information, derive residual samples of the current block by performing at least one of dequantization or inverse transform on the transform coefficients of the current block, and reconstruct the current block based on the residual samples of the current block. Here, the inverse transform can be performed based on a non-separable primary transform (NSPT), and the NSPT can be applied based on at least one of a size of the current block, a tree type, or a component type.
[0010] In the image decoding method and apparatus according to the present disclosure, a pre-defined allowed transform block size can be divided into a first group of a block size set to which the NSPT can be applied and a second group of a block size set to which the NSPT is not applied.
[0011] In the image decoding method and apparatus according to the present disclosure, when the size of the current block belongs to the first group, the inverse transform of the current block can be performed based on the NSPT.
[0012] In the image decoding method and apparatus according to the present disclosure, when the size of the current block belongs to the second group, inverse transform of the current block can be performed based on a separable primary transform.
[0013] In the image decoding method and apparatus according to the present disclosure, when the size of the current block belongs to the second group, inverse transform of the current block can be performed based on a non-separable secondary transform and a separable primary transform.
[0014] In the image decoding method and apparatus according to the present disclosure, the first group can include 4x4, and the second group can include 8x8.
[0015] In the image decoding method and apparatus according to the present disclosure, the first group can include 4x8 or 8x4, and the second group can include 16x16.
[0016] In the image decoding method and apparatus according to the present disclosure, the first group can include 4x16 or 16x4, and the second group can include 16x32 or 32x16.
[0017] In the image decoding method and apparatus according to the present disclosure, the first group can include 4x32 or 32x4, and the second group can include 32x32.
[0018] In the image decoding method and apparatus according to the present disclosure, the first group can include 8x32 or 32x8, and the second group can include 32x32.
[0019] The image encoding method and apparatus according to the present disclosure can derive residual samples of a current block, derive transform coefficients of the current block by performing at least one of transform or quantization on the residual samples of the current block, and encode the transform coefficients of the current block. Here, the transform can be performed based on a non-separable primary transform (NSPT), and the NSPT can be applied based on at least one of a size of the current block, a tree type, or a component type.
[0020] A computer-readable digital storage medium storing encoded video / image information is provided, which enables a decoding apparatus according to the present disclosure to perform an image decoding method.
[0021] A computer-readable digital storage medium storing video / image information generated according to an image encoding method according to the present disclosure is provided.
[0022] A method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure are provided.
[0023] Advantageous effects
[0024] The present disclosure can improve transform performance by using a non-separable primary transform as a primary transform.
[0025] This disclosure can improve transformation performance by using a dimensionally reduced, non-separable master transform kernel to perform the transformation.
[0026] This disclosure can improve coding efficiency by effectively determining and / or signaling the inseparable transform kernel based on coding parameters. Attached Figure Description
[0027] Figure 1 A video / image encoding system according to this disclosure is shown.
[0028] Figure 2 A schematic block diagram of an encoding apparatus to which embodiments of the present disclosure are applicable and which performs encoding of video / image signals is shown.
[0029] Figure 3 A schematic block diagram of a decoding apparatus to which embodiments of the present disclosure are applicable and which performs decoding of video / image signals is shown.
[0030] Figure 4 An image decoding method is shown, performed by a decoding device (300) according to an embodiment of the present disclosure.
[0031] Figure 5 An intra-frame prediction mode and its prediction direction according to this disclosure are illustrated by way of example.
[0032] Figure 6 A schematic configuration of a decoding device (300) performing an image decoding method according to the present disclosure is shown.
[0033] Figure 7 An image encoding method is shown performed by an encoding device (200) according to an embodiment of the present disclosure.
[0034] Figure 8 A schematic configuration of an encoding device (200) performing an image encoding method according to the present disclosure is shown.
[0035] Figure 9 Examples of content streaming systems to which embodiments of this disclosure can be applied are shown. Detailed Implementation
[0036] Because this disclosure can be modified in various ways and has multiple embodiments, specific embodiments will be shown in the accompanying drawings and described in detail in the specific embodiments. However, this disclosure is not intended to be limited to the specific embodiments, but should be understood to include all changes, equivalents, and substitutions included within the spirit and scope of this disclosure. Similar reference numerals are used for similar components in the description of the various figures.
[0037] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. Terms are used only to distinguish one component from others. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. Terms include any combination of one or more of the associated terms.
[0038] When a component is described as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but the other component can exist in between. On the other hand, when a component is described as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in between.
[0039] The terminology used in this application is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that terms such as “comprising” or “having” are intended to specify the presence of the features, quantities, steps, operations, components, portions, or combinations thereof described in this specification, but do not preclude the possibility of the presence or addition of one or more other features, quantities, steps, operations, components, portions, or combinations thereof.
[0040] This disclosure relates to video / image coding. For example, the methods / implementations disclosed herein can be applied to methods disclosed in the Multifunctional Video Coding (VVC) standard. Additionally, the methods / implementations disclosed herein can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).
[0041] This specification sets forth various implementations of video / image encoding, and unless otherwise specified, these implementations may be combined with each other.
[0042] In this article, video can refer to a collection of images over time. A frame typically refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a frame during encoding. A slice / tile can include at least one Code Tree Unit (CTU). A frame can consist of at least one slice / tile. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a frame. A tile column is a rectangular area of CTUs with a height equal to the height of the frame and a width specified by the syntax requirements of the frame parameter set. A tile row is a rectangular area of CTUs with a height specified by the frame parameter set and a width equal to the width of the frame. CTUs within a tile can be arranged continuously according to CTU raster scans, and tiles within a frame can be arranged continuously according to tile raster scans. A slice can include an integer number of complete tiles of a frame that can be exclusively included in a single NAL unit, or an integer number of consecutive complete CTU rows within a tile. Furthermore, a frame can be divided into at least two sub-frames. A sub-frame can be a rectangular area of at least one slice within a frame.
[0043] A pixel, or pelin, can represent the smallest unit that makes up a frame (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0044] A unit can represent the basic unit of image processing. A unit may include a specific region of the image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region". In general, an M×N block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0045] In this document, “A or B” can mean “A only”, “B only”, or “both A and B”. In other words, “A or B” can be interpreted as “A and / or B”. For example, “A, B or C” can mean “A only”, “B only”, “C only”, or “any combination of A, B and C”.
[0046] The forward slash ( / ) or comma used in this article can indicate "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0047] In this document, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0048] Additionally, in this document, "at least one of A, B, and C" can mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0049] Additionally, the parentheses used in this document can indicate "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction". In other words, "prediction" in this document is not limited to "intra-frame prediction", and "intra-frame prediction" can be cited as an example of "prediction". Furthermore, even when the indication is "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction".
[0050] In this article, the technical features described individually in a single diagram can be implemented individually or simultaneously.
[0051] Figure 1 A video / image encoding system according to this disclosure is shown.
[0052] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device).
[0053] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or stream via a digital storage medium or network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may consist of a separate device or external components.
[0054] A video source can acquire video / images through processes that capture, synthesize, or generate video / images. A video source may include means for capturing video / images and means for generating video / images. Means for capturing video / images may include at least one camera, a video / image archive containing previously captured video / images, etc. Means for generating video / images may include a computer, tablet computer, smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., and in this case, the process of capturing video / images can be replaced by a process of generating related data.
[0055] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0056] The transmitting unit can send encoded video / image information or data, output in bitstream form, to the receiving unit of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files according to a predetermined file format, and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and send it to a decoding device.
[0057] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.
[0058] The renderer can render decoded video / images. The rendered video / images can be displayed through a display unit.
[0059] Figure 2 A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.
[0060] Reference Figure 2The encoding device 200 may consist of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0061] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into at least one processing unit. As an example, a processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively segmented from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure.
[0062] For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree, binary tree, and / or ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree and / or ternary tree structure. Alternatively, a binary tree structure can be applied before the quadtree structure. The coding process according to this specification can be performed based on the final coding unit that is no longer segmented. In this case, based on image characteristics, coding efficiency, etc., the largest coding unit can be directly used as the final coding unit, or, if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, as described later.
[0063] As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be divided or segmented from the aforementioned final encoding unit, respectively. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0064] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an M×N block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term to form a frame (or image) corresponding to a pixel or cell.
[0065] Encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to converter 232. In this case, the unit in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231.
[0066] Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes prediction samples of the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction on a per-block or per-unit basis. Predictor 220 can generate various information about the prediction (e.g., prediction mode information) and send it to entropy encoder 240, as described later in the description of the various prediction modes. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.
[0067] Intra-predictor 222 can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples can be located near the current block or positioned at a specific distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include 33 or 65 directional modes. However, this is just an example; more or fewer directional modes can be used depending on the configuration. Intra-predictor 222 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0068] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block, a co-located CU (colCU), etc., and the reference frame including the temporally neighboring block may be referred to as a co-located frame (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0069] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply both intra-frame and inter-frame prediction simultaneously. This can be referred to as the Inter-intra-frame Combined Prediction (CIIP) mode. Alternatively, the predictor can predict blocks based on the Intra-Block Copy (IBC) prediction mode, or it can predict blocks based on a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as Screen Content Coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values within the frame can be signaled based on information about the palette table and palette index. The prediction signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal.
[0070] Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen–Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT represents the transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT represents the transform obtained based on generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or it can be applied to non-square blocks of variable size.
[0071] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. This information about the quantized transform coefficients can be called residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a 1D vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on this 1D vector form.
[0072] The entropy encoder 240 can perform various encoding methods such as Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can encode information required for video / image reconstruction other than quantization transform coefficients (e.g., values of syntax elements, etc.) together or separately.
[0073] Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) unit level. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. Information and / or syntax elements transmitted from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded and included in the bitstream through the encoding process described above. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be configured as internal / external components of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.
[0074] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 234 and inverse transformer 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed (similar to when a skip mode is applied), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can also be used for inter-frame prediction of the next frame by filtering, as described later. Furthermore, luminance mapping and chroma scaling (LMCS) can be applied in frame encoding and / or reconstruction processing.
[0075] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270 (specifically, the DPB of memory 270). Various filtering methods can include deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various information about the filtering and send it to entropy encoder 240. The information about the filtering can be encoded in entropy encoder 240 and output as a bitstream.
[0076] The modified reconstructed frame sent to memory 270 can be used as a reference frame in inter-frame predictor 221. When inter-frame prediction is applied through it, the encoding device can avoid prediction mismatch between the encoding device 200 and the decoding device, and can also improve encoding efficiency.
[0077] The DPB of memory 270 can store modified reconstructed frames for use as reference frames in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current frame from which motion information is derived (or encoded) and / or of blocks in previously reconstructed frames. The stored motion information can be sent to inter-frame predictor 221 as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current frame and send them to intra-frame predictor 222.
[0078] Figure 3 A rough block diagram of a decoding device that can be implemented using embodiments of the present disclosure and perform decoding of video / image signals is shown.
[0079] ReferenceFigure 3 The decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322.
[0080] According to the implementation, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded screen buffer (DPB) and can be configured by a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.
[0081] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The encoding device processes video / image information to reconstruct the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied in the encoding device. Therefore, the decoding processing unit can be an encoding unit, and the encoding unit can be segmented from the encoding tree unit or a larger encoding unit according to a quadtree structure, binary tree structure, and / or ternary tree structure. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be played back by a playback device.
[0082] Decoding device 300 can receive data in bitstream form from... Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements that are signaled / received, as described later herein, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of neighboring blocks and the block to be decoded, or information about symbols / bins decoded in previous steps, perform arithmetic decoding of bins by predicting the occurrence probability of bins based on the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model of the next symbol / bin. Among the information decoded in the entropy decoder 310, prediction information is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients and related parameter information) from which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, filtering information among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiving unit can be a component of the entropy decoder 310.
[0083] Furthermore, the decoding device according to this specification may be referred to as a video / image / screen decoding device, and the decoding device may be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0084] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0085] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0086] Predictor 320 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. Predictor 320 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine the specific intra-frame / inter-frame prediction mode.
[0087] Predictor 320 can generate prediction signals based on various prediction methods described later. For example, predictor 320 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as the Inter-Frame Intra-Frame Combined Prediction (CIIP) mode. Alternatively, the predictor can predict blocks based on the Intra-Frame Block Copy (IBC) prediction mode, or it can predict blocks based on a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as Screen Content Coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0088] Intra-predictor 331 can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples can be located near the current block or at a specific distance away. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0089] Inter-frame predictor 332 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0090] Adder 340 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array). When there is no residual for the block to be processed (similar to when a skip mode is applied), the prediction block can be used as a reconstruction block.
[0091] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, can be output through filtering as described later, or can be used for inter-frame prediction of the next frame. In addition, luminance mapping and chroma scaling (LMCS) can be applied in the frame decoding process.
[0092] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and send the modified reconstructed image to memory 360 (specifically, the DPB of memory 360). Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0093] The (modified) reconstructed frame stored in the DPB of memory 360 can be used as a reference frame in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current frame from which motion information is derived (or decoded) and / or motion information of blocks in previously reconstructed frames. The stored motion information can be sent to inter-frame predictor 332 as motion information of spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current frame and send them to intra-frame predictor 331.
[0094] The embodiments described in this document in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can also be applied equivalently or correspondingly to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300, respectively.
[0095] Figure 4 An image decoding method is shown, performed by a decoding device (300) according to an embodiment of the present disclosure.
[0096] Reference Figure 4 The transform coefficients of the current block can be derived from the bit stream (S400). That is, the bit stream may include residual information of the current block, and the transform coefficients of the current block can be derived by decoding the residual information.
[0097] Reference Figure 4 The residual sample of the current block can be derived by performing at least one of dequantization and inverse transformation on the transform coefficients of the current block (S410).
[0098] When applying Adaptive Multiple Transform Selection (MTS), the inverse transform can be performed based on at least one of DCT-2, DST-7, or DCT-8. Here, DCT-2, DST-7, DCT-8, etc., can be referred to as transform type, transform kernel, or transform core.
[0099] In this disclosure, the inverse transform can refer to a separable transform. However, it is not limited to this; the inverse transform can refer to an inseparable transform, or it can be a concept that includes both separable and inseparable transforms. Furthermore, the inverse transform in this disclosure refers to the principal transform, but is not limited to this; it can be applied to a second transform by modifying it to the same / similar form.
[0100] For example, as an inverse transform method, DCT-2 and the non-separable transform can be used alone, or the non-separable transform can be used in addition to at least one of DCT-2, DST-7 or DCT-8, or the non-separable transform can replace one or more of the transform kernels of DCT-2, DST-7 or DCT-8.
[0101] As a more specific implementation, when (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) are present as transform kernel candidates for separable transforms, a non-separable transform can replace or add to one or more of these five transform kernel candidates. Here, the notation (transformer 1, transformer 2) indicates that transform 1 is applied in the horizontal direction and transform 2 is applied in the vertical direction. When a non-separable transform replaces some transform kernel candidates, the remaining transform kernel candidates other than (DCT-2, DCT-2) and (DST-7, DST-7) can be replaced with a non-separable transform. However, the above transform kernel candidates are merely examples and may include other types of DCTs and / or DSTs, and may include transform skipping as a transform kernel candidate.
[0102] An inseparable transformation can refer to a transformation or inverse transformation based on an inseparable transformation matrix. That is, unlike a separable transformation, which performs horizontal and vertical transformations independently by separating the vertical and horizontal transformations, an inseparable transformation can perform both horizontal and vertical transformations simultaneously.
[0103] For example, when performing an inseparable transformation on a 4×4 block, the input data X of the inseparable transformation is shown in Equation 1 below.
[0104] [Formula 1]
[0105]
[0106] When the input data X is represented in vector form, the vector X' can be represented as follows.
[0107] [Equation 2]
[0108]
[0109] In this case, the inseparable transformation can be performed as shown in Equation 3 below.
[0110] [Formula 3]
[0111]
[0112] In Equation 3, F represents the transformation coefficient vector, T represents the 16×16 inseparable transformation matrix, and • represents the multiplication of the matrix and the vector.
[0113] The 16×1 transformation coefficient vector F can be derived using Equation 3. F can be reconfigured into 4×4 blocks according to a predetermined scanning order. The scanning order can be horizontal scanning, vertical scanning, diagonal scanning, z-scanning, raster scanning, or a predefined scanning.
[0114] The set of inseparable transforms and / or the transform kernel of inseparable transforms can be configured differently based on the prediction mode (e.g., intra-frame mode, inter-frame mode, etc.), the width, height or number of pixels of the current block, the position of the sub-blocks within the current block, the syntax elements explicitly notified by the signal, the statistical characteristics of the neighboring samples, and whether a quadratic transform or quantization parameter (QP) is used.
[0115] Specifically, for intra-frame modes, predefined intra-frame prediction modes can be grouped into n inseparable transform sets, each of which can include k transform kernel candidates. Here, n and k can be arbitrary constants, according to the same rules (conditions) defined for both the encoding and decoding devices.
[0116] The number of inseparable transform sets and / or the number of transform kernel candidates included in each inseparable transform set can be configured differently based on the width and / or height of the current block. For example, for a 4×4 block, n1 inseparable transform sets and k1 transform kernel candidates can be configured. For a 4×8 block, n2 inseparable transform sets and k2 transform kernel candidates can be configured. Furthermore, the number of inseparable transform sets and the number of transform kernel candidates included in each inseparable transform set can be configured differently based on the product of the width and height of the current block. For example, when the product of the width and height of the current block is equal to or greater than 256, n3 inseparable transform sets and k3 transform kernel candidates can be configured; otherwise, n4 inseparable transform sets and k4 transform kernel candidates can be configured. That is, since the degree of variation in the statistical characteristics of the residual signal varies with the block size, the number of inseparable transform sets and transform kernel candidates can be configured differently to reflect this.
[0117] When the current block is divided into multiple sub-blocks, the statistical characteristics of the residual signal can differ for each sub-block. Therefore, the number of inseparable transform sets and transform kernel candidates can be configured differently. For example, when a 4×8 or 8×4 block is divided into two 4×4 sub-blocks and an inseparable transform is applied to each sub-block, n5 inseparable transform sets and k5 transform kernel candidates can be configured for the top-left 4×4 sub-block, and n6 inseparable transform sets and k6 transform kernel candidates can be configured for the other 4×4 sub-blocks.
[0118] Based on syntax elements explicitly signaled by a signal, the number of inseparable transform sets and transform kernel candidates can be configured differently. As syntax elements, information indicating one of multiple inseparable transform configurations can be used. For example, when three inseparable transform configurations are supported (i.e., n7 inseparable transform sets and k7 transform kernel candidates, n8 inseparable transform sets and k8 transform kernel candidates, and n9 inseparable transform sets and k9 transform kernel candidates), syntax elements can have values of 0, 1, and 2, and the inseparable transform configuration applied to the current block can be determined based on the value of the syntax element signaled by the signal.
[0119] The number of inseparable transform sets and transform kernel candidates can be configured differently depending on whether a quadratic transform is applied and / or which quadratic transform is applied. For example, when no quadratic transform is applied, a set including n... 10 A set of inseparable transformations and k 10 An inseparable transformation configuration of n transformation kernel candidates. When applying a quadratic transformation, it is possible to apply n... 11 A set of inseparable transformations and k 11 Inseparable transformation configuration of a transform kernel candidate.
[0120] Based on the quantization parameter (QP) and / or the range to which the QP value belongs, different configurations of the inseparable transformation can be applied. For example, when the QP value is small, configurations including n can be applied. 12 A set of inseparable transformations and k 12 An inseparable transformation configuration of n transformation kernel candidates. On the other hand, when the QP value is large, an application including n... 13 A set of inseparable transformations and k 13 The non-separable transform configuration of each transform kernel candidate. When the QP value is less than or equal to a threshold (e.g., 32), the case is classified as having a smaller QP value; otherwise, the case is classified as having a larger QP value. Alternatively, the range of QP values can be divided into three or more, and different non-separable transform configurations can be applied to each range.
[0121] For relatively large blocks, instead of using an inseparable transformation corresponding to the block's width and height, the block can be divided into multiple sub-blocks, and an inseparable transformation corresponding to the width and height of each sub-block can be used. For example, when performing an inseparable transformation on a 4×8 block, the 4×8 block can be divided into two 4×4 sub-blocks, and an inseparable transformation based on the 4×4 block can be used for each 4×4 sub-block. Alternatively, an 8×16 block can be divided into two 8×8 sub-blocks, and an inseparable transformation based on the 8×8 block can be used.
[0122] The set of non-separable transforms can be determined based on the intra-prediction modes and mapping table of the current block. The mapping table defines the mapping relationship between predefined intra-prediction modes and the set of non-separable transforms. The predefined intra-prediction modes can include two non-directional modes and 65 directional modes. Typically, non-separable transforms have a larger transform kernel size than separable transforms. This means that the computational complexity required for transform processing is higher, and the memory required to store the transform kernel is larger. Furthermore, while separable transforms may only consider statistical properties existing in the horizontal and / or vertical directions, non-separable transforms can consider statistical properties in a two-dimensional space including both the horizontal and vertical directions, thus providing better compression efficiency. Since the statistical properties and residual diversity vary depending on the directionality of the intra-prediction mode, there may be cases where non-separable transforms are absolutely necessary, and there may be intra-prediction modes whose residual properties can be identified solely by separable transforms. Therefore, by predefining which transform to use in the encoding and decoding devices based on the intra-prediction modes, transform processing can be designed with optimized complexity and memory requirements. Non-directional modes can include the planar mode numbered 0 and the DC mode numbered 1, while directional modes can include intra-prediction modes numbered 2 to 66. However, this is just an example, and this disclosure can also be applied to situations where the number of predefined intra-prediction modes varies.
[0123] Due to the application of Wide Angle Intra Prediction (WAIP), the predefined intra prediction modes can also include intra prediction modes from -14 to -1 and intra prediction modes from 67 to 80.
[0124] Figure 5 An intra-frame prediction mode and its prediction direction according to this disclosure are illustrated exemplarily. (Refer to...) Figure 5 Patterns -14 to -1, 2 to 33, and 35 to 80 are symmetrical about pattern 34 in terms of prediction direction. For example, patterns 10 and 58 are symmetrical about the direction corresponding to pattern 34, and pattern -1 is symmetrical about pattern 67. Therefore, for vertical directional patterns that are symmetrical about pattern 34 and the horizontal directional patterns, the input data can be transposed and used. Transposing the input data means that the rows and columns in the M×N input data of the two-dimensional block are transformed into columns and rows, respectively, to form N×M data.
[0125] For example, when using 4×4 blocks, the 16 data points forming the 4×4 blocks can be appropriately arranged to form a 16×1 one-dimensional vector for the inseparable transformation. In this case, the one-dimensional vector can be formed in row-major or column-major order. The residual samples obtained from the inseparable transformation can be arranged in the above order to form a two-dimensional block.
[0126] For modes -14 to -1 and 2 to 33, the data arrangement order for forming a 16×1 input vector is row-major. For modes 35 to 80, the input vector can be formed according to column-major.
[0127] Pattern 34 cannot be considered either a horizontal or vertical orientation pattern, but in this disclosure, it is classified as a horizontal orientation pattern. That is, for patterns -14 to -1 and 2 to 33, the input data arrangement method for horizontal orientation patterns (i.e., row priority order) is used, and for vertical orientation patterns symmetrical about pattern 34, the input data can be transposed and used.
[0128] For non-square blocks, the symmetry in square blocks cannot be utilized (i.e., the symmetry between mode P and mode (68-P) in an N×N block (2<=P<=33) or the symmetry between mode Q and mode (66-Q) (-14<=Q<=-1)). Therefore, in addition to relying solely on the symmetry of intra-frame prediction modes, the symmetry between block shapes that are transposes of each other can also be utilized, i.e., the symmetry between K×L blocks and L×K blocks. Specifically, there is a symmetry relationship between a K×L block predicted by mode P and an L×K block predicted by mode (68-P). Alternatively, there is a symmetry relationship between a K×L block predicted by mode Q and an L×K block predicted by mode (66-Q).
[0129] Since a K×L block with mode 2 and an L×K block with mode 66 can be considered symmetrical to each other, the same transform kernel can be applied to both K×L and L×K blocks. If the set of non-separable transforms for the intra-prediction modes of a K×L block is mapped, then in order to apply the non-separable transform to an L×K block, the set of non-separable transforms can be derived through a mapping table corresponding to a K×L block based on mode (68-P) (rather than mode P applied to the L×K block). Alternatively, the set of non-separable transforms can be derived through a mapping table corresponding to a K×L block based on mode (66-Q) (rather than mode Q applied to the L×K block).
[0130] For example, to apply an inseparable transformation to an L×K block, the set of inseparable transformations can be selected based on mode 2 instead of mode 66. Furthermore, for a K×L block, the input data can be read in a predetermined order (e.g., row-major or column-major order) to form a 1D vector, and then the corresponding inseparable transformation can be applied. For an L×K block, the input data can be read in transposed order to form a 1D vector, and then the corresponding inseparable transformation can be applied. That is, when a K×L block is read in row-major order, an L×K block can be read in column-major order. Conversely, when a K×L block is read in column-major order, an L×K block can be read in row-major order.
[0131] Furthermore, when applying mode 34 to a K×L block, the set of inseparable transformations can be determined based on mode 34, and the input data can be read in a predetermined order to form a 1D vector and perform the corresponding inseparable transformation. When applying mode 34 to an L×K block, the set of inseparable transformations can be determined based on mode 34, but the input data can be read in transposed order to form a 1D vector and perform the corresponding inseparable transformation.
[0132] In this disclosure, a method for determining the set of inseparable transformations and a method for forming input data are described based on K×L blocks. However, the aforementioned symmetry of K×L blocks can be utilized to perform inseparable transformations based on L×K blocks. Alternatively, blocks with a width greater than their height can be restricted to being used as reference blocks. Alternatively, the symmetry can be restricted to not being used in the case of non-square blocks. In this case, non-square blocks can use a set of inseparable transformations and / or transformation kernel candidates with different numbers than those of square blocks, and a different mapping table can be used to select the set of inseparable transformations.
[0133] An example of a mapping table used to select sets of inseparable transforms is as follows:
[0134] [Table 1]
[0135]
[0136] Table 1 shows an example of assigning non-separable transform sets to each intra-prediction mode when five non-separable transform sets are available. The value of `predModeIntra` indicates the value of the intra-prediction mode considering WAIP, and `TrSetIdx` is the index indicating a specific non-separable transform set. In Table 1, it can be confirmed that the same non-separable transform set is applied to modes located in symmetrical directions according to the intra-prediction mode. Table 1 is merely an example of using five non-separable transform sets and does not limit the total number of non-separable transform sets used for non-separable transforms.
[0137] Alternatively, as shown in Table 2, for compression performance, non-separable transformations may not be applied to WAIP.
[0138] [Table 2]
[0139]
[0140] Alternatively, as shown in Table 3, instead of configuring a separate indivisible transform set for WAIP, the indivisible transform set corresponding to the prediction modes in adjacent frames can be shared.
[0141] [Table 3]
[0142]
[0143] The set of inseparable transforms can include multiple transform kernel candidates, and one of these candidates can be used selectively. This can be achieved using an index signaled via a bitstream. Alternatively, one of the multiple transform kernel candidates can be implicitly determined based on the context information of the current block. Here, the context information can refer to the size of the current block or whether an inseparable transform is applied to neighboring blocks. The size of the current block can be defined as its width, height, maximum / minimum width and height, the sum of width and height, or the product of width and height.
[0144] The method for determining the transform kernel for the inverse transform of the current block will be described in detail below.
[0145] Embodiment 1
[0146] As described above, inverse transforms can be divided into separable transforms and non-separable transforms. A separable transform means performing a transform on a two-dimensional block in the horizontal and vertical directions, respectively, while a non-separable transform means performing a single transform on a sample that constitutes the entire or part of the two-dimensional block. When representing a separable transform, it can be represented as a pair of horizontal and vertical transforms, which in this disclosure will be represented as (horizontal transform, vertical transform).
[0147] Multiple transformation sets can be defined for the inverse transform of the current block. Each transformation set can include one or more transform kernel candidates.
[0148] For example, one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), or (DCT-8, DCT-8) can be applied as a separable transform, and the above four transform kernel candidates can be considered as a transform set. Additionally, (DCT-2, DCT-2) can be considered as a transform set. A transform skip without applying a transform can also be considered as a transform set; (DCT-2, DCT-2) and the transform skip can be considered as a transform set. In this disclosure, a transform kernel can refer to a single transform (e.g., DCT-2, DST-7) or two transform pairs (e.g., (DCT-2, DCT-2)).
[0149] As another example of a transform set, the aforementioned inseparable transform set may exist. In this disclosure, the inseparable transform applied as the master transform can be represented as the Inseparable Master Transform (NSPT). In the NSPT, multiple inseparable transform sets can be configured, and each inseparable transform set may include one or more transform kernels as transform kernel candidates. In the case of the NSPT, one of multiple inseparable transform sets is selected based on the intra-frame prediction mode, and the multiple inseparable transform sets used for the NSPT can be represented as a list of NSPT sets. This is as described above, and its detailed description will be omitted here.
[0150] A group of one or more transform sets that can be used for the current block can be configured from multiple predefined transform sets. The group of one or more transform sets can be configured by a predetermined regional unit to which the current block belongs, hereinafter referred to as a set. Here, the predetermined regional unit can be at least one of a picture, a slice, a coding tree unit row (CTU row), or a coding tree unit (CTU).
[0151] For example, the transform set consisting of (DCT-2, DCT-2) is called S1, and the transform set consisting of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) is called S2. Furthermore, the above list of NSPT sets can include N inseparable transform sets, which are respectively called S... 3,1 S 3,2 ... S 3,N Here, N can be 35, but is not limited to this.
[0152] When the intra-prediction mode based on the current block is selected as S 3,13 When used as an inseparable transform set for NSPT, the transform kernel applicable to the current block can belong to S1, S2, or S... 3,13 One of them. In this case, the set available for the current block can be represented as {S1, S2, S...} 3,13}
[0153] As described above, since the set according to this disclosure is a group of one or more transform sets available for the current block, the set can be configured differently based on the context of the current block. Here, the context can include at least one of shape, size, or intra-prediction mode. If a total of K contexts are defined, K sets can be generated, and each set can be represented as C. i (i=1, 2, ..., N). For example, when the block size to which NSPT is applicable is 4×4, 8×8, 16×16 and 32×32 and one of a total of 35 inseparable transform sets is selected based on the intra-prediction mode, a total of 4×35=140 contexts can be defined if different transform kernels are applied for each block size.
[0154] The set can be configured based on the context of the current block, and in this case, the processes of selecting one of multiple transform sets belonging to the set and selecting one of multiple transform kernel candidates belonging to the selected transform set can be performed. Here, the selection of transform sets and transform kernel candidates can be performed implicitly based on the context of the current block, or it can be performed based on an index explicitly indicated by a signal. Alternatively, the processes of selecting one of multiple transform sets belonging to the set and selecting one of multiple transform kernel candidates belonging to the selected transform set can be performed separately. For example, an index for selecting a transform set can be first indicated by a signal, and one of multiple transform sets belonging to the set can be selected based on that index. Then, an index indicating one of the multiple transform kernel candidates belonging to the transform set can be indicated by a signal, and one of the transform kernel candidates can be selected from the transform set based on the index indicated by the signal. The transform kernel of the current block can be determined based on the selected transform kernel candidate. Alternatively, selecting a transform set from the set can be performed implicitly based on the context of the current block, and selecting a transform kernel candidate from the selected transform set can be performed based on an index indicated by a signal. Alternatively, selecting a transform set from the set can be performed based on an index signaled by a signal, and selecting a transform kernel candidate from the selected transform set can be implicitly performed based on the context of the current block. Alternatively, selecting a transform set from the set can be implicitly performed based on the context of the current block, and selecting a transform kernel candidate from the selected transform set can also be implicitly performed based on the context of the current block. Of course, when the number of transform sets belonging to the set is 1, signaling for the index of the selected transform set is not required. Similarly, when the number of transform kernel candidates belonging to the selected transform set is 1, signaling for the index of the transform kernel candidate is not required. Alternatively, signaling can be used to indicate the index of one of all transform kernel candidates belonging to the current set. In this case, the process of selecting a transform set from the set can be omitted. In this case, priority can be considered when shuffling all transform sets belonging to the set. For example, when assigning a smaller-length binary code (e.g., truncated unary code) to a smaller-value index, assigning a smaller-value index to a transform kernel candidate that is more conducive to improving coding performance may be advantageous. When shuffling all transformation kernel candidates belonging to a set according to priority, different shuffling can be applied to each set. Alternatively, instead of shuffling all transformation kernel candidates belonging to a set, it is possible to selectively shuffle only some of them.
[0155] Embodiment 2
[0156] The transform kernel for the inverse transform of the current block can be determined based on MTS (Multiple Transform Selection).
[0157] The MTS according to this disclosure can use at least one of DST-7, DCT-8, DCT-5, DST-4, DST-1 or IDT (identity transformation) as the transformation kernel. Additionally, the MTS according to this disclosure may also include a DCT-2 transformation kernel.
[0158] In this disclosure, multiple MTS sets can be defined for MTS. One of the multiple MTS sets can be determined based on the current block size and / or intra-prediction mode. For example, when determining an MTS set, 16 transform block sizes can be considered, and for directional modes, the shape of the transform blocks and the symmetry between the intra-prediction modes can be considered. For WAIP (Wide Angle Intra-Prediction) modes (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), the MTS set corresponding to mode 2 can be applied for modes -1 to -14 (or -15), and the MTS set corresponding to mode 66 can be applied for modes 67 to 80 (or 81). A separate MTS set can be assigned to MIP (Matrix-Based Intra-Prediction) modes.
[0159] For example, as shown in Table 4 below, an MTS set can be assigned / defined based on the transform block size and intra-prediction mode.
[0160] [Table 4]
[0161]
[0162] Table 4 shows the assignment of MTS sets based on 16 transform block sizes and intra-prediction modes. The predefined number of MTS sets is 80, and the index indicating one of the 80 MTS sets can have values from 0 to 79, as shown in Table 4.
[0163] [Table 5]
[0164]
[0165]
[0166]
[0167] Table 5 shows the transform kernel candidates included in the various MTS sets described in Table 4. Each MTS set can consist of six transform kernel candidates. The transform kernel candidate index has a value of one of 0 to 5 and can indicate one of the six transform kernel candidates. Here, each transform kernel candidate can be a combination of horizontal and vertical transform kernels for separable transforms, and 25 transform kernel candidates with indices of 0 to 24 can be defined.
[0168] [Table 6]
[0169]
[0170] Table 6 provides examples of the 25 transform kernel candidates described in Table 5. Specifically, the horizontal and vertical transforms of the transform kernel candidates are represented as (horizontal transform, vertical transform). For each transform kernel candidate index, the horizontal / vertical transform when the intra-prediction mode is less than 35 can be the opposite of the horizontal / vertical transform when the intra-prediction mode is greater than or equal to 35. When the value of the intra-prediction mode is greater than or equal to 35, a mode symmetric about mode 34 can be derived, and the MTS set can be selected from Table 4 based on this mode. Additionally, the symmetry of the block shape can be considered. When the original transform block has a size of W×H, by symmetry, the original transform block can be considered to have a size of H×W, and the MTS set can be selected from Table 4. Here, the value of the intra-prediction mode can be a modified value of the intra-prediction mode. That is, as the mode value of WAIP, for values from -14 (or -15) to -1, it is modified to mode 2; for values from 67 to 80 (or 81), it is modified to mode 66; and for the remaining modes, the value of the original intra-prediction mode can be set to the value of the modified intra-prediction mode. In this case, since the extended mode of WAIP is also symmetrically configured with respect to mode 34, the symmetry with respect to mode 34 can be applied to all orientation modes except for planar mode and DC mode.
[0171] For example, when predicting a 16×32 block based on pattern 54, pattern 14 (=68-54) can be derived as a pattern symmetric to pattern 54, and the block size can be considered as 32×16. In this case, an MTS set with index 72 can be selected, as defined in Table 4.
[0172] When applying MIP mode, the MTS set assigned to the MIP mode can be selected based on the current block size, regardless of the block shape symmetry. Alternatively, when applying MIP mode, the block shape symmetry can be considered when selecting the MTS set assigned to the MIP mode based on the size of the symmetrical block. For example, when applying MIP mode to an 8×16 block, the 8×16 block can be considered as a symmetrical 16×8 block, and the MTS set with index 49 can be selected as defined in Table 4. Alternatively, when applying MIP mode, the intra-prediction mode can be considered as a planar mode. In this case, the MTS set assigned to the MIP mode can be selected based on the current block size, regardless of the block shape symmetry. Alternatively, the block shape symmetry can be considered when selecting the MTS set assigned to the MIP mode based on the size of the symmetrical block.
[0173] For MIP mode, a flag can be used to indicate whether MIP mode is applied in transposed mode. When MIP mode is applied to an M×N current block and the flag indicates that transposed mode is applied, the intra-prediction mode can be treated as a planar mode, and the M×N current block can be treated as an N×M block. That is, from Table 4, the MTS set corresponding to an N×M block size and a planar mode can be selected. As described in Table 6, when the value of the intra-prediction mode is greater than or equal to 35, the horizontal and vertical transforms are swapped, but since the intra-prediction mode of the current block is treated as a planar mode, the horizontal and vertical transforms of the transform kernel candidates may not be swapped. Alternatively, when MIP mode is applied to an M×N current block and the flag indicates that transposed mode is applied, the intra-prediction mode may not be treated as a planar mode, and the M×N current block can be treated as an N×M block. That is, from Table 4, the MTS set corresponding to an N×M block size and MIP mode can be selected.
[0174] In Table 5, the transform kernel candidate selected by the transform kernel candidate index can be set as the transform kernel of the current block. Alternatively, based on the size of the current block, at least one of the horizontal or vertical transforms of the selected transform kernel candidate can be changed to another transform kernel. For example, when the transform kernel candidate index is 3 and both the width and height of the current block are less than or equal to 16, at least one of the horizontal or vertical transforms of the transform kernel candidate corresponding to transform kernel candidate index 3 can be changed to another transform kernel. In this case, the horizontal and vertical transforms can be changed independently of each other. When the difference (or the absolute value of the difference) between the value of the intra-prediction mode and the value of the horizontal mode of the current block is less than or equal to a predetermined threshold, the vertical transform of the selected transform kernel candidate can be changed to IDT (identity transform). When the difference (or the absolute value of the difference) between the value of the intra-prediction mode and the value of the vertical mode of the current block is less than or equal to a predetermined threshold, the horizontal transform of the selected transform kernel candidate can be changed to IDT (identity transform). Here, the threshold can be determined based on the width and height of the current block, as shown in Table 7 below.
[0175] [Table 7]
[0176]
[0177] Table 7 is used to change the horizontal and / or vertical transforms of transform kernel candidates selected by the transform kernel candidate index to another transform kernel, and the threshold is defined according to the size of the transform block.
[0178] The six transform kernel candidates that make up an MTS set can be distinguished by transform kernel candidate indices from 0 to 5, as defined in Table 5. The transform kernel candidate indices can be signaled via a bitstream. A flag indicating whether the MTS set is available / applied (MTS enable flag or MTS flag) can be signaled, and when this flag indicates that the MTS set is available / applied, the transform kernel candidate indices can be signaled. The MTS flag can consist of a bin, and one or more contexts (hereinafter referred to as CABAC contexts) can be assigned to this bin for CABAC-based entropy coding. For example, different CABAC contexts can be assigned to non-MIP mode and MIP mode respectively.
[0179] Based on the context of the current block, the number of transform kernel candidates available for the current block can be set differently. For example, as the context of the current block, the sum of the absolute values of all or some transform coefficients in the current block can be considered. The sum of the absolute values of the transform coefficients is called AbsSum. When AbsSum is less than or equal to T1, only one transform kernel candidate corresponding to transform kernel candidate index 0 may be available. When AbsSum is greater than T1 and less than or equal to T2, four transform kernel candidates corresponding to transform kernel candidate indices 0 to 3 may be available. When AbsSum is greater than T2, six transform kernel candidates corresponding to transform kernel candidate indices 0 to 5 may be available. Here, T1 can be 6 and T2 can be 32, but this is only an example.
[0180] When AbsSum is less than or equal to T1, since the number of transform kernel candidates available for the current block is 1, the transform kernel candidate corresponding to transform kernel candidate index 0 can be set as the transform kernel for the current block without signaling the transform kernel candidate index. When AbsSum is greater than T1 and less than or equal to T2, since four transform kernel candidates are available, one of the four transform kernel candidates can be selected based on the transform kernel candidate index with two bins. That is, transform kernel candidate indices 0 to 3 can be signaled as 00, 01, 10, and 11 respectively. For these two bins, the MSB (most significant bit) can be signaled first, and the LSB (least significant bit) can be signaled later. Different CABAC contexts can be assigned to each bin. For example, a CABAC context other than the CABAC context assigned for the MTS flag can be assigned to each bin of the two bins. Alternatively, bypass coding can be applied without assigning CABAC contexts to the two bins. When AbsSum is greater than T2, the transform kernel candidate index has values from 0 to 5, making it impossible to represent the transform kernel candidate index with only two bins. In this case, the transform kernel candidate index can be represented by assigning two or more bins, for example, by truncating binary encoding. For each bin assigned by the truncated binary encoding method, a CABAC context can be assigned, or bypass encoding can be applied without assigning a CABAC context. Alternatively, a CABAC context can be assigned to some of the multiple bins (e.g., the first bin, or the first and second bins), and bypass encoding can be applied to the remaining bins.
[0181] Embodiment 3
[0182] The transform kernel of the current block can be determined based on a transform set that includes one or more transform kernel candidates. The transform kernel of the current block can be derived as one of the one or more transform kernel candidates belonging to the transform set.
[0183] The process of determining the transform kernel of the current block may include at least one of 1) determining the transform set of the current block or 2) selecting a transform kernel candidate from the transform set of the current block. The process of determining the transform set may be the process of selecting one of a plurality of identical predefined transform sets in the encoding and decoding devices. Alternatively, the process of determining the transform set may be the process of configuring one or more transform sets available for the current block from a plurality of identical predefined transform sets in the encoding and decoding devices, and selecting one of the configured transform sets. Alternatively, the process of determining the transform set may be the process of configuring a transform set based on transform kernel candidates available for the current block from a plurality of identical predefined transform kernel candidates in the encoding and decoding devices.
[0184] When the transform set of the current block includes multiple transform kernel candidates, the process of selecting one of the multiple transform kernel candidates for the current block can be performed. However, when the transform set of the current block includes only one transform kernel candidate (i.e., when the number of transform kernel candidates available for the current block is 1), the transform kernel of the current block can be set as the corresponding transform kernel candidate.
[0185] The transform set according to this disclosure may refer to the (inseparable) transform set in Embodiment 1 above, or it may refer to the MTS set in Embodiment 2. Alternatively, the transform set may be defined separately from the (inseparable) transform set in Embodiment 1 or the MTS set in Embodiment 2. In this case, the transform set may include one or more specific transform kernels as transform kernel candidates. A specific transform kernel may be defined as a pair of transform kernels for horizontal transformation and a transform kernel for vertical transformation, or it may be defined as a single transform kernel applied equally to both horizontal and vertical transformations.
[0186] In embodiments of this disclosure, the process of applying NSPT (an inseparable transform applied as the master transform) is described in detail. NSPT can be applied to the entire or a portion of a transform block. Based on forward NSPT, residual samples existing in the region where NSPT is applied can be used as 1D vector inputs to NSPT. In other words, residual samples existing in the entirety or a portion of a single transform block (referred to in this disclosure as the region of interest (ROI)) can be collected as 1D vectors and configured as inputs. Then, when forward NSPT is applied, the master transform coefficients can be obtained. Conversely, when backward NSPT is applied to the master transform coefficients, a 1D vector output can be obtained. The residual samples of the ROI can be obtained by arranging the element values of the corresponding output vector at defined positions within the 2D transform block.
[0187] For a non-separable transform kernel used in NSPT, the matrix dimension can be determined based on the size of the Region of Interest (ROI). In this disclosure, the transform kernel can be referred to as a transform type or a transform matrix, and a non-separable transform kernel used in NSPT can be referred to as an NSPT kernel. For example, when the current block is an M×N transform block, the ROI is the entire region of the M×N transform block, and a square NSPT is applied, the dimension of the corresponding transform matrix can be MN×MN. For example, when the ROI is the entire region of an 8×8 transform block, the dimension of the NSPT kernel can be 64×64.
[0188] According to embodiments of this disclosure, when applying NSPT to residuals generated by intra-frame prediction, the NSPT kernel can be adaptively determined based on the intra-frame prediction mode. Since the statistical characteristics of the residual block can vary according to the intra-frame prediction mode, compression efficiency can be improved by adaptively determining the NSPT kernel based on the intra-frame prediction mode.
[0189] A shared NSPT kernel can be configured to be applied to at least one intra-prediction mode. As described above, the set of inseparable transforms can be determined based on the intra-prediction mode and mapping table of the current block. The mapping table can define the mapping relationship between predefined intra-prediction modes and the set of inseparable transforms. The predefined intra-prediction modes can include two non-directional modes and 65 directional modes.
[0190] As an implementation method, intra-prediction modes can be grouped into intra-prediction mode groups. One NSPT kernel can be assigned to an intra-prediction mode group, or multiple NSPT kernels can be assigned. In other words, an inseparable transform set (NSPT set) including at least one NSPT kernel can be assigned to an intra-prediction mode group. The inseparable transform set can be mapped to an intra-prediction mode, and one of the N NSPT kernels included in the inseparable transform set can be selected.
[0191] As an example, intra-prediction groups may include adjacent prediction modes (e.g., modes 17, 18, and 19). Additionally, intra-prediction groups may include modes with symmetry. For example, in the above... Figure 5 In this context, directional modes can be symmetrical about a diagonal mode (i.e., intra-prediction mode 34). In this case, two symmetrical modes can be configured as a set (or a pair). For example, modes 18 and 50 can be included in the same set because they are symmetrical about mode 34. However, for modes with symmetry, a process of transposing the 2D input block and then configuring the one-dimensional input vector can be added before applying the feedforward NSPT kernel. For example, when the intra-prediction mode is less than or equal to 34, the one-dimensional input vector can be derived from the corresponding input block in row-major order without transposing the 2D input block. When the intra-prediction mode is greater than 34, the one-dimensional input vector can be configured either by first transposing the 2D input block and then reading the corresponding input block in row-major order, or by keeping the 2D input block as is and reading the corresponding input block in column-major order.
[0192] Table 8 below shows the mapping table for allocating NSPT sets according to intra-prediction modes. Referring to Table 8, a total of 35 NSPT sets from 0 to 34 can be defined. The NSPT set assigned to the most recent general directional mode can be assigned to the extended WAIP mode (i.e., Figure 5 (Modes -14 to -1 and modes 67 to 80 in the dataset). In other words, NSPT set 2 can be assigned to extended WAIP modes.
[0193] [Table 8]
[0194]
[0195] The NSPT set may include at least one NSPT core (or core candidate). In other words, the NSPT set may include N NSPT core candidates. As an example, N may be set to a value equal to or greater than 1, such as 1, 2, 3, 4, etc. The core applicable to the current block among the at least one NSPT core included in the NSPT set can be signaled using an index. In this disclosure, the corresponding index may be referred to as the NSPT index. As an example, the NSPT index may have values of 0, 1, 2, ..., N-1.
[0196] Furthermore, as an implementation, when the number of NSPT core candidates is 1, the NSPT index value can be fixed at 0. In this case, the NSPT index can be inferred without separate signal notification. Additionally, the flag indicating whether to apply NSPT can be signaled separately from the NSPT index. In this disclosure, the corresponding flag can be referred to as the NSPT flag.
[0197] NSPT can be applied when the NSPT flag value is 1. NSPT can be omitted when the NSPT flag value is 0. The NSPT flag value can be inferred as 0 when it is not signaled. As an example, the NSPT index can be applied when the NSPT flag value is 1. One of the N kernel candidates included in the NSPT set selected by the intra-prediction mode can be specified based on the signaled NSPT index.
[0198] In implementation, the entropy encoding method for the NSPT index can be defined in various ways by taking into account the number (N) of NSPT kernels included in the NSPT set. For example, as a method of mapping values from 0 to N-1 to bin strings (i.e., binarization method), truncated unary binarization, truncated binarization, and fixed-length binarization methods can be used.
[0199] For example, when the number N of kernel candidates in the configured NSPT set is 2, one bin can be used to specify one of the two candidates. For example, 0 can indicate the first candidate, and 1 can indicate the second candidate. Alternatively, when N is 3 and truncated unary binarization is applied, two bins can be used to specify the candidate. For example, the first, second, and third candidates can be binarized to 0, 10, and 11 respectively and signaled. As an implementation, the binarization bins can be encoded using context coding or bypass coding.
[0200] This disclosure describes a Reduced Principal Transform (RPT) method using a dimensionality reduction transform kernel as the principal transform. As described above, when applying forward NSPT, samples belonging to a 2D residual block can be arranged (or rearranged) into 1D vectors according to row-major order (or column-major order). The transformation matrix used for NSPT can then be multiplied by the arranged vectors. When the corresponding 2D residual block is an M×N block (M is the horizontal length, N is the vertical length), the length of the rearranged 1D vector can be M*N. In other words, the corresponding 2D residual block can also be represented as an M*N×1 dimensional column vector. In this disclosure, for convenience, M*N can be represented as MN. In this case, the dimension of the corresponding transformation matrix can be MN×MN. In summary, forward NSPT can be performed by multiplying the left side of the MN×1 vector by the corresponding MN×MN transformation matrix to obtain an MN×1 transformation coefficient vector.
[0201] When applying RPT, the r transformation coefficients can be obtained by multiplying by an r×MN matrix instead of an MN×MN matrix as the aforementioned forward NSPT transformation matrix. Here, r represents the number of rows in the transformation matrix, and MN represents the number of columns in the transformation matrix. According to the embodiments of this disclosure, the value of r can be set to be less than or equal to MN. In other words, the existing forward NSPT transformation matrix includes MN rows, and each row consists of 1×MN row vectors and the corresponding transformation basis vectors of the NSPT transformation matrix. The corresponding transformation coefficients can be obtained by multiplying each transformation basis vector by MN×1 sample column vectors.
[0202] Since the existing forward NSPT transformation matrix consists of MN row vectors, MN transformation coefficients (i.e., MN×1 transformation coefficient column vectors) can be obtained by applying forward NSPT. Furthermore, for forward RPT, the transformation matrix can consist of r transformation basis vectors instead of MN transformation basis vectors. Therefore, when applying forward RPT, r transformation coefficients (r×1 transformation coefficient column vectors) can be obtained instead of MN.
[0203] The RPT kernel can be configured by selecting r transform basis vectors as partial transform basis vectors for configuring the MN×MN forward NSPT kernel. In this disclosure, the transform kernel can be referred to as a transform type or transform matrix, and the inseparable transform kernel used for NSPT can be referred to as the RPT kernel. In other words, when selecting r 1×MN row vectors from the MN×MN forward NSPT kernel, it may be advantageous from a coding performance perspective to select the most important transform basis vectors. Specifically, in terms of energy concentration through the transform, by multiplying by the forward NSPT transform matrix, more energy can be concentrated on the first-appearing transform coefficients. In other words, the transform basis vectors located at the top of the forward NSPT transform matrix can generate transform coefficients with greater energy. With this in mind, the r×MN forward RPT kernel can be configured (or derived) by taking r from the top of the forward NSPT kernel.
[0204] According to this disclosure, the RPT only takes a portion (i.e., r) of the transform coefficients obtained by applying the existing NSPT; therefore, the energy of the original signal may be partially lost. In other words, distortion between the original signal and the natural signal may occur through correspondence processing. However, since only r transform coefficients are generated by applying the RPT instead of MN, the number of bits required to encode the corresponding transform coefficients can be reduced. Therefore, for signals with a large amount of energy concentrated on a small number of transform coefficients (e.g., image residual signals), the gain obtained by reducing signaling bits can be significantly large, thereby improving coding performance.
[0205] The backward NSPT is a transformation matrix, and can be the transpose of the aforementioned forward NSPT kernel. In this case, the input data can be the transform coefficient signal, rather than a sample signal such as a residual signal. Specifically, when the forward NSPT transformation matrix is G and the sample signal rearranged into a 1D vector is x, the transform coefficient vector obtained by multiplying the corresponding transformation matrix by the left side can be represented as shown in Equation 4 below.
[0206] [Formula 4]
[0207] y=Gx
[0208] Referring to Equation 4, x and y can be MN×1 column vectors. G can be in the form of an MN×MN matrix. The backward NSPT process can be represented using the same variables as in Equation 5 below.
[0209] [Formula 5]
[0210] x=G T y
[0211] In Equation 5, G TThis refers to the transpose of G. The forward RPT and backward RPT operations according to this disclosure can also be represented by these two equations. However, when RPT is applied, y is an r×1 column vector instead of an MN×1 column vector, and G is an r×MN matrix instead of an MM×MN matrix. In other words, even when RPT is applied instead of NSPT, the dimension of the sample signal (e.g., the image residual signal) does not change, which may mean that the original number of sample signals (i.e., the MN sample signal) can be reconstructed using only r transform coefficients via backward RPT. In other words, the original MN sample signal can be reconstructed by encoding only r transform coefficients less than MN, which improves coding performance.
[0212] In embodiments of this disclosure, an RPT structure is proposed that defines the value of r by considering the statistical properties of the residual block and derives a residual block of the existing transform block size from a residual block of reduced size determined according to the defined value of r. If another additional transformation (i.e., a quadratic transformation) is applied to predict the statistical distribution of the master transform coefficients, quantization is applied to the master transform coefficients, so that the quantized non-zero coefficients can be concentrated in a relatively low-frequency domain. Therefore, the reduced quadratic transformation of the statistical distribution of the master transform coefficients can define the statistical properties of the master transform coefficients relatively simply by setting the value of r for a given low-frequency domain. However, the RPT according to this disclosure differs fundamentally from the reduced quadratic transformation, which defines the value of r by considering the statistical properties of samples within the residual block whose properties are very different from the distribution of the master transform coefficients. Hereinafter, various embodiments for determining the RPT kernel as the dimension reduction transform matrix are described. In other words, methods for determining or defining the value of r in the RPT are described below.
[0213] In embodiments of this disclosure, the value of r in the RPT can be determined by considering the worst-case complexity allowed by the transformation system. As an implementation, the worst-case complexity can be calculated based on the number of multiplications per sample. MN*r multiplications are required to apply the RPT to both the forward and backward directions based on an M×N block. Since a 2D block consists of a total of MN samples, the number of multiplications per sample can be calculated as (MN*r) / MN = r. Therefore, the value of r can be configured to remain less than or equal to the maximum allowed number of multiplications per sample. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16×16 block, the value of r can be determined to be less than or equal to 16. In other words, the forward RPT kernel can be set to 16×256.
[0214] In another implementation, memory usage can be considered a measure of worst-case complexity. As an example, the allowed memory size per core can be set. For instance, when each core coefficient requires p bytes (in this disclosure, the elements configuring the transform core are referred to as core coefficients) and memory usage is set to be less than or equal to q bytes per core, the value of r can be set to be less than or equal to q / (MN*p). For example, when for a 16×16 block forward RPT core, p is 1 byte, and memory usage is set to be less than or equal to 8KB per core (q = 8 KB = 2...). 13 When the value of r is less than or equal to 32 bytes, the value of r can be set to less than or equal to 32 bytes.
[0215] Additionally, as another example, memory usage and / or the number of multiplications per sample can be considered as a measure of worst-case complexity. For instance, when the maximum possible number of multiplications per sample is set to 16 for a 16×16 block, and memory usage is set to less than or equal to 8KB per core (the core coefficient is represented as 1 byte), the value of r can be set to less than or equal to 16.
[0216] Furthermore, in the implementation, the value of 'r' for configuring the RPT core can be determined by specific information. In other words, the value of 'r' for configuring the RPT core can be determined based on predefined coding parameters. For example, the value of 'r' can be determined based on the block size. In other words, the RPT core can be variably determined based on the block size. Here, the block can be at least one of a coding block, a transform block, and a prediction block. Additionally, for example, the value of 'r' can be determined based on prediction information. Here, prediction information can include information about inter-frame / intra-frame prediction, intra-frame prediction mode information, etc. Additionally, for example, the value of 'r' can be determined based on information communicated by signals (the values of syntax elements). For example, the value of 'r' can be variably determined based on quantization parameter values. Furthermore, regarding complexity improvement, a fixed value predefined as the value of 'r' can be used, and this predefined fixed value can be determined based on information communicated by signals.
[0217] When the sample signal is multiplied by the RPT kernel r×MN, r transform coefficients are obtained. These r transform coefficients can be arranged according to a predefined scan order (e.g., forward / backward zigzag scan order, forward / backward horizontal scan order, forward / backward vertical scan order, forward / backward diagonal scan order, scan order specified based on the intra-frame prediction mode, etc.). When the transform coefficients obtained by applying the forward RPT are arranged according to this scan order (e.g., a scan order in units of coefficient groups (CGs) can also be applied), if the value of r is less than MN, the interior of the M×N block may not be completely filled by the r transform coefficients, thus potentially resulting in blank spaces. As an embodiment of this disclosure, the characteristics of the residual signal can be considered to predict the aforementioned blank spaces in the following manner.
[0218] - You can fill the blank space with the values of available neighboring pixels.
[0219] - The value of the blank space can be filled based on the values of available neighboring pixels and the intra-prediction mode. For example, the value of the blank space can be predicted by performing intra-prediction based on the values of available neighboring pixels and the intra-prediction mode.
[0220] - You can fill empty spaces with values that are predefined fixed values (e.g., 0).
[0221] - Values can be used to fill the blank space from available neighboring pixels by using a predetermined intra-frame prediction mode (e.g., planar mode).
[0222] In this disclosure, filling the blank space with 0 in the above example can be referred to as zeroing. When filling the blank space with 0, the following implementation can be applied. When a non-zero transform coefficient is detected (or resolved) in the corresponding blank space portion during the resolution of transform coefficients on the decoding device side, it can be considered (or inferred) that RPT is not applied. In other words, when a non-zero transform coefficient exists in a predefined region representing the corresponding blank space, it can be considered that RPT is not applied. In this case, signaling (or resolution) indicating whether RPT is applied and / or specifying an index of one of a plurality of RPT core candidates may not be executed. As an example, when a non-zero transform coefficient exists in a predefined region representing the corresponding blank space, the predefined variable value can be updated, and it can be inferred that RPT is not applied based on the updated variable value.
[0223] In embodiments of this disclosure, the application of RPT can be determined based on the size and / or form of the block. Furthermore, the RPT kernel can be determined differently depending on the size and / or form of the block. Since the value of r can vary depending on the size and / or form of the block (i.e., for individual M×N blocks), the blank space can also vary depending on the size and / or form of the block. Therefore, regions for checking whether non-zero transform coefficients are detected can be defined differently for blocks of different sizes and / or forms. In other words, zeroing regions can be determined differently.
[0224] As an example, when a 16×64 matrix is applied as the forward RPT matrix for an 8×8 block, the value of r can be 16. In this case, when CG is a 4×4 sub-block, only the top-left 4×4 block can be filled with non-zero RPT transform coefficients, while the remaining three 4×4 sub-blocks (i.e., the top-right, bottom-left, and bottom-right sub-blocks) can be filled with values of 0. In this case, when non-zero transform coefficients are detected in the corresponding remaining three 4×4 sub-block regions during decoding, it can be considered that RPT is not applied. Furthermore, as mentioned above, a flag indicating whether RPT is applied or an index of one of multiple RPT core candidates can be specified without signaling.
[0225] Additionally, as an example, when a 32×128 matrix is applied as the forward RPT matrix for a 16×8 block (i.e., the value of r is 32) and the CG is a 4×4 sub-block, only the two CGs in the scan order can be filled with non-zero RPT transform coefficients. For example, the top-left 4×4 sub-block and the 4×4 sub-block adjacent to the bottom of the top-left sub-block can be filled with the corresponding RPT transform coefficients. The area filled with 0 as blank space can be determined as the remaining area besides the corresponding two 4×4 sub-blocks. The RPT kernel can be determined differently depending on the size and / or form of the block, and as described, the blank space can be determined differently for 8×8 blocks and 16×8 blocks.
[0226] As an implementation method, when the value of r is a multiple of the CG size and the transform coefficients are scanned in units of CGs, if a non-zero transform coefficient is detected in a CG belonging to a blank space, the flags and / or indices related to RPT do not need to be signaled. In other words, the transform coefficients within each CG can be scanned in a specified order, and the scan sequence can be moved to the next CG in units of CG and the transform coefficients within the CG can be scanned in the same way. In existing image compression techniques, since a flag indicating the presence of non-zero transform coefficients in the corresponding CG is first signaled for each CG, it is possible to determine whether to apply RPT using only the corresponding information, which reduces signaling overhead and related implementation complexity.
[0227] As described above, when applying RPT, if a non-zero transform coefficient is detected in a zero-filled blank space region, RPT may not be applied. In this case, signaling related to RPT information can be omitted. However, since it is impossible to determine whether to apply RPT when no non-zero transform coefficient is detected in the corresponding blank space region, the flag indicating whether RPT is applied can be resolved after resolving (or signaling) the relevant transform coefficients to ultimately determine whether RPT is applied.
[0228] As an implementation method, a forward quadratic transform can be applied to the transform coefficients generated by applying the RPT. Alternatively, a forward quadratic transform can be applied to the region containing the generated transform coefficients in the M×N block. In this disclosure, from the perspective of the forward quadratic transform, the corresponding region or a portion of the corresponding region can be referred to as a Region of Interest (ROI). For the backward direction, a backward quadratic transform can be applied first, followed by a backward RPT. Specifically, a region containing r transform coefficients generated by applying the forward RPT, or a portion of the corresponding region, can be designated as an ROI to apply the forward quadratic transform. In this case, when a 16×64 forward RPT transform matrix is applied to an 8×8 region, the 16 generated transform coefficients can be located in the upper left 4×4 sub-block, and the corresponding sub-block region can be designated as an ROI to apply the forward quadratic transform to the corresponding ROI.
[0229] Furthermore, the RPT kernel can adjust coefficient values by incorporating operations such as integer or fixed-point arithmetic. In other words, the RPT kernel can be configured to perform transformations in a practical encoding / decoding system via integer (or fixed-point) arithmetic by appropriately scaling the kernel coefficients belonging to the corresponding kernel (rather than theoretically orthogonal or non-orthogonal transformations (where orthogonal and non-orthogonal transformations refer to transformations where the norm of each transform basis vector is 1)). Even when applying RPT, it can be reflected as equivalently as multiplying by a scaling factor when applying separable transformations in existing image compression techniques. In this case, separable or non-separable transformations (including RPT) can be performed while maintaining other processing besides the transformation (e.g., quantization and dequantization).
[0230] The integer coefficients of the RPT kernel can be obtained by multiplying the transform basis vector by the scaling value described above. As an implementation, multiplying by the scaling value may include applying operations such as rounding, flooring, and flooring to each kernel coefficient. In other words, an integer RPT kernel obtained by the above method can be defined and used for transform / inverse transform processing. As mentioned above, when the scaled integer kernel coefficients are obtained through operations such as rounding, flooring, and flooring, the maximum and minimum values can be obtained for all kernel coefficients, thus providing a sufficient number of bits to represent all kernel coefficients. For example, when the maximum value is less than or equal to 127 and the minimum value is greater than or equal to -128, all integer kernel coefficients can be represented in 8 bits (especially through a complement expression of 2, etc.).
[0231] Typically, when the maximum value is less than or equal to (2) (N-1) -1) and the minimum value is greater than or equal to -2. (N-1) When all integer kernel coefficients can be represented by N bits, when the maximum value is greater than (2^N, ... (N-1) -1) or the minimum value is less than -2 (N-1) At this point, it may be impossible to represent all integer kernel coefficients with N bits. In this case, 1) all kernel coefficients can be multiplied by a scaling value to adjust them to fall within the N-bit range, or 2) the number of bits required to represent the kernel coefficients can be increased (i.e., N+1 bits or more). When all kernel coefficients need to be multiplied by 2... -p (p>=1) When representing them with N bits, it can be done by multiplying by 2 afterwards. -p To compensate for them, so that they can be integrated into the existing encoding / decoding process. As an implementation, multiply by 2. p This can be achieved by performing an additional left shift operation of p bits or by reducing the right shift amount applied in the quantization or dequantization process by p.
[0232] The above method can be used to represent all kernel coefficients in 8 bits, 9 bits, 10 bits, etc. Of course, the scaling value of the kernel coefficients can be set for different block sizes or kernels, and the number of bits used to represent the kernel coefficients can be set differently.
[0233] The above-described NSPT can be applied based on at least one of the current block size, tree type, or component type. As an example, it can be determined whether to apply NSPT based on at least one of the current block size, tree type, or component type. The NSPT index can be signaled based on at least one of the current block size, tree type, or component type. An NSPT set or NSPT kernel can be derived based on at least one of the current block size, tree type, or component type.
[0234] The predefined allowed transform block sizes in the decoding device can be broadly divided into two groups. Either group (hereinafter referred to as the first group) can refer to the set of block sizes to which NSPT applies. The first group can consist of any one allowed transform block size, or it can consist of two or more allowed block sizes. The block size to which NSPT applies can be defined as a block size where at least one of the width and height is less than or equal to a predetermined threshold. Alternatively, the block size to which NSPT applies can be defined as a block size where the product of the width and height is less than or equal to a predetermined threshold. Alternatively, the block size to which NSPT applies can be defined as a block size where the maximum value of the width and height is less than or equal to a predetermined threshold. The threshold can be an integer of 4, 8, 16, 32, 64, 128, or greater.
[0235] The other group (hereinafter referred to as the second group) may refer to the set of block sizes for which NSPT is not applied. The aforementioned separable principal transformation can be applied to the block sizes belonging to the second group. Alternatively, the non-separable quadratic transformation can be applied to all or some of the block sizes belonging to the second group.
[0236] For example, when the current block size belongs to the first group, a backward NSPT can be applied to the (dequantized) transform coefficients of the current block. When the current block size belongs to the second group, a backward separable principal transform can be applied to the (dequantized) transform coefficients of the current block. Alternatively, when the current block size belongs to the second group, a backward non-separable quadratic transform (e.g., the low-frequency non-separable transform LFNST) can be applied to the (dequantized) transform coefficients of the current block first, and a backward separable principal transform (e.g., DCT-2) can be applied to the transform coefficients obtained therefrom.
[0237] For example, as a set of block sizes to which NSPT applies, the first group can be defined as a set of 4×4, 4×8, 8×4, and 8×8. Alternatively, the first group can be defined as a set of 4×8, 8×4, and 8×8. Alternatively, the first group can be defined as a set of 4×8 and 8×4. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 8×4, 8×8, and 16×4. Alternatively, the first group can be defined as a set of 4×8, 4×16, 8×4, 8×8, and 16×4. Alternatively, the first group can be defined as a set of 4×8, 4×16, 8×4, and 16×4. Alternatively, the first group can be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, and 16×16. Alternatively, the first group can be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, and 16×8. Alternatively, the first group can be defined as a set of 4×8, 8×4, 8×8, 8×16, and 16×8. Alternatively, the first group can be defined as a set of 4×8, 8×4, 8×16, 16×8, 16×16, 16×32, 32×16, and 32×32. Alternatively, the first group can be defined as a set of 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, and 32×16. Alternatively, the first group can be defined as a set of 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, and 32×16. Alternatively, the first group can be defined as a set of 4×8, 8×4, 8×16, 16×8, 16×16, 16×32, and 32×16. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 8×4, and 16×4. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 8×4, 8×8, 8×16, 16×4, and 16×8. Alternatively, the first group can be defined as a set of 4×8, 4×16, 8×4, 8×8, 8×16, 16×4, and 16×8. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 8×4, 8×16, 16×4, and 16×8. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 8×4, 8×8, 8×16, 16×4, 16×8, and 16×16. Alternatively, the first group can be defined as a set of 4×8, 4×16, 8×4, 8×8, 8×16, 16×4, 16×8, and 16×16.Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 8×4, 8×16, 16×4, 16×8, and 16×16. Alternatively, the first group can be defined as a set of 4×8, 4×16, 8×4, 8×16, 16×4, 16×8, and 16×16. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 4×32, 8×4, 8×16, 8×32, 16×4, 16×8, 32×4, and 32×8. Alternatively, the first group can be defined as a set of 4×8, 4×16, 4×32, 8×4, 8×16, 8×32, 16×4, 16×8, 32×4, and 32×8. The first group can be defined as a set of 4×4, 4×8, 4×16, 4×32, 8×4, 8×8, 8×16, 8×32, 16×4, 16×8, 16×16, 32×4, and 32×8. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 4×32, 8×4, 8×8, 8×16, 8×32, 16×4, 16×8, 16×32, 32×4, 32×8, and 32×16. Alternatively, the first group can be defined as a set of 4×4, 4×8, 4×16, 4×32, 8×4, 8×8, 8×16, 8×32, 16×4, 16×8, 16×16, 16×32, 32×4, 32×8, and 32×16. Alternatively, the first group can be defined as the set of 4×8, 4×16, 4×32, 8×4, 8×16, 8×32, 16×4, 16×8, 16×32, 32×4, 32×8, and 32×16. Alternatively, the first group can be defined as the set of 4×4, 4×8, 4×16, 4×32, 8×4, 8×8, 8×16, 8×32, 16×4, 16×8, 16×16, 16×32, 32×4, 32×8, 32×16, and 32×32.
[0238] An NSPT matrix (or NSPT kernel) with a predetermined dimension can be applied to the block size belonging to the first group. Here, the NSPT matrix can be represented as a P×Q dimensional matrix as the backward transformation matrix, where the P×Q matrix represents a matrix with P rows and Q columns, respectively.
[0239] As examples, a 16×16 NSPT matrix can be applied to a 4×4 block. A 32×20 NSPT matrix can be applied to at least one of a 4×8 block or an 8×4 block. A 64×24 NSPT matrix can be applied to at least one of a 4×16 block or a 16×4 block. A 64×32 NSPT matrix can be applied to an 8×8 block. A 128×40 NSPT matrix can be applied to at least one of an 8×16 block or a 16×8 block. A 256×44 NSPT matrix can be applied to a 16×16 block. A 128×36, 128×38, or 128×40 NSPT matrix can be applied to a 4×32 block or a 32×4 block. A 256×48 NSPT matrix can be applied to an 8×32 block or a 32×8 block. A 512×52 or 512×54 NSPT matrix can be applied to a 16×32 block or a 32×16 block.
[0240] Since the P×Q matrix is a backward NSPT matrix, the P×1 output vector can be obtained by applying the P×Q matrix to the Q×1 input vector (i.e., (P×Q matrix) × (Q×1 input vector)). Here, the Q×1 input vector can correspond to the (dequantized) transform coefficients for which NSPT is applied within the current block. In this case, the value of Q can refer to the number of transform coefficients for which NSPT is applied, and can be less than or equal to the product of the width and height of the current block. The value of Q can be variably determined based on the size of the current block within the first group of block sizes mentioned above. Alternatively, the value of Q can be set to be equal to the block sizes within the first group. The P×1 output vector can correspond to the residual signal (or decoded residual sample). The value of P can be equal to the product of the width and height of the current block.
[0241] Conversely, the forward NSPT matrix can be represented as a Q×P matrix, which is the transpose of the P×Q matrix. A Q×1 output vector can be obtained by applying the Q×P matrix to the P×1 input vector (i.e., (Q×P matrix)×(P×1 input vector)). Here, the P×1 input vector can correspond to the residual samples in the current block to which NSPT is applied. The value of P can be equal to the product of the width and height of the current block. The Q×1 output vector can correspond to the transform coefficients derived by NSPT in the current block. In this case, the value of Q can refer to the number of transform coefficients output by NSPT and can be less than or equal to the product of the width and height of the current block. Similarly, the value of Q can be variably determined based on the size of the current block belonging to the first group of block sizes described above. Alternatively, the value of Q can be set for blocks of equal size belonging to the first group.
[0242] As in the example, NSPT can be applied to M×N blocks and N×M blocks that are non-square blocks. For example, NSPT can be applied to 4×8 blocks and 8×4 blocks. Alternatively, NSPT can be applied to 4×16 blocks and 16×4 blocks, or NSPT can be applied to 8×16 blocks and 16×8 blocks, or NSPT can be applied to 16×32 blocks and 32×16 blocks.
[0243] By applying NSPT to a specific block size belonging to the first group, the transformation can be performed in a more complex manner, and coding performance can be improved. When applying forward LFNST, the transform coefficients of the principal transform in the remaining region except the region where LFNST is applied (i.e., the region of interest, ROI) can be zeroed. Furthermore, LFNST can consist of a small number of transform basis vectors. In this case, performance degradation may occur when separable principal transforms such as DCT-2 and non-separable quadratic transforms such as LFNST are applied to the corresponding block size instead of NSPT. When applying NSPT instead of LFNST for the corresponding case, the zeroing process is omitted, and coding performance is improved compared to the case of applying LFNST. Additionally, performance improvement can be expected with the application of NSPT. NSPT or LFNST can be applied using the following symmetry. Here, for LFNST, the symmetry is used only for the ROI region to perform the transpose operation on the corresponding input block. On the other hand, for NSPT, the symmetry is used to perform the transpose operation on the entire block. Therefore, for NSPT, a more complex symmetry can be used to train and apply the corresponding NSPT kernel, thus performance improvement can be expected.
[0244] Furthermore, when applying LFNST instead of NSPT to an 8×8 block, a 32×64 transform matrix can be applied instead of a 16×64 transform matrix from the perspective of forward transform. Here, a 16×64 transform matrix can be configured by sampling the top 16 rows of the 32×64 transform matrix. When LFNST based on a 16×64 transform matrix is applied to an 8×8 block, 16 multiplications are required per sample to apply LFNST, but when using a 32×64 transform matrix, 32 multiplications are required per sample to apply LFNST. However, when using a 32×64 transform matrix in this way, an improvement in coding performance can be expected.
[0245] When the current block's tree type is single-tree, NSPT can be applied to the luma component of the current block, but not to the chroma component. When the current block's tree type is dual-tree, NSPT can be applied to both the luma and chroma components of the current block.
[0246] Alternatively, regardless of whether the current block's tree type is single-tree, NSPT can be applied to the luma component of the current block, but not to the chroma component. Alternatively, regardless of whether the current block's tree type is single-tree, NSPT can be applied to both the luma and chroma components of the current block.
[0247] As an example, when the current block's tree type is single-tree, NSPT is allowed for both the luma and chroma components, and the current block size belongs to the first group, an NSPT index can be signaled, and the luma and chroma components of the current block can share the corresponding NSPT index. Here, the NSPT index can be an index used to select any transform kernel candidate for NSPT. When the luma and chroma block sizes of the current block belong to the first group, the transform kernel candidate selected by the same NSPT index can be applied to both the luma and chroma components. When the current block's tree type is single-tree and NSPT is applied only to the luma component, LFNST can be omitted from the application of LFNST to the chroma component of the current block, and a separate transform can be applied. Alternatively, when the current block's tree type is single-tree and NSPT is applied only to the luma component, LFNST can be applied to the chroma component of the current block.
[0248] In a single-tree configuration, there may be a high correlation between the luma and chroma components. In this case, unnecessary signaling can be reduced and compression efficiency can be improved by applying NSPT only to the luma component or by applying transform kernel candidates selected by a single NSPT index to both the luma and chroma components. On the other hand, for non-single-tree configurations, the luma and chroma components have independent partitioning and coding structures. In this case, by signaling the NSPT index of each component, the characteristics of each component can be reflected and compression efficiency can be improved.
[0249] The NSPT kernel for NSPT can be derived based on at least one of symmetry between intra-prediction modes or symmetry between block shapes. As an example, the NSPT kernel can be derived as an NSPT kernel corresponding to at least one of a mode symmetric to the intra-prediction mode of the current block or a block shape symmetric to the block shape of the current block. Alternatively, the NSPT kernel can be derived based on an NSPT set including one or more NSPT kernel candidates, wherein the NSPT set can be derived as an NSPT set corresponding to at least one of a mode symmetric to the intra-prediction mode of the current block or a block shape symmetric to the block shape of the current block. Any one of the one or more NSPT kernel candidates belonging to the NSPT set can be set as the NSPT kernel for the current block. For this purpose, an NSPT index specifying any one of the one or more NSPT kernel candidates belonging to the NSPT set can be used. The NSPT index can be signaled via the bitstream or can be derived based on the aforementioned symmetries.
[0250] Symmetry may exist between at least two intra-prediction modes predefined in the decoding device. For ease of description, the symmetry around the top-left diagonal mode (i.e., mode 34) will be described below. (See also...) Figure 5 Symmetry exists among the orientation patterns. All patterns, except for the planar pattern (number 0) and the DC pattern (number 1), have a predicted direction. Patterns 2 through 66 can be referred to as normal orientation patterns (represented as [2, 66]), patterns -14 through -1 (represented as [-14, -1]), and patterns 67 through 80 (represented as [67, 80]) can be referred to as wide orientation patterns. Wide orientation patterns can include at least one of a value less than -14 or a value greater than 80. (See reference...) Figure 5 All patterns except pattern 0 and pattern 1 are symmetric about pattern 34. Specifically, for pattern [2, 66], pattern x is symmetric to pattern (68-x), and between patterns [-14, -1] and [67, 80], pattern x is symmetric to pattern (66-x). The same symmetry relation can be established between patterns [N, -1] and [67, 66-N]. Here, N can be an integer less than or equal to -14.
[0251] Furthermore, regarding the symmetry between block shapes, M×N blocks and N×M blocks can be defined as blocks that are symmetrical to each other. Here, M and N can be the same or different from each other. Alternatively, M1×N1 blocks and M2×N2 blocks can be defined as blocks that are symmetrical to each other when the width-to-height ratio (M1 / N1) of M1×N1 block is the same as the height-to-width ratio (N2 / M2) of M2×N2 block. Alternatively, M1×N1 blocks and M2×N2 blocks can be defined as blocks that are symmetrical to each other when the width-to-height ratio (M1 / N1) of M1×N1 block is the same as the height-to-width ratio (M2 / N2) of M2×N2 block.
[0252] Within a square block, mutually symmetrical patterns can share at least one of the NSPT set, NSPT index, or NSPT kernel. In other words, at least one of the NSPT set, NSPT index, or NSPT kernel of any symmetrical pattern can be equally applied to another symmetrical pattern.
[0253] As an example, mutually symmetrical patterns can share a single NSPT kernel. However, for one symmetrical pattern, the corresponding NSPT kernel can be applied to the input data, while for another symmetrical pattern, the corresponding NSPT kernel can be applied after applying a transpose operation to the input data. Specifically, when pattern x belongs to pattern [2, 33], a 1D vector can be configured for the M×M block of input data in column-major order for pattern x, and the NSPT kernel can be applied to the corresponding 1D vector. Here, configuring the 1D vector in column-major order involves reading the input data column by column from the M×M block of input data to obtain M columns, and arranging them sequentially to configure the 1D vector. On the other hand, a 1D vector can be configured in row-major order for the pattern (68-x) that is symmetrical to pattern x, and the same corresponding NSPT kernel can be applied to the corresponding 1D vector. Here, configuring the 1D vector in row-major order involves reading the input data row by row from the M×M block of input data to obtain M rows, and arranging them sequentially to configure the 1D vector. When mode x belongs to mode [N, -1] (N≤-14), 1D vectors can be configured in row priority order for the mode (66-x) that is symmetrical to mode x, and the same NSPT kernel as mode x can be applied to the corresponding 1D vector. Column priority or row priority order can be applied to modes 0 and 1, and can also be applied to mode 34. Additionally, row priority order can be applied to intra-prediction modes belonging to mode [2, 33], and column priority order can be applied to modes symmetrical to the corresponding intra-prediction modes. Row priority order can be applied to intra-prediction modes belonging to mode [N, -1], and column priority order can be applied to modes symmetrical to them.
[0254] For non-square blocks, in addition to the symmetry between intra-prediction modes, the symmetry between block shapes can also be considered. A non-square block with width and height of M and N can be considered to have a symmetric relationship with other non-square blocks with width and height of N and M. As an example, in mode [2, 66], symmetry can exist between mode x of an M×N block and mode (68-x) of an N×M block. Similarly, when mode x of an M×N block belongs to mode [N, -1] (N≤-14), symmetry can exist between mode x of an M×N block and mode (66-x) of an N×M block.
[0255] The method for configuring a 1D vector from the input data block is as described above. In other words, when column priority is applied to pattern x, row priority can be applied to the pattern that is symmetrical to it. Alternatively, when row priority is applied to pattern x, column priority can be applied to the pattern that is symmetrical to it. Specifically, when column priority is applied to pattern x, M columns can be obtained by reading input data from the M×N block as input data in column units, and these columns can be arranged sequentially to configure a 1D vector. Here, each column can have a length N. For the pattern symmetrical to pattern x, N rows can be obtained by reading input data from the M×N block as input data in row units, and these rows can be arranged sequentially to configure a 1D vector. Here, each row can have a length M. Alternatively, when row priority is applied to pattern x, N rows can be obtained by reading input data from the M×N block as input data in row units, and these rows can be arranged sequentially to configure a 1D vector. Here, each row can have a length M. For a pattern symmetric to pattern x, M columns can be obtained by reading input data column by column from an M×N block of input data, and these columns can be arranged sequentially to configure a 1D vector. Here, each column can have a length N.
[0256] When the current block is an M×N block with mode x and the aforementioned symmetry is applied to the current block, the NSPT set and / or NSPT kernel of the current block can be determined based on at least one of an intra-prediction mode symmetric to mode x or an N×M block size symmetric to the M×N block size. Here, the NSPT kernel can be set to the NSPT kernel of an N×M block, rather than the NSPT kernel of an M×N block. In other words, when the symmetry is applied to the current block, the NSPT set and / or NSPT kernel of blocks symmetric to the current block can be used in the same manner. As described above, 1D vectors can be configured from the input data block according to a predetermined priority that corresponds to the input of the NSPT kernel.
[0257] Additionally, there may be a restriction that symmetry is only used when the value of the intra-prediction mode of the current block is greater than 34. In other words, when the value of the intra-prediction mode of the current block is greater than 34, a transpose operation can be applied when configuring 1D vectors from the input data block, and an NSPT set or NSPT kernel corresponding to the block shape and / or mode with symmetry to the current block can be used. Specifically, when the intra-prediction mode of the current block belongs to mode [N, -1] and mode [2, 34], symmetry may not be used for the current block. On the other hand, when the intra-prediction mode of the current block belongs to mode [35, 66] and mode [67, 66-N], symmetry may be used for the current block. Here, N can be an integer less than or equal to -14.
[0258] The symmetry-based derivation of the NSPT set or NSPT kernel can be performed adaptively based on the size of the current block. As an example, for 4×4 blocks and 8×8 blocks, the symmetry-based derivation of the NSPT set or NSPT kernel is possible, but for 4×8 blocks and 8×4 blocks, it may not be possible to derive the NSPT set or NSPT kernel based on the symmetry.
[0259] The number of available NSPT sets can vary depending on whether symmetry is used. For example, when symmetry is used, the number of available NSPT sets can be 35, while when symmetry is not used, the number of available NSPT sets can be 67.
[0260] Table 9 below provides examples of using symmetry to determine the NSPT set and shows the mapping between the NSPT set and the intra-frame prediction mode when the number of available NSPT sets is 35.
[0261] [Table 9]
[0262]
[0263] Referring to Table 9, when the value (X) of the intra-prediction mode of the current block is less than 0, the NSPT set of the current block can be determined as the NSPT set with NSPT set index 2 out of 35 NSPT sets. When the value (X) of the intra-prediction mode of the current block is greater than or equal to 0 and less than or equal to 34, the NSPT set of the current block can be determined as the NSPT set with NSPT set index X out of 35 NSPT sets. When the value (X) of the intra-prediction mode of the current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block can be determined as the NSPT set with NSPT set index (68-X) out of 35 NSPT sets. When the value (X) of the intra-prediction mode of the current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block can be the same as the NSPT set with value (68-X) corresponding to the mode symmetrical to the intra-prediction mode of the current block. Similarly, when the value (X) of the intra-prediction mode of the current block is greater than 66, the NSPT set of the current block can be determined as the NSPT set with NSPT set index 2 among the 35 NSPT sets. When the value (X) of the intra-prediction mode of the current block is greater than 66, the NSPT set of the current block can be the same as the NSPT set corresponding to the mode symmetrical to the intra-prediction mode of the current block.
[0264] Table 10 below provides an example of determining the NSPT set without using symmetry, and shows the mapping between the NSPT set and the intra-prediction mode when the number of available NSPT sets is 67.
[0265] [Table 10]
[0266]
[0267] Referring to Table 10, when the value (X) of the intra-prediction mode of the current block is less than 0, the NSPT set of the current block can be determined as the NSPT set with NSPT set index 2 out of 67 NSPT sets. When the value (X) of the intra-prediction mode of the current block is greater than or equal to 0 and less than or equal to 66, the NSPT set of the current block can be determined as the NSPT set with NSPT set index X out of 67 NSPT sets. Similarly, when the value (X) of the intra-prediction mode of the current block is greater than 66, the NSPT set of the current block can be determined as the NSPT set with NSPT set index 66 out of 67 NSPT sets.
[0268] Symmetry can be used to save memory size required to store transform kernels while maintaining performance depending on the application of the transform. For example, when using 35 NSPT sets instead of 67 NSPT sets, the memory size required to store NSPT kernels can be significantly reduced.
[0269] The number of available NSPT sets and / or the number of NSPT kernel candidates belonging to an NSPT set can vary depending on the block size. For example, the number of available NSPT sets for a 4×4 block can be 35, for 4×8 and 8×4 blocks it can be 19, and for an 8×8 block it can be 10. A 4×4 block NSPT set can consist of three NSPT kernel candidates, a 4×8 and 8×4 block NSPT set can consist of three or two NSPT kernel candidates, and an 8×8 block NSPT set can consist of one NSPT kernel candidate.
[0270] As the block size increases, the transform kernel size can also increase. Therefore, the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to those sets can be reduced, saving memory space required to store the transform kernels. Furthermore, as the block size increases, the characteristics of the residual signal within the corresponding block tend to become more generalized. Therefore, reducing the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to those sets can help maintain compression efficiency while reducing implementation complexity by reflecting these statistical characteristics.
[0271] The NSPT core can be configured with 8-bit precision. The coefficients within the NSPT core can range from -128 to 127. When the precision increases beyond 8 bits, the result obtained through matrix multiplication can be shifted right by the increased precision. For example, if the value obtained after matrix multiplication based on an 8-bit precision NSPT core is shifted right by S bits and stored in a buffer, then if the core coefficients are configured with N-bit precision, they can be shifted right by (S + (N-8)) bits and stored in a buffer.
[0272] When the NSPT core is configured with 8-bit precision, it prevents excessive increase in internal precision in the encoder / decoder that performs the transformation, thereby reducing implementation complexity in terms of memory requirements and operations, while minimizing the reduction in compression efficiency.
[0273] When a backward NSPT is applied to a current block of size N×N, the size of the NSPT kernel (or NSPT matrix) can be represented as MN×r. Here, MN can refer to the product of the width and height of the current block. This can refer to the output length of the NSPT or the number of residual samples generated by the NSPT. Additionally, r can refer to the input length of the NSPT or the number of (dequantized) transform coefficients to which the NSPT is applied. r can be an integer greater than or equal to 0 and less than or equal to MN. Below is an example of an NSPT matrix of MN×r based on the block size.
[0274] A 4×4 block NSPT matrix can be composed of a 16×16 matrix. 4×8 and 8×4 block NSPT matrices can be composed of 32×20, 32×16, 32×24, 32×28, or 32×32 matrices. An 8×8 block NSPT matrix can be composed of 64×16, 64×24, 64×32, 64×40, 64×48, 64×56, or 64×64 matrices. 4×16 and 16×4 block NSPT matrices can be composed of 64×16, 64×24, 64×32, 64×40, 64×48, 64×56, or 64×64 matrices. The NSPT matrix for 8×16 and 16×8 blocks can be composed of 128×96, 128×64, 128×48, or 128×32 matrices. The NSPT matrix for 16×16 blocks can be composed of 256×128, 256×96, or 256×64 matrices. The NSPT matrix for 16×32 and 32×16 blocks can be composed of 512×256 or 512×128 matrices. The NSPT matrix for 32×32 blocks can be composed of 1024×512, 1024×256, or 1024×128 matrices.
[0275] Alternatively, a 16×16 matrix can be applied to 4×N blocks and N×4 blocks. Here, N can be an integer greater than or equal to 4. A 64×16 matrix can be applied to 8×8 blocks. A 64×32 matrix can be applied to 8×N blocks and N×8 blocks. Here, N can be an integer greater than or equal to 16. A 96×32 matrix can be applied to 16×N blocks and N×16 blocks. Here, N can be an integer greater than or equal to 16.
[0276] Alternatively, the value of r in the MN×r NSPT matrix can be determined according to predetermined criteria. These criteria may be (1) ensuring that the sum of the computational cost of the primary transformation and the secondary transformation is less than or equal to a certain level, and (2) ensuring that the number of multiplications per sample required for the NSPT operation is less than or equal to a certain number.
[0277] Based on the inverse transform, when performing a separable principal transform via matrix multiplication on an M×N block, each sample requires (M+N) multiplications to perform the corresponding principal transform. Furthermore, when applying LFNST to a specific region of interest (ROI), assuming the LFNST matrix of the inverse transform is a P×Q matrix, each sample requires (P*Q) / (M*N) multiplications. Here, the P×Q matrix can refer to a matrix with P rows and Q columns.
[0278] When applying NSPT instead of DCT-2 transform (or separable transform, such as KLT) and LFNST to an M×N block, the value of r that ensures the number of multiplications per sample in the case of applying the corresponding NSPT is less than or equal to the number of multiplications per sample in the case of applying DCT-2 transform and LFNST can be determined as follows.
[0279] [Formula 6]
[0280]
[0281] When the value of r is set to its maximum value (i.e., r = M + N + (P*Q) / (M*N)) while satisfying Equation 6 above, the value of r in the NSPT matrices of each block size can be set as follows. In Equation 6, when the value of (P*Q) / (M*N) is not an integer, an integer close to the value of (P*Q) / (M*N) can be used. As an example, a floor operation can be applied to the value of (P*Q) / (M*N). In this case, r can be set as (M + N + floor((P*Q) / (M*N))). Here, floor(x) can refer to the largest integer not greater than x. Alternatively, a rounding operation can be applied to the value of (P*Q) / (M*N). In this case, r can be set as (M + N + round((P*Q) / (M*N))). Here, round(x) can refer to the value obtained by rounding x. Alternatively, the value of (P*Q) / (M*N) can be rounded up. In this case, r can be set as (M+N+ceil((P*Q) / (M*N))). Here, ceil(x) can refer to the smallest integer greater than or equal to x. When the rounding down operation is applied, the inequality in Equation 6 above is satisfied. However, when the rounding or rounding up operation is applied, the inequality in Equation 6 above may not be satisfied.
[0282] For a 4×4 block NSPT, the maximum value of r is 24. However, since the value of r must be less than or equal to 16, the value of r can be set to 16. For 4×8 and 8×4 block NSPTs, the maximum value of r is 20. The value of r can be set to 20. For an 8×8 block NSPT, the maximum value of r is 32. The value of r can be set to 32. For 4×16 and 16×4 block NSPTs, the maximum value of r is 24. The value of r can be set to 24. For 8×16 and 16×8 block NSPTs, the maximum value of r is 40. The value of r can be set to 40. For a 16×16 block NSPT, the maximum value of r is 44. The value of r can be set to 44. For 16×32 and 32×16 block NSPTs, the maximum value of r is 54. The value of r can be set to 54. For a 32×32 block NSPT, the maximum value of r is 67. The value of r can be set to 67. For NSPTs with 4×32 blocks and 32×4 blocks, the maximum value of r is 38. The value of r can be set to 38. Alternatively, the value of r can be set to 20. For NSPTs with 8×32 blocks and 32×8 blocks, the maximum value of r is 48. The value of r can be set to 48. Alternatively, the value of r can be set to 24.
[0283] In the NSPT matrices of the various block sizes mentioned above, there may be cases where the value of r is not a multiple of 4. For ease of implementation, it may be advantageous to set the value of r to a multiple of 4. For example, when implementing parallel processing using Single Instruction Multiple Data (SIMD) instructions, it may be advantageous to set the value of r to a multiple of 4 when processing the inner product of four transformation basis vectors simultaneously (i.e., generating four transformation coefficients simultaneously) during the application of forward NSPT.
[0284] As an example, for NSPTs with 16×32 blocks and 32×16 blocks, the value of r can be set to 52 or 56 instead of 54. For NSPTs with 32×32 blocks, the value of r can be set to 64 or 68 instead of 67. For NSPTs with 4×32 blocks and 32×4 blocks, the value of r can be set to 36 or 40 instead of 38.
[0285] More generally, the value of r can be set to a multiple of K. Here, K can be an integer greater than or equal to 1. As an example, the value of r can be set to a multiple of K that satisfies the inequality in Equation 7 below.
[0286] [Formula 7]
[0287]
[0288] In Equation 7 above, func() can be either rounded down, rounded to the nearest integer, or rounded up.
[0289] The r value according to the aforementioned predetermined standard does not consider the case of zeroing. In other words, when applying forward LFNST, the transform coefficients of the principal transform in the remaining region outside the region where LFNST is applied are set to zero, so the actual computational cost required to apply DCT-2 and LFNST can be less than the aforementioned computational cost. Therefore, when considering zeroing, the value of r can be set to a value less than the r value according to the predetermined standard.
[0290] Since zeroing is not performed for 4×4 blocks, the value of r can be set to a value less than or equal to 16.
[0291] For a 4×8 block, the remaining region except for the top-left 4×4 block can be zeroed out based on the forward transform, and a 16×16 matrix as the forward LFNST matrix can be applied to the top-left 4×4 block. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 8 ((4×4×8)+(4×8×4)) / (4×8)=8), and the number of per-sample multiplications required in the LFNST is 8 ((16×16) / 32=8). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 16, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 16.
[0292] For an 8×4 block, the remaining region except for the top-left 4×4 block can be zeroed out based on the forward transform, and the 16×16 matrix, which serves as the forward LFNST matrix, can be applied to the top-left 4×4 block. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 6 ((4×8×4)+(4×4×4) / (8×4)=6), and the number of per-sample multiplications required in the LFNST is 8 ((16×16) / 32=8). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 14, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 14.
[0293] For an 8×8 block, zeroing may not be performed for the separable master transformation, and in this case, the value of r can be set to a value less than or equal to 32.
[0294] For an 8×16 block, the remaining region except for the top-left 8×8 block can be zeroed out based on the forward transform, and a 64×32 matrix, which serves as the forward LFNST matrix, can be applied to the top-left 8×8 block. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 16 ((8×8×16)+(8×16×8) / (8×16)=16), and the number of per-sample multiplications required in the LFNST is 16 ((64×32) / 128=16). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 32, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 32.
[0295] For a 16×8 block, the remaining region except for the top-left 8×8 block can be zeroed out based on the forward transform, and the 64×32 matrix, which serves as the forward LFNST matrix, can be applied to the top-left 8×8 block. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 12 ((8×16×8)+(8×8×8) / (16×8)=12), and the number of per-sample multiplications required in the LFNST is 16 ((64×32) / 128=16). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 28, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 28.
[0296] For a 16×16 block, the remaining region except for the top-left 12×12 block can be zeroed out based on the forward transform, and the 96×32 matrix, which serves as the forward LFNST matrix, can be applied to the top-left 12×12 block. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 21 ((12×16×16)+(12×16×12) / (16×16)=21), and the number of per-sample multiplications required in the LFNST is 12 ((96×32) / 256=12). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 33, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 33.
[0297] For a 4×16 block, the remaining region except for the top-left 4×4 block can be zeroed based on the forward transform, and a 16×16 matrix can be applied to the top-left 4×4 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 8 ((4×4×16)+(4×16×4) / (4×16)=8), and the number of per-sample multiplications required in the LFNST is 4 ((16×16) / 64=4). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 12, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 12.
[0298] For a 16×4 block, the remaining region except for the top-left 4×4 block can be zeroed out based on the forward transform, and a 16×16 matrix can be applied to the top-left 4×4 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 5 (((4×16×4)+(4×4×4)) / (16×4)=5), and the number of per-sample multiplications required in the LFNST is 4 ((16×16) / 64=4). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 9, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 9.
[0299] For a 4×32 block, the remaining region except for the top-left 4×4 block can be zeroed out based on the forward transform, and a 16×16 matrix can be applied to the top-left 4×4 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 8 ((4×4×32)+(4×32×4) / (4×32)=8), and the number of per-sample multiplications required in the LFNST is 2 ((16×16) / 128=2). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 10, which is the sum of the number of per-sample multiplications in the separable master transform and the number of per-sample multiplications in the LFNST. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can be set to less than or equal to 10.
[0300] For a 32×4 block, the remaining region except for the top-left 4×4 block can be zeroed out based on the forward transformation, and a 16×16 matrix can be applied to the top-left 4×4 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transformation is 4.5 ((4×32×4)+(4×4×4) / (32×4)=4.5), and the number of per-sample multiplications required in the LFNST is 2 ((16×16) / 128=2). Therefore, when the separable master transformation and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 6.5. Since the same amount of computation is required even when applying the backward separable master transformation and LFNST, the value of r can be set to less than or equal to 6.5. Here, the value of r is the value of the configuration matrix dimension, so it can be set to an integer of 6 or 7, rather than 6.5.
[0301] For an 8×32 block, the remaining region except for the top-left 8×8 block can be zeroed out based on the forward transform, and a 64×32 matrix can be applied to the top-left 8×8 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 16 ((8×8×32)+(8×32×8) / (8×32)=16), and the number of per-sample multiplications required in the LFNST is 8 ((64×32) / 256=8). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 24. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can also be set to less than or equal to 24.
[0302] For a 32×8 block, the remaining region except for the top-left 8×8 block can be zeroed out based on the forward transform, and a 64×32 matrix can be applied to the top-left 8×8 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 10 ((8×32×8)+(8×8×8) / (32×8)=10), and the number of per-sample multiplications required in the LFNST is 8 ((64×32) / 256=8). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 18. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can also be set to less than or equal to 18.
[0303] For a 16×32 block, the remaining region except for the top-left 12×12 block can be zeroed out based on the forward transform, and a 96×32 matrix can be applied to the top-left 12×12 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transform is 21 ((12×16×32)+(12×32×12) / (16×32)=21), and the number of per-sample multiplications required in the LFNST is 6 ((96×32) / 512=6). Therefore, when the separable master transform and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 27. Since the same amount of computation is required even when applying the backward separable master transform and LFNST, the value of r can also be set to less than or equal to 27.
[0304] For a 32×16 block, the remaining region except for the top-left 12×12 block can be zeroed out based on the forward transformation, and a 96×32 matrix can be applied to the top-left 12×12 block as the forward LFNST matrix. When performing this zeroing, the number of per-sample multiplications required in the forward separable master transformation is 13.5 ((12×32×12)+(12×16×12) / (32×16)=13.5), and the number of per-sample multiplications required in the LFNST is 6 ((96×32) / 512=6). Therefore, when the separable master transformation and LFNST are replaced with NSPT, the value of r can be set to less than or equal to 19.5. Since the same amount of computation is required even when applying the backward separable master transformation and LFNST, the value of r can be set to less than or equal to 19.5. Here, the value of r is the value of the configuration matrix dimension, so it can be set to an integer of 19 or 20, rather than 19.5.
[0305] In the NSPT matrices for each block size described above, there may be cases where the value of r is not a multiple of K. Here, K can be 4. In this case, for ease of implementation, the value of r can be set to a multiple of K. When the value of r is preset using the above method... prev When r is a multiple of K, the value of r can be set as follows.
[0306] [Formula 8]
[0307] r=func(r prev / K)×K
[0308] In Equation 8 above, func() can be either rounded down, rounded to the nearest integer, or rounded up.
[0309] As mentioned above, the r value in the NSPT matrix of an M×N block can be different from the r value in the NSPT matrix of an N×M block. For example, the backward NSPT matrix of a 4×8 block can be a 32×16 matrix, and the backward NSPT matrix of an 8×4 block can be a 32×14 matrix. In this case, the symmetry between the M×N and N×M blocks can be used to determine the NSPT matrix.
[0310] Suppose the current block is an M×N block with pattern x. When the aforementioned symmetry is applied to the current block, instead of applying an NSPT matrix corresponding to an M×N block size or pattern x, an NSPT matrix corresponding to at least one of a pattern symmetric to pattern x or an N×M block size symmetric to an M×N block size can be applied. In this case, the NSPT matrix corresponding to an N×M block size can be applied to the current block as is. Alternatively, an NSPT matrix corresponding to an N×M block size can be applied, but for the value of r, the value of r in the NSPT matrix corresponding to an M×N block size can be used.
[0311] As an example, the backward NSPT matrix for a 4×8 block can be a 32×16 matrix (i.e., the r value in the NSPT matrix is 16), and the backward NSPT matrix for an 8×4 block can be a 32×14 matrix (i.e., the r value in the NSPT matrix is 14). When the current block is an 8×4 block with pattern x, the NSPT matrix of at least one of the pattern symmetric to pattern x or the 4×8 block symmetric to the 8×4 block can be applied to the current block. In this case, the 32×16 matrix can be used as is as the backward NSPT matrix for the 4×8 block, or a 32×14 matrix with the r value in the backward NSPT matrix of the 8×4 block can be used. Here, the 32×14 matrix can be derived by sampling the 14 rows from the left in the 32×16 matrix. Thus, when the 32×14 matrix is applied to the current block with a block size of 8×4, the above-mentioned predetermined criteria are satisfied.
[0312] Conversely, when the current block is a 4×8 block with pattern x, the NSPT matrix of at least one of the patterns symmetric to pattern x or 8×4 blocks symmetric to the 4×8 block can be applied to the current block. In this case, a 32×14 matrix of an 8×4 block can be applied to the current block instead of a 32×16 matrix of a 4×8 block. Thus, NSPT can be performed using fewer multiplications than is allowed for a 4×8 block.
[0313] For the NSPT of M×N blocks and N×M blocks, when the r values satisfying the above predetermined conditions are r1 and r2 respectively, the backward NSPT matrix of M×N blocks and N×M blocks can be set as MN×max(r1, r2). Here, max(r1, r2) can refer to selecting values greater than or equal to r1 and r2.
[0314] As an example, the backward NSPT matrix for a 4×8 block can be a 32×16 matrix (i.e., r=16 in the NSPT matrix), and the backward NSPT matrix for an 8×4 block can be a 32×14 matrix (i.e., r=14 in the NSPT matrix). When the current block is a 4×8 block with pattern x, the NSPT matrix of at least one of the patterns symmetric to pattern x or the 8×4 blocks symmetric to the 4×8 block can be applied to the current block. In this case, a 32×16 matrix can be used as the backward NSPT matrix for the 8×4 block. When the backward NSPT matrix is not configured as MN×max(r1, r2), the 32×14 matrix will be used as the NSPT matrix for the 8×4 block. However, when the NSPT matrices for the 4×8 block and the NSPT matrix for the 8×4 block are configured as 32×max(16, 14) matrices, the 32×16 matrix can be fully utilized.
[0315] Conversely, when the current block is an 8×4 block with pattern x, the NSPT matrix of at least one of the patterns symmetric to pattern x or symmetric to the 8×4 block can be applied to the current block. In this case, a 32×16 matrix can be used as the backward NSPT matrix of the 4×8 block, or a 32×14 matrix can be used. Here, the 32×14 matrix can be derived by sampling the 14 rows from the left in the 32×16 matrix.
[0316] When the NSPT matrix is configured as described above, a transformation consisting of the maximum number of transformation basis vectors can be applied while satisfying predetermined conditions, thereby maximizing coding performance.
[0317] In the above implementation, the value of r can be set to a multiple of 16. For example, for backward NSPTs of 4×8 and 8×4 blocks, a 32×16 matrix can be applied instead of a 32×20 matrix. The transform coefficients of the transform block can be encoded in predetermined coefficient groups (CGs). Here, a CG can be defined as a group of 16 transform coefficients; for example, a CG can be a sub-block of size such as 4×4, 2×8, or 8×2. A CG may not contain any non-zero transform coefficients, and in this case, the encoding of the transform coefficients for the corresponding CG can be skipped. Therefore, setting the value of r to a multiple of 16 has the advantage of reducing implementation complexity.
[0318] Transform coefficients can be derived by applying a forward NSPT to M×N blocks of residual samples. In this case, due to the zeroing, the number of derived transform coefficients can be less than or equal to the value (M*N). In other words, the forward NSPT matrix can be defined as an r×(M*N) matrix, where r can represent the output length of the NSPT or the number of transform coefficients derived by the NSPT, and (M*N) can represent the input length of the NSPT or the number of residual samples to which the NSPT is applied.
[0319] The derived transform coefficients can be arranged in an M×N block according to a predetermined scan order, and areas without transformed coefficients can be filled with 0 (i.e., set to zero). Therefore, during the scanning of transformed coefficients in the decoding device, when a non-zero transformed coefficient is found in a region that should be filled with 0 if NSPT is applied (or when the scan position of the last valid coefficient in the M×N block is greater than or equal to r), it is considered that NSPT is not applied to the corresponding M×N block, and the NSPT index can be notified without a signal. An index representing the scan position of 0 can be assigned to the top-left coefficient (i.e., DC component coefficient) in the M×N block, and indices incremented by 1 can be assigned to the remaining coefficients in the M×N block in a predetermined order.
[0320] One or more r values can be defined for the block sizes to which NSPT applies. As an example, one or more r values can be defined for each block size to which NSPT applies. Alternatively, one r value can be defined for each block size to which NSPT applies, and the r value for any block size to which NSPT applies may be different from the r value for another block size. Alternatively, one r value can be defined for some block sizes to which NSPT applies, and at least two r values can be defined for the remaining block sizes.
[0321] When multiple r values are available, a signal can be used to specify the index of any one of the r values or r itself. The corresponding index can be signaled in a high-level syntax (HLS) such as VPS, SPS, PPS, PH, or SH, or at the block level such as CTU, CU, or TU. When the value of r falls within a specific range, sufficient bits to encompass the corresponding range can be allocated and signaled. For example, when the value of r is in the range of 1 to 256, 8 bits can be specified as a fixed length and signaled.
[0322] The transformation kernel of the current block can be determined based on any one of embodiments 1 to 3 described above. Alternatively, the transformation kernel of the current block can be determined based on a combination of at least two of embodiments 1 to 3, to the extent that the inventions according to embodiments 1 to 3 described above do not conflict with each other.
[0323] The transform index used for the inverse transform of the current block can be signaled. Here, the transform index can specify any one of one or more transform kernels (or transform matrices) belonging to the transform set. Alternatively, the transform index can specify an NSPT index belonging to one or more NSPT kernels belonging to the NSPT set. Or, the transform index can specify an LFNST index belonging to one or more LFNST kernels belonging to the LFNST set.
[0324] Whether a transformation index corresponds to an NSPT index can be determined based on whether the size of the current block is one of the block sizes belonging to the first group mentioned above. Assume that the block sizes applicable to NSPT and LFNST are different from each other. In this case, when the size of the current block belongs to the first group, the transformation index signaled for the current block can correspond to the NSPT index, and the NSPT kernel can be determined from the NSPT set based on the corresponding transformation index. On the other hand, when the size of the current block does not belong to the first group, the transformation index signaled for the current block can correspond to the LFNST index, and the LFNST kernel can be determined from the LFNST set based on the corresponding transformation index. When the size of the current block does not belong to the first group, it may mean that the size of the current block belongs to the second group mentioned above. Alternatively, when the size of the current block does not belong to the first group, it may mean that the size of the current block corresponds to a block size applicable to LFNST among the block sizes belonging to the second group. Thus, the NSPT index and the LFNST index can be configured as an integrated syntax, rather than separate syntaxes.
[0325] As an example, suppose the block sizes for which NSPT applies to the first group are 4×4, 4×8, 8×4, and 8×8. For block sizes belonging to the first group, NSPT can be applied instead of LFNST. Specifically, NSPT can be applied instead of a combination of separable master transforms (e.g., DCT-2, separable KLT) and LFNST. The NSPT index can be signaled for the four block sizes belonging to the first group, and the LFNST index can be signaled for the remaining block sizes (where LFNST is allowed).
[0326] In this way, the amount of encoded information can be reduced when the NSPT and LFNST indices are signaled as a syntax. Furthermore, implementation complexity can be reduced by applying at least one of the initial values of binarization, CABAC context, or entropy encoding equally to the NSPT / LFNST indices.
[0327] Alternatively, NSPT and LFNST indexes can be signaled separately as individual syntaxes. In this case, implementation complexity may increase somewhat, but compression performance can be improved by performing optimized entropy coding for each index.
[0328] When the number of LFNST kernel candidates belonging to the LFNST set is the same as the number of NSPT kernel candidates belonging to the NSPT set, the same binarization can be applied to the LFNST index and the NSPT index. The same CABAC context (or, CABAC context increment) can be assigned to the bins of the LFNST index and the NSPT index.
[0329] Different binarization and / or CABAC contexts can be used for LFNST and NSPT indices. Different CABAC initialization values can be assigned to LFNST and NSPT indices. As an example, either the LFNST or NSPT index can be binarized based on fixed-length binarization, and the other can be binarized based on truncated unary binarization. Different CABAC contexts and / or CABAC initialization values can be assigned even when the binarization of the LFNST and NSPT indices is the same. Different binarization and / or CABAC contexts can be used for LFNST and NSPT indices when the number of LFNST kernel candidates belonging to the LFNST set and the number of NSPT kernel candidates belonging to the NSPT set are different from each other.
[0330] The number of NSPT kernel candidates belonging to the NSPT set can be set differently for each block size. Alternatively, the block sizes belonging to the first group can be divided into multiple subgroups. In this case, the number of NSPT kernel candidates belonging to the NSPT set can be set differently for each of the multiple subgroups. Here, at least one of the multiple subgroups can include multiple different block sizes.
[0331] The binarization applied to the NSPT index can vary depending on the number of NSPT kernel candidates belonging to the NSPT set.
[0332] As an example, when the number of NSPT kernel candidates in an NSPT set of a specific block size is 3, the NSPT index can have any value from 0 to 3. When the NSPT index value is 0, it indicates that NSPT is not applied to the current block. When the NSPT index value is not 0, it indicates the NSPT kernel candidate corresponding to the given NSPT index among the three NSPT kernel candidates. A bin can be assigned to distinguish between the case where NSPT is applied and the case where NSPT is not applied. A bin value of 0 corresponds to a NSPT index value of 0. On the other hand, a bin value of 1 corresponds to a NSPT index value of 1, 2, or 3. In this case, truncated univariate binarization can be applied to distinguish the three NSPT kernel candidates. In other words, two bins can be assigned to distinguish the three NSPT kernel candidates as 0, 10, and 11.
[0333] When the number of NSPT kernel candidates in an NSPT set of a specific block size is 2, the NSPT index can have any value from 0 to 2. A value of 0 indicates that NSPT is not applied to the current block. A value other than 0 indicates that the NSPT kernel candidate is the one corresponding to the given NSPT index. A bin can be assigned to distinguish between applying and not applying NSPT. Two NSPT kernel candidates can be distinguished by assigning a bin representing either one of them.
[0334] When the number of NSPT kernel candidates in the NSPT set for a specific block size is 1, the NSPT index can have either a value of 0 or 1. A value of 0 indicates that NSPT is not applied to the current block. A value of 1 indicates an NSPT kernel candidate. In this case, only one bin can be used to specify whether to apply NSPT and the NSPT kernel candidate.
[0335] The inverse transform of the current block can be a separable principal transform and / or an inverse transform based on LFNST. In other words, backward LFNST can be applied to all or part of the (dequantized) transform coefficients of the current block, and then backward separable principal transform can be applied to the transform coefficients derived by LFNST to derive the residual samples. As an example, backward LFNST can be applied to the (dequantized) transform coefficients belonging to a portion of the current block. Here, the portion refers to the region to which forward LFNST is applied, and is referred to below as the region of interest (ROI). The transform coefficients derived by LFNST can be arranged in the ROI region according to a predetermined scan order. The predetermined scan order can be row-major or column-major. Backward separable principal transform can be applied to the transform coefficients derived by LFNST and the transform coefficients belonging to the remaining regions within the current block excluding the ROI region. Alternatively, during the forward transform process, when zeroing is performed on the remaining regions within the current block excluding the ROI region (i.e., when the transform coefficients in the remaining regions are set to 0), the backward separable master transform can be applied to the transform coefficients derived by LFNST.
[0336] The following section describes a method for signaling the transform index of the inverse transform of the current block. Here, the inverse transform can refer to a backward inseparable transform. The inseparable transform can refer to the aforementioned LFNST or NSPT, and the transform index can refer to the LFNST index or NSPT index.
[0337] When the current block is encoded as a single tree, an inseparable transform can be applied to the luma component of the current block, but not to the chroma component. In this case, the transform index of the current block can be signaled based on at least one of a first condition for the luma component or a second condition for the chroma component. For example, the transform index of the current block can be signaled when both the first condition for the luma component and the second condition for the chroma component are satisfied; otherwise (i.e., when either the first condition for the luma component or the second condition for the chroma component is satisfied), no signal is required. Alternatively, the transform index of the current block can be signaled when the first condition for the luma component is satisfied, without checking whether the second condition for the chroma component is satisfied; otherwise, no signal is required. When the transform index is not signaled, it can be set so that an inseparable transform is not applied to the current block; for example, the transform index of the current block can be derived as 0.
[0338] According to the first condition of the luminance component of this disclosure, it can mean that there are no non-zero transform coefficients in a predetermined region within the luminance component block of the current block. Here, the predetermined region within the luminance component block can be defined as the remaining region in the luminance component block excluding the first region, which can be referred to as the second region to distinguish it from the first region. The first region can be defined as a region consisting of samples (or sample positions) with the same number as the number of transform coefficients input to the backward inseparable transform (or the number of transform coefficients output through the forward inseparable transform). The first region may include the upper left sample position of the luminance component block. The first region may include at least one non-zero transform coefficient. The first region may be the region to which the last valid coefficient in the luminance component block belongs. The width and height of the first region may be less than or equal to the width and height of the luminance component block, respectively.
[0339] As described above, the number of transform coefficients for applying the backward non-separable transform can be determined based on the size of the current block. The size of the current block can be the same as that of the luma component block. The size of the current block can be defined as a combination of width (W) and height (H), such as W×H. However, it is not limited to this; the size of the current block can be defined as any one of width or height, minimum / maximum of width and height, or the product of width and height.
[0340] As an example, in an encoding device, when an inseparable transform is applied to an M×N block, r transform coefficients less than or equal to (M*N) can be output. This may mean that the forward inseparable transform has r output lengths. The r output transform coefficients can be arranged sequentially within the M×N block from the top-left sample position (i.e., the DC position) according to a predetermined scan order. The region within the current block consisting of r transform coefficients can correspond to the first region described above. Then, zeroing can be applied to the sample positions after the r-th position within the M×N block (i.e., the second region within the current block). Zeroing can assign 0 to sample positions within the second region. Thus, when an inseparable transform is applied in the encoding device, there are no non-zero transform coefficients for the sample positions where zeroing is applied.
[0341] In the process of deriving transform coefficients based on residual information in the decoding device, if a non-zero transform coefficient is found at a sample position that should have been set to zero (i.e., a sample position after the r-th or second region) when an inseparable transform has been applied in the encoding device, it means that the inseparable transform has not been applied to the M×N block. In this case, the transform index of the inseparable transform can be notified without signaling for the M×N block. In this case, it can be set not to apply the inseparable transform to the M×N block; as an example, the corresponding transform index can be derived as 0.
[0342] According to the second condition of the chroma component of this disclosure, there are no non-zero transform coefficients in a predetermined region within the chroma component block of the current block. Here, the predetermined region within the chroma component block can be defined as the remaining region in the chroma component block excluding the first region, which can be referred to as the second region to distinguish it from the first region. The first region can be defined as a region consisting of samples (or sample positions) of the same number as the input length of the backward inseparable transform corresponding to the size of the chroma component block. The corresponding input length can refer to the number of transform coefficients input to the backward inseparable transform. The first region may include the upper left sample position of the chroma component block. The first region may include at least one non-zero transform coefficient. The first region may be the region to which the last valid coefficient in the chroma component block belongs. The width and height of the first region may be less than or equal to the width and height of the chroma component block, respectively.
[0343] Although the inseparable transform is not applied to the chroma component blocks, the input length of the backward inseparable transform can be determined based on the size of the chroma component blocks. Here, the input length of the inseparable transform can be determined as the input length of the inseparable transform based on the size of the chroma component blocks. Alternatively, the matrix size (or input / output length) of the inseparable transform can be determined based on the size of the luma component blocks corresponding to the chroma component blocks, and half of the input length of the inseparable transform can be determined as the input length of the inseparable transform. The method for determining the inseparable transform matrix based on the block size is as described above, and a detailed description is omitted here.
[0344] Specifically, instead of applying an inseparable transform to the chroma components, the transform coefficients belonging to predetermined regions within the chroma component block can be set to zero in the same manner as when applying an inseparable transform to the chroma components. In other words, the encoding device can retain only the transform coefficients belonging to some regions within the chroma component block and set the transform coefficients belonging to the remaining regions to 0. Here, "some regions" can refer to the regions where the transform coefficients output by the corresponding inseparable transform are arranged if an inseparable transform is applied to the chroma component block.
[0345] As an example, suppose the color is encoded in a 4:2:0 color format, the current block's tree type is a single tree, and the luma component block and chroma component block sizes are 8×16 and 4×8 respectively. Both 8×16 and 4×8 are block sizes that allow for non-separable transformations. In this case, a non-separable transformation can be applied to the luma component block, but not to the chroma component block. However, the chroma component block can also be zeroed out in the same way as when applying a non-separable transformation.
[0346] When performing forward NSPT on a 4×8 block based on a 20×32 NSPT matrix, only 20 transform coefficients can be output via NSPT. These 20 transform coefficients are arranged sequentially in the 4×8 block from the top-left sample position according to a predetermined scan order, and the remaining 12 sample positions are zeroed out. When zeroing is applied to the chrominance component block in the same way, for the 20th sample position in the chrominance component block according to the predetermined scan order from the top-left sample position, the pre-output transform coefficients can be retained as is, and zeros can be assigned from the 21st to the 32nd sample positions. Furthermore, the chrominance component block can consist of a Cb component block and a Cr component block, and both component blocks can be zeroed out in the same way.
[0347] The size of the first region within the chroma component block can be determined based on the size of the chroma component block. The first region can be defined as a region consisting of r sample positions. Alternatively, the first region can be defined as a region consisting of min(threshold, r) sample positions. min(threshold, r) is a function of the minimum value between the output threshold and r. The threshold is a predefined value in the encoding and decoding devices and can be an integer of 4, 8, 16, 32, 64, or larger. The threshold can vary based on the size of the chroma component block or the size of the corresponding luma component block (i.e., the size of the current block).
[0348] r can be determined based on the input length of the backward inseparable transform corresponding to the size of the chroma component block. As an example, r can be the same as the input length of the backward inseparable transform corresponding to the size of the chroma component block. Here, the input length can refer to the number of transform coefficients applied by the backward inseparable transform. This could mean that the backward inseparable transform corresponding to the size of the chroma component block is a P×r matrix. Alternatively, r can be determined based on the output length of the forward inseparable transform corresponding to the size of the chroma component block. As an example, r can be the same as the output length of the forward inseparable transform corresponding to the size of the chroma component block. Here, the output length can refer to the number of transform coefficients output by the forward inseparable transform. This could mean that the forward inseparable transform corresponding to the size of the chroma component block is an r×P matrix. In the inseparable transform matrix, the value of P can be the product of the width and height of the chroma component block. Alternatively, the value of P can be the number of transform coefficients (or residual samples) derived through the backward inseparable transform. The value of P can also be the number of samples belonging to the region within the chroma component block to which the forward inseparable transform is applied. As an example, when the inseparable transform is NSPT or LFNST, the value of P can be the same as the product of the width and height of the chrominance component block. When the inseparable transform is LFNST, the value of P can be the same as the number of transform coefficients belonging to the region of interest (ROI) to which the forward inseparable transform is applied.
[0349] As described above, when the inseparable transformation is an NSPT, an NSPT matrix with a predetermined dimension can be determined / mapped based on the block size of the applicable NSPT. For example, the forward NSPT matrices corresponding to 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, and 16×8 blocks can be 16×16, 20×32, 20×32, 32×64, 24×64, 24×64, 40×128, and 40×128 matrices, respectively. In other words, when the chroma component block is 4×4, the value of r can be 16. When the chroma component block is 4×8 or 8×4, the value of r can be 20. When the chroma component block is 8×8, the value of r can be 32. When the chroma component block is 4×16 or 16×4, the value of r can be 24. When the chroma component blocks are 8×16 blocks or 16×8 blocks, the value of r can be 40.
[0350] As described above, when the inseparable transformation is LFNST, an LFNST matrix with a predetermined dimension can be determined / mapped based on the block size. As an example, the forward LFNST matrix corresponding to 4×N blocks and / or N×4 blocks can be a 16×16 matrix. In other words, the value of r can be 16 when the chroma component blocks are 4×N or N×4 blocks. Here, N can be an integer greater than or equal to 4. Alternatively, the forward LFNST matrix corresponding to 8×8 blocks can be a 16×64 matrix. In other words, the value of r can be 16 when the chroma component blocks are 8×8 blocks. Alternatively, the forward LFNST matrix corresponding to 8×N blocks and / or N×8 blocks can be a 32×64 matrix. In other words, the value of r can be 32 when the chroma component blocks are 8×N or N×8 blocks. Here, N can be an integer greater than or equal to 16. A 16×64 matrix corresponding to an 8×8 block can be obtained by sampling the top 16 rows of a 32×64 matrix that serves as the forward LFNST matrix corresponding to an 8×N block or an N×8 block. Alternatively, the forward LFNST matrix corresponding to a 16×N block and / or an N×16 block can be a 32×96 matrix. In other words, when the chroma component blocks are 16×N blocks or N×16 blocks, the value of r can be 32. Here, N can be an integer greater than or equal to 16.
[0351] As mentioned above, the predefined allowed transform block sizes can be divided into a first group of block sizes for which NSPT can be applied and a second group of block sizes for which NSPT cannot be applied.
[0352] As an example, the first group can be defined as including at least one of 4×4, 4×8, 8×4, 8×8, 4×16, 16×4, 8×16, or 16×8, and the second group can be defined as including the remaining block size. In this case, NSPT can be applied to the block size belonging to the first group, and LFNST can be applied to all or part of the block size belonging to the second group. Alternatively, the first group can be defined as including at least one of 4×4, 4×8, 8×4, or 8×8, and the second group can be defined as including the remaining block size. In this case, NSPT can be applied to the block size belonging to the first group, and LFNST can be applied to all or part of the block size belonging to the second group. Alternatively, the first group can be defined as including at least one of 4×4, 4×8, 8×4, 8×8, 4×16, or 16×4, and the second group can be defined as including the remaining block size. In this case, NSPT can be applied to the block size belonging to the first group, and LFNST can be applied to all or part of the block size belonging to the second group.
[0353] Alternatively, LFNST can be applied to 4×4 blocks. This could mean that a 4×4 block corresponds to a block size for which both NSPT and LFNST can be applied. Alternatively, it could mean that a 4×4 block corresponds to a block size for which LFNST can be applied, but not for which NSPT can be applied. In other words, it could mean excluding 4×4 block sizes from the first group.
[0354] Alternatively, LFNST can be applied to 4×4 blocks and 8×8 blocks. This could mean that 4×4 blocks and 8×8 blocks correspond to block sizes for which both NSPT and LFNST can be applied. Alternatively, it could mean that 4×4 blocks and 8×8 blocks correspond to block sizes for which LFNST can be applied, but not block sizes for which NSPT can be applied. In other words, it could mean excluding 4×4 and 8×8 block sizes from the first group.
[0355] Alternatively, LFNST can be applied to 8×8 blocks. This could mean that an 8×8 block corresponds to a block size for which both NSPT and LFNST can be applied. Alternatively, it could mean that an 8×8 block corresponds to a block size for which LFNST can be applied, but not a block size for which NSPT can be applied. In other words, it could mean excluding 8×8 block sizes from the first group.
[0356] The value of r for the chroma component block can be set according to the method described above. The value of r can be set according to the method described above regardless of whether an inseparable transformation (e.g., NSPT and / or LFNST) is applied to the chroma component block. Alternatively, the value of r can be set according to the method described above when no inseparable transformation (e.g., NSPT and / or LFNST) is applied to the chroma component block. Alternatively, the value of r can be set according to the method described above when an inseparable transformation (e.g., NSPT and / or LFNST) is applied to the chroma component block.
[0357] Even when no inseparable transformation is applied to the chroma component blocks, the value of r can be set according to the above method when the size of the corresponding chroma component block corresponds to the block size for which NSPT or LFNST can be applied. In this case, when the size of the corresponding chroma component block corresponds to the block size for which NSPT can be applied, the value of r can be set based on the input length (or output length of the forward NSPT matrix) of the corresponding block size. When the size of the corresponding chroma component block corresponds to the block size for which LFNST can be applied, the value of r can be set based on the input length (or output length of the forward LFNST matrix) of the corresponding block size. In the following text, the method for zeroing the chroma component blocks is described when the tree type of the current block is a single tree.
[0358] When the chroma component block is 4×N or N×4 (N is an integer of 4 or greater), only a predetermined number of transform coefficients from the transform coefficients of the chroma component block can be retained, and the remaining transform coefficients can be set to 0. The transform coefficients of the chroma component block can be derived using a forward separable master transform. Alternatively, the transform coefficients of the chroma component block can be derived using at least one of forward NSPT or LFNST. The predetermined number can be r or min(16, r). The retained transform coefficients can be arranged within the chroma component block according to a predetermined scan order. 0 can be assigned to the remaining sample positions where the retained transform coefficients are not filled. In the decoding process, an inverse transform can be applied to r or min(16, r) transform coefficients within the chroma component block.
[0359] When the chroma component block is an M×N block (M and N are integers of 16 or greater), a predetermined number of transform coefficients from the chroma component block can be retained, and the remaining transform coefficients can be set to 0. The transform coefficients of the chroma component block can be derived using a forward separable master transform. Alternatively, the transform coefficients of the chroma component block can be derived using at least one of forward NSPT or LFNST. The predetermined number can be r or min(256, r). The retained transform coefficients can be arranged within the chroma component block according to a predetermined scan order. 0 can be assigned to the remaining sample positions where the retained transform coefficients are not filled. In the decoding process, an inverse transform can be applied to r or min(256, r) transform coefficients within the chroma component block.
[0360] When the chroma component block is 8×N or N×8 (N is an integer of 8 or greater), only a predetermined number of transform coefficients from the chroma component block's transform coefficients can be retained, and the remaining transform coefficients can be set to 0. The transform coefficients of the chroma component block can be derived using a forward separable master transform. Alternatively, the transform coefficients of the chroma component block can be derived using at least one of forward NSPT or LFNST. The predetermined number can be r or min(64, r). The retained transform coefficients can be arranged within the chroma component block according to a predetermined scan order. 0 can be assigned to the remaining sample positions where the retained transform coefficients are not filled. In the decoding process, an inverse transform can be applied to r or min(64, r) transform coefficients within the chroma component block.
[0361] The chromaticity component blocks can be zeroed out based on NSPT or LFNST.
[0362] Specifically, when the size of the chroma component block (or the size of the corresponding luminance component block) corresponds to the block size for which NSPT can be applied, the chroma component block can be zeroed according to NSPT. Specifically, when the size of the chroma component block (or the size of the corresponding luminance component block) corresponds to the block size for which LFNST can be applied, the chroma component block can be zeroed according to LFNST.
[0363] Depending on the tree type of the current block, you can selectively use either the size of the luma component block or the size of the chroma component block to determine whether to apply the zeroing method based on NSPT or the zeroing method based on LFNST to the chroma component block. For example, when the current block's tree type is single-tree, the determination can be based on the size of the luma component block; when the current block's tree type is two-tree, the determination can be based on the size of the chroma component block.
[0364] When it is determined that both NSPT and LFNST can be applied to a chroma component block, zeroing can be applied based on either the output length of NSPT or the output length of LFNST. When both NSPT and LFNST can be applied to a chroma component block, zeroing can be applied based on the output length of the inseparable transform with a preset priority. For example, NSPT can have a higher priority than LFNST. LFNST can have a higher priority than NSPT. Alternatively, when both NSPT and LFNST can be applied to a chroma component block, zeroing can also be applied based on the minimum or maximum value of the output lengths of NSPT and LFNST. For example, when both NSPT and LFNST are applied to a chroma component block, and the output length of the forward NSPT corresponding to the chroma component block is 20, and the output length of the forward LFNST corresponding to the chroma component block is 16, only the transform coefficients at the 16th sample position from the top left sample position within the chroma component block can be retained, and the transform coefficients at the remaining sample positions can be set to 0.
[0365] As described above, when the current block is encoded in a single tree, only the luminance component is subjected to an inseparable transform, while the chrominance component is not subjected to an inseparable transform and a zeroing is applied, the first condition of the luminance component and the second condition of the chrominance component can be checked, and when the first and second conditions are met, the transform index of the current block (in particular, the luminance component block) can be notified by a signal.
[0366] Alternatively, when the tree type of the current block is a single tree and an inseparable transform is applied only to the luminance component, the aforementioned zeroing of the chrominance component may not be applied. In this case, the first condition for the luminance component can be checked without checking the second condition for the chrominance component. When the first condition for the luminance component is met, the transform index of the current block (specifically, the luminance component block) can be signaled. On the other hand, when the first condition for the luminance component is not met (i.e., when there are non-zero transform coefficients in the second region within the luminance component block), the transform index of the current block may not be signaled. In this case, the corresponding transform index can be derived as 0.
[0367] Alternatively, when the current block's tree type is a single tree and an inseparable transform is applied to the luminance and chrominance components, both the first condition for the luminance component and the second condition for the chrominance component can be checked. In this case, if both the first and second conditions are met, the transform index of the current block can be signaled. If either the first or second condition is not met, the transform index of the current block can be signaled without notification.
[0368] When the color format is 4:2:0, the current block's tree type is single-tree, and the luma component block size is M×N, the chroma component block size can be (M / 2)×(N / 2). Assume that the intra-fraction sub-partition (ISP) mode is not applied to the current block. When applying an inseparable transform to the M×N luma component, an inseparable transform such as NSPT or LFNST can be applied to the chroma component, or neither NSPT nor LFNST can be applied. Specifically, it can be partitioned as follows.
[0369] 1) When applying LFNST to an M×N block of luminance components
[0370] 1-a) Apply LFNST to the (M / 2)×(N / 2) transform block of the chromaticity components.
[0371] 1-b) Apply NSPT to the (M / 2)×(N / 2) transform block of the chromaticity components
[0372] 1-c) Do not apply both NSPT and LFNST to the (M / 2)×(N / 2) transform block of the chrominance components (in this case, apply the DCT-2 based horizontal / vertical master transform to the chrominance components).
[0373] 2) When applying NSPT to the M×N transform block of the luminance component
[0374] 2-a) Apply LFNST to the (M / 2)×(N / 2) transform block of the chromaticity components.
[0375] 2-b) Apply NSPT to the (M / 2)×(N / 2) transform block of the chromaticity components.
[0376] 2-c) Do not apply both NSPT and LFNST to the (M / 2)×(N / 2) transform block of the chrominance components (in this case, apply the DCT-2 based horizontal / vertical master transform to the chrominance components).
[0377] Assuming that both M and N are greater than or equal to 4, NSPT and LFNST can be applied, and NSPT can be applied to 4×4 blocks, 4×8 blocks, 8×4 blocks, and 8×8 blocks. In this case, when the tree type of the current block is a single tree, the following situation may occur.
[0378] When the luma component block is 16×8 blocks, the chroma component block can be 8×4 blocks. It can be configured to ensure that LFNST is applied to the luma component and NSPT is applied to the chroma component. Alternatively, it can be configured to ensure that LFNST is applied to the luma component and not apply either NSPT or LFNST to the chroma component. In this case, even if neither NSPT nor LFNST is applied to the chroma component, zeroing can be applied in the same manner as when applying NSPT or LFNST as described above. Furthermore, since the chroma component block is 8×4 blocks in this example, zeroing based on NSPT can be applied.
[0379] When the luma component block is an 8×8 block, the chroma component block can be a 4×4 block. It can be configured to ensure that NSPT is applied to both the luma and chroma components. Alternatively, it can be configured to ensure that NSPT is applied to the luma component but not to apply either NSPT or LFNST to the chroma component. In this case, even if neither NSPT nor LFNST is applied to the chroma component, zeroing can be applied in the same manner as when applying NSPT or LFNST as described above. Furthermore, since the chroma component block is a 4×4 block in this example, zeroing based on NSPT can be applied.
[0380] When the luma component block is 8×4, the chroma component block can be 4×2. It can be configured to ensure that NSPT is applied to the luma component but neither NSPT nor LFNST is applied to the chroma component. In this case, zeroing based on NSPT or LFNST may not be applied to the chroma component. Conversely, the transformation coefficients from the top-left sample position to the nth sample position in the chroma component block according to a predetermined scanning order can be left as is, and zeroing can be applied to sample positions after the nth position. Here, n is a predefined value equally for both the encoding and decoding devices, and can be an integer greater than or equal to 2. As an example, n can be 4. The fourth sample position from the top-left sample position can belong to a 2×2 region that is the area including the top-left sample of the chroma component block.
[0381] When the luma component block is 32×32, the chroma component block can be 16×16. It can be configured to ensure that LFNST is applied to both the luma and chroma components. Alternatively, it can be configured to ensure that NFNST is applied to the luma component and not to apply either NSPT or LFNST to the chroma component. In this case, even if neither NSPT nor LFNST is applied to the chroma component, zeroing can be applied in the same manner as when applying NSPT or LFNST as described above. Furthermore, since the chroma component block is 16×16 in this example, zeroing based on LFNST can be applied.
[0382] When the current block's tree type is single-tree and Intra-Segmentation Partitioning (ISP) mode is applied to the current block, or when the current block's tree type is dual-tree, the width and height of the chroma component block will not be half the width and height of the luma component block, respectively. In this case, for the luma component, whether to apply NSPT or LFNST and / or the matrix size (or input / output length) of the corresponding non-separable transform can be determined based on the size of the corresponding transform block. For the chroma component, whether to apply NSPT or LFNST and / or the matrix size (or input / output length) of the corresponding non-separable transform can be determined based on the size of the corresponding transform block. Additionally, when the current block's tree type is single-tree, as described above, it can be configured to ensure that NSPT or LFNST is not applied to the chroma component, and of course, zeroing can be applied in the same manner as when applying NSPT or LFNST.
[0383] When an inseparable transform is applied to the luminance or chrominance component, the transform index of the inseparable transform can be signaled when the following conditions are met. The conditions for signaling the transform index can be considered as additional conditions to the first condition for the luminance component and the second condition for the chrominance component. In other words, the transform index can be signaled only when both the first and second conditions are met, and at least one of the following conditions is also met. Alternatively, the following conditions can be considered as independent conditions unrelated to the first and second conditions. In other words, the transform index can be signaled when at least one of the following conditions is met, regardless of whether the first and second conditions are met.
[0384] 1) When there are non-zero transform coefficients at positions other than the top left sample position of at least one transform block of the color components (i.e., Y, Cb, Cr), the transform index can be notified by a signal; otherwise (i.e., when there are no non-zero transform coefficients at positions other than the top left sample position of the transform blocks of all color components), the transform index can be notified by a signal.
[0385] As an example, when the tree type of the current block is a dual-tree and the current block is a transform block of the chrominance component, the corresponding transform index can be signaled only if there are non-zero transform coefficients at a position other than the top-left sample position of at least one of the transform blocks of the Cb and Cr components.
[0386] Alternatively, when the tree type of the current block is a single tree, the corresponding transform index can be signaled only if there are non-zero transform coefficients at a position other than the top-left sample position of at least one of the transform blocks of the luminance component and the chrominance component (the chrominance component can be composed of Cb components and Cr components).
[0387] Alternatively, when the tree type of the current block is a single tree and an inseparable transform is applied only to the luminance component, the transform index can be signaled when there are non-zero transform coefficients at positions other than the top-left sample position of at least one transform block of the luminance and chrominance components, and the first condition for the luminance component and the second condition for the chrominance component are satisfied. Even when there are no non-zero transform coefficients in the transform block of the luminance component, the transform index can still be signaled when the first and second conditions are satisfied.
[0388] 2) When at least one transform block of the color components (i.e., Y, Cb, Cr) of the current block skips the encoding via transform (i.e., when the transform skip flag of the corresponding component is 1), the non-separable transform may not be applied to the transform blocks of all color components. When any color component skips the encoding via transform, the transform index of the current block may not be notified by signaling. The corresponding transform index can be derived as 0.
[0389] Alternatively, transform skipping can be applied to the transform blocks of some of the color components in the current block, while transform skipping can be omitted from the transform blocks of other components. In this case, non-separable transforms can be omitted from the transform blocks of the color components to which transform skipping is applied, while non-separable transforms can be applied to the transform blocks of the remaining color components.
[0390] Based on whether the color components for which transform skipping is applied meet predetermined conditions, an inseparable transform can be applied to color components for which transform skipping is not applied, and the transform index of the corresponding color component can be signaled. As an example, the transform index can be signaled for color components for which transform skipping is not applied only if the color components for which transform skipping is applied meet the predetermined conditions. Here, the predetermined conditions may include at least one of the following: a first condition for the luminance component, a second condition for the chrominance component, or a third condition that a non-zero transform coefficient exists at a position other than the top-left sample position in the transform block of at least one color component. Alternatively, the predetermined conditions for color components for which transform skipping is applied can be disregarded.
[0391] Reference Figure 4 The current block S420 can be reconstructed based on the residual samples of the current block.
[0392] The predicted samples for the current block can be derived based on the intra-prediction mode of the current block. The reconstructed samples for the current block can be generated based on the predicted samples and residual samples of the current block.
[0393] Figure 6 A schematic configuration of a decoding device (300) performing an image decoding method according to the present disclosure is shown.
[0394] Reference Figure 6The decoding apparatus (300) according to this disclosure may include a transform coefficient deriver (600), a residual sample deriver (610), and a reconstruction block generator (620). The transform coefficient deriver (600) may be configured in Figure 3 In the entropy decoder (310), the residual sample deriver (610) can be configured in Figure 3 In the residual processor (320), the reconstructed block generator (620) can be configured in Figure 3 In the adder (340).
[0395] The transform coefficient derivator (600) can obtain the residual information of the current block from the bit stream and decode it to derive the transform coefficients of the current block.
[0396] The residual sample derivator (610) can derive the residual sample of the current block by performing at least one of dequantization or inverse transformation on the transform coefficients of the current block.
[0397] The residual sample derivator (610) can determine the transformation kernel of the inverse transform of the current block using a predetermined transformation kernel determination method, and derive the residual samples of the current block based on this kernel. This is different from the reference... Figure 4 The descriptions are the same, so their detailed descriptions will be omitted here.
[0398] The reconstructed block generator (620) can reconstruct the current block based on the residual samples of the current block.
[0399] Figure 7 An image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure is shown.
[0400] Reference Figure 7 The residual sample of the current block (S700) can be derived.
[0401] The residual samples of the current block can be derived by subtracting the predicted samples from the original samples of the current block. Here, the predicted samples can be derived based on a predetermined intra-frame prediction mode.
[0402] Reference Figure 7 The transformation coefficients of the current block can be derived by performing at least one of transformation or quantization on the residual samples of the current block (S710).
[0403] The transformation method according to this disclosure can be understood as referring to Figure 4 The inverse process of the described inverse transform. Methods and references for determining the transform kernel. Figure 4 The description is the same. A detailed description will be omitted here.
[0404] For example, one or more transformation sets can be defined / configured for the current block, and each transformation set can include one or more transformation kernel candidates. In this case, one of the multiple transformation sets can be selected as the transformation set for the current block. One of the multiple transformation kernel candidates belonging to the transformation set of the current block can be selected. This selection can be performed implicitly based on the context of the current block. Alternatively, the optimal transformation set and / or transformation kernel candidate for the current block can be selected, and its index can be indicated by a signal.
[0405] Alternatively, the transform kernel for the current block can be determined based on an MTS set. One of multiple MTS sets can be selected based on at least one of the current block size or intra-prediction mode. The selected MTS set may include one or more transform kernel candidates. One or more transform kernel candidates can be selected, and the transform kernel for the current block can be determined based on the selected transform kernel candidate. The selection of transform kernel candidates can be performed using a transform kernel candidate index derived from the context of the current block. Alternatively, the optimal transform kernel candidate for the current block can be selected, and the transform kernel candidate index indicating the selected transform kernel candidate can be signaled.
[0406] Alternatively, the transform kernel for the current block can be determined based on the Non-Separable Principal Transform (NSPT) kernel. When the size of the current block belongs to the first group of block sizes to which NSPT is applicable, forward NSPT can be applied to the current block; when the size of the current block belongs to the second group, forward NSPT may not be applied. When the size of the current block belongs to the second group, forward separable principal transforms (e.g., DCT-2) can be applied to the residual samples of the current block to derive the transform coefficients. Forward LFNST can be additionally applied to all or some of the transform coefficients derived through the separable principal transform.
[0407] Additionally, NSPT can be applied based on at least one of the tree type or component type of the current block. The NSPT kernel (or NSPT matrix) used for NSPT can be determined using symmetry between intra-prediction modes or symmetry between block shapes. When applying forward NSPT to an M×N current block, the NSPT kernel can be represented as r×MN. Here, r represents the output length of the NSPT or the number of transform coefficients generated by the NSPT, and MN is the product of the width and height of the current block, which can represent the input length of the NSPT or the number of residual samples to which the NSPT is applied. The method for determining the size of this NSPT kernel is referenced. Figure 4 The description is the same.
[0408] The LFNST and / or NSPT indices used for transformations can be encoded into an integrated grammar, or the LFNST and NSPT indices can be encoded separately and inserted into the bitstream. This includes binarization of the LFNST and NSPT indices, and the allocation and referencing of the CABAC context and initial values. Figure 4 The description is the same.
[0409] Additionally, refer to Figure 4 A method for signaling the transformation index is described, which can also be applied to methods for encoding the transformation index.
[0410] Reference Figure 7 A bitstream can be generated by encoding the transform coefficients of the current block (S720).
[0411] Residual information about the transform coefficients can be generated based on the transform coefficients of the current block, and a bit stream can be generated by encoding the residual information.
[0412] Figure 8 A schematic configuration of an encoding device (200) performing an image encoding method according to the present disclosure is shown.
[0413] Reference Figure 8 The encoding device (200) according to this disclosure may include a residual sample derivative (800), a transform coefficient derivative (810), and a transform coefficient encoder (820). The residual sample derivative (800) and the transform coefficient derivative (810) may be configured in... Figure 2 In the residual processor (230), the transform coefficient encoder (820) can be configured in Figure 2 In the entropy encoder (240).
[0414] The residual sample derivator (800) can derive the residual sample of the current block by subtracting the predicted sample from the original sample of the current block. Here, the predicted sample can be derived based on a predetermined intra-frame prediction mode.
[0415] The transform coefficient derivator (810) can derive the transform coefficients of the current block by performing at least one of transform or quantization on the residual samples of the current block. The transform coefficient derivator 810 can determine the transform kernel of the current block based on at least one of the embodiments 1 to 3 described above and apply the transform kernel to the residual samples of the current block to derive the transform coefficients.
[0416] The transform coefficient encoder (820) can encode the transform coefficients of the current block to generate a bit stream.
[0417] In the above embodiments, the method is described based on a flowchart as a series of steps or blocks. However, the corresponding embodiments are not limited to this order of steps. Some steps may occur simultaneously with other steps or in a different order, as described above. In addition, those skilled in the art will understand that the steps shown in the flowchart are not exclusive. Other steps may be included, or one or more steps in the flowchart may be deleted, without affecting the scope of the embodiments of this disclosure.
[0418] The methods described above according to embodiments of the present disclosure can be implemented in software, and the encoding and / or decoding devices according to the present disclosure can be included in an apparatus for performing image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0419] In this disclosure, when the implementation is implemented as software, the above-described method can be implemented as a module (process, function, etc.) performing the above-described functions. The module can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, logic circuitry, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the implementations described herein can be executed by implementation on a processor, microprocessor, controller, or chip. For example, the functional units shown in the various figures can be executed by implementation on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0420] Furthermore, decoding and encoding devices employing embodiments of this disclosure can be included in multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices similar to video communication, mobile streaming devices, storage media, cameras, devices for providing video-on-demand (VOD) services, OTT (over-the-top) video devices, devices for providing internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and can be used to process video signals or data signals. For example, OTT (over-the-top) video devices can include game consoles, Blu-ray players, internet-connected televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0421] Furthermore, the processing methods applying the embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to the embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical media storage devices. Additionally, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission via the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired / wireless communication networks.
[0422] Furthermore, the embodiments of this disclosure can be implemented by a computer program product using program code, and the program code can be executed on a computer using the embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0423] Figure 9 Examples of content streaming systems to which embodiments of this disclosure can be applied are shown.
[0424] Reference Figure 9 A content streaming system that applies embodiments of the present disclosure may generally include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0425] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0426] A bitstream can be generated by an encoding method or bitstream generation method that applies the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0427] A streaming server sends multimedia data to a user device via a web server based on a user request, and the web server acts as a medium to inform the user what services are available. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which controls the commands / responses between the various devices in the content streaming system.
[0428] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a specific time period.
[0429] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touchscreen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0430] In a content streaming system, each server can operate as a distributed server, and in this case, the data received from each server can be distributed and processed.
[0431] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as an apparatus, and the technical features of the apparatus claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the apparatus claims of this disclosure can be combined and implemented as an apparatus, and the technical features of the method claims and the apparatus claims of this disclosure can be combined and implemented as a method.
Claims
1. An image decoding method, the image decoding method comprising the steps of: obtaining residual information from a bitstream; deriving transform coefficients of a current block based on the residual information; deriving residual samples of the current block by performing at least one of dequantization or inverse transform on the transform coefficients of the current block; and reconstructing the current block based on the residual samples of the current block, wherein the inverse transform is performed based on a non-separable primary transform (NSPT), and wherein the NSPT is applied based on at least one of a size of the current block, a tree type, or a component type. A pre-defined allowed transform block size is divided into a first group as a set of block sizes to which the NSPT is applicable and a second group as a set of block sizes for which the NSPT is not applied.
2. The image decoding method of claim 1, wherein, When the size of the current block belongs to the first group, the inverse transform of the current block is performed based on the NSPT.
3. The image decoding method according to claim 2, wherein When the size of the current block belongs to the second group, the inverse transform of the current block is performed based on a separable primary transform.
4. The image decoding method of claim 2, wherein, When the size of the current block belongs to the second group, the inverse transform of the current block is performed based on a non-separable secondary transform and a separable primary transform.
5. The image decoding method of claim 2, wherein, The first group includes 4×4, and the second group includes 8×8.
6. The image decoding method of claim 2, wherein, The first group includes 4×8 or 8×4, and the second group includes 16×16.
7. The image decoding method of claim 2, wherein, The first group includes 4×16 or 16×4, and the second group includes 16×32 or 32×16.
8. The image decoding method of claim 2, wherein, The first group includes 4×32 or 32×4, and the second group includes 32×32.
9. The image decoding method of claim 2, wherein, The first group includes 8×32 or 32×8, and the second group includes 32×32.
10. The image decoding method of claim 2, wherein, 11.An image encoding method, the image encoding method comprising the steps of: deriving residual samples of a current block; deriving transform coefficients of the current block by performing at least one of a transform or quantization on the residual samples of the current block; and encoding the transform coefficients of the current block, wherein the transform is performed based on a non-separable primary transform (NSPT), and wherein the NSPT is applied based on at least one of a size of the current block, a tree type, or a component type. 12.A computer-readable storage medium storing a bitstream generated by the image encoding method according to claim 11. 13.A method for transmitting data, the method comprising the steps of: obtaining a bitstream for image information, wherein the bitstream is generated by deriving residual samples of a current block, deriving transform coefficients by performing at least one of a transform or quantization on the residual samples of the current block, and encoding the transform coefficients of the current block; and transmitting data including the bitstream, wherein the transform is performed based on a non-separable primary transform (NSPT), and wherein the NSPT is applied based on at least one of a size of the current block, a tree type, or a component type.