Image encoding / decoding method and apparatus and recording means store bit streams
Patent Information
- Application Number
- BR112025022009
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Publication Date
- 2026-09-15
Smart Images

Figure 00000000_0000_ABST
Description
1 / 135 “METHOD AND APPARATUS FOR ENCODING / DECODING IMAGES, AND MEANS OF RECORDING STORING BIT STREAMS” Technical Field
[001] This disclosure relates to a method and apparatus for encoding / decoding images and a means of storing a bit stream. Antecedent Technique
[002] Recently, the demand for high-resolution, high-quality images, such as HD (high definition) and UHD (ultra-high definition) images, has increased in various fields of application, and consequently, highly efficient image compression technologies are being discussed.
[003] There are a variety of technologies, such as interprediction technology which predicts a pixel value included in a current photo from a previous or subsequent photo to a current photo with video compression technology, intraprediction technology which predicts a pixel value included in a current photo using pixel information in a current photo, entropy coding technology which allocates a short signal to a value with a high frequency of appearance and a long signal to a value with a low frequency of appearance, etc. and these image compression technologies can be used to effectively compress image data and transmit or store it. Disclosure Technical Problem
[004] The present disclosure aims to provide a method and apparatus for performing a transform using a non-separable primary transform.
[005] The present disclosure aims to provide a method and apparatus for performing a transform using a non-separable primary transform kernel of reduced dimension.
[006] The purpose of this disclosure is to provide a method and apparatus for Petition 870250092804, dated 10 / 10 / 2025, page 10 / 157 2 / 135 determine / signal a non-separable transform kernel based on an encoding parameter. Technical Solution
[007] An image decoding method and apparatus according to the present disclosure can obtain residual information from a bitstream, derive transform coefficients of a current block based on the residual information, derive residual samples of the current block by performing at least one dequantization or inverse transform on the transform coefficients of the current block, and reconstruct the current block based on residual samples of the current block. Herein, the inverse transform can be performed based on a non-separable primary transform (NSPT), and the NSPT can be applied based on at least one size, tree type, or component type of the current block.
[008] In an image decoding method and apparatus according to the present disclosure, predefined allowed transform block sizes can be divided into a first group which is a set of block sizes to which NSPT can be applied and a second group which is a set of block sizes to which NSPT is not applied.
[009] In an image decoding method and apparatus according to the present disclosure, when a current block size belongs to the first group, the inverse current block transform can be performed based on the NSPT.
[010] In an image decoding method and apparatus according to the present disclosure, when a current block size belongs to the second group, the inverse transform of the current block can be performed based on a separable primary transform.
[011] In an image decoding method and apparatus according to the present disclosure, when a current block size belongs to the second group, the inverse current block transform can be performed based on a Petition 870250092804, dated 10 / 10 / 2025, page 11 / 157 3 / 135 a non-separable secondary transform and a separable primary transform.
[012] In a method and apparatus for decoding images according to the present disclosure, the first group may include 4x4, and the second group may include 8x8.
[013] In a method and apparatus for decoding images according to the present disclosure, the first group may include 4x8 or 8x4, and the second group may include 16x16.
[014] In a method and apparatus for decoding images according to the present disclosure, the first group may include 4x16 or 16x4, and the second group may include 16x32 or 32x16.
[015] In a method and apparatus for decoding images according to the present disclosure, the first group may include 4x32 or 32x4, and the second group may include 32x32.
[016] In a method and apparatus for decoding images according to the present disclosure, the first group may include 8x32 or 32x8, and the second group may include 32x32.
[017] An image coding method and apparatus according to the present disclosure can derive residual samples from a current block, derive transform coefficients from the current block by performing at least one transform or quantization on residual samples from the current block, and encode transform coefficients from the current block. Herein, the transform can be performed based on a non-separable primary transform (NSPT), and the NSPT can be applied based on at least one size, tree type, or component type of the current block.
[018] A computer-readable digital storage medium is provided that stores encoded video / image information, which causes the image decoding method to be performed by a decoding device. Petition 870250092804, dated 10 / 10 / 2025, page 12 / 157 4 / 135 according to this disclosure.
[019] A computer-readable digital storage medium is provided that stores video / image information generated in accordance with the image encoding method as disclosed herein.
[020] A method and apparatus are provided for transmitting video / image information generated in accordance with an image encoding method as disclosed herein. Advantageous Effects
[021] The present disclosure can improve the performance of a transform by using a non-separable primary transform as a primary transform.
[022] The present disclosure can improve the performance of a transform by performing a transform using a reduced-dimensional non-separable primary transform kernel.
[023] This disclosure can improve encoding efficiency by effectively determining and / or signaling a non-separable transform kernel based on an encoding parameter. Brief Description of the Drawings
[024] Figure 1 shows a video / image encoding system according to the present disclosure.
[025] Figure 2 shows a schematic block diagram of an encoding apparatus to which an embodiment of the present disclosure is applicable and the encoding of video / image signals is performed.
[026] Figure 3 shows a schematic block diagram of a decoding apparatus to which an embodiment of the present disclosure is applicable and the decoding of video / image signals is performed.
[027] Figure 4 illustrates an image decoding method. Petition 870250092804, dated 10 / 10 / 2025, page 13 / 157 5 / 135 performed by a decoding apparatus (300) as an embodiment in accordance with the present disclosure.
[028] Figure 5 shows examples of intraprediction modes and their prediction directions according to this disclosure.
[029] Figure 6 illustrates a schematic configuration of a decoding apparatus (300) that performs an image decoding method according to the present disclosure.
[030] Figure 7 illustrates an image encoding method performed by an encoding apparatus (200) according to an embodiment of the present disclosure.
[031] Figure 8 illustrates a schematic configuration of an encoding apparatus (200) that performs an image encoding method according to the present disclosure.
[032] Figure 9 shows an example of a continuous content streaming system to which the modalities of the present disclosure can be applied. Best Way
[033] As the present disclosure may undergo several changes and have several embodiments, specific embodiments will be illustrated in a drawing and described in detail in a detailed description. However, the present disclosure is not intended to limit the present embodiment to a specific embodiment and should be understood as including all alterations, equivalents and substitutes included in the spirit and technical scope of the present disclosure. In describing each drawing, similar reference numbers are used for similar components.
[034] A term such as first, second, etc. may be used to describe several components, but the components should not be limited by the terms. The terms are used only to distinguish one component from other components. Petition 870250092804, dated 10 / 10 / 2025, page 14 / 157 6 / 135 For example, the first component may be referred to as the second component without departing from the scope of a right under this disclosure, and similarly, the second component may also be referred to as the first component. An "and / or" term includes any one of a plurality of related stated items or a combination of a plurality of related stated items.
[035] When a component is referred to as being connected or being linked to another component, it should be understood that it may be directly connected or linked to another component, but there may be another component in between. Conversely, when a component is referred to as “being directly connected” or “being directly linked” to another component, it should be understood that there is no other component in between.
[036] The term used in this application is used only to describe a specific embodiment and is not intended to limit the present disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that a term such as “include” or “have”, etc., is intended to designate the presence of attributes, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but does not in advance exclude the possibility of the presence or addition of one or more other attributes, numbers, steps, operations, components, parts, or combinations thereof.
[037] This disclosure relates to video / image coding. For example, a method / embodiment disclosed herein may be applied to a method disclosed in the Versatile Video Coding Standard (VVC). In addition, a method / embodiment disclosed herein may be applied to a method disclosed in the Essential Video Coding Standard (EVC), the AOMedia Video 1 (AV1) standard, the 2nd generation Audio Video Coding Standard (AVS2), or the standard of Petition 870250092804, dated 10 / 10 / 2025, page 15 / 157 7 / 135 next-generation video / image encoding (e.g., H.267 or H.268, etc.).
[038] This specification proposes several video / image encoding modes and, unless otherwise specified, the modes may be implemented in combination with each other.
[039] Here, a video can refer to a set of a series of images over time. A photo generally refers to a unit that represents an image at a specific time period, and a slice / tile is a unit that is part of a photo in the encoding. A slice / tile can include at least one encoding tree unit (CTU). A photo can consist of at least one slice / tile. A tile is a rectangular region composed of a plurality of CTUs within a specific tile column and a specific tile row of a photo. A tile column is a rectangular region of CTUs that has the same height as a photo and a width designated by a syntax requirement of a photo parameter set. A tile row is a rectangular region of CTUs that has a height designated by a photo parameter set and the same width as a photo.CTUs within a tile can be arranged consecutively according to the CTU raster scan, just as tiles within a photo can be arranged consecutively according to the tile raster scan. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a photo tile that can be uniquely included in a single NAL unit. Meanwhile, a photo can be divided into at least two sub-photos. A sub-photo can be a rectangular region of at least one slice within a photo.
[040] A pixel, pixel, or pel can refer to the smallest unit that constitutes a photo (or image). Additionally, 'sample' can be used as a term corresponding to a pixel. A sample can generally represent a pixel or Petition 870250092804, dated 10 / 10 / 2025, page 16 / 157 8 / 135 is a pixel value and can represent just one pixel / a pixel value of a luma component or just one pixel / a pixel value of a chroma component.
[041] A unit may represent a basic unit of image processing. A unit may include at least one specific region of a photo and information related to the corresponding region. A unit may include one luma block and two chroma blocks (e.g., cb, cr). In some cases, a unit may be used interchangeably with a term such as a block or a region, etc. In a general case, an MxN block may include a set (or array) of transform coefficients or samples (or arrays of samples) consisting of M columns and N rows.
[042] Here, “A or B” can refer to “only A”, “only B” or “both A and B”. In other words, here, “A or B” can be interpreted as “A and / or B”. For example, here, “A, B or C” can refer to “only A”, “only B”, “only C” or “any combination of A, B and C)”.
[043] A slash ( / ) or a comma used here can refer to “and / or”. For example, “A / B” can refer to “A and / or B”. Consequently, “A / B” can refer to “only A”, “only B” or “both A and B”. For example, “A, B, C” can refer to “A, B or C”.
[044] Here, “at least one of A and B” can refer to “only A”, “only B” or “both A and B”. Furthermore, here, an expression such as “at least one of A or B” or “at least one of A and / or B” can be interpreted in the same way as “at least one of A and B”.
[045] Furthermore, here, “at least one of A, B and C” may refer to “only A”, “only B”, “only C” or “any combination of A, B and C”. Furthermore, “at least one of A, B or C” or “at least one of A, B and / or C” may refer to “at least one of A, B and C”. Petition 870250092804, dated 10 / 10 / 2025, p. 17 / 157 9 / 135
[046] Furthermore, a parenthesis used here may refer to “for example”. Specifically, when indicated as “prediction (intraprediction)”, “intraprediction” may be proposed as an example of “prediction”. In other words, “prediction” here is not limited to “intraprediction” and “intraprediction” may be proposed as an example of “prediction”. Furthermore, even when indicated as “prediction (i.e., intraprediction)”, “intraprediction” may be proposed as an example of “prediction”.
[047] Here, a technical attribute described individually in a drawing can be implemented individually or simultaneously.
[048] Figure 1 shows a video / image encoding system according to the present disclosure.
[049] With reference to Figure 1, a video / image encoding system may include the first device (a source device) and the second device (a receiving device).
[050] A source device may transmit encoded video / image information or data in file form or streaming format to a receiving device via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device 12 may be referred to as a video / image encoding device, and the encoding device 22 may be referred to as a video / image encoding device. A transmitter may be included in an encoding device. A receiver may be included in a decoding device. A renderer may include a display unit, and a display unit may consist of a separate device or an external component.
[051] A video source can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. A source of Petition 870250092804, dated 10 / 10 / 2025, page 18 / 157 10 / 135 Video may include a video / image capture device and a video / image generation device. A device for capturing a video / image may include at least one camera, a video / image file including previously captured videos / images, etc. A device for generating a video / image may include a computer, a tablet, a smartphone, etc., and may generate (electronically) a video / image. For example, a virtual video / image may be generated by a computer, etc., and in this case, a process of capturing a video / image may be replaced by a process of generating related data.
[052] An encoding device can encode an input video / image. The encoding device can perform a series of procedures, such as prediction, transformation, quantization, etc., for compression and encoding efficiency. Encoded data (encoded video / image information) can be output in the form of a bitstream.
[053] A transmission unit can transmit encoded video / image information or output data in the form of a bitstream to a receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming transmission. A digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit may include an element to generate a media file via a predetermined file format and may include an element for transmission via a broadcast / communication network. A receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[054] An encoding device can decode a video / image by performing a series of procedures such as dequantization, inverse transform, prediction, etc. corresponding to the operation of an encoding device. Petition 870250092804, dated 10 / 10 / 2025, page 19 / 157 11 / 135
[055] A renderer can render a decoded video / image. A rendered video / image can be displayed through a display unit.
[056] Figure 2 shows an approximate block diagram of an encoding apparatus to which an embodiment of the present disclosure can be applied and the encoding of a video / image signal is performed.
[057] With reference to Figure 2, a coding apparatus 200 may consist of an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. A predictor 220 may include an interpredictor 221 and an intrapredictor 222. A residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. A residual processor 230 may additionally include a subtractor 231. An adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above can be configured by at least one hardware component (e.g., an encoder chip set or a processor) according to an embodiment.In addition, a 270 memory may include a photodecoded buffer (DPB) and may be configured by a digital storage medium. The hardware component may additionally include a 270 memory as an internal / external component.
[058] An image partitioner 210 can partition an input image (or photo, frame) inserted into an image encoding device 200 into at least one processing unit. As an example, the processing unit can be referred to as an encoding unit (CU). In this case, an encoding unit can be recursively partitioned according to a quaternary tree (quad-tree), binary tree, and ternary tree (QTBTTT) structure from an encoding tree unit. Petition 870250092804, dated 10 / 10 / 2025, page 20 / 157 12 / 135 (CTU) or the largest coding unit (LCU).
[059] For example, a coding unit can be partitioned into a plurality of coding units with greater depth based on a quaternary tree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quaternary tree structure can be applied first, and a binary tree structure and / or a ternary structure can be applied later. Alternatively, a binary tree structure can be applied before a quaternary tree structure. A coding procedure according to this specification can be performed based on a final coding unit that is no longer partitioned. In this case, based on coding efficiency, etc.According to an image characteristic, the largest encoding unit can be used directly as a final encoding unit, or, if necessary, an encoding unit can be recursively partitioned into encoding units of greater depth, and an encoding unit with an ideal size can be used as a final encoding unit. Here, an encoding procedure can include procedures such as prediction, transformation, and reconstruction, etc., described later.
[060] As another example, the processing unit may additionally include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or partitioned from a final encoding unit described above, respectively. The prediction unit may be a sample prediction unit, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[061] In some cases, a unit can be used interchangeably Petition 870250092804, dated 10 / 10 / 2025, page 21 / 157 13 / 135 with a term such as a block or a region, etc. In a general case, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value and can represent only a pixel / a pixel value of a luma component or only a pixel / a pixel value of a chroma component. A sample can be used as a term that makes a photo (or image) correspond to a pixel or a pel.
[062] A coding device 200 can subtract a prediction signal (a prediction block, a prediction sample array) emitted from an interpredictor 221 or an intrapredictor 222 from an input image signal (an original block, an original sample array) to generate a residual signal (a residual signal, a residual sample array), and a generated residual signal is transmitted to a transformer 232. In this case, a unit that subtracts a prediction signal (a prediction block, a prediction sample array) from an input image signal (an original block, an original sample array) within a coding device 200 can be referred to as a subtractor 231.
[063] A 220 predictor can perform a prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block including prediction samples for the current block. A 220 predictor can determine whether intra- or interprediction is applied to a unit of a current block or a CU. A 220 predictor can generate various prediction information, such as prediction mode information, etc., and transmit it to a 240 entropy encoder, as described later in a description of each prediction mode. Prediction information can be encoded in a 240 entropy encoder and output in the form of a bitstream.
[064] An intrapredictor 222 can predict a current block by referring to samples within a current photo. The samples mentioned can be Petition 870250092804, dated 10 / 10 / 2025, page 22 / 157 14 / 135 are positioned in the vicinity of the current block or can be positioned at a certain distance from the current block according to a prediction mode. In intraprediction, prediction modes can include at least one non-directional mode and a plurality of directional modes. A non-directional mode can include at least one DC mode or a planar mode. A directional mode can include 33 directional modes or 65 directional modes according to the level of detail of a prediction direction. However, this is an example, and more or less directional modes can be used according to a configuration. An intrapredictor 222 can also determine a prediction mode applied to a current block using the prediction mode applied to a neighboring block.
[065] An interpredictor 221 can derive a prediction block for a current block based on a reference block (a reference sample array) specified by a motion vector in a reference photo. In this case, in order to reduce the amount of motion information transmitted in an interprediction mode, motion information can be predicted at a unit of one block, one sub-block, or one sample, based on the motion information correlation between a neighboring block and a current block. The motion information can include a motion vector and a reference photo index. The motion information can additionally include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For interprediction, a neighboring block can include an existing spatial neighboring block in a current photo and an existing temporal neighboring block in a reference photo.A reference photo including the reference block and a reference photo including the neighboring time block may be the same or different. The neighboring temporal block may be referred to as a colocalized reference block, colocalized CU (colCU), etc., and a reference photo including the neighboring temporal block may be referred to as a colocalized photo (colPic). For example, an interpreter 221 may set up a list of candidates. Petition 870250092804, dated 10 / 10 / 2025, p. 23 / 157 15 / 135 of motion information based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference photo index of the current block. Interprediction can be performed based on various prediction modes and, for example, for a jump mode and a merge mode, an interpredictor 221 can use motion information from a neighboring block as motion information from a current block. For a jump mode, unlike the merge mode, a residual signal may not be transmitted. For a motion vector prediction (MVP) mode, a motion vector from a neighboring block is used as a motion vector predictor and a motion vector difference is signaled to indicate a motion vector from a current block.
[066] A 220 predictor can generate a prediction signal based on several prediction methods described later. For example, a predictor can not only apply intraprediction or interprediction for block prediction, but can also apply intraprediction and interprediction simultaneously. This can be referred to as a combined inter- and intraprediction (CIIP) mode. Furthermore, a predictor can be based on an intrablock copy (IBC) prediction mode or it can be based on a palette mode for block prediction. The IBC prediction mode or palette mode can be used for image / video content encoding of a game, etc., such as screen content encoding (SCC), etc. IBC basically performs prediction within a current frame, but it can be performed similarly to interprediction as it derives a reference block within a current frame. In other words, IBC can use at least one of the interprediction techniques described here.A palette mode can be considered an example of intracoding or intraprediction. When a palette mode is applied, a sample value within a photo can be signaled based on information in a palette table and a palette index. A prediction signal is generated through this. Petition 870250092804, dated 10 / 10 / 2025, page 24 / 157 The 16 / 135 predictor can be used to generate a reconstructed signal or a residual signal.
[067] A 232 transformer can generate transform coefficients by applying a transform technique to a residual signal. For example, a transform technique can include at least one of the following: Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditionally Non-Linear Transform (CNT). Here, GBT refers to a transform obtained from this graph when the relationship information between pixels is expressed as a graph. CNT refers to the transform obtained based on the generation of a prediction signal using all previously reconstructed pixels. Furthermore, a transform process can be applied to a block of square pixels of the same size or it can be applied to a non-square block of variable size.
[068] A 233 quantizer can quantize transform coefficients and transmit them to a 240 entropy encoder, and a 240 entropy encoder can encode a quantized signal (information about quantized transform coefficients) and generate it as a bit stream. Information about the quantized transform coefficients can be called residual information. A 233 quantizer can rearrange quantized transform coefficients in a block form into a 1D vector form based on the scanning order of the coefficients and can generate information about the quantized transform coefficients based on the quantized transform coefficients in 1D vector form.
[069] A 240 entropy encoder can perform various encoding methods, such as exponential Golomb, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. A 240 entropy encoder can encode information needed for video / image reconstruction (e.g., a syntax element value, Petition 870250092804, dated 10 / 10 / 2025, page 25 / 157 17 / 135 etc.) different from quantized transform coefficients, either together or separately.
[070] Encoded information (e.g., encoded video / image information) can be transmitted or stored in a Network Abstraction Layer (NAL) unit in a bitstream format. The video / image information may additionally include information about various parameter sets, such as an adaptation parameter set (APS), a photo parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS), etc. Furthermore, the video / image information may additionally include general restriction information. Here, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded using the encoding procedure described above and included in the bitstream.The bit stream can be transmitted over a network or stored on a digital storage medium. Here, a network may include a transmission network and / or a communication network, etc., and a digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing a signal output from a 240 entropy encoder may be configured as an internal / external element of a 200 encoding device, or a transmission unit may also be included in a 240 entropy encoder.
[071] Quantized transform coefficients emitted by a quantizer 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be reconstructed by applying dequantization and inverse transform to quantized transform coefficients via a dequantizer 234 and an inverse transformer 235. A summer Petition 870250092804, dated 10 / 10 / 2025, page 26 / 157 18 / 135 A 250 adder can add a reconstructed residual signal to a prediction signal output from an interpredictor 221 or an intrapredictor 222 to generate a reconstructed signal (a reconstructed photo, a reconstructed block, a reconstructed sample array). When there is no residual for a block to be processed, such as when a jump mode is applied, a predicted block can be used as a reconstructed block. A 250 adder can be referred to as a reconstructed block reconstructor or generator. A generated reconstructed signal can be used for intraprediction of a next block to be processed within a current photo, and can also be used for interprediction of a next photo through filtering, as described later. Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in a photo encoding and / or reconstruction process.
[072] A 260 filter can enhance the subjective / objective quality of the image by applying filtering to a reconstructed signal. For example, a 260 filter can generate a modified reconstructed photo by applying various filtering methods to a reconstructed photo and can store the modified reconstructed photo in a 270 memory, specifically in a DPB of a 270 memory. The various filtering methods can include deblock filtering, adaptive sample shift, adaptive loop filter, bilateral filter, etc. A 260 filter can generate various filtering information and transmit it to a 240 entropy encoder. Filtering information can be encoded in a 240 entropy encoder and output in the form of a bitstream.
[073] A modified reconstructed photo transmitted to a memory 270 can be used as a reference photo in an interpredictor 221. When interprediction is applied through it, an encoding device can avoid prediction mismatch in an encoding device 200 and a decoding device, and can also improve encoding efficiency.
[074] A DPB of a 270 memory can store a reconstructed photo Petition 870250092804, dated 10 / 10 / 2025, page 27 / 157 19 / 135 modified to use it as a reference photo in an interpredictor 221. A memory 270 can store motion information of a block from which motion information in a current photo is derived (or encoded) and / or motion information of blocks in a pre-reconstructed photo. The stored motion information can be transmitted to an interpredictor 221 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. A memory 270 can store reconstructed samples of the reconstructed blocks in a current photo and transmit them to an intrapredictor 222.
[075] Figure 3 shows an approximate block diagram of a decoding device to which an embodiment of the present disclosure can be applied and the decoding of a video / image signal is performed.
[076] With reference to Figure 3, a decoding apparatus 300 can be configured including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350 and a memory 360. A predictor 330 can include an interpredictor 331 and an intrapredictor 332. A residual processor 320 can include a dequantizer 321 and an inverse transformer 321.
[077] According to one embodiment, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be configured by a hardware component (e.g., a decoder chip set or a processor). Furthermore, a memory 360 may include a photodecoded buffer (DPB) and may be configured by a digital storage medium. The hardware component may additionally include a memory 360 as an internal / external component.
[078] When a bitstream including video / image information is entered, a decoding device 300 can reconstruct an image in response to a process in which the video / image information is processed. Petition 870250092804, dated 10 / 10 / 2025, page 28 / 157 20 / 135 in an encoding device of Figure 2. For example, a decoding device 300 can derive units / blocks based on block partitioning information obtained from the bitstream. A decoding device 300 can perform decoding using a processing unit applied in an encoding device. Consequently, a decoding processing unit can be an encoding unit, and an encoding unit can be partitioned from an encoding tree unit or the larger encoding unit according to a quaternary tree structure, a binary tree structure, and / or a ternary tree structure. At least one transform unit can be derived from an encoding unit. And a reconstructed, decoded image signal emitted through a decoding device 300 can be reproduced through a playback device.
[079] A decoding device 300 can receive an output signal from an encoding device of Figure 2 in the form of a bitstream, and a received signal can be decoded by means of an entropy decoder 310. For example, an entropy decoder 310 can analyze the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or photo reconstruction). The video / image information may additionally include information about various parameter sets, such as an adaptation parameter set (APS), a photo parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. A decoding device can decode a photo based on information about the parameter set and / or the general constraint information.The signaled / received information and / or syntax elements described further below can be decoded through... Petition 870250092804, dated 10 / 10 / 2025, page 29 / 157 21 / 135 decoding procedure and obtained from the bit stream. For example, an entropy decoder 310 can decode information in a bit stream based on an encoding method, such as Golomb exponential encoding, CAVLC, CABAC, etc., and can output a syntax element value necessary for image reconstruction and quantized values of a transform coefficient related to a residue.In more detail, a CABAC entropy decoding method can receive a bin corresponding to each syntax element of a bitstream, determine a context model using information from the syntax element to be decoded, decoding information from a neighboring block and a block to be decoded, or information from a symbol / bin decoded in a previous step, perform the arithmetic decoding of a bin predicting a probability of occurrence of a bin according to a determined context model, and generate a symbol corresponding to a value of each syntax element. In this case, a CABAC entropy decoding method can update a context model using information about a decoded symbol / bin to a context model of a next symbol / bin after determining a context model.Among the information decoded in an entropy decoder 310, prediction information is provided to a predictor (an interpredictor 332 and an intrapredictor 331), and a residual value at which entropy decoding was performed in an entropy decoder 310, i.e., quantized transform coefficients and related parameter information, can be fed into a residual processor 320. A residual processor 320 can derive a residual signal (a residual block, residual samples, a residual sample array). Furthermore, filtering information between information decoded in an entropy decoder 310 can be provided to a filter 350. Meanwhile, a receiving unit (not shown) receives an output signal from one. Petition 870250092804, dated 10 / 10 / 2025, page 30 / 157 22 / 135 encoding apparatus may be additionally configured as an internal / external element of a decoding apparatus 300 or a receiving unit may be a component of an entropy decoder 310.
[080] Meanwhile, a decoding apparatus according to this specification may be called a video / image / photo decoding apparatus, and the decoding apparatus may be divided into an information decoder (a video / image / photo information decoder) and a sample decoder (a video / image / photo sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the dequantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the interpredictor 332 or the intrapredictor 331.
[081] A 321 dequantizer can dequantize quantized transform coefficients and output transform coefficients. A 321 dequantizer can rearrange quantized transform coefficients into a two-dimensional block. In this case, the rearrangement can be performed based on the scanning order of the coefficients performed in an encoding apparatus. A 321 dequantizer can perform dequantization on quantized transform coefficients using a quantization parameter (e.g., quantization step size information) and obtain transform coefficients.
[082] An inverse transformer 322 inversely transforms transform coefficients to obtain a residual signal (a residual block, a residual sample array).
[083] A 320 predictor can perform prediction on a current block and generate a predicted block including prediction samples for the current block. A 320 predictor can determine whether intra- or interprediction is applied to the current block based on prediction information emitted from a 310 entropy decoder and can determine a specific intraprediction / interprediction mode. Petition 870250092804, dated 10 / 10 / 2025, page 31 / 157 23 / 135
[084] A 320 predictor can generate a prediction signal based on several prediction methods described later. For example, a 320 predictor can not only apply intraprediction or interprediction for block prediction, but can also apply intraprediction and interprediction simultaneously. This can be referred to as a combined inter- and intraprediction (CIIP) mode. Furthermore, a predictor can be based on an intrablock copy (IBC) prediction mode or it can be based on a palette mode for block prediction. The IBC prediction mode or palette mode can be used for image / video content encoding of a game, etc., such as screen content encoding (SCC), etc. IBC basically performs prediction within a current frame, but it can be performed similarly to interprediction as it derives a reference block within a current frame. In other words, IBC can use at least one of the interprediction techniques described here.A palette mode can be considered an example of intracoding or intraprediction. When a palette mode is applied, information in a palette table and a palette index can be included in the video / image information and flagged.
[085] A 331 intrapredictor can predict a current block by referring to samples within a current photo. The referenced samples can be positioned in the vicinity of the current block or can be positioned at a certain distance from the current block according to a prediction mode. In intraprediction, prediction modes can include at least one non-directional mode and a plurality of directional modes. A 331 intrapredictor can also determine a prediction mode applied to a current block using the prediction mode applied to a neighboring block.
[086] A 332 interpreter can derive a prediction block for a current block based on a reference block (a reference sample array) specified by a motion vector in a reference photo. In this case, the Petition 870250092804, dated 10 / 10 / 2025, page 32 / 157 24 / 135 In order to reduce the amount of motion information transmitted in an interprediction mode, motion information can be predicted at a unit level—a block, a sub-block, or a sample—based on the correlation of motion information between a neighboring block and a current block. Motion information can include a motion vector and a reference photo index. Motion information can additionally include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For interprediction, a neighboring block can include an existing spatial neighboring block in a current photo and an existing temporal neighboring block in a reference photo. For example, an interpredictor 332 can set up a list of motion information candidates based on neighboring blocks and derive a motion vector and / or a reference photo index from the current block based on received candidate selection information.Interprediction can be performed based on various prediction modes, and prediction information may include information indicating an interprediction mode for the current block.
[087] A 340 adder can add a residual signal obtained to a prediction signal (a prediction block, a prediction sample array) emitted from a predictor (including an interpredictor 332 and / or an intrapredictor 331) to generate a reconstructed signal (a reconstructed photo, a reconstructed block, a reconstructed sample array). When there is no residual for a block to be processed, such as when a jump mode is applied, a prediction block can be used as a reconstructed block.
[088] A 340 adder can be referred to as a reconstructor or a reconstructed block generator. A generated reconstructed signal can be used for intraprediction of a next block to be processed in a current photo, can be sent through filtering as described later, or can be used for interprediction of a next photo. Meanwhile, luma mapping with Petition 870250092804, dated 10 / 10 / 2025, page 33 / 157 The 25 / 135 chroma scale (LMCS) can be applied in a photo decoding process.
[089] A 350 filter can enhance the subjective / objective quality of the image by applying filtering to a reconstructed signal. For example, a 350 filter can generate a modified reconstructed photo by applying various filtering methods to a reconstructed photo and transmit the modified reconstructed photo to a 360 memory, specifically a DPB of a 360 memory. The various filtering methods can include unlock filtering, adaptive sample shift, adaptive loop filter, bilateral filter, etc.
[090] The reconstructed (modified) photo stored in the DPB of memory 360 can be used as a reference photo in interpredictor 332. A memory 360 can store motion information of a block from which motion information in a current photo is derived (or decoded) and / or motion information of blocks in a pre-reconstructed photo. The stored motion information can be transmitted to an interpredictor 260 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. A memory 360 can store reconstructed samples of the reconstructed blocks in a current photo and transmit them to an intrapredictor 331.
[091] Here, the modalities described in a filter 260, an interpredictor 221 and an intrapredictor 222 of a coding apparatus 200 can also be applied equally or correspondingly to a filter 350, an interpredictor 332 and an intrapredictor 331 of a decoding apparatus 300, respectively.
[092] Figure 4 illustrates an image decoding method performed by a decoding apparatus (300) as an embodiment according to the present disclosure.
[093] With reference to Figure 4, the transform coefficients of the current block can be derived from the bit stream (S400). That is, the bit stream can include Petition 870250092804, dated 10 / 10 / 2025, page 34 / 157 26 / 135 residual information from the current block, and the current block transform coefficients can be derived by decoding the residual information.
[094] With reference to Figure 4, residual samples of the current block can be derived by performing at least one of the following operations: dequantization and inverse transform on the transform coefficients of the current block (S410).
[095] When Adaptive Multiple Transform Selection (MTS) is applied, the inverse transform can be performed based on at least one of DCT-2, DST-7, or DCT-8. Here, DCT-2, DST-7, DCT-8, etc. can be referred to as transform type, transform kernel, or transform core.
[096] In the present disclosure, the inverse transform may mean a separable transform. However, it is not limited to this, and the inverse transform may mean a non-separable transform, or it may be a concept that includes both a separable and a non-separable transform. Furthermore, the inverse transform in the present disclosure means a primary transform, but is not limited to this, and may be applied to a secondary transform by being modified in an identical / similar way.
[097] For example, as a method for inverse transform, only DCT2 and a non-separable transform can be used, or a non-separable transform can be used in addition to at least one of DCT-2, DST-7 or DCT-8, or a non-separable transform can replace the transform kernel of one or more DCT-2, DST-7 or DCT-8.
[098] As a more specific embodiment, when there are (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), (DCT-8, DCT-8) as transform kernel candidates for a separable transform, a non-separable transform can replace or be added to one or more of the five transform kernel candidates. Here, the notation (transform1, transform2) indicates that transform1 is applied in the horizontal direction and transform2 is applied in the directional direction. Petition 870250092804, dated 10 / 10 / 2025, page 35 / 157 27 / 135 vertical. When the non-separable transform replaces some of the transform kernel candidates, the remaining transform kernel candidates, except (DCT-2, DCT-2) and (DST-7, DST-7), can be replaced by the non-separable transform. However, the transform kernel candidates described above are only examples, and other types of DCT and / or DST can be included, and a transform jump can be included as a transform kernel candidate.
[099] A non-separable transform can mean a transform or inverse transform based on an arrangement of non-separable transforms. That is, unlike a separable transform that performs horizontal and vertical transforms independently, separating vertical and horizontal transforms, a non-separable transform can perform horizontal and vertical transforms at the same time.
[0100] For example, when a non-separable transform is performed on a 4x4 block, the input data X for the non-separable transform are shown in the following equation 1. [Equation 1] X = fX00k 10k20k 30 I
[0101] When the input data X is expressed in vector form, the vector X' can be expressed as follows. [Equation 2] X = [^00,-^01, -^02,^03,^10,^11,-^12,^13,-^20, -^21,-^22,-^23,-^30,-^31,-^32 A33]T
[0102] In this case, the non-separable transform can be performed as in the following equation 3. [Equation 3] F = T · X'
[0103] In equation 3, F represents a transform coefficient vector, Petition 870250092804, dated 10 / 10 / 2025, page 36 / 157 28 / 135 T represents a 16x16 non-separable transformed matrix, and · represents the multiplication of a matrix and a vector.
[0104] A 16x1 transform coefficient vector F can be derived using equation 3, and F can be reconfigured into 4x4 blocks according to a predetermined scan order. The scan order can be a horizontal scan, a vertical scan, a diagonal scan, a z-scan, a raster scan, or a predefined scan.
[0105] The non-separable transform set and / or transform kernel for the non-separable transform can be configured in several ways based on at least one prediction mode (e.g., intra mode, inter mode, etc.), the width, height, or number of pixels of the current block, the position of a sub-block within the current block, explicitly signaled syntax elements, statistical characteristics of neighboring samples, whether a secondary transform is used, or a quantization parameter (QP).
[0106] Specifically, for intra mode, the predefined intraprediction modes can be grouped to match an sets of non-separable transforms, and each set of non-separable transforms can include k transform kernel candidates. Here, nek can be arbitrary constants according to rules (conditions) defined identically for the encoding apparatus and the decoding apparatus.
[0107] The number of non-separable transform sets and / or the number of transform kernel candidates included in the non-separable transform set can be configured differently depending on the width and / or height of the current block. For example, for a 4x4 block, m non-separable transform sets and k1 transform kernel candidates can be configured. For a 4x8 block, n2 non-separable transform sets and k2 transform kernel candidates can be configured. Furthermore, the number of sets Petition 870250092804, dated 10 / 10 / 2025, page 37 / 157 29 / 135 of non-separable transform sets and the number of transform kernel candidates included in each non-separable transform set can be configured differently depending on the product of the width and height of the current block. For example, when the product of the width and height of the current block is equal to or greater than 256, n3 non-separable transform sets and k3 transform kernel candidates can be configured, and otherwise, n4 non-separable transform sets and k4 transform kernel candidates can be configured. That is, since the degree of change in the statistical characteristics of the residual signal varies depending on the block size, the number of non-separable transform sets and transform kernel candidates can be configured differently to reflect this.
[0108] When the current block is divided into a plurality of sub-blocks, the statistical characteristics of the residual signal may be different for each sub-block and therefore the number of non-separable transform sets and transform kernel candidates may be configured differently. For example, when a 4x8 or 8x4 block is divided into two 4x4 sub-blocks and a non-separable transform is applied to each sub-block, n5 non-separable transform sets and k5 transform kernel candidates may be configured for the upper left 4x4 sub-block and n6 non-separable transform sets and k6 transform kernel candidates may be configured for other 4x4 sub-blocks.
[0109] Based on the explicitly signaled syntax element, the number of non-separable transform sets and transform kernel candidates can be configured differently. As a syntax element, information indicating one of a plurality of non-separable transform configurations can be used. For example, when three types of non-separable transform configurations are supported (i.e., n7 transform sets) Petition 870250092804, dated 10 / 10 / 2025, page 38 / 157 30 / 135 non-separable and k7 transform kernel candidates, n8 non-separable transform sets and k8 transform kernel candidates, n9 non-separable transform sets and k9 transform kernel candidates), the syntax element can have values of 0, 1, and 2, and the non-separable transform configuration applied to the current block can be determined based on the value of the signaled syntax element.
[0110] Based on the application of a secondary transform and / or which secondary transform is applied, the number of non-separable transform sets and transform kernel candidates can be configured differently. For example, when a secondary transform is not applied, a non-separable transform configuration including n10 non-separable transform sets and kw transform kernel candidates can be applied. When a secondary transform is applied, a non-separable transform configuration including nu non-separable transform sets and kn transform kernel candidates can be applied.
[0111] Based on the quantization parameter (QP) and / or the range to which the QP value belongs, different non-separable transform configurations can be applied. For example, when the QP value has a small value, a non-separable transform configuration including n12 non-separable transform sets and k12 transform kernel candidates can be applied. Conversely, when the QP value has a large value, a non-separable transform configuration including m3 non-separable transform sets and k13 transform kernel candidates can be applied. When the QP value is less than or equal to a threshold (e.g., 32), the case is classified as having a small QP value, and otherwise, the case is classified as having a large QP value. Alternatively, the range of QP values can be divided into three or more, and different non-separable transform configurations can be applied. Petition 870250092804, dated 10 / 10 / 2025, page 39 / 157 31 / 135 applied for each interval.
[0112] For relatively large blocks, instead of using a non-separable transform corresponding to the width and height of the block, the block can be divided into a plurality of sub-blocks and a non-separable transform corresponding to the width and height of the sub-block can be used. For example, when performing a non-separable transform for a 4x8 block, the 4x8 block can be divided into two 4x4 sub-blocks and a 4x4 block-based non-separable transform can be used for each of the 4x4 sub-blocks. Alternatively, an 8x16 block can be divided into two 8x8 sub-blocks and an 8x8 block-based non-separable transform can be used.
[0113] The set of non-separable transforms can be determined based on the intraprediction mode of the current block and the mapping table. The mapping table can define a mapping relationship between the predefined intraprediction modes and the sets of non-separable transforms. The predefined intraprediction modes can include two non-directional modes and 65 directional modes. In general, the non-separable transform has a larger transform kernel size than the separable transform. This means that the computational complexity required for the transform process is high and the memory required to store the transform kernel is large.Meanwhile, while the separable transform can only consider statistical features existing in the horizontal and / or vertical directions, the non-separable transform can simultaneously consider statistical features in a two-dimensional space, including both horizontal and vertical directions, thus providing better compression efficiency. Since statistical features and residual diversity differ depending on the directionality of the intraprediction mode, there may be cases where the non-separable transform is absolutely necessary, and there may be intraprediction modes where residual features can be... Petition 870250092804, dated 10 / 10 / 2025, page 40 / 157 32 / 135 sufficiently identified only by the separable transform. Therefore, by predefining which transform to use based on the intraprediction mode in the encoding and decoding apparatus, the transform process can be designed with optimized complexity and memory requirements. The non-directional mode can include the planar mode of the number 0 and the DC mode of the number 1, and the directional mode can include the intraprediction modes of the numbers 2 to 66. However, this is only an example, and the present disclosure can also be applied to cases where the number of predefined intraprediction modes is different.
[0114] Due to the application of wide-angle intraprediction (WAIP), the predefined intraprediction modes may additionally include intraprediction modes from -14 to -1 and intraprediction modes from 67 to 80.
[0115] Figure 5 shows examples of intraprediction modes and their prediction directions according to this disclosure. With reference to Figure 5, modes -14 to -1 and 2 to 33 and modes 35 to 80 are symmetrical with respect to mode 34 in terms of prediction direction. For example, modes 10 and 58 are symmetrical with respect to the direction corresponding to mode 34, and mode 1 is symmetrical with respect to mode 67. Consequently, for a vertical directionality mode symmetrical to a horizontal directionality mode with respect to mode 34, the input data can be transposed and used. Transposing the input data means that the rows and columns in the MxN input data of a two-dimensional block become columns and rows, respectively, to form NxM data.
[0116] For example, when a 4x4 block is used, 16 data points forming a 4x4 block can be arranged appropriately to form a one-dimensional 16x1 vector for non-separable transformation. In this case, the one-dimensional vector can be formed in first-row order or first-column order. The residual samples resulting from the non-separable transformation can be arranged in the order above to form a two-dimensional block. Petition 870250092804, dated 10 / 10 / 2025, page 41 / 157 33 / 135
[0117] For modes -14 to -1 and 2 to 33, when the data arrangement order to form a 16x1 input vector is first row order, for modes 35 to 80, the input vector can be formed according to first column order.
[0118] Mode 34 can be considered neither a horizontal directionality mode nor a vertical directionality mode, but in this disclosure, it is classified as belonging to a horizontal directionality mode. That is, for modes -14 to -1 and 2 to 33, the input data arrangement method for the horizontal directionality mode, i.e., the order of the first row, is used, and for the vertical directionality mode symmetrical with respect to mode 34, the input data can be transposed and used.
[0119] For non-square blocks, the symmetry in square blocks (i.e., the symmetry between mode P and mode (68-P) in an NxN block (2<=P<=33) or the symmetry between mode Q and mode (66-Q) (-14<=Q<=-1)) cannot be used. Therefore, in addition to the symmetry based solely on the intraprediction mode, the symmetry between block shapes that are in a transposition relation to each other, i.e., the symmetry between a KxL block and an LxK block, can also be used. Specifically, there is a symmetry relation between a KxL block predicted by mode P and an LxK block predicted by mode (68-P). Alternatively, there is a symmetry relation between a KxL block predicted by mode Q and an LxK block predicted by mode (66-Q).
[0120] Since a KxL block with mode 2 and an LxK block with mode 66 can be viewed as symmetrical to each other, the same transform kernel can be applied to the KxL block and the LxK block. If a non-separable transform set for the intraprediction mode of the KxL block is mapped, to apply a non-separable transform to the LxK block, the non-separable transform set can be derived through a mapping table corresponding to the KxL block based on the (68-P) mode instead of the P mode applied to the LxK block. Alternatively, the set Petition 870250092804, dated 10 / 10 / 2025, page 42 / 157 34 / 135 of non-separable transform can be derived via a mapping table corresponding to the KxL block based on the (66-Q) mode instead of the Q mode applied to the LxK block.
[0121] For example, to apply a non-separable transform to an LxK block, the set of non-separable transforms can be selected based on mode 2 instead of mode 66. Furthermore, for a KxL block, the input data can be read in a predetermined order (e.g., row-first order or column-first order) to form a 1D vector, and then the corresponding non-separable transform can be applied. For an LxK block, the input data can be read in transposed order to form a 1D vector, and then the corresponding non-separable transform can be applied. That is, when the KxL block is read in row-first order, the LxK block can be read in column-first order. Conversely, when the KxL block is read in column-first order, the LxK block can be read in row-first order.
[0122] Furthermore, when mode 34 is applied to the KxL block, a set of non-separable transforms can be determined based on mode 34, and the input data can be read in a predetermined order to form a 1D vector and perform the corresponding non-separable transform. When mode 34 is applied to the LxK block, a set of non-separable transforms can be determined based on mode 34, but the input data can be read in a transposed order to form a 1D vector and perform the corresponding non-separable transform.
[0123] In this disclosure, a method for determining a set of non-separable transforms and a method for forming input data are described based on a KxL block. However, the non-separable transform can be performed based on an LxK block using the symmetry described above for a KxL block. Alternatively, a block with width greater than height can be Petition 870250092804, dated 10 / 10 / 2025, page 43 / 157 35 / 135 restricted to be used as a reference block. Alternatively, symmetry can be restricted to not be used in the case of non-square blocks. In this case, a non-square block can use a different number of non-separable transform sets and / or transform kernel candidates than a square block, and can select a non-separable transform set using a different mapping table than a square block.
[0124] An example of a mapping table for selecting a non-separable transform set is as follows: [Table 1] predModeIntra TrSetIdx predModeIntra < 0 4 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 66 1 67 <= predModeIntra <= 80 4
[0125] Table 1 shows an example of assigning a set of non-separable transforms to each intraprediction mode when there are five sets of non-separable transforms. The value of predModeIntra signifies the value of the intraprediction mode considering WAIP, and TrSetIdx is an index indicating a specific set of non-separable transforms. In Table 1, it can be confirmed that the same set of non-separable transforms is applied to modes located in symmetrical directions according to the intraprediction mode. Table 1 is only an example of the use of five sets of non-separable transforms and does not limit the total number of sets of non-separable transforms for Petition 870250092804, dated 10 / 10 / 2025, page 44 / 157 36 / 135 non-separable transforms.
[0126] Alternatively, as shown in Table 2, the non-separable transform may not be applied to WAIP for compression performance. [Table 2] predModeIntra TrSetIdx 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 66 1
[0127] Alternatively, as shown in Table 3, instead of setting up a separate non-separable transform set for WAIP, a non-separable transform set corresponding to an adjacent intraprediction mode can be shared. [Table 3] predModeIntra TrSetIdx predModeIntra < 0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 80 1
[0128] The non-separable transform set may include a plurality of transform kernel candidates, and one of the plurality of transform kernel candidates may be used selectively. For this purpose, an index signed via a bitstream may be used. Alternatively, one of several Petition 870250092804, dated 10 / 10 / 2025, p. 45 / 157 37 / 135 transform kernel candidates can be implicitly determined based on the context information of a current block. Here, context information can mean the size of a current block or whether a non-separable transform is applied to a neighboring block. Here, the size of the current block can be set as a width, a height, a maximum / minimum value of width and height, a sum of width and height, or a product of width and height.
[0129] Below, a method for determining the transform kernel for the inverse transform of the current block will be described in detail. Mode 1
[0130] As described above, the inverse transform can be divided into a separable transform and a non-separable transform. The separable transform means performing the transform in the horizontal and vertical directions, respectively, for a two-dimensional block, and the non-separable transform may mean performing a single transform for samples that constitute all or part of the two-dimensional block. When expressing a separable transform, it can be expressed as a pair of a horizontal transform and a vertical transform and, in the present disclosure, will be expressed as (horizontal transform, vertical transform).
[0131] A plurality of transform sets can be defined for the inverse transform of the current block. Each transform set can include one or more transform kernel candidates.
[0132] For example, one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8) or (DCT-8, DCT-8) can be applied as a separable transform, and the four transform kernel candidates above can be considered as a transform set. Furthermore, (DCT-2, DCT-2) can be considered as a transform set. A transform jump that does not apply a transform can also be considered as a transform set, and Petition 870250092804, dated 10 / 10 / 2025, page 46 / 157 38 / 135 (DCT-2, DCT-2), and the transform jump can be considered as a transform set. In the present disclosure, a transform kernel can refer to one transform (e.g., DCT-2, DST-7) or it can refer to two pairs of transforms (e.g., (DCT-2, DCT-2)).
[0133] As another example of a transform set, there may be the non-separable transform set described above. In the present disclosure, a non-separable transform applied as a primary transform may be denoted as a Non-Separable Primary Transform (NSPT). In the NSPT, a plurality of non-separable transform sets may be configured, and each non-separable transform set may include one or more transform kernels as transform kernel candidates. In the case of the NSPT, one of the plurality of non-separable transform sets is selected based on the intraprediction mode, and the plurality of non-separable transform sets for NSPT may be denoted as a list of NSPT sets. This is as described above, and a detailed description thereof will be omitted here.
[0134] A group of one or more sets of transforms available for a current block can be configured from a plurality of predefined sets of transforms. The group of one or more sets of transforms can be configured in a predetermined area unit to which the current block belongs and is hereinafter referred to as a collection. Here, the predetermined area unit can be at least one of a picture, a slice, a coding tree unit line (CTU line), or a coding tree unit (CTU).
[0135] For example, a transform set consisting of (DCT-2, DCT-2) is called Si, and a transform set consisting of (DST-7, DST7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) is called S2. Furthermore, the list of NSPT sets described above may include N sets of non-separable transforms, and the N sets of non-separable transforms are called Petition 870250092804, dated 10 / 10 / 2025, p. 47 / 157 39 / 135 S3,i, S3,2, ..., S3,n, respectively. Here, N can be 35, but it is not limited to that.
[0136] When S3,13 is selected as the non-separable transform set for NSPT based on the intraprediction mode of the current block, the transform kernel applicable to the current block can belong to one of S1, S2, or S3,13. In this case, the collection available for the current block can be denoted as {S1, S2, S3,13}.
[0137] As described above, since a collection according to the present disclosure is a group of one or more sets of transforms available for a current block, the collection can be configured differently based on the context of the current block. Here, the context can include at least one format, one size, or one intraprediction mode. If a total of K contexts are defined, K collections can be generated, and each collection can be denoted as Ci (i=1, 2, ..., N). For example, when the block sizes to which NSPT is applicable are 4x4, 8x8, 16x16, and 32x32, and one of a total of 35 sets of non-separable transforms is selected based on the intraprediction mode, a total of 4 x 35 = 140 contexts can be defined if different transform kernels are applied for each block size.
[0138] A collection can be configured based on the context of the current block, and in this case, a process of selecting one from a plurality of transform sets belonging to the collection and selecting one from a plurality of transform kernel candidates belonging to the selected transform set can be performed. Here, the selection of the transform set and the transform kernel candidate can be performed implicitly based on the context of the current block or can be performed based on an explicitly signed index. Alternatively, the process of selecting one from a plurality of transform sets belonging to the collection and the process of selecting one from a plurality of transform kernel candidates Petition 870250092804, dated 10 / 10 / 2025, page 48 / 157 40 / 135 belonging to the selected transform set can be played separately. For example, an index to select a transform set can be signaled first, and one of a plurality of transform sets belonging to the collection can be selected based on the index. Then, an index indicating one of a plurality of transform kernel candidates belonging to the transform set can be signaled, and one of the transform kernel candidates can be selected from the transform set based on the signaled index. The transform kernel of the current block can be determined based on the selected transform kernel candidate.Alternatively, the selection of a transform set from the collection can be performed implicitly based on the context of the current block, and the selection of a candidate transform kernel from the selected transform set can be performed based on a signed index. Alternatively, the selection of a transform set from the collection can be performed based on a signed index, and the selection of a candidate transform kernel from the selected transform set can be performed implicitly based on the context of the current block. Alternatively, the selection of a transform set from the collection can be performed implicitly based on the context of the current block, and the selection of a candidate transform kernel from the selected transform set can also be performed implicitly based on the context of the current block.Of course, when the number of transform sets belonging to the collection is 1, the index to select the transform set may not be signed. Similarly, when the number of transform kernel candidates belonging to the selected transform set is 1, the index to indicate the transform kernel candidate may not be signed. Alternatively, an index indicating one of all transform kernel candidates belonging to the current collection may be. Petition 870250092804, dated 10 / 10 / 2025, page 49 / 157 41 / 135 flagged. In this case, the process of selecting a transform set from the collection can be omitted. In this case, all transform sets belonging to the collection can be shuffled considering priorities. For example, in the case of assigning short-length binary codes to small-value indices, such as truncated unary codes, it may be advantageous to assign small-value indices to transform kernel candidates that are more advantageous for improving encoding performance. By shuffling all transform kernel candidates belonging to the collection according to priorities, different shuffles can be applied to each collection. Furthermore, instead of shuffling all transform kernel candidates belonging to the collection, only some of them can be selectively shuffled. Mode 2
[0139] The transform kernel for the inverse transform of the current block can be determined based on MTS (Multiple Transform Selection).
[0140] The MTS according to this disclosure may use at least one of the DST-7, DCT-8, DCT-5, DST-4, DST-1, or IDT (identity transform) as a transform kernel. In addition, the MTS according to this disclosure may additionally include a DCT-2 transform kernel.
[0141] In the present disclosure, a plurality of MTS sets for MTS can be defined. Based on the size and / or an intraprediction mode of a current block, one of the plurality of MTS sets can be determined. For example, when determining an MTS set, 16 transform block sizes can be considered and, for a directional mode, the shape of the transform block and the symmetry between the intraprediction modes can be considered. For the WAIP (Wide Angle Intra Prediction) mode (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), an MTS set corresponding to mode 2 can be applied for modes -1 to -14 (or -15), and an MTS set Petition 870250092804, dated 10 / 10 / 2025, p. 50 / 157 42 / 135 corresponding to mode 66 can be applied to modes 67 to 80 (or 81). A separate set of MTS can be assigned to the MIP (Matrix-based Intraprediction) mode.
[0142] For example, a set of MTS according to the transform block size and intraprediction mode can be assigned / defined as shown in Table 4 below. [Table 4] Block Size Intraprediction Mode Width Height [0, 1] [2, 12] [13, 23] [24, 34] MIP 4 4 0 1 2 3 4 4 8 5 6 7 8 9 4 16 10 11 12 13 14 4 32 15 16 17 18 19 8 4 20 21 22 23 24 8 8 25 26 27 28 29 8 16 30 31 32 33 34 8 32 35 36 37 38 39 16 4 40 41 42 43 44 16 8 45 46 47 48 49 16 16 50 51 52 53 54 16 32 55 56 57 58 59 32 4 60 61 62 63 64 32 8 65 66 67 68 69 32 16 70 71 72 73 74 32 32 75 76 77 78 79
[0143] Table 4 shows the allocation of MTS sets according to 16 transform block sizes and intraprediction modes. The number of Petition 870250092804, dated 10 / 10 / 2025, page 51 / 157 43 / 135 sets of predefined MTS is 80, and the index indicates one of the 80 sets of MTS can have a value from 0 to 79, as shown in Table 4. [Table 5] MTS set index Transform kernel candidate index 0 1 2 3 4 5 0 18 24 17 23 8 12 1 18 3 7 22 0 16 2 18 2 17 22 3 23 3 18 3 15 17 12 23 4 18 12 3 19 10 13 5 18 12 19 23 13 24 6 18 12 17 2 3 23 7 18 2 17 22 12 23 8 18 2 11 17 22 23 9 18 12 19 23 3 10 10 16 12 13 24 7 8 11 16 2 11 23 12 18 12 13 17 2 22 12 18 13 17 11 2 21 12 18 14 16 13 19 22 3 10 15 18 12 13 7 14 22 16 16 12 11 1 18 22 17 17 13 3 22 12 18 18 6 12 1 22 13 17 19 16 12 13 15 2 23 20 18 24 23 19 12 17 21 18 24 2 17 0 23 Petition 870250092804, dated 10 / 10 / 2025, page 52 / 157 44 / 135 22 17 3 4 22 2 13 23 18 12 19 23 3 15 24 18 12 19 23 3 10 25 6 12 18 24 13 19 26 6 12 2 21 13 18 27 17 11 1 22 2 18 28 16 17 3 11 12 23 29 8 12 19 23 11 24 30 16 13 7 23 12 19 31 6 12 1 11 18 22 32 17 11 1 21 12 18 33 6 11 17 21 12 18 34 8 11 14 17 12 22 35 6 12 11 21 14 16 36 6 12 11 1 17 21 37 6 12 11 2 17 21 38 6 11 21 1 12 17 39 16 12 11 7 1 5 40 8 12 19 24 11 17 41 18 13 1 22 2 24 42 6 2 17 21 19 22 43 16 12 11 19 8 15 44 8 12 17 24 13 15 45 6 12 19 21 17 18 46 6 12 13 21 2 18 47 16 2 17 21 1 11 48 6 17 19 23 12 16 Petition 870250092804, dated 10 / 10 / 2025, p. 53 / 157 45 / 135 49 6 12 14 17 8 22 50 6 7 11 21 9 12 51 16 12 11 1 7 21 52 6 12 11 1 17 21 53 6 12 11 21 1 16 54 8 7 9 11 12 21 55 6 12 7 11 14 21 56 6 12 7 11 1 21 57 16 12 11 1 2 21 58 6 11 17 21 1 12 59 6 12 7 11 9 21 60 18 12 14 21 6 21 61 16 11 1 22 2 17 62 16 11 1 22 2 17 63 16 13 15 7 14 19 64 8 12 1 19 16 23 65 6 12 7 9 13 21 66 6 12 13 2 7 18 67 16 12 1 21 11 17 68 16 11 7 19 12 15 69 8 12 7 11 14 21 70 6 12 7 11 8 9 71 6 12 7 11 2 21 72 6 12 1 11 21 22 73 6 7 11 16 9 12 74 6 12 7 11 9 21 75 6 12 7 11 13 17 Petition 870250092804, dated 10 / 10 / 2025, p. 54 / 157 46 / 135 76 6 12 11 21 2 7 77 6 12 1 11 2 7 78 6 12 7 11 16 21 79 6 12 7 11 9 16
[0144] Table 5 shows the transform kernel candidates included in each set of MTS described in Table 4. Each set of MTS can be composed of six transform kernel candidates. The transform kernel candidate index has a value from 0 to 5 and can indicate one of the six transform kernel candidates. Here, each transform kernel candidate can be a combination of a horizontal transform kernel and a vertical transform kernel for a separable transform, and 25 transform kernel candidates with indices from 0 to 24 can be defined. [Table 6] Kernel Index Combination when the intraprediction mode value is less than 35 when the intraprediction mode value is greater than or equal to 35 0 (DCT-8, DCT-8) (DCT-8, DCT-8) 1 (DST-7, DCT-8) (DCT-8, DST-7) 2 (DCT-5, DCT-8) (DCT-8, DCT-5) 3 (DST-4, DCT-8) (DCT-8, DST-4) 4 (DST-1, DCT-8) (DCT-8, DST-1) 5 (DCT-8, DST-7) (DST-7, DCT-8) 6 (DST-7, DST-7) (DST-7, DST-7) 7 (DCT-5, DST-7) (DST-7, DCT-5) 8 (DST-4, DST-7) (DST-7, DST-4) 9 (DST-1, DST-7) (DST-7, DST-1) 10 (DCT-8, DCT-5) (DCT-5, DCT-8) Petition 870250092804, dated 10 / 10 / 2025, page 55 / 157 47 / 135 11 (DST-7, DCT-5) (DCT-5, DST-7) 12 (DCT-5, DCT-5) (DCT-5, DCT-5) 13 (DST-4, DCT-5) (DCT-5, DST-4) 14 (DST-1, DCT-5) (DCT-5, DST-1) 15 (DCT-8, DST-4) (DST-4, DCT-8) 16 (DST-7, DST-4) (DST-4, DST-7) 17 (DCT-5, DST-4) (DST-4, DCT-5) 18 (DST-4, DST-4) (DST-4, DST-4) 19 (DST-1, DST-4) (DST-4, DST-1) 20 (DCT-8, DST-1) (DST-1, DCT-8) 21 (DST-7, DST-1) (DST-1, DST-7) 22 (DCT-5, DST-1) (DST-1, DCT-5) 23 (DST-4, DST-1) (DST-1, DST-4) 24 (DST-1, DST-1) (DST-1, DST-1)
[0145] Table 6 is an example of the 25 transform kernel candidates described in Table 5. Specifically, the horizontal transform and vertical transform of the transform kernel candidate are expressed as (horizontal transform, vertical transform). For each transform kernel candidate index, the horizontal / vertical transform when the intraprediction mode is less than 35 can be the opposite of the horizontal / vertical transform when the intraprediction mode is greater than or equal to 35. When the intraprediction mode value is greater than or equal to 35, a mode symmetrical with respect to mode 34 can be derived, and a set of MTS can be selected from Table 4 based on the mode. In addition, the symmetry of the block shape can be further considered.When the original transform block has a size LxA, the original transform block can be considered as having a size AxL through its symmetrization, and a set of MTS can be selected from Table 4. Here, the... Petition 870250092804, dated 10 / 10 / 2025, page 56 / 157 48 / 135, the intraprediction mode value, can be the modified intraprediction mode value. That is, as a mode value for WAIP, from -14 (or -15) to -1, it is modified to mode 2, from 67 to 80 (or 81), it is modified to mode 66, and for the remaining modes, the original intraprediction mode value can be adjusted as the modified intraprediction mode value. In this case, since the extended modes for WAIP are also configured symmetrically with respect to mode 34, the symmetry with respect to mode 34 can be used for all directional modes except for Planar mode and DC mode.
[0146] For example, when a 16x32 block is predicted based on mode 54, mode 14 (=68-54) can be derived as a symmetric mode to mode 54, and the block size can be considered as 32x16. In this case, an MTS set with index 72 can be selected, as defined in Table 4.
[0147] When MIP mode is applied, the MTS set assigned to MIP mode can be selected based on the current block size, without considering the symmetry of the block shape. Alternatively, when MIP mode is applied, the MTS set assigned to MIP mode can be selected based on the symmetrical block size, considering the symmetry of the block shape. For example, when MIP mode is applied to an 8x16 block, the 8x16 block can be considered a symmetrical 16x8 block, and an MTS set with an index of 49 can be selected as defined in Table 4. Alternatively, when MIP mode is applied, the intraprediction mode can be considered the Planar mode. In this case, the MTS set assigned to MIP mode can be selected based on the current block size, without considering the symmetry of the block shape.Alternatively, the set of MTS assigned to MIP mode can be selected based on the size of the symmetrical block, considering the symmetry of the block shape.
[0148] For MIP mode, a flag can be used to indicate whether MIP mode is applied in transposition mode. When MIP mode is applied to the current block. Petition 870250092804, dated 10 / 10 / 2025, page 57 / 157 49 / 135 of MxN and the flag indicates that the transposition mode is applied, the intraprediction mode can be considered as Planar mode, and the current MxN block can be considered as an NxM block. That is, from Table 4, one can select a set of MTS corresponding to the size of the NxM block and the Planar mode. As described in Table 6, when the value of the intraprediction mode is greater than or equal to 35, the horizontal and vertical transforms are swapped, but since the intraprediction mode of the current block is considered the Planar mode, the horizontal and vertical transforms of the candidate transform kernel cannot be swapped. Alternatively, when the MIP mode is applied to the current MxN block and the flag indicates that the transposition mode is applied, the intraprediction mode may not be considered as Planar mode, and the current MxN block can be considered as an NxM block.In other words, from Table 4, one can select a set of MTS corresponding to the NxM block size and MIP mode.
[0149] In Table 5, a candidate transform kernel selected by a candidate transform kernel index can be set as a transform kernel for the current block. Alternatively, based on the size of the current block, at least one horizontal transform or one vertical transform of the selected candidate transform kernel can be changed to another transform kernel. For example, when the candidate transform kernel index is 3 and both the width and height of the current block are less than or equal to 16, at least one horizontal transform or one vertical transform of the candidate transform kernel corresponding to the candidate transform kernel index of 3 can be changed to another transform kernel. In this case, the horizontal transform and the vertical transform can be changed independently of each other.When the difference (or the absolute value of the difference) between the intraprediction mode value of the current block and the mode value. Petition 870250092804, dated 10 / 10 / 2025, page 58 / 157 If the horizontal difference (or the absolute value of the difference) between the intraprediction mode value of the current block and the vertical mode value is less than or equal to a predetermined threshold, the vertical transform of the selected candidate transform kernel can be changed to an Identity Transform (IDT). Here, the threshold can be determined based on the width and height of the current block, as shown in Table 7 below. [Table 7] Block size Limit width height 4 4 8 4 8 6 4 16 4 8 4 8 8 8 8 8 16 6 16 4 4 16 8 2 16 16 -1
[0150] Table 7 is for changing the horizontal and / or vertical transform of a candidate transform kernel selected by a candidate transform kernel index to another transform kernel and sets limits according to the size of a transform block.
[0151] Six transform kernel candidates that make up an MTS set can be distinguished by transform kernel candidate indices from 0 to 5, as defined in Table 5. The transform kernel candidate index Petition 870250092804, dated 10 / 10 / 2025, page 59 / 157 51 / 135 can be signaled via a bitstream. A flag indicating whether the MTS set is available / applied (MTS-enabled flag or MTS flag) can be signaled, and a candidate index to the transform kernel can be signaled when the flag indicates that the MTS set is available / applied. The MTS flag can be composed of a bin, and one or more contexts for CABAC-based entropy encoding (hereinafter referred to as CABAC contexts) can be assigned to the bin. For example, different CABAC contexts can be assigned to non-MIP mode and MIP mode, respectively.
[0152] Based on the current block context described above, the number of transform kernel candidates available for the current block can be defined differently. For example, as the current block context, the sum of the absolute values of all or some of the transform coefficients in the current block can be considered. The sum of the absolute values of the transform coefficients is referred to as AbsSum. When AbsSum is less than or equal to T1, only one transform kernel candidate corresponding to a transform kernel candidate index of 0 can be available. When AbsSum is greater than T1 and less than or equal to T2, four transform kernel candidates corresponding to transform kernel candidate indices from 0 to 3 can be available. When AbsSum is greater than T2, six transform kernel candidates corresponding to transform kernel candidate indices from 0 to 5 can be available.Here, T1 could be 6 and T2 could be 32, but that's just an example.
[0153] When AbsSum is less than or equal to T1, since the number of transform kernel candidates available for the current block is 1, the transform kernel candidate corresponding to the transform kernel candidate index of 0 can be set as the transform kernel of the current block without signaling the transform kernel candidate index. When AbsSum is Petition 870250092804, dated 10 / 10 / 2025, page 60 / 157 52 / 135 greater than T1 and less than or equal to T2, as four transform kernel candidates are available, one of the four transform kernel candidates can be selected based on the transform kernel candidate index with two bins. That is, the transform kernel candidate indices from 0 to 3 can be signaled as 00, 01, 10, and 11, respectively. For the two bins, the MSB (Most Significant Bit) can be signaled first and the LSB (Least Significant Bit) can be signaled later. Different CABAC contexts can be assigned to each bin. For example, a CABAC context different from the CABAC context assigned to the MTS flag can be assigned to each bin for the two bins. Alternatively, bypass coding can be applied without assigning a CABAC context to the two bins.When AbsSum is greater than T2, the candidate index of the transform kernel has a value from 0 to 5; therefore, the candidate index of the transform kernel cannot be expressed with only two bins. In this case, the candidate index of the transform kernel can be expressed by assigning two or more bins, as in truncated binary encoding. For each bin assigned by the truncated binary encoding method, a CABAC context can be assigned, or bypass encoding can be applied without assigning a CABAC context. Alternatively, a CABAC context can be assigned to some of the plurality of bins (e.g., the first bin or the first and second bins), and bypass encoding can be applied to the remaining bins. Mode 3
[0154] The transform kernel of the current block can be determined based on a transform set including one or more transform kernel candidates. The transform kernel of the current block can be derived as one of one or more transform kernel candidates belonging to the set of Petition 870250092804, dated 10 / 10 / 2025, p. 61 / 157 53 / 135 transformed.
[0155] The process of determining the transform kernel of the current block may include at least one of the following: 1) the process of determining the transform set of the current block or 2) the process of selecting a candidate transform kernel from the transform set of the current block. The process of determining the transform set may be the process of selecting one from a plurality of transform sets that are identically predefined in the encoding and decoding apparatus. Alternatively, the process of determining the transform set may be a process of configuring one or more transform sets available to the current block from a plurality of transform sets that are identically predefined in the encoding and decoding apparatus, and selecting one of the configured transform sets.Alternatively, the transform set determination process can be a transform set configuration process based on a transform kernel candidate available for the current block from among a plurality of transform kernel candidates that are identically predefined in the encoding and decoding apparatus.
[0156] When the current block transform set includes a plurality of transform kernel candidates, a selection process of one from the plurality of transform kernel candidates for the current block can be performed. However, when the current block transform set includes one transform kernel candidate (that is, when the number of transform kernel candidates available for the current block is 1), the current block transform kernel can be set to the corresponding transform kernel candidate.
[0157] The transform set according to the present disclosure may Petition 870250092804, dated 10 / 10 / 2025, p. 62 / 157 54 / 135 means the (non-separable) transform set in Mode 1 described above, or it may mean the MTS set in Mode 2. Alternatively, the transform set may be defined separately from the (non-separable) transform set in Mode 1 or the MTS set in Mode 2. In that case, the transform set may include one or more specific transform kernels as transform kernel candidates. A specific transform kernel may be defined as a pair of a transform kernel for horizontal transform and a transform kernel for vertical transform, or it may be defined as a transform kernel that is applied equally to horizontal and vertical transforms.
[0158] In the embodiment of the present disclosure, the process of applying NSPT, a non-separable transform applied as a primary transform, is described in detail. NSPT can be applied to all or part of a transform block. Based on the forward NSPT, a residual sample existing in a region where the NSPT is applied can be entered as a 1D NSPT vector. In other words, residual samples existing in all or part of a single transform block (referred to as Region of Interest (ROI) in the present disclosure) can be collected as a 1D vector and configured as an input. Then, when the forward NSPT is applied, a primary transform coefficient can be obtained. Conversely, when the reverse NSPT is applied to a primary transform coefficient, a 1D vector output can be obtained.A residual sample for a ROI can be obtained by arranging each element value by setting up a corresponding output vector at a given position within a 2D transform block.
[0159] For a non-separable transform kernel for NSPT, a matrix dimension can be determined according to the size of an ROI. In the present disclosure, a transform kernel can be referred to as a type of Petition 870250092804, dated 10 / 10 / 2025, page 63 / 157 55 / 135 transform or transform matrix, and a non-separable transform kernel for NSPT can be referred to as an NSPT kernel. For example, when a current block is an MxN transform block, a ROI is a region for the entire MxN transform block, and square NSPT is applied, the dimension of a corresponding transform matrix can be MN x MN. For example, when an ROI is a region for the entire 8x8 transform block, the dimension of an NSPT kernel can be 64 x 64.
[0160] According to an embodiment of the present disclosure, when NSPT is applied to a residual generated by intraprediction, an NSPT kernel can be adaptively determined according to an intraprediction mode. Since a statistical characteristic of a residual block can vary depending on an intraprediction mode, compression efficiency can be improved by adaptively determining an NSPT kernel according to an intraprediction mode.
[0161] It can be configured to share an NSPT kernel applied to at least one intraprediction mode. As described above, a non-separable transform set can be determined based on the intraprediction mode of a current block and a mapping table. A mapping table can define a mapping relationship between predefined intraprediction modes and non-separable transform sets. The predefined intraprediction modes can include two non-directional modes and 65 directional modes.
[0162] As a modality, intraprediction modes can be grouped into an intraprediction mode group. An NSPT kernel can be allocated, or a plurality of NSPT kernels can be allocated, to an intraprediction mode group. In other words, a non-separable transform set (NSPT set) including at least one NSPT kernel can be allocated to an intraprediction mode group. A non-separable transform set can Petition 870250092804, dated 10 / 10 / 2025, p. 64 / 157 56 / 135 can be mapped to an intraprediction mode, and one of the N NSPT kernels included in a non-separable transform set can be selected.
[0163] As an example, an intraprediction group can include adjacent prediction modes (e.g., modes 17, 18, and 19). Furthermore, an intraprediction group can include modes with symmetry. For example, directional modes can be symmetrical around a diagonal mode (i.e., intraprediction mode 34) in Figure 5 described above. In this case, two symmetrical modes can form a group (or pair). For example, mode 18 and mode 50 can be included in the same group because they are symmetrical around mode 34. However, for modes with symmetry, the process of transposing a 2D input block and then setting up a one-dimensional input vector can be added before applying a direct NSPT kernel. For example, when an intraprediction mode is less than or equal to 34, a one-dimensional input vector can be derived from a corresponding input block in the first order of the row without transposing a 2D input block.When an intraprediction mode is greater than 34, a one-dimensional input vector can be set up by first transposing a 2D input block and then reading a corresponding input block in the first order of the row, or by leaving a 2D input block as is and reading a corresponding input block in the first order of the column.
[0164] Table 8 below illustrates a mapping table for allocating a set of NSPTs according to an intraprediction mode. Referring to Table 8, a total of 35 sets of NSPTs from 0 to 34 can be defined. An NSPT set allocated to the nearest general directional mode can be allocated to an extended WAIP mode (i.e., modes -14 to -1 and modes 67 to 80 in Figure 5). In other words, NSPT set 2 can be allocated to an extended WAIP mode. [Table 8] Petition 870250092804, dated 10 / 10 / 2025, p. 65 / 157 57 / 135 Meow irrtrayw InlCkyiNSTF Intraprp Mode; [jj^onj-NSTj Fear Inüipra ΜΙ ϋΙΠΙ Ü i* I* EPI JS J1 >1 tf «4 H 111· ]| II 419 111! H B 13 Ai H □ ΒΒΒ QDOB EQ EJI3 ΟΪ9ΕΠ£3 El ΕΙΙΕ3 □ □ ΒΕ11Β ΕΙΕΙΒ El Β Ε1Ε1ΠΪ ϋΐπ ιίΠ*11Φ£ί1£^21ί!1ΜΤ^
[0165] An NSPT set may include at least one NSPT kernel (or kernel candidate). In other words, an NSPT set may include N NSPT kernel candidates. As an example, N may be set to a value equal to or greater than 1, such as 1, 2, 3, 4, etc. A kernel applied to a current block among at least one NSPT kernel included in an NSPT set may be signaled using an index. In the present disclosure, a corresponding index may be referred to as an NSPT index. As an example, an NSPT index may have a value of 0, 1, 2, ..., N - 1.
[0166] Furthermore, as a modality, when the number of NSPT kernel candidates is 1, an NSPT index value can be set to 0. In this case, an NSPT index can be inferred without being separately flagged. Additionally, a flag indicating whether NSPT should be applied can be flagged separately from an NSPT index. In the present disclosure, a corresponding flag can be referred to as the NSPT flag.
[0167] When an NSPT flag value is 1, NSPT can be applied. When an NSPT flag value is 0, NSPT cannot be applied. When an NSPT flag is not flagged, an NSPT flag value can be inferred as 0. For example, when an NSPT flag value is 1, an NSPT index can be applied. One of the N kernel candidates included in an NSPT set selected by an intraprediction mode can be specified based on a flagged NSPT index.
[0168] In one embodiment, the entropy encoding method of an NSPT index can be defined in several ways, considering the number (N) of NSPT kernels included in an NSPT set. For example, as a method for Petition 870250092804, dated 10 / 10 / 2025, page 66 / 157 58 / 135 mapping a value from 0 to N-1 to a compartment sequence (i.e., a binarization method), unary truncated binarization, truncated binarization, and fixed-length binarization methods can be used.
[0169] For example, when the value of N, the number of kernel candidates that configure an NSPT set, is 2, one of the two candidates can be specified with a bin. For example, 0 can indicate the first candidate and 1 can indicate the second candidate. Furthermore, when the value of N is 3 and truncated unary binarization is applied, a candidate can be specified with two bins. For example, the first, second, and third candidates can be binarized to 0, 10, and 11, respectively, and signed. As a modality, the binarized bins can be encoded using context encoding or bypass encoding.
[0170] In this disclosure, a reduced primary transform (RPT) method is described using a reduced-dimensional transform kernel as a primary transform. As described above, when the direct NSPT is applied, samples belonging to a 2D residual block can be arranged (or rearranged) into a 1D vector according to the order of the first row (or the order of the first column). Subsequently, a transform matrix for NSPT can be multiplied by a arranged vector. When a corresponding 2D residual block is an M x N block (M is a horizontal length and N is a vertical length), the length of a rearranged 1D vector can be M*N. In other words, a corresponding 2D residual block can also be represented as a column vector with a dimension of M*N x 1. In this disclosure, M*N can be expressed as MN for convenience. In this case, the dimension of a corresponding transform matrix can be MN x MN.In summary, the direct NSPT can operate in such a way that a transform coefficient vector MN x 1 is obtained by multiplying the left side of an MN x 1 vector by an MN x MN transform matrix. Petition 870250092804, dated 10 / 10 / 2025, page 67 / 157 59 / 135 corresponding.
[0171] When RPT is applied, the transform coefficients r can be obtained by multiplying an r x MN matrix, instead of multiplying an MN x MN matrix, as a direct NSPT transform matrix described above. Here, r represents the number of rows of a transform matrix, and MN represents the number of columns of a transform matrix. According to an embodiment of the present disclosure, the value of r can be set to less than or equal to MN. In other words, the existing direct NSPT transform matrix includes MN rows, and each row is a 1 x MN row vector and a transform basis vector of a corresponding NSPT transform matrix. A corresponding transform coefficient can be obtained by multiplying each transform basis vector by a sample column vector MN x 1.
[0172] Since the existing direct NSPT transform matrix is composed of MN row vectors, the transform coefficients MN (i.e., MN x 1 transform coefficient column vectors) can be obtained by applying direct NSPT. Meanwhile, for direct RPT, a transform matrix can be composed of transform basis vectors r instead of transform basis vectors MN. Consequently, when direct RPT is applied, transform coefficients r, instead of MN, (i.e., transform coefficient column vectors rx 1) can be obtained.
[0173] An RPT kernel can be configured by selecting r transform basis vectors that are part of the transform basis vectors that configure a direct MN x MN NSPT kernel. In the present disclosure, a transform kernel may be referred to as a transform type or transform matrix, and a non-separable transform kernel for NSPT may be referred to as an RPT kernel. In other words, when r 1 x MN row vectors are selected from a direct MN x MN NSPT kernel, it may be advantageous Petition 870250092804, dated 10 / 10 / 2025, page 68 / 157 60 / 135 select the most important transform basis vectors from the perspective of coding performance. Specifically, in terms of energy compression through transform, more energy can be concentrated in transform coefficients that appear first by multiplying a direct NSPT transform matrix. In other words, a transform basis vector positioned on the upper side of a direct NSPT transform matrix can generate a transform coefficient with higher energy. Considering this, a direct RPT kernel rx MN can be configured (or derived) by taking r from the upper side of a direct NSPT kernel.
[0174] The RPT according to the present disclosure takes only a part (i.e., r) of the transform coefficients obtained by applying the existing NSPT and, consequently, the energy of the original signal may be partially lost. In other words, distortion between an original signal and a natural signal may occur through a corresponding process. However, since only r, instead of MN, transform coefficients are generated by applying RPT, the number of bits required to encode the corresponding transform coefficients can be reduced. Therefore, for a signal where a large amount of energy is concentrated in a small number of transform coefficients (e.g., a residual image signal), the gains obtained by reducing signaling bits can be significantly large, thus improving encoding performance.
[0175] The reverse NSPT is a transform matrix and can be a transpose matrix of a direct NSPT kernel described above. In this case, the input data can be a transform coefficient signal instead of a sample signal, such as a residual signal. Specifically, when a direct NSPT transform matrix is G and a sample signal rearranged into a 1D vector is x, a transform coefficient vector is obtained by multiplying a matrix of Petition 870250092804, dated 10 / 10 / 2025, page 69 / 157 The 61 / 135 transformation corresponding to the left-hand side can be expressed as in Equation 4 below. [Equation 4] y = Gx
[0176] Referring to Equation 4, x and y can be an MN x 1 column vector. G can take the form of an MN x MN matrix. A reverse NSPT process can be expressed as in Equation 5 below using the same variable. [Equation 5] x = GTy
[0177] In Equation 5, GT signifies a transpose matrix of G. A forward RPT operation and a reverse RPT operation according to the present disclosure can also be expressed by the two equations. However, when RPT is applied, y is an r x 1 column vector instead of an MN x 1 column vector, and G is an r x MN matrix instead of an MM x MN matrix. In other words, even when RPT is applied instead of NSPT, the dimension of a sample signal (e.g., an image residual signal) is not altered, which may mean that the original number of sample signals (i.e., MN sample signals) can be reconstructed with only r transform coefficients via reverse RPT. In other words, the original MN sample signals can be reconstructed by encoding only r transform coefficients that are smaller than MN, which may improve encoding performance.
[0178] In one embodiment of the present disclosure, an RPT structure is proposed that defines the value of r considering the statistical characteristics of a residual block and derives a residual block of the size of the existing transform block from a reduced-size residual block determined according to the defined value of r. If another additional transform (i.e., a secondary transform) is applied to predict the statistical distribution of a transform coefficient Petition 870250092804, dated 10 / 10 / 2025, page 70 / 157 62 / 135 primary, a quantization process is applied to a primary transform coefficient, so that non-zero quantized coefficients can be concentrated in a relatively low frequency domain. Consequently, a reduced secondary transform for the statistical distribution of a primary transform coefficient can define the statistical characteristics of a primary transform coefficient in a relatively simple way by fitting the value of r to a given low frequency domain. However, the RPT according to the present disclosure has a fundamental difference from a reduced secondary transform as a technology for defining the value of r when considering the statistical characteristics of samples within a residual block that has a very different characteristic from the distribution of a primary transform coefficient.The following describes various methods for determining an RPT kernel that is a reduced-dimensional transform matrix. In other words, a method for determining or defining the value of r in RPT is described below.
[0179] In one embodiment of the present disclosure, the value of r in RPT can be determined by considering the worst-case complexity allowed by a transform system. As an embodiment, the worst-case complexity can be calculated based on the number of multiplications per sample. Multiplications MN * r are required to apply RPT in both the forward and reverse directions based on an M x N block. Since a 2D block is composed of a total of MN samples, the number of multiplications per sample can be calculated as (MN * r) / MN = r. Consequently, the value of r can be set to be less than or equal to the maximum number of multiplications per sample allowed. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16x16 block, the value of r can be determined to be less than or equal to 16. In other words, a forward RPT kernel can be set to 16 x 256.
[0180] In another modality, the use of memory can be considered as Petition 870250092804, dated 10 / 10 / 2025, p. 71 / 157 63 / 135 is a measure of the worst possible complexity. As an example, a memory size allowed per kernel can be adjusted. For example, when p bytes are required per kernel coefficient (in the present disclosure, each element that configures a transform kernel is referred to as a kernel coefficient) and memory usage is defined as less than or equal to q bytes per kernel, the value of r can be adjusted as less than or equal to q / (MN * p). For example, when p is 1 byte for an RPT advance kernel for a 16x16 block and memory usage is defined as less than or equal to 8 KB per kernel (q = 8 KB = 213 bytes), the value of r can be adjusted as less than or equal to 32.
[0181] Furthermore, as another example, memory usage and / or the number of multiplications per sample can be considered as a measure of worst-case complexity. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16x16 block and memory usage is set to less than or equal to 8 KB per kernel (a kernel coefficient is expressed as 1 byte), the value of r can be set to less than or equal to 16.
[0182] Furthermore, in one embodiment, the value of r that configures an RPT kernel can be determined by specific information. In other words, the value of r that configures an RPT kernel can be determined based on a predefined encoding parameter. For example, the value of r can be determined based on the size of a block. In other words, an RPT kernel can be variably determined based on the size of a block. Here, a block can be at least one encoding block, one transform block, and one prediction block. Additionally, for example, the value of r can be determined based on prediction information. Here, prediction information can include information about interprediction / intraprediction, information about the intraprediction mode, etc. Furthermore, for example, the value of r can be determined based on signaled information (the value of a syntax element). For example, the value of r can be Petition 870250092804, dated 10 / 10 / 2025, page 72 / 157 64 / 135 determined variably according to a quantization parameter value. Furthermore, in terms of complexity improvement, a predefined fixed value such as the r-value can be used, and the predefined fixed value can be determined based on signaled information.
[0183] When a sample signal is multiplied by an RPT kernel r x MN, r transform coefficients can be obtained. The resulting transform coefficients r can be arranged according to the predefined scan order of the transform coefficients (e.g., forward / reverse zigzag scan order, forward / reverse horizontal scan order, forward / reverse vertical scan order, forward / reverse diagonal scan order, scan order specified based on an intraprediction mode, etc.). When the transform coefficients obtained by applying forward RPT are arranged according to this scan order (e.g., the scan order in the unit of a coefficient group (CG) can also be applied), if the value of r is less than MN, the interior of an M x N block may not be completely filled with r transform coefficients, so an empty space may occur.As a manifestation of the present disclosure, an empty space described above can be predicted in the following way, considering the characteristics of a residual signal.
[0184] - The value of an empty space can be filled using the value of an available neighboring pixel.
[0185] - The value of an empty space can be filled based on the value of an available neighboring pixel and an intraprediction mode. For example, the value of an empty space can be predicted by performing intraprediction based on the value of an available neighboring pixel and an intraprediction mode.
[0186] - The value of an empty space can be filled using a predefined fixed value (for example, 0).
[0187] - The value of an empty space can be filled starting from one pixel Petition 870250092804, dated 10 / 10 / 2025, page 73 / 157 65 / 135 neighbor available using a predetermined intraprediction mode (e.g., a planar mode).
[0188] In the present disclosure, filling an empty space with 0 between the examples described above can be referred to as the zeroing process. When an empty space is filled with 0, the following arrangement can be applied. When a non-zero transform coefficient is detected (or parsed) in a corresponding empty space portion during the parsing of a transform coefficient on the decoding device side, it can be considered (or inferred) that RPT is not applied. In other words, when a non-zero transform coefficient exists in a predefined region representing the corresponding empty space, it can be considered that RPT is not applied. In this case, signaling (or parsing) a flag indicating whether to apply RPT and / or an index designating one of a plurality of RPT kernel candidates may not be performed.For example, when a non-zero transform coefficient exists in a predefined region representing the corresponding empty space, a predefined variable value can be updated, and it can be inferred that the RPT is not applied based on an updated variable value.
[0189] In one embodiment of the present disclosure, the application of RPT can be determined based on the size and / or shape of a block. Furthermore, an RPT kernel can be variably determined according to the size and / or shape of a block. Since the value of r can be different according to the size and / or shape of a block (i.e., for each M x N block), an empty space can be different according to the size and / or shape of a block. Consequently, a region to check if a non-zero transform coefficient is detected can be defined differently for each size and / or shape of a block. In other words, a zeroing region can be variably determined.
[0190] For example, when a 16x64 matrix is applied as a matrix Petition 870250092804, dated 10 / 10 / 2025, page 74 / 157 66 / 135 For a direct RPT for an 8x8 block, the value of r can be 16. In this case, when a CG is a 4x4 subblock, only the upper-left 4x4 block can be filled with non-zero RPT transform coefficients, and the three remaining 4x4 subblocks (i.e., upper-right, lower-left, and lower-right subblocks) that are empty space can be filled with the value 0. In this case, when a non-zero transform coefficient is detected in the three corresponding remaining 4x4 subblock regions during a decoding process, it can be considered that the RPT is not applied. And, as described above, a flag indicating whether to apply RPT or an index designating one of a plurality of RPT kernel candidates may not be flagged.
[0191] Furthermore, as an example, when a 32x128 matrix is applied as a direct RPT matrix to a 16x8 block (i.e., the value of r is 32) and a CG is a 4x4 subblock, only two CGs in the scan order can be filled with non-zero RPT transform coefficients. For example, an upper-left 4x4 subblock and a 4x4 subblock adjacent to the bottom of an upper-left subblock can be filled with corresponding RPT transform coefficients. A region filled with 0 as an empty space can be determined as the remaining region excluding two corresponding 4x4 subblocks. An RPT kernel can be variably determined according to the size and / or shape of a block, and as described, an empty space can be determined differently for an 8x8 block and a 16x8 block.
[0192] As a modality, when the value of r is a multiple of a CG size and a transform coefficient is scanned in the unit of a CG, if a non-zero transform coefficient is detected in CGs belonging to an empty space, a flag and / or an index related to the RPT may not be flagged. In other words, the transform coefficients within a CG may be scanned in the order designated for each CG, and the transform coefficients Petition 870250092804, dated 10 / 10 / 2025, page 75 / 157 67 / 135 within a CG can be scanned in the same way after moving to the next CG in the scanning order for a CG unit. In existing image compression technology, since a flag for the existence of a non-zero transform coefficient in a corresponding CG is signaled first for each CG, the application of RPT can be determined only with matching information, which can reduce signaling overhead and related implementation complexity.
[0193] As described above, when RPT is applied, if a non-zero transform coefficient is detected in an empty space region filled with 0, RPT may not be applied. In this case, the RPT-related information flag may be omitted. However, since it is not possible to determine whether to apply RPT when a non-zero transform coefficient is not detected in a corresponding empty space region, a flag representing whether to apply RPT may be analyzed after analyzing (or flagging) a related transform coefficient to finally determine whether to apply RPT.
[0194] As an embodiment, a direct secondary transform can be additionally applied to the transform coefficients generated by applying RPT. Alternatively, a direct secondary transform can be additionally applied to a region where the corresponding generated transform coefficients are positioned in an M x N block. In the present disclosure, a corresponding region or a part of a corresponding region can be referred to as an ROI from the perspective of a direct secondary transform. For a reverse direction, a reverse secondary transform can be applied first and then a reverse RPT can be applied. Specifically, a region where the transform coefficients r generated by applying direct RPT are positioned or a part of a corresponding region can be defined as an ROI for Petition 870250092804, dated 10 / 10 / 2025, page 76 / 157 68 / 135 apply a direct secondary transform. In this case, when a 16x64 direct RPT transform matrix is applied to an 8x8 region, the 16 generated transform coefficients can be positioned in an upper-left 4x4 subblock, and a corresponding subblock region can be defined as a ROI to apply a direct secondary transform to a corresponding ROI.
[0195] Furthermore, an RPT kernel can adjust a coefficient value by considering an operation as either an integer operation or a fixed-point operation. In other words, an RPT kernel can be configured to perform the transform through an integer operation (or a fixed-point operation) on a real codec system by appropriately scaling the kernel coefficients pertaining to a corresponding kernel, not a theoretical orthogonal transform or a non-orthogonal transform (here, an orthogonal transform and a non-orthogonal transform represent a transform in which the norm of each basis vector of the transform is 1). It can be reflected equally even when the RPT is applied, as well as a multiplied scaling factor when applying a separable transform in existing image compression technology.In this case, a separable transform or a non-separable transform (including RPT) can be performed while maintaining other processes (e.g., quantization and dequantization processes) different from the transform.
[0196] The integer coefficient of a PTR kernel can be obtained by multiplying a transform basis vector by a scaling value described above. As an embodiment, the multiplication by a scaling value may include applying an operation such as rounding, flooring, ceilinging, etc. to each kernel coefficient. In other words, an integer PTR kernel obtained through the method described above can be defined and used for a transform / inverse transform process. As described above, when the kernel coefficients of a scaled integer value are obtained through an operation such as rounding, flooring, ceilinging, Petition 870250092804, dated 10 / 10 / 2025, page 77 / 157 69 / 135 etc., the maximum and minimum values can be obtained for all kernel coefficients, so the number of bits sufficient to express all kernel coefficients can be obtained from the maximum and minimum values. For example, when the maximum value is less than or equal to 127 and the minimum value is greater than or equal to -128, all the entire kernel coefficients can be expressed in 8 bits (especially through two's complement expression, etc.).
[0197] Generally, when the maximum value is less than or equal to (2(N-1)- 1) and the minimum value is greater than or equal to -2(N-1), all the coefficients of the entire kernel can be expressed in N bits. When the maximum value is greater than (2(N-1)- 1) or the minimum value is less than -2(^), not all the coefficients of the entire kernel can be expressed in N bits. In this case, 1) all the kernel coefficients can be additionally multiplied by a scaling value to adjust them so that they are within the range of N bits, or 2) the number of bits needed to express a kernel coefficient can be increased (i.e., N+1 bits or more). When all the kernel coefficients need to be multiplied by 2-p(p >= 1) to express them with N bits, they can be compensated by multiplying by 2-p subsequently so that they are merged into the existing encoding / decoding process.As a modality, multiplying by 2p can be implemented by additionally performing a left shift operation of p bits or by reducing the amount of right shift applied in a quantization or dequantization process in p.
[0198] All kernel coefficients can be expressed in 8 bits, 9 bits, 10 bits, etc. using a method described above and, of course, the scaling value of a kernel coefficient can be set differently for each block or kernel size and the number of bits to express a kernel coefficient can be adjusted differently.
[0199] The NSPT described above can be applied based on at least Petition 870250092804, dated 10 / 10 / 2025, page 78 / 157 70 / 135 one of the following: size, tree type, or component type of a current block. For example, the application of NSPT can be determined based on at least one of the following factors: size, tree type, or component type of a current block. An NSPT index can be signaled based on at least one of the following: size, tree type, or component type of a current block. An NSPT set or an NSPT kernel can be derived based on at least one of the following: size, tree type, or component type of a current block.
[0200] The predefined allowed transform block sizes in a decoding device can be broadly divided into two groups. Either group (hereinafter referred to as the first group) can mean a set of block sizes to which NSPT is applicable. The first group can consist of any of the allowed transform block sizes or can consist of two or more block sizes among the allowed block sizes. A block size to which NSPT is applicable can be defined as a block size in which at least one width and one height is less than or equal to a predetermined threshold value. Alternatively, a block size to which NSPT is applicable can be defined as a block size in which the product of a width and a height is less than or equal to a predetermined threshold value.Alternatively, a block size to which NSPT applies can be set to a block size where the maximum value of a width and a height is less than or equal to a predetermined threshold value. The threshold value can be an integer of 4, 8, 16, 32, 64, 128 or higher.
[0201] Another of the two groups (hereinafter referred to as the second group) may mean a set of block sizes to which the NSPT is not applied. The separable primary transform described above may be applied to block sizes belonging to the second group. In addition, a non-separable secondary transform may be applied to all or part of the block sizes belonging to the Petition 870250092804, dated 10 / 10 / 2025, page 79 / 157 71 / 135 second group.
[0202] For example, when the size of a current block belongs to the first group, the reverse NSPT can be applied to the (dequantized) transform coefficient of a current block. When the size of a current block belongs to the second group, a reverse separable primary transform can be applied to the (dequantized) transform coefficient of a current block. Alternatively, when the size of a current block belongs to the second group, a reverse non-separable secondary transform (e.g., low-frequency non-separable transform, LFNST) can be applied first to the (dequantized) transform coefficient of a current block and a reverse separable primary transform (e.g., DCT-2) can be applied to a transform coefficient obtained from it.
[0203] As an example, the first group, which is a set of block sizes to which the NSPT is applicable, can be adjusted as a set of 4x4, 4x8, 8x4 and 8x8. Alternatively, the first group can be adjusted as a set of 4x8, 8x4 and 8x8. Alternatively, the first group can be adjusted as a set of 4x8 and 8x4. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 8x4, 8x8 and 16x4. Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 8x4, 8x8 and 16x4. Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 8x4 and 16x4. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, and 16x16. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 8x4, 8x8, 8x16, and 16x8. Alternatively, the first group can be adjusted as a set of 4x8, 8x4, 8x8, 8x16, and 16x8.Alternatively, the first group can be adjusted as a set of 4x8, 8x4, 8x16, and 16x8. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, 32x16, and 32x32. Alternatively, the... Petition 870250092804, dated 10 / 10 / 2025, p. 80 / 157 72 / 135 The first group can be configured as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, and 32x16. Alternatively, the first group can be configured as a set of 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, and 32x16. Alternatively, the first group can be configured as a set of 4x8, 8x4, 8x16, 16x8, 16x16, 16x32, and 32x16. Alternatively, the first group can be configured as a set of 4x8, 8x4, 8x16, 16x8, 16x32, and 32x16. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 8x4, and 16x4. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, and 16x8. Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, and 16x8. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 8x4, 8x16, 16x4, and 16x8.Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 8x4, 8x16, 16x4, and 16x8. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, and 16x16. Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, and 16x16. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 8x4, 8x16, 16x4, 16x8, and 16x16. Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 8x4, 8x16, 16x4, 16x8, and 16x16. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x16, 8x32, 16x4, 16x8, 32x4, and 32x8. Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 4x32, 8x4, 8x16, 8x32, 16x4, 16x8, 32x4, and 32x8. The first group can be arranged as a set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x16, 32x4, and 32x8.Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x32, 32x4, 32x8, and 32x16. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x16, 16x32, 32x4, 32x8, and 32x16. Petition 870250092804, dated 10 / 10 / 2025, page 81 / 157 73 / 135 Alternatively, the first group can be adjusted as a set of 4x8, 4x16, 4x32, 8x4, 8x16, 8x32, 16x4, 16x8, 16x32, 32x4, 32x8, and 32x16. Alternatively, the first group can be adjusted as a set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x16, 16x32, 32x4, 32x8, 32x16, and 32x32.
[0204] An NSPT matrix (or an NSPT kernel) with a predetermined dimension can be applied to block sizes belonging to the first group. Here, an NSPT matrix can be expressed as a matrix with dimension PxQ as an inverse transform matrix, and a PxQ matrix represents a matrix in which the number of rows and the number of columns are P and Q, respectively.
[0205] As an example, a 16x16 NSPT matrix can be applied to a 4x4 block. A 32x20 NSPT matrix can be applied to at least one 4x8 block or one 8x4 block. A 64x24 NSPT matrix can be applied to at least one 4x16 block or one 16x4 block. A 64x32 NSPT matrix can be applied to an 8x8 block. A 128x40 NSPT matrix can be applied to at least one 8x16 block or one 16x8 block. A 256x44 NSPT matrix can be applied to a 16x16 block. A 128x36, 128x38, or 128x40 NSPT matrix can be applied to a 4x32 or 32x4 block. An NSPT matrix of 256x48 can be applied to an 8x32 or 32x8 block. An NSPT matrix of 512x52 or 512x54 can be applied to a 16x32 or 32x16 block.
[0206] Since the matrix PxQ is a reverse NSPT matrix, an output vector Px1 can be obtained by applying a matrix PxQ to an input vector Qx1 (i.e., (matrix PxQ)x(input vector Qx1)). Here, an input vector Qx1 can correspond to transform coefficients (dequantized) within a current block to which the NSPT is applied. In this case, the value of Q can signify the number of transform coefficients to which the NSPT is applied and can be less than or equal to the product of the width and height of a current block. The value of Q can be variably determined based on the size of a current block between the Petition 870250092804, dated 10 / 10 / 2025, page 82 / 157 74 / 135 block sizes belonging to the first group described above. Alternatively, the value of Q can be set equally for block sizes belonging to the first group. An output vector Px1 can correspond to a residual signal (or decoded residual samples). The value of P can be equal to the product of the width and height of a current block.
[0207] On the other hand, a direct NSPT matrix can be expressed as a QxP matrix, which is the transpose of the PxQ matrix. An output vector Qx1 can be obtained by applying a QxP matrix to an input vector Px1 (i.e., (QxP matrix)x(input vector Px1)). Here, an input vector Px1 can correspond to residual samples within a current block to which NSPT is applied. The value of P can be equal to the product of the width and height of a current block. An output vector Qx1 can correspond to transform coefficients within a current block derived through NSPT. In this case, the value of Q can mean the number of transform coefficients produced by NSPT and can be less than or equal to the product of the width and height of a current block. Similarly, the value of Q can be variably determined based on the size of a current block among the block sizes belonging to the first group described above.Alternatively, the value of Q can be set similarly for block sizes belonging to the first group.
[0208] As in the example, NSPT can be applied to an MxN block and an NxM block, which are non-square blocks. For example, NSPT can be applied to a 4x8 block and an 8x4 block. Alternatively, NSPT can be applied to a 4x16 block and a 16x4 block, or NSPT can be applied to an 8x16 block and a 16x8 block, or NSPT can be applied to a 16x32 block and a 32x16 block.
[0209] By applying NSPT to specific block sizes belonging to the first group, the transformation can be performed in a more sophisticated way and the encoding performance can be improved. When direct LFNST is Petition 870250092804, dated 10 / 10 / 2025, p. 83 / 157 75 / 135 applied, the primary transform coefficients of the remaining region, except for a region to which the LFNST is applied (i.e., Region of Interest, ROI), can be zeroed. Furthermore, the LFNST can be composed of a small number of transform basis vectors. In this case, when a separable primary transform, such as DCT-2, and a non-separable secondary transform, such as LFNST, are applied to corresponding block sizes instead of NSPT, performance degradation can occur. When NSPT is applied instead of LFNST for a corresponding case, the zeroing process is omitted, and the coding performance can be improved compared to an LFNST application case. Additionally, performance improvement can be expected through an NSPT application method. NSPT or LFNST can be applied using the symmetry described below.Here, for LFNST, a transposition operation is performed on a matching input block using symmetry only for a region ROI. On the other hand, for NSPT, a transposition operation is performed on the entire block using symmetry. Thus, for NSPT, symmetry can be used in a more sophisticated way to train and apply a matching NSPT kernel, therefore, an improvement in performance can be expected.
[0210] Furthermore, when LFNST, not NSPT, is applied to an 8x8 block, a 32x64 transform matrix, instead of a 16x64 transform matrix, can be applied from the forward transform perspective. Here, a 16x64 transform matrix can be set up by sampling the top 16 rows into a 32x64 transform matrix. When the 16x64 transform matrix-based LFNST is applied to an 8x8 block, 16 multiplications are needed per sample to apply the LFNST, but when a 32x64 transform matrix is used, 32 multiplications are needed per sample to apply the LFNST. However, when a 32x64 transform matrix is used in this way, an improvement in coding performance can be expected. Petition 870250092804, dated 10 / 10 / 2025, page 84 / 157 76 / 135
[0211] When the tree type of a current block is a single tree, NSPT can be applied to the luma component of a current block, but NSPT cannot be applied to the chroma component of a current block. When the tree type of a current block is a double tree, NSPT can be applied to both the luma component and the chroma component of a current block.
[0212] Alternatively, regardless of whether a current block's tree type is a single tree, NSPT can be applied to the luma component of a current block and NSPT cannot be applied to the chroma component of a current block. Alternatively, regardless of whether a current block's tree type is a single tree, NSPT can be applied to both the luma component and the chroma component of a current block.
[0213] For example, when the tree type of a current block is a single tree, NSPT is allowed for one luma component and one chroma component, and the size of a current block belongs to the first group, an NSPT index can be assigned, and the luma component and the chroma component of a current block can share a corresponding NSPT index. Here, an NSPT index can be an index to select any of the transform kernel candidates for NSPT. When the size of the luma block and the chroma block of a current block belong to the first group, a transform kernel candidate selected by the same NSPT index can be applied to one luma component and one chroma component. When the tree type of a current block is a single tree and NSPT is applied only to one luma component, LFNST may not be applied, and a separate transform can be applied to the chroma component of a current block.Alternatively, when the tree type of a current block is a single tree and NSPT is applied only to a luma component, LFNST can be applied to the chroma component of a current block.
[0214] A single tree can have a high correlation between a component Petition 870250092804, dated 10 / 10 / 2025, page 85 / 157 77 / 135 of luma and one chroma component. In this case, unnecessary signaling can be reduced and compression efficiency can be improved by applying NSPT only to one luma component or, more commonly, by applying a candidate transform kernel selected by an NSPT index to both a luma component and a chroma component. On the other hand, for a non-unique tree, a luma component and a chroma component have independent partitioning and encoding structures, respectively. In this case, the characteristics of each component can be reflected, and compression efficiency can be improved by signaling an NSPT index for each component.
[0215] An NSPT kernel for NSPT can be derived based on at least one of the following: symmetry between intraprediction modes or symmetry between block shapes. As an example, an NSPT kernel can be derived as an NSPT kernel corresponding to at least one mode symmetrical to the intraprediction mode of a current block or a block shape symmetrical to the block shape of a current block. Alternatively, an NSPT kernel can be derived based on an NSPT set including one or more NSPT kernel candidates, wherein an NSPT set can be derived as an NSPT set corresponding to at least one mode symmetrical to the intraprediction mode of a current block or a block shape symmetrical to the block shape of a current block. Any one of one or more NSPT kernel candidates belonging to the NSPT set can be fitted as the NSPT kernel of a current block.For this purpose, an NSPT index specifying any one or more NSPT kernel candidates belonging to an NSPT set can be used. An NSPT index can be signaled via a bitstream or can be derived based on the symmetry described above.
[0216] There may be symmetry between at least two intraprediction modes among the predefined intraprediction modes in a decoding device. A Petition 870250092804, dated 10 / 10 / 2025, page 86 / 157 78 / 135 For the sake of descriptive convenience, the symmetry around an upper-left diagonal mode (i.e., mode 34) is described. With reference to Figure 5, there is symmetry between the directional modes. Excluding a planar mode, which is on° 0, and a DC mode, which is on° 1, all modes have a predictive direction. Mode 2 to mode 66 can be referred to as the normal directional mode (which can be expressed as [2, 66]), and mode -14 to mode -1 (which can be expressed as [-14, -1]) and mode 67 to mode 80 (which can be expressed as [67, 80]) can be called the wide directional mode. A broad directional mode can include at least one mode with a value less than -14 or one mode with a value greater than 80. With reference to Figure 5, all modes except mode 0 and mode 1 are symmetrical around mode 34. Specifically, the x and (68 - x) modes are symmetrical to the [2, 66] mode, and the x and (66 - x) modes are symmetrical between the [-14, -1] mode and the [67, 80] mode.The same symmetry relationship can be established between the mode [N, -1] and the mode [67, 66 - N]. Here, N can be an integer less than or equal to -14.
[0217] Meanwhile, with regard to the symmetry between the shapes of the blocks, an MxN block and an NxM block can be defined as blocks with symmetry to each other. Here, M and N can be the same or different from each other. Alternatively, when the ratio between the width and height of an M1xN1 block (M1 / N1) and the ratio between the height and width of an M2xN2 block (N2 / M2) are the same, an M1xN1 block and an M2xN2 block can be defined as blocks with symmetry to each other. Alternatively, when the ratio between the width and height of an M1xN1 block (M1 / N1) and the ratio between the height and width of an M2xN2 block (M2 / N2) are the same, an M1xN1 block and an M2xN2 block can be defined as blocks with symmetry to each other.
[0218] In a square block, modes that are symmetric to each other can share at least one NSPT set, one NSPT index, or one NSPT kernel. In other words, at least one NSPT set, one NSPT index, or one NSPT kernel for any of the symmetric modes can be applied equally to another of the Petition 870250092804, dated 10 / 10 / 2025, p. 87 / 157 79 / 135 symmetrical modes.
[0219] For example, symmetric modes can share an NSPT kernel. However, for any of the symmetric modes, a corresponding NSPT kernel can be applied to the input data, and for another, a corresponding NSPT kernel can be applied after applying a transposition operation to the input data. Specifically, when mode x belongs to mode [2, 33], a 1D vector can be set up for an MxM block that is input data in the order of the first column for mode x, and an NSPT kernel can be applied to a corresponding 1D vector. Here, setting up a 1D vector according to the first column order can read input data in the unit of a column from an MxM block, which is input data, to obtain M columns, and arrange them sequentially to set up a 1D vector.On the other hand, a 1D vector can be set in row-first order for the (68 - x) mode symmetric to the x mode, and the same corresponding NSPT kernel can be applied to a corresponding 1D vector. Here, setting a 1D vector according to row-first order can read input data in the unit of one row from an MxM block, which are input data to obtain M rows, and arrange them sequentially to set up a 1D vector. When the x mode belongs to the [N, -1] mode (N < -14), a 1D vector can be set in row-first order for the (66 - x) mode symmetric to the x mode, and the same NSPT kernel of the x mode can be applied to a corresponding 1D vector. Column-first order or row-first order can be applied to mode 0 and mode 1, and column-first order or row-first order can also be applied to mode 34.Furthermore, the first row order can be applied to an intraprediction mode belonging to the [2, 33] mode, and the first column order can be applied to a mode symmetric to a corresponding intraprediction mode. The first row order can be applied to an intraprediction mode belonging to the [N, -1] mode, and the first column order. Petition 870250092804, dated 10 / 10 / 2025, p. 88 / 157 80 / 135 can be applied in a symmetrical way to it.
[0220] For a non-square block, in addition to the symmetry between the intraprediction modes, the symmetry between the shapes of the blocks can also be considered. A non-square block whose width and height are M and N, respectively, can be considered as having a symmetry relation with a non-square block whose width and height are N and M, respectively. As an example, in the [2, 66] mode, there may be symmetry between the x-mode of an MxN block and the (68 - x)-mode of an NxM block. Similarly, when the x-mode of an MxN block belongs to the [N, -1]-mode (N < -14), there may be symmetry between the x-mode of an MxN block and the (66 - x)-mode of an NxM block.
[0221] A method for setting up a 1D vector from a block of input data is described above. In other words, when column-first order is applied to x-mode, row-first order can be applied to a mode symmetric to the same. Alternatively, when row-first order is applied to x-mode, column-first order can be applied to a mode symmetric to the same. Specifically, when column-first order is applied to x-mode, M columns can be obtained by reading input data in the unit of one column from an MxN block that is input data, and they can be arranged sequentially to set up a 1D vector. Here, each column can have a length of N. For a mode symmetric to x-mode, N rows can be obtained by reading input data in the unit of one row from an MxN block that is input data, and they can be arranged sequentially to set up a 1D vector. Here, each row can have a length of M.Alternatively, when first-row order is applied to x-mode, N rows can be obtained by reading input data in the unit of one row from an MxN block that is input data, and they can be arranged sequentially to set up a 1D vector. Here, each row can have a length of M. For a mode symmetrical to x-mode,... Petition 870250092804, dated 10 / 10 / 2025, page 89 / 157 81 / 135 M columns can be obtained by reading input data in the unit of one column from an MxN block that is input data, and they can be arranged sequentially to set up a 1D vector. Here, each column can have a length of N.
[0222] When a current block is an MxN block with x-mode and the symmetry described above is used for a current block, the NSPT set and / or the NSPT kernel of a current block can be determined based on at least one intraprediction mode symmetric to x-mode or an NxM block size symmetric to an MxN block size. Here, an NSPT kernel can be set as an NSPT kernel for an NxM block, not an NSPT kernel for an MxN block. In other words, when symmetry is used for a current block, an NSPT set and / or an NSPT kernel for a block that has symmetry with a current block can be used in the same way. As described above, a 1D vector can be set from an input data block according to a predetermined priority, which can correspond to the input NSPT kernel.
[0223] Furthermore, there may be a limit to the use of symmetry only when the intraprediction mode value of a current block is greater than 34. In other words, when the intraprediction mode value of a current block is greater than 34, a transposition operation can be applied when setting up a 1D vector of an input data block, and an NSPT set or an NSPT kernel corresponding to a block shape and / or a mode symmetrical with a current block can be used. Specifically, when the intraprediction mode of a current block belongs to the [N, -1] mode and the [2, 34] mode, symmetry cannot be used for a current block. On the other hand, when the intraprediction mode of a current block belongs to the [35, 66] mode and the [67, 66 - N] mode, symmetry can be used for a current block. Here, N can be an integer less than or equal to -14.
[0224] Derivation of an NSPT set or an NSPT kernel with Petition 870250092804, dated 10 / 10 / 2025, page 90 / 157 82 / 135 based on symmetry can be performed adaptively based on the size of a current block. For example, for a 4x4 block and an 8x8 block, an NSPT set or an NSPT kernel can be derived based on symmetry, and for a 4x8 block and an 8x4 block, an NSPT set or an NSPT kernel may not be derived based on symmetry.
[0225] Depending on whether symmetry is used, the number of available NSPT sets may be different. For example, when symmetry is used, the number of available NSPT sets may be 35, and when symmetry is not used, the number of available NSPT sets may be 67.
[0226] Table 9 below refers to an example where an NSPT set is determined using symmetry and shows a mapping relationship between an NSPT set and an intraprediction mode when the number of available NSPT sets is 35. [Table 9] Intraprediction Mode NSPT Set Index X < 0 2 0 < X < 34 X 35 < X < 66 68 - XX > 66 2
[0227] Referring to Table 9, when the value (X) of the intraprediction mode of a current block is less than 0, the NSPT set of a current block can be determined as an NSPT set with an NSPT set index of 2 among the 35 NSPT sets. When the value (X) of the intraprediction mode of a current block is greater than or equal to 0 and less than or equal to 34, the NSPT set of a current block can be determined as an NSPT set with an NSPT set index of X among the 35 NSPT sets. When the value (X) of the intraprediction mode of a current block is greater than or equal to 35 and less than or equal to 66, the Petition 870250092804, dated 10 / 10 / 2025, p. 91 / 157 The 83 / 135 NSPT set of a current block can be determined as an NSPT set with an NSPT set index of (68-X) among the 35 NSPT sets. When the value (X) of the intraprediction mode of a current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of a current block can be the same as an NSPT set corresponding to the value (68-X) of a mode symmetrical to the intraprediction mode of a current block. Similarly, when the value (X) of the intraprediction mode of a current block is greater than 66, the NSPT set of a current block can be determined as an NSPT set with an NSPT set index of 2 among the 35 NSPT sets. When the value (X) of the intraprediction mode of a current block is greater than 66, the NSPT set of a current block may be the same as an NSPT set corresponding to a mode symmetric to the intraprediction mode of a current block.
[0228] Table 10 below refers to an example in which an NSPT set is determined without using symmetry and shows a mapping relationship between an NSPT set and an intraprediction mode when the number of available NSPT sets is 67. [Table 10] Intraprediction Mode NSPT Set Index X < 0 2 0 < X < 66 XX > 66 66
[0229] Referring to Table 10, when the value (X) of the intraprediction mode of a current block is less than 0, the NSPT set of a current block can be determined as an NSPT set with an NSPT set index of 2 among the 67 NSPT sets. When the value (X) of the intraprediction mode of a current block is greater than or equal to 0 and less than or equal to 66, the NSPT set of a current block can be determined as an NSPT set with an index of Petition 870250092804, dated 10 / 10 / 2025, page 92 / 157 84 / 135 NSPT set of X among the 67 NSPT sets. Similarly, when the value (X) of the intraprediction mode of a current block is greater than 66, the NSPT set of a current block can be determined as an NSPT set with an NSPT set index of 66 among the 67 NSPT sets.
[0230] Symmetry can be used to save memory space required to store a transform kernel, while maintaining performance according to the transform application. For example, when 35 sets of NSPTs, instead of 67 sets of NSPTs, are used using symmetry, the memory space required to store an NSPT kernel can be significantly reduced.
[0231] The number of available NSPT sets and / or the number of NSPT kernel candidates belonging to an NSPT set may vary depending on the block size. For example, the number of available NSPT sets for a 4x4 block may be 35, the number of available NSPT sets for a 4x8 block and an 8x4 block may be 19, and the number of available NSPT sets for an 8x8 block may be 10. An NSPT set for a 4x4 block may consist of three NSPT kernel candidates, an NSPT set for a 4x8 block and an 8x4 block may consist of three or two NSPT kernel candidates, and an NSPT set for an 8x8 block may consist of one NSPT kernel candidate.
[0232] The size of a transform kernel can increase as the block size increases. Consequently, the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to an NSPT set can be reduced to save memory space needed to store a transform kernel. Furthermore, as the block size increases, the characteristics of a residual signal within a corresponding block tend to become more generalized. Consequently, reducing Petition 870250092804, dated 10 / 10 / 2025, p. 93 / 157 85 / 135 The number of available NSPT sets and / or the number of NSPT kernel candidates belonging to an NSPT set can help maintain compression efficiency while reducing implementation complexity by reflecting these statistical features.
[0233] An NSPT kernel can be configured with 8-bit precision. The coefficient range within an NSPT kernel can be greater than or equal to -128 and less than or equal to 127. When the corresponding precision is increased to more than 8 bits, a resulting value obtained through matrix multiplication can be shifted to the right due to the increased precision. For example, when a value obtained after matrix multiplication based on an 8-bit precision NSPT kernel is shifted to the right by S bits and stored in a buffer, if a kernel coefficient is configured with N-bit precision, it can be shifted to the right by (S (N-8)) bits and stored in a buffer.
[0234] When an NSPT kernel is configured with 8-bit precision, it is possible to avoid an excessive increase in internal precision in an encoder / decoder that performs the transform, thus reducing implementation complexity in terms of memory requirements and number of operations, while simultaneously minimizing a decrease in compression efficiency.
[0235] When reverse NSPT is applied to a current block of size NxN, the size of an NSPT kernel (or an NSPT matrix) can be expressed as MN x r. Here, MN can mean the product of the width and height of a current block. This can mean the output length of the NSPT or the number of residual samples generated by the NSPT. Additionally, r can mean the input length of the NSPT or the number of transform (dequantized) coefficients to which the NSPT is applied. r can be an integer greater than or equal to 0 and less than or equal to MN. A Petition 870250092804, dated 10 / 10 / 2025, p. 94 / 157 86 / 135 below is an example of an NSPT matrix of MN xr according to block size.
[0236] An NSPT matrix for a 4x4 block can be composed of a 16x16 matrix. An NSPT matrix for a 4x8 block and an 8x4 block can be composed of a 32x20 matrix, a 32x16 matrix, a 32x24 matrix, a 32x28 matrix, or a 32x32 matrix. An NSPT matrix for an 8x8 block can be composed of a 64x16 matrix, a 64x24 matrix, a 64x32 matrix, a 64x40 matrix, a 64x48 matrix, a 64x56 matrix, or a 64x64 matrix. An NSPT matrix for a 4x16 block and a 16x4 block can be composed of a 64x16 matrix, a 64x24 matrix, a 64x32 matrix, a 64x40 matrix, a 64x48 matrix, a 64x56 matrix, or a 64x64 matrix. An NSPT matrix for an 8x16 block and a 16x8 block can be composed of a 128x96 matrix, a 128x64 matrix, a 128x48 matrix, or a 128x32 matrix. An NSPT matrix for a 16x16 block can be composed of a 256x128 matrix, a 256x96 matrix, or a 256x64 matrix.An NSPT matrix for a 16x32 block and a 32x16 block can be composed of a 512x256 matrix or a 512x128 matrix. An NSPT matrix for a 32x32 block can be composed of a 1024x512 matrix, a 1024x256 matrix, or a 1024x128 matrix.
[0237] Alternatively, a 16x16 matrix can be applied to a 4xN block and an Nx4 block. Here, N can be an integer greater than or equal to 4. A 64x16 matrix can be applied to an 8x8 block. A 64x32 matrix can be applied to an 8xN block and an Nx8 block. Here, N can be an integer greater than or equal to 16. A 96x32 matrix can be applied to a 16xN block and an Nx16 block. Here, N can be an integer greater than or equal to 16.
[0238] Alternatively, the value of r in the NSPT matrix of MN xr can be determined according to a predetermined criterion. One criterion here could be (1) ensuring that the sum of the amount of calculation for the primary transform and the amount of calculation for the secondary transform is less than or equal to a certain Petition 870250092804, dated 10 / 10 / 2025, page 95 / 157 87 / 135 level, and (2) ensure that the number of multiplications per sample required for an NSPT operation is less than or equal to a certain number.
[0239] Based on the inverse transform, when a separable primary transform is performed by matrix multiplication for an MxN block, (MN) multiplications are required per sample to perform a corresponding primary transform. Furthermore, when the LFNST is applied to a certain region of interest (ROI), (P*Q) / (M*N) multiplications are required per sample when it is assumed that the LFNST matrix of the inverse transform is a PxQ matrix. Here, a PxQ matrix can mean a matrix whose number of rows is P and the number of columns is Q.
[0240] When NSPT, instead of DCT-2 transform (or separable transform such as KLT) and LFNST, is applied to an MxN block, the value of r ensuring that the number of multiplications per sample for a case where corresponding NSPT is applied is less than or equal to the number of multiplications per sample for a case where DCT-2 transform and LFNST are applied can be determined as follows. [Equation 6] r < (M + N + (P*Q) / (M*N))
[0241] When the value of r is set to the maximum value when satisfying Equation 6 above (i.e., r = M + N + (P*Q) / (M*N)), the value of r in an NSPT array for each block size can be set as follows. In Equation 6, when the value of (P*Q) / (M*N) is not an integer, an integer close to the value of (P*Q) / (M*N) can be used. As an example, a floor operation can be applied to the value of (P*Q) / (M*N). In this case, r can be set as (M + N + floor((P*Q) / (M*N))). Here, floor(x) can refer to the largest integer not greater than x. Alternatively, a rounding operation can be applied to the value of (P*Q) / (M*N). In this case, r can be adjusted as (M + N + round((P*Q) / (M*N))). Petition 870250092804, dated 10 / 10 / 2025, p. 96 / 157 88 / 135 Here, round(x) can refer to a value obtained by rounding x. Alternatively, a ceiling operation can be applied to the value of (P*Q) / (M*N). In this case, r can be adjusted as (M + N + ceil((P*Q) / (M*N))). Here, ceil(x) can refer to the smallest integer greater than or equal to x. When a floor operation is applied, an inequality in Equation 6 above may be satisfied. However, when a rounding operation or a ceiling operation is applied, an inequality in Equation 6 above may not be satisfied.
[0242] For NSPT for a 4x4 block, the maximum value of r is 24. However, since the value of r must be less than or equal to 16, the value of r can be adjusted to 16. For NSPT for a 4x8 block and an 8x4 block, the maximum value of r is 20. The value of r can be adjusted to 20. For NSPT for an 8x8 block, the maximum value of r is 32. The value of r can be adjusted to 32. For NSPT for a 4x16 block and a 16x4 block, the maximum value of r is 24. The value of r can be adjusted to 24. For NSPT for an 8x16 block and a 16x8 block, the maximum value of r is 40. The value of r can be adjusted to 40. For NSPT for a 16x16 block, the maximum value The value of r is 44. The value of r can be adjusted to 44. For NSPT for a 16x32 block and a 32x16 block, the maximum value of r is 54. The value of r can be adjusted to 54. For NSPT for a 32x32 block, the maximum value of r is 67. The value of r can be adjusted to 67. For NSPT for a 4x32 block and a 32x4 block, the maximum value of r is 38.The value of r can be set to 38. Alternatively, the value of r can be set to 20. For NSPT for an 8x32 block and a 32x8 block, the maximum value of r is 48. The value of r can be set to 48. Alternatively, the value of r can be set to 24.
[0243] In an NSPT array for each block size described above, there may be a case where the value of r is not a multiple of 4. To facilitate implementation, it may be advantageous for the value of r to be a multiple of 4. For example, when implementing parallel processing via Single Instruction, Petition 870250092804, dated 10 / 10 / 2025, page 97 / 157 89 / 135 Multiple Data (Single Instruction Multiple Data (SIMD)), etc., when the inner product of four transform basis vectors is processed simultaneously (i.e., four transform coefficients are generated simultaneously) in the direct NSPT application process, it may be advantageous to set the value of r as a multiple of 4.
[0244] For example, for NSPT for a 16x32 block and a 32x16 block, the value of r can be adjusted as 52 or 56 instead of 54. For NSPT for a 32x32 block, the value of r can be adjusted as 64 or 68 instead of 67. For NSPT for a 4x32 block and a 32x4 block, the value of r can be adjusted as 36 or 40 instead of 38.
[0245] More generally, the value of r can be set to a multiple of K. Here, K can be an integer greater than or equal to 1. As an example, the value of r can be set to a multiple of K that satisfies an inequality in Equation 7 below. [Equation 7] r < func ( ( M + N + (P*Q) / (M*N) ) / K ) * K
[0246] In Equation 7 above, func() can be floor, rounding, or ceiling as described above.
[0247] The value of r according to a predetermined pattern described above refers to a case where zeroing is not considered. In other words, when the direct LFNST is applied, the primary transform coefficients in the remaining region, except for a region where the LFNST is applied, are zeroed, so the actual amount of calculation required to apply DCT-2 and LFNST may be less than the amount of calculation above. Thus, when zeroing is considered, the value of r may be adjusted to a value less than the value of r according to the predetermined pattern.
[0248] Since zeroing is not performed for a 4x4 block, the value of Petition 870250092804, dated 10 / 10 / 2025, page 98 / 157 90 / 135 r can be adjusted to a value less than or equal to 16.
[0249] For a 4x8 block, zeroing can be performed for the remaining region, except for an upper-left 4x4 block based on the direct transform, and a 16x16 matrix, which is a direct LFNST matrix, can be applied to an upper-left 4x4 block. When this zeroing is performed, the number of sample multiplications required in the direct separable primary transform is 8 (((4x4x8) (4x8x4)) / (4x8)=8), and the number of sample multiplications required in the LFNST is 8 ((16x16) / 32=8). Thus, when the separable primary transform and the LFNST are replaced by NSPT, the value of r can be set to a value less than or equal to 16, which is the sum of the number of sample multiplications in the separable primary transform and the number of sample multiplications in the LFNST. Since the same amount of calculation is required even when applying the reverse separable primary transform and the LFNST, the value of r can be adjusted to a value less than or equal to 16.
[0250] For an 8x4 block, zeroing can be performed for the remaining region, except for an upper-left 4x4 block based on the direct transform, and a 16x16 matrix, which is a direct LFNST matrix, can be applied to an upper-left 4x4 block. When this zeroing is performed, the number of sample multiplications required in the direct separable primary transform is 6 ((4x8x4) (4x4x4) / (8x4)=6), and the number of sample multiplications required in the LFNST is 8 ((16x16) / 32=8). Thus, when the separable primary transform and the LFNST are replaced by NSPT, the value of r can be set to a value less than or equal to 14, which is the sum of the number of sample multiplications in the separable primary transform and the number of sample multiplications in the LFNST. Since the same amount of calculation is required even when applying the reverse separable primary transform and the LFNST, the value of r can be adjusted to a value less than or equal to 14. Petition 870250092804, dated 10 / 10 / 2025, p. 99 / 157 91 / 135
[0251] For an 8x8 block, zeroing cannot be performed for a separable primary transform and, in this case, the value of r can be set to a value less than or equal to 32.
[0252] For an 8x16 block, zeroing can be performed for the remaining region, except for an upper-left 8x8 block based on the direct transform, and a 64x32 matrix, which is a direct LFNST matrix, can be applied to an upper-left 8x8 block. When this zeroing is performed, the number of sample multiplications required in the direct separable primary transform is 16 ((8x8x16) (8x16x8) / (8x16)=16), and the number of sample multiplications required in the LFNST is 16 ((64x32) / 128=16). Thus, when the separable primary transform and the LFNST are replaced by NSPT, the value of r can be set to a value less than or equal to 32, which is the sum of the number of sample multiplications in the separable primary transform and the number of sample multiplications in the LFNST. Since the same amount of calculation is required even when applying the reverse separable primary transform and the LFNST, the value of r can be adjusted to a value less than or equal to 32.
[0253] For a 16x8 block, zeroing can be performed for the remaining region, except for an upper-left 8x8 block based on the direct transform, and a 64x32 matrix, which is a direct LFNST matrix, can be applied to an upper-left 8x8 block. When this zeroing is performed, the number of sample multiplications required in the direct separable primary transform is 12 ((8x16x8) (8x8x8) / (16x8)=12), and the number of sample multiplications required in the LFNST is 16 ((64x32) / 128=16). Thus, when the separable primary transform and the LFNST are replaced by NSPT, the value of r can be set to a value less than or equal to 28, which is the sum of the number of sample multiplications in the separable primary transform and the number of sample multiplications in the LFNST. Since the same amount of calculation is required even when applying the Petition 870250092804, dated 10 / 10 / 2025, pp. 100 / 157 92 / 135 reverse separable primary transform and LFNST, the value of r can be adjusted to a value less than or equal to 28.
[0254] For a 16x16 block, zeroing can be performed for the remaining region, except for a 12x12 upper left block based on the direct transform, and a 96x32 matrix, which is a direct LFNST matrix, can be applied to a 12x12 upper left block. When this zeroing is performed, the number of sample multiplications required in the direct separable primary transform is 21 ((12x16x16) (12x16x12) / (16x16)=21), and the number of sample multiplications required in the LFNST is 12 ((96x32) / 256=12). Thus, when the separable primary transform and the LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 33, which is the sum of the number of sample multiplications in the separable primary transform and the number of sample multiplications in the LFNST. Since the same amount of calculation is required even when applying the reverse separable primary transform and the LFNST, the value of r can be adjusted to a value less than or equal to 33.
[0255] For a 4x16 block, zeroing can be performed for the remaining regions, except for an upper left 4x4 block based on a direct transform, and a 16x16 matrix, which is a direct LFNST matrix, can be applied to an upper left 4x4 block. When this zeroing is performed, the number of sample multiplications required in a forward separable primary transform is 8 (((4x4x16) (4x16x4) / (4x16)=8), and the number of sample multiplications required in LFNST is 4 ((16x16) / 64=4). Thus, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 12, which is the sum of the number of sample multiplications in a separable primary transform and the number of sample multiplications in LFNST. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be Petition 870250092804, dated 10 / 10 / 2025, p. 101 / 157 93 / 135 adjusted to a value less than or equal to 12.
[0256] For a 16x4 block, zeroing can be performed for the remaining regions, except for a 4x4 upper left block based on a direct transform, and a 16x16 matrix, which is a direct LFNST matrix, can be applied to a 4x4 upper left block. When this zeroing is performed, the number of sample multiplications required in a direct separable primary transform is 5 (((4x16x4)+(4x4x4)) / (16x4)=5), and the number of sample multiplications required in LFNST is 4 ((16x16) / 64=4). Thus, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be set to a value less than or equal to 9, which is the sum of the number of sample multiplications in a separable primary transform and the number of sample multiplications in LFNST. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 9.
[0257] For a 4x32 block, zeroing can be performed for the remaining regions, except for a top-left 4x4 block based on a direct transform, and a 16x16 matrix, which is a direct LFNST matrix, can be applied to a top-left 4x4 block. When this zeroing is performed, the number of sample multiplications required in a forward separable primary transform is 8 (((4x4x32)+(4x32x4) / (4x32)=8), and the number of sample multiplications required in LFNST is 2 ((16x16) / 128=2). Thus, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be set to a value less than or equal to 10, which is the sum of the number of sample multiplications in a separable primary transform and the number of sample multiplications in LFNST. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 10. Petition 870250092804, dated 10 / 10 / 2025, p. 102 / 157 94 / 135
[0258] For a 32x4 block, zeroing can be performed for the remaining regions, except for a 4x4 upper left block based on a direct transform, and a 16x16 matrix, which is a direct LFNST matrix, can be applied to a 4x4 upper left block. When this zeroing is performed, the number of sample multiplications required in a forward separable primary transform is 4.5 (((4x32x4)+(4x4x4) / (32x4)=4.5), and the number of sample multiplications required in LFNST is 2 ((16x16) / 128=2). Consequently, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 6.5. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be adjusted to a value less than or equal to 6.5.Here, the value of r is a value that sets a matrix dimension, therefore it can be set to an integer of 6 or 7 instead of 6.5.
[0259] For an 8x32 block, zeroing can be performed for the remaining regions, except for an upper left 8x8 block based on a direct transform, and a 64x32 matrix, which is a direct LFNST matrix, can be applied to an upper left 8x8 block. When this zeroing is performed, the number of sample multiplications required in a forward separable primary transform is 16 (((8x8x32)+(8x32x8) / (8x32)=16), and the number of sample multiplications required in LFNST is 8 ((64x32) / 256=8). Consequently, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 24. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be adjusted to a value less than or equal to 24.
[0260] For a 32x8 block, zeroing can be performed for the remaining regions, except for an upper-left 8x8 block based on a direct transform, and a 64x32 matrix, which is a direct LFNST matrix, can be Petition 870250092804, dated 10 / 10 / 2025, page 103 / 157 95 / 135 applied to an upper left 8x8 block. When this zeroing is performed, the number of sample multiplications required in a forward separable primary transform is 10 (((8x32x8)+(8x8x8) / (32x8)=10), and the number of sample multiplications required in LFNST is 8 ((64x32) / 256=8). Consequently, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 18. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be adjusted to a value less than or equal to 18.
[0261] For a 16x32 block, zeroing can be performed for the remaining regions, except for a 12x12 upper left block based on a direct transform, and a 96x32 matrix, which is a direct LFNST matrix, can be applied to a 12x12 upper left block. When this zeroing is performed, the number of sample multiplications required in a forward separable primary transform is 21 (((12x16x32)+(12x32x12) / (16x32)=21), and the number of sample multiplications required in LFNST is 6 ((96x32) / 512=6). Consequently, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 27. Since the same amount of calculation is required even when applying a reverse separable primary transform and LFNST, the value of r can be adjusted to a value less than or equal to 27.
[0262] For a 32x16 block, zeroing can be performed for the remaining regions, except for a 12x12 upper left block based on a direct transform, and a 96x32 matrix, which is a direct LFNST matrix, can be applied to a 12x12 upper left block. When this zeroing is performed, the number of sample multiplications required in a direct separable primary transform is 13.5 (((12x32x12)+(12x16x12) / (32x16)=13.5), and the number of sample multiplications Petition 870250092804, dated 10 / 10 / 2025, page 104 / 157 The 96 / 135 required in LFNST is 6 ((96x32) / 512=6). Consequently, when a separable primary transform and LFNST are replaced by NSPT, the value of r can be adjusted to a value less than or equal to 19.5. Since the same amount of calculation is required even when applying the reverse separable primary transform and LFNST, the value of r can be adjusted to a value less than or equal to 19.5. Here, the value of r is a value that sets a matrix dimension, therefore it can be adjusted to an integer of 19 or 20 instead of 19.5.
[0263] In an NSPT array for each block size described above, there may be a case where the value of r is not a multiple of K. Here, K may be 4. In this case, to facilitate implementation, the value of r can be adjusted as a multiple of K. When the value of r that is predefined through the method described above is rprev, the value of r that is a multiple of K can be adjusted as follows. [Equation 8] r = func(rprev / K) x K
[0264] In Equation 8 above, func() can be floor, rounding, or ceiling as described above.
[0265] As described above, the value of r in an NSPT matrix for an MxN block may be different from the value of r in an NSPT matrix for an NxM block. For example, a reverse NSPT matrix for a 4x8 block may be a 32x16 matrix, and a reverse NSPT matrix for an 8x4 block may be a 32x14 matrix. In this case, an NSPT matrix can be determined using the symmetry between an MxN block and an NxM block.
[0266] It is assumed that a current block is an MxN block with mode x. When the symmetry described above is used for a current block, instead of applying an NSPT matrix corresponding to the size of the MxN block or mode x, an NSPT matrix corresponding to at least one mode symmetric to mode x or the size of the NxM block symmetric to the size of the MxN block can be applied. In this case, a Petition 870250092804, dated 10 / 10 / 2025, pp. 105 / 157 The 97 / 135 NSPT matrix corresponding to the NxM block size can be applied to an existing block as is. Alternatively, an NSPT matrix corresponding to the NxM block size is applied, but for the value of r, the value of r in an NSPT matrix corresponding to the MxN block size can be used.
[0267] For example, a reverse NSPT matrix for a 4x8 block can be a 32x16 matrix (i.e., the value of r in an NSPT matrix is 16), and a reverse NSPT matrix for an 8x4 block can be a 32x14 matrix (i.e., the value of r in an NSPT matrix is 14). When a current block is an 8x4 block with mode x, an NSPT matrix for at least one mode symmetric to mode x or a 4x8 block symmetric to an 8x4 block can be applied to a current block. In this case, a 32x16 matrix, which is a reverse NSPT matrix for a 4x8 block, can be used as is, or a 32x14 matrix with the value of r in a reverse NSPT matrix for an 8x4 block can be used. Here, a 32x14 matrix can be derived by sampling 14 rows from the left into a 32x16 matrix. In this way, when a 32x14 matrix is applied to a current block with a block size of 8x4, the predetermined pattern described above is satisfied.
[0268] On the other hand, when a current block is a 4x8 block with mode x, an NSPT matrix for at least one mode symmetric to mode x or an 8x4 block symmetric to a 4x8 block can be applied to a current block. In this case, a 32x14 matrix for an 8x4 block, not a 32x16 matrix for a 4x8 block, can be applied to a current block. With this, the NSPT can be executed using fewer multiplications than the number of multiplications allowed for a 4x8 block.
[0269] For NSPT for an MxN block and an NxM block, when the values of r that satisfy the predetermined condition described above are r1 and r2, respectively, a regression NSPT matrix for an MxN block and an NxM block can be fitted as MN x max(r1, r2). Here, max(r1, r2) can refer to selecting a value greater than or equal to r1 and r2. Petition 870250092804, dated 10 / 10 / 2025, pp. 106 / 157 98 / 135
[0270] For example, a reverse NSPT matrix for a 4x8 block can be a 32x16 matrix (i.e., the value of r in an NSPT matrix is 16), and a reverse NSPT matrix for an 8x4 block can be a 32x14 matrix (i.e., the value of r in an NSPT matrix is 14). When a current block is a 4x8 block with x mode, an NSPT matrix for at least one mode symmetric to x mode or an 8x4 block symmetric to a 4x8 block can be applied to a current block. In this case, a 32x16 matrix, which is a reverse NSPT matrix for an 8x4 block, can be used. When a reverse NSPT matrix is not set to MN x max(r1, r2), a 32x14 matrix will be used as an NSPT matrix for an 8x4 block. However, when an NSPT matrix for a 4x8 block and an NSPT matrix for an 8x4 block are configured as a 32 x max(16, 14) matrix, a 32x16 matrix can be fully applied.
[0271] On the other hand, when a current block is an 8x4 block with mode x, an NSPT matrix for at least one mode symmetric to mode x or a 4x8 block symmetric to an 8x4 block can be applied to a current block. In this case, a 32x16 matrix, which is a reverse NSPT matrix for a 4x8 block, can be used, or a 32x14 matrix can be used. Here, a 32x14 matrix can be derived by sampling 14 rows from the left in a 32x16 matrix.
[0272] When an NSPT matrix is configured as described above, the transform consisting of the maximum number of transform basis vectors can be applied as long as it satisfies the predetermined condition, thus maximizing encoding performance.
[0273] In the embodiment described above, the value of r can be set as a multiple of 16. For example, for reverse NSPT for a 4x8 block and an 8x4 block, a 32x16 matrix, not a 32x20 matrix, can be applied. The transform coefficients of a transform block can be encoded in the unit of a predetermined coefficient group (CG). Here, a CG can be defined as a Petition 870250092804, dated 10 / 10 / 2025, pp. 107 / 157 99 / 135 group of 16 transform coefficients and, as an example, a CG can be a subblock of size such as 4x4, 2x8 or 8x2. A non-zero transform coefficient may not exist within a CG and, in that case, the process of encoding a transform coefficient can be skipped to a corresponding CG. Thus, when the value of r is set as a multiple of 16, there is an advantage in reducing the complexity of the implementation.
[0274] The transform coefficients can be derived by applying direct NSPT to residual samples from an MxN block. In this case, the number of derived transform coefficients may be less than or equal to the value of (M*N) due to zeroing. In other words, a direct NSPT matrix can be defined as an rx (M*N) matrix, where r can mean the output length of the NSPT or the number of transform coefficients derived through the NSPT, and (M*N) can mean the input length of the NSPT or the number of residual samples to which the NSPT is applied.
[0275] The derived transform coefficients can be arranged in an MxN block according to the predetermined scan order, and a region where the transform coefficients are not filled can be filled with 0 (i.e., zeroing). Thus, in the process of scanning transform coefficients in a decoding device, when a non-zero transform coefficient is found in a region that would have been filled with 0 if the NSPT had been applied (or when the scan position of the last valid coefficient in an MxN block is greater than or equal to ar), the NSPT is considered not to be applied to a corresponding MxN block, and an NSPT index may not be signed.An index representing a sweep position of 0 can be allocated to an upper-left coefficient (i.e., a DC component coefficient) within an MxN block, and an index increased by 1 can be allocated to the remaining coefficients within an MxN block in a predetermined order. Petition 870250092804, dated 10 / 10 / 2025, pp. 108 / 157 100 / 135
[0276] One or more r values available for block sizes to which NSPT applies can be defined. For example, one or more r values can be defined for each block size to which NSPT applies. Alternatively, an r value can be defined for each block size to which NSPT applies, and the r value for any of the block sizes to which NSPT applies can be different from the r value for another. Alternatively, an r value can be defined for some of the block sizes to which NSPT applies, and at least two r values can be defined for the rest.
[0277] When a plurality of r values is available, an index specifying any one of a plurality of r values or r itself can be signaled. A corresponding index can be signaled in a high-level syntax (HLS), such as VPS, SPS, PPS, PH, or SH, or it can be signaled at a block level, such as CTU, CU, or TU. When the value of r is a value belonging to a specific range, enough bits to encompass a corresponding range can be allocated and signaled. For example, when the value of r belongs to the range 1 to 256, 8 bits can be designated as fixed-length and signaled.
[0278] The transform kernel of a current block can be determined based on any of the embodiments described above from 1 to 3. Alternatively, the transform kernel of a current block can be determined based on a combination of at least two of the embodiments 1 to 3 within a range where an invention according to the embodiments 1 to 3 described above does not conflict with each other.
[0279] A transform index for the inverse transform of a current block can be signaled. Here, a transform index can specify any one of one or more transform kernels (or transform matrices) belonging to a transform set. Here, a transform index can mean an NSPT index specifying any one of one or more NSPT kernels. Petition 870250092804, dated 10 / 10 / 2025, pp. 109 / 157 101 / 135 belonging to a set of NSPTs. Alternatively, the transform index may mean an LFNST index specifying any one of one or more LFNST kernels belonging to a set of LFNSTs.
[0280] The correspondence between a transform index and an NSPT index can be determined based on whether the size of a current block is one of the block sizes belonging to the first group described above. It assumes that the block sizes to which NSPT is applicable and the block sizes to which LFNST is applicable are distinct from each other. In this case, when the size of a current block belongs to the first group, a signed transform index for a current block can correspond to an NSPT index, and an NSPT kernel can be determined from an NSPT set based on a corresponding transform index. Conversely, when the size of a current block does not belong to the first group, a signed transform index for a current block can correspond to an LFNST index, and an LFNST kernel can be determined from an LFNST set based on a corresponding transform index.When the size of a current block does not belong to the first group, it may mean that the size of a current block belongs to the second group described above. Alternatively, when the size of a current block does not belong to the first group, it may mean that the size of a current block corresponds to a block size to which LFNST is applicable among the block sizes belonging to the second group. In this way, an NSPT index and an LFNST index can be configured as an integrated syntax, not as a separate syntax.
[0281] As an example, it is assumed that the block sizes to which NSPT is applicable, belonging to the first group, are 4x4, 4x8, 8x4 and 8x8. For block sizes belonging to the first group, NSPT can be applied instead of LFNST. Specifically, NSPT can be applied instead of a combination Petition 870250092804, dated 10 / 10 / 2025, page 110 / 157 102 / 135 of separable primary transform (e.g., DCT-2, separable KLT) and LFNST. An NSPT index can be signed for four block sizes belonging to the first group, and an LFNST index can be signed for the remaining block sizes (where LFNST is allowed).
[0282] Thus, when an NSPT index and an LFNST index are flagged as a syntax, the amount of encoded information can be reduced. Furthermore, the implementation complexity can be reduced by applying at least one of the following: binarization, a CABAC context, or an initial value for entropy encoding to an NSPT / LFNST index alike.
[0283] Alternatively, an NSPT index and an LFNST index can be flagged separately as a separate syntax. In this case, the implementation complexity may be partially increased, but compression performance can be improved by optimized entropy encoding performance for each index.
[0284] When the number of LFNST kernel candidates belonging to an LFNST set and the number of NSPT kernel candidates belonging to an NSPT set are the same, the same binarization can be applied to an LFNST index and an NSPT index. The same CABAC context (or CABAC context increment) can be allocated to the bins of an LFNST index and an NSPT index.
[0285] Different binarization and / or CABAC contexts can be used for an LFNST index and an NSPT index. A different initial CABAC value can be allocated to an LFNST index and an NSPT index. For example, either the LFNST and NSPT indexes can be binarized based on fixed-length binarization, and the other can be binarized based on truncated unary binarization. Even when the binarization of an LFNST index and an NSPT index are the same, a different CABAC context and / or initial CABAC value can be allocated. When the Petition 870250092804, dated 10 / 10 / 2025, p. 111 / 157 103 / 135 The number of LFNST kernel candidates belonging to an LFNST set and the number of NSPT kernel candidates belonging to an NSPT set are different from each other; different binarizations and / or CABAC contexts may be used for an LFNST index and an NSPT index.
[0286] The number of NSPT kernel candidates belonging to an NSPT set can be adjusted differently for each block size. Alternatively, the block sizes belonging to the first group can be divided into a plurality of subgroups. In this case, the number of NSPT kernel candidates belonging to an NSPT set can be adjusted differently for each of the various subgroups. Here, at least one of a plurality of subgroups can include a plurality of different block sizes.
[0287] Depending on the number of NSPT kernel candidates belonging to an NSPT set, the binarization applied to an NSPT index may be different.
[0288] For example, when the number of NSPT kernel candidates in an NSPT set for a specific block size is 3, an NSPT index can have a value from 0 to 3. When the value of an NSPT index is 0, this may represent that NSPT is not applied to a current block. When the value of an NSPT index is not 0, it may represent an NSPT kernel candidate corresponding to a matching NSPT index among the three NSPT kernel candidates. A bin can be allocated to distinguish between a case where NSPT is applied and a case where it is not applied. A case where the value of a matching bin is 0 may correspond to a case where the value of an NSPT index is 0. Conversely, a case where the value of a matching bin is 1 may correspond to a case where the value of an NSPT index is 1, 2, or 3. In this case, truncated unary binarization can be applied to Petition 870250092804, dated 10 / 10 / 2025, page 112 / 157 104 / 135 distinguish three NSPT kernel candidates. In other words, two compartments can be allocated to distinguish three NSPT kernel candidates as 0, 10, and 11.
[0289] When the number of NSPT kernel candidates in an NSPT set for a specific block size is 2, an NSPT index can have a value from 0 to 2. When the value of an NSPT index is 0, this may represent that NSPT is not applied to a current block. When the value of an NSPT index is not 0, it may represent an NSPT kernel candidate corresponding to a matching NSPT index between the two NSPT kernel candidates. A compartment can be allocated to distinguish between a case where NSPT is applied and a case where it is not applied. Two NSPT kernel candidates can be distinguished by allocating a compartment representing either of the two NSPT kernel candidates.
[0290] When the number of NSPT kernel candidates in an NSPT pool for a given block size is 1, an NSPT index can have a value of 0 or 1. When the value of an NSPT index is 0, this may indicate that NSPT is not applied to the current block. When the value of an NSPT index is 1, it may indicate an NSPT kernel candidate. In this case, the application of NSPT and an NSPT kernel candidate can be specified with only one compartment.
[0291] The inverse transform of the current block can be a separable primary transform and / or an inverse transform based on LFNST. In other words, the reverse LFNST can be applied to all or part of the transform coefficients (dequantized) of a current block, and then a reverse separable primary transform can be applied to the transform coefficients derived through the LFNST to derive residual samples. As an example, the reverse LFNST can be applied to transform coefficients (dequantized) Petition 870250092804, dated 10 / 10 / 2025, page 113 / 157 105 / 135 belonging to a partial region of a current block. Here, a partial region refers to a region to which the direct LFNST is applied and is henceforth referred to as a region of interest (ROI). The transform coefficients derived through the LFNST can be arranged in an ROI region according to the predetermined scan order. The predetermined scan order can be first row or first column. A reverse separable primary transform can be applied to transform coefficients derived through the LFNST and transform coefficients belonging to the remaining region, except for an ROI region within the current block.Alternatively, in a forward transform process, when zeroing is performed on the remaining region except for a region ROI within the current block (i.e., when the transform coefficients within the remaining region are set to 0), a reverse separable primary transform can be applied to the transform coefficients derived through LFNST.
[0292] Next, a method for assigning a transform index to the inverse transform of a current block will be described. Inverse transform here may refer to a non-separable inverse transform. A non-separable transform may refer to LFNST or NSPT as described above, and a transform index may refer to an LFNST or NSPT index.
[0293] When a current block is encoded as a single tree, a non-separable transform can be applied to the luma component of a current block, and a non-separable transform can be applied to the chroma component of a current block. In this case, a transform index for a current block can be signaled based on at least one of the first conditions for a luma component or the second condition for a chroma component. As an example, a transform index for a current block can be signaled when the first condition for a luma component and the second condition for a chroma component are satisfied, and otherwise (i.e., when either one Petition 870250092804, dated 10 / 10 / 2025, pp. 114 / 157 106 / 135 of the first conditions for a luma component or the second condition for a chroma component is not met), it may not be signed. Alternatively, a transform index for a current block may be signed when the first condition for a luma component is met, without checking whether the second condition for a chroma component is met; otherwise, it may not be signed. When a transform index is not signed, a non-separable transform may be set to not apply to a current block, and as an example, a transform index for a current block may be derived as 0.
[0294] The first condition for a luma component according to the present disclosure may mean that there is no non-zero transform coefficient in a predetermined region within the luma component block of a current block. Herein, a predetermined region within a luma component block may be defined as the remaining regions, excluding the first region in a luma component block, which may be referred to as the second region to distinguish it from the first region. The first region may be defined as a region composed of the same number of samples (or sample positions) as the number of input transform coefficients for a reverse non-separable transform (or the number of transform coefficients emitted through a forward non-separable transform). The first region may include the upper-left sample position of a luma component block.The first region may include at least one non-zero transform coefficient. The first region may be a region to which the last significant coefficient in a block of luma components belongs. The width and height of the first region may be less than or equal to the width and height of a block of luma components, respectively.
[0295] As described above, the number of transform coefficients Petition 870250092804, dated 10 / 10 / 2025, pp. 115 / 157 107 / 135 to which the reverse non-separable transform is applied can be determined based on the size of a current block. The size of a current block and a luma component block can be the same. The size of a current block can be set as a combination of width (L) and height (A), such as LxA. However, it is not limited to this, and the size of a current block can be defined as any of the following: width or height, minimum / maximum width and height value, or product of width and height.
[0296] As an example, in an encoding device, when a non-separable transform is applied to an MxN block, r transform coefficients less than or equal to (M*N) can be produced. This can mean that a non-separable direct transform has r output lengths. The output transform coefficients r can be arranged sequentially within an MxN block according to the predetermined scan order starting from the upper left sample position (i.e., DC position) of an MxN block. A region composed of r transform coefficients within a current block can correspond to the first region described above. Then, zeroing can be applied to a sample position after the r-th within an MxN block (i.e., the second region within a current block). Zeroing can refer to allocating 0 to a sample position within the second region.Therefore, when a non-separable transform is applied in an encoding device, a non-zero transform coefficient does not exist for sample positions to which zeroing is applied.
[0297] In the process of deriving a transform coefficient based on residual information in a decoding device, when a non-zero transform coefficient is found at a sample position (i.e., a sample position after the r-th or second region) where zeroing would have been applied if a non-separable transform had been Petition 870250092804, dated 10 / 10 / 2025, pp. 116 / 157 108 / 135 applied in an encoding device, this means that a non-separable transform was not applied to an MxN block. In this case, a transform index for a non-separable transform may not be assigned to an MxN block. In this case, a non-separable transform can be defined not to be applied to an MxN block and, as an example, a corresponding transform index can be derived as 0.
[0298] The second condition for a chroma component according to the present disclosure may mean that there is no non-zero transform coefficient in a predetermined region within the chroma component block of a current block. Here, a predetermined region within a chroma component block may be defined as the remaining regions, excluding the first region in a chroma component block, which may be referred to as the second region to distinguish it from the first region. The first region may be defined as a region composed of the same number of samples (or sample positions) as the input length of a non-separable reverse transform corresponding to the size of a chroma component block. A corresponding input length may mean the number of transform coefficients entered in a non-separable reverse transform.The first region may include the upper-left sample position of a chroma component block. The first region may include at least one non-zero transform coefficient. The first region may be a region to which the last significant coefficient in a chroma component block belongs. The width and height of the first region may be less than or equal to the width and height of a chroma component block, respectively.
[0299] Although a non-separable transform is not applied to a block of chroma components, the input length of a reverse non-separable transform can be determined based on the size of a block of Petition 870250092804, dated 10 / 10 / 2025, pp. 117 / 157 109 / 135 chroma components. Here, the input length of a non-separable transform can be determined as the input length of a non-separable transform determined based on the size of a chroma component block. Alternatively, the matrix size (or input / output length) of a non-separable transform can be determined based on the size of a luma component block corresponding to a chroma component block, and half the input length of a corresponding non-separable transform can be determined as the input length of a non-separable transform. A method for determining a non-separable transform matrix according to block size is described above, and a detailed description is omitted here.
[0300] Specifically, a non-separable transform is not applied to a chroma component, but zeroing can be applied to transform coefficients belonging to a predetermined region within a block of chroma components in the same way as when a non-separable transform is applied to a chroma component. In other words, an encoding device can leave only transform coefficients belonging to some regions within a block of chroma components and set transform coefficients belonging to the remaining regions to 0. Here, some regions above may mean a region where the transform coefficients produced through a corresponding non-separable transform are disposed if a non-separable transform is applied to a block of chroma components.
[0301] As an example, assume that it is encoded in a 4:2:0 color format, the tree type of a current block is a single tree, the size of a luma component block and a chroma component block are 8x16 and 4x8, respectively, and 8x16 and 4x8 are block sizes that allow a non-separable transform. In this case, a non-separable transform can be Petition 870250092804, dated 10 / 10 / 2025, pp. 118 / 157 110 / 135 applied to a block of luminance components, and a non-separable transform cannot be applied to a block of chroma components. However, zeroing can also be applied to a block of chroma components in the same way as when a non-separable transform is applied.
[0302] When it is assumed that the direct NSPT is performed on a 4x8 block based on a 20x32 NSPT matrix, only 20 transform coefficients can be output through the NSPT. 20 transform coefficients will be arranged sequentially from a top-left sample position in a 4x8 block according to the predetermined scan order, and zeroing will be applied to the remaining 12 sample positions. When zeroing is applied to a chroma component block in the same way, the pre-output transform coefficients can be left as is for the 20th sample position, according to the predetermined scan order from a top-left sample position in a chroma component block, and 0 can be allocated from the 21st sample position to the 32nd sample position.Furthermore, a chroma component block can be composed of a Cb component block and a Cr component block, and zeroing can be applied to two component blocks in the same way.
[0303] The size of the first region within the chroma component block described above can be determined based on the size of a chroma component block. The first region can be defined as a region consisting of r sample positions. Alternatively, the first region can be defined as a region consisting of min(threshold, r) sample positions. min(threshold, r) is a function that generates the minimum value between threshold and r. The threshold is a predefined value in an encoding and decoding device, and can be an integer of 4, 8, 16, 32, 64 or more. The threshold can be variable based on the size of a chroma component block or the size of a block. Petition 870250092804, dated 10 / 10 / 2025, pp. 119 / 157 111 / 135 of the luma component corresponds to a block of chroma component (i.e., the size of a current block).
[0304] r can be determined based on the input length of a non-separable reverse transform corresponding to the size of a chroma component block. For example, r can be the same as the input length of a non-separable reverse transform corresponding to the size of a chroma component block. Here, an input length can refer to the number of transform coefficients to which a non-separable reverse transform is applied. This can mean that a non-separable reverse transform corresponding to the size of a chroma component block is a P x r matrix. Alternatively, r can be determined based on the output length of a non-separable direct transform corresponding to the size of a chroma component block. For example, r can be the same as the output length of a non-separable direct transform corresponding to the size of a chroma component block.Here, an output length can refer to the number of transform coefficients produced through a non-separable forward transform. This can mean that a non-separable forward transform corresponding to the size of a block of chroma components is an rxP matrix. In the non-separable transform matrix, the value of P can be the product of the width and height of a block of chroma components. Alternatively, the value of P can be the number of transform coefficients (or residual samples) derived through a non-separable reverse transform. The value of P can be the number of samples belonging to a region to which a non-separable forward transform is applied within a block of chroma components. For example, when a non-separable transform is NSPT or LFNST, the value of P can be the same as the product of the width and height of a block of chroma components.When a non-separable transform is LFNST, the value of P can be the same. Petition 870250092804, dated 10 / 10 / 2025, pp. 120 / 157 112 / 135 is the number of transform coefficients belonging to a region belonging to a region (ROI) to which a direct non-separable transform is applied.
[0305] As described above, when the non-separable transform is NSPT, an NSPT matrix with a predetermined dimension can be determined / mapped according to a block size to which NSPT can be applied. As an example, a direct NSPT matrix corresponding to a 4x4 block, a 4x8 block, an 8x4 block, an 8x8 block, a 4x16 block, a 16x4 block, an 8x16 block, and a 16x8 block can be a 16x16 matrix, a 20x32 matrix, a 20x32 matrix, a 32x64 matrix, a 24x64 matrix, a 24x64 matrix, a 40x128 matrix, and a 40x128 matrix, respectively. In other words, when a chroma component block is a 4x4 block, the value of r can be 16. When a chroma component block is a 4x8 block or an 8x4 block, the value of r can be 20. When a chroma component block is an 8x8 block, the value of r can be 32. When a chroma component block is a 4x16 block or a 16x4 block, the value of r can be 24.When a chroma component block is an 8x16 block or a 16x8 block, the value of r can be 40.
[0306] As described above, when the non-separable transform is LFNST, an LFNST matrix with a predetermined dimension can be determined / mapped according to a block size. As an example, a direct LFNST matrix corresponding to a 4xN block and / or an Nx4 block can be a 16x16 matrix. In other words, when a chroma component block is a 4xN block or an Nx4 block, the value of r can be 16. Here, N can be an integer greater than or equal to 4. Alternatively, a direct LFNST matrix corresponding to an 8x8 block can be a 16x64 matrix. In other words, when a chroma component block is an 8x8 block, the value of r can be 16. Alternatively, a direct LFNST matrix corresponding to an 8xN block and / or an Nx8 block can be a 32x64 matrix. In other words, when a chroma component block Petition 870250092804, dated 10 / 10 / 2025, pp. 121 / 157 113 / 135 is an 8xN block or an Nx8 block, the value of r can be 32. Here, N can be an integer greater than or equal to 16. A 16x64 matrix, which is a direct LFNST matrix corresponding to the 8x8 block, can be obtained by sampling 16 rows from the highest row in a 32x64 matrix, which is a direct LFNST matrix corresponding to an 8xN block or an Nx8 block. Alternatively, a direct LFNST matrix corresponding to a 16xN block and / or an Nx16 block can be a 32x96 matrix. In other words, when a chroma component block is a 16xN block or an Nx16 block, the value of r can be 32. Here, N can be an integer greater than or equal to 16.
[0307] As described above, the predefined allowed transform block sizes can be divided into the first group, which is a set of block sizes to which NSPT can be applied, and the second group, which is a set of block sizes to which NSPT is not applied.
[0308] As an example, the first group can be defined to include at least one of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, or 16x8, and the second group can be defined to include the remaining block sizes. In this case, NSPT can be applied to block sizes belonging to the first group, and LFNST can be applied to all or part of the block sizes belonging to the second group. Alternatively, the first group can be defined to include at least one of 4x4, 4x8, 8x4, or 8x8, and the second group can be defined to include the remaining block sizes. In this case, NSPT can be applied to block sizes belonging to the first group, and LFNST can be applied to all or part of the block sizes belonging to the second group. Alternatively, the first group can be defined to include at least one 4x4, 4x8, 8x4, 8x8, 4x16, or 16x4 block, and the second group can be defined to include the remaining block sizes.In this case, NSPT can be applied to block sizes belonging to the first group, and LFNST can be applied to all or part of the block sizes. Petition 870250092804, dated 10 / 10 / 2025, pp. 122 / 157 114 / 135 belong to the second group.
[0309] Meanwhile, LFNST can also be applied to a 4x4 block. This may mean that a 4x4 block corresponds to a block size to which both NSPT and LFNST can be applied. Alternatively, this may mean that a 4x4 block corresponds to a block size to which LFNST can be applied, but does not correspond to a block size to which NSPT can be applied. In other words, this may mean that a 4x4 block size is excluded from the first group.
[0310] Alternatively, LFNST can also be applied to a 4x4 block and an 8x8 block. This may mean that a 4x4 block and an 8x8 block correspond to a block size to which both NSPT and LFNST can be applied. Alternatively, this may mean that a 4x4 block and an 8x8 block correspond to a block size to which LFNST can be applied, but do not correspond to a block size to which NSPT can be applied. In other words, this may mean that a 4x4 and 8x8 block size is excluded from the first group.
[0311] Alternatively, LFNST can also be applied to an 8x8 block. This may mean that an 8x8 block corresponds to a block size to which both NSPT and LFNST can be applied. Alternatively, this may mean that an 8x8 block corresponds to a block size to which LFNST can be applied, but does not correspond to a block size to which NSPT can be applied. In other words, this may mean that an 8x8 block size is excluded from the first group.
[0312] The value of r for a block of chroma components can be adjusted according to the method described above. The value of r can be adjusted according to the method described above, regardless of whether a non-separable transform (e.g., NSPT and / or LFNST) is applied to a block of chroma components. Alternatively, the value of r can be adjusted according to the method Petition 870250092804, dated 10 / 10 / 2025, pp. 123 / 157 115 / 135 described above when a non-separable transform (e.g., NSPT and / or LFNST) is not applied to a chroma component block. Alternatively, the value of r can be adjusted according to the method described above when a non-separable transform (e.g., NSPT and / or LFNST) is applied to a chroma component block.
[0313] Even when a non-separable transform is not applied to a chroma component block, when the size of a corresponding chroma component block matches a block size to which NSPT or LFNST can be applied, the value of r can be adjusted according to the method described above. In this case, when the size of a corresponding chroma component block matches a block size to which NSPT can be applied, the value of r can be adjusted based on the input length of a reverse NSPT matrix (or the output length of a direct NSPT matrix) corresponding to a corresponding block size. When the size of a corresponding chroma component block matches a block size to which LFNST can be applied, the value of r can be adjusted based on the input length of a reverse LFNST matrix (or the output length of a direct LFNST matrix) corresponding to a corresponding block size.Next, when the tree type of a current block is a single tree, a zeroing method for a chroma component block will be described.
[0314] When a chroma component block is a 4xN block or an Nx4 block (N is an integer of 4 or more), only a predetermined number of transform coefficients among the transform coefficients of a chroma component block can be left, and the remaining transform coefficients can be set to 0. The transform coefficients of the chroma component block can be derived via a direct separable primary transform. Alternatively, the transform coefficients of the block of Petition 870250092804, dated 10 / 10 / 2025, pp. 124 / 157 116 / 135 chroma components can be derived through at least one direct NSPT or LFNST. The predetermined number can be r or min(16, r). The left transform coefficients can be arranged within a block of chroma components according to the predetermined scan order. 0 can be allocated to the remaining sample positions that are not filled with the left transform coefficients. In a decoding process, an inverse transform can be applied to the transform coefficients r or min(16, r) within a block of chroma components.
[0315] When a chroma component block is an MxN block (M and N are integers of 16 or more, respectively), only a predetermined number of transform coefficients among the transform coefficients of a chroma component block can be left, and the remaining transform coefficients can be set to 0. The transform coefficients of the chroma component block can be derived via a direct separable primary transform. Alternatively, the transform coefficients of the chroma component block can be derived via at least one direct NSPT or LFNST. The predetermined number can be r or min(256, r). The left transform coefficients can be arranged within a chroma component block according to the predetermined scan order. 0 can be allocated to the remaining sample positions that are not filled with the left transform coefficients.In a decoding process, an inverse transform can be applied to the transform coefficients r or min(256, r) within a block of chroma components.
[0316] When a chroma component block is an 8xN block or an Nx8 block (N is an integer of 8 or more), only a predetermined number of transform coefficients between the transform coefficients of a chroma component block can be left and the transform coefficients Petition 870250092804, dated 10 / 10 / 2025, pp. 125 / 157 The remaining 117 / 135 can be set to 0. The transform coefficients of the chroma component block can be derived via a direct separable primary transform. Alternatively, the transform coefficients of the chroma component block can be derived via at least one direct NSPT or LFNST. The predetermined number can be r or min(64, r). The left transform coefficients can be arranged within a chroma component block according to the predetermined scan order. 0 can be allocated to the remaining sample positions that are not filled with the left transform coefficients. In a decoding process, an inverse transform can be applied to the transform coefficients r or min(64, r) within a chroma component block.
[0317] Zeroing for the chroma component block can be zeroing according to NSPT or zeroing according to LFNST.
[0318] Specifically, when the size of a chroma component block (or the size of a luma component block corresponding to a chroma component block) matches a block size to which NSPT can be applied, a zeroing method according to NSPT can be applied to a chroma component block. Specifically, when the size of a chroma component block (or the size of a luma component block corresponding to a chroma component block) matches a block size to which LFNST can be applied, a zeroing method according to LFNST can be applied to a chroma component block.
[0319] Depending on the tree type of a current block, either luma component block size or chroma component block size can be selectively used to determine whether an NSPT-based zeroing method or an LFNST-based zeroing method is applied to a chroma component block. For example, when the tree type of a block Petition 870250092804, dated 10 / 10 / 2025, pp. 126 / 157 118 / 135 currently is a single tree, it can be determined based on the size of a luma component block, and when the tree type of a current block is a double tree, it can be determined based on the size of a chroma component block.
[0320] When it is determined that both NSPT and LFNST can be applied to a block of chroma components, zeroing can be applied based on either the output length of NSPT or LFNST. When NSPT and LFNST can be applied to a block of chroma components, zeroing can be applied based on the output length of a non-separable transform with a predefined priority. For example, NSPT can have a higher priority than LFNST. Alternatively, when NSPT and LFNST can be applied to a block of chroma components, zeroing can also be applied based on the minimum or maximum value of the output length of NSPT and the output length of LFNST.For example, when both NSPT and LFNST are applied to a block of chroma components, the output length of the direct NSPT corresponding to a block of chroma components is 20 and the output length of the direct LFNST corresponding to a block of chroma components is 16, only the transform coefficients at the 16th sample position of an upper left sample position within a block of chroma components can be left, and the transform coefficients at the remaining sample positions can be set to 0.
[0321] As described above, when a current block is encoded in a single tree, a non-separable transform is applied only to a luma component and a non-separable transform is not applied to a chroma component, but zeroing is applied, the first condition for a luma component and the second condition for a chroma component can be checked and, Petition 870250092804, dated 10 / 10 / 2025, pp. 127 / 157 119 / 135 when the first and second conditions are met, a transform index for a current block (in particular, a luma component block) can be signaled.
[0322] Alternatively, when the tree type of a current block is a single tree and a non-separable transform is applied only to a luma component, the zeroing described above may not be applied to a chroma component. In this case, the first condition for a luma component may be checked without checking the second condition for a chroma component. When the first condition for a luma component is satisfied, a transform index for a current block (in particular, a luma component block) may be signed. Conversely, when the first condition for a luma component is not satisfied (i.e., when there is a non-zero transform coefficient in the second region within a luma component block), a transform index may not be signed for a current block. In this case, a corresponding transform index may be derived as 0.
[0323] Alternatively, when the tree type of a current block is a single tree and a non-separable transform is applied to a luma component and a chroma component, both the first condition for a luma component and the second condition for a chroma component described above can be verified. In this case, when both the first and second conditions are satisfied, a transform index for a current block can be signaled. When either the first or second condition is not satisfied, a transform index for a current block may not be signaled.
[0324] When a color format is 4:2:0, the tree type of a current block is a single tree and the size of a luma component block is MxN, the size of a chroma component block can be (M / 2)x(N / 2). It assumes that an Intra-Subpartition (ISP) mode is not applied to a current block. When Petition 870250092804, dated 10 / 10 / 2025, pp. 128 / 157 120 / 135 A non-separable transform is applied to an MxN block for a luma component; a non-separable transform, such as NSPT or LFNST, can be applied to a chroma component, or NSPT and LFNST may not be applied. Specifically, it can be broken down as follows. 1) When LFNST is applied to an MxN block for a luma component 1-a) LFNST is applied to a (M / 2)x(N / 2) transform block for a chroma component. 1-b) NSPT is applied to a (M / 2)x(N / 2) transform block for a chroma component. 1-c) Both NSPT and LFNST are not applied to a (M / 2)x(N / 2) transform block for a chroma component (in this case, a DCT-2 based horizontal / vertical primary transform is applied to a chroma component). 2) When NSPT is applied to an MxN transform block for a luma component 2-a) LFNST is applied to a (M / 2)x(N / 2) transform block for a chroma component. 2-b) NSPT is applied to a (M / 2)x(N / 2) transform block for a chroma component. 2-c) Both NSPT and LFNST are not applied to a (M / 2)x(N / 2) transform block for a chroma component (in this case, a DCT-2 based horizontal / vertical primary transform is applied to a chroma component).
[0325] It is assumed that when M and N are greater than or equal to 4, NSPT and LFNST can be applied and NSPT can be applied to 4x4, 4x8, 8x4 and 8x8 blocks. In this case, when the tree type of a current block is a single tree, the following cases may occur. Petition 870250092804, dated 10 / 10 / 2025, pp. 129 / 157 121 / 135
[0326] When a luma component block is a 16x8 block, a chroma component block can be an 8x4 block. It can be configured to ensure that LFNST is applied to a luma component and NSPT is applied to a chroma component. Alternatively, it can be configured to ensure that LFNST is applied to a luma component and that both NSPT and LFNST are not applied to a chroma component. In this case, even when both NSPT and LFNST are not applied to a chroma component, zeroing can be applied in the same way as when NSPT or LFNST are applied as described above. Furthermore, since a chroma component block is an 8x4 block in this example, zeroing according to NSPT can be applied.
[0327] When a luma component block is an 8x8 block, a chroma component block can be a 4x4 block. It can be configured to ensure that NSPT is applied to a luma component and NSPT is applied to a chroma component. Alternatively, it can be configured to ensure that NSPT is applied to a luma component and that both NSPT and LFNST are not applied to a chroma component. In this case, even when both NSPT and LFNST are not applied to a chroma component, zeroing can be applied in the same way as when NSPT or LFNST are applied as described above. Furthermore, since a chroma component block is a 4x4 block in this example, zeroing according to NSPT can be applied.
[0328] When a luma component block is an 8x4 block, a chroma component block can be a 4x2 block. It can be configured to ensure that NSPT is applied to a luma component and that both NSPT and LFNST are not applied to a chroma component. In this case, zeroing according to NSPT or LFNST cannot be applied to a chroma component. Instead, a transform coefficient from an upper-left sample position in a chroma component block to the nth sample position, Petition 870250092804, dated 10 / 10 / 2025, pp. 130 / 157 122 / 135 according to the predetermined scan order, can be left as is, and zeroing can be applied to a sample position after the nth. Here, n is a predefined value equally for an encoding device and a decoding device, and can be an integer greater than or equal to 2. For example, n can be 4. The fourth sample position from the top left sample position can belong to a 2x2 region, such as a region that includes the top left sample of a chroma component block.
[0329] When a luma component block is a 32x32 block, a chroma component block can be a 16x16 block. It can be configured to ensure that LFNST is applied to a luma component and LFNST is applied to a chroma component. Alternatively, it can be configured to ensure that NFNST is applied to a luma component and that both NSPT and LFNST are not applied to a chroma component. In this case, even when both NSPT and LFNST are not applied to a chroma component, zeroing can be applied in the same way as when NSPT or LFNST are applied as described above. Furthermore, since a chroma component block is a 16x16 block in this example, zeroing according to LFNST can be applied.
[0330] When the tree type of a current block is a single tree and an Intra-Subpartition (ISP) mode is applied to a current block, or when the tree type of a current block is a double tree, the width and height of a chroma component block may not be half the width and height of a luma component block, respectively. In this case, for a luma component, the application of NSPT or LFNST and / or the matrix size (or input / output length) of a corresponding non-separable transform can be determined based on the size of a corresponding transform block. For a chroma component, the application of NSPT or LFNST and / or the matrix size (or length) Petition 870250092804, dated 10 / 10 / 2025, pp. 131 / 157 The input / output (123 / 135) of a corresponding non-separable transform can be determined based on the size of a corresponding transform block. Furthermore, when the tree type of a current block is a single tree, as described above, it can be configured to ensure that NSPT or LFNST is not applied to a chroma component, and of course, zeroing can be applied in the same way as when NSPT or LFNST is applied.
[0331] When a non-separable transform is applied to a luma component or a chroma component, a transform index for the non-separable transform can be signaled when the following condition is met. The conditions for signaling a transform index described below can be considered as an additional condition to the first condition for a luma component and the second condition for a chroma component. In other words, a transform index can be signaled only when the first and second conditions are met and at least one of the conditions described below is met. Alternatively, the conditions described below can be considered as an independent condition, independent of the first and second conditions.In other words, a transform index can be signaled when at least one of the conditions described below is met, regardless of whether both the first and second conditions are met.
[0332] 1) A transform index may be signed when a non-zero transform coefficient exists in a position other than a top-left sample position for at least one transform block among the color components (i.e., Y, Cb, Cr) and otherwise (i.e., when a non-zero transform coefficient does not exist in a position other than a top-left sample position for the transform blocks of all color components), a transform index may not be signed.
[0333] For example, when the tree type of a current block is a tree Petition 870250092804, dated 10 / 10 / 2025, pp. 132 / 157 124 / 135 double and a current block is the transform block of a chroma component, a corresponding transform index can be signaled only when a non-zero transform coefficient exists in a position other than an upper-left sample position for at least one transform block between a Cb component and a Cr component.
[0334] Alternatively, when the tree type of a current block is a single tree, a corresponding transform index can be signaled only when a non-zero transform coefficient exists in a position other than a top-left sample position for at least one transform block between a luma component and a chroma component (a chroma component can be composed of a Cb component and a Cr component).
[0335] Alternatively, when the tree type of a current block is a single tree and a non-separable transform is applied only to a luma component, a transform index may be signaled when a non-zero transform coefficient exists in a position other than a top-left sample position for at least one transform block between a luma component and a chroma component and the first condition for a luma component and the second condition for a chroma component are satisfied. Even when a non-zero transform coefficient does not exist in the luma component's transform block, a transform index may be signaled when the first and second conditions are satisfied.
[0336] 2) When at least one transform block among the color components (i.e., Y, Cb, Cr) of a current block is encoded by transform jump (i.e., when a transform jump flag for a corresponding component is 1), a non-separable transform cannot be applied to the transform blocks of all color components. When any of the Petition 870250092804, dated 10 / 10 / 2025, pp. 133 / 157 125 / 135 color components are encoded by the transform jump; a transform index for a current block may not be signed. A corresponding transform index may be derived as 0.
[0337] Alternatively, a transform jump can be applied to the transform block of some components among the color components of a current block, and a transform jump can be applied to the transform block of other components. In this case, a non-separable transform cannot be applied to the transform block of a color component to which a transform jump is applied, and a non-separable transform can be applied to the transform block of the remaining color components.
[0338] Based on the satisfaction of a predetermined condition for a color component to which the transform jump is applied, a non-separable transform can be applied to a color component to which a transform jump is not applied, and a transform index for a corresponding color component can be signed. As an example, a transform index can be signed for a color component to which a transform jump is not applied only when a predetermined condition is satisfied for a color component to which the transform jump is applied. Here, a predetermined condition can include at least one of the first conditions described above for a luma component or the second condition for a chroma component, or the third condition that a non-zero transform coefficient exists at a position other than a top-left sample position in a transform block for at least one color component.Alternatively, it can be configured to not check the predetermined condition for a color component to which a transform jump is applied.
[0339] With reference to Figure 4, a current block can be reconstructed based on the residual sample of a current S420 block. Petition 870250092804, dated 10 / 10 / 2025, pp. 134 / 157 126 / 135
[0340] Prediction samples for the current block can be derived based on the intraprediction mode of the current block. Reconstructed samples for the current block can be generated based on prediction samples and residual samples from the current block.
[0341] Figure 6 illustrates a schematic configuration of a decoding apparatus (300) that performs an image decoding method according to the present disclosure.
[0342] With reference to Figure 6, the decoding apparatus (300) according to the present disclosure may include a transform coefficient differentiator (600), a residual sample differentiator (610) and a reconstructed block generator (620). The transform coefficient differentiator (600) may be configured in the entropy decoder (310) of Figure 3, the residual sample differentiator (610) may be configured in the residual processor (320) of Figure 3, and the reconstructed block generator (620) may be configured in the adder (340) of Figure 3.
[0343] The transform coefficient differentiator (600) can obtain residual information from the current block of the bit stream and decode it to derive the transform coefficients of the current block.
[0344] The residual sample derivative (610) can derive residual samples from the current block by performing at least one dequantization or inverse transform on the transform coefficients of the current block.
[0345] The residual sample derivative (610) can determine a transform kernel for the inverse transform of the current block by means of a predetermined transform kernel determination method and derive the residual samples of the current block based on that. It is the same as described by referring to Figure 4, and a detailed description of that will be omitted here.
[0346] The rebuilt block generator (620) can rebuild the current block Petition 870250092804, dated 10 / 10 / 2025, pp. 135 / 157 127 / 135 based on residual samples from the current block.
[0347] Figure 7 illustrates an image encoding method performed by an encoding apparatus (200) according to an embodiment of the present disclosure.
[0348] With reference to Figure 7, residual samples from a current block can be derived (S700).
[0349] Residual samples from the current block can be derived by subtracting prediction samples from the original samples of the current block. Here, prediction samples can be derived based on a predetermined intraprediction mode.
[0350] With reference to Figure 7, the transform coefficients of the current block can be derived by performing at least one transform or quantization on the residual samples of the current block (S710).
[0351] A transform method according to the present disclosure can be understood as the reverse inverse transform process described with reference to Figure 4. A method for determining a transform kernel for the transform is the same as described with reference to Figure 4. A detailed description of this will be omitted here.
[0352] For example, one or more transform sets for the current block can be defined / configured, and each transform set can include one or more transform kernel candidates. In this case, one of the several transform sets can be selected as the transform set for the current block. One of the several transform kernel candidates belonging to the current block transform set can be selected. The selection can be performed implicitly based on the context of the current block. Alternatively, an ideal transform set and / or transform kernel candidate for the current block can be selected, and an index indicating this. Petition 870250092804, dated 10 / 10 / 2025, pp. 136 / 157 128 / 135 can be signaled.
[0353] Alternatively, the transform kernel of the current block can be determined based on a set of MTSs. One of a plurality of MTS sets can be selected based on at least one size or intraprediction mode of the current block. The selected MTS set can include one or more transform kernel candidates. One or more transform kernel candidates can be selected, and the transform kernel of the current block can be determined based on the selected transform kernel candidate. The selection of the transform kernel candidate can be performed using a transform kernel candidate index derived based on the context of the current block. Alternatively, an ideal transform kernel candidate for the current block can be selected, and a transform kernel candidate index indicating the selected transform kernel candidate can be signaled.
[0354] Alternatively, the transform kernel of a current block can be determined based on a non-separable primary transform kernel (NSPT). When the size of a current block belongs to the first group, which is a set of block sizes to which the NSPT is applicable, the direct NSPT can be applied to a current block, and when the size of a current block belongs to the second group, the direct NSPT cannot be applied to a current block. When the size of a current block belongs to the second group, a direct separable primary transform (e.g., DCT-2) can be applied to the residual sample of a current block to derive transform coefficients. The direct LFNST can additionally be applied to all or part of the transform coefficients derived through the separable primary transform.
[0355] Furthermore, NSPT can be applied based on at least one tree type or component type of a current block. An NSPT kernel (or an NSPT matrix) for NSPT can be determined using the symmetry between modes of Petition 870250092804, dated 10 / 10 / 2025, pp. 137 / 157 129 / 135 intraprediction or the symmetry between block shapes. When direct NSPT is applied to a current block of MxN, an NSPT kernel can be expressed as r x MN. Here, r means the output length of the NSPT or the number of transform coefficients generated by the NSPT, and MN is the product of the width and height of a current block, which can mean the input length of the NSPT or the number of residual samples to which the NSPT is applied. One method for determining the size of this NSPT kernel is the same as described by referring to Figure 4.
[0356] An LFNST index and / or an NSPT index for the transform can be encoded as an integrated syntax, or an LFNST index and an NSPT index can be encoded and inserted into a bitstream, respectively. The binarization for an LFNST index and an NSPT index, and the allocation of a CABAC context and an initial value, are the same as described with reference to Figure 4.
[0357] In addition, a method for signaling a transform index referring to Figure 4 was described, which can be applied equally to a method for encoding a transform index.
[0358] With reference to Figure 7, a bitstream can be generated by encoding the transform coefficients of the current block (S720).
[0359] Residual information about the transform coefficients can be generated based on the transform coefficients of the current block, and a bitstream can be generated by encoding the residual information.
[0360] Figure 8 illustrates a schematic configuration of an encoding apparatus (200) that performs an image encoding method according to the present disclosure.
[0361] With reference to Figure 8, the encoding apparatus (200) according to the present disclosure may include a residual sample differentiator (800), a transform coefficient differentiator (810) and a transform coefficient encoder (820). The residual sample differentiator (800) and the transform coefficient differentiator Petition 870250092804, dated 10 / 10 / 2025, pp. 138 / 157 130 / 135 transform (810) can be configured in the residual processor (230) of Figure 2, and the transform coefficient encoder (820) can be configured in the entropy encoder (240) of Figure 2.
[0362] The residual sample derailleur (800) can derive residual samples from the current block by subtracting prediction samples from the original samples of the current block. Here, prediction samples can be derived based on a predetermined intraprediction mode.
[0363] The transform coefficient derivative (810) can derive the transform coefficients of the current block by performing at least one transform or quantization on the residual samples of the current block. A transform coefficient derivative 810 can determine the transform kernel of a current block based on at least one of the embodiments described above 1 to 3 and apply the transform kernel to the residual sample of a current block to derive a transform coefficient.
[0364] The transform coefficient encoder (820) can encode the transform coefficients of the current block to generate a bit stream.
[0365] In the embodiment described above, the methods are described based on a flowchart as a series of steps or blocks, but a corresponding embodiment is not limited to the order of the steps, and some steps may occur simultaneously or in a different order from other steps, as described above. Furthermore, those skilled in the art may understand that the steps shown in a flowchart are not exclusive and that other steps may be included or one or more steps in a flowchart may be excluded without affecting the scope of the embodiments of this disclosure.
[0366] The method described above, according to the embodiments of this disclosure, may be implemented in a software form, and an encoding apparatus and / or a decoding apparatus, according to this disclosure, Petition 870250092804, dated 10 / 10 / 2025, pp. 139 / 157 131 / 135 can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0367] In the present disclosure, when embodiments are implemented as software, the method described above can be implemented as a module (a process, a function, etc.) that performs the function described above. A module can be stored in memory and can be executed by a processor. Memory can be internal or external to a processor and can be connected to a processor by a variety of well-known means. A processor can include an application-specific integrated circuit (ASIC), other chip assemblies, a logic circuit, and / or a data processing device. Memory can include read-only memory (ROM), random-access memory (RAM), flash memory, a memory card, a storage medium, and / or another storage device. In other words, the embodiments described herein can be performed by means of implementation in a processor, a microprocessor, a controller, or a chip.For example, functional units shown in each diagram can be implemented in a computer, processor, microprocessor, controller, or chip. In this case, implementation information (e.g., instruction information) or an algorithm can be stored in a digital storage medium.
[0368] In addition, a decoding apparatus and an encoding apparatus to which the embodiment(s) of the present disclosure apply may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device. Petition 870250092804, dated 10 / 10 / 2025, pp. 140 / 157 132 / 135 such as a video communication device, a mobile streaming transmission device, a storage medium, a camcorder, a device for providing video-on-demand (VoD) service, an over-the-top (OTT) video device, a device for providing streaming service over the Internet, a three-dimensional (3D) video device, a virtual reality (VR) device, a virtual reality (AR) device, a video phone device, a transport terminal (e.g., a vehicle terminal (including an autonomous vehicle), an aircraft terminal, a ship terminal, etc.) and a medical video device, etc., and may be used to process a video signal or a data signal.For example, an over-the-top (OTT) video device might include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0369] In addition, a processing method to which the embodiment(s) of the present disclosure apply may be produced in the form of a program executed by a computer and may be stored on a computer-readable recording medium. Multimedia data with a data structure in accordance with the embodiment(s) of the present disclosure may also be stored on a computer-readable recording medium. Computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CDROM, a magnetic tape, a floppy disk, and an optical media storage device.Furthermore, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmission via the Internet). Additionally, a bitstream generated by an encoding method may be... Petition 870250092804, dated 10 / 10 / 2025, pp. 141 / 157 133 / 135 stored on a computer-readable recording medium or can be transmitted via a wired or wireless communication network.
[0370] In addition, the embodiment(s) of this disclosure may be implemented by a computer program product by program code, and the program code may be executed on a computer by the embodiment(s) of this disclosure. The program code may be stored on a computer-readable carrier.
[0371] Figure 9 shows an example of a continuous content streaming system to which the modalities of the present disclosure can be applied.
[0372] With reference to Figure 9, the content streaming system to which modality(ies) of the present disclosure is / are applied may broadly include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0373] The encoding server generates a bitstream by compressing the input content from multimedia input devices, such as a smartphone, camera, camcorder, etc., into digital data and transmits it to the streaming server. As another example, when multimedia input devices, such as a smartphone, camera, or camcorder, directly generate a bitstream, the encoding server can be omitted.
[0374] The bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of this disclosure are applied, and the streaming server can temporarily store the bitstream in a bitstream transmission or reception process. Petition 870250092804, dated 10 / 10 / 2025, pp. 142 / 157 134 / 135
[0375] A streaming server transmits multimedia data to a user device based on the user's request via a web server, and the web server serves as a means to inform a user about which service is available. When a user requests a desired service from the web server, the web server delivers it to a streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, and in that case, the control server manages a command / response between each device in the content streaming system.
[0376] The streaming server can receive content from a media storage and / or encoding server. For example, when content is received from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0377] An example of a user device may include a mobile phone, a smartphone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDAs), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, smart glasses, a head-mounted display (HMD)), a digital TV, a desktop computer, and digital signage, etc.
[0378] Each server in the content streaming system can be operated as a distributed server, and in this case, data received from each server can be distributed and processed.
[0379] The claims set forth herein may be combined in various ways. For example, a technical feature of a method claim. Petition 870250092804, dated 10 / 10 / 2025, pp. 143 / 157 135 / 135 of this disclosure can be combined to be implemented as a device, and a technical feature of a device claim of this disclosure can be combined to be implemented as a method. Furthermore, a technical feature of a method claim of this disclosure and a technical feature of a device method claim can be combined and implemented as a device, and a technical feature of a method claim of this disclosure and a technical feature of a device method claim of this disclosure can be combined and implemented as a method. Petition 870250092804, dated 10 / 10 / 2025, pp. 144 / 157
Claims
1 / 3 CLAIMS 1. Image decoding method, CHARACTERIZED in that it comprises: obtaining residual information from a bitstream; deriving transform coefficients of a current block based on the residual information; deriving residual samples of the current block by performing at least one dequantization or inverse transform on the transform coefficients of the current block; and reconstructing the current block based on the residual samples of the current block, wherein the inverse transform is performed based on a non-separable primary transform (NSPT), and wherein the NSPT is applied based on at least one size, tree type, or component type of the current block.
2. Method, according to claim 1, CHARACTERIZED in that predefined allowed transform block sizes are divided into a first group which is a set of block sizes to which NSPT is applicable and a second group which is a set of block sizes to which NSPT is not applicable.
3. Method, according to claim 2, CHARACTERIZED in that when the size of the current block belongs to the first group, the inverse transform of the current block is performed based on the NSPT.
4. Method, according to claim 2, CHARACTERIZED in that when the size of the current block belongs to the second group, the inverse transform of the current block is performed based on a separable primary transform.
5. Method, according to claim 2, CHARACTERIZED by the fact that when the size of the current block belongs to the second group, the inverse transform of the current block is performed based on a non-separable secondary transform and a separable primary transform.
6. Method according to claim 2, characterized in that the first group includes 4x4 and the second group includes 8x8.
7. Method according to claim 2, characterized in that the first group includes 4x8 or 8x4, and the second group includes 16x16.
8. Method according to claim 2, characterized in that the first group includes 4x16 or 16x4, and the second group includes 16x32 or 32x16.
9. Method according to claim 2, characterized in that the first group includes 4x32 or 32x4, and the second group includes 32x32.
10. Method according to claim 2, characterized in that the first group includes 8x32 or 32x8, and the second group includes 32x32.
11. Image decoding method, CHARACTERIZED in that it comprises: deriving residual samples from a current block; deriving transform coefficients from the current block by performing at least one transform or quantization on the residual samples from the current block; and encoding the transform coefficients from the current block, wherein the transform is performed based on a non-separable primary transform (NSPT), and wherein the NSPT is applied based on at least one size, tree type, or component type of the current block.
12. Computer-readable storage medium CHARACTERIZED by storing a bitstream generated by the image encoding method as defined in claim 11. Petition 870250092804, dated 10 / 10 / 2025, pp. 146 / 157 3 / 3 13. Method for data transmission, CHARACTERIZED in that it comprises: obtaining a bitstream for image information, wherein the bitstream is generated by deriving residual samples from a current block, deriving transform coefficients by performing at least one transform or quantization on the residual samples of the current block, and encoding the transform coefficients of the current block; and transmitting data, including the bitstream, wherein the transform is performed based on a non-separable primary transform (NSPT), and wherein the NSPT is applied based on at least one size, tree type, or component type of the current block. Petition 870250092804, dated 10 / 10 / 2025, pp. 147 / 157