Transform-based image encoding method and apparatus therefor

By using LFNST and scaling lists, image encoding is processed based on the tree type and tag information of the blocks, solving the problem of high-cost transmission and storage of high-resolution images and videos, and improving image compression efficiency and quantization efficiency, especially the quantization efficiency of chroma components.

CN120956899APending Publication Date: 2025-11-14LG ELECTRONICS INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511225157.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-01-10
Filing Date
2021-01-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies suffer from high costs due to increased information volume when transmitting and storing high-resolution, high-quality images and videos, especially in immersive media such as virtual reality, artificial reality, and holograms, where image characteristics differ from real images, necessitating efficient image and video compression technologies.

Method used

The LFNST and scaling list method is used to process image encoding. The application of the scaling list is determined based on the tree type and label information of the block, which improves quantization efficiency, especially for the chroma component of single tree type.

Benefits of technology

It improves the overall image/video compression and quantization efficiency, especially the quantization efficiency of the chroma component in single-tree types, and is suitable for encoding and decoding high-resolution and high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956899A_ABST
    Figure CN120956899A_ABST
Patent Text Reader

Abstract

The invention relates to a transform-based image encoding method and an apparatus therefor. An image decoding method according to the present document comprises the steps of: applying LFNST to a transform coefficient in order to derive a modified transform coefficient; and deriving a residual sample for the target block based on an inverse primary transform of the modified transform coefficient, wherein the step of deriving the transform coefficient includes: determining whether a scaling list is applied to the current block based on a tree type of the current block and whether LFNST is applied; and deriving a transform coefficient for the current block from the residual information on the basis of the determination result, and enabling the scaling list to be applied when the tree type of the current block is a single tree and a chroma component.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application No. 202180013799.6 (International Application No. PCT / KR2021 / 000339), filed on August 10, 2022, with an international application date of January 11, 2021, entitled "Transformation-based Image Compilation Method and Apparatus". Technical Field

[0002] This disclosure relates to image coding technology, and more specifically, to a method and apparatus for encoding images based on transformations in an image coding system. Background Technology

[0003] There has been a growing demand for high-resolution, high-quality images and videos, such as ultra-high-definition (HUD) images and 4K or 8K or higher video, across various fields. As image and video data becomes higher resolution and higher quality, the relative amount of information or bits transmitted increases compared to existing image and video data. Therefore, transmission and storage costs increase if media such as existing wired or wireless broadband lines are used to transmit image data or if existing storage media are used to store image and video data.

[0004] Furthermore, there has been a growing interest in and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms. The broadcasting of images and videos, such as game graphics, whose image characteristics differ from those of real-world images, is also on the rise.

[0005] Therefore, efficient image and video compression technologies are needed to effectively compress, transmit, store, and play back high-resolution and high-quality images and videos with such various characteristics. Summary of the Invention

[0006] Technical issues

[0007] The technical aspect of this disclosure is to provide a method and apparatus for increasing image coding efficiency.

[0008] Another technical aspect of this disclosure is to provide a method and apparatus for increasing quantization efficiency.

[0009] Another aspect of this disclosure is to provide a method and apparatus for increasing the quantization efficiency of chromaticity components in a single-tree type.

[0010] Technical solution

[0011] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided. The method may include: deriving transform coefficients for a current block based on residual information received from a bitstream; deriving modified transform coefficients by applying LFNST to the transform coefficients; and deriving residual samples for a target block based on the inverse primary transform of the modified transform coefficients. The deriving of the transform coefficients may include determining whether a scaling list is applied to the current block based on whether LFNST is applied and the tree type of the current block; and deriving transform coefficients for the current block from the residual information based on the determination result. The scaling list may be applied when the tree type of the current block is a single tree and the current block is a chroma component.

[0012] LFNST may not be applied to the chroma components of the current block.

[0013] When the tree type of the current block is single tree and LFNST is performed on the current block, the scaling list may not be applied to the luminance component of the current block.

[0014] The method may further include receiving flag information indicating whether the scaling list is available when LFNST is performed.

[0015] When the flag indicates that the scaling list is unavailable and the LFNST index is greater than 0, the scaling list may not be applied to the luminance component.

[0016] When the flag indicates that the scaling list is unavailable and the LFNST index is greater than 0, the scaling list may not be applied to the chroma components if the current block's tree type is dual-tree chroma.

[0017] When the flag indicates that the scaling list is unavailable and the LFNST index is greater than 0, the scaling list may not be applied to the luminance component if the tree type of the current block is dual-tree luminance.

[0018] According to another embodiment of this disclosure, an image encoding method performed by an encoding device is provided. The method may include deriving transform coefficients of the current block from residual samples of the current block based on a transform process; determining whether a scaling list should be applied to the current block based on whether LFNST was performed during the transform process and the tree type of the current block; and quantizing the transform coefficients based on the determination, wherein the scaling list can be applied when the tree type of the current block is a single tree and the current block is a chroma component.

[0019] According to another embodiment of the present disclosure, a digital storage medium for storing image data can be provided, the image data including encoded image information and bit stream generated according to an image encoding method performed by an encoding device.

[0020] According to another embodiment of this disclosure, a digital storage medium can be provided that stores image data including encoded image information and bitstream, so that a decoding device can perform an image decoding method.

[0021] Beneficial effects

[0022] According to this disclosure, the overall image / video compression efficiency can be increased.

[0023] According to this disclosure, quantification efficiency can be increased.

[0024] According to this disclosure, the quantization efficiency of chromaticity components in a single-tree type can be increased.

[0025] The effects achievable through the specific examples of this disclosure are not limited to those listed above. For example, various technical effects may exist that can be understood or derived from this disclosure by one of ordinary skill in the art. Therefore, the specific effects of this disclosure are not limited to those explicitly described in this disclosure, but may include various effects that can be understood or derived from the technical features of this disclosure. Attached Figure Description

[0026] Figure 1 The illustrations represent examples of video / image coding systems to which this disclosure applies.

[0027] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding device to which this disclosure applies.

[0028] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device to which this disclosure applies.

[0029] Figure 4 The illustrations schematically depict multiple transformation schemes according to embodiments of this document.

[0030] Figure 5 An example is shown of intra-frame orientation patterns with 65 predicted orientations.

[0031] Figure 6 This is a diagram used to explain the RST according to embodiments of this disclosure.

[0032] Figure 7 This is a diagram illustrating the order in which the output data of the forward primary transform are arranged into a one-dimensional vector, according to the example.

[0033] Figure 8 This is a diagram illustrating the order in which the output data of the forward secondary transformation is arranged into two-dimensional blocks, according to an example.

[0034] Figure 9 This is a diagram illustrating a wide-angle intra-frame prediction mode according to an embodiment of this document.

[0035] Figure 10 This is a diagram illustrating the block shape applied using LFNST.

[0036] Figure 11 This is a diagram illustrating the arrangement of the output data of the positive LFNST according to the example.

[0037] Figure 12 The diagram illustrates how the amount of output data used for the forward LFNST is limited to a maximum of 16, based on the example.

[0038] Figure 13 This is a diagram illustrating the zeroing of a block in which 4x4 LFNST has been applied, based on an example.

[0039] Figure 14 This is a diagram illustrating the zeroing of a block in an 8x8 LFNST block, based on an example.

[0040] Figure 15 This is a diagram illustrating the image decoding method based on the example.

[0041] Figure 16 This is a diagram illustrating the image encoding method based on the example.

[0042] Figure 17 The diagram illustrates the structure of a content streaming system that applies this disclosure. Detailed Implementation

[0043] This document can be modified in various ways and can have various embodiments, and specific embodiments are illustrated and described in detail in the accompanying drawings. However, this is not intended to limit the document to the specific embodiments. The terminology generally used in this specification is used to describe specific embodiments and not to limit the technical spirit of the document. Unless otherwise expressly indicated in the context, singular expressions include plural expressions. Terms such as "comprising" or "having" in this specification should be understood to indicate the presence of the features, numbers, steps, operations, elements, components or combinations thereof described in this specification, without excluding the possibility of the presence or addition of one or more features, numbers, steps, operations, elements, components or combinations thereof.

[0044] Furthermore, for ease of description in relation to different features and functions, the elements in the accompanying drawings described in this document are illustrated independently. This does not imply that the individual elements are implemented as separate hardware or separate software. For example, at least two elements may be combined to form a single element, or a single element may be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of the claims of this document, unless they depart from the spirit of this document.

[0045] Hereinafter, preferred embodiments of the document are described in more detail with reference to the accompanying drawings. In the drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.

[0046] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Recommendation H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Recommendation H.265), EVC (essential video coding) standard, AVS2 standard, etc.).

[0047] Various embodiments relating to video / image encoding may be provided in this document, and unless otherwise specified, the embodiments may be combined and performed together.

[0048] In this document, video can refer to a collection of images over time. Generally, an image refers to a unit representing an image within a specific time period, and a tile is a unit that constitutes a part of an image. A tile may include one or more coding tree units (CTUs). An image may consist of one or more tiles. An image may consist of one or more groups of tiles. A group of tiles may include one or more tiles.

[0049] A pixel or cell (pel) can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when that pixel value is converted to the frequency domain, it can refer to a transform coefficient in the frequency domain.

[0050] A unit can represent a basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as blocks and regions may be used interchangeably. Typically, an MxN block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0051] In this document, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expression “A / B” can mean “A and / or B”. Additionally, “A, B” can mean “A and / or B”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”. Additionally, “A / B / C” can mean “at least one of A, B, and / or C”.

[0052] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" could include 1) "A only", 2) "B only", and / or 3) both "A and B". In other words, the term "or" in this document should be interpreted as indicating "alternatively or alternatively".

[0053] In this disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".

[0054] Additionally, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only", or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0055] Additionally, the parentheses used in this disclosure may mean "for example". Specifically, when indicated as "prediction (intra-frame prediction)", this may mean that "intra-frame prediction" is presented as an example of "prediction". That is, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be presented as an example of "prediction". Additionally, when indicated as "prediction (i.e., intra-frame prediction)", this may also mean that "intra-frame prediction" is presented as an example of "prediction".

[0056] The technical features described individually in one of the accompanying drawings of this disclosure may be implemented individually or simultaneously.

[0057] Figure 1 An example of a video / image encoding system that can be applied to an embodiment of this document is illustrated schematically.

[0058] refer to Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.

[0059] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0060] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.

[0061] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.

[0062] A transmitter can send encoded video / image information or data, output as a bitstream, to a receiver via a digital storage medium or network, either as a file or a stream. Digital storage media can include various media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.

[0063] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.

[0064] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.

[0065] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding device to which this document can be applied. In the following text, the term "video encoding device" may include an image encoding device.

[0066] refer to Figure 2The encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0067] Image partitioner 210 partitions the input image (or picture or frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or maximum coding unit (LCU), coding units can be recursively partitioned according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be partitioned into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. Alternatively, a binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding unit that has not been further partitioned. In this case, the maximum coding unit can be directly used as the final coding unit based on the encoding efficiency according to the image characteristics. Alternatively, coding units can be recursively partitioned into coding units of even greater depth as needed, such that the optimally sized coding unit can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can be divided or segmented from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving the transformation coefficients and / or a unit for deriving the residual signal from the transformation coefficients.

[0068] Depending on the context, units and terms such as blocks and regions can be used interchangeably. Typically, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can usually represent pixels or pixel values, and can represent pixel / pixel values ​​only for the luminance component, or only for the chrominance component. Samples can be used as a term corresponding to pixels or pels of a picture (or image).

[0069] Subtractor 231 subtracts the predicted signal (predicted block, predicted sample array) output from predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to converter 232. Predictor 220 can perform prediction on the processing target block (hereinafter referred to as the "current block") and can generate a prediction block that includes the prediction samples of the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to the prediction, such as prediction mode information, and send the generated information to entropy encoder 240. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.

[0070] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.

[0071] Inter-frame predictor 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, at the block, sub-block, or sample level. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information of the current block. In skip mode, unlike merge mode, residual signals cannot be sent. In motion information prediction (motion vector prediction, MVP) mode, motion vectors of neighboring blocks can be used as motion vector prediction terms, and the motion vector of the current block can be indicated by sending the motion vector difference as a signal.

[0072] Predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to predict a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Furthermore, the predictor can perform prediction on a block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games. Although IBC essentially performs prediction within the current block, the aspect of deriving a reference block within the current block can be similar to inter-frame prediction. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.

[0073] The predicted signals generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate reconstructed signals or residual signals. The transformer 232 can generate transform coefficients by applying transform techniques to the residual signals. For example, transform techniques may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when representing the relationship information between pixels. CNT refers to a transform obtained based on the predicted signals generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to blocks of variable size instead of square blocks.

[0074] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form. Entropy encoder 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 can encode information necessary for video / image reconstruction, other than the quantized transform coefficients (e.g., values ​​of syntax elements), together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored in bitstream form on a unit-by-unit basis in the Network Abstraction Layer (NAL). The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. Furthermore, the video / image information may also include general constraint information. In this disclosure, information and / or syntax elements transmitted from the encoding device / transmitted as signals to the decoding device may be included in the video / image information. The video / image information can be encoded by the encoding process described above and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include broadcast networks, communication networks, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage device (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.

[0075] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 234 and inverse transformer 235, a residual signal (residual block or residual sample) can be reconstructed. Adder 250 adds the reconstructed residual signal to the prediction signal output from predictor 220, thereby generating a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample, or array of reconstructed samples). When the target block has no residual, as in the case of applying skip mode, the prediction block can be used as a reconstructed block. The generated reconstructed signal can be used for intra-frame prediction of the next processed target block in the current block, and, as described later, can be used for inter-frame prediction of the next image by filtering.

[0076] Meanwhile, Luminance Mapping and Chromaticity Scaling (LMCS) can be applied during image encoding and / or reconstruction.

[0077] Filter 260 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. As discussed later in the description of each filtering method, filter 260 can generate various information related to filtering and send the generated information to entropy encoder 290. The information about filtering can be encoded in entropy encoder 290 and output as a bitstream.

[0078] The modified reconstructed image, which has been sent to memory 270, can be used as a reference image in inter-frame predictor 280. This allows the encoding device to avoid prediction mismatch between the encoding and decoding devices when applying inter-frame prediction, and also improves encoding efficiency.

[0079] Memory 270 DPB can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current image from which motion information has been derived (or encoded), and / or motion information of blocks in reconstructed images. The stored motion information can be sent to inter-frame predictor 221 to be used as motion information for neighboring blocks or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 222.

[0080] Figure 3 This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this document can be applied.

[0081] refer to Figure 3 The video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0082] When the input includes a bitstream containing video / image information, the decoding device 300 can interact with data already prepared therein. Figure 2 The processing of video / image information in the encoding device correspondingly reconstructs the image. For example, the decoding device 300 can derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit, which can be segmented along a quadtree, binary tree, and / or ternary tree structure using encoding tree units or maximum encoding units. One or more transform units can be derived from the encoding units. And, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.

[0083] Decoding device 300 is capable of receiving data from... in bitstream form. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this disclosure, the information and / or syntax elements transmitted / received by signals, which will be described later, can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax elements and decoding information of neighboring and decoded target blocks, or information about symbols / bins decoded in the previous step, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model after determining it using information about symbols / bins decoded for the next symbol / bin using the context model. Information about prediction from the information decoded in the entropy decoder 310 can be provided to the predictor 330, and information about the residuals (i.e., the quantized transform coefficients) and associated parameter information from the entropy decoder 310, which has already undergone entropy decoding, can be input to the dequantizer 321. Furthermore, information about filtering from the information decoded in the entropy decoder 310 can be provided to the filter 350. Simultaneously, a receiver (not shown) that receives the signal output from the encoding device can further configure the decoding device 300 as an internal / external element, and the receiver can be a component of the entropy decoder 310. Furthermore, the decoding device according to this disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be classified into information decoders (video / image / picture information decoders) and sample decoders (video / image / picture sample decoders). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, a predictor 330, an adder 340, a filter 350, and a memory 360.

[0084] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the order of coefficient scans already performed in the encoding device. The dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.

[0085] The inverse converter 322 obtains the residual signal (residual block, residual sample array) by performing an inverse transformation on the transformation coefficients.

[0086] The predictor can perform predictions on the current block and generate a prediction block that includes prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and specifically, can determine the intra-frame / inter-frame prediction mode.

[0087] The predictor can generate a predicted signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to predict a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-block copying (IBC) to predict a block. Intra-frame block copying can be used for content image / video coding in games, such as screen content coding (SCC). Although IBC essentially performs prediction within the current block, the aspect of deriving a reference block within the current block can be similar to that performed by inter-frame prediction. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.

[0088] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples can be located in the neighbors of the current block or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.

[0089] Inter-frame predictor 332 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction may include information indicating the inter-frame prediction mode used for the current block.

[0090] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from predictor 330. When the target block has no residual, as in the case of applying skip mode, the prediction block can be used as the reconstruction block.

[0091] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be output by filtering as described below, or it can be used for inter-frame prediction of the next image.

[0092] In addition, Luminance Mapping and Chromatography Scaling (LMCS) can be applied to image decoding processing.

[0093] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.

[0094] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current image from which motion information has been derived (or decoded) and / or motion information of blocks in a reconstructed image. The stored motion information can be sent to inter-frame predictor 260 to be used as motion information of neighboring blocks or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.

[0095] The examples described in this specification in the predictor 330, dequantizer 321, inverse transformer 322 and filter 350 of the decoding device 300 can be similarly or accordingly applied to the predictor 220, dequantizer 234, inverse transformer 235 and filter 260 of the encoding device 200.

[0096] As described above, during video encoding, prediction is performed to improve compression efficiency. A prediction block, i.e., a target coding block, can be generated by prediction, including prediction samples for the current block. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived similarly in both the encoding and decoding devices. The encoding device can improve image coding efficiency by signaling information (residual information) about the residual between the original block (not the original block) and the prediction block and the original sample values. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.

[0097] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, derive quantized transform coefficients by performing a quantization process on the transform coefficients, and send the relevant residual information to the decoding device as a signal (via bitstream). In this case, the residual information can include information such as value information, position information, transform scheme, transform kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform inverse quantization / inverse transform processes based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. Furthermore, the encoding device can derive residual blocks by performing inverse quantization / inverse transform on the quantized transform coefficients for inter-frame prediction reference in subsequent images and can generate a reconstructed image.

[0098] Figure 4 The illustration schematically depicts a multi-transformation technique according to an embodiment of the present disclosure.

[0099] refer to Figure 4 The converter can correspond to the above. Figure 2 The converter in the encoding device, and the inverse converter can correspond to the above. Figure 2 The inverse converter in the encoding device, or corresponding to Figure 3 The inverse converter in the decoding device.

[0100] The transformer can derive (primary) transform coefficients (S410) by performing a primary transform based on residual samples (residual sample array) in the residual block. This primary transform can be referred to as the core transform. In this paper, the primary transform can be based on multiple transform selection (MTS), and when the multiple transform is applied as the primary transform, it can be referred to as the multi-core transform.

[0101] Multi-core transform can represent a method of transforming a signal by additionally using Discrete Cosine Transform (DCT) Type 2 and Discrete Sine Transform (DST) Type 7, DCT Type 8, and / or DST Type 1. In other words, multi-core transform can represent a method of transforming a spatial domain residual signal (or residual block) into frequency domain transform coefficients (or primary transform coefficients) based on multiple transform kernels selected from DCT Type 2, DST Type 7, DCT Type 8, and DST Type 1. In this paper, from the perspective of the transformer, the primary transform coefficients can be referred to as time transform coefficients.

[0102] In other words, when applying conventional transform methods, transform coefficients can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2. However, when applying multi-core transforms, transform coefficients (or primary transform coefficients) can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. Here, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transform types, transform kernels, or transform cores. These DCT / DST types can be defined based on basis functions.

[0103] If a multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel can be selected from the transform kernels for the target block. A vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can represent the transform of the horizontal components of the target block, while the vertical transform can represent the transform of the vertical components of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target block (CU or sub-block), including the residual block.

[0104] Furthermore, according to one example, if a primary transform is performed by applying an MTS, the mapping relationship of the transform kernels can be set by setting specific basis functions to predetermined values ​​and combining the basis functions to be applied in the vertical or horizontal transform. For example, when the horizontal transform kernel is expressed as trTypeHor and the vertical transform kernel is expressed as trTypeVer, the value of trTypeHor or trTypeVer of 0 can be set to DCT2, the value of trTypeHor or trTypeVer of 1 can be set to DST-7, and the value of trTypeHor or trTypeVer of 2 can be set to DCT-8.

[0105] In this scenario, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transform cores. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values ​​are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values ​​are 1, MTS index 2 can indicate that trTypeHor is 2 and trTypeVer is 1, MTS index 3 can indicate that trTypeHor is 1 and trTypeVer is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values ​​are 2.

[0106] In one example, the transformation kernel set based on MTS index information is illustrated in the table below.

[0107] [Table 1]

[0108]

[0109] The transformer can derive modified (secondary) transform coefficients by performing a secondary transform based on the (primary) transform coefficients (S420). The primary transform is a transform from the spatial domain to the frequency domain, while the secondary transform refers to a transform performed using the correlation between the (primary) transform coefficients to achieve a more compressed expression. The secondary transform can include an inseparable transform. In this case, the secondary transform can be called an inseparable secondary transform (NSST) or a mode-correlated inseparable secondary transform (MDNSST). An inseparable secondary transform can be represented as generating modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived from the primary transform based on an inseparable transform matrix. In this case, instead of applying the vertical and horizontal transforms separately (or possibly not independently), the transform can be applied all at once based on the inseparable transform matrix. In other words, an inseparable secondary transform can represent a transform method in which the (primary) transform coefficients are not applied separately in the vertical and horizontal directions. For example, a two-dimensional signal (transform coefficients) is rearranged into a one-dimensional signal through a defined direction (e.g., row-major or column-major direction), and then modified transform coefficients (or secondary transform coefficients) are generated based on the inseparable transform matrix. For example, according to row-major order, M×N blocks are arranged in a row in the order of first row, second row, ..., and Nth row. According to column-major order, M×N blocks are arranged in a row in the order of first column, second column, ..., and Nth column. The inseparable secondary transform can be applied to the upper left region of a block containing (primary) transform coefficients (hereinafter referred to as a transform coefficient block). For example, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 8, an 8×8 inseparable secondary transform can be applied to the upper left 8×8 region of the transform coefficient block. Furthermore, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 4, and the width (W) or height (H) of the transform coefficient block is less than 8, then a 4×4 inseparable secondary transformation can be applied to the upper left min(8,W) × min(8,H) region of the transform coefficient block. However, this embodiment is not limited to this, and for example, even if only the width (W) or height (H) of the transform coefficient block is equal to or greater than 4, a 4×4 inseparable secondary transformation can still be applied to the upper left min(8,W) × min(8,H) region of the transform coefficient block.

[0110] Specifically, for example, if a 4×4 input block is used, the inseparable secondary transformation can be performed as follows.

[0111] A 4×4 input block X can be represented as follows.

[0112] [Formula 1]

[0113]

[0114] If X is represented as a vector, then the vector It can be represented as follows.

[0115] [Equation 2]

[0116]

[0117] In Equation 2, the vector It is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row priority.

[0118] In this case, the inseparable secondary transformation can be calculated as follows.

[0119] [Formula 3]

[0120]

[0121] In this formula, Let T denote the transformation coefficient vector, and let T denote the 16×16 (inseparable) transformation matrix.

[0122] Using Equation 3 above, the 16×1 transformation coefficient vector can be derived. And the vector can be scanned in order (horizontal, vertical, and diagonal, etc.). Reorganize into 4×4 blocks. However, the above calculation is an example, and hypercube-Givens transform (HyGT) and similar methods can also be used to calculate inseparable secondary transformations in order to reduce the computational complexity of inseparable secondary transformations.

[0123] Furthermore, in inseparable secondary transforms, the transform kernel (or transform type) can be selected as mode-dependent. In this case, the mode can include intra-frame prediction mode and / or inter-frame prediction mode.

[0124] As described above, inseparable secondary transformations can be performed based on 8×8 or 4×4 transformations determined by the width (W) and height (H) of the transform coefficient block. An 8×8 transformation is a transformation applicable to an 8×8 region contained within the transform coefficient block when both W and H are equal to or greater than 8, and this 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, a 4×4 transformation is a transformation applicable to a 4×4 region contained within the transform coefficient block when both W and H are equal to or greater than 4, and this 4×4 region can be the top-left 4×4 region within the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.

[0125] Here, to select mode-dependent transform kernels, two inseparable secondary transform kernels can be configured per transform set for both 8×8 and 4×4 transforms for inseparable secondary transforms, and four transform sets can exist. That is, four transform sets can be configured for 8×8 transforms, and four transform sets can be configured for 4×4 transforms. In this case, each transform set in the four transform sets for 8×8 transforms can include two 8×8 transform kernels, and each transform set in the four transform sets for 4×4 transforms can include two 4×4 transform kernels.

[0126] However, as the size of the transformation (i.e., the size of the region to which the transformation is applied) can be, for example, a size other than 8×8 or 4×4, the number of sets can be n, and the number of transformation kernels in each set can be k.

[0127] The transform set can be referred to as the NSST set or the LFNST set. A specific set within the transform set can be selected, for example, based on the intra-prediction mode of the current block (CU or sub-block). The Low-Frequency Inseparable Transform (LFNST) can be an example of a reduced inseparable transform, which will be described later, and represents an inseparable transform for low-frequency components.

[0128] For reference, for example, intra-prediction modes may include two non-directional (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. Non-directional intra-prediction modes may include planar intra-prediction mode number 0 and DC intra-prediction mode number 1, and directional intra-prediction modes may include 65 intra-prediction modes numbered 2 through 66. However, this is an example, and this document can be applied even if the number of intra-prediction modes differs. Furthermore, in some cases, intra-prediction mode number 67 may be used, and intra-prediction mode number 67 may represent a linear model (LM) mode.

[0129] Figure 5 The intra-frame orientation patterns for 65 predicted directions are schematically shown.

[0130] refer to Figure 5 Based on the intra-prediction mode 34 with a left-top diagonal prediction direction, intra-prediction modes can be divided into intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality. Figure 5In the diagram, H and V represent horizontal and vertical orientations, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 unit at the sample grid position. These numbers can represent the offset for the mode index value. Intra-prediction modes 2 to 33 are horizontally oriented, and intra-prediction modes 34 to 66 are vertically oriented. Strictly speaking, intra-prediction mode 34 can be considered neither horizontal nor vertical, but it can be classified as horizontally oriented when determining the transform set of the secondary transform. This is because the input data is transposed for a vertical orientation mode symmetric to intra-prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data to N×M data. Intra-prediction modes 18 and 50 can represent the horizontal and vertical intra-prediction modes, respectively, and intra-prediction mode 2 can be called the upper-right diagonal intra-prediction mode because it has a left reference pixel and performs prediction in the upper-right direction. Similarly, intra-prediction mode 34 can be referred to as the bottom-right diagonal intra-prediction mode, while intra-prediction mode 66 can be referred to as the bottom-left diagonal intra-prediction mode.

[0131] Based on the example, four transform sets can be mapped according to the intra-frame prediction mode, as shown in the table below.

[0132] [Table 2]

[0133]

[0134] As shown in Table 2, any one of the four transform sets, i.e., lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra-frame prediction mode.

[0135] When a specific set is determined to be used for an inseparable secondary transform, one of the k transform kernels in that set can be selected using the inseparable secondary transform index. The encoding device can derive the inseparable secondary transform index indicating the specific transform kernel based on rate-distortion (RD) check and can signal the inseparable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the inseparable secondary transform index. For example, lfnst index 0 can refer to the first inseparable secondary transform kernel, lfnst index 1 can refer to the second inseparable secondary transform kernel, and lfnst index 2 can refer to the third inseparable secondary transform kernel. Alternatively, lfnst index 0 can indicate that the first inseparable secondary transform is not applied to the target block, and lfnst indexes 1 through 3 can indicate three transform kernels.

[0136] The converter can perform an inseparable secondary transformation based on the selected transform core and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by a quantizer, and can be encoded and sent as a signal to the decoding device, and transmitted to the dequantizer / inverse converter in the encoding device.

[0137] Furthermore, as described above, if the secondary transformation is omitted, the (primary) transform coefficients, which are the output of the primary (separable) transform, can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and sent to the decoding device as a signal, and transmitted to the dequantizer / inverse transformer in the encoding device.

[0138] The inverse transformer can perform a series of processes in the reverse order of those already executed in the aforementioned transformers. The inverse transformer can receive (dequantized) transform coefficients and derive (primary) transform coefficients by performing a secondary (inverse) transform (S450), and obtain residual blocks (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding and decoding devices can generate reconstructed blocks based on the residual blocks and the prediction blocks, and can generate reconstructed images based on the reconstructed blocks.

[0139] The decoding device may further include a secondary inverse transform application determiner (or a component for determining whether to apply the secondary inverse transform) and a secondary inverse transform determiner (or a component for determining the secondary inverse transform). The secondary inverse transform application determiner can determine whether to apply the secondary inverse transform. For example, the secondary inverse transform can be NSST, RST, or LFNST, and the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on a secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on the transform coefficients of the residual block.

[0140] The secondary inverse transform determiner can determine the secondary inverse transform. In this case, the secondary inverse transform determiner can determine the secondary inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-prediction mode. In embodiments, the secondary transform determination method can be determined depending on the primary transform determination method. Various combinations of primary and secondary transforms can be determined according to the intra-prediction mode. Furthermore, in the example, the secondary inverse transform determiner can determine the region where the secondary inverse transform is applied based on the size of the current block.

[0141] Furthermore, as mentioned above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, the primary (separable) inverse transform can be performed, and residual blocks (residual samples) can be obtained. As mentioned above, the encoding and decoding devices can generate reconstructed blocks based on the residual blocks and the prediction blocks, and can generate reconstructed images based on the reconstructed blocks.

[0142] Furthermore, in this disclosure, the concept of reduced secondary transformation (RST), in which the size of the transformation matrix (kernel) is reduced, can be applied to reduce the computational and storage requirements of the inseparable secondary transformation.

[0143] Furthermore, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in this disclosure, i.e., kernel coefficients or matrix coefficients, can be represented in 8 bits. This is feasible in decoding and encoding devices, and compared to existing 9-bit or 10-bit methods, it reduces the amount of storage required to store the transform kernel and can reasonably accommodate performance degradation. Additionally, representing the kernel matrix in 8 bits allows for the use of smaller multipliers and is more suitable for Single Instruction Multiple Data (SIMD) instructions for optimal software implementation.

[0144] In this specification, the term "RST" can refer to a transformation performed on the residual samples of a target block based on a transformation matrix whose size is reduced according to a reduction factor. When performing a reduction transformation, the computational cost required for the transformation can be reduced due to the smaller size of the transformation matrix. In other words, RST can be used to address computational complexity issues that arise when transforming large blocks or when transforming indivisible blocks.

[0145] RST can be referred to by various terms such as reduced transform, reduced secondary transform, diminished transform, simplified transform, and simple transform, and the names that RST can be called are not limited to the examples listed. Alternatively, since RST is performed primarily in the low-frequency region of the transform block that includes non-zero coefficients, it can be called low-frequency inseparable transform (LFNST). The transform index can be called the LFNST index.

[0146] Simultaneously, when performing a secondary inverse transform based on RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include an inverse reduced secondary transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients, and an inverse primary transformer that derives residual samples for the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In this disclosure, deriving transform coefficients based on a transform can mean deriving the transform coefficients by applying a transform.

[0147] Figure 6This is a diagram illustrating an RST according to an embodiment of the present disclosure.

[0148] In this disclosure, "target block" may refer to the current block, residual block, or transform block to be encoded.

[0149] In the example RST, an N-dimensional vector can be mapped to an R-dimensional vector in another space, such that a reduced transformation matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to the R / N value. The reduction factor can be called a reduction factor, shrinkage factor, simplification factor, or other various terms. Furthermore, R can be called a reduction coefficient, but depending on the situation, the reduction factor can refer to R. Additionally, depending on the situation, the reduction factor can refer to the N / R value.

[0150] In this example, the reduction factor or reduction coefficient can be sent by signaling via a bitstream, but the example is not limited to this. For example, predefined values ​​for the reduction factor or reduction coefficient can be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient does not need to be sent by signaling separately.

[0151] The size of the reduced transformation matrix, as shown in the example, can be less than N×N (the size of the regular transformation matrix) and can be limited as shown in Equation 4 below.

[0152] [Formula 4]

[0153]

[0154] Figure 6 The matrix T in the reduced transformation block shown in (a) can refer to the matrix T in Equation 4. R×N .like Figure 6 As shown in (a), when the reduced transformation matrix T R×N When multiplied by the residual sample used for the target block, the transformation coefficients used for the current block can be derived.

[0155] In the example, if the size of the block to which the transformation is applied is 8×8 and R=16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 6 The RST in (a) can be represented as the matrix operation shown in Equation 5. In this case, the storage and multiplication computations can be reduced to approximately 1 / 4 by a reduction factor.

[0156] In this disclosure, matrix operations can be understood as operations on column vectors obtained by multiplying a column vector by a matrix placed to the left of the column vector.

[0157] [Formula 5]

[0158]

[0159] In Equation 5, r1 to r 64 The residual samples used for the target block can be represented, and specifically, they can be the transformation coefficients generated by applying the primary transformation. As a result of the calculation in Equation 5, the transformation coefficients ci of the target block can be derived, and the process of deriving ci can be shown in Equation 6.

[0160] [Formula 6]

[0161]

[0162] As a result of Equation 6, the transformation coefficients c1 to c of the target block can be derived. R In other words, when R = 16, the transformation coefficients c1 to c of the target block can be derived. 16 If a regular transform is applied instead of an RST, and a 64×64 (N×N) transform matrix is ​​multiplied by a 64×1 (N×1) residual sample, then although 64 (N) transform coefficients are derived for the target block, only 16 (R) transform coefficients are derived because an RST is applied. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data sent from the encoding device 200 to the decoding device 300 is reduced, thus improving the transmission efficiency between the encoding device 200 and the decoding device 300.

[0163] When considering the size of the transformation matrix, the size of a regular transformation matrix is ​​64×64 (N×N), but the size of a reduced transformation matrix is ​​reduced to 16×64 (R×N). Therefore, compared to performing a regular transformation, the storage usage ratio of performing an RST can be reduced. Furthermore, compared to the number of multiplications (N×N) when using a regular transformation matrix, using a reduced transformation matrix can reduce the number of multiplications (R×N) by the R / N ratio.

[0164] In the example, the transformer 232 of the encoding device 200 can derive transform coefficients for the target block by performing a primary transform and an RST-based secondary transform on the residual samples for the target block. These transform coefficients can be passed to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive modified transform coefficients based on the inverse reduced secondary transform (RST) for the transform coefficients, and can derive residual samples for the target block based on the inverse primary transform for the modified transform coefficients.

[0165] Based on the example inverse RST matrix T N×R The size is N×R, which is smaller than the size of the conventional inverse transformation matrix N×N, and is similar to the reduced transformation matrix T shown in Equation 4.R×N It has a transpose relationship.

[0166] Figure 6 (b) shows the matrix T in the reduced inverse transform block. t It can refer to the inverse RST matrix T N×R T (The superscript T indicates transpose). For example... Figure 6 As shown in (b), when the inverse RST matrix T RxN T Multiplying by the transform coefficients of the target block yields either the modified transform coefficients of the target block or a residual sample of the target block. The inverse RST matrix T RxN T It can be represented as (T) RxN T ) NxR .

[0167] More specifically, when the inverse RST is used as a secondary inverse transformation, when the inverse RST matrix T RxN T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. Furthermore, the inverse RST can be used as the inverse primary transform, and in this case, when the inverse RST matrix T... RxN T When multiplied by the transformation coefficients of the target block, the residual samples of the target block can be derived.

[0168] In the example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 6 (b) The RST can be represented as the matrix operation shown in Equation 7.

[0169] [Formula 7]

[0170]

[0171] In Equation 7, c1 to c 16 The transformation coefficients of the target block can be represented. As a result of Equation 7, the transformation coefficients representing the modifications to the target block or the r values ​​of the residual samples of the target block can be derived. i And export r i The process can be shown in Equation 8.

[0172] [Formula 8]

[0173]

[0174] As a result of the calculation in Equation 8, the transformation coefficients representing the modification of the target block or the residual samples of the target block, r1 to r2, can be derived. NFrom the perspective of the size of the inverse transformation matrix, the size of the regular inverse transformation matrix is ​​64×64 (N×N), but the size of the inverse reduced transformation matrix is ​​reduced to 64×16 (R×N). Therefore, compared with performing the regular inverse transformation, the storage utilization rate of performing the inverse RST can be reduced by the R / N ratio. In addition, when comparing the number of multiplications N×N when using the regular inverse transformation matrix, using the inverse reduced transformation matrix can reduce the number of multiplications (N×R) by the R / N ratio.

[0175] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied based on the transform sets in Table 2. Since a transform set includes two or three transforms (kernels) depending on the intra-prediction mode, it can be configured to select one of up to four transforms, including those without applying secondary transforms. In the transforms without applying secondary transforms, the application of an identity matrix can be considered. Assuming indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is ​​applied, i.e., without applying secondary transforms), the transform index or lfnst index, which is a syntax element, can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the top-left 8×8 block, the 8×8 NSST in the RST configuration can be specified via the transform index, or the 8×8 lfnst can be specified when applying LFNST. 8×8 lfnst and 8×8 RST refer to transformations of 8×8 regions within a transform coefficient block when both W and H of the target block are equal to or greater than 8, and the 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to transformations of 4×4 regions within a transform coefficient block when both W and H of the target block are equal to or greater than 4, and the 4×4 region can be the top-left 4×4 region within the transform coefficient block.

[0176] According to embodiments of this disclosure, for the transformation during the encoding process, only 48 data points can be selected, and a maximum 16×48 transformation kernel matrix can be applied to them instead of applying a 16×64 transformation kernel matrix to the 64 data points forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transformation kernel matrix to generate m coefficients. That is, when performing RST by applying an m×48 transformation kernel matrix (m≤16) to an 8×8 region, 48 data points are input, and m coefficients are generated. When m is 16, 48 data points are input, and 16 coefficients are generated. That is, assuming 48 data points form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied sequentially, thereby generating a 16×1 vector. Here, the 48 data points forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 data points constituting the region other than the lower right 4×4 region within the 8×8 region. Here, when matrix operations are performed by applying a maximum 16×48 transformation kernel matrix, 16 modified transformation coefficients are generated. These 16 modified transformation coefficients can be arranged in the upper left 4×4 region according to the scan order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.

[0177] For the inverse transform in the decoding process, the transpose of the aforementioned transform kernel matrix can be used. That is, when performing inverse RST or LFNST during the inverse transform performed by the decoding device, the input coefficient data for applying inverse RST is arranged in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector with the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged into a two-dimensional block according to a predetermined arrangement order.

[0178] In summary, during the transformation process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on the 48 transform coefficients in the upper left, upper right, and lower left regions of the 8×8 region (excluding the lower right region) with a 16×48 transform kernel matrix. For matrix operations, the 48 transform coefficients are input as a one-dimensional array. When performing matrix operations, 16 modified transform coefficients are derived, and these modified coefficients can be arranged in the upper left region of the 8×8 region.

[0179] Conversely, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 region, the 16 transform coefficients corresponding to the upper left region of the 8×8 region can be input as a one-dimensional array according to the scan order, and matrix operations can be performed with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix, and therefore can be represented as an n×1 column vector. Furthermore, * denotes matrix multiplication. When performing matrix operations, 48 ​​modified transform coefficients can be derived, and these 48 modified transform coefficients can be arranged in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region.

[0180] When the secondary inverse transform is based on RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse reduced secondary transformer for deriving modified transform coefficients from the transform coefficients based on the inverse RST; and an inverse primary transformer for deriving residual samples for the target block from the modified transform coefficients based on the inverse primary transform. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In this disclosure, deriving transform coefficients based on a transform can mean deriving the transform coefficients by applying a transform.

[0181] The inseparable transform (LFNST) described above will be described in detail below. LFNST may include a forward transform performed by the encoding device and an inverse transform performed by the decoding device.

[0182] The encoding device receives the result (or part of the result) derived after applying the primary (core) transform as input and applies a forward secondary transform (secondary transform).

[0183] [Formula 9]

[0184]

[0185] In Equation 9, x and y are the input and output of the secondary transformation, respectively, and G is the matrix representing the secondary transformation, with the transformation basis vectors consisting of column vectors. In the case of inverse LFNST, when the dimension of the transformation matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of matrix G becomes Gin. T Dimensions.

[0186] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of the eight transformed basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.

[0187] On the other hand, for a positive LFNST, matrix G T The dimensions are [16×48], [8×48], [16×16], and [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper part of the [16×48] matrix and the [16×16] matrix, respectively.

[0188] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or an [8×1] vector can be used as output y. In video encoding and decoding, the output of the forward primary transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by properly arranging the 2D data that is the output of the forward transform.

[0189] Figure 7 This is a diagram illustrating the order in which the output data of the forward primary transform are arranged into a one-dimensional vector, according to the example. Figure 7 The left figures of (a) and (b) show the order used to construct the [48×1] vector, and Figure 7 The right figures (a) and (b) illustrate the order used to construct the [16×1] vector. In the case of LFNST, this can be achieved by combining 2D data with... Figure 7 Arrange the same order in (a) and (b) to obtain a one-dimensional vector x.

[0190] The orientation of the output data of the forward primary transform can be determined based on the intra-prediction mode of the current block. For example, when the intra-prediction mode of the current block is horizontal relative to the diagonal direction, the orientation can be determined by... Figure 7 The output data of the forward primary transform are arranged in the order of (a), and when the intra-prediction mode of the current block is perpendicular to the diagonal direction, it can be arranged according to... Figure 7 The output data of the forward primary transformation are arranged in the order of (b).

[0191] Based on the example, different methods can be applied. Figure 7 The order of (a) and (b), and in order to derive and apply Figure 7The order of (a) and (b) results in the same outcome (y vector), and the column vectors of matrix G can be rearranged according to the order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.

[0192] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data in the process of using the result of the forward secondary transformation as input (e.g., in the process of performing quantization or residual coding), the output y vector of Equation 9 needs to be properly arranged into 2D data again.

[0193] Figure 8 This is a diagram illustrating the order in which the output data of the forward secondary transformation is arranged into two-dimensional blocks, according to an example.

[0194] In the case of LFNST, the output values ​​can be arranged in 2D blocks according to a predetermined scan order. Figure 8 (a) shows how the output values ​​are arranged in 16 positions of a 2D block according to the diagonal scan order when the output y is a [16×1] vector. Figure 8 (b) shows that when the output y is an [8×1] vector, the output values ​​are arranged in 8 positions of the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 8 In (b), X indicates that it is filled with zeros.

[0195] According to another example, since the order in which the output vector y is processed during quantization or residual coding can be preset, the output vector y does not need to be arranged in a specific order. Figure 8 In the 2D block shown. However, in the case of residual coding, data encoding can be performed in 2D block (e.g., 4×4) cells (e.g., CG (coefficient group)), and in this case, according to as Figure 8 The data is arranged in a specific order within the diagonal scan sequence.

[0196] Simultaneously, the decoding device can configure the one-dimensional input vector y by arranging the two-dimensional data output from the dequantization process according to a preset scan order used for the inverse transform. The input vector y can be output as the output vector x using the following formula.

[0197] [Formula 10]

[0198]

[0199] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.

[0200] The output vector x is based on Figure 7 The sequence shown is arranged in a two-dimensional block and is arranged as two-dimensional data, which becomes the input data (or part of the input data) for the inverse primary transformation.

[0201] Therefore, the inverse secondary transformation is the reverse of the forward secondary transformation process as a whole, and in the case of the inverse transformation, unlike in the forward direction, the inverse secondary transformation is applied first, followed by the inverse primary transformation.

[0202] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be chosen as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.

[0203] Additionally, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can consist of two matrices. The choice of which of the four transform sets to use is determined based on the intra-prediction mode, and more specifically, based on the values ​​of the intra-prediction mode extended by taking into account wide-angle intra-prediction (WAIP). The choice of which matrix to select from the two matrices constituting the selected transform set is derived via index signaling. More specifically, 0, 1, and 2 can be used as index values ​​for transmission; 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra-prediction mode values.

[0204] Figure 9 This is a diagram illustrating a wide-angle intra-frame prediction mode according to an embodiment of this document.

[0205] Typical intra-prediction mode values ​​can have values ​​from 0 to 66 and from 81 to 83, and intra-prediction mode values ​​extended due to WAIP can have values ​​from -14 to 83 as shown. Values ​​from 81 to 83 indicate CCLM (Cross-Component Linear Model) mode, and values ​​from -14 to -1 and from 67 to 80 indicate intra-prediction mode extended due to WAIP application.

[0206] When predicting a block whose width is greater than its height, the top reference pixel is typically closer to the block's interior. Therefore, predictions in the bottom-left direction are more accurate than those in the top-right direction. Conversely, when the block's height is greater than its width, the left reference pixel is typically closer to the block's interior. Therefore, predictions in the top-right direction are more accurate than those in the bottom-left direction. Thus, applying remapping (i.e., mode index modification) to the index of the wide-angle intra-frame prediction mode can be advantageous.

[0207] When applying wide-angle intra-prediction, information about existing intra-prediction patterns can be sent using signals, and after the information is parsed, it can be remapped to the index of the wide-angle intra-prediction pattern. Therefore, the total number of intra-prediction patterns used for a specific block (e.g., a non-square block of a specific size) can remain unchanged; that is, the total number of intra-prediction patterns is 67, and the encoding of the intra-prediction patterns used for a specific block can remain unchanged.

[0208] Table 3 below illustrates the process of deriving the modified intra-frame mode by remapping the intra-frame prediction mode to the wide-angle intra-frame prediction mode.

[0209] [Table 3]

[0210]

[0211] In Table 3, the extended intra-prediction mode values ​​are ultimately stored in the `predModeIntra` variable, and `ISP_NO_SPLIT` indicates that the CU block is not divided into sub-partitions using the intra-segmentation (ISP) technique currently used in the VVC standard. The `cIdx` variable values ​​of 0, 1, and 2 indicate the cases for the luma, Cb, and Cr components, respectively. The `log2` function shown in Table 3 returns a base-2 log value, and the `Abs` function returns the absolute value.

[0212] The variable `predModeIntra`, which indicates the intra-prediction mode, along with the height and width of the transform block, are used as input values ​​for the wide-angle intra-prediction mode mapping process, and the output value is the modified intra-prediction mode `predModeIntra`. The height and width of the transform block or coded block can be the height and width of the current block used for intra-prediction mode remapping. In this case, the variable `whRatio`, which reflects the width-to-width ratio, can be set to `Abs(Log2(nW / nH))`.

[0213] For non-square blocks, the intra-prediction mode can be divided into two cases and modified accordingly.

[0214] First, if all conditions (1) to (3) are met, (1) the width of the current block is greater than its height, (2) the intra-prediction mode before modification is equal to or greater than 2, and (3) the intra-prediction mode is less than the value derived from (8+2*whRatio) when the variable whRatio is greater than 1 and less than 8 when the variable whRatio is less than or equal to 1 [predModeIntra is less than (whRatio > 1) ? (8 + 2 * whRatio) : 8], then the intra-prediction mode is set to a value 65 greater than the intra-prediction mode [predModeIntra is set to equal to (predModeIntra+65)].

[0215] If the above is different, that is, if the following conditions (1) to (3) are met, (1) the height of the current block is greater than the width, (2) the intra-prediction mode before modification is less than or equal to 66, and (3) the intra-prediction mode is greater than the value derived from (60-2*whRatio) when the variable whRatio is greater than 1 and is greater than 60 when the variable whRatio is less than or equal to 1 [predModeIntra is greater than (whRatio > 1) ? (60 − 2 * whRatio) : 60], then the intra-prediction mode is set to a value 67 smaller than the intra-prediction mode [predModeIntra is set to equal to (predModeIntra-67)].

[0216] Table 2 above illustrates how to select the transform set in LFNST based on intra-prediction mode values ​​extended by WAIP. For example... Figure 9 As shown, modes 14 to 33 and modes 35 to 80 are symmetrical about the prediction directions around mode 34. For example, modes 14 and 54 are symmetrical about the direction corresponding to mode 34. Therefore, the same set of transformations is applied to modes located in mutually symmetrical directions, and this symmetry is also reflected in Table 2.

[0217] Meanwhile, it is assumed that the positive LFNST input data of mode 54 is symmetrical to the positive LFNST input data of mode 14. For example, for modes 14 and 54, according to Figure 7 (a) and Figure 7 The arrangement shown in (b) rearranges the two-dimensional data into one-dimensional data. Furthermore, it can be seen that... Figure 7 (a) and Figure 7 The pattern in the sequence shown in (b) is symmetrical about the direction indicated by pattern 34 (diagonal direction).

[0218] Meanwhile, as mentioned above, the size and shape of the target block determine which transformation matrix, either the [48×16] matrix or the [16×16] matrix, will be applied to the LFNST.

[0219] Figure 10 This is a diagram illustrating the block shape to which LFNST is applied. Figure 10 (a) shows a 4×4 block, (b) shows a 4×8 block and an 8×4 block, (c) shows a 4×N block or an N×4 block, where N is 16 or greater, (d) shows an 8×8 block, and (e) shows an M×N block, where M≥8, N≥8 and N>8 or M>8.

[0220] exist Figure 10 In the diagram, blocks with thick boundaries indicate the area where LFNST is applied. For Figure 10For blocks (a) and (b), LFNST is applied to the top-left 4×4 region, and for Figure 10 In block (c), LFNST is applied individually to two consecutive top-left 4×4 regions. Figure 10 In (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, it will be referred to as "4×4 LFNST" in the following text. A [16×16] or [16×8] matrix can be applied based on the matrix dimension of G used in Equations 9 and 10.

[0221] More specifically, a [16×8] matrix is ​​applied to Figure 10 (a) 4×4 blocks (4×4 TU or 4×4 CU), and a [16×16] matrix is ​​applied to Figure 10 The blocks in (b) and (c). This is to adjust the worst-case computational complexity to 8 multiplications per sample.

[0222] about Figure 10 In (d) and (e), LFNST is applied to the top-left 8×8 region, and this LFNST is referred to as "8×8 LFNST" below. As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since the [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values ​​from the top-left 8×8 region are used as input values ​​for the forward LFNST. That is, as can be obtained from... Figure 7 (a) left-hand order or Figure 7 As can be seen from the left-hand order in (b), a [48×1] vector can be constructed based on samples belonging to the other three 4×4 blocks while leaving the bottom right 4×4 block as is.

[0223] A [48×8] matrix can be applied to Figure 10 (d) 8×8 blocks (8×8 TU or 8×8 CU), and a [48×16] matrix can be applied. Figure 10 The 8×8 block in (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.

[0224] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 9, [8×1] or [16×1] vectors) are generated. In the forward LFNST, due to matrix G... T Due to its characteristic, the amount of output data is equal to or less than the amount of input data.

[0225] Figure 11This is a diagram illustrating the arrangement of the output data of the forward LFNST according to an example, and showing the blocks of the output data of the forward LFNST arranged according to the block shape.

[0226] exist Figure 11 The shaded area in the upper left corner of the block shown corresponds to the region where the output data of the forward LFNST is located. The positions marked with 0 indicate samples filled with 0 values, and the remaining areas represent regions that were not altered by the forward LFNST. In regions not altered by LFNST, the output data of the forward primary transform remains unchanged.

[0227] As mentioned above, since the dimensions of the applied transformation matrix vary depending on the shape of the block, the amount of output data also varies. Figure 11 The output data of a forward LFNST may not completely fill the top-left 4×4 block. Figure 11 In cases (a) and (d), the [16×8] matrix and the [48×8] matrix are applied to the block indicated by the thick line or a portion of the area inside the block, respectively, and an [8×1] vector is generated as the output of the positive LFNST. That is, according to Figure 8 The scan order shown in (b) can fill only 8 output data, such as Figure 11 As shown in (a) and (d), zeros can be filled in the remaining 8 positions. Figure 10 (d) In the case of the LFNST application block, such as Figure 11 As shown in (d), the two 4×4 blocks adjacent to the top-left 4×4 block, the top-right and bottom-left blocks, are also filled with 0 values.

[0228] As described above, essentially, by sending the LFNST index via a signal, it is specified whether LFNST is applied and the transformation matrix to be applied. Figure 11 As shown, when LFNST is applied, since the number of output data of the positive LFNST can be equal to or less than the number of input data, the following area filled with zero values ​​appears.

[0229] 1) such as Figure 11 As shown in (a), the samples in the top left 4×4 block are those from the eighth position onwards in the scanning order, i.e., from the ninth to the sixteenth.

[0230] 2) such as Figure 11 As shown in (d) and (e), when applying a [48×16] matrix or an [8×8] matrix, the two 4×4 blocks adjacent to the top left 4×4 block or the second and third 4×4 blocks in the scan order.

[0231] Therefore, if non-zero data is found in regions 1) and 2), it is determined that LFNST has not been applied, so the signaling for the corresponding LFNST index can be omitted.

[0232] Meanwhile, the following simplification method can be applied to the LFNST used.

[0233] (i) As shown in the example, the number of output data for a forward LFNST can be limited to a maximum of 16.

[0234] exist Figure 10 In case (c), the 4×4 LFNST can be applied to two adjacent 4×4 regions, and in this case, a maximum of 32 LFNST output data can be generated. When the number of positive LFNST output data is limited to a maximum of 16, in the case of 4×N / N×4 (N≥16) blocks (TU or CU), the 4×4 LFNST is applied only to one 4×4 region in the upper left, and the LFNST can be applied only to... Figure 10 All blocks are processed at once. This simplifies the implementation of image encoding.

[0235] Figure 12 The example shows that the amount of output data used for the positive LFNST is limited to a maximum of 16. (See example.) Figure 12 When LFNST is applied to the top left 4x4 region of a 4xN or Nx4 block where N is 16 or greater, the output data of the positive LFNST becomes 16.

[0236] (ii) As in the example, zeroing can be additionally applied to regions to which LFNST has not been applied. In this document, zeroing can mean filling all positions belonging to a particular region with a value of 0. That is, zeroing can be applied to regions that have not changed due to LFNST and maintain the result of the positive primary transformation. As mentioned above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, zeroing can be divided into two types as follows ((ii)-(A) and (ii)-(B)).

[0237] (ii)-(A) When 4×4 LFNST is applied, areas to which 4×4 LFNST is not applied can be zeroed. Figure 13 This is a diagram illustrating the zeroing of a block in which a 4×4 LFNST is applied, based on an example.

[0238] like Figure 13 As shown, regarding the block to which 4×4 LFNST was applied, that is, for Figure 11 All blocks in (a), (b) and (c) whose entire regions are not subject to LFNST can be filled with zeros.

[0239] on the other hand, Figure 13 (d) shows when... Figure 12When the maximum number of output data for the positive LFNST shown is limited to 16, the remaining blocks that have not been subjected to the 4×4 LFNST are zeroed out.

[0240] (ii)-(B) When 8×8 LFNST is applied, areas where 8×8 LFNST is not applied can be zeroed. Figure 14 This is a diagram illustrating the zeroing process in an 8×8 LFNST block, based on the example application.

[0241] like Figure 14 As shown, regarding the application of 8×8 LFNST to the block, that is, for Figure 11 In all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.

[0242] (iii) Due to the zeroing presented in (ii) above, the zero-filled areas may differ when LFNST is applied. Therefore, it is possible to compare... Figure 11 In the case of LFNST, a wider area is used to perform zeroing as proposed in (ii) to check for the presence of non-zero data.

[0243] For example, when (ii)-(B) is applied, in the examination Figure 11 After checking whether there is non-zero data in the zero-filled regions in (d) and (e), additional checks are performed. Figure 14 The presence of non-zero data in the region filled with 0s can be checked, and signaling for the LFNST index can be executed only if no non-zero data exists.

[0244] Of course, even with the zeroing proposed in application (ii), the existence of non-zero data can be checked in the same way as existing LFNST index signaling. That is, when checking... Figure 11 After confirming the presence of non-zero data within the zero-padded block, LFNST index signaling can be applied. In this case, the encoding device only performs zeroing and the decoding device does not assume zeroing; that is, it only checks whether non-zero data exists within the zero-padded block. Figure 11 In regions explicitly marked as 0, LFNST index resolution can be performed.

[0245] Various embodiments of combinations of simplified methods for applying LFNST ((i), (ii)-(A), (ii)-(B), (iii)) can be derived. Of course, the combinations of the above simplified methods are not limited to the following embodiments, and any combination can be applied to LFNST.

[0246] Example

[0247] - Limit the number of output data for the positive LFNST to a maximum of 16 (i).

[0248] - When 4×4 LFNST is applied, all areas where 4×4 LFNST is not applied are zeroed (ii)- (A)

[0249] - When 8×8 LFNST is applied, all areas where 8×8 LFNST is not applied are zeroed (ii)- (B)

[0250] - After checking whether non-zero data also exists in the existing area filled with zero values ​​and in the area filled with zero due to additional clearing ((ii)-(A), (ii)-(B)), the LFNST index is signaled only if no non-zero data exists (iii).

[0251] In the implementation example, when LFNST is applied, the region containing non-zero output data is limited to the upper left 4×4 region. More specifically, in Figure 13 (a) and Figure 14 In case (a), the eighth position in the scan order is the last position where non-zero data can exist. Figure 13 (b) and (c) and Figure 14 In case (b), the sixteenth position in the scan order (i.e., the position at the bottom right edge of the top left 4×4 block) is the last position where data other than 0 can exist.

[0252] Therefore, when applying LFNST, after checking whether non-zero data exists at a position where the residual coding process is not allowed (a position beyond the last position), it can be determined whether to send the LFNST index with a signal.

[0253] In the case of the zeroing method proposed in (ii), the computational cost required to perform the entire transformation process can be reduced because the amount of data ultimately generated when both the primary transform and LFNST are applied is reduced. That is, when LFNST is applied, since zeroing is applied to the forward primary transform output data existing in the region where LFNST is not applied, it is not necessary to generate data for the region that is zeroed during the performance of the forward primary transform. Therefore, the computational cost required to generate the corresponding data can be reduced. The additional effects of the zeroing method proposed in (ii) are summarized below.

[0254] First, as mentioned above, reduce the amount of computation required to perform the entire transformation process.

[0255] Specifically, when (ii)-(B) is applied, the worst-case computational cost is reduced, making the transformation process lightweight. In other words, in general, a large amount of computation is required to perform large-scale primary transformations. By applying (ii)-(B), the amount of data derived as a result of performing a forward LFNST can be reduced to 16 or less. Furthermore, the effect of reducing the number of transformation operations increases further as the size of the entire block (TU or CU) increases.

[0256] Secondly, it can reduce the amount of computation required for the entire transformation process, thereby reducing the power consumption required to perform the transformation.

[0257] Third, it reduces the delay involved in the transformation process.

[0258] Secondary transforms, such as LFNST, add computational complexity to existing primary transforms, thus increasing the overall latency involved in performing the transforms. Specifically, in the case of intra-frame prediction, the increased latency due to secondary transforms during encoding leads to an increase in latency until reconstruction because reconstructed data from neighboring blocks is used during prediction. This can result in an increase in the overall latency of intra-frame predictive coding.

[0259] However, if the zeroing suggested in application (ii) is applied, the delay time for performing the primary transformation can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transformation, making it easier to implement the encoding device.

[0260] Meanwhile, in traditional intra-frame prediction, the target block to be encoded is treated as a single coding unit and encoding is performed without partitioning. However, ISP (Intra-Frame Sub-Partition) coding refers to performing intra-frame prediction coding by partitioning the target block into blocks in the horizontal or vertical direction. In this case, reconstructed blocks can be generated by performing encoding / decoding on a block-by-block basis, and the reconstructed blocks can be used as reference blocks for the next partitioned block. For example, in ISP coding, a coding block can be partitioned into two or four sub-blocks and encoded, and in ISP, intra-frame prediction is performed on a sub-block by referencing the reconstructed pixel values ​​of the sub-blocks located to its left or top. In the following text, the term "coding" can be used as a concept encompassing both encoding performed by the encoding device and decoding performed by the decoding device.

[0261] The following describes an example of applying LFNST only to the luminance component in a single tree.

[0262] The following is a syntax table of coding units related to signaling with LFNST and MTS indices, based on an example.

[0263] [Table 4]

[0264]

[0265] The meanings of the main variables in the table above are as follows.

[0266] 1. cbWidth, cbHeight: Width and height of the current coding block.

[0267] 2. log2TbWidth, log2TbHeight: The base-2 logarithmic values ​​of the width and height of the current transform block. By applying zeros, the size of the current transform block can be reduced to the upper-left region where non-zero coefficients may exist.

[0268] 3. `sps_lfnst_enabled_flag`: A flag indicating whether LFNST is enabled. A value of 0 indicates that LFNST is not enabled, and a value of 1 indicates that LFNST is enabled. This flag is defined in the Sequence Parameter Set (SPS).

[0269] 4. `CuPredMode[chType][x0][y0]`: The prediction mode corresponding to the variable `chType` and the position (x0, y0). `chType` can have values ​​of 0 and 1, where 0 represents the luma component and 1 represents the chroma component. The position (x0, y0) indicates the location on the image, and using the value of `CuPredMode[chType][x0][y0]`, `MODE_INTRA` (intra-frame prediction) and `MODE_INTER` (inter-frame prediction) are possible.

[0270] 5. IntraSubPartitionsSplitType: Indicates which ISP is applied to the current coding unit, and ISP_NO_SPLIT indicates that the coding unit is not split into partitions.

[0271] 6. intra_mip_flag[ x0 ][ y0 ]: Position (x0 , y0) as described in section 4 above. intra_mip_flag is a flag indicating whether matrix-based intra-prediction (MIP) mode is applied. A value of 0 indicates that MIP is not applied, and a value of 1 indicates that MIP is applied.

[0272] 7.cIdx: A value of 0 indicates luminance, while values ​​of 1 and 2 indicate the chromaticity components Cb and Cr, respectively.

[0273] 8. treeType: Indicates whether it is a single-tree or dual-tree configuration. (SINGLE_TREE: Single-tree; DUAL_TREE_LUMA: Dual-tree for the luminance component; DUAL_TREE_CHROMA: Dual-tree for the chrominance component)

[0274] 9.lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If it is not parsed, the element is inferred to have a value of 0. That is, the default value is set to 0, which indicates that LFNST is not applied.

[0275] The descriptions of the aforementioned syntax elements can be applied to the syntax elements shown in the table below.

[0276] In Table 4, transform_skip_flag[x0][y0][0]==0 is a condition used to determine whether to signal the lfnst index relative to the luminance component based on whether to skip the transformation.

[0277] Therefore, based on the example, the following coding unit syntax table is proposed to remove the correlation between the signaling for the transform skip flag for the luminance component and the signaling for the LFNST index for the chrominance component.

[0278] [Table 5]

[0279]

[0280] In the embodiments shown in Table 5, the signaling for the LFNST index of the luma component depends only on the transform skip flag of the luma component for both dual-tree and single-tree partitioning modes. In dual-tree mode, the LFNST index of the chroma component can be signaled based solely on the transform skip flag of the chroma component. In single-tree partitioning mode, LFNST is not applied to the chroma component to reduce worst-case latency.

[0281] The variable LfnstTransformNotSkipFlag shown in Table 5 is set according to the tree type of the current block and the transformation skip flag value used for the color components, and the LFNST index can only be signaled when the value of the variable is 1.

[0282] When the tree type is not dual-tree chroma (treeType != DUAL_TREE_CHROMA), that is, when the tree type is single-tree or dual-tree luminance, if the transform skip flag used for the luminance component is 0, the variable LfnstTransformNotSkipFlag can be set to 1. treeType != DUAL_TREE_CHROMA ? transform_skip_flag[ x0 ][ y0 ][ 0 ] == 0: (transform_skip_flag[ x0 ][ y0 ][ 1 ] == 0 || transform_skip_flag[ x0 ][ y0 ][ 2 ] == 0)), and when the tree type is dual-tree chroma, if the transform skip flag value (transform_skip_flag[ x0 ][ y0 ][ 1 ]) for chroma component Cb is 0 or the transform skip flag value (transform_skip_flag[ x0][y0][1]) for chroma component Cr is 0, then the variable LfnstTransformNotSkipFlag can be set to 1 ( treeType != DUAL_TREE_CHROMA ? transform_skip_flag[ x0 ][ y0 ][ 0 ] == 0 : (transform_skip_flag[ x0 ][ y0 ] [ 1 ] = = 0 || transform_skip_flag[ x0 ][ y0 ][ 2 ] = = 0 ) ).

[0283] In this disclosure, the operator “x ?y : z” indicates that if x is true, then x is y, otherwise x is z (if x is true, it is evaluated as the value of y, otherwise it is evaluated as the value of z).

[0284] The standard text for the transformation process in Table 5 is as follows.

[0285] [Table 6]

[0286]

[0287]

[0288] If a valid coefficient exists at the zeroing position when applying LFNST, the variable LfnstZeroOutSigCoeffFlag in Table 4 is 0; otherwise, it is 1. The variable LfnstZeroOutSigCoeffFlag can be set according to several conditions shown in Table 11 below.

[0289] The variable `LfnstZeroOutSigCoeffFlag` indicates whether a valid coefficient exists in a second region besides the top-left first region of the current block. The variable's value is initially set to 1 and can be changed to 0 if a valid coefficient exists in the second region. The LFNST index can be resolved only while maintaining the initial setting of the variable `LfnstZeroOutSigCoeffFlag`. When it is determined and the value of the variable `LfnstZeroOutSigCoeffFlag` is 1, LFNST can be applied to the luma component or all chroma components of the current block, thus the color index of the current block is uncertain.

[0290] According to the example, when all the last valid coefficients of a transform block with a corresponding coded block flag (CBF, 1 if at least one valid coefficient exists in the corresponding block, 0 otherwise) set to 1 are in the DC position (top left position), the variable LfnstDcOnly in Table 4 is 1; otherwise, it is 0. Specifically, the position of the last valid coefficients is checked relative to one luminance transform block in a two-tree luminance system, and relative to both the Cb and Cr transform blocks in a two-tree chrominance system. In a single-tree system, the position of the last valid coefficients can be checked relative to the luminance, Cb, and Cr transform blocks.

[0291] The syntax table for signaling the encoding unit of the LFNST index, based on another example, is as follows.

[0292] [Table 7]

[0293]

[0294] In Table 7, the variable `transform_skip_flag[x0][y0][cIdx]` indicates whether transform skipping is applied to the coded block for the component indicated by `cIdx`. `cIdx` can have values ​​of 0, 1, and 2, where 0 indicates the luminance component, and 1 and 2 indicate the Cb and Cr components, respectively. A value of 1 for `transform_skip_flag[x0][y0][cIdx]` indicates that transform skipping is applied, while 0 indicates that transform skipping is not applied.

[0295] In Table 7, the variable LfnstNotSkipFlag can be set to 1 only when the transform is not applied to all components (all coding blocks) forming the current coding unit, and can be set to 0 in other cases. The LFNST index (lfnst_idx in Table 7) is signaled only when LfnstNotSkipFlag is 1.

[0296] When the current coding unit is encoded in a single-tree structure (when treeType is SINGLE_TREE in Table 7), all components include Y, Cb, and Cr. When the current coding unit is encoded in a separate tree structure for luma (when treeType is DUAL_TREE_LUMA in Table 7), all components include only Y. When the current coding unit is encoded in a separate tree structure for chroma (when treeType is DUAL_TREE_CHROMA in Table 7), all components include Cb and Cr.

[0297] In other words, even if one of the components consisting of the current coding unit skips coding through a transform, no signal is sent to the LFNST index, and the value of the LFNST index is inferred to be 0. That is, LFNST is not applied.

[0298] In structures where LFNST is not applied when even one of the components (Y, Cb, and Cr) is encoded via transform skip as shown in Table 7, when multiple components (Y, Cb, and Cr) are encoded consecutively in a single tree and it is determined during parsing that a component has undergone transform skip (e.g., when transform skip is determined for Cb and Cr), the corresponding transform coefficients can be configured not to be additionally buffered with respect to the corresponding component encoded via transform skip until the value of the LFNST index is parsed.

[0299] For example, when it is found that any component is skipped in encoding by a transform, it is determined that LFNST should not be applied, so inverse quantization, inverse transform, etc. can be performed immediately thereafter.

[0300] Instead of Table 7, the grammar table can be described as concisely as in Table 8.

[0301] [Table 8]

[0302]

[0303] When determining whether to signal the LFNST index by checking only whether the transform skip is applied to the luminance component in a single tree, the syntax table for the encoding unit can be configured as follows.

[0304] [Table 9]

[0305]

[0306] In Table 9, LfnstNotSkipFlag is determined to be the same as in Table 7 or Table 8 in the individual trees (i.e., in the luma-only tree, LfnstNotSkipFlag is set to 1 when no transform skip is applied to the luma component, and 0 otherwise; while in the chroma-only tree, LfnstNotSkipFlag is set to 1 when no transform skip is applied to both the Cb and Cr components, and 0 otherwise). In the single tree, LfnstNotSkipFlag is set to 1 when no transform skip is applied to the luma-only component, and 0 otherwise. Instead of Table 8, the syntax table shown in Table 9 can be applied.

[0307] [Table 10]

[0308]

[0309] The following describes an example of determining LFNST index signaling conditions when LFNST is applied only to the luminance component in a single tree.

[0310] In the coding unit syntax tables of Tables 4 and 5, 7, 8, 9, and 10, the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag are used as conditions for signaling the LFNST index. Essentially, both LfnstDcOnly and LfnstZeroOutSigCoeffFlag are initialized to 1, as shown in Table 7, and their values ​​can be updated to 0 in the syntax table used for residual coding, as shown in Table 11. For reference, when skipping the encoding of components (which can be Y, Cb, or Cr) using a transform, a different syntax table (transform_ts_coding) is imported instead of the residual coding in Table 11; therefore, when parsing the LFNST index for the component, the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag are not updated.

[0311] [Table 11]

[0312]

[0313]

[0314] In Table 11, lastSubBlock indicates the position of the subblock (coefficient group (CG)), where the last valid (non-zero) coefficient is located in scan order. 0 indicates a subblock that includes the DC component, while a value greater than 0 indicates a subblock that does not include the DC component.

[0315] `lastScanPos` indicates the position of the last valid coefficient within a sub-block in scan order. When a sub-block contains 16 positions, values ​​from 0 to 15 are possible.

[0316] LastSignificantCoeffX and LastSignificantCoeffY indicate the x and y coordinates of the last effective coefficient located in the transform block. The x-coordinate starts from 0 and increases from left to right, and the y-coordinate starts from 0 and increases from top to bottom. A value of 0 for both variables means that the last effective coefficient is located at DC.

[0317] Essentially, Table 11 is applied to the embodiment in Table 5 to determine the values ​​of the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag. When Table 11 is applied and the coded blocks are encoded in a single tree, the residual coding presented in Table 11 can be imported for all components. For example, residual coding can be performed on each component when all Y, Cb, and Cr components are skipped from encoding without using a transform.

[0318] Therefore, when applying Table 11 and encoding the current coding block using a single tree, if the last non-zero coefficient of even a component is located outside the DC position (the top left position of the transform block), the value of the variable LfnstDcOnly can be updated to 0. Furthermore, if the last non-zero coefficient of even a component is located in a region where the transform coefficients cannot be located when LFNST is applied (i.e., in a region outside the first to eighth positions in a 4×4 or 8×8 transform block according to the forward transform coefficient scan order, or in a region outside the top left 4×4 region in the transform block to which LFNST applies in the current VVC standard), the value of the variable LfnstZeroOutSigCoeffFlag can be updated to 0. As shown in the encoding unit syntax tables in Tables 5, 7, 8, 9 and 10, the LFNST index can be signaled only when the value of the variable LfnstZeroOutSigCoeffFlag is 1, while in modes other than ISP mode, the LFNST index can be signaled only when the value of the variable LfnstDcOnly is 0.

[0319] However, when LFNST is applied only to the luminance component in a single tree, updates to the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag may not be permitted when residual coding is performed on components for which LFNST has not been applied (chrominance components Cb or Cr). This is because it is logically inappropriate not to determine whether to signal the LFNST index (i.e., whether LFNST is applied) based on the arrangement or distribution of the transform coefficients of the components for which LFNST has been applied.

[0320] Table 12 restricts the updates of variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag to the luminance component only within a single tree. In cases other than a single tree, variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag can be updated for all components (as shown in Table 11).

[0321] [Table 12]

[0322]

[0323]

[0324] In the residual coding presented in Table 12, since treeType is added as a parameter compared to Table 11, the syntax table used for the transformation unit can be modified as shown in Table 13.

[0325] [Table 13]

[0326]

[0327] Based on the details in Table 5, some details can be replaced by the embodiments in Tables 7 through 10, or the details in Table 11 or 12 can be applied. The following possible combinations can be configured based on Tables 7 through 10, Table 11, or Table 12.

[0328] 1. Table 7 (or Table 8) + Table 11

[0329] 2. Table 7 (or Table 8) + Table 12

[0330] 3. Table 9 (or Table 10) + Table 11

[0331] 4. Table 9 (or Table 10) + Table 12

[0332] The following describes a method for applying a scaling list to the chromaticity component when LFNST is applied only to the luminance component in a single tree.

[0333] Currently, in VVC WD, a syntax element called `scaling_matrix_for_lfnst_disabled_flag` is defined. When `scaling_matrix_for_lfnst_disabled_flag` is 1, the scaling list is not applied when applying LFNST, while when `scaling_matrix_for_lfnst_disabled_flag` is 0, the scaling list can be applied when applying LFNST.

[0334] Here, the scaling list is a matrix used to specify specific weights (weights) for each transform coefficient position in the transform block, and dequantization or quantization is achieved by multiplying by the weights for each transform coefficient, thereby enabling the application of differentiated dequantization or quantization based on the importance of the transform coefficient.

[0335] In a single tree, as in the embodiments in Table 5, LFNST can be applied only to the luma component, and the scaling list is not applied to the luma component when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST is applied while encoding the coded block in the single tree. Here, the scaling list can be applied to the chroma component where LFNST is not applied.

[0336] Table 14 shows an example of the dequantization process (scaling process) to achieve the above situation.

[0337] [Table 14]

[0338]

[0339]

[0340]

[0341]

[0342] In Table 14, treeType indicates the tree type of the coding unit to which the currently processed transform block belongs, and single_tree, dual_tree_luma, and dual_tree_chroma indicate a single tree, a single tree for luma, and a single tree for chroma (dual-tree chroma), respectively.

[0343] In this embodiment, since LFNST can be applied only to the luminance component in a single tree, the scaling list is not applied to the luminance component when scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST is applied (lfnst_idx[xTby][yTby] is greater than 0), and (when cIdx is 0).

[0344] However, for the chroma component (when the cIdx value is greater than 0), it can be determined whether to apply the scaling list by further checking different conditions (e.g., checking transform_skip_flag[xTby][yTby][cIdx]).

[0345] In a single tree, such as in the case of the luminance component in a single tree, when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST is applied (lfnst_idx[xTby][yTby] value is greater than 0), the scaling list is not applied to the luminance and chrominance components.

[0346] Furthermore, in individual trees, such as in the case of chroma components in a single tree, it can be determined whether to apply the scaling list by further examining different conditions (e.g., examining transform_skip_flag[xTby][yTby][cIdx]).

[0347] Therefore, when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST can be applied only to the luminance component in a single tree, the scaling list can be applied to the chrominance component but not to the luminance component.

[0348] According to the example, a combination of Table 14 and the above embodiments can be applied (based on the details of Table 5, some details can be replaced with the embodiments of Tables 7 to 10, or a combination of the details of Table 11 or Table 12 can be applied).

[0349] In this case, as described in the specification text “Transformation process for scaling transform coefficients” in Table 6, LFNST can be configured to be applied only to the luminance component in a single tree.

[0350] The following figures are provided to illustrate specific examples of this disclosure. Because the figures are provided for illustration purposes using specific terminology for apparatus or for signals / messages / fields, the technical features of this disclosure are not limited to the specific terminology used in the following figures.

[0351] Figure 15 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.

[0352] Figure 15 Each process disclosed in the document is based on a reference. Figures 4 to 14 Some details described. Therefore, with reference Figures 3 to 14 The description of overlapping specific details will be omitted or will be shown schematically.

[0353] According to the embodiment, the decoding device 300 can receive from the bitstream flag information indicating whether the scaling list is available when LFNST is performed, the LFNST index for the current block, and residual information (S1510).

[0354] Specifically, the decoding device 300 can decode information about the quantization transform coefficients used for the current block from the bitstream, and can derive the quantization transform coefficients of the target block based on the information about the quantization transform coefficients used for the current block. The information about the quantization transform coefficients of the target block can be included in the Sequence Parameter Set (SPS) or slice header, and can include at least one of the following: information about whether RST is applied, information about the reduction factor, information about the minimum transform size for applying RST, information about the maximum transform size for applying RST, the inverse RST size, and information about the transform index indicating any transform kernel matrix included in the transform set.

[0355] The decoding device can further receive information about the intra-prediction mode used for the current block and information about whether ISP is applied to the current block. The decoding device can receive and parse flag information indicating whether ISP coding or the ISP mode is applied, thereby deriving whether the current block is divided into a predetermined number of sub-partition transform blocks. Here, the current block can be a coded block. Furthermore, the decoding device can derive the size and number of the sub-partition blocks by flag information indicating the direction in which the current block is divided.

[0356] The LFNST index is used to specify the value of the LFNST matrix when LFNST is applied as the inverse secondary inseparable transformation, and can have values ​​ranging from 0 to 2. For example, an LFNST index value of 0 can indicate that no LFNST is applied to the current block, an LFNST index value of 1 can indicate the first LFNST matrix, and an LFNST index value of 2 can indicate the second LFNST matrix.

[0357] Information about the ISP and LFNST indices can be received at the coding unit level.

[0358] The flag information received by the decoding device indicating whether the scaling list is available when performing LFNST can be represented by `scaling_matrix_for_lfnst_disabled_flag` or `sps_scaling_matrix_for_lfnst_disabled_flag`, and can be sent as a signal in the sequence parameter set. A value of 1 indicates that the scaling list is not applied when LFNST is applied, and a value of 0 indicates that the scaling list is applicable when LFNST is applied. The scaling list is a matrix used to assign specific weights (weighting values) to each transform coefficient position in the transform block, and can be dequantized or quantized by multiplying it by the weights used for each transform coefficient, thereby enabling differential dequantization or quantization to be applied according to the importance of the transform coefficient.

[0359] The decoding device 300 can determine whether the scaling list is applied to the current block based on whether LFNST is applied and the tree type of the current block, so as to dequantize the transform coefficients used for the current block (S1520).

[0360] Whether to apply the scaling list can be determined based on the flag information and the value of the LFNST index.

[0361] When the current block's tree type is single-tree, the color components of the current block may include a luminance component, a first chromaticity component indicating chromaticity Cb, and a second chromaticity component indicating chromaticity Cr. When the current block's tree type is two-tree luminance, the current block may include a luminance component. When the current block's tree type is two-tree chromaticity, the color components of the current block may include a first chromaticity component and a second chromaticity component.

[0362] Here, the current block can be a transform block, which is a transform unit. When the tree type of the current block is single-tree, the current block can include transform blocks for the luma component, transform blocks for the first chroma component, and transform blocks for the second chroma component. When the tree type of the current block is two-tree luma, the current block can include transform blocks for the luma component. And when the tree type of the current block is two-tree chroma, the current block can include transform blocks for the first chroma component and transform blocks for the second chroma component.

[0363] According to the example, when the current block type is single-tree, LFNST can be applied only to the luma component, and when the current block is encoded in a single-tree, the value of scaling_matrix_for_lfnst_disabled_flag is 1, and LFNST is applied; the scaling list is not applied to the luma component. However, the scaling list can be applied to the chroma component where LFNST is not applied.

[0364] In summary, when the flag information regarding the scaling list indicates that the scaling list is unavailable and the LFNST index is greater than 0 (i.e., LFNST is applied), the scaling list may not be applied when the tree type of the current block is single-tree and the current block is a luma component, and the scaling list may be applied when the tree type of the current block is single-tree and the current block is a chroma component.

[0365] According to the example, if the flag information regarding the scaling list indicates that the scaling list is unavailable and the LFNST index is greater than 0, then the LFNST can be applied to the current block when the tree type of the current block is dual-tree chroma, and therefore the scaling list is not applied to the chroma components.

[0366] According to the example, if the flag information regarding the scaling list indicates that the scaling list is unavailable and the LFNST index is greater than 0, then when the tree type of the current block is dual-tree luminance, the LFNST can be applied to the current block, and therefore the scaling list is not applied to the luminance component.

[0367] Subsequently, the decoding device derives the transform coefficients for the current block from the residual information based on the determined result (S1530).

[0368] The derived transform coefficients can be arranged in 4×4 blocks according to a reverse diagonal scan order, and the transform coefficients in the 4×4 blocks can also be arranged according to a reverse diagonal scan order. In other words, the dequantized transform coefficients can be arranged according to the reverse scan order applied in a video codec, such as in VVC or HEVC.

[0369] The decoding device can derive modified transform coefficients from the transform coefficients based on the LFNST matrix and LFNST index used for LFNST, that is, by applying LFNST (S1540).

[0370] LFNST is an inseparable transform where the transform is applied to the coefficients without separating them in a specific direction, unlike a primary transform where the coefficients to be transformed are separated vertically or horizontally and then transformed. This inseparable transform can be a low-frequency inseparable transform that applies the forward transform only to the low-frequency region rather than the entire region of the block.

[0371] The decoding device can derive various variables to apply LFNST, and can determine whether to apply LFNST based on the tree type and size of the current block.

[0372] The decoding device can derive a first variable (variable LfnstDcOnly) indicating whether a valid coefficient exists at a location outside the DC component in the current block and a second variable (variable LfnstZeroOutSigCoeffFlag) indicating whether a transform coefficient exists in a second region outside the first region at the top left of the current block.

[0373] The first and second variables are initially set to 1, wherein the first variable can be updated to 0 when the effective coefficients are located outside the location of the DC component in the current block, and the second variable can be updated to 0 when the transform coefficients are located in the second region.

[0374] When the first variable is updated to 0 and the second variable remains 1, LFNST can be applied to the current block.

[0375] For the luminance component applicable to Intra-Frame Sub-Partition (ISP) mode, the LFNST index can be parsed without deriving the variable LfnstDcOnly.

[0376] Specifically, when the ISP mode is applied and the transform skip flag for the luminance component (i.e., transform_skip_flag[x0][y0][0]) is 0, the LFNST index can be signaled regardless of the value of the variable LfnstDcOnly, when the tree type of the current block is a single tree or a double tree for luminance.

[0377] However, for chrominance components that do not apply ISP mode, the value of the variable LfnstDcOnly can be set to 0 based on the values ​​of the transform skip flags for chrominance component Cb (transform_skip_flag[x0][y0][1]) and chrominance component Cr (transform_skip_flag[x0][y0][2]). That is, when the value of cIdx in transform_skip_flag[x0][y0][cIdx] is 1, the value of the variable LfnstDcOnly can be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0; and when the value of cIdx is 2, the value of the variable LfnstDcOnly can be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0. When the value of the variable LfnstDcOnly is 0, the decoding device can resolve the LFNST index; otherwise, the LFNST index can be inferred as 0 without signaling.

[0378] The second variable can be LfnstZeroOutSigCoeffFlag, which indicates that zeroing should be performed when LFNST is applied. The second variable can be initially set to 1 and can be changed to 0 when valid coefficients are present in the second region.

[0379] The variable LfnstZeroOutSigCoeffFlag can be exported as 0 when the index of the sub-block with the last non-zero coefficient is greater than 0 and the width and height of the transform block are both equal to or greater than 4, or when the position of the last non-zero coefficient in the sub-block with the last non-zero coefficient is greater than 7 and the size of the transform block is 4×4 or 8×8. A sub-block refers to a 4×4 block used as a coding unit in residual coding and can be called a coefficient group (CG). Sub-block index 0 refers to the top-left 4×4 sub-block.

[0380] In other words, when non-zero coefficients are derived in a region other than the top left region (where LFNST transform coefficients may exist in the transform block), or when non-zero coefficients exist in a 4×4 block or 8×8 block at a position other than the eighth position of the scan order, the variable LfnstZeroOutSigCoeffFlag is set to 0.

[0381] The decoding device can determine the set of LFNSTs, including the LFNST matrix, based on the intra-prediction mode derived from information about the intra-prediction mode, and can select any one of the multiple LFNST matrices based on the LFNST set and the LFNST index.

[0382] Here, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks into which the current block is divided. That is, since the same intra-prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra-prediction mode can also be applied equivalently to all sub-partition transform blocks. Furthermore, since the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partition transform blocks into which the current block is divided.

[0383] As described above, the transform set can be determined based on the intra-prediction mode used for the transform block to be transformed, and the inverse LFNST can be performed based on the transform kernel matrix (i.e., any one of the LFNST matrices) included in the transform set indicated by the LFNST index. The matrix applied to the inverse LFNST can be called the inverse LFNST matrix or the LFNST matrix, and can be referred to by any term, as long as the matrix is ​​the transpose of the matrix used for the forward LFNST.

[0384] In the example, the inverse LFNST matrix can be a non-square matrix, where the number of columns is less than the number of rows.

[0385] The decoding device can derive residual samples for the current block based on the primary inverse transform of the transform coefficients used for modification (S1550).

[0386] Here, as the primary inverse transform, a general separable transform can be used, or the aforementioned MTS can be used.

[0387] Subsequently, the decoding device 300 can generate a reconstruction sample based on the residual sample used for the current block and the prediction sample used for the current block.

[0388] The following figures are provided to illustrate specific examples of this disclosure. Because specific terminology used in the apparatuses or signals / messages / fields illustrated in the figures is provided for illustrative purposes, the technical features of this disclosure are not limited to the specific terminology used in the following figures.

[0389] Figure 16 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.

[0390] exist Figure 16 Each process disclosed in the document is based on a reference. Figures 4 to 14 Some details described. Therefore, with reference Figure 2 and Figures 4 to 14 The description of overlapping specific details will be omitted or will be shown schematically.

[0391] According to an embodiment, the encoding device 200 can derive prediction samples for the current block based on the intra-prediction mode applied to the current block.

[0392] When the ISP is applied to the current block, the encoding device can perform prediction by transforming the block for each sub-partition.

[0393] The encoding device can determine whether to apply ISP encoding or ISP mode to the current block, i.e., the encoded block, and can determine the direction in which the current block is divided, and can derive the size and number of the sub-blocks based on the determination result.

[0394] The same intra-prediction mode can be applied to the sub-partition transform blocks into which the current block is divided, and the encoding device can derive prediction samples for each sub-partition transform block. That is, the encoding device performs intra-prediction sequentially according to the partitioning form of the sub-partition transform blocks, for example, horizontally or vertically, or from left to right or from top to bottom. For the leftmost or topmost sub-block, the reconstructed pixels of the already encoded coded blocks are referred to in the conventional intra-prediction method. Further, for each side of a subsequent inner sub-partition transform block that is not adjacent to the previous sub-partition transform block, in order to derive the reference pixels adjacent to that side, the reconstructed pixels of the already encoded adjacent coded blocks are referenced as in the conventional intra-prediction method.

[0395] The encoding device 200 can derive residual samples for the current block based on the predicted samples (S1610).

[0396] The encoding device 200 can derive transform coefficients for the current block by applying at least one of LFNST or MTS to the residual sample, and can arrange the transform coefficients according to a predetermined scan order.

[0397] The encoding device can derive transform coefficients for the current block based on transform processes such as primary transforms and / or secondary transforms on the residual samples. LFNST can be applied when the current block is a single-tree type and a luminance component, and LFNST can be omitted when the current block is a single-tree type and a chrominance component (S1620).

[0398] The primary transform can be performed using multiple transform kernels, as in MTS, where the transform kernel can be selected based on the intra-prediction mode.

[0399] The encoding device 200 can determine whether to perform a secondary transform or a non-separable transform, specifically LFNST, on the transform coefficients used for the current block, and can derive modified transform coefficients by applying LFNST to the transform coefficients.

[0400] LFNST is an inseparable transform where the transform is applied to the coefficients without separating them in a specific direction, unlike a primary transform where the coefficients to be transformed are separated vertically or horizontally and then transformed. This inseparable transform can also be a low-frequency inseparable transform that applies the transform only to the low-frequency region rather than the entire target block to be transformed.

[0401] The encoding device can derive various variables to apply LFNST, and can determine whether to apply LFNST based on the tree type and size of the current block.

[0402] The encoding device can derive a first variable (variable LfnstDcOnly) indicating whether a valid coefficient exists at a location other than the location of the DC component in the current block, and a second variable (variable LfnstZeroOutSigCoeffFlag) indicating whether a transform coefficient exists in a second region other than the first region in the upper left of the current block.

[0403] The first and second variables are initially set to 1, wherein the first variable can be updated to 0 when the effective coefficients are located outside the location of the DC component in the current block, and the second variable can be updated to 0 when the transform coefficients are located in the second region.

[0404] When the first variable is updated to 0 and the second variable remains 1, LFNST can be applied to the current block.

[0405] For the luminance component applicable to Intra-Frame Sub-Partition (ISP) mode, LFNST can be applied without deriving the variable LfnstDcOnly.

[0406] Specifically, when the ISP mode is applied and the transform skip flag for the luma component (i.e., transform_skip_flag[x0][y0][0]) is 0, LFNST can be applied regardless of the value of the variable LfnstDcOnly, provided that the tree type of the current block is a single tree or a double tree for luma.

[0407] However, for chroma components that do not apply ISP mode, the value of the variable LfnstDcOnly can be set to 0 based on the values ​​of the transform skip flags for chroma component Cb (transform_skip_flag[x0][y0][1]) and chroma component Cr (transform_skip_flag[x0][y0][2]). That is, when the value of cIdx in transform_skip_flag[x0][y0][cIdx] is 1, the value of the variable LfnstDcOnly can be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0; and when the value of cIdx is 2, the value of the variable LfnstDcOnly can be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0. When the value of the variable LfnstDcOnly is 0, the encoding device can apply LFNST; otherwise, the encoding device can choose not to apply LFNST.

[0408] The second variable can be LfnstZeroOutSigCoeffFlag, which indicates that zeroing should be performed when LFNST is applied. The second variable can be initially set to 1 and can be changed to 0 when valid coefficients are present in the second region.

[0409] The variable LfnstZeroOutSigCoeffFlag can be exported as 0 when the index of the sub-block with the last non-zero coefficient is greater than 0 and the width and height of the transform block are both equal to or greater than 4, or when the position of the last non-zero coefficient in the sub-block with the last non-zero coefficient is greater than 7 and the size of the transform block is 4×4 or 8×8. A sub-block refers to a 4×4 block used as a coding unit in residual coding and can be called a coefficient group (CG). Sub-block index 0 refers to the top-left 4×4 sub-block.

[0410] In other words, when non-zero coefficients are derived in a region other than the top left region (where LFNST transform coefficients may exist in the transform block), or when non-zero coefficients exist in a 4×4 block or 8×8 block at a position other than the eighth position of the scan order, the variable LfnstZeroOutSigCoeffFlag is set to 0.

[0411] The coding device can determine the set of LFNSTs, including the LFNST matrix, based on the intra-prediction mode derived from information about the intra-prediction mode, and can select any one of the multiple LFNST matrices.

[0412] Here, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks into which the current block is divided. That is, since the same intra-prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra-prediction mode can also be applied equivalently to all sub-partition transform blocks. Furthermore, since the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partition transform blocks into which the current block is divided.

[0413] As described above, the transform set can be determined based on the intra-prediction mode used for the transform block to be transformed, and LFNST can be performed based on the transform kernel matrices (i.e., any of the LFNST matrices) included in the LFNST transform set. The matrix applied to LFNST can be called an LFNST matrix and can be referred to by any term, as long as the matrix is ​​the transpose of the matrix used for inverse LFNST.

[0414] In the example, the LFNST matrix can be a non-square matrix with fewer rows than columns.

[0415] The encoding device can determine whether the scaling list is applied to the current block based on whether LFNST is performed during the transformation process and the tree type of the current block (S1630).

[0416] The scaling list is a matrix used to assign specific weights (weights) to each transform coefficient position in a transform block, and can be dequantized or quantized by multiplying it by the weights used for each transform coefficient, thereby enabling differential dequantization or quantization based on the importance of the transform coefficient.

[0417] According to the example, when the tree type of the current block is single-tree and the current block is a luma component, the encoding device may not apply a scaling list, and when the tree type of the current block is single-tree and the current block is a chroma component, the encoding device may apply a scaling list.

[0418] When the current block's tree type is single-tree, the color components of the current block may include a luminance component, a first chromaticity component indicating chromaticity Cb, and a second chromaticity component indicating chromaticity Cr. When the current block's tree type is two-tree luminance, the current block may include a luminance component. When the current block's tree type is two-tree chromaticity, the color components of the current block may include a first chromaticity component and a second chromaticity component.

[0419] Here, the current block can be a transform block, which is a transform unit. When the tree type of the current block is single-tree, the current block can include transform blocks for the luma component, transform blocks for the first chroma component, and transform blocks for the second chroma component. When the tree type of the current block is two-tree luma, the current block can include transform blocks for the luma component. And when the tree type of the current block is two-tree chroma, the current block can include transform blocks for the first chroma component and transform blocks for the second chroma component.

[0420] According to the example, when the current block is a single tree, the encoding device can apply LFNST only to the luma component, and when LFNST is applied, the encoding device does not apply the scaling list to the luma component. However, the encoding device can apply the scaling list to the chroma component where LFNST is not applied.

[0421] In summary, when the LFNST index is greater than 0 (i.e., LFNST is applied), the encoding device may not apply the scaling list when the tree type of the current block is single-tree and the current block is a luminance component, and the encoding device may apply the scaling list when the tree type of the current block is single-tree and the current block is a chrominance component.

[0422] According to the example, when the LFNST index is greater than 0, LFNST can be applied to the current block when the tree type of the current block is dual-tree chroma, and therefore the encoding device does not apply the scaling list.

[0423] According to the example, when the LFNST index is greater than 0, LFNST can be applied to the current block when the tree type of the current block is dual-tree luminance, and therefore the encoding device does not apply the scaling list.

[0424] The encoding device can quantize the transform coefficients based on a determination, namely whether the scaling list is applied to the current block (S1640).

[0425] In other words, the encoding device can use a scaling list to quantize the transform coefficients of a transform block that has not applied LFNST, and can quantize the transform coefficients of a transform block that has applied LFNST without using a scaling list.

[0426] The encoding device can encode and output residual information and flag information indicating whether the scaling list is available when LFNST is performed (S1650).

[0427] The flag indicating whether the scaling list is available when LFNST is executed can be represented by `scaling_matrix_for_lfnst_disabled_flag` or `sps_scaling_matrix_for_lfnst_disabled_flag`, and can be sent as a signal in the sequence parameter set. A value of 1 indicates that the scaling list is not applied when LFNST is applied, and a value of 0 indicates that the scaling list is applicable when LFNST is applied.

[0428] When the LFNST index is greater than 0 and the current block is a single tree, LFNST can be applied to the luminance component, and therefore the encoding device can encode the value of the flag as 1.

[0429] However, when the LFNST index is greater than 0 and the current block is a single tree, LFNST is not applied to the chroma component, and therefore the encoding device can construct image information that allows the scaling list to be applied.

[0430] When the LFNST index is greater than 0 and the tree type of the current block is dual-tree chroma, the LFNST can be applied to the current block, and therefore the encoding device can encode the value of the flag as 1 so that the scaling list is not applied to the chroma components.

[0431] According to the example, when the LFNST index is greater than 0 and the tree type of the current block is dual-tree luminance, LFNST can be applied to the current block, and therefore the encoding device can encode the value of the flag as 1, so that the scaling list is not applied to the luminance component.

[0432] The encoding device can derive quantized transform coefficients by quantizing the transform coefficients used for modifications to the current block, and can encode the LFNST index.

[0433] The encoding device can generate residual information that includes information about the quantized transform coefficients. The residual information may include the aforementioned transform-related information / syntax elements. The encoding device can encode image / video information including the residual information and can output the image / video information as a bitstream.

[0434] Specifically, the encoding device 200 can generate information about the quantized transform coefficients and can encode the information about the quantized transform coefficients.

[0435] According to this embodiment, the syntax elements of the LFNST index can indicate whether (inverse) LFNST and any of the LFNST matrices included in the LFNST set are applied, and when the LFNST set includes two transformation kernel matrices, the syntax elements of the LFNST index can have three values.

[0436] According to the example, when the tree structure of the current block is a dual-tree type, the LFNST index can be encoded for each pair in the luma block and chroma block.

[0437] According to an embodiment, the value of the syntax element of the transform index can include: 0, which indicates that no (inverse) LFNST is applied to the current block; 1, which indicates the first LFNST matrix in the LFNST matrix; and 2, which indicates the second LFNST matrix in the LFNST matrix.

[0438] In this disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the transform coefficients of the quantization may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or for the sake of consistency, they may still be referred to as transform coefficients.

[0439] Furthermore, in this disclosure, the quantized transform coefficients and the transform coefficients can be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information can include information about the transform coefficients, and this information can be signaled via residual coding syntax. Transform coefficients can be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients can be derived via the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaled transform coefficients. These details can also be applied / expressed in other parts of this disclosure.

[0440] In the above embodiments, the method is described based on a flowchart using a series of steps or blocks. However, this disclosure is not limited to the order of the steps, and a step may be performed in a different order or in steps than described above, or simultaneously with another step. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and one or more steps of the flowchart may be incorporated or removed without affecting the scope of this disclosure.

[0441] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0442] When the embodiments of this disclosure are implemented in software, the above methods can be specifically embodied as modules (processes, functions, etc.) for performing the above functions. Modules can be stored in memory and can be executed by a processor. Memory can be internal or external to the processor and can be connected to the processor in various known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this disclosure can be specifically implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be specifically implemented and executed on a computer, processor, microprocessor, controller, or chip.

[0443] Furthermore, the decoding and encoding devices using this disclosure can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, Internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, Internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0444] Furthermore, the processing methods of this disclosure can be generated in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having the data structure according to this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices in which computer-readable data is stored. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media embodied in the form of a carrier wave (e.g., transmission via the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks. Additionally, embodiments of this disclosure can be embodied as computer program products by program code, and the program code can be executed on a computer by embodiments of this disclosure. The program code can be stored on a computer-readable medium.

[0445] Figure 17 The diagram illustrates the structure of a content streaming system that applies this disclosure.

[0446] Furthermore, the content streaming system using this disclosure may primarily include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0447] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmits the bitstream to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted. The bitstream can be generated using the encoding method or bitstream generation method described in the embodiments of this document. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.

[0448] A streaming server sends multimedia data to a user's device via a web server based on a user's request, and the web server acts as a medium for informing the user of services. When a user requests a desired service from the web server, the web server forwards it to the streaming server, which then sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to manage the commands / responses between devices within the content streaming system.

[0449] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bit stream for a predetermined period of time.

[0450] For example, user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

[0451] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims can be combined to be implemented or performed in a device, and the technical features of the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features of the method claims and the device claims can be combined to be implemented or performed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or performed in a method.

Claims

1. An apparatus comprising: Memory; as well as At least one processor is connected to the memory, and the at least one processor is configured to: Based on the residual information received from the bitstream, the transform coefficients for the current block are derived, and Based on the inverse transform used for the transform coefficients, residual samples for the current block are derived. The transformation coefficients are derived based on the following: Based on the residual information, transform coefficients for quantization of the current block are derived, and The quantized transform coefficients are dequantized to derive the transform coefficients. Specifically, based on whether a scaling list is applied to the current block, the quantized transform coefficients are dequantized. Specifically, based on the Low Frequency Inseparable Transform (LFNST) index of the current block and the tree type of the current block, it is determined whether the scaling list is applied to the current block. Specifically, based on the fact that the tree type of the current block is a single tree and the color component of the current block is a chroma component, LFNST is not applied to the chroma component of the current block, and the scaling list is applied to the chroma component of the current block. Specifically, based on the flag information indicating that the scaling list cannot be used in the block to which the LFNST is applied, the tree type of the current block is dual-tree chroma, and the LFNST index is greater than 0, the scaling list is not applied to the chroma components of the current block.

2. The device according to claim 1, wherein, Based on the fact that the tree type of the current block is a single tree and the LFNST index is greater than 0, the scaling list is not applied to the luminance component of the current block.

3. The device according to claim 1, wherein, Based on the flag information indicating that the scaling list is unavailable, the LFNST index is greater than 0, and the tree type of the current block is the single tree, the scaling list is not applied to the luminance component of the current block.

4. The device according to claim 1, wherein, Based on the flag information indicating that the scaling list is unavailable, the LFNST index is greater than 0, and the tree type of the current block is the dual-tree chroma, the scaling list is not applied to the chroma components of the current block.

5. The device according to claim 1, wherein, Based on the flag information indicating that the scaling list is unavailable, the LFNST index is greater than 0, and the tree type of the current block is dual-tree luminance, the scaling list is not applied to the luminance component of the current block.

6. An apparatus comprising: Memory; as well as At least one processor is connected to the memory, and the at least one processor is configured to: Based on the transformation process, the transformation coefficients for the current block are derived from the residual samples used for the current block. The transform coefficients are quantized based on whether a scaling list is applied to the current block. Generate residual information including information about the quantized transform coefficients, and The image information, including the residual information, is encoded. Specifically, based on the Low Frequency Inseparable Transform (LFNST) index of the current block and the tree type of the current block, it is determined whether the scaling list is applied to the current block. Wherein, based on the fact that the tree type of the current block is a single tree and the color component of the current block is a chroma component, the scaling list is applied to the chroma component of the current block, and Wherein, based on the fact that the tree type of the current block is dual-tree chroma and the LFNST index is greater than 0, the scaling list is not applied to the chroma components of the current block.

7. The device according to claim 6, wherein, Based on the fact that the tree type of the current block is the single tree and the color component of the current block is the chroma component, the LFNST is not performed on the chroma component of the current block.

8. The device according to claim 6, wherein, Based on the fact that the tree type of the current block is a single tree and the LFNST index is greater than 0, the scaling list is not applied to the luminance component of the current block.

9. The device according to claim 6, wherein, Based on the fact that the tree type of the current block is dual-tree luminance and the LFNST index is greater than 0, the scaling list is not applied to the luminance component of the current block.

10. The device according to claim 6, wherein, The at least one processor is further configured to: The encoding indicates whether the scaling list is available for the flag information of the block to which the LFNST is applied.

11. The device according to claim 6, wherein, The current block includes the transform block.

12. An apparatus comprising: At least one processor is configured to generate a bitstream for image information, wherein the bitstream is generated based on: deriving transform coefficients for the current block from residual samples for the current block based on a transform process, quantizing the transform coefficients based on whether a scaling list is applied to the current block, generating residual information including information about the quantized transform coefficients, and encoding the image information including the residual information. as well as A transmitter configured to transmit the data comprising the bit stream. Specifically, based on the Low Frequency Inseparable Transform (LFNST) index of the current block and the tree type of the current block, it is determined whether the scaling list is applied to the current block. Wherein, based on the fact that the tree type of the current block is a single tree and the color component of the current block is a chroma component, the scaling list is applied to the chroma component of the current block, and Wherein, based on the fact that the tree type of the current block is dual-tree chroma and the LFNST index is greater than 0, the scaling list is not applied to the chroma components of the current block.

Citation Information

Cited By

  • A boron-nitrogen-phosphorus co-doped bismuth-based flow battery membrane electrode and a preparation method thereof

    CN122696697A