Determining intra-prediction modes for indexing non-separable transform kernels

The LFNST tool addresses the complexity and storage issues of non-separable transforms in video coding by employing non-spanning transforms and explicit signaling, enhancing coding efficiency and performance in video coding standards.

JP2026504981APending Publication Date: 2026-02-10GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025542399
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-24
Filing Date
2024-01-08
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing video coding techniques face challenges in efficiently utilizing non-separable transforms for intra-prediction due to high computational complexity and large storage requirements, particularly in video coding standards like H.266/VVC, which limits their effectiveness in achieving optimal coding performance.

Method used

The introduction of a low-frequency non-separable secondary transform (LFNST) tool in video coding systems, which applies to a range of block sizes with reduced complexity by using non-spanning transforms and explicit signaling mechanisms, allowing for efficient selection of transform matrices based on intra-prediction modes.

Benefits of technology

The LFNST tool reduces computational complexity and storage needs while maintaining coding efficiency by adaptively applying non-separable transforms, improving video coding performance with minimal reconstruction loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026504981000001_ABST
    Figure 2026504981000001_ABST
Patent Text Reader

Abstract

According to one aspect of the present disclosure, a decoding method is provided that is executed by a decoder. The method may include analyzing a bitstream by a processor. The method may include determining, by the processor, at least one non-separable transform and an intra-prediction mode to be used to decode a coding unit (CU) of the bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. The method may include selecting, by the processor, a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. The method may include decoding, by the processor, the CU based on the intra-prediction method and the transform matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 440,899, filed January 24, 2023, entitled "DETERMINATION OF INTRA PREDICTION MODE FOR INDEXATION INTO NON-SEPARABLE TRANSFORM KERNELS," the entire contents of which are incorporated herein by reference. [Background technology]

[0002] TECHNICAL FIELD The embodiments of the present disclosure relate to video encoding and decoding.

[0003] Digital video has become mainstream and is used in a wide range of applications, including digital television, video telephony, and video conferencing. Advances in computing and communication technologies, as well as efficient video coding techniques, have made these digital video applications possible. Various video coding techniques can be used to compress video data, whereby encoding of the video data may be performed using one or more video coding standards. Exemplary video coding standards may include, but are not limited to, Versatile Video Coding (H.266 / VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, and the like. Recent video coding techniques have also been evaluated in exploratory video coding models, such as the enhanced compression model (ECM). Summary of the Invention

[0004] According to one aspect of the present disclosure, a decoding method is provided that is executed by a decoder. The method may include analyzing a bitstream by a processor. The method may include determining, by the processor, at least one non-separable transform and an intra-prediction mode to be used to decode a coding unit (CU) of the bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. The method may include selecting, by the processor, a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. The method may include decoding, by the processor, the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to intra block copy (IBC), matrix weighted intra prediction (MIP), or intra template matching prediction (IntraTMP) being used for the CU. The at least one non-separable transform may include a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

[0005] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and a memory having instructions stored therein. The memory stores instructions that, when executed by the processor, cause the processor to analyze a bitstream. In response to enabling at least one non-separable transform for an intra-prediction method, the memory stores instructions that, when executed by the processor, cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to decode a CU of the bitstream. The at least one non-separable transform may be associated with multiple transform matrix sets. The memory stores instructions that, when executed by the processor, cause the processor to select a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. The memory stores instructions that, when executed by the processor, cause the processor to decode a CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to IBC, MIP, or IntraTMP. The at least one non-separable transform may include LFNST or NSPT.

[0006] According to another aspect of the present disclosure, a non-transitory computer-readable medium for storing instructions for a decoder is provided. The instructions, when executed by a processor, cause the processor to analyze a bitstream. When executed by the processor, the instructions, when executed by the processor, cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to decode a CU of the bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. When executed by the processor, the instructions, when executed by the processor, cause the processor to select a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. When executed by the processor, the instructions, when executed by the processor, cause the processor to decode a CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to IBC, MIP, or IntraTMP. The at least one non-separable transform may include LFNST or NSPT.

[0007] According to yet another aspect of the present disclosure, an encoding method performed by an encoder is provided. The method may include, in response to enabling at least one non-separable transform for an intra-prediction method, determining, by a processor, an intra-prediction mode and the at least one non-separable transform to be used to encode a CU into a bitstream, where the at least one non-separable transform is associated with a plurality of transform matrix sets. The method may include selecting, by the processor, a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode. The method may include encoding, by the processor, the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST or NSPT.

[0008] According to yet another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and a memory having instructions stored thereon. The memory stores instructions that, when executed by the processor, cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to encode a CU into a bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. The memory stores instructions that, when executed by the processor, cause the processor to select a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. The memory stores instructions that, when executed by the processor, cause the processor to encode the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST or NSPT.

[0009] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions is provided. When executed by a processor, the instructions cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to encode a CU into a bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. When executed by a processor, the instructions cause the processor to select a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. When executed by a processor, the instructions cause the processor to encode the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST or NSPT.

[0010] These illustrative examples are not intended to limit or define the present disclosure, but rather to provide examples to aid in understanding the present disclosure. Specific embodiments provide further explanations, with other examples being described in specific embodiments. [Brief explanation of the drawings]

[0011] [Figure 1] 1 illustrates a block diagram of an exemplary encoding system according to some embodiments of the present disclosure. [Figure 2] 1 illustrates a block diagram of an exemplary decoding system according to some embodiments of the present disclosure. [Figure 3] 2 shows a detailed block diagram of an example encoder in the encoding system of FIG. 1, in accordance with some embodiments of the present disclosure. [Figure 4] 3 shows a detailed block diagram of an example decoder in the decoding system of FIG. 2, in accordance with some embodiments of the present disclosure. [Figure 5]1 illustrates an example image divided into coding tree units (CTUs), according to some embodiments of the present disclosure. [Figure 6] 1 illustrates an example CTU divided into coding units (CUs), according to some embodiments of the present disclosure. [Figure 7A] 1 illustrates a low-frequency non-separable transform (LFNST) kernel for 4×N and N×4 block sizes in VVC, according to some embodiments of the present disclosure. [Figure 7B] 10 illustrates an LFNST kernel for 8×N and N×8 block sizes in VVC, in accordance with some embodiments of the present disclosure. [Figure 8] 1 illustrates various intra-angle prediction modes according to some embodiments of the present disclosure. [Figure 9A] 10 illustrates LFNST kernels for 4×N and N×4 block sizes in ECM, according to some embodiments of the present disclosure. [Figure 9B] 10 illustrates LFNST kernels for 8×N and N×8 block sizes in ECM, according to some embodiments of the present disclosure. [Figure 9C] 10 illustrates an LFNST kernel for a 16x16 block size in ECM, in accordance with some embodiments of the present disclosure. [Figure 10] 1 illustrates a matrix-weighted intra-prediction (MIP) mode for ECM, according to some embodiments of the present disclosure. [Figure 11] 1 illustrates an intra-block copy (IBC) mode for ECM according to some embodiments of the present disclosure. [Figure 12] 1 illustrates an intra-template matching prediction (IntraTMP) mode for ECM, according to some embodiments of the present disclosure. [Figure 13] 1 illustrates a decoder-side intra mode derivation (DIMD) mode for ECM, according to some embodiments of the present disclosure. [Figure 14]1 illustrates a template-based intra mode derivation (TIMD) mode for ECM, according to some embodiments of the present disclosure. [Figure 15] 1 shows a flowchart of an example video decoding method according to some embodiments of the present disclosure. [Figure 16] 1 illustrates a flowchart of an exemplary video encoding method according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012] The drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to further explain the principles of the present disclosure and to enable those skilled in the art to make and use the disclosure.

[0013] An embodiment of the present disclosure will be described with reference to the drawings.

[0014] While several configurations and arrangements are discussed, it should be understood that this is done for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. Those skilled in the art will recognize that the present disclosure can be used in a variety of other applications.

[0015] It should be noted that references in the specification to "one embodiment," "embodiment," "exemplary embodiment," "some embodiments," "an embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in combination with an embodiment, it is within the knowledge of one of ordinary skill in the relevant art to implement such feature, structure, or characteristic in combination with other embodiments, whether or not explicitly described.

[0016] Generally, understanding can be achieved, at least in part, from contextual usage. For example, the term "one or more" as used herein can be used to describe any singular feature, structure, or characteristic, or can be used to describe a plural combination of features, structures, or characteristics, depending at least in part on the context. Similarly, terms such as "a," "an," or "the" can also be understood as conveying a singular or plural usage, depending at least in part on the context. Furthermore, the term "based on" is not necessarily intended to convey an exclusive set of factors, but can be understood to allow for the presence of additional factors not necessarily explicitly recited, depending at least in part on the context.

[0017] Aspects of a video encoding and decoding system will now be described with reference to various apparatus and methods. These apparatus and methods are described in the following specific embodiments and illustrated in the figures by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as "elements"). These elements may be realized using electronic hardware, firmware, computer software, or any other combination. Whether these elements are implemented as hardware, firmware, or software depends on the particular application and design constraints imposed on the overall system.

[0018] The techniques described herein can be used for various video encoding and decoding applications. As described herein, video encoding and decoding includes video encoding and video decoding. Video encoding and decoding can be performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, and reconstruction can be performed on a coding block, a transform block, or a prediction block. As described herein, a block to be coded / decoded is referred to as a "current block." For example, the current block can represent a coding block, a transform block, or a prediction block according to a current coding / decoding process. Furthermore, it should be understood that the term "unit" used in this disclosure refers to a basic unit for performing a particular coding / decoding process, and the term "block" refers to an array of samples of a predetermined size. Unless otherwise specified, "block" and "unit" can be used interchangeably.

[0019] FIG. 1 illustrates a block diagram of an exemplary encoding system 100 according to some embodiments of the present disclosure. FIG. 2 illustrates a block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure. Either of the systems 100 or 200 may be applied to or integrated into various systems and devices capable of data processing, such as computers and wireless communication devices. For example, the system 100 or 200 may be all or part of a mobile phone, desktop computer, laptop computer, tablet computer, in-vehicle computer, game console, printer, positioning device, wearable electronic device, smart sensor, virtual reality (VR) device, augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. As illustrated in FIGS. 7 and 8, the system 100 or 200 may include a processor 102, memory 104, and interface 106. While these components are shown as interconnected by a bus, other connection types are also acceptable. It should be understood that the system 100 or 200 may include any other suitable components for performing the functions described herein.

[0020] The processor 102 may include a microprocessor, such as a graphics processing unit (GPU), image signal processor (ISP), central processing unit (CPU), digital signal processor (DSP), tensor processing unit (TPU), vision processing unit (VPU), neural processing unit (NPU), synergistic processing unit (SPU), or physics processing unit (PPU), microcontroller units (MCU), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gate logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described in this disclosure. While only one processor is shown in FIGS. 7 and 8, it should be understood that multiple processors may be included. The processor 102 may be a hardware device having one or more processing cores, and may be capable of executing software.Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, processes, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Software may include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for instructing hardware are also included in the broad category of software.

[0021] Memory 104 may broadly include memory (also referred to as primary / system memory) and storage (also referred to as secondary memory). For example, memory 104 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM), or other optical disk storage, magnetic disk storage or other magnetic storage devices such as hard disk drives (HDDs), flash drives, solid-state drives (SSDs), or any other medium that can be used to carry or store desired program code in the form of instructions that can be accessed and executed by processor 102. Broadly, memory 104 may be embodied by any computer-readable medium, such as a non-transitory computer-readable medium. Although only one memory is shown in Figures 7 and 8, it should be understood that multiple memories may be included.

[0022] The interface 106 may broadly include a data interface and a communication interface, which is configured to send and receive signals in the process of transmitting and receiving information to and from other external network elements. For example, the interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although only one interface is shown in Figures 7 and 8, it should be understood that multiple interfaces may be included.

[0023] The processor 102, memory 104, and interface 106 may be implemented in various forms in the system 100 or 200 to perform video encoding and decoding functions. In some embodiments, the processor 102, memory 104, and interface 106 of the system 100 or 200 are implemented on (e.g., integrated onto) one or more system-on-chip (SoCs). In one example, the processor 102, memory 104, and interface 106 may be integrated onto an application processor (AP) SoC, which handles application processing in an operating system (OS) environment, including running video encoding and decoding applications. In another example, the processor 102, memory 104, and interface 106 may be integrated onto a dedicated processor chip for video encoding and decoding, such as a GPU or ISP chip specialized for image and video processing in a real-time operating system (RTOS).

[0024] As shown in FIG. 1, in the encoding system 100, the processor 102 may include one or more modules, such as an encoder 101. While FIG. 1 illustrates the encoder 101 within one processor 102, it should be understood that the encoder 101 may include one or more sub-modules that may be implemented on different processors that are proximate or remote from each other. The encoder 101 (and any corresponding sub-modules or sub-units) may be a hardware unit (e.g., part of an integrated circuit) of the processor 102 that is designed to be used with other components or software units that are implemented by the processor 102 executing at least some programs (e.g., instructions). The program instructions may be stored in a computer-readable medium, such as the memory 104, and when executed by the processor 102, the instructions may cause the processor to perform a process having one or more functions related to video encoding (e.g., image segmentation, inter-prediction, intra-prediction, transform, quantization, filtering, entropy coding, etc.), as described in detail below.

[0025] Similarly, as shown in FIG. 2, in the decoding system 200, the processor 102 may include one or more modules, such as a decoder 201. While FIG. 2 illustrates the decoder 201 being located within one processor 102, it should be understood that the decoder 201 may include one or more sub-modules that may be implemented on different processors that are proximate or remote from each other. The decoder 201 (and any corresponding sub-modules or sub-units) may be a hardware unit (e.g., part of an integrated circuit) of the processor 102 that is designed to be used with other components or software units that are implemented by the processor 102 executing at least some programs (e.g., instructions). The program instructions may be stored in a computer-readable medium, such as the memory 104, and when executed by the processor 102, the instructions may cause the processor to perform a process having one or more functions related to video decoding (e.g., entropy decoding, inverse quantization, inverse transform, inter-prediction, intra-prediction, filtering), as described in detail below.

[0026] FIG. 3 illustrates a detailed block diagram of an exemplary encoder 101 in the encoding system 100 of FIG. 1 , according to some embodiments of the present disclosure. As illustrated in FIG. 3 , the encoder 101 may include a partitioning module 302, an inter-prediction module 304, an intra-prediction module 306, a transform module 308, a quantization module 310, an inverse quantization module 312, an inverse transform module 314, a filter module 316, a buffer module 318, and an encoding module 320. It should be understood that each element illustrated in FIG. 3 is shown independently to represent different characteristic functions in a video encoder. This does not mean that each component is formed by a separate hardware unit or a single piece of software. That is, for convenience of explanation, each element is described as an independent element, and at least two elements may be combined to form a single element, or one element may be divided into multiple elements to perform a function. It should also be understood that some elements are not essential elements for performing the functions described in the present disclosure, but are optional elements for improving performance. Furthermore, it should be understood that these elements may be realized using electronic hardware, firmware, computer software, or any other combination. Whether these elements are implemented as hardware, firmware, or software depends on the particular application and design constraints imposed on the encoder 101.

[0027] The splitting module 302 may be configured to split an input image of a video into at least one processing unit. An image may be a frame of a video or a field of a video. In some embodiments, the image includes an array of luma samples in a monochrome format, or an array of luma samples and two corresponding arrays of chroma samples.

[0028] Similar to H.265 / HEVC, H.266 / VVC is a block-based hybrid spatial and temporal predictive coding scheme. As shown in FIG. 5, during encoding, an input image 500 is first divided into multiple square blocks (CTUs 502) by a partitioning module 302. For example, the CTUs 502 may be blocks of 128×128 pixels. As shown in FIG. 6, each CTU 502 in the image 500 can be divided into one or more CUs 602 by the partitioning module 302, and the CUs 602 can be used for prediction and transformation. Unlike H.265 / HEVC, in H.266 / VVC, the CUs 602 may be rectangular or square and can be coded without further division into prediction units or transform units. For example, as shown in FIG. 6, dividing a CTU 502 into multiple CUs 602 may include quadtree partitioning (shown by solid lines), binary tree partitioning (shown by dashed lines), and ternary tree partitioning (shown by dashed lines). According to some embodiments, each CU 602 may be the same size as its root CTU 502, or each CU 602 may be a subdivision of the root CTU 502 as small as a 4x4 block.

[0029] At this time, the division module 302 can divide the image into multiple coding units. Each coding unit can be further subdivided into one or more prediction units (PUs), and each prediction unit can be subdivided into one or more transform units (TUs). However, in the general case, each coding unit corresponds to one prediction unit and one transform unit of the same size.

[0030] Referring to FIG. 4, the inter prediction module 304 may be configured to perform inter prediction on the prediction unit, and the intra prediction module 306 may be configured to perform intra prediction on the prediction unit. Whether to use inter prediction or intra prediction on the prediction unit is determined, and specific information (e.g., intra prediction mode, motion vector, reference image, etc.) may be determined according to various prediction methods. Residual values ​​in a residual block between the generated prediction block and the original block may be input to the transform module 308. Furthermore, the encoding module 320 may encode prediction mode information, motion vector information, etc. used for prediction into a bitstream along with quantization levels of transformed or non-transformed coefficients. It should be understood that in some encoding modes, transform and / or quantization may be skipped.

[0031] In some embodiments, the inter prediction module 304 can predict a prediction unit based on information of at least one of previously coded images. In some cases, the inter prediction module 304 can predict a prediction unit based on information of a partial region already coded in a current image. The inter prediction module 304 may include sub-modules such as a reference image interpolation module, a motion prediction module, and a motion compensation module (not shown). For example, the reference image interpolation module can receive reference image information from the buffer module 318 and generate pixel information for an integer number of pixels or fewer pixels based on the reference image. For luma pixels, a discrete cosine transform (DCT)-based 8-tap interpolation filter with variable filter coefficients can be used to generate pixel information for an integer number of pixels or fewer pixels in 1 / 4 pixel increments. For chroma signals, a DCT-based 4-tap interpolation filter with variable filter coefficients can be used to generate pixel information for an integer number of pixels or fewer pixels in 1 / 8 pixel increments. The motion prediction module can perform motion prediction based on a reference image interpolated by a reference image interpolator. Various methods (e.g., full search-based block matching algorithm (FBMA), three-step search (TSS), and new three-step search algorithm (NTS)) can be used to calculate motion vectors. The motion vectors can have motion vector values ​​in 1 / 2, 1 / 4, or 1 / 16 pixel units, or integer number of pixels, based on interpolated pixels. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods (e.g., skip method, merge method, advanced motion vector prediction (AMVP) method, intra block copy method, etc.) can be used as motion prediction methods.

[0032] Continuing to refer to FIG. 3 , in some embodiments, the intra prediction module 306 may generate a prediction unit based on information about reference pixels surrounding the current block, which may be pixel information within the current image. The reference pixels may be located on a reference line adjacent or not adjacent to the current block. If a block adjacent to the current prediction unit is an intra-predicted block and thus the reference pixels are intra-predicted pixels, the reference pixels in the intra-predicted block may be used instead of the reference pixel information of the adjacent intra-predicted block. That is, if reference pixels are unavailable, at least one reference pixel among available reference pixels may be used instead of the unavailable reference pixel information. In intra prediction, prediction modes may include angular prediction modes and non-angular prediction modes. The angular prediction modes use reference pixel information based on the prediction direction, and the non-angular prediction modes do not use direction information when performing prediction. The mode for predicting luma information may be different from the mode for predicting chroma information. The chroma information may be predicted using intra-prediction mode information used to predict luma information or predicted luma signal information. When performing intra prediction, if the size of the prediction unit is the same as the size of the transform unit, intra prediction can be performed on the prediction unit based on the left pixel, the upper left pixel, and the top pixel. However, when performing intra prediction, if the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed using reference pixels based on the transform unit.

[0033] The intra prediction method may generate a predicted block after applying an adaptive intra smoothing (AIS) filter to reference pixels based on the prediction mode. The type of AIS filter applied to the reference pixels may be different. To perform the intra prediction method, the intra prediction mode of the current prediction unit may be predicted based on the intra prediction mode of a prediction unit existing in the vicinity of the current prediction unit. When predicting the prediction mode of the current prediction unit using mode information of neighboring prediction units, if the intra prediction mode of the current prediction unit is the same as that of the neighboring prediction unit, predetermined flag information may be used to transmit information indicating that the prediction mode of the current prediction unit is the same as that of the neighboring prediction unit. If the prediction mode of the current prediction unit is different from that of the neighboring prediction unit, additional flag information may be used to encode the prediction mode information of the current block.

[0034] 3, a residual block may be generated, the residual block including a prediction unit, which is predicted based on the prediction unit generated by the prediction module 304 or 306 and residual coefficient information (also referred to herein as "residual"), which is a difference value between the prediction unit and the original block. The generated residual block may be input to a transform module 308. Additional details regarding residuals and transforms for video coding are provided below.

[0035] In a hybrid video coding system, redundancy in a video signal is first mined by applying inter- or intra-prediction tools to each CU. The difference between the original samples of a CU and its predicted block is generally called the residual. Even after prediction, the residual may still have high spatial correlation. Although conditional entropy coding can capture some of the spatial dependence between neighboring samples, it is computationally impractical to create an entropy coding statistical model that can completely mine the spatial correlation in the residual. In contrast, transform coding is a practical and effective method for removing the spatial correlation of the residual.

[0036] For example, the transform module 308 may transform the residual using an integerized version of a two-dimensional DCT and may apply the transform in the horizontal and vertical directions, respectively. For an M×N residual sample block (where M is the width of the block and N is the height of the block), the transform module 308 may obtain intermediate transform coefficients by applying a one-dimensional DCT to each row, and then applying a one-dimensional DCT to each column of the intermediate transform coefficients.

[0037] The benefit of applying a transform can be estimated by the transform coding gain, which is the distortion (D SQ ), and the distortion (D TC ) is defined as the ratio of the residual to the transform coding gain G TC is further transformed into the variance

number

number

[0038]

number

[0039] Based on this interpretation, transform coding gain, and correspondingly the overall coding gain of the video coding system, can be achieved if the resulting transform coefficients exhibit energy compaction characteristics. In other words, the variance distribution of the transform coefficients is concentrated on a small number of transform coefficients, compared to the likely original residual samples, whose variance is evenly distributed.

[0040] The use of one-dimensional transforms applied separately in the horizontal and vertical directions scales computationally well as block sizes increase. In the above example, the transform coefficients are obtained by a matrix realization of the DCT, which requires (M+N) multiplications per sample. "Butterfly" factorization can reduce the number of multiplications per sample, at the expense of higher computational delay. Separable transforms can achieve optimal energy concentration for spatial features along Cartesian directions (e.g., vertical or horizontal). For example, vertical edges are perfectly concentrated by a vertical DCT. However, separable transforms cannot optimally utilize spatial features along non-Cartesian directions. In such cases, a well-designed non-separable transform can achieve better coding performance.

[0041] For example, a two-dimensional non-separable transform is applied directly to an input sample block while a separable transform applies one-dimensional transforms in the horizontal and vertical directions, respectively. One desirable property of a transform is that the transform vector spans the space of input samples. This means that any input vector (e.g., any combination of input sample values) can be represented by a weighted sum of transform vectors. For a transform to have spanning properties, the number of transform vectors must be at least equal to the dimension of the input space; in other words, the number of output transform coefficients must be at least equal to the number of input samples. For example, the one-dimensional DCT in VVC is a spanning transform. And for a spanning non-separable transform, if the input sample block is an M×N residual, the transform also outputs an M×N transform coefficient block, which can be realized by an (M×N)×(M×N) matrix multiplication implementation.

[0042] To derive a non-separable transform that produces coding gain for a particular directional feature, the transform can be trained. For example, a set of representative residual blocks corresponding to the directional feature of interest can be grouped, and then a Karhunen-Loeve Transform (KLT) can be calculated based on the covariance matrix of the set of residual blocks. The process can be repeated for K different sets of residual blocks. An overall transform kernel is then derived, with dimensions (M×N)×(M×N)×K in this example.

[0043] As described in this section, spanning non-separable transforms have two problems. First, they have high computational complexity. Because non-separable transforms are usually obtained by learning, they are generally not decomposable. The matrix implementation of spanning non-separable transforms in the above example results in a complexity of M × N multiplications per sample. The second problem is that the (multiple) transform kernels occupy a large amount of storage in the encoder 101 and decoder 201. In the above example, a single kernel that can adapt to K different directional features has (M × N) × (M × N) × K weights. This kernel can only be applied to residual blocks of size M × N. To make a non-separable transform applicable to multiple block sizes, one transform kernel must be learned for each block size.

[0044] To address the above spanning non-separable transformation issues, the LFNST tool was introduced into VVC and several modifications were made.

[0045] First, according to the first change, the LFNST tool applies to a wide range of block sizes, but only two LFNST kernels are defined. For example, the small LFNST kernel is applied to blocks of size 4xN or Nx4 (where N ≥ 4). The large LFNST kernel is applied to all other relatively large block sizes (e.g., 8x8 and above).

[0046] 7A shows a schematic diagram 700 of LFNST kernels for 4×N and N×4 block sizes in VVC, according to some embodiments of the present disclosure. FIG. 7B shows a schematic diagram 701 of LFNST kernels for 8×N and N×8 block sizes in VVC, according to some embodiments of the present disclosure.

[0047] Figures 7A and 7B show the sample positions that LFNST operates on. For example, from the encoder's perspective, for block sizes of 4xN and Nx4, the top-left 4x4 sample position (shown as the shaded region in Figure 7A) is transformed by a small LFNST. The remaining sample positions (shown as the white region in Figure 7A) are ignored, or "zeroed out." From the decoder's perspective, an inverse LFNST is applied to generate the top-left 4x4 sample, and the remaining samples are filled with zeros. A similar policy applies for larger block sizes, where LFNST operates on the three top-left 4x4 sample position blocks (shown as the shaded region in Figure 7B). The remaining sample positions are zeroed out.

[0048] As a result of the "zero-out" policy, the size of the LFNST is significantly reduced compared to a full-size transform applied to all sample positions. However, it is inherently lossy and cannot recover values ​​at sample positions ignored by the LFNST. If the LFNST tool were applied directly to the residual samples, such losses would be too great for the LFNST tool to use. However, the reason why the LFNST is called a secondary transform is because the LFNST is applied after the separable DCT has already been performed in the encoder and operates on the primary transform coefficients to generate secondary transform coefficients. In other words, the DCT can be considered a primary transform. According to embodiments of the present disclosure, the leftmost sample position in a primary transform coefficient block corresponds to the horizontal low-frequency portion of the DCT, while the topmost sample position corresponds to the vertical low-frequency portion of the DCT. By preferentially transforming and reconstructing the top-left sample position in the decoder 201, the LFNST can reconstruct low-frequency information based on the original residual. As mentioned above, coding gains arise due to the energy-concentrated nature of transforms, and it is well-documented that the variance (energy) of camera-captured images and video signals is concentrated in the low-frequency DCT coefficients. Therefore, while "zeroing out" prevents LFNST from reversibly reconstructing any residual block, in practice, for most types of images and video signals, the loss is likely to be minimal.

[0049] The second modification is that the applied transforms for the small LFNST kernel and the large LFNST kernel are not spanning transforms. From the perspective of the encoder 101, the number of output (secondary transform) coefficients is smaller than the number of input (primary transform) coefficients. For example, the small LFNST kernel takes 4 × 4 = 16 primary transform coefficients as input but produces only 8 output secondary transform coefficients. The large LFNST kernel takes 3 × 4 × 4 = 48 input primary transform coefficients and outputs 8 secondary transform coefficients. The use of a non-spanning transform introduces additional reconstruction loss. However, this loss can be traded off in a controllable manner and with reduced implementation complexity. First, a spanning non-separable transform can be designed using the KLT method described above. By following this method, the basis vectors of the transform correspond to feature vectors of a covariance matrix calculated based on a set of representative residual blocks. These feature vectors can be ranked by their importance according to their corresponding feature values, and the most important feature vectors are selected to reconstruct the non-spanning non-separable transform. For example, the eight feature vectors with the largest feature values ​​can be selected to form a non-spanning transform for a small LFNST kernel.

[0050] In summary, compared to spanning non-separable transforms, the above two modifications significantly reduce the complexity of the LFNST kernel. For relatively small blocks, using a small LFNST kernel reduces the complexity of each transform block from (4 × N) × (4 × N) multiplications (for N ≥ 4) to 16 × 8 multiplications. For relatively large blocks, using a large LFNST kernel reduces the complexity of each transform block from (8 × N) × (8 × N) multiplications (for N ≥ 8) to 48 × 8 multiplications.

[0051] The LFNST kernel does not contain only one transformation matrix. Multiple transformation matrices are learned to achieve better coding gain for various image and video signals. The number of different transformation matrices is the product of the third and fourth dimensions of the LFNST kernel. The dimensions of a small LFNST kernel are 16x8x2x4, while the dimensions of a large LFNST kernel are 48x8x2x4. The LFNST kernel is expressed with two additional dimensions because the specific transformation matrix of a transform block is selected through a combination of explicit signaling and implicit selection.

[0052] Explicit signaling is performed by an LFNST index written into the bitstream, which can take values ​​of 0, 1, or 2, where 0 indicates not using LFNST for the transform block, and values ​​1 or 2 indicate selecting from the third dimension of the LFNST kernel. The drawback of potential reconstruction loss due to zero-out and non-spanning simplifications is ameliorated by the explicit signaling mechanism. While using LFNST can result in excessive reconstruction loss for a transform block, the LFNST tool can be disabled by writing 0 to the LFNST index.

[0053] Implicit selection can be enabled by restricting LFNST to only target coding units that use intra prediction. Intra prediction generates a predicted block for a coding unit based on neighboring reference samples adjacent to the top and left of the current block. The intra prediction mode specifies a specific method for constructing the predicted block in the bitstream. Simple methods for intra prediction include taking the average of reference samples ("DC" mode) or constructing an affine interpolation between several reference samples ("planar" mode). However, most intra prediction modes are reserved for specifying an intra angle direction, in which the predicted block is constructed assuming that the values ​​of the reference samples are replicated along a specific direction. When an intra angle direction is used, it can be a strong hint about the directional characteristics of the residual block. Implicit selection of the LFNST transform is performed by mapping the intra prediction mode to one of four possible values ​​of a "transform set index," which is used to index the fourth dimension of the LFNST kernel. The mapping used in VVC is shown in Table 1.

[0054] [Table 1]

[0055] Intra-prediction modes 0 and 1 correspond to intra-prediction plane mode and intra-DC prediction mode, respectively. These modes are treated as special cases by mapping to transform set index 0. Otherwise, the remaining intra-prediction modes correspond to intra-angle directions 800, partially shown in FIG. 8. Intra-prediction mode 2 corresponds to diagonal intra-prediction from the lower-left corner. As the intra-prediction mode number increases, the intra-prediction direction rotates clockwise, where intra-prediction mode 34 corresponds to diagonal intra-prediction from the upper-left corner and intra-prediction mode 66 corresponds to diagonal intra-prediction from the upper-right corner.

[0056] For intra-prediction modes greater than 34 (corresponding to an intra-angular prediction direction clockwise from the diagonal in the upper-left corner), the selected LFNST transform matrix is ​​applied to the primary transform coefficients in a transposed manner. In one implementation, this may be performed by scanning the primary transform coefficients in a transposed direction before applying the LFNST transform. For example, from the perspective of encoder 101, if the current block is predicted using intra-prediction mode 2, the primary transform coefficients may be reordered from a two-dimensional mode within the block to a one-dimensional vector using a row-major scan before applying the selected LFNST transform matrix T. And, in that example, if the current block is instead predicted using intra-prediction mode 66 and the same written LFNST indices are used, the primary transform coefficients are instead reordered to a one-dimensional vector using a column-major scan before applying the same LFNST transform matrix T. In another implementation, the same current block with intra-prediction mode 66 may still be equivalently transformed by performing a row-major scan on the primary transform coefficients, but instead reordering the rows of transform matrix T.

[0057] More generally, the application of the LFNST transform matrix can be described as follows: Let the primary transform coefficient located at the Yth row and Xth column be p x、y and the dimension of the LFNST transform matrix T is A × B, where A is the number of secondary transform coefficients and B is the number of non-zeroed primary transform coefficients. Then, for 34 or less intra prediction modes, the primary transform coefficients p x、y A one-dimensional vector P can be constructed by any scan order operating on

[0058]

number

[0059] For intra prediction modes greater than 34, P can alternatively be constructed using the transposed scan order defined by equation (3) below.

[0060]

number

[0061] The forward LFNST transform can be understood as a matrix multiplication S = TP, where S is a one-dimensional vector of secondary transform coefficients. In practice, this transform is realized as S = n(TP), because all multiplications are implemented using integer arithmetic, and n() represents the normalization operation required for the integerized LFNST to approximate the ideal transform represented in floating-point. The secondary transform coefficients are written back into the transform block in forward diagonal scan order. Using the same sign notation introduced above and the convention that the (0,0) location corresponds to "low frequency" or "DC" in the conventional DCT, we can define the scan order s according to equation (4) below.

[0062]

number

[0063] From the decoder 201's perspective, the inverse LFNST transform is a matrix multiplication P=N(T T S), in other words, the inverse transformation is performed by transposing the matrix T.

[0064] By transposing the primary transform coefficients for intra prediction modes greater than 34, the same LFNST transform matrix can be shared in symmetric intra angular prediction directions.

[0065] In post-VVC research efforts, extensions to LFNST have been proposed and integrated into ECM. The LFNST tool in ECM achieves improved coding gain by relaxing some of the complexity reduction restrictions imposed on the original LFNST tool employed in VVC.

[0066] ECM has three LFNST kernels. Similar to the LFNST tool in VVC, in most cases, most of the transform block is zeroed out, as shown in FIGS. 9A-9C. For example, FIG. 9A shows a schematic diagram 900 of LFNST kernels for 4×N and N×4 block sizes in ECM, according to some embodiments of the present disclosure. FIG. 9B shows a schematic diagram 901 of LFNST kernels for 8×N and N×8 block sizes in ECM, according to some embodiments of the present disclosure. FIG. 9C shows a schematic diagram 903 of LFNST kernels for 16×16 block sizes in ECM, according to some embodiments of the present disclosure.

[0067] 9A-9C, the shaded areas represent primary transform coefficient positions in the ECM that are acted upon by LFNST, while the white areas represent transform coefficient positions that are zeroed out. For blocks of size 4xN or Nx4 (where N≧4), a small LFNST kernel is used for the primary transform coefficients of the 4x4 block in the upper left corner. For blocks of size 8xN or Nx8 (where N≧8), a medium LFNST kernel is used for the primary transform coefficients of the four 4x4 blocks in the upper left corner. For blocks of 16x16 or larger, a large LFNST kernel is used for the primary transform coefficients of the six 4x4 blocks in the upper left corner.

[0068] For the small LFNST kernel, the size of the LFNST kernel in ECM is 16 × 16 × 3 × 35; for the medium LFNST kernel, the size of the LFNST kernel in ECM is 64 × 32 × 3 × 35; and for the large LFNST kernel, the size of the LFNST kernel in ECM is 96 × 32 × 3 × 35. Compared to the LFNST tool in VVC, the range of LFNST indices written increases from 2 to 3, and the number of LFNST transform sets increases from 4 to 35. This means that there are 35 LFNST transform matrices for each of the three indices. The mapping from intra prediction modes to LFNST transform set indices is shown in Table 2. Similar to the LFNST tool in VVC, when the intra prediction mode is greater than 34, the primary transform coefficients are transposed.

[0069] [Table 2] The complexity burden of the LFNST tool can be evaluated in three ways. First, there is an additional storage burden imposed on the decoder 201 because the decoder 201 must store the LFNST kernel. Second, there is the worst-case number of multiplications per sample that the decoder 201 must perform when the LFNST tool is enabled. Third, there is the additional number of multiplications per sample that the encoder 101 must perform when a full search is performed for the LFNST tool. According to these three ways, the enhanced LFNST proposed in ECM is more complex than the LFNST of VVC. However, in terms of the total number of multiplications per sample, the worst-case decoder complexity may still be lower than the worst-case decoder complexity of other transform options.

[0070] In a matrix multiplication implementation of the DCT applied to each MxN transform, the number of multiplications per sample is (M + N). Therefore, the worst-case complexity occurs when (M + N) is at its maximum value. In practice, the complexity can be reduced by alternative implementations of the DCT (e.g., butterfly factorization), but it is still useful to evaluate the complexity of matrix multiplication implementations. In ECM, the separable DCT is extended to a 128-point DCT, and the largest transform is the 128-point DCT. The worst-case complexity of the separable DCT can then be 128 + 128 = 256 multiplications per sample.

[0071] The worst-case decoder complexity of LFNST for ECM can be evaluated by considering a variety of different block sizes. For a fair comparison, the evaluation includes the cost of performing the primary transforms. For a 4x4 block, the primary transforms involve 4 + 4 = 8 multiplications per sample. The LFNST involves 16x16 matrix multiplications, i.e., 16 multiplications per sample. Thus, the overall cost of LFNST for a 4x4 block is 24 multiplications per sample.

[0072] For a 4x8 block, a simple implementation of the DCT primary transform typically requires eight 4x4 transforms along the short dimension and four 8x8 transforms along the long dimension, resulting in a total of 4 + 8 = 12 multiplications per sample. However, because the LFNST reconstructs nonzero coefficient values ​​for primary transform coefficient positions in only the upper-left 4x4 block, an optimized decoder can take advantage of this by performing only four 4x4 transforms along the short dimension and then four 4x8 transforms along the long dimension, resulting in a total of 2 + 4 = 6 multiplications per sample. Because the order of separable transforms is generally fixed, in the worst-case scenario, the decoder 201 can first perform four 4x8 transforms along the long dimension. Then, the decoder 201 can perform eight 4x4 transforms along the short dimension, resulting in only 4 + 4 = 8 multiplications per sample. The LFNST is still a 16x16 matrix multiplication, but the cost is spread across larger blocks, resulting in eight multiplications per sample. The worst-case cost of an LFNST for a 4x8 block is then 16 multiplications per sample. The same principle applies in general to 4xN or Nx4 block sizes. Therefore, the number of multiplications per sample for a 4xN or Nx4 block will always be less than or equal to the number of multiplications per sample for a 4x4 block.

[0073] For an 8x8 block, the primary transform involves 8 + 8 = 16 multiplications per sample. The LFNST involves 64x32 matrix multiplications, i.e., 32 multiplications per sample, so the overall cost of the LFNST for an 8x8 block is 48 multiplications per sample.

[0074] For an 8x16 block, assume again that decoder 201 exploits the zero-out property of LFNST reconstruction. Only the primary transform coefficient position of the 8x8 block in the upper left corner is nonzero. Decoder 201 can exploit this, i.e., perform only eight 8x8 transforms along the short dimension. Next, decoder 201 can perform eight 8x16 transforms along the long dimension. This results in a total of 4 + 8 = 12 multiplications per sample. Alternatively, decoder 201 can first perform eight 8x16 transforms along the long dimension. Then, decoder 201 can perform 16 8x8 transforms along the short dimension. This results in 8 + 8 = 16 multiplications per sample. LFNST further adds an additional (64 x 32) / (8 x 16) = 16 multiplications per sample, resulting in an overall worst-case complexity of 32 multiplications per sample. As mentioned above, the number of multiplications per sample for an 8xN or Nx8 block is always less than or equal to the number of multiplications per sample for an 8x8 block.

[0075] For 16x16 blocks, the zero-out property of LFNST reconstruction means that only six 4x4 blocks of primary transform coefficients (as shown in Figure 9C) have nonzero values ​​within a mode. For simplicity, we assume a more relaxed mode in which the primary transform position of the 12x12 block in the upper left corner can have nonzero values. First, the decoder can exploit this fact to perform 12 12x16 transforms along only one dimension. Then, the decoder 201 can perform 16 12x16 transforms along the second dimension, which involves 9 + 12 = 21 multiplications per sample. LFNST involves (96 x 32) / (16 x 16) = 12 multiplications per sample, resulting in an overall complexity of 33 multiplications per sample.

[0076] For an M×N block (M, N≧16), decoder 201 can first perform twelve 12×M transforms along one dimension. Then, decoder 201 can perform M 12×N transforms along the second dimension, resulting in (12×12) / N+12 multiplications per sample, resulting in a separable DCT. The worst-case complexity, when N takes its minimum value of 16, requires 21 multiplications per sample, which is equivalent to the complexity of a 16×16 block. LFNST also adds an additional (96×32) / (M×N) multiplications per sample, which is always less than or equal to the number of multiplications per sample for a 16×16 block. Therefore, for larger M×N block sizes, the overall complexity of LFNST in ECM is always less than or equal to the number of multiplications per sample for a 16×16 block.

[0077] After comprehensively evaluating the decoder complexity for different block sizes of the LFNST for ECM, we show that the worst-case complexity is 48 multiplications per sample (occurring in the case of an 8 × 8 block). This worst-case complexity includes the cost of implementing the separable DCT using a matrix multiplication implementation. However, due to the zero-out optimization of the LFNST, this worst-case complexity is significantly lower than the worst-case complexity of implementing only the separable DCT (estimated to be 256 multiplications per sample). Assuming a more practical implementation of the separable DCT using butterfly factorization is used, the worst-case complexity of the LFNST still occurs in the case of an 8 × 8 block, and the cost is the sum of the 32 multiplications per sample for the LFNST plus the cost of the butterfly DCT. In this case, the LFNST is probably worst-case compared to the cost of the butterfly DCT applied to each 256 × 256 block.

[0078] As mentioned above, the use of non-separable quadratic transforms can significantly reduce complexity by using zero-outs in selected primary transform coefficient regions. However, further coding is possible using non-separable primary transforms (NSPTs). In early work on non-separable primary transforms, the realized transforms were complex and the kernel weights were obtained by overfitting a test dataset, but significant gains (average rate reduction of 3.43% based on the Bjontegaard index) were achievable.

[0079] We propose a practical implementation of NSPT. For example, NSPT can be applied only to a small set of block sizes: 4x4, 4x8, 8x4, and 8x8. For these block sizes, NSPT replaces the primary transform and LFNST. Similar to LFNST, we use both the written index and the implicit selection of intra-prediction modes to guide the selection of an appropriate matrix for a particular block and train the NSPT kernel. We propose four NSPT kernels: a small NSPT kernel with dimensions 16x16x3x35 for 4x4 blocks; a medium NSPT kernel with dimensions 32x20x3x35 for 4x8 and 8x4 blocks; and a large NSPT kernel with dimensions 64x32x3x35 for 8x8 blocks.

[0080] According to the present disclosure, zeroing can be defined as a reduction in the input dimension of a transform kernel, which corresponds to the first dimension in the notation for transform kernel dimensions in the present disclosure. Reducing the input dimension of a forward transform is equivalent to reducing the support of the transform. For example, zeroing out of an LFNST corresponds to a reduction in the number of DCT primary transform coefficients that the forward LFNST operates on to generate secondary transform coefficients. In the proposed NSPT, the transform operates directly on the residual coefficients, so reducing the first dimension of the NSPT kernel involves reducing the number of residual coefficients that the forward NSPT operates on to generate primary transform coefficients. In the above-mentioned known proposal, the size of the first dimension of each NSPT kernel is always equal to the number of samples in a block, so zeroing out as defined in the present disclosure is not used. However, in the known proposal, zeroing out is instead defined as a reduction in the output dimension of a transform kernel, which corresponds to the second dimension in the notation for transform kernel dimensions in the present disclosure. This definition is unambiguous in the proposal, because no reduction is yet performed on the input side of the NSPT kernel, providing an example of the more commonly used term "zeroing out." However, for consistency and clarity in this disclosure, "zero-out" is defined to describe the reduction of the input dimension of a transform kernel, while the reduction of the output dimension is denoted as a non-spanning or lossy transform. For medium-sized NSPT kernels and large-sized NSPT kernels, the second dimension is smaller than the first dimension, which means that the NSPT in these cases is a lossy transform.

[0081] Similar to LFNST, an NSPT index is written to the bitstream, which can take values ​​of 0, 1, 2, or 3. 0 indicates no NSPT is used for the transform block, and values ​​1 through 3 indicate the selection along the third dimension within the corresponding NSPT kernel. The selection along the fourth dimension of the NSPT kernel is determined by the intra prediction mode mapping, as shown in Table 3. The method is the same as that for extended LFNST in ECM. Similar to LFNST, if the intra prediction mode is greater than 34 (which means the intra angle direction is clockwise relative to the upper-left diagonal), the input of the transform is transposed. However, in the case of NSPT, the input contains residual coefficients instead of primary transform coefficients.

[0082] [Table 3]

[0083] For an M×N residual block, if the intra prediction mode is 34 or less, the residual sample located at the y-th row and x-th column is denoted by r x,y Let us denote the NSPT transform matrix T selected from the NSPT kernel for an M×N block as A×B in dimension, where A is the number of primary transform coefficients and B=M×N is the number of residual samples in the block. Then, a one-dimensional vector R can be constructed according to any scan order operating on the residual samples, as defined in Equation (5) below.

[0084]

number

[0085] For more than 34 intra-prediction modes, R may alternatively be constructed using a transposed scan order based on equation (6).

[0086]

number

[0087] Additionally, if the intra prediction mode is greater than 34, the NSPT transform matrix T is alternatively selected from the NSPT kernel for blocks of NxM shape. For square block shapes, this is the same kernel. However, if the block size is 4x8 or 8x4, the transform matrix is ​​selected from a different NSPT kernel.

[0088] The forward NSPT transform can be realized as P = n(TR), where P is a one-dimensional vector of NSPT transform coefficients and n() represents the normalization operation required to approximate the ideal transform represented in floating-point. The transform coefficients are written back into the transform block in forward diagonal scan order. Following the same sign notation introduced above and the convention that the (0,0) location corresponds to "low frequency" or "DC" in the conventional DCT, we describe the scan order based on equation (7).

[0089]

number

[0090] From the decoder 201's perspective, the inverse NSPT transform is a matrix multiplication R=n(T T P), in other words the inverse transformation is performed by transposing the matrix T.

[0091] The kernel size and block size for which NSPT is enabled are designed to ensure that NSPT is practical. This can be confirmed by comparing the complexity of NSPT for each block size with the complexity of the corresponding LFNST it replaces. For 4x4 blocks, the complexity of NSPT is 16 multiplications per sample, while the complexity of LFNST is 24 multiplications per sample. For 4x8 and 8x4 blocks, the complexity of NSPT is 20 multiplications per sample, while the complexity of LFNST is 16 multiplications per sample. For 8x8 blocks, the complexity of NSPT is 32 multiplications per sample, while the complexity of LFNST is 48 multiplications per sample. In some cases, the complexity of NSPT is lower than the complexity of the LFNST it replaces, and in other cases, NSPT is more complex. Therefore, the complexity of NSPT is designed so that it does not increase the burden on the encoder. From the decoder 201's perspective, the worst-case complexity of NSPT, which requires 32 multiplications per sample, is still lower than the worst-case complexity of all transform options the decoder must support. Therefore, the worst-case decoder complexity does not increase.

[0092] To improve the coding performance of the NSPT tool, extensions to larger block sizes are envisioned. NSPT can be extended beyond the original set of block sizes: 4x4, 4x8, 8x4, and 8x8. For block sizes of 4x16, 16x4, 8x16, and 16x8, NSPT is additionally applied instead of LFNST. For block sizes of 4x16 and 16x4, an NSPT kernel with dimensions 64x24x3x35 is used. For block sizes of 8x16 and 16x8, an NSPT kernel with dimensions 128x40x3x35 is used.

[0093] As shown in Tables 1, 2, and 3 above, both the LFNST and NSPT tools use a mapping from intra prediction modes to transform set indices, allowing the selection of a particular transform to be guided by the directionality of the intra prediction mode. When using simple intra prediction modes, such as planar (numbered 0) or DC (numbered 1) modes, this mapping supports default transform selection. However, most benefit is gained from mapping from intra angle prediction modes numbered 2 through 66 and wide-angle intra angle prediction modes numbered less than 0 or greater than 66.

[0094] In addition to the above intra prediction modes, VVC also includes several intra prediction tools, including intra block copy (IBC) and matrix-weighted intra prediction (MIP). Because IBC and MIP are not related to a specific intra angle prediction direction, when combining either of these prediction methods with LFNST or NSPT, the intra prediction mode must be explicitly determined to select an LFNST or NSPT matrix. In VVC, when IBC is selected or MIP is selected for a small CU size, LFNST is disabled. For a large CU size, LFNST is enabled for MIP, and the intra prediction mode is set to planar to select an LFNST matrix.

[0095] ECM's research efforts include more intra prediction tools, such as intra template matching prediction (IntraTMP), decoder-side intra mode derivation (DIMD), and template-based intra mode derivation (TIMD).

[0096] MIP is an intra prediction method that generates a predictor block for a CU using trained weights. Similar to the intra-angle prediction mode, MIP also generates a prediction for the current CU using an upper neighboring reference sample and a left neighboring reference sample. Intra-angle prediction generates a prediction block by replicating reference samples along the direction of the angle mode, making prediction relatively simple. In contrast, MIP can predict textures and patterns by utilizing offline training and how to generate a prediction block with higher degrees of freedom based on reference samples.

[0097] To limit the complexity of the implementation, the MIP operation 1000 includes averaging and interpolation as shown in Figure 10. Averaging (at 1001) averages the reference samples to obtain up to eight downsampled reference samples bdry red The matrix-vector multiplication (at 1003) considers only the case where MIP is applied to large block sizes, and produces a reduced prediction signal pred with dimensions 8×8 based on equation (8). red is generated.

[0098] pred red =A k bdry red+ b k (8) where A k is a matrix of dimensions 64x8, and b k is an offset vector of length 64.

[0099] Interpolation (at 1005) uses linear interpolation to expand the 8x8 reduced prediction signal to the overall dimension of the CU.

[0100] A particular matrix A is selected from the set by writing the value k into the bitstream as the MIP syntax element modeId. k and offset b k When the MIP isTransposed flag is written, the MIP matrices are also reused in transposed form.

[0101] In summary, there are three sets containing MIP matrices and offset vectors, covering a range of block sizes. Set S0 contains 16 matrices:

number

number

number

number

number

number

[0102] FIG. 11 shows a schematic diagram 1100 of an IBC method according to some embodiments of the present disclosure. Referring to FIG. 11, when predicting a CU 1102 using intra block copy mode, a block vector 1104 is written to indicate which block in the same picture is copied and used as the predictor for the current block. Writing the block vector 1104 can be performed by writing a block vector difference (BVD) to the bitstream, allowing the decoder 201 to determine the block vector 1104 by adding the BVD to the block vector predictor. Alternatively, if a block vector from a previous CU exactly matches the current block vector (block vector 1104), the block vector can be written using a merge flag. Regardless of the signaling mechanism, the block vector 1104 points to a location in the same picture and indicates a sample block whose size is equal to that of the current CU, which is used as the IBC predictor block 1106 for the current CU (CU 1102). Some restrictions can be applied to the block vector 1104, since it must point to a position in the current image that was decoded before the current CU.

[0103] FIG. 12 shows a schematic diagram 1200 of an intra-template matching prediction (IntraTMP) method according to some embodiments of the present disclosure. Referring to FIG. 12, IntraTMP is similar to IBC because the current CU is also predicted by a sample block from the current image. However, unlike IBC, block vectors are not written to the bitstream. Instead, the decoder 201 compares L-shaped or other shaped templates of reconstructed samples adjacent to the current CU with L-shaped templates of candidate predictors within a predetermined search area. The IntraTMP predictor block is determined by finding the best candidate template that matches the current CU template. The best match can be determined by finding the template that minimizes the sum of absolute differences (SAD) or the sum of absolute transformed differences (SATD), or by comparing hashes between the templates. The traversal search algorithm of the search space can be exhaustive (e.g., scanning the template in the search space with a shift in sample resolution) or fast (e.g., first performing a coarse search, then performing a local fine search with a best match of the coarse search). In either case, the encoder 101 and the decoder 201 can perform the search algorithm in the same manner, whereby the IntraTMP predictor is implicitly known to the encoder 101 and the decoder 201 without signaling in the bitstream. In Figure 12, the current CU template 1202 and the best matching template 1204 are shown with hatched shading, and the other templates 1206 in the search space are shown with dashed shading.

[0104] FIG. 13 illustrates a schematic diagram 1300 of a decoder-side intra-mode derivation (DIMD) method according to some embodiments of the present disclosure.

[0105] Referring to FIG. 13, in DIMD, an intra-prediction mode (or multiple intra-prediction modes) is implicitly derived from an L-shaped template 1304 (hereinafter referred to as "template 1304") of reconstructed samples neighboring the current CU 1302. The size of the template 1304 is three samples wide. The decoder 201 moves a 3x3 gradient analysis window 1306 over the template 1304. At each position, a local gradient is calculated by applying a Sobel filter. The 3x3 sample set at a position within the template 1304 is called T K Assuming this, we can write the Sobel filter according to equation (9).

[0106]

number

[0107] Then, the horizontal gradient G is calculated by taking the dot product shown in the following equations (10) and (11). k、x and the vertical gradient G k、y Estimate.

[0108] G k,x =T k M x (10) G k,y =T k M y (11)

[0109] The local gradient size G according to equations (12) and (13), respectively. k and the local gradient angle θ k can be estimated.

[0110]

number

[0111] Local gradient angle θ k is the intra-angle prediction direction IPM kFor example, an angle of 0 degrees corresponds to horizontal intra prediction mode 18. In fact, the decoder 201 k、x and G k、y IPM using a fast implementation method (e.g., table lookup) based on k can be directly estimated. At the start of the DIMD method, an empty histogram H (with each entry filled with zeros) is initialized with a size equal to the number of intra-prediction modes. As the DIMD method performs gradient analysis at each local window, the histogram H is updated according to Equation (14).

[0112] H[IPM k ]+=G k (14)

[0113] Thus, each local gradient analysis "votes" for an intra-prediction mode. At the end of the DIMD method, the intra-prediction mode with the highest count in H may be selected as the single representative intra-prediction mode for the current CU. Alternatively, multiple intra-prediction modes may be derived from H, in order of highest count.

[0114] 14 shows a schematic diagram 1400 of a template-based intra-mode derivation (TIMD) method according to some embodiments of the present disclosure. Referring to FIG. 14, in TIMD, an intra-prediction mode (or multiple intra-prediction modes) is implicitly derived from templates 1404 above and to the left of reconstructed samples neighboring a current CU 1402.

[0115] A set of candidate intra-prediction modes is searched from a most probable mode (MPM) list, which is constructed from the intra-prediction modes used by neighboring CUs. Then, for each candidate intra-prediction mode, an intra-angle prediction method generates a prediction of the template 1404 based on the template reference sample 1406. The candidate intra-prediction mode that generates a template predictor that best matches the template 1404 is selected as the TIMD intra-prediction mode. The best match can be determined by finding the predictor that minimizes the sum of absolute errors (SAD) or the sum of absolute transformed errors (SATD), or by comparing hashes between the predictor and the template. Alternatively, multiple intra-prediction modes may be obtained in increasing order of SAD / SATD.

[0116] The MIP method, the IBC method, and the IntraTMP method can be effective intra prediction tools. Greater gains can be achieved by using a combination of MIP with NSPT and LFNST, a combination of IBC with LFNST and NSPT, and a combination of IntraTMP with LFNST and NSPT. Here, the intra prediction mode can be derived by the solution shown in FIG. 4.

[0117] The transform module 308 can convert the video signal in the residual block from the pixel domain to a transform domain (e.g., the frequency domain, depending on the transform method). In some examples, the transform module 308 can be skipped, and the video signal is not converted to the transform domain.

[0118] The quantization module 310 may be configured to quantize a coefficient at each position within a coding block to generate a quantization level for the position. The current block may be a residual block. That is, the quantization module 310 may perform a quantization process on each residual block. The residual block may include N×M positions (samples), each associated with a transformed or untransformed video signal / data (e.g., luma and / or chroma information), where N and M are positive integers. In this disclosure, before quantization, a transformed or untransformed video signal at a particular position is referred to herein as a "coefficient." After quantization, the quantized value of the coefficient is referred to herein as a "quantization level" or "level."

[0119] Quantization can be used to reduce the dynamic range of a transformed or untransformed video signal, thereby using fewer bits to represent the video signal. Quantization typically involves dividing by a quantization step size followed by rounding, while inverse quantization involves multiplying by the quantization step size. The quantization step size may be indicated by a quantization parameter (QP). This type of quantization process is called scalar quantization. Quantization of all coefficients within a coding block can be performed independently, and this type of quantization method is used in several existing video compression standards (e.g., H.264 / AVC and H.265 / HEVC). The QP in quantization can affect the bitrate used to encode / decode the video images. For example, a higher QP may result in a lower bitrate, and a lower QP may result in a higher bitrate.

[0120] For an N×M coding block, a specific coding scan order can be used to convert the two-dimensional (2D) coefficients of the block into a one-dimensional (1D) order for quantization and encoding of the coefficients. Typically, the coding scan starts from the upper-left corner of the coding block and stops at the lower-right corner or at the last non-zero coefficient / level in the lower-right direction. It should be understood that the coding scan order can include any suitable order, such as a zig-zag scan order, a vertical (column) scan order, a horizontal (row) scan order, a diagonal scan order, or any combination thereof. Quantization of coefficients in the coding block can utilize information from the coding scan order. For example, the quantization can depend on the state of the previous quantization level along the coding scan order. To further improve coding efficiency, the quantization module 310 can use multiple quantizers (e.g., two scalar quantizers). Which quantizer to use to quantize the current coefficient can be determined based on information from the previous coefficient in the coding scan order. This quantization process is called dependent quantization.

[0121] Referring to FIG. 3, the encoding module 320 may be configured to encode the quantization level of each position within the coding block into a bitstream. In some embodiments, the encoding module 320 may perform entropy encoding on the coding block. The entropy encoding may convert each quantization level into a corresponding binary representation (e.g., binary bin) using various binarization methods (e.g., Golomb-Rice binarization). The binary representation may then be further compressed using an entropy encoding algorithm. The compressed data may be added to the bitstream. In addition to the quantization levels, the encoding module 320 may further encode various information, such as block type information of the coding unit, prediction mode information, partition unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information, input from the prediction modules 304 and 306. In some embodiments, the encoding module 320 may perform residual encoding on the coding block to convert the quantization levels into a bitstream. For example, after quantization, there may be N×M quantization levels in an N×M block. These NxM levels can be zero or non-zero values. If the non-zero levels are not binary, they can be further binarized into binary bins, for example, using unified Truncated Rice (TR) and restricted EGk binarization.

[0122] Non-binary syntax elements can be mapped to binary codewords. The bijective mapping between codes and codewords (usually using simple structured codes) is called binarization. Binary arithmetic coding can encode binary syntax elements and binary symbols (also called bins) for non-binary data codewords. The core coding engine of context-adaptive binary arithmetic coding (CABAC) can support two modes of operation: a context coding mode that encodes bins with an adaptive probability model, and a relatively low-complexity bypass mode that uses a fixed probability of 1 / 2. The adaptive probability model is also called a context, and the assignment of a probability model to each bin is called context modeling.

[0123] 3, the inverse quantization module 312 may be configured to inverse quantize the quantization levels by the inverse quantization module 312, and the inverse transform module 314 may be configured to inverse transform the coefficients transformed by the transform module 308. The reconstructed residual blocks generated by the inverse quantization module 312 and the inverse transform module 314 may be combined with the prediction units predicted by the prediction modules 304 or 306 to generate reconstructed blocks.

[0124] The filter module 316 may include at least one deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can remove block artifacts generated by boundaries between blocks in the reconstructed image. The SAO module can compensate for the offset of the deblocked video relative to the original video on a pixel-by-pixel basis. The ALF can be performed based on a value obtained by comparing the reconstructed, filtered video with the original video. The buffer module 318 can be configured to store the reconstructed blocks or images calculated by the filter module 316 and can provide the reconstructed, stored blocks or images to the inter prediction module 304 when performing inter prediction.

[0125] FIG. 4 illustrates a detailed block diagram of an exemplary decoder 201 in the decoding system 200 of FIG. 2 , according to some embodiments of the present disclosure. As illustrated in FIG. 4 , the decoder 201 may include a decoding module 402, an inverse quantization module 404, an inverse transform module 406, an inter-prediction module 408, an intra-prediction module 410, a filter module 412, and a buffer module 414. It should be understood that each element illustrated in FIG. 4 is shown independently to represent different characteristic functions in a video decoder, and is not intended to imply that each component is formed by a separate hardware or software unit. That is, for convenience of explanation, each element is described as an independent element. At least two elements may be combined to form a single element, and one element may be divided into multiple elements to perform a function. Furthermore, it should be understood that some elements are not essential elements for performing the functions described in the present disclosure, but are optional elements for improving performance. It should also be understood that these elements may be implemented using electronic hardware, firmware, computer software, or any other combination. Whether these elements are implemented as hardware, firmware, or software depends on the particular application and design constraints imposed on the decoder 201 .

[0126] When receiving a video bitstream from a video encoder (e.g., encoder 101), the input bitstream may be decoded by decoder 201 in a procedure complementary to that of the video encoder. Therefore, for convenience of explanation, some decoding details described above with respect to encoding may be omitted. The decoding module 402 may be configured to decode the bitstream to obtain various information encoded in the bitstream, such as the quantization level for each position within a coding block. In some embodiments, the decoding module 402 may perform entropy decoding (decompression) corresponding to the entropy encoding (compression) performed by the encoder, and may obtain a binary representation (e.g., binary bins) using, for example, video local-area network (VideoLAN) coding (VLC), context-adaptive variable-length coding (CAVLC), CABAC, syntax-based binary arithmetic coding (SBAC), PIPE coding, etc. The decoding module 402 can further convert the binary representation into quantization levels using Golomb-Rice binarization (e.g., including EGk binarization and combined TR and constrained EGk binarization). In addition to the quantization levels of the positions within the transform unit, the decoding module 402 can decode various other information, such as parameters used for Golomb-Rice binarization (e.g., Rice parameters), block type information of the coding unit, prediction mode information, partition unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. During the decoding process, the decoding module 402 can perform reordering on the bitstream to reconstruct and rearrange the data from a 1D order into 2D reordered blocks in a reverse scan manner based on the coding scan order used by the encoder.

[0127] The inverse quantization module 404 may be configured to inverse quantize the quantization level of each position of the coding block (e.g., a 2D reconstruction block) to obtain a coefficient for each position. In some embodiments, the inverse quantization module 404 may perform dependent inverse quantization based on quantization parameters provided by an encoder, the quantization parameters including information related to the quantizers used in the dependent quantization, such as the quantization step size used by each quantizer.

[0128] The inverse transform module 406 may be configured to perform an inverse transform, performing an inverse DCT, an inverse DST, an inverse KLT, an inverse LFNST, or an inverse NSPT for the DCT, the discrete sine transform (DST), the KLT, the LFNST, or the NSPT, respectively, performed by the encoder to transform data back from the transform domain (e.g., coefficients) to the pixel domain (e.g., luma and / or chroma information). In some embodiments, the inverse transform module 406 can selectively perform a transform operation (e.g., DCT, DST, KLT, LFNST, NSPT) based on multiple pieces of information (e.g., prediction method, size of the current block, prediction direction, etc.).

[0129] In some implementations, whether to apply LFNST and / or NSPT to transform the residual block may be determined based on the intra-prediction mode information of the prediction unit used to generate the residual block and an index in the bitstream. For example, the inverse transform module 406 may use the index to select a transform matrix set from the three matrix transform sets described above. The inverse transform module 406 may use the intra-prediction mode information to select a transform matrix from the selected transform matrix set (selecting from among 35 transform matrices). Non-limiting examples are not provided.

[0130] For example, in some implementations, the inverse transform module 406 may determine to enable LFNST for a CU predicted by IBC. Here, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an LFNST matrix corresponding to the IBC intra-prediction mode from the LFNST matrix set. In this example and the following examples, index 1 may correspond to the first transform matrix set, index 2 may correspond to the second transform matrix set, and index 3 may correspond to the third transform matrix set.

[0131] In some implementations, when enabling LFNST for a CU predicted by IBC, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index, and select from the LFNST matrix set an LFNST matrix that corresponds to the intra-prediction mode derived by the identified DIMD.

[0132] In some implementations, when enabling LFNST for a CU predicted by IBC, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index, and select from the LFNST matrix set an LFNST matrix that corresponds to the intra-prediction mode derived by the identified TIMD.

[0133] In some implementations, the inverse transform module 406 may determine to enable NSPT for a CU predicted by IBC, where the inverse transform module 406 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an NSPT matrix corresponding to IBC from the NSPT matrix set.

[0134] In some implementations, when NSPT is enabled for a CU predicted by IBC, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to the intra-prediction mode derived by the identified DIMD.

[0135] In some implementations, when NSPT is enabled for a CU predicted by IBC, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set based on an intra-prediction mode derived by the identified TIMD.

[0136] In some implementations, if NSPT is not enabled for a CU predicted by IBC, the inverse transform module 406 may estimate the index as 0 without writing the NSPT index.

[0137] In some implementations, if NSPT is not enabled for a CU predicted by IBC, an LFNST index is written instead, where the inverse transform module 406 may use any one of the aforementioned implementations for selecting an LFNST matrix.

[0138] In some implementations, the inverse transform module 406 may determine to enable LFNST for a CU predicted by IntraTMP, where the inverse transform module 406 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an LFNST matrix from the LFNST matrix set that corresponds to an identified IntraTMP intra-prediction mode (e.g., planar mode).

[0139] In some implementations, when enabling LFNST for a CU predicted by IntraTMP, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index, and select from the LFNST matrix set an LFNST matrix that corresponds to the intra prediction mode derived by the identified TIMD.

[0140] In some implementations, if LFNST is not enabled for a CU predicted by IntraTMP, the LFNST index may not be included in the bitstream and the inverse transform module 406 may infer the index as 0.

[0141] In some implementations, the inverse transform module 406 may determine to enable NSPT for a CU predicted by IntraTMP, where the inverse transform module 406 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an NSPT matrix from the NSPT matrix set that corresponds to an identified IntraTMP intra-prediction mode (e.g., planar mode).

[0142] In some implementations, when NSPT is enabled for a CU predicted by IntraTMP, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to the intra prediction mode derived by the identified DIMD.

[0143] In some implementations, when NSPT is enabled for a CU predicted by IntraTMP, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to the intra prediction mode derived by the identified TIMD.

[0144] In some implementations, if NSPT is not enabled for a CU predicted by IntraTMP, the NSPT index may not be included in the bitstream and the inverse transform module 406 may infer the index as 0.

[0145] In some implementations, if NSPT is not enabled for a CU predicted by IntraTMP, an LFNST index can be written instead, where the inverse transform module 406 may use any one of the aforementioned implementations for selecting an LFNST matrix.

[0146] In some implementations, the inverse transform module 406 may determine to enable LFNST for a CU predicted by MIP, where the inverse transform module 406 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3) and an intra-prediction mode derived by the identified TIMD.

[0147] In some implementations, if LFNST is not enabled for a CU predicted by MIP, the LFNST index may not be included in the bitstream and the inverse transform module 406 may infer the index as 0.

[0148] In some implementations, when NSPT is enabled for a CU predicted by MIP, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to an identified MIP intra-prediction mode (e.g., a planar mode).

[0149] In some implementations, the inverse transform module 406 may determine to enable NSPT for a MIP-predicted CU, where the inverse transform module 406 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3) and an identified MIP intra-prediction mode (e.g., planar mode).

[0150] In some implementations, when NSPT is enabled for a CU predicted by MIP, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to the intra-prediction mode derived by the identified DIMD.

[0151] In some implementations, when NSPT is enabled for a CU predicted by MIP, the inverse transform module 406 can select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to the intra-prediction mode derived by the identified TIMD.

[0152] In some implementations, if NSPT is not enabled for a CU predicted by MIP, the NSPT index may not be included in the bitstream and the inverse transform module 406 may infer the index as 0.

[0153] In some implementations, if NSPT is not enabled for a CU predicted by MIP, an LFNST index can be written instead, where the inverse transform module 406 may use any one of the aforementioned implementations for selecting an LFNST matrix.

[0154] In some implementations, LFNST and / or NSPT may be enabled for all allowed CU sizes predicted by IBC, IntraTMP, and MIP, where the LFNST / NSPT matrices can be selected based on any of the aforementioned implementations.

[0155] In some implementations, LFNST and NSPT may be enabled for CUs predicted by IBC, IntraTMP, and MIP. Here, the inverse transform module 406 may determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is any of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the inverse transform module 406 may determine to use NSPT; otherwise, it may use LFNST. Similarly, the LFNST / NSPT matrix may be selected based on any of the aforementioned implementations.

[0156] The inter prediction module 408 and the intra prediction module 410 may be configured to generate a prediction block based on information related to the generation of the prediction block provided by the decoding module 402 and information of a previously decoded block or image provided by the buffer module 414. As described above, when performing intra prediction in a manner similar to the operation of an encoder, if the size of the prediction unit and the size of the transform unit are the same, intra prediction may be performed on the prediction unit based on pixels located to the left, upper left, and top of the prediction unit. However, if the size of the prediction unit and the size of the transform unit are different when performing intra prediction, intra prediction may be performed using reference pixels by the transform unit.

[0157] A reconstructed block or image that combines the outputs of the inverse transform module 406 and the prediction module 408 or 410 may be provided to a filter module 412. The filter module 412 may include a deblocking filter, an offset correction module, and an ALF. A buffer module 414 may store the reconstructed image or block and use it as a reference image or block for the inter prediction module 408, which may output the reconstructed image.

[0158] Consistent with the scope of the present disclosure, the encoding module 320 and the decoding module 402 may be configured to adopt a quantization level binarization scheme, in which the Rice parameter may adapt to the bit depth and / or bit rate at which the video image is encoded to improve encoding efficiency.

[0159] 15 shows a flowchart of an example video decoding method 1500 according to some embodiments of the present disclosure. Method 1500 may be performed by a system such as decoding system 200, decoder 201, inverse transform module 406, or intra prediction module 410, for example. Method 1500 may include operations 1502-1512, as described below. It should be understood that some steps are optional and that some steps may be performed simultaneously or in a different order than that shown in FIG. 15.

[0160] 15, at 1502, the system may analyze a bitstream. For example, referring to FIGS. 2 and 4, the decoder 201 may analyze a bitstream encoded by the encoder 101. By analyzing the bitstream, the decoder 201 may obtain information indicating whether LFNST and / or NSPT are enabled for IBC, IntraTMP, and / or MIP.

[0161] In response to enabling at least one non-separable transform for the intra prediction method at 1504, the system may determine at least one non-separable transform and an intra prediction mode to be used to decode the CU of the bitstream. For example, with reference to Figures 2 and 4, the inverse transform module 406 may determine the intra prediction mode and at least one non-separable transform (e.g., LFNST and / or NSPT) to be used to decode the CU of the bitstream based on information obtained by analyzing the bitstream.

[0162] At 1506, the system may select a transform matrix from a plurality of transform matrix sets based on a bitstream index and an intra-prediction mode. For example, referring to FIG. 4, in some implementations, the inverse transform module 406 may determine to enable LFNST for a CU predicted by IBC. Here, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3) and select an LFNST matrix corresponding to the IBC intra-prediction mode from the LFNST matrix set. In this example and the following examples, index 1 may correspond to the first transform matrix set, index 2 may correspond to the second transform matrix set, and index 3 may correspond to the third transform matrix set. In some implementations, when enabling LFNST for a CU predicted by IBC, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index and select an LFNST matrix corresponding to the intra-prediction mode derived by the identified DIMD from the LFNST matrix set. In some implementations, when enabling LFNST for a CU predicted by IBC, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index, and select an LFNST matrix from the LFNST matrix set that corresponds to the intra-prediction mode derived by the identified TIMD. In some implementations, the inverse transform module 406 may determine to enable NSPT for a CU predicted by IBC. Here, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an NSPT matrix from the NSPT matrix set that corresponds to IBC.In some implementations, when NSPT is enabled for a CU predicted by IBC, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set that corresponds to the intra-prediction mode derived by the identified DIMD. In some implementations, when NSPT is enabled for a CU predicted by IBC, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set based on the intra-prediction mode derived by the identified TIMD. In some implementations, when NSPT is not enabled for a CU predicted by IBC, the inverse transform module 406 may not write an NSPT index and may estimate the index as 0. In some implementations, when NSPT is not enabled for a CU predicted by IBC, an LFNST index is instead written. Here, any one of the above implementations for selecting an LFNST matrix may be used by the inverse transform module 406. In some implementations, the inverse transform module 406 may determine to enable LFNST for a CU predicted by IntraTMP. Here, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3) and select an LFNST matrix from the LFNST matrix set that corresponds to an identified IntraTMP intra-prediction mode (e.g., planar mode). In some implementations, when enabling LFNST for a CU predicted by IntraTMP, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index and select an LFNST matrix from the LFNST matrix set that corresponds to an identified TIMD-derived intra-prediction mode.In some implementations, when LFNST is not enabled for a CU predicted by IntraTMP, the LFNST index is not included in the bitstream, and the inverse transform module 406 may estimate the index as 0. In some implementations, the inverse transform module 406 may determine to enable NSPT for a CU predicted by IntraTMP. Here, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an NSPT matrix from the NSPT matrix set that corresponds to an identified IntraTMP intra-prediction mode (e.g., planar mode). In some implementations, when NSPT is enabled for a CU predicted by IntraTMP, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index, and select an NSPT matrix from the NSPT matrix set that corresponds to an identified DIMD-derived intra-prediction mode. In some implementations, when NSPT is enabled for a CU predicted by IntraTMP, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index, and select an NSPT matrix from the NSPT matrix set that corresponds to the intra-prediction mode derived by the identified TIMD. In some implementations, when NSPT is not enabled for a CU predicted by IntraTMP, the NSPT index may not be included in the bitstream, and the inverse transform module 406 may estimate the index as 0. In some implementations, when NSPT is not enabled for a CU predicted by IntraTMP, an LFNST index may instead be written. Here, the inverse transform module 406 may use any one of the aforementioned implementations for selecting an LFNST matrix. In some implementations, the inverse transform module 406 may determine to enable LFNST for a CU predicted by MIP.Here, the inverse transform module 406 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3) and an intra-prediction mode derived by the identified TIMD. In some implementations, if LFNST is not enabled for a MIP-predicted CU, the LFNST index may not be included in the bitstream, and the inverse transform module 406 may estimate the index as 0. In some implementations, if NSPT is enabled for a MIP-predicted CU, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index, and select an NSPT matrix from the NSPT matrix set that corresponds to the identified MIP intra-prediction mode (e.g., planar mode). In some implementations, the inverse transform module 406 may determine to enable NSPT for a MIP-predicted CU. Here, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3) and an identified MIP intra-prediction mode (e.g., planar mode). In some implementations, when NSPT is enabled for a MIP-predicted CU, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set that corresponds to the intra-prediction mode derived by the identified TIMD. In some implementations, when NSPT is enabled for a MIP-predicted CU, the inverse transform module 406 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set that corresponds to the intra-prediction mode derived by the identified TIMD. In some implementations, when NSPT is not enabled for a MIP-predicted CU, the NSPT index may not be included in the bitstream and the inverse transform module 406 may estimate the index as 0.In some implementations, if NSPT is not enabled for a CU predicted by MIP, an LFNST index can be written instead. Here, the inverse transform module 406 may use any one of the above-mentioned implementations for selecting an LFNST matrix. In some implementations, LFNST and NSPT may be enabled for CUs predicted by IBC, IntraTMP, and MIP. Here, an LFNST or NSPT matrix can be selected based on any of the above-mentioned implementations. Here, the inverse transform module 406 may determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is any of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the inverse transform module 406 may determine to use NSPT; otherwise, it may use LFNST. Similarly, an LFNST / NSPT matrix can be selected based on any of the above-mentioned implementations.

[0163] At 1508, the system can decode the CU based on the intra prediction method and the transform matrix. For example, referring to Figure 4, the decoder 201 can decode the CU based on the intra prediction method and the non-separable transform.

[0164] In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method at 1510, the system may select a first non-separable transform to be used to decode a CU if the CU is of a first size. For example, referring to FIG. 4, LFNST and / or NSPT may be enabled for CUs predicted by IBC, IntraTMP, and MIP. Here, the inverse transform module 406 may determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is any of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the inverse transform module 406 may determine to use NSPT; otherwise, it may use LFNST. Similarly, the LFNST / NSPT matrix may be selected based on any of the above implementation schemes.

[0165] In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method at 1512, the system can select a second non-separable transform to be used to decode a CU when the CU is a second size different from the first size. For example, referring to FIG. 4, LFNST and NSPT may be enabled for CUs predicted by IBC, IntraTMP, and MIP. Here, the inverse transform module 406 can determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is one of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the inverse transform module 406 can determine to use NSPT; otherwise, it can use LFNST. Similarly, the LFNST / NSPT matrix can be selected based on any of the above implementation schemes.

[0166] 16 shows a flowchart of an example video encoding method 1600 according to some embodiments of the present disclosure. Method 1600 may be performed by a system such as encoding system 100, encoder 101, transform module 308, or intra-prediction module 306. Method 1600 may include operations 1602-1610, as described below. It should be understood that some steps are optional and some steps may be performed simultaneously or in a different order than that shown in FIG. 16.

[0167] 16, in response to enabling at least one non-separable transform for an intra prediction method at 1602, the system may determine at least one non-separable transform and an intra prediction mode to be used to encode the CU into a bitstream. Referring to FIGS. 1 and 3, the transform module 308 may determine the intra prediction mode and at least one non-separable transform (e.g., LFNST and / or NSPT) to be used to encode the CU into a bitstream based on a rate-distortion optimization algorithm.

[0168] At 1604, the system may select a transform matrix from a plurality of transform matrix sets based on a bitstream index and an intra-prediction mode. For example, referring to FIG. 3, in some implementations, the transform module 308 may determine to enable LFNST for a CU predicted by IBC. Here, the transform module 308 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3) and select an LFNST matrix from the LFNST matrix set corresponding to the intra-prediction mode. In this example and the following examples, index 1 may correspond to the first transform matrix set, index 2 may correspond to the second transform matrix set, and index 3 may correspond to the third transform matrix set. In some implementations, when enabling LFNST for a CU predicted by IBC, the transform module 308 may select an LFNST matrix set based on a bitstream index and select an LFNST matrix from the LFNST matrix set corresponding to the intra-prediction mode derived by the identified DIMD. In some implementations, when enabling LFNST for a CU predicted by IBC, the transform module 308 may select an LFNST matrix set based on a bitstream index and select an LFNST matrix from the LFNST matrix set that corresponds to the intra-prediction mode derived by the identified TIMD. In some implementations, the transform module 308 may determine to enable NSPT for a CU predicted by IBC. Here, the transform module 308 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3) and select an NSPT matrix from the NSPT matrix set that corresponds to IBC. In some implementations, when enabling NSPT for a CU predicted by IBC, the transform module 308 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set that corresponds to the intra-prediction mode derived by the identified TIMD.In some implementations, when enabling NSPT for a CU predicted by IBC, the transform module 308 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set based on the intra-prediction mode derived by the identified TIMD. In some implementations, when not enabling NSPT for a CU predicted by IBC, the transform module 308 may not write an NSPT index and may estimate the index as 0. In some implementations, when not enabling NSPT for a CU predicted by IBC, an LFNST index is instead written. Here, the transform module 308 may use any one of the above-mentioned implementations for selecting an LFNST matrix. In some implementations, the transform module 308 may determine to enable LFNST for a CU predicted by IntraTMP. Here, the transform module 308 may select an LFNST matrix set based on a bitstream index (e.g., 1, 2, or 3) and select an LFNST matrix from the LFNST matrix set that corresponds to the identified IntraTMP intra-prediction mode (e.g., planar mode). In some implementations, when enabling LFNST for a CU predicted by IntraTMP, the transform module 308 may select an LFNST matrix set based on a bitstream index, and select an LFNST matrix from the LFNST matrix set that corresponds to the intra-prediction mode derived by the identified TIMD. In some implementations, when not enabling LFNST for a CU predicted by IntraTMP, the LFNST index may not be included in the bitstream, and the transform module 308 may estimate the index as 0. In some implementations, the transform module 308 may determine to enable NSPT for a CU predicted by IntraTMP.Here, the transform module 308 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3), and select an NSPT matrix from the NSPT matrix set that corresponds to the identified IntraTMP intra-prediction mode (e.g., planar mode). In some implementations, when enabling NSPT for a CU predicted by IntraTMP, the transform module 308 may select an NSPT matrix set based on a bitstream index, and select an NSPT matrix from the NSPT matrix set that corresponds to the identified TIMD-derived intra-prediction mode. In some implementations, when enabling NSPT for a CU predicted by IntraTMP, the transform module 308 may select an NSPT matrix set based on a bitstream index, and select an NSPT matrix from the NSPT matrix set that corresponds to the identified TIMD-derived intra-prediction mode. In some implementations, when not enabling NSPT for a CU predicted by IntraTMP, the NSPT index is not included in the bitstream, and the transform module 308 may estimate the index as 0. In some implementations, if NSPT is not enabled for a CU predicted by IntraTMP, an LFNST index may alternatively be written. Here, the transform module 308 may use any one of the aforementioned implementations for selecting an LFNST matrix. In some implementations, the transform module 308 may determine to enable LFNST for a CU predicted by MIP. Here, the transform module 308 may select an LFNST matrix set based on the bitstream index (e.g., 1, 2, or 3) and the intra-prediction mode derived by the identified TIMD. In some implementations, if LFNST is not enabled for a CU predicted by MIP, the LFNST index may not be included in the bitstream, and the transform module 308 may estimate the index as 0.In some implementations, when enabling NSPT for a MIP-predicted CU, the transform module 308 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set that corresponds to an identified MIP intra-prediction mode (e.g., planar mode). In some implementations, the transform module 308 may determine to enable NSPT for a MIP-predicted CU. Here, the transform module 308 may select an NSPT matrix set based on a bitstream index (e.g., 1, 2, or 3) and an identified MIP intra-prediction mode (e.g., planar mode). In some implementations, when enabling NSPT for a MIP-predicted CU, the transform module 308 may select an NSPT matrix set based on a bitstream index and select an NSPT matrix from the NSPT matrix set that corresponds to an identified DIMD-derived intra-prediction mode. In some implementations, when NSPT is enabled for a MIP-predicted CU, the transform module 308 may select an NSPT matrix set based on a bitstream index, and from the NSPT matrix set, select an NSPT matrix corresponding to the intra-prediction mode derived by the identified TIMD. In some implementations, when NSPT is not enabled for a MIP-predicted CU, the NSPT index may not be included in the bitstream, and the transform module 308 may estimate the index as 0. In some implementations, when NSPT is not enabled for a MIP-predicted CU, an LFNST index may be written instead. Here, any one of the above-mentioned implementations for selecting an LFNST matrix may be used by the transform module 308. In some implementations, LFNST and NSPT may be enabled for IBC, IntraTMP, and MIP-predicted CUs. Here, the LFNST / NSPT matrix may be selected based on any of the above-mentioned implementations.Here, the transform module 308 may determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is one of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the transform module 308 may determine to use NSPT; otherwise, it may use LFNST. Similarly, the LFNST / NSPT matrix may be selected based on any of the aforementioned implementation schemes.

[0169] At 1606, the system can encode the CU based on the intra prediction method and the transform matrix. For example, referring to Figure 3, the encoder 101 can encode the CU based on the intra prediction method and the non-separable transform.

[0170] In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method at 1608, the system may select a first non-separable transform to be used to encode a CU if the CU is of a first size. For example, referring to FIG. 3, LFNST and / or NSPT may be enabled for CUs predicted by IBC, IntraTMP, and MIP. Here, the transform module 308 may determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is any of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the transform module 308 may determine to use NSPT; otherwise, it may use LFNST. Similarly, the LFNST / NSPT matrix may be selected based on any of the above implementations.

[0171] In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method at 1610, the system can select a second non-separable transform to be used to encode a CU when the CU is a second size different from the first size. For example, referring to FIG. 3, LFNST and NSPT may be enabled for CUs predicted by IBC, IntraTMP, and MIP. Here, the transform module 308 can determine whether to use LFNST or NSPT for the current CU based on the size of the CU. By way of example and not limitation, if the size of the CU is any of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, and 16x8, the transform module 308 can determine to use NSPT; otherwise, it can use LFNST. Similarly, the LFNST / NSPT matrix can be selected based on any of the above implementations.

[0172] In each aspect of the present disclosure, the functions described herein can be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. Storage media can be any available medium that can be accessed by a processor (e.g., processor 102 in FIGS. 1 and 2). By way of example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other magnetic storage devices, flash drives, SSDs, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system (e.g., a mobile device or computer). As used herein, magnetic disks and optical disks include CDs, laser disks, optical disks, digital video disks (DVDs), and floppy disks, where magnetic disks typically reproduce data magnetically while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0173] According to one aspect of the present disclosure, a decoding method is provided that is executed by a decoder. The method may include analyzing a bitstream by a processor. The method may include determining, by the processor, at least one non-separable transform and an intra-prediction mode to be used to decode a CU of the bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. The method may include selecting, by the processor, a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. The method may include decoding, by the processor, the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, matrix MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST and NSPT.

[0174] In some implementations, selecting, by the processor, a transform matrix from a plurality of transform matrix sets based on the bitstream index and the intra-prediction mode may include selecting, by the processor, one transform matrix set from the plurality of transform matrix sets based on the bitstream index. In some implementations, selecting, by the processor, a transform matrix from a plurality of transform matrix sets based on the bitstream index and the intra-prediction mode may include selecting, by the processor, a transform matrix from the one transform matrix set based on the intra-prediction mode.

[0175] In some implementations, selecting, by the processor, a transformation matrix from the one set of transformation matrices based on an intra-prediction mode may include selecting, by the processor, a transformation matrix from the one set of transformation matrices based on an intra-prediction mode derived by DIMD or TIMD.

[0176] In some implementations, if at least one non-separable transform is not enabled for an intra prediction method, the bitstream index may be omitted from the bitstream.

[0177] In some implementations, at least one non-separable transform may be enabled for a CU of any size.

[0178] In some implementations, the at least one non-separable transform may include a first non-separable transform and a second non-separable transform.

[0179] In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the method may further include selecting, by the processor, a first non-separable transform to be used to decode the CU if the CU is a first size. In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the method may further include selecting, by the processor, a second non-separable transform to be used to decode the CU if the CU is a second size different from the first size.

[0180] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and a memory having instructions stored therein. The memory stores instructions that, when executed by the processor, cause the processor to analyze a bitstream. The memory stores instructions that, when executed by the processor, cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to decode a CU of the bitstream in response to enabling at least one non-separable transform for an intra-prediction method, where the at least one non-separable transform is associated with a plurality of transform matrix sets. The memory stores instructions that, when executed by the processor, cause the processor to select a transform matrix from the plurality of transform matrix sets based on a bitstream index and an intra-prediction mode. The memory stores instructions that, when executed by the processor, cause the processor to decode the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST or NSPT.

[0181] In some implementations, the memory stores instructions that, when executed by a processor, cause the processor to select one transformation matrix set from the plurality of transformation matrix sets based on the bitstream index and the intra-prediction mode, for selecting a transformation matrix from the plurality of transformation matrix sets based on the bitstream index. In some implementations, the memory stores instructions that, when executed by a processor, cause the processor to select a transformation matrix from the one transformation matrix set based on the intra-prediction mode, for selecting a transformation matrix from the plurality of transformation matrix sets based on the bitstream index and the intra-prediction mode.

[0182] In some implementations, to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode, the memory stores instructions that, when executed by the processor, cause the processor to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode derived by DIMD or TIMD.

[0183] In some implementations, if at least one non-separable transform is not enabled for an intra prediction method, the bitstream index may be omitted from the bitstream.

[0184] In some implementations, at least one non-separable transform may be enabled for a CU of any size.

[0185] In some implementations, the at least one non-separable transform includes a first non-separable transform and a second non-separable transform.

[0186] In some implementations, in response to enabling a first non-separable transform and a second non-separable transform for an intra prediction method, the memory stores instructions that, when executed by the processor, cause the processor to select a first non-separable transform to be used to decode a CU if the CU is of a first size. In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for an intra prediction method, the memory stores instructions that, when executed by the processor, cause the processor to select a second non-separable transform to be used to decode a CU if the CU is of a second size different from the first size.

[0187] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions. The instructions, when executed by a processor, can cause the processor to analyze a bitstream. When executed by the processor, the instructions can cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to decode a CU of the bitstream in response to enabling at least one non-separable transform for an intra-prediction method, where the at least one non-separable transform is associated with multiple transform matrix sets. When executed by the processor, the instructions can cause the processor to select a transform matrix from multiple transform matrix sets based on a bitstream index and an intra-prediction mode. When executed by the processor, the instructions can cause the processor to decode the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode can be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST or NSPT.

[0188] In some implementations, the instructions, when executed by a processor, for selecting a transformation matrix from a plurality of transformation matrix sets based on a bitstream index and an intra-prediction mode cause the processor to select one transformation matrix from a plurality of transformation matrix sets based on a bitstream index. In some implementations, the instructions, when executed by a processor, for selecting a transformation matrix from a plurality of transformation matrix sets based on a bitstream index and an intra-prediction mode cause the processor to select a transformation matrix from the one transformation matrix set based on the intra-prediction mode.

[0189] In some implementations, the instructions, when executed by a processor, cause the processor to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode derived by DIMD or TIMD in order to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode.

[0190] In some implementations, if at least one non-separable transform is not enabled for an intra prediction method, the bitstream index may be omitted from the bitstream.

[0191] In some implementations, at least one non-separable transform may be enabled for a CU of any size.

[0192] In some implementations, the at least one non-separable transform includes a first non-separable transform and a second non-separable transform.

[0193] In some implementations, in response to enabling a first non-separable transform and a second non-separable transform for an intra prediction method, the instructions, when executed by a processor, cause the processor to select a first non-separable transform to be used to decode a CU when the CU is a first size. In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for an intra prediction method, the instructions, when executed by a processor, cause the processor to select a second non-separable transform to be used to decode a CU when the CU is a second size different from the first size.

[0194] According to yet another aspect of the present disclosure, there is provided an encoding method performed by an encoder. The method may include, in response to enabling at least one non-separable transform for an intra-prediction method, determining, by a processor, an intra-prediction mode and the at least one non-separable transform to be used to encode a CU into a bitstream, where the at least one non-separable transform is associated with a plurality of transform matrix sets. The method may include, by the processor, selecting, based on a bitstream index and the intra-prediction mode, a transform matrix from the plurality of transform matrix sets. The method may include, by the processor, encoding the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST or NSPT.

[0195] In some implementations, selecting, by the processor, a transform matrix from a plurality of transform matrix sets based on the bitstream index and the intra-prediction mode may include selecting, by the processor, one transform matrix from the plurality of transform matrix sets based on the bitstream index. In some implementations, selecting, by the processor, a transform matrix from a plurality of transform matrix sets based on the bitstream index and the intra-prediction mode may include selecting, by the processor, a transform matrix from the one transform matrix set based on the intra-prediction mode.

[0196] In some implementations, selecting a transformation matrix from the one set of transformation matrices by the processor based on an intra-prediction mode may include selecting a transformation matrix from the one set of transformation matrices by the processor based on an intra-prediction mode derived by DIMD or TIMD.

[0197] In some implementations, if at least one non-separable transform is not enabled for an intra prediction method, the bitstream index may be omitted from the bitstream.

[0198] In some implementations, at least one non-separable transform may be enabled for a CU of any size.

[0199] In some implementations, the at least one non-separable transform includes a first non-separable transform and a second non-separable transform.

[0200] In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the method may include selecting, by the processor, a first non-separable transform to be used to encode the CU if the CU is a first size. In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the method may include selecting, by the processor, a second non-separable transform to be used to encode the CU if the CU is a second size different from the first size.

[0201] According to yet another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and a memory having instructions stored thereon. The memory stores instructions that, when executed by the processor, cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to encode a CU into a bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. The memory stores instructions that, when executed by the processor, cause the processor to select a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. The memory stores instructions that, when executed by the processor, cause the processor to encode the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST and NSPT.

[0202] In some implementations, the memory stores instructions that, when executed by a processor, cause the processor to select one transformation matrix from the plurality of transformation matrix sets based on the bitstream index and the intra-prediction mode. In some implementations, the memory stores instructions that, when executed by a processor, cause the processor to select one transformation matrix from the plurality of transformation matrix sets based on the bitstream index and the intra-prediction mode. In some implementations, the memory stores instructions that, when executed by a processor, cause the processor to select one transformation matrix from the plurality of transformation matrix sets based on the bitstream index and the intra-prediction mode.

[0203] In some implementations, to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode, the memory stores instructions that, when executed by the processor, cause the processor to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode derived by DIMD or TIMD.

[0204] In some implementations, if at least one non-separable transform is not enabled for an intra prediction method, the bitstream index may be omitted from the bitstream.

[0205] In some implementations, at least one non-separable transform may be enabled for a CU of any size.

[0206] In some implementations, the at least one non-separable transform includes a first non-separable transform and a second non-separable transform.

[0207] In some implementations, in response to enabling a first non-separable transform and a second non-separable transform for an intra prediction method, the memory stores instructions that, when executed by the processor, cause the processor to select a first non-separable transform to be used to encode a CU when the CU is a first size. In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for an intra prediction method, the memory stores instructions that, when executed by the processor, cause the processor to select a second non-separable transform to be used to encode the CU when the CU is a second size different from the first size.

[0208] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions. When executed by a processor, the instructions cause the processor to determine at least one non-separable transform and an intra-prediction mode to be used to encode a CU into a bitstream in response to enabling at least one non-separable transform for an intra-prediction method. The at least one non-separable transform may be associated with multiple transform matrix sets. When executed by a processor, the instructions cause the processor to select a transform matrix from the multiple transform matrix sets based on a bitstream index and an intra-prediction mode. When executed by a processor, the instructions cause the processor to encode the CU based on the intra-prediction method and the transform matrix. The intra-prediction mode may be determined in response to using one or more of IBC, MIP, or IntraTMP for the CU. The at least one non-separable transform may include one or more of LFNST and NSPT.

[0209] In some implementations, the instructions, when executed by a processor, for selecting a transformation matrix from a plurality of transformation matrix sets based on a bitstream index and an intra-prediction mode cause the processor to select one transformation matrix from a plurality of transformation matrix sets based on a bitstream index. In some implementations, the instructions, when executed by a processor, for selecting a transformation matrix from a plurality of transformation matrix sets based on a bitstream index and an intra-prediction mode cause the processor to select a transformation matrix from the one transformation matrix set based on the intra-prediction mode. In some implementations, the instructions, when executed by a processor, cause the processor to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode derived by DIMD or TIMD in order to select a transformation matrix from the one transformation matrix set based on an intra-prediction mode.

[0210] In some implementations, if at least one non-separable transform is not enabled for an intra prediction method, the bitstream index may be omitted from the bitstream.

[0211] In some implementations, at least one non-separable transform may be enabled for a CU of any size.

[0212] In some implementations, the at least one non-separable transform includes a first non-separable transform and a second non-separable transform.

[0213] In some implementations, in response to enabling a first non-separable transform and a second non-separable transform for an intra prediction method, the instructions, when executed by a processor, cause the processor to select a first non-separable transform to be used to encode a CU when the CU is a first size. In some implementations, in response to enabling the first non-separable transform and the second non-separable transform for an intra prediction method, the instructions, when executed by a processor, cause the processor to select a second non-separable transform to be used to encode the CU when the CU is a second size different from the first size.

[0214] The above description of each example discloses the general nature of the disclosure, thereby enabling those skilled in the art to readily modify and / or adapt such embodiments for various applications by applying knowledge within the art without undue experimentation and without departing from the general concept of the disclosure. Such adaptations and modifications are therefore intended to be within the meaning and range of equivalents of the disclosed examples, based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology herein is for the purpose of description, not limitation, and thus should be interpreted by one of ordinary skill in the art in light of the teaching and guidance.

[0215] The embodiments of the present disclosure have been described above with the help of functional components that illustrate the implementation manner of specific functions and their relationships. For the convenience of explanation, the boundaries of these functional components have been arbitrarily defined herein. Alternative boundaries can be defined as long as the specific functions and their relationships are appropriately performed.

[0216] The Description and Abstract of the Invention may describe one or more exemplary embodiments of the present disclosure but not all of the embodiments contemplated by the inventors, and therefore, the Description and Abstract of the Invention are not intended to limit the scope of the disclosure and claims in any way.

[0217] Various functional blocks, modules, and steps are disclosed above. The configurations provided are exemplary and not limiting. Thus, the functional blocks, modules, and steps may be rearranged or combined in ways different from the examples provided above. Similarly, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permissible.

[0218] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A decoding method performed by a decoder, comprising: parsing the bitstream by a processor; determining, by the processor, in response to enabling at least one non-separable transform for an intra-prediction method, the at least one non-separable transform and an intra-prediction mode to be used to decode a coding unit (CU) of the bitstream, wherein the at least one non-separable transform is associated with a plurality of transform matrix sets; selecting, by the processor, a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode; decoding, by the processor, the CU based on the intra prediction method and the transform matrix; the intra-prediction mode is determined in response to using one or more of intra-block copy (IBC), matrix-weighted intra-prediction (MIP), or intra-template matching prediction (IntraTMP) for the CU; 10. A method of decoding, wherein the at least one non-separable transform comprises one or more of a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

2. Selecting, by the processor, the transform matrix from the plurality of transform matrix sets based on the bitstream index and the intra-prediction mode, comprises: selecting, by the processor, a set of transformation matrices from the plurality of sets of transformation matrices based on the bitstream index; selecting, by the processor, the transformation matrix from the set of transformation matrices based on the intra-prediction mode. The decoding method of claim 1 .

3. Selecting, by the processor, the transform matrix from the one set of transform matrices based on the intra-prediction mode, comprises: selecting, by the processor, the transform matrix from the set of transform matrices based on an intra-prediction mode derived by decoder-side intra-mode derivation (DIMD) or template-based intra-mode derivation (TIMD). The decoding method according to claim 2 .

4. If the at least one non-separable transform is not enabled for the intra prediction method, the bitstream index is omitted from the bitstream. The decoding method of claim 1 .

5. enabling the at least one non-separable transform for coding units (CUs) of any size; The decoding method of claim 1 .

6. the at least one non-separable transform includes a first non-separable transform and a second non-separable transform; The decoding method of claim 1 .

7. In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the decoding method includes: if the CU is of a first size, selecting, by the processor, the first non-separable transform to be used to decode the CU; or and if the CU is a second size different from the first size, selecting, by the processor, the second non-separable transform to be used to decode the CU. The decoding method according to claim 6.

8. A decoder comprising: a processor; When executed by the processor, the processor Parsing the bitstream; determining an intra-prediction mode and the at least one non-separable transform to be used to decode a coding unit (CU) of the bitstream in response to enabling the at least one non-separable transform for an intra-prediction method, the at least one non-separable transform being associated with a plurality of transform matrix sets; selecting a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode; decoding the CU based on the intra prediction method and the transform matrix; and a memory having instructions stored therein; the intra-prediction mode is determined in response to using one or more of intra-block copy (IBC), matrix-weighted intra-prediction (MIP), or intra-template matching prediction (IntraTMP) for the CU; 10. The decoder, wherein the at least one non-separable transform comprises one or more of a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

9. and, when executed by the processor, to the memory, to select the transform matrix from the plurality of transform matrix sets based on the bitstream index and the intra-prediction mode. selecting a set of transformation matrices from the plurality of sets of transformation matrices based on the bitstream index; selecting the transformation matrix from the set of transformation matrices based on the intra prediction mode.

9. A decoder according to claim 8.

10. and, when executed by the processor, to the memory, to select the transform matrix from the one transform matrix set based on the intra-prediction mode. instructions to select the transform matrix from the set of transform matrices based on an intra-prediction mode derived by decoder-side intra-mode derivation (DIMD) or template-based intra-mode derivation (TIMD); 10. A decoder according to claim 9.

11. If the at least one non-separable transform is not enabled for the intra prediction method, the bitstream index is omitted from the bitstream.

9. A decoder according to claim 8.

12. enabling the at least one non-separable transform for coding units (CUs) of any size; 9. A decoder according to claim 8.

13. the at least one non-separable transform includes a first non-separable transform and a second non-separable transform; 9. A decoder according to claim 8.

14. In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the memory includes: If the CU is of a first size, selecting the first non-separable transform used to decode the CU; or and storing instructions for performing, if the CU is a second size different from the first size, selecting the second non-separable transform to be used to decode the CU.

14. A decoder according to claim 13.

15. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to: Parsing the bitstream; determining an intra-prediction mode and the at least one non-separable transform to be used to decode a coding unit (CU) of the bitstream in response to enabling the at least one non-separable transform for an intra-prediction method, the at least one non-separable transform being associated with a plurality of transform matrix sets; selecting a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode; decoding the CU based on the intra prediction method and the transform matrix; the intra-prediction mode is determined in response to using one or more of intra-block copy (IBC), matrix-weighted intra-prediction (MIP), or intra-template matching prediction (IntraTMP) for the CU; 10. The non-transitory computer-readable medium, wherein the at least one non-separable transform comprises one or more of a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

16. To select the transform matrix from the plurality of transform matrix sets based on the bitstream index and the intra-prediction mode, the instructions, when executed by the processor, cause the processor to: selecting a set of transformation matrices from the plurality of sets of transformation matrices based on the bitstream index; selecting the transformation matrix from the set of transformation matrices based on the intra-prediction mode.

16. The non-transitory computer-readable medium of claim 15.

17. To select the transformation matrix from the one transformation matrix set based on the intra prediction mode, the instructions, when executed by the processor, cause the processor to: selecting the transform matrix from the set of transform matrices based on an intra-prediction mode derived by decoder-side intra-mode derivation (DIMD) or template-based intra-mode derivation (TIMD); 17. The non-transitory computer-readable medium of claim 16.

18. If the at least one non-separable transform is not enabled for the intra prediction method, the bitstream index is omitted from the bitstream.

16. The non-transitory computer-readable medium of claim 15.

19. enabling the at least one non-separable transform for coding units (CUs) of any size; 16. The non-transitory computer-readable medium of claim 15.

20. the at least one non-separable transform includes a first non-separable transform and a second non-separable transform; 16. The non-transitory computer-readable medium of claim 15.

21. In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the instructions, when executed by the processor, cause the processor to: If the CU is of a first size, selecting the first non-separable transform used to decode the CU; or selecting the second non-separable transform to be used to decode the CU if the CU is a second size different from the first size.

21. The non-transitory computer-readable medium of claim 20.

22. 1. An encoding method performed by an encoder, comprising: determining, by a processor, in response to enabling at least one non-separable transform for an intra-prediction method, the at least one non-separable transform and an intra-prediction mode to be used to encode a coding unit (CU) into a bitstream, wherein the at least one non-separable transform is associated with a plurality of transform matrix sets; selecting, by the processor, a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode; encoding, by the processor, the CU based on the intra prediction mode and the transform matrix; the intra-prediction mode is determined in response to using one or more of intra-block copy (IBC), matrix-weighted intra-prediction (MIP), or intra-template matching prediction (IntraTMP) for the CU; 10. A method of encoding, wherein the at least one non-separable transform comprises one or more of a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

23. Selecting, by the processor, the transform matrix from the plurality of transform matrix sets based on the bitstream index and the intra-prediction mode, comprises: selecting, by the processor, a set of transformation matrices from the plurality of sets of transformation matrices based on the bitstream index; selecting, by the processor, the transformation matrix from the set of transformation matrices based on the intra-prediction mode.

23. The encoding method of claim 22.

24. Selecting, by the processor, the transform matrix from the one set of transform matrices based on the intra-prediction mode, comprises: selecting, by the processor, the transform matrix from the set of transform matrices based on an intra-prediction mode derived by decoder-side intra-mode derivation (DIMD) or template-based intra-mode derivation (TIMD).

24. The encoding method of claim 23.

25. If the at least one non-separable transform is not enabled for the intra prediction method, the bitstream index is omitted from the bitstream.

23. The encoding method of claim 22.

26. enabling the at least one non-separable transform for coding units (CUs) of any size; 23. The encoding method of claim 22.

27. the at least one non-separable transform includes a first non-separable transform and a second non-separable transform; 23. The encoding method of claim 22.

28. In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the encoding method further comprises: if the CU is of a first size, selecting, by the processor, the first non-separable transform used to encode the CU; or and if the CU is a second size different from the first size, selecting, by the processor, the second non-separable transform used to encode the CU.

28. The encoding method of claim 27.

29. 1. An encoder comprising: a processor; When executed by the processor, it causes the processor to: In response to enabling at least one non-separable transform for an intra-prediction method, determining an intra-prediction mode and the at least one non-separable transform to be used to encode a coding unit (CU) into a bitstream, wherein the at least one non-separable transform is associated with a plurality of transform matrix sets; selecting a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode; encoding the CU based on the intra prediction method and the transform matrix. the intra-prediction mode is determined in response to using one or more of intra-block copy (IBC), matrix-weighted intra-prediction (MIP), or intra-template matching prediction (IntraTMP) for the CU; The encoder, wherein the at least one non-separable transform comprises one or more of a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

30. and, when executed by the processor, to the memory, to select the transform matrix from the plurality of transform matrix sets based on the bitstream index and the intra-prediction mode. selecting a set of transformation matrices from the plurality of sets of transformation matrices based on the bitstream index; selecting the transformation matrix from the set of transformation matrices based on the intra-prediction mode.

30. The encoder of claim 29.

31. and, when executed by the processor, to the memory, to select the transform matrix from the one transform matrix set based on the intra-prediction mode. and storing instructions for selecting the transform matrix from the set of transform matrices based on an intra-prediction mode derived by decoder-side intra-mode derivation (DIMD) or template-based intra-mode derivation (TIMD).

31. The encoder of claim 30.

32. If the at least one non-separable transform is not enabled for a coding intra-prediction method, the bitstream index is omitted from the bitstream.

30. The encoder of claim 29.

33. enabling the at least one non-separable transform for coding units (CUs) of any size; 30. The encoder of claim 29.

34. the at least one non-separable transform includes a first non-separable transform and a second non-separable transform; 30. The encoder of claim 29.

35. In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the memory includes: If the CU is of a first size, selecting the first non-separable transform used to encode the CU; or and storing instructions for performing, if the CU is a second size different from the first size, selecting the second non-separable transform to be used to encode the CU.

35. The encoder of claim 34.

36. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to: In response to enabling at least one non-separable transform for an intra-prediction method, determining an intra-prediction mode and the at least one non-separable transform to be used to encode a coding unit (CU) into a bitstream, wherein the at least one non-separable transform is associated with a plurality of transform matrix sets; selecting a transform matrix from the plurality of transform matrix sets based on a bitstream index and the intra-prediction mode; encoding the CU based on the intra prediction method and the transform matrix; the intra-prediction mode is determined in response to using one or more of intra-block copy (IBC), matrix-weighted intra-prediction (MIP), or intra-template matching prediction (IntraTMP) for the CU; 10. The non-transitory computer-readable medium, wherein the at least one non-separable transform comprises one or more of a low frequency non-separable secondary transform (LFNST) or a non-separable primary transform (NSPT).

37. To select the transform matrix from the plurality of transform matrix sets based on the bitstream index and the intra-prediction mode, the instructions, when executed by the processor, cause the processor to: selecting a set of transformation matrices from the plurality of sets of transformation matrices based on the bitstream index; selecting the transformation matrix from the set of transformation matrices based on the intra-prediction mode.

37. The non-transitory computer-readable medium of claim 36.

38. To select the transformation matrix from the one transformation matrix set based on the intra prediction mode, the instructions, when executed by the processor, cause the processor to: selecting the transform matrix from the set of transform matrices based on an intra-prediction mode derived by decoder-side intra-mode derivation (DIMD) or template-based intra-mode derivation (TIMD); 38. The non-transitory computer-readable medium of claim 37.

39. If the at least one non-separable transform is not enabled for the intra prediction method, the bitstream index is omitted from the bitstream.

37. The non-transitory computer-readable medium of claim 36.

40. enabling the at least one non-separable transform for coding units (CUs) of any size; 37. The non-transitory computer-readable medium of claim 36.

41. the at least one non-separable transform includes a first non-separable transform and a second non-separable transform; 37. The non-transitory computer-readable medium of claim 36.

42. In response to enabling the first non-separable transform and the second non-separable transform for the intra prediction method, the instructions, when executed by the processor, cause the processor to: If the CU is of a first size, selecting the first non-separable transform used to encode the CU; or selecting the second non-separable transform to be used to encode the CU if the CU is a second size different from the first size.

42. The non-transitory computer-readable medium of claim 41.