Signaling transform sets

The residual blocks are encoded and decoded by an improved transform set method, which solves the video coding redundancy problem in the existing technology, improves the video coding efficiency, reduces the transmission bandwidth requirement, and optimizes the storage and streaming process.

CN120642327APending Publication Date: 2025-09-12TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093086.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2023-10-31
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing video coding technologies have difficulty in effectively reducing redundancy when compressing uncompressed video signals, resulting in high transmission bandwidth requirements, especially in storage and streaming applications.

Method used

An improved transform set method is used to encode and decode the residual block, the video block is reconstructed by inverse transform, and the encoding process is optimized by using transform coefficient sets and intra/inter prediction modes.

Benefits of technology

It improves video encoding efficiency, reduces transmission bandwidth requirements, and optimizes the amount of data during storage and streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120642327A_ABST
    Figure CN120642327A_ABST
Patent Text Reader

Abstract

The present application generally relates to video encoding / decoding, and in particular to improvements in signaling a set of transforms. A method includes: receiving an encoded video bitstream; extracting a transformation coefficient set of the current block based on the coded video code stream; determining a transform set index for the current block in the intra prediction mode based on syntax elements explicitly signaled in the encoded video bitstream, the transform set index indicating a transform set of the at least two transform sets; determining a transformation set according to the transformation set index; performing inverse transform using the transform coefficient set and the determined transform set to obtain a residual block of the current block; and the device reconstructs the current block based on the residual block.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference

[0001] This application is based upon and claims the benefit of priority of U.S. Provisional Application No. 63 / 470,763, filed on June 2, 2023, the entire contents of which are incorporated herein by reference. This application is also based upon and claims the benefit of priority of U.S. Non-Provisional Patent Application No. 18 / 497,778, filed on October 30, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application describes a set of advanced video / streaming encoding / decoding techniques. More specifically, the disclosed techniques relate to signaling a set of transforms for residual blocks. Background Art

[0003] Uncompressed digital video may include a series of images and may have specific bit rate requirements for transmission bandwidth in storage, data processing, and streaming applications. One goal of video encoding and decoding may be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention

[0004] Various embodiments of methods, apparatus, and computer-readable storage media are described herein for improving the transform set used to signal a residual block.

[0005] According to one aspect, an embodiment of the present application provides a method for decoding a current block of a current frame in an encoded video stream. The method includes a device receiving an encoded video stream. The device includes a memory storing instructions and a processor communicating with the memory. The method also includes: the device extracting a transform coefficient set of the current block based on the encoded video stream; the device determining a transform set index of the current block in an intra-frame prediction mode based on a syntax element explicitly signaled in the encoded video stream, the transform set index indicating a transform set in at least two transform sets; the device determining a transform set according to the transform set index; the device performing an inverse transform using the transform coefficient set and the determined transform set to obtain a residual block of the current block; and the device reconstructing the current block based on the residual block.

[0006] According to another aspect, embodiments of the present application provide an apparatus for processing a current block of a current frame in an encoded video stream. The apparatus includes a memory storing instructions; and a processor in communication with the memory. When the processor executes the instructions, the processor causes the apparatus to perform the above-described method for video decoding and / or encoding.

[0007] In another aspect, an embodiment of the present application provides a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above method for video decoding and / or encoding.

[0008] The above aspects and other aspects and embodiments thereof are described in more detail in the drawings, the description and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0010] Figure 1 A schematic diagram showing a simplified block diagram of a communication system (100) according to an example embodiment;

[0011] Figure 2 A schematic diagram illustrating a simplified block diagram of a communication system (200) according to an example embodiment;

[0012] Figure 3 a schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment;

[0013] Figure 4 A schematic diagram illustrating a simplified block diagram of a video encoder according to an example embodiment;

[0014] Figure 5 shows a block diagram of a video encoder according to another example embodiment;

[0015] Figure 6 shows a block diagram of a video decoder according to another example embodiment;

[0016] Figure 7 A scheme for encoding block partitioning according to an exemplary embodiment of the present application is shown;

[0017] Figure 8 Another scheme of coding block partitioning according to an exemplary embodiment of the present application is shown;

[0018] Figure 9 Another scheme of coding block partitioning according to an exemplary embodiment of the present application is shown;

[0019] Figure 10 shows an example of fine angles in directional intra prediction;

[0020] Figure 11 The nominal angles in directional intra prediction are shown;

[0021] Figure 12 A schematic diagram showing the use of a secondary transform in the encoding and decoding process;

[0022] Figure 13 shows a low frequency inseparable transformation process according to an exemplary embodiment of the present application;

[0023] Figure 14 An example of the logic flow of the method of the present application is shown; and

[0024] Figure 15 A schematic diagram of a computer system according to an example embodiment of the present application is shown. DETAILED DESCRIPTION

[0025] The present invention will now be described in detail below with reference to the accompanying drawings. The accompanying drawings are part of the present invention and illustrate specific examples of embodiments by way of illustration. However, it should be noted that the present invention may be embodied in a variety of different forms, and therefore, any embodiment described below should not be interpreted as intended to limit the subject matter described or claimed. It should also be noted that the present invention may be embodied as a method, device, component, or system. Accordingly, embodiments of the present invention may, for example, take the form of hardware, software, firmware, or any combination thereof.

[0026] Throughout the specification and claims, terms may have subtle meanings that are suggested or implied by the context beyond their explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Likewise, the phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include all or part of any combination of exemplary embodiments / embodiments.

[0027] In general, terms can be understood at least in part based on their use in the context. For example, terms such as "and", "or", or "and / or" as used herein can include multiple meanings, which can depend at least in part on the context in which such terms are used. Typically, "or", if used in an association list, such as A, B, or C, is intended to mean A, B, and C, which are used herein in an inclusive sense, and A, B, or C, which are used herein in an exclusive sense. In addition, the terms "one or more" or "at least one" as used herein, at least in part depending on the context, can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as "one", "an", or "the" can also be understood to express singular usage or plural usage, which depends at least in part on the context. In addition, the term "based on" or "determined by..." can be understood to not necessarily be intended to express an exclusive set of factors, but can allow for the presence of additional factors that are not necessarily explicitly described again, which depends at least in part on the context.

[0028] like Figure 1 As shown, the terminal device can be implemented by a server, a personal computer and a smart phone, but the application of the basic principles of the present application may not be limited thereto. Embodiments of the present application can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (150) represents any number or type of network for transmitting encoded video data between terminal devices, including, for example, a wired connection (wired) communication network and / or a wireless communication network. The communication network (150) can exchange data in a circuit switching channel, a packet switching channel and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks and / or the Internet.

[0029] Figure 2 The following illustrates the deployment of a video encoder and video decoder in a video streaming environment, illustrating an example application of the subject matter of the present application. The disclosed subject matter can be equally applied to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, and storing compressed video on digital media such as CDs, DVDs, and memory sticks.

[0030] like Figure 2As shown, a streaming system may include a video capture subsystem (213), which may include a video source (201), such as a digital camera. The video source is used to create an uncompressed video image stream (202). In the example, the video image stream (202) includes samples recorded by the digital camera. The video image stream (202) is depicted with bold lines to emphasize that the video image stream (202) has a larger data size than the encoded video data (204) (or encoded video code stream). The video image stream (202) can be processed by an electronic device (220). The electronic device (220) includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter described in detail below. The encoded video data (204) (or the encoded video code stream (204)) is depicted with thin lines to emphasize that, compared to the uncompressed video image stream (202), the encoded video data (204) (or the encoded video code stream (204)) has a smaller data size and can be stored on the streaming server (205) for future use or directly transmitted to a downstream video device (not shown). One or more streaming client subsystems, such as Figure 2 The client subsystem (206) and the client subsystem (208) in the embodiment of the present invention can access the streaming server (205) to retrieve the copy (207) and the copy (209) of the encoded video data (204). The client subsystem (206) can include, for example, a video decoder (210) in the electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and produces an output video image stream (211), which can be presented on a display (212) (e.g., a display screen) or other presentation device (not depicted).

[0031] Figure 3 A block diagram of a video decoder (310) of an electronic device (330) according to any of the following embodiments of the present application is shown. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video decoder (310) may be used to replace Figure 2 A video decoder (210) in an example of FIG.

[0032] like Figure 3As shown, a receiver (331) can receive one or more encoded video sequences from a channel (301). To combat network jitter and / or handle playback timing, a buffer memory (315) can be set between the receiver (331) and the entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) can reconstruct symbols (321) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (310) and potentially for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) can parse / entropy decode the encoded video sequence. The parser (320) can extract a subgroup parameter set for at least one pixel subgroup in the video decoder from the encoded video sequence. The subgroup can include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (320) may also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the coded video sequence. The reconstruction of the symbols (321) may involve a number of different processing or functional units. The units involved and how they are involved may be controlled by subgroup control information parsed by the parser (320) from the coded video sequence.

[0033] The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive quantized transform coefficients in the form of symbols (321) and control information from the parser (320), including information indicating which transform method to use, block size, quantization factors / parameters, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values, which may be input to the aggregator (355).

[0034] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed image, but may use predictive information from a previously reconstructed portion of the current image. Such predictive information may be provided by the intra-image prediction unit (352). In some cases, the intra-image prediction unit (352) may use reconstructed information extracted from the current image buffer (358) to generate surrounding blocks of the same size and shape as the block being reconstructed. For example, the current image buffer (358) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (355) may add the prediction information generated by the intra-frame prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0035] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to inter-frame coded and potentially motion compensated blocks. In this case, the motion compensated prediction unit (353) may access the reference picture memory (357) based on the motion vector to extract samples for prediction. After motion compensation is performed on the extracted reference samples according to the symbols (321), these samples may be added to the output of the scaler / inverse transform unit (351) by an aggregator (355) (the output of the unit (351) is called residual samples or residual signal), thereby generating output sample information.

[0036] The output samples of the aggregator (355) can be used by various loop filtering techniques in the loop filter unit (356). The loop filter unit (356) includes multiple types of loop filters. The output of the loop filter unit (356) can be a sample stream that can be output to the display device (312) and stored in the reference image memory (357) for subsequent inter-frame image prediction.

[0037] Figure 4 The block diagram of the video encoder (403) according to the embodiment of the present application is shown. The video encoder (403) can be set in the electronic device (420). The electronic device (420) can also include a transmitter (440) (e.g., a transmission circuit). The video encoder (403) can be used to replace Figure 4 A video encoder (403) in an example embodiment.

[0038] The video encoder (403) may receive video samples from a video source. According to some examples, the video encoder (403) may encode and compress images of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of the controller (450). In some embodiments, the controller (450) may be functionally coupled to and control other functional units as described below. Parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0039] In some example embodiments, the video encoder (403) can be configured to operate in an encoding loop. The encoding loop can include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols and creates sample data in a similar manner to the (remote) decoder, even though the embedded decoder 433 processes the video stream encoded by the source encoder 430 without entropy coding (because any compression between symbols and the encoded video code stream in entropy coding may be lossless in the video compression techniques considered in the disclosed subject matter). It can be observed at this point that any decoder technology other than parsing / entropy decoding that may only exist in the decoder may also necessarily need to exist in a corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on the decoder operation, which is related to the decoding portion of the encoder. Therefore, the description of the encoder technology may be abbreviated because it is the inverse process of the fully described decoder technology. A more detailed description of the encoder is provided below only in certain areas or aspects.

[0040] During operation, in some example embodiments, the source encoder (430) may perform motion-compensated predictive coding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures."

[0041] A local video decoder (433) can decode encoded video data of a picture that can be designated as a reference picture. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference picture and can cause the reconstructed reference picture to be stored in a reference picture cache (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the far-end (remote) video decoder.

[0042] The predictor (435) may perform prediction search for the encoding engine (432). That is, for a new image to be encoded, the predictor (435) may search the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image.

[0043] The controller (450) may manage encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.

[0044] The outputs of all of the above functional units may be entropy encoded in an entropy encoder (445). A transmitter (440) may buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission over a communication channel (460), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) may combine the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0045] The controller (450) can manage the operation of the video encoder (403). During encoding, the controller (450) can assign a coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types: intra picture (I picture), predicted picture (P picture), bidirectionally predicted picture (B picture), multi-predicted picture. The source picture can generally be spatially subdivided into a plurality of sample coding blocks, as described in further detail below.

[0046] Figure 5 is a diagram of a video encoder (503) according to another embodiment disclosed herein. The video encoder (503) is configured to receive a processed block (e.g., a prediction block) of sample values ​​within a current video image in a video image sequence and to encode the processed block into an encoded image that is part of an encoded video sequence. The example video encoder (503) may be used in place of Figure 4 The video encoder (403) in the example.

[0047] For example, the video encoder (503) receives a matrix of sample values ​​for a processing block. The video encoder (503) uses, for example, rate-distortion optimization (RDO) to determine whether to use intra mode, inter mode, or bi-prediction mode to encode the processing block.

[0048] exist Figure 5 In an embodiment of the present invention, the video encoder (503) includes Figure 5 Shown are an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), ​​a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together.

[0049] The inter-frame encoder (530) is used to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in a previous image and a subsequent image in display order), generate inter-frame prediction information (e.g., redundant information description, motion vectors, merge mode information according to an inter-frame coding technique), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique.

[0050] The intra encoder (522) is configured to receive samples of a current block (e.g., a processing block), compare the block to previously encoded blocks in the same image, generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques).

[0051] The general controller (521) may be configured to determine general control data and control other components of the video encoder (503) based on the general control data, thereby, for example, determining a prediction mode for a block and providing a control signal to the switch (526) based on the prediction mode.

[0052] The residual calculator (523) may be configured to calculate the difference (residual data) between the received block and the prediction result of the block selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) may be configured to encode the residual data to generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be configured to format the code stream to include the encoded blocks and perform entropy encoding.

[0053] Figure 6 FIG is a diagram of an example video decoder (610) according to another embodiment disclosed herein. The video decoder (610) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In an embodiment, the video decoder (610) may be used instead of Figure 4 A video decoder (410) is shown in the example.

[0054] exist Figure 6 In one embodiment, the video decoder (610) includes Figure 6 The deployment shown in FIG 7 shows an entropy decoder ( 671 ), an inter-frame decoder ( 680 ), a residual decoder ( 673 ), a reconstruction module ( 674 ) and an intra-frame decoder ( 672 ) coupled together.

[0055] The entropy decoder (671) is configured to reconstruct certain symbols from an encoded image, the symbols representing syntax elements constituting the encoded image. The inter-frame decoder (680) is configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) is configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) is configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) is configured to combine the residual output by the residual decoder (673) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which may constitute a part of a reconstructed image, which is a part of a reconstructed video.

[0056] It should be noted that the video encoder (203), video encoder (403), and video encoder (503), as well as the video decoder (210), video decoder (310), and video decoder (610) may be implemented using any suitable technology. In some example embodiments, the video encoder (203), video encoder (403), and video encoder (503), as well as the video decoder (210), video decoder (310), and video decoder (610) may be implemented using one or more integrated circuits. In another embodiment, the video encoder (203), video encoder (403), and video encoder (503), as well as the video decoder (210), video decoder (310), and video decoder (410) may be implemented using one or more processors executing software instructions.

[0057] Turning now to block partitioning for encoding and decoding, partitioning can generally start from a basic block and can follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. Partitioning can be hierarchical and recursive. After dividing or partitioning the basic block following any of the example partitioning procedures or other procedures described below, or a combination thereof, a final set of partitions or coding blocks can be obtained. Each of these partitions can be at one of the partition levels in the partition hierarchy and can have various shapes. Each of these partitions can be referred to as a coding block (CB). For the various example partitioning implementations further described below, each CB generated can have any allowed size and partition level. Such partitions are called coding blocks because they can form units on which some basic encoding / decoding decisions can be made, and encoding / decoding parameters can be optimized, determined, and signaled in the encoded video stream. The highest or deepest level in the final partition represents the depth of the coding block partition structure of the tree. The coding block can be a luminance coding block or a chrominance coding block. The CB tree structure for each color can be referred to as a coding block tree (CBT). The coding blocks of all color channels can be collectively referred to as coding units (CUs). The hierarchical structure of all color channels may be collectively referred to as a coding tree unit (CTU). The partitioning modes or structures of various color channels in a CTU may be the same or different.

[0058] In some embodiments, the partition tree schemes or structures used for the luma channel and the chroma channels may not have to be the same. In other words, the luma channel and the chroma channels may have independent coding tree structures or modes. In addition, whether the luma channel and the chroma channels use the same or different coding partition tree structures and the actual coding partition tree structure to be used may depend on whether the slice being coded is a P slice, a B slice or an I slice. For example, for an I slice, the chroma channel and the luma channel may have independent coding partition tree structures or coding partition tree structure modes, while for a P slice or a B slice, the luma channel and the chroma channels may share the same coding partition tree scheme. When independent coding partition tree structures or modes are applied, the luma channel may be partitioned into CBs by one coding partition tree structure, while the chroma channels may be partitioned into chroma CBs by another coding partition tree structure.

[0059] Figure 7 An example of a predefined 10-way partitioning structure / pattern is shown that allows recursive partitioning to form a partition tree. A root block may start from a predefined level (eg, from a basic block at the 128x128 or 64x64 level). Figure 7 Example partition structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some example embodiments, Figure 7The coding tree depth can be further defined to indicate the partition depth starting from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and Figure 7 After one further split of the root block, the coding tree depth increases by 1. In some embodiments, it may be possible to only allow all square partitions in 710 to be split according to Figure 7 The pattern is recursively partitioned to the next level of the partition tree.

[0060] In some other example implementations of coding block partitioning, a quadtree structure may be used. This quadtree partitioning can be applied hierarchically and recursively to any square-shaped partition. Whether to perform further quadtree partitioning on a basic block or intermediate block or partition can be adaptively determined based on various local characteristics of the basic block or intermediate block / partition.

[0061] In yet other examples, a ternary tree partitioning scheme may be used to partition basic blocks or any intermediate blocks, such as Figure 8 The ternary tree pattern can be performed in a vertical direction, as shown in 802, or in a horizontal direction, as shown in 804. Figure 8 In the example of , the split ratio is shown as 1:2:1, but other ratios can be predefined. In some embodiments, two or more different ratios can be predefined. In some embodiments, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid extra transformations.

[0062] The above partitioning schemes can be combined in any manner at different partitioning levels. As an example, the above quadtree and binary tree partitioning schemes can be combined to partition a basic block into a quadtree-binary tree (QTBT) structure. In this scheme, a basic block or intermediate block / partition can be partitioned by a quadtree or a binary tree (if a set of predefined conditions are specified) sequentially following the set of predefined conditions. Figure 9, where a basic block is first partitioned into four partitions by a quadtree, as shown in 902, 904, 906, and 908. Thereafter, each resulting partition is partitioned into four further partitions by a quadtree at the next level (as shown in 908), or is partitioned into two further partitions by a binary tree (e.g., horizontally or vertically, as shown in 902 or 906, both of which are symmetrical), or is not partitioned (such as 904). For square-shaped partitions, recursive binary or quadtree partitioning can be allowed, as shown in the overall example of the partitioning pattern in 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree partitioning and dashed lines represent binary tree partitioning. A flag can be used to indicate whether the binary tree partitioning of each binary tree partition node (non-leaf binary tree partition) is horizontal or vertical. For example, as shown in 920, consistent with the partitioning structure of 910, a flag "0" can represent a horizontal binary tree partitioning, while a flag "1" can represent a vertical binary tree partitioning. For quadtree partitions, there is no need to indicate the partition type, because quadtree partitioning always splits the block or partition both horizontally and vertically to produce four sub-blocks / partitions of equal size. In some embodiments, the flag "1" can indicate horizontal binary tree partitioning, while the flag "0" can indicate vertical binary tree partitioning.

[0063] In some example implementations of QTBT, the quadtree and binary tree partitioning rule sets may be represented by the following predefined parameters and their associated corresponding functions: –CTU size: the root node size of the quadtree (the size of the basic block) –MinQTSize: Minimum allowed quadtree leaf node size –MaxBTSize: Maximum allowed binary tree root node size –MaxBTDepth: Maximum allowed binary tree depth –MinBTSize: Minimum allowed binary tree leaf node size

[0064] In some example implementations of the QTBT partition structure, the CTU size can be set to 128×128 luma samples with two corresponding blocks of 64×64 chroma samples (when an example chroma subsampling is considered and used), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf node can be within the minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the node is 128×128, it will not be binary tree partitioned first because the size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes that do not exceed MaxBTSize can be binary tree partitioned. Figure 9 In the example of , the basic block is 128×128. According to a predefined set of rules, the basic block can only be split by a quadtree. The partition depth of the basic block is 0. Each of the four partitions obtained is 64×64, not exceeding MaxBTSize, and can be further quadtree split or binary tree split at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be ignored. When the width of the binary tree node is equal to MinBTSize (i.e., 4), further horizontal splitting can be ignored. Similarly, when the height of the binary tree node is equal to MinBTSize, further vertical splitting is not considered.

[0065] In some example embodiments, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure or independent QTBT structures for luma and chroma. For example, for P slices and B slices, the luma and chroma CTBs in one CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CBs according to one QTBT structure, and the chroma CTB can be partitioned into chroma CBs according to another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice, for example, an I slice can consist of coding blocks of one luma component or coding blocks of two chroma components, while a CU in a P slice or B slice can consist of all coding blocks of the three color components.

[0066] The above various CB partitioning schemes and further partitioning of CB to PB can be combined in any manner. The following specific embodiments are provided as non-limiting examples.

[0067] In some embodiments, a set of intra-frame prediction modes (interchangeably referred to as "intra-frame modes") may include a predefined number of directional intra-frame prediction modes. These intra-frame prediction modes may correspond to a predefined number of directions along which out-of-block samples are selected as prediction samples for the currently predicted sample in a particular block. In another specific embodiment example, eight (8) main directional modes corresponding to multiple angles between 45 degrees and 207 degrees from the horizontal axis may be predefined and supported. In some other embodiments of intra-frame prediction, in order to further exploit a wider variety of spatial redundancies in directional textures, the directional intra-frame modes may be further extended to a set of angles with a finer granularity. For example, the above 8-angle embodiment may be configured to provide eight nominal angles, referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as Figure 11 As shown, for each nominal angle, a predefined number (e.g., 7) of finer angles can be added. With this extension, corresponding to the same number of predefined directional intra modes, more (e.g., 56 in this example) directional angles can be used for intra prediction. The predicted angle can be represented by the nominal intra angle plus the angle increment. For the specific example above, where each nominal angle has 7 finer angle directions, the angle increment can be a step size of -3 to 3 times 3 degrees. Some angle schemes can be used, such as Figure 10 As shown, there are 65 different prediction angles. In some embodiments, 8 nominal modes and 5 non-angle smoothing modes are first signaled. If the current mode is an angle mode, an index is further signaled to indicate the angle increment to the corresponding nominal angle. In some embodiments, in order to implement directional prediction modes in a common way, all 56 directional intra-frame prediction modes can be implemented with a unified directional predictor that projects each pixel to a reference sub-pixel position and interpolates the reference pixel through a 2-tap bilinear filter.

[0068] The residual transform of the intra-frame prediction block or inter-frame prediction block can then be performed, followed by quantization of the transform coefficients. For the purpose of performing the transform, both intra-frame and inter-frame coding blocks can be further partitioned into multiple transform blocks (sometimes used interchangeably as "transform units", even though the term "unit" is generally used to refer to a set of three color channels, for example, a "coding unit" includes a luma coding block and a chroma coding block) before the transform. In some embodiments, a maximum partition depth of the coded block (or prediction block) can be specified (the term "coded block" can be used interchangeably with "coding block"). For example, such partitioning may not exceed 2 levels. The partitioning of prediction blocks into transform blocks can be handled differently for intra-frame prediction blocks and inter-frame prediction blocks. However, in some embodiments, such partitioning can be similar for intra-frame prediction blocks and inter-frame prediction blocks.

[0069] In some example embodiments, and for inter-frame coded blocks, transform unit partitioning can be done in a recursive manner, where the maximum value of the partition depth can be a predefined number of levels (e.g., 2 levels). Partitioning can stop or continue recursively at any sub-partition and at any level. For example, a block is partitioned into four quadtree sub-blocks, and one of the sub-blocks is further partitioned into four second-level transform blocks, while the partitioning of the other sub-blocks stops after the first level, resulting in a total of 7 transform blocks of two different sizes. In some embodiments, transform partitioning can support transform block shapes of 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1 and transform block sizes ranging from 4×4 to 64×64. In some example embodiments, if the coding block is less than or equal to 64×64, then the transform block partitioning can be applied only to the luma component (in other words, the chroma transform blocks will be the same as the coding block under this condition). Otherwise, if the coding block width or height is greater than 64, the luma and chroma coding blocks can be implicitly split into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.

[0070] Each of the above transform blocks can then be subjected to a primary transform. The primary transform essentially moves the residuals in the transform block from the spatial domain to the frequency domain. In some implementations of the actual primary transform, to support the above extended coding block partitioning example, multiple transform sizes (ranging from 4 to 64 points for each of the two dimensions) and multiple transform shapes (square; rectangular with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) can be allowed.

[0071] Now, introducing the actual main transform, in some example embodiments, the 2-D transform process may involve the use of a hybrid transform kernel (e.g., which may be composed of a different 1-D transform for each dimension of the encoded residual transform block). Examples of 1-D transform kernels may include, but are not limited to: a) 4-point, 8-point, 16-point, 32-point, and 64-point DCT-2; b) 4-point, 8-point, and 16-point asymmetric DST (DST-4, DST-7) and their flipped versions; and c) 4-point, 8-point, 16-point, and 32-point identity transforms. The choice of transform kernel for each dimension may be based on a rate-distortion (RD) criterion.

[0072] In some example embodiments, the availability of a hybrid transform kernel for a particular primary transform implementation may be based on the transform block size and prediction mode. For chroma components, transform type selection may be performed implicitly. For example, for intra-prediction residuals, the transform type may be selected based on the intra-prediction mode. For inter-prediction residuals, the transform type for the chroma block may be selected based on the transform type selection for the co-located luma block. Therefore, for chroma components, transform type signaling is not present in the bitstream.

[0073] In some embodiments, for the intra prediction residual block of the luma color component, a secondary transform method, i.e., intra secondary transform (IST), may be applied to the primary transform coefficient block before quantization is performed at the encoder. Accordingly, a secondary inverse transform may be applied to the dequantized transform coefficient block before the inverse primary transform is performed at the decoder. IST is not applied to the chroma color components. Figure 12 The use of IST in the encoding and decoding process is shown in FIG.

[0074] In some embodiments using IST, a non-separable transform process is applied. To apply a forward non-separable transform to a specific region of an input transform coefficient block consisting of N samples, the N samples are first arranged into an N×1 vector in coefficient scan order according to the relative coordinates of each sample in the input. An N×N transform kernel (K) is then selected and the non-separable transform is performed using the following arithmetic operation: in is the output Nx1 vector that replaces a specific region of a transform coefficient block using the coefficient scan order.

[0075] In some embodiments of performing an inverse non-separable transform, a block of dequantized transform coefficients is first given as input, and then, based on the size of the transform block, a specific region of the dequantized transform coefficient block is identified. The input vector is formed by rearranging the transform coefficients using a coefficient scanning order. Given a selected N×N transform kernel and an input vector Perform an inverse nonseparable transformation using the following arithmetic operation: in(·)T Refers to the matrix transpose operation. Output is an N×1 vector that replaces a specific region of the input block in a specific coefficient scanning order.

[0076] In some embodiments, the input to the forward IST is a coefficient vector consisting of low-frequency main transform coefficients in a zigzag scan. Depending on the block size, a 16-point or 64-point non-separable secondary transform can be selected. When the minimum value of the main transform width and the main transform height is less than 8, a 16-point IST is used, and the low-frequency main transform coefficients refer to the first 16 main transform coefficients in the zigzag scan order. When the main transform width and height are both greater than or equal to 8, a 64-point IST is applied, and the low-frequency main transform coefficients refer to the first 64 main transform coefficients in the zigzag scan order. The 16-point non-separable transform uses an 8×16 transform kernel, and the 64-point non-separable transform uses a 32×64 transform kernel. In addition, when IST is applied, the high-frequency transform coefficients that have not been processed by the secondary transform are cleared to zero.

[0077] In some embodiments, a total of 12 secondary transform sets (or IST sets) can be defined, where each set contains 3 secondary transform kernels. For each intra-coded transform block, the nominal intra prediction mode and primary transform type are first identified, and then the IST set is selected based on Table 1. In some embodiments, for Paeth prediction mode and recursive intra prediction mode, the IST is neither applied nor signaled. Table 1: Mapping from intra nominal mode and main transform type to IST set

[0078] In some embodiments, given an IST set, four encoder options may exist: 1) not performing a secondary transform, 2) performing a secondary transform using the first transform kernel in the given IST set, 3) performing a secondary transform kernel using the second transform kernel in the given IST set, and / or 4) performing a secondary transform kernel using the third transform kernel in the given IST set. In some embodiments, the encoder uses the syntax element ist_idx to signal this selection. At the decoder, the value of the syntax element ist_idx is first parsed. Then, given the IST set and the value associated with ist_idx, the secondary transform kernel is determined. After signaling the primary transform type, the syntax element ist_idx is signaled for each luma transform block. Signaling ist_idx is performed when all of the following conditions are true: the current block is an intra-coded luma transform block; the primary transform type is either DCT in both dimensions or ADST in both dimensions; the intra prediction mode is neither Pareto prediction mode nor recursive intra prediction mode; the transform partition depth is 0; and the end-of-block (EOB) position falls within the low-frequency transform coefficient region to which the secondary transform is applicable. In some embodiments, the entropy coding context of ist_idx is derived based on the transform block size.

[0079] In some embodiments, a secondary transform may be performed on the primary transform coefficients. For example, a low-frequency non-separable transform (LFNST), referred to as a reduced secondary transform, may be performed between the forward primary transform and quantization (at the encoder) and between dequantization and the inverse primary transform (at the decoder side), as shown in FIG. Figure 13 , to further remove the correlation of the main transform coefficients. Therefore, LFNST can take a portion of the main transform coefficients, such as the low-frequency portion (thereby "reducing" from the complete set of main transform coefficients of the transform block) for secondary transform. In the LFNST example, a 4×4 non-separable transform or an 8×8 non-separable transform can be performed depending on the transform block size. For example, 4×4 LFNST can be applied to small transform blocks (e.g., min(width, height)<8), while 8×8 LFNST can be applied to larger transform blocks (e.g., min(width, height)>8). For example, if an 8×8 transform block undergoes 4×4 LFNST, only the low-frequency 4×4 portion of the 8×8 main transform coefficients is further subjected to secondary transform.

[0080] like Figure 13As specifically shown in FIG, the transform block can be 8×8 (or 16×16). Therefore, the forward main transform 1305 of the transform block produces an 8×8 (or 16×16) main transform coefficient matrix 1304, where each square cell represents a 2×2 (or 4×4) portion. For example, the input to the forward LFNST may not be the entire 8×8 (or 16×16) main transform coefficients. For example, a 4×4 (or 8×8) LFNST may be used for the secondary transform. Therefore, only the 4×4 (or 8×8) low-frequency main transform coefficients of the main transform coefficient matrix 1304 (as indicated by the shaded portion 1306 (upper left)) may be used as input to the LFNST. The remaining portion of the main transform coefficient matrix may not undergo the secondary transform. Thus, after the secondary transform, the portion of the main transform coefficients that has undergone LFNST processing becomes the secondary transform coefficients, while the remaining portion that has not undergone LFNST processing (e.g., the unshaded portion of matrix 1304) remains the corresponding main transform coefficients. In some example embodiments, the remaining portion that has not undergone secondary transformation may all be set to zero coefficients.

[0081] The following describes an application example of the non-separable transform used in LFNST. To apply the example of 4×4 LFNST, a 4×4 input block X (representing, for example, the 4×4 low-frequency portion of the main transform coefficient block, such as Figure 13 The shaded portion 1306 of the main transformation matrix 1304 can be expressed as:

[0082] This 2-D input matrix can first be linearized or scanned in the order of examples to become the vector Then, the non-separable transform for 4×4 LFNST can be calculated as in Indicates the output transform coefficient vector, and T is the 16×16 transform matrix. Then, the resulting 16×1 coefficient vector The block is then scanned in reverse order (e.g., horizontally, vertically, or diagonally) into 4×4 blocks. Coefficients with smaller indices can be placed together with smaller scan indices into the 4×4 coefficient block. In this way, redundancy in the primary transform coefficients X can be further reduced by the second transform T, thereby further improving compression.

[0083] The above LFNST example performs the non-separable transform based on the direct matrix multiplication method, enabling it to be implemented once without multiple iterations. In some other example embodiments, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST example can be further reduced to minimize the computational complexity and the requirement for the memory space to store the transform coefficients. Such embodiments can be referred to as reduced non-separable transform (RST). More specifically, the main idea of RST is to map an N-dimensional vector (where N is 4×4 = 16 in the above example, but for an 8×8 block, it can be equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) represents the dimensionality reduction factor. Therefore, the RST matrix is no longer an N×N transform matrix but becomes an R×N matrix, as follows:

[0084] In the above matrix, the R rows of the transform matrix are the R bases after the reduction of the N-dimensional space. Therefore, the transform converts the N-dimensional input vector into a reduced R-dimensional output vector. Thus, as shown in FIG. 23, the secondary transform coefficients (square shaded area 1308) obtained from the primary coefficient 1306 are reduced in dimension to N / R times. The three squares (1309, 1310, 1311) around 1308 can be filled with zeros.

[0085] The inverse transform matrix of RST can be the transpose of its forward transform. For the 8×8 LFNST example (for more diverse descriptions in this article compared to the above 4×4 LFNST), the example reduction factor 4 can be applied, and thus the 64×64 direct non-separable transform matrix is correspondingly reduced to a 16×64 direct matrix. Additionally, in some embodiments, only a part, rather than all, of the input primary coefficients can be linearized into the input vector of the LFNST. For example, only a part of the example 8×8 input primary transform coefficients can be linearized into the above X vector. For a specific example, in the four 4×4 quadrants of the 8×8 primary transform coefficient matrix, the lower right (high-frequency coefficients) can be omitted, and only the other three quadrants are linearized into a 48×1 vector instead of a 64×1 vector using a predefined scanning order. In such embodiments, the non-separable transform matrix can be further reduced from 16×64 to 16×48.

[0086] Therefore, the example reduced 48×16 inverse RST matrix can be used on the decoder side to generate the upper left, upper right and lower left 4×4 quadrants of the 8×8 core (main) transform coefficients. Specifically, when a further reduced 16×48 RST matrix is ​​used instead of a 16×64 RST with the same transform set configuration, the inseparable secondary transform takes as input the vectorized 48 matrix elements of the three 4×4 quadrant blocks (excluding the lower right 4×4 block) from the 8×8 main coefficient block. In such an embodiment, the omitted lower right 4×4 main transform coefficients will be ignored in the secondary transform. This further reduced transform converts the 48×1 vector into a 16×1 output vector, which is scanned back into a 4×4 matrix to fill Figure 13 1308. The three squares (1309, 1310, 1311) surrounding the secondary transform coefficients of 1308 may be filled with zeros.

[0087] With this dimensionality reduction in RST, the memory usage for storing all LFNST matrices is reduced. In the above example, for example, the memory usage can be reduced from 10KB to 8KB with a fairly small performance degradation compared to an implementation without dimensionality reduction.

[0088] In some embodiments, to reduce complexity, LFNST can be further restricted to only be performed outside the portion of the main transform coefficients that are to undergo LFNST (e.g., Figure 13This applies when all coefficients (except for the portion 1306 of 1304 in the matrix 1306) are non-significant. Therefore, when LFNST is applied, all coefficients that are only subjected to the main transform (e.g., the unshaded portion of the main coefficient matrix 1304) can be close to zero. This restriction allows the LFNST index signaling at the last significant position to be adjusted and thus avoids some additional coefficient scanning. When this restriction is not applied, additional coefficient scanning may be required to check for significant coefficients at specific positions. In some embodiments, the worst-case treatment of LFNST (in terms of the number of multiplications per pixel) can limit the non-separable transforms for 4×4 and 8×8 blocks to 8×16 and 8×48 transforms, respectively. In these cases, when LFNST is applied, the last significant scan position must be less than 8 for other sizes less than 16. For blocks of shape 4×N and N×4 with N>8, the above restriction means that LFNST is now only performed once on the top left 4×4 region. Since all main transform-only coefficients are zero when LFNST is applied, the number of operations required for the main transform is reduced in this case. From the encoder's perspective, the quantization of the coefficients can be simplified when testing the LFNST transform. For the first 16 coefficients (in scan order), rate-distortion optimized quantization (RDO) must be performed to the greatest extent possible; the remaining coefficients can be forced to zero.

[0089] In some example embodiments, the available RST kernels can be specified as multiple transform sets, each of which includes at least two inseparable transform matrices. For example, there can be a total of 4 transform sets, each with 2 inseparable transform matrices (kernels) for LFNST. These kernels can be pre-trained offline, and therefore they are data-driven. The offline trained transform kernels can be stored in a memory or hard-coded in an encoding device or decoding device for use during the encoding / decoding process. The selection of a transform set during the encoding or decoding process can be determined by the intra-frame prediction mode. The mapping from the intra-frame prediction mode to the transform set can be predefined. Table 2 shows an example of such a predefined mapping. For example, when one of the three cross-component linear model (CCLM) modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (i.e., 81<=predModeIntra<=83), transform set 0 can be selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate can be further specified by an explicitly signaled LFNST index. For example, one such index can be signaled for each intra CU in the codestream after the transform coefficients. Table 2: Transformation selection table IntraPredMode Transformation Collection Index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0

[0090] Because LFNST is limited to being applicable only when all coefficients outside the first coefficient subgroup or portion in the above example embodiments are insignificant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index can be context-coded, but does not depend on the intra prediction mode, and only the first bit can be context-coded. In addition, LFNST can be applied to intra CUs in intra and inter slices, and for luma and chroma. If dual-tree is enabled, the LFNST indices for luma and chroma can be signaled separately. For inter slices (dual-tree disabled), a single LFNST index can be signaled and used for luma and chroma.

[0091] In some embodiments, there are two approaches for block-based 2-D data transformations: i) separable transforms, and ii) non-separable transforms. In a separable transform, each column and row of the block is considered to be a 1-D signal, and a 1-D transform is used to map the data block to a set of coefficients. The 1-D transform used in each direction can be the same, but can also be different. For a non-separable transform, the blocks are typically sorted into 1-D vectors by ordering the columns or rows of the blocks in a lexicographical manner. The disadvantage of this is that non-separable transforms may require more memory to store the entries of the transformation matrix, and large matrix multiplications are generally too complex for hardware implementation. Therefore, in some embodiments, separable transforms are attractive. However, separable transforms have a cost because they only exploit correlations with columns or rows; therefore, separable transforms have lower compression performance around directional edges compared to non-separable transforms.

[0092] In some example embodiments, when Intra Sub-Partitioning (ISP) mode is selected, LFNST may be disabled and the RST index may not be signaled, as even if RST is applied to every feasible partition block, the performance improvement may be minimal. Furthermore, disabling RST for the ISP prediction residual may reduce coding complexity. In some other embodiments, when Multiple Linear Regression Intra Sub-Partitioning (MIP) mode is selected, LFNST may also be disabled and the RST index may not be signaled.

[0093] Considering that large CUs larger than 64×64 (or any other predefined size representing the maximum transform block size) are implicitly partitioned (e.g., TU tiling) due to existing maximum transform size limits (e.g., 64×64), LFNST index searches can quadruple the data buffering for a certain number of decoding pipeline stages. Therefore, in some embodiments, the maximum size allowed for LFNST can be limited to, for example, 64×64. In some embodiments, LFNST can be enabled only when DCT2 is used as the primary transform.

[0094] In some other embodiments, an intra secondary transform (IST) is provided for the luma component by defining, for example, 12 secondary transform sets, each with, for example, 3 kernels. An intra mode-dependent index can be used for transform set selection. Kernel selection within a set can be based on signaled syntax elements. The IST can be enabled when DCT2 or ADST is used as the horizontal and vertical primary transforms.

[0095] In some embodiments, depending on the block size, a 4×4 non-separable transform or an 8×8 non-separable transform can be selected. If min(tx_width, tx_height) < 8, a 4×4 IST can be selected. For larger blocks, an 8×8 IST can be used. Here, tx_width and tx_height correspond to the transform block width and height, respectively. The input to the IST can be the low-frequency main transform coefficients in a zigzag scan order.

[0096] In some other example embodiments, without restriction, the number of core sets may not be 12, for example, there may be 14 secondary transform core sets. The mapping between the secondary core sets and the various intra prediction modes may be predetermined or preconfigured. For example, a particular intra prediction mode may be mapped to one of the 12 or 14 sets. Each core set may contain at least two (or a certain amount of) secondary cores instead of 3. For example, there may be 6 secondary cores in each secondary transform core set. The number of cores in each set may be determined based on a trade-off between the overhead associated with increasing the number of secondary cores and the potential additional coding gain. In some embodiments, the number of cores in each core set may be 3 to 6. From a practical and statistical perspective, after there are 6 optional secondary transform cores in each core set, the coding gain may become sufficiently small compared to the increase in coding overhead.

[0097] In some embodiments, there are some issues or problems associated with signaling transform sets that can be addressed or improved to increase encoding / decoding efficiency. For a non-limiting example, the choice of transform set for the secondary transform depends on the intra prediction mode, which limits the flexibility of the encoder in selecting the best transform set for the residual block. This application describes various embodiments for improving the signaling of transform sets, wherein for any given intra mode, at least one or more transform sets are selectable, thereby addressing at least one of the issues or problems discussed above, improving encoding / decoding efficiency, and improving video codec technology.

[0098] The various embodiments and / or implementations described in this application may be performed individually or in combination in any order and may be applicable to decoding, encoding, or streaming. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). At least one processor executes a program stored in a non-volatile computer-readable medium. In this application, the term block may be interpreted as a prediction block, a coding block, or a coding unit (CU).

[0099] Figure 14A flowchart 1400 is shown for an exemplary method for improving signaling of transform sets, which follows the basic principles of the above embodiments. The exemplary decoding method begins at 1401 and may include some or all of the following steps: S1410, receiving an encoded video stream; S1420, extracting a transform coefficient set for a current block based on the encoded video stream; S1430, determining a transform set index for the current block in intra-prediction mode based on syntax elements explicitly signaled in the encoded video stream, the transform set index indicating one of a plurality of transform sets; S1440, determining a transform set based on the transform set index; S1450, performing an inverse transform using the transform coefficient set and the determined transform set to obtain a residual block for the current block; and / or S1460, reconstructing the current block based on the residual block by a device. The exemplary method stops at S1499.

[0100] In the present application, a transform set index may be an integer used to index a transform set in a list / group of transform sets. For example, when there are N transform sets in the transform set list, the transform set index may be 0, indicating the first transform set in the transform set list, 1, indicating the second transform set in the transform set list, ..., or N-1, indicating the Nth transform set in the transform set list. In some embodiments, a transform set may include one or more transform kernels, and when there are more than one transform kernel within a transform set, the selection of the transform kernel within the transform set may be signaled or implicitly derived (e.g., based on encoded information).

[0101] Another embodiment may include another exemplary method for improving signaling of a transform set. The method may include some or all of the following steps: receiving a coded video stream; extracting a transform coefficient set of a current block based on the coded video stream; extracting a transform set index for the current block in any given intra prediction mode based on the coded video stream, the transform set index indicating a transform set from at least two transform sets; determining a transform set based on the transform set index; determining a transform kernel from the determined transform set based on the coded video stream; performing an inverse transform on the transform coefficient set based on the determined transform kernel to obtain a residual block of the current block; and / or reconstructing the current block based on the residual block.

[0102] The present application describes another method for improving the signaling of transform sets. The method may include some or all of the following steps: transforming a residual block of an intra-predicted video block (or a current block) using at least one primary transform kernel to generate a primary transform coefficient set; for any given intra-prediction mode, selecting a secondary transform set from a set of secondary transform sets; selecting a transform kernel from the secondary transform sets; performing a transform based on the primary transform coefficient set to generate a secondary transform coefficient set; and / or encoding the secondary transform coefficient set and a transform set index corresponding to the secondary transform set into a coded video stream associated with the current block.

[0103] In any part or combination of the above embodiments, the determined transform set includes a secondary transform set; and / or performing an inverse transform of the transform coefficient set based on the determined transform kernel to obtain a residual block of the current block includes: performing an inverse secondary transform of the transform coefficient set based on the transform kernel determined in the secondary transform set to generate a main transform coefficient set of the current block, and / or performing an inverse main transform of the main transform coefficient set based on the main transform kernel to generate a residual block of the current block.

[0104] In any part or combination of the above embodiments, extracting the transform set index of the current block includes: extracting transform set indexes corresponding to multiple transform sets in response to the current block being in an intra prediction mode.

[0105] In any part or combination of the above embodiments, extracting the transform set index of the current block based on the encoded video stream includes: performing entropy decoding on the encoded video stream based on the context of the encoded information to obtain the transform set index.

[0106] In any portion or combination of the above embodiments, the encoded information includes at least one of: intra-frame prediction mode, inter-frame prediction mode, transform partition mode or depth, transform coefficient value, last non-zero position or end of block (EOB), number of non-zero transform coefficients or quantization parameter.

[0107] In any part or combination of the above embodiments, extracting the transform set index of the current block based on the encoded video stream includes: in response to a condition being met, entropy decoding the encoded video stream to obtain the transform set index, wherein the condition depends on the encoded information.

[0108] In any portion or combination of the above embodiments, the encoded information includes at least one of: intra-frame prediction mode, inter-frame prediction mode, transform partition mode or depth, transform coefficient value, last non-zero position or end of block (EOB), number of non-zero transform coefficients or quantization parameter.

[0109] In any portion or combination of the above embodiments, a first plurality of transform sets for a first block overlaps with a second plurality of transform sets for a second block, wherein the first block and the second block have at least one of: different intra prediction modes, different transform partitioning modes or depths, different transform coefficient values, different last non-zero positions or EOBs, or different numbers of non-zero transform coefficients.

[0110] In any part or combination of the above embodiments, the plurality of transform sets includes at least one of: all separable transform sets, all inseparable transform sets, or a list of separable transform sets and inseparable transform sets.

[0111] In any part or combination of the above embodiments, the plurality of transform sets include at least one of: all secondary transform sets, all primary transform sets, or a combination of the secondary transform sets and the primary transform sets.

[0112] In any part or combination of the above embodiments, the multiple transform sets include all primary transform sets; in response to the transform set index belonging to the first subset, the secondary transform is implicitly enabled; and / or in response to the transform set index belonging to the second subset, the secondary transform is implicitly disabled.

[0113] In any part or combination of the above embodiments, determining the transform kernels in the determined transform set includes: determining the transform kernels in the determined transform set, wherein the transform kernels in the determined transform set are reordered based on the encoded information.

[0114] In any part or combination of the above embodiments, determining the transform set according to the transform set index includes: determining the transform set according to the transform set index, wherein the transform sets are reordered based on the encoded information.

[0115] In various embodiments of the present application, a transform set indicates a group of multiple transform kernels / basis and one transform kernel / basis.

[0116] In various embodiments, when selecting a transform set for a residual block in any given intra mode, multiple transform sets are available and the index of the selected transform set may be signaled, eg, the transform set index (tx_set_idx).

[0117] In some embodiments, the method of signaling a transform set index to indicate one transform set among at least two transform sets may be used only when the current coding block is intra-coded.

[0118] In some embodiments, the transform set index (tx_set_idx) can be entropy encoded using a context value that depends on the encoded information, which may include some or all of the following: intra-frame prediction mode, inter-frame prediction mode, transform partition mode / depth, transform coefficient value, last non-zero position (or EOB), number of non-zero transform coefficients, or quantization parameter.

[0119] In some embodiments, the transform set index (tx_set_idx) is conditionally entropy coded (meaning that the syntax is entropy coded under some conditions and not under other conditions), and the condition depends on the encoded information, which may include some or all of the following: intra prediction mode, transform partition mode / depth, transform coefficient value, last non-zero position (or EOB), number of non-zero transform coefficients. For example, when the intra prediction mode of the current block is the first predefined mode, the condition is met and the transform set index (tx_set_idx) is entropy coded; and when the intra prediction mode of the current block is the second predefined mode, the condition is not met and the transform set index (tx_set_idx) is not entropy coded.

[0120] In some embodiments, for different intra prediction modes, different transform partition modes / depths, and / or different transform coefficient values, and / or different last non-zero positions (or EOBs), and / or different numbers of non-zero transform coefficients, the transform set candidates may have overlapping transform sets. For example, when the intra prediction mode of the current block is the first mode, the transform set candidates for the current block may include 1, 3, 5, and 6; and when the intra prediction mode of the current block is the second mode, the transform set candidates for the current block may include 4, 5, 6, and 7, where 5 and 6 are overlapping transform sets for the first mode and the second mode.

[0121] In some embodiments, the method is applied when the transform set is a non-separable transform set. Alternatively, the method is applied when the transform set is a separable transform set. Alternatively, the method is applied when the transform set is a mixture of separable and non-separable transform sets. In some embodiments, the transform set is only the secondary transform set, or only the primary transform set, or a combination of the primary transform set and the secondary transform set. In some embodiments, where the transform set is the primary transform set, the secondary transform is implicitly enabled or disabled for a signaled index belonging to the set.

[0122] In some embodiments, the transform candidates in the transform set can be reordered for better entropy coding given the encoded information. The encoded information may include some or all of the following: intra-frame prediction mode, inter-frame prediction mode, transform partition mode / depth, transform coefficient values, last non-zero position (or EOB), number of non-zero transform coefficients, quantization parameter. In some embodiments, the transform set can be reordered for better entropy coding given the encoded information. The encoded information may include some or all of the following: intra-frame prediction mode, inter-frame prediction mode, transform partition mode / depth, transform coefficient values, last non-zero position (or EOB), number of non-zero transform coefficients, quantization parameter.

[0123] Various embodiments of the present application may include a method for encoding a current block into a video code stream, which is performed by an encoder, including the inverse of any part or all of the process described for a decoder.

[0124] Various embodiments herein may include a method for encoding a current block for video streaming performed by one or more electronic devices (e.g., a streaming media player), including any part or all of the process for a decoder and / or any part or all of the process described for an encoder.

[0125] The above operations can be combined or arranged in any number or order as needed. At least two of the steps and / or operations can be performed in parallel. The embodiments and implementation methods in this application can be used alone or in combination in any order. In addition, each of the methods (or embodiments), encoders and decoders can be implemented by processing circuits (for example, one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium. The embodiments in this application can be applied to luminance blocks or chrominance blocks. The term block can be interpreted as a prediction block, a coding block or a coding unit, i.e., a CU. The term "block" can also be used here to refer to a transform block. In the following items, when a block size (or block dimension) is mentioned, it can refer to the block width or height, or the maximum value of the width and height, or the minimum value of the width and height, or the area size (width*height), or the aspect ratio of the block (width:height, or height:width).

[0126] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 15 A computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0127] Computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or similar mechanisms to create code comprising instructions that may be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly or through interpretation, microcode execution, etc.

[0128] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, and the like.

[0129] Figure 15 The components shown for the computer system (1800) are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the application. Neither should the arrangement of components be interpreted as constituting any dependency or requirement on any one component or combination of components illustrated in the exemplary embodiment of the computer system (1800).

[0130] The computer system (1800) may include certain human interface input devices. The input human interface devices may include one or more of the following (only one of each is depicted): keyboard (1801), mouse (1802), trackpad (1803), touch screen (1810), data gloves (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).

[0131] The computer system (1800) may also include certain human-computer interface output devices. Such human-computer interface output devices can stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1810), a data glove (not shown), or a joystick (1805), but there may also be tactile feedback devices that are not used as input devices), audio output devices (such as: speakers (1809), headphones (not depicted)), visual output devices, and printers (not depicted). Visual output devices include screens (1810), virtual reality glasses (not depicted), holographic displays, and smoke canisters (not depicted). Screens (1810) include CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output in a manner such as stereo output.

[0132] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including media (1821) such as CD / DVD ROM / RW (1820) including CD / DVD, thumb drives (1822), removable hard drives or solid-state drives (1823), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), and the like.

[0133] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.

[0134] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, wireless, wired, or optical. The network may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide-area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial networks including CAN buses, etc.

[0135] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel (1840) of the computer system (1800).

[0136] The core (1840) may include one or more central processing units (CPUs) (1841), graphics processing units (GPUs) (1842), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (1843), hardware accelerators (1844) for certain tasks, a graphics adapter (1850), and the like. These devices, along with read-only memory (ROM) (1845), random access memory (1846), and internal mass storage devices (1847) such as internal non-user accessible hard drives, SSDs, and the like, may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to enable expansion with additional CPUs, GPUs, and the like. Peripheral devices may be attached directly to the core's system bus (1848) or to the core's system bus (1848) via a peripheral bus (1849). In an example, a screen (1810) may be connected to a graphics adapter (1850). Architectures for peripheral buses include PCI, USB, and the like.

[0137] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of this application, or they may be of a type well known and available to those skilled in the art of computer software.

[0138] Although several exemplary embodiments have been described herein, there are changes, permutations, and various substitute equivalents that fall within the scope of the present application. It will therefore be appreciated that one skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and are therefore within the spirit and scope of the present application.

Claims

1. A method for decoding a current block of a current frame in an encoded video stream, characterized in that: The method comprises: A device receives the encoded video stream, the device comprising a memory storing instructions and a processor communicating with the memory; The device extracts the transform coefficient set of the current block based on the encoded video stream; The device determines, based on a syntax element explicitly signaled in the encoded video stream, a transform set index of the current block in an intra prediction mode, the transform set index indicating a transform set in at least two transform sets; The device determines the transform set according to the transform set index; The apparatus performs inverse transform using the transform coefficient set and the determined transform set to obtain a residual block of the current block; and The device reconstructs the current block based on the residual block.

2. The method according to claim 1, wherein: The determined transform set includes a secondary transform set; and Performing the inverse transform using the transform coefficient set and the determined transform set to obtain the residual block of the current block includes: An inverse secondary transform of the transform coefficient set is performed based on the secondary transform set to generate a primary transform coefficient set for the current block.

3. The method according to claim 1, characterized in that Determining the transform set index of the current block includes: In response to the current block being in the intra prediction mode, the transform set indexes corresponding to the at least two transform sets are extracted.

4. The method according to claim 1, wherein Determining the transform set index of the current block includes: The coded video stream is entropy decoded based on a context dependent on coded information to obtain the transform set index.

5. The method according to claim 4, characterized in that The encoded information includes at least one of: intra prediction mode, inter prediction mode, transform partition mode or depth, transform coefficient values, last non-zero position or end of block (EOB), number of non-zero transform coefficients, or quantization parameter.

6. The method according to claim 1, characterized in that Determining the transform set index of the current block includes: In response to a condition being met, entropy decoding is performed on the encoded video stream to obtain the transform set index, wherein the condition depends on the encoded information.

7. The method according to claim 6, characterized in that The encoded information includes at least one of: intra prediction mode, inter prediction mode, transform partition mode or depth, transform coefficient values, last non-zero position or end of block (EOB), number of non-zero transform coefficients, or quantization parameter.

8. The method according to claim 1, wherein: A first at least two transform sets for a first block overlap with a second at least two transform sets for a second block, wherein the first block and the second block have at least one of: different intra prediction modes, different transform partitioning modes or depths, different transform coefficient values, different last non-zero positions or EOBs, or different numbers of non-zero transform coefficients.

9. The method according to claim 1, wherein: The at least two transform sets include at least one of: all separable transform sets, all non-separable transform sets, or a list of separable transform sets and non-separable transform sets.

10. The method according to claim 1, wherein: The at least two transform sets include at least one of: all secondary transform sets, all primary transform sets, or a combination of the secondary transform sets and the primary transform sets.

11. The method according to claim 1, wherein: The at least two transform sets include all primary transform sets; In response to the transform set index belonging to the first subset, implicitly enabling a secondary transform; or In response to the transform set index belonging to the second subset, a secondary transform is implicitly disabled.

12. The method according to claim 1, characterized in that Further including: Transform kernels are determined in the determined transform set, wherein the transform kernels in the determined transform set are reordered based on the encoded information.

13. The method according to claim 1, wherein Determining the transform set according to the transform set index includes: The transform set is determined according to the transform set index, wherein the transform sets are reordered based on the encoded information.

14. A device for decoding a current block of a current frame in an encoded video stream, characterized in that: The device comprises: a memory for storing instructions; and A processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1 to 13.

15. A non-volatile computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by a processor, the instructions are used to cause the processor to perform the method according to any one of claims 1 to 13.