Transforming dictionaries

By using transform dictionary to improve transform set signaling of residual blocks, the problem of redundancy reduction in the existing video encoding technology is solved, and more efficient video data compression and transmission is achieved.

CN120303941APending Publication Date: 2025-07-11TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082655.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2023-11-30
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing video encoding techniques have inefficient problems in reducing redundancy in uncompressed input video signals, especially in storage and streaming requirements for bandwidth are not effectively met.

Method used

The signaling method of improving the transform set of residual blocks by using a transform dictionary includes determining the transform heap, transform set, and transform cores to achieve more efficient video decoding and encoding.

Benefits of technology

Improves the compression efficiency during video encoding and decoding, reduces the amount of data, and makes more efficient use of storage and transmission bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303941A_ABST
    Figure CN120303941A_ABST
Patent Text Reader

Abstract

The present application generally relates to video encoding / decoding, in particular to transform dictionaries. The method comprises the following steps: receiving an encoded video code stream; determining a transform stack of the current block based on the encoded video stream, the transform stack being one of a transform dictionary, the transform dictionary comprising a plurality of transform stacks, and the transform stack comprising a plurality of transform sets; determining a transform set in the transform stack; determining a transformation kernel in the transformation set; performing inverse transform on the current block based on the transform kernel to obtain a residual block of the current block; and reconstructing the current block based on the residual block.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by Reference

[0001] This application claims priority to U.S. Provisional Application No. 63 / 533,571, filed on August 18, 2023, the entire content of which is incorporated herein by reference. This application also claims priority to U.S. Non - Provisional Patent Application No. 18 / 521,504, filed on November 28, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application describes a series of advanced video / streaming encoding / decoding techniques. More specifically, the disclosed techniques relate to a transform dictionary. Background Art

[0003] Uncompressed digital video may include a series of images and may have specific bit - rate requirements for storage, data processing, and transmission bandwidth in streaming applications. One purpose of video encoding and decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention

[0004] This application describes various embodiments of methods, apparatuses, and computer - readable storage media for improving the signaling of the transform set of residual blocks by using a transform dictionary.

[0005] According to one aspect, an embodiment of this application provides a method for decoding a current block of a current frame in an encoded video bitstream. The method includes: receiving, by a device, the encoded video bitstream. The device includes a memory storing instructions and a processor communicating with the memory. The method further includes: determining, by the device according to the encoded video bitstream, a transform heap of the current block, the transform heap being one of the transform heaps in a transform dictionary, the transform dictionary including a plurality of transform heaps, and a transform heap including a plurality of transform sets; determining, by the device, the transform set in the transform heap; determining, by the device, the transform kernel in the transform set; performing, by the device, an inverse transform on the current block based on the transform kernel to obtain a residual block of the current block; and reconstructing, by the device, the current block based on the residual block.

[0006] According to another aspect, an embodiment of this application provides an apparatus for processing a current block of a current frame in an encoded video bitstream. The apparatus includes a memory storing instructions and a processor communicating with the memory. When the processor executes these instructions, the processor is configured to cause the apparatus to perform the above - mentioned video decoding and / or encoding method.

[0007] In another aspect, an embodiment of this application provides a non - transitory computer - readable medium storing instructions that, when executed by a computer, cause the computer to perform the above - mentioned video decoding and / or encoding method.

[0008] The above aspects and other aspects and their implementations will be described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Other features, properties, and various advantages of the subject matter disclosed in this application will become more apparent from the following detailed description and the drawings, in which:

[0010] Figure 1 A simplified block diagram schematic of a communication system (100) according to an exemplary embodiment is shown;

[0011] Figure 2 A simplified block diagram schematic of a communication system (200) according to an exemplary embodiment is shown;

[0012] Figure 3 A simplified block diagram schematic of a video decoder according to an exemplary embodiment is shown;

[0013] Figure 4 A simplified block diagram schematic of a video encoder according to an exemplary embodiment is shown;

[0014] Figure 5 A block diagram of a video encoder according to another exemplary embodiment is shown;

[0015] Figure 6 A block diagram of a video decoder according to another exemplary embodiment is shown;

[0016] Figure 7 A coding block partitioning scheme according to an exemplary embodiment of this application is shown;

[0017] Figure 8 Another coding block partitioning scheme according to an exemplary embodiment of this application is shown;

[0018] Figure 9 Another coding block partitioning scheme according to an exemplary embodiment of this application is shown;

[0019] Figure 10 Exemplary fine angles in directional intra-prediction are shown;

[0020] Figure 11 A nominal angle in directional intra-prediction is shown;

[0021] Figure 12 A schematic diagram showing the use of a secondary transform during encoding and decoding is shown;

[0022] Figure 13 An exemplary hierarchy of a transform dictionary is shown;

[0023] Figure 14 shows an exemplary logic flow diagram of the method of the present application;

[0024] Figure 15 shows a schematic diagram of a computer system according to an exemplary embodiment of the present application. Detailed implementation manners

[0025] The present invention will be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of the embodiments in a diagrammatic manner. However, it should be noted that the present invention can be implemented in many different forms, and thus the subject matter covered or claimed should be understood as not being limited to any of the embodiments described below. It should also be noted that the present invention can be implemented as a method, a device, a component, or a system. Therefore, the embodiments of the present invention can, for example, take the form of hardware, software, firmware, or any combination thereof.

[0026] Throughout the specification and the claims, the meaning of terms may have subtle meanings (beyond their explicit literal meanings) that can be represented or implied according to the context. The terms "in one embodiment" or "in some embodiments" do not necessarily refer to the same embodiment, and the terms "in another embodiment" or "in other embodiments" do not necessarily refer to different embodiments. Similarly, the terms "in one implementation" or "in some implementations" do not necessarily refer to the same implementation, and the terms "in another implementation" or "in other implementations" do not necessarily refer to different implementations. For example, the claimed subject matter is intended to cover combinations of all or part of the exemplary embodiments / implementations.

[0027] Generally, the terms are understood at least in part based on their use in the context. For example, terms such as "and", "or", or "and / or" used herein may have multiple meanings, which may at least in part depend on the context in which this term is used. Generally, if "or" is used to associate a list such as A, B, or C, it is meant to represent A, B, and C (inclusive meaning here), as well as A, B, or C (exclusive meaning here). In addition, at least in part depending on the context, the terms "one or more" or "at least one" used herein can be used to describe any feature, structure, or property in a single way, or can be used to describe a combination of multiple features, structures, or properties. Similarly, at least in part depending on the context, terms such as "a", "an", or "the" may also be understood to express singular or plural usage. It should also be noted that at least in part depending on the context, the terms "based on" or "determined by" can be understood as not necessarily meaning to express a set of exclusive factors, but can allow for the existence of other factors not explicitly described.

[0028] As Figure 1As shown, the terminal device can be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles of the present application is not limited thereto. Embodiments of the present application can be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, and / or similar devices. The network (150) represents any number or type of network for transmitting encoded video data between terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0029] As an example of the application of the subject matter disclosed in the present application, Figure 2 shows the placement of a video encoder and a video decoder in a video streaming environment. The subject matter disclosed in the present application is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0030] As Figure 2 shown, a video streaming system can include a video capture subsystem (213), and the video capture subsystem (213) can include, for example, a video source (201) of a digital camera for creating an uncompressed video picture or image stream (202). In an embodiment, the video picture stream (202) includes samples recorded by the digital camera of the video source (201). Compared with the encoded video data (204) (or encoded video bitstream), the video picture stream (202) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (202) can be processed by an electronic device (220) that includes a video encoder (203) coupled to the video source (201). The video encoder (203) can include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (202), the encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (204) (or encoded video bitstream (204)), which can be stored on a streaming server (205) for future use. One or more streaming client subsystems, such as Figure 2The client subsystems (206) and client subsystem (208) therein can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) can include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and generates an output video picture stream (211) that can be presented on a display (212) (such as a display screen) or another presentation device (not depicted).

[0031] Figure 3 A block diagram of a video decoder (310) in an electronic device (330) is shown according to the following embodiments of the present application. The electronic device (330) can include a receiver (331) (such as a receiving circuit). The video decoder (310) can be used to replace Figure 2 the video decoder (210) in the embodiment.

[0032] As Figure 3 shown, the receiver (331) can receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or handle the playback time, a buffer memory (315) can be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) can reconstruct symbols (321) according to the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (310) and potential information for controlling a display device (312) (such as a display screen), etc. The parser (320) can parse / entropy decode the encoded video sequences. The parser (320) can extract subgroup parameter sets for at least one subgroup among subgroups of pixels in the video decoder from the encoded video sequences. Subgroups can include Group of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (320) can also extract information from the encoded video sequences, such as transform coefficients (such as Fourier transform coefficients), quantizer parameter values, motion vectors, and so on. The reconstruction of the symbols (321) can involve multiple different processing processes and functional units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.

[0033] The first unit may be a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive quantized transform coefficients as symbols (521) and control information including information indicating which type of inverse transform, block size, quantization factor / parameter, and quantization scaling matrix to use. The scaler / inverse transform unit (351) may output a block including sample values, and the sample values may be input into an aggregator (355).

[0034] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some embodiments, the aggregator (355) adds the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0035] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In this case, a motion compensation prediction unit (353) may access a reference picture memory (357) based on a motion vector to extract samples for inter-picture prediction. After motion-compensating the extracted reference samples according to the symbols (321) belonging to the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of the scaler / inverse transform unit (351) is referred to as a residual sample or a residual signal) to generate output sample information.

[0036] The output samples of the aggregator (355) may be employed by various loop filtering techniques in a loop filter unit (356), and the loop filter unit (356) includes several types of loop filters. The output of the loop filter unit (356) may be a sample stream, and the sample stream may be output to a rendering device (312) and stored in the reference picture memory (357) for subsequent inter-picture prediction.

[0037] Figure 4 is a block diagram of a video encoder (403) according to an exemplary embodiment disclosed in the present application. The video encoder (403) may be disposed in an electronic device (420). The electronic device (420) may further include a transmitter (440) (e.g., a transmission circuit). The video encoder (403) may be used to replaceFigure 4 The video encoder (403) in the embodiment.

[0038] The video encoder (403) may receive video samples from a video source (401). According to some exemplary embodiments, the video encoder (403) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Performing an appropriate encoding speed constitutes a function of the rate control component (450). In some embodiments, the component (450) may be functionally coupled to other functional units as described below and control these functional units. The set of parameters set by the component (450) may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, and so on.

[0039] In some exemplary embodiments, the video encoder (403) may be configured to operate in an encoding loop. The encoding loop may include a source encoder (430) and an (in - loop) decoder (433) embedded in the video encoder (403). Although the in - loop decoder (433) processes the encoded video stream through the source encoder 430 without entropy encoding, the decoder (433) reconstructs symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (because in the video compression techniques contemplated by the subject matter disclosed in this application, any compression between the symbols in entropy encoding and the encoded video bitstream can be lossless). At this point, it can be observed that any decoder technique other than parsing / entropy decoding that exists only in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the subject matter disclosed in this application may sometimes focus on decoder operations, which correspond to the decoding part of the encoder. Thus, the description of encoder techniques can be simplified because encoder techniques are inverse to the decoder techniques described in full. A more detailed description of the encoder is only needed in certain areas or aspects and is provided below.

[0040] During operation, in some exemplary embodiments, the source encoder (430) may perform motion - compensated predictive encoding, referring to one or more previously encoded pictures in the video sequence designated as "reference pictures", and the motion - compensated predictive encoding performs predictive encoding on the input picture.

[0041] The local video decoder (433) can decode the encoded video data of a picture that can be specified as a reference picture. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference picture, and can store the reconstructed reference picture in the reference picture cache (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture, which has the same content (no transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0042] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) can search in the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata that can be used as an appropriate prediction reference for the new picture, such as a reference picture motion vector, block shape, etc.

[0043] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0044] The outputs of all the above functional units can be entropy encoded in the entropy encoder (445). The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission through the communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0045] The controller (450) can manage the operations of the video encoder (403). During encoding, the controller (450) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types: intra picture (I picture), predictive picture (P picture), bi-predictive picture (B picture), and multi-predictive picture. The source picture can generally be spatially subdivided into a plurality of sample coding blocks, which will be described in further detail below.

[0046] Figure 5 is a block diagram of a video encoder (503) according to another exemplary embodiment disclosed in the present application. The video encoder (503) is used to receive sample values (such as prediction blocks) within a current video picture in a video picture sequence, and encode the processing blocks into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (503) can be used to replaceFigure 4 The video encoder (403) in the embodiment.

[0047] In an embodiment, the video encoder (503) receives a matrix of sample values for processing blocks. Then, the video encoder (503) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-predictive mode to optimally encode the processing blocks.

[0048] In Figure 5 an embodiment of, the video encoder (503) includes an inter encoder (530), an intra encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in the exemplary arrangement of Figure 5 .

[0049] The inter encoder (530) is used to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in the display order), generate inter prediction information (e.g., a description of redundancy information, a motion vector, merge mode information according to inter coding techniques), and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique.

[0050] The intra encoder (522) is used to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, generate quantization coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques).

[0051] The general controller (521) can be used to determine general control data and control other components of the video encoder (503) based on the general control data to, for example, determine a prediction mode of a block and provide a control signal to the switch (526) based on the prediction mode.

[0052] The residual calculator (523) can be used to calculate the difference (residual data) between the received block and a prediction result for the block selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) can be used to encode the residual data to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is used to perform an inverse transform and generate decoded residual data. The entropy encoder (525) can be used to format a bitstream to produce an encoded block and perform entropy coding.

[0053] Figure 6 is a block diagram of an exemplary video decoder (610) according to another embodiment disclosed in the present application. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an embodiment, the video decoder (610) can be used to replace Figure 4 the video decoder (410) in the embodiment.

[0054] In Figure 6 the embodiment of, the video decoder (610) includes an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together as shown in the exemplary arrangement of Figure 6 .

[0055] The entropy decoder (671) can be used to reconstruct certain symbols according to the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. The inter-frame decoder (680) can be used to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) can be used to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) can be used to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The reconstruction module (674) can be used to combine the residual output by the residual decoder (673) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, and the reconstructed block forms part of a reconstructed picture, and the reconstructed picture is part of a reconstructed video.

[0056] It should be noted that any suitable technology can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In some exemplary embodiments, one or more integrated circuits can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610).

[0057] For block partitioning for encoding and decoding, the normal partitioning can start from a base block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After partitioning or splitting the base block according to any of the exemplary partitioning processes described below or other processes or a combination thereof, a set of final partitions or coding blocks can be obtained. Each of these partitions can be at one of the partitioning levels in the partitioning hierarchy and can be of various shapes. Each partition can be called a coding block (CB). For the various exemplary partitioning embodiments described further below, each resulting CB can have any allowed size and partitioning level. Such partitions are called coding blocks because these partition areas can form units for which some basic encoding / decoding decisions can be made and for which encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning tree structure. A coding block can be a luminance coding block or a chrominance coding block. The CB tree structure for each color can be called a coding block tree (CBT). The coding blocks for all color channels can be collectively referred to as coding units (CUs). The hierarchies for all color channels can be collectively referred to as coding tree units (CTUs). The partitioning pattern or structure for each color channel in a CTU can be the same or different.

[0058] In some embodiments, the partitioning tree scheme or structure for the luminance channel and the chrominance channel may not have to be the same. In other words, the luminance channel and the chrominance channel can have independent coding tree structures or patterns. Additionally, whether the luminance channel and the chrominance channel use the same or different coding partitioning tree structures, and the actual coding partitioning tree structure used, may depend on whether the slice being encoded is a P slice, a B slice, or an I slice. For example, for an I slice, the chrominance channel and the luminance channel can have independent coding partitioning tree structures or coding partitioning tree structure patterns, while for a P slice or a B slice, the luminance channel and the chrominance channel can share the same coding partitioning tree scheme. When an independent coding partitioning tree structure or pattern is applied, the luminance channel can be partitioned into CBs by one coding partitioning tree structure, while the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.

[0059] Figure 7 An exemplary predefined 10-way partitioning structure / pattern is shown, which allows recursive partitioning to form a partitioning tree. The root block can start at a predefined level (e.g., starting from a base block of 128×128 or 64×64). Figure 7Exemplary partitioning structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some exemplary embodiments, Figure 7 the rectangular partitions in cannot be further subdivided. The coding tree depth can be further defined to indicate the depth of splitting from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after further splitting the root block once according to Figure 7 the coding tree depth increases by 1. In some embodiments, it may be possible to recursively partition all square partitions 710 into the next level of the partition tree only according to the pattern in Figure 7 .

[0060] In some other exemplary embodiments of the coding block partitioning, a quadtree structure can be used. Such quadtree splitting can be applied hierarchically and recursively to any square-shaped partition. Whether to perform further quadtree splitting on the base block or intermediate block or partition can be adjusted according to the local characteristics of the base block or intermediate block / partition.

[0061] In some other examples, a ternary partitioning scheme can be used to partition the base block or any intermediate block, as shown in Figure 8 . The ternary partitioning pattern can be implemented vertically (as shown in 802) or horizontally (as shown in 804). Although the exemplary split ratio in Figure 8 is shown as 1:2:1, it can also be other predefined ratios. In some embodiments, two or more different ratios can be predefined. In some embodiments, the width and height of the partitions in the exemplary ternary tree are always powers of 2 to avoid additional transformations.

[0062] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the quadtree partitioning and binary partitioning schemes described above can be combined to partition the base block into a quadtree-binary-tree (QTBT) structure. In such a scheme, the base block or intermediate block / partition can be selected for quadtree splitting or binary splitting according to a set of predefined conditions (if specified). Figure 9Shows a specific example where the base block quadtree is first divided into four partitions, as shown at 902, 904, 906, and 908. Subsequently, for each resulting partition, either it is further quadtree-divided into four smaller partitions (as shown at 908), or it is binary-divided into two smaller partitions at the next level (which can be horizontal or vertical, as shown at 902 or 906, both of which are symmetric), or it is not divided (as shown at 904). For square-shaped partitions, binary division or quadtree division is allowed recursively, as shown by the overall exemplary division pattern at 910 and the corresponding tree structure / representation at 920, where solid lines represent quadtree division and dashed lines represent binary division. A flag can be used for each binary division node (non-leaf binary partition) to indicate whether the binary division is horizontal or vertical. For example, as shown at 920 and consistent with the division structure of 910, the flag "0" can represent a horizontal binary division, and the flag "1" can represent a vertical binary division. For quadtree-divided partitions, generally there is no need to indicate the division type, because quadtree division always divides the block or partition both horizontally and vertically into four sub-blocks / partitions of equal size. In some embodiments, the flag "1" can represent a horizontal binary division, while the flag "0" can represent a vertical binary division.

[0063] In some exemplary embodiments of the QTBT, the rule sets for quadtree division and binary division can be represented by the following predefined parameters and their associated functions: - CTU size: The size of the root node of the quadtree (the size of the base block) - MinQTSize: The minimum allowable size of a quadtree leaf node - MaxBTSize: The maximum allowable size of a binary tree root node - MaxBTDepth: The maximum allowable depth of a binary tree - MinBTSize: The minimum allowable size of a binary tree leaf node

[0064] In some exemplary embodiments of the QTBT partitioning structure, the CTU size can be set to a block of 128×128 luma samples and two corresponding 64×64 chroma samples (when considering the use of exemplary chroma subsampling). MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be applied to the CTU first to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from its allowed minimum size of 16×6 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the node is 128×128, it will not be split by binary tree first because its size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes with a size not exceeding MaxBTSize can be split by binary tree. In Figure 9 the example of Figure 9 , the base block is 128×128. The base block can only be split by quadtree according to a predefined set of rules. The partitioning depth of the base block is 0. Each of the four resulting partitions is 64×64, which does not exceed MaxBTSize, and can be further split by quadtree or binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further splitting will be considered. When the width of a binary tree node is equal to MinBTSize (i.e., 4), no further horizontal splitting will be considered. Similarly, when the height of a binary tree node is equal to MinBTSize, no further vertical splitting will be considered.

[0065] In some exemplary embodiments, the above QTBT scheme can be configured to support the same QTBT structure or independent QTBT structures for the luma channel and the chroma channel. For example, for P slices and B slices, the luma CTB and the chroma CTB in a CTU may share the same QTBT structure. However, for I slices, the luma CTB may be partitioned into coding blocks (CBs) through a QTBT structure, while the chroma CTB may be partitioned into chroma coding blocks (chroma CBs) through another QTBT structure. This means that: CUs in I slices can be used to represent different color channels. For example, an I slice may consist of coding blocks of the luma component or coding blocks of two chroma components, while CUs in P slices or B slices may consist of coding blocks of all three color components.

[0066] The above various CB partitioning schemes and the schemes for further partitioning CBs into PBs can be combined in any way. Specific embodiments are provided below as non-limiting examples.

[0067] In some embodiments, a set of intra prediction modes (interchangeably referred to as "intra modes") may include a predefined number of directional intra prediction modes. These intra prediction modes may correspond to a predefined number of directions along which samples outside the block are selected to predict the samples to be predicted in a particular block. In another specific exemplary embodiment, 8 main direction modes may be supported and predefined, and these main direction modes correspond to angles from 45 degrees to 207 degrees with respect to the horizontal axis. In some other embodiments of intra prediction, in order to further exploit more variations of spatial redundancy in the directional texture, the directional intra modes may be further extended to have a finer-grained set of angles. For example, as Figure 11 shown, the implementation of the above 8 angles may be configured to provide 8 nominal angles, namely V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED respectively, and for each nominal angle, a predefined number (e.g., 7) of finer angles may be added. Through this kind of extension, a larger total number of direction angles (e.g., 56 in this example) can be provided for intra prediction, and these numbers of angles correspond to the same number of predefined directional intra prediction modes. The prediction angle can be expressed as the nominal intra angle plus the angle increment. For the above specific example, each nominal angle has 7 finer angle directions, and the angle increment can be in steps of -3 to 3 multiplied by 3 degrees. Some angle schemes can be used, such as Figure 10 shown, using 65 different prediction angles. In some embodiments, the 8 nominal modes and 5 non-angle smoothing modes are first signaled, and then if the current mode is an angle mode, an index is further signaled to indicate the angle increment of the corresponding nominal angle. In some embodiments, in order to implement the directional prediction mode in a general way, all 56 directional intra prediction modes can be implemented using a unified direction predictor that projects each pixel to a reference sub-pixel position and interpolates the reference pixels through a 2-tap bilinear filter.

[0068] Then, the residuals of the intra-prediction block or the inter-prediction block can be transformed, and then the transform coefficients can be quantized. To perform the transform, before the transform is carried out, both the intra-coded block and the inter-coded block can be further divided into a plurality of transform blocks (sometimes also referred to as "transformation units", although "unit" is usually used to represent a set of three color channels, for example, an "encoding unit" will include a luminance coding block and a chrominance coding block). In some embodiments, the maximum partitioning depth of the coding block (or prediction block) can be specified (the terms "coded block" and "coding block" can be used interchangeably). For example, such partitioning may not exceed 2 levels. Between the intra-prediction block and the inter-prediction block, the prediction block can be divided into transform blocks in different ways. However, in some embodiments, such partitioning between the intra-prediction block and the inter-prediction block can be similar.

[0069] In some exemplary embodiments, for an inter-coded block, the transform unit partitioning can be performed in a recursive manner to a predefined maximum number of levels (e.g., level 2). For any sub-partition and at any level, the splitting may stop or continue recursively. For example, a block is split into four quadtree sub-blocks, one of these sub-blocks is further split into four transform blocks of the second level, while the splitting of the other sub-blocks will stop after the first level, resulting in a total of 7 transform blocks with two different sizes. In some embodiments, the transform partitioning can support transform block shapes of 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, with a size range from 4×4 to 64×64. In some exemplary embodiments, if the coding block is less than or equal to 64×64, the transform block partitioning can be applied only to the luminance component (in other words, under this condition, the chrominance transform block will be the same as the coding block). Otherwise, if the width or height of the coding block is greater than 64, the luminance coding block and the chrominance coding block can be implicitly divided into a plurality of min(W, 64)×min(H, 64) and min(W, 32)×min(H, 32) transform blocks, respectively.

[0070] Then each of the above transform blocks can be subjected to a primary transform. The primary transform essentially converts the residuals in the transform block from the spatial domain to the frequency domain. In some embodiments of the actual primary transform, to support the above-exemplified extended coding block partitioning, multiple transform sizes (from 4 points to 64 points for each of the two dimensions) and transform shapes (square; rectangles with an aspect ratio of 2:1 / 1:2 and 4:1 / 1:4) can be allowed.

[0071] For an actual primary transform, in some exemplary embodiments, a two-dimensional (2-D) transform process may involve using a hybrid transform kernel (e.g., for each dimension of an encoded residual transform block, the transform kernel may consist of different one-dimensional (1-D) transforms). Exemplary 1-D transform kernels may include, but are not limited to: a) 4-point, 8-point, 16-point, 32-point, 64-point DCT-2; b) 4-point, 8-point, 16-point asymmetric DST (DST-4, DST-7) and their flipped versions; c) 4-point, 8-point, 16-point, 32-point identity transform. The transform kernel to be used for each dimension may be selected based on the rate-distortion (RD) criterion.

[0072] In some exemplary embodiments, the availability of a hybrid transform kernel for a particular primary transform embodiment may be based on the transform block size and the prediction mode. For the chrominance component, the selection of the transform type may be performed in an implicit manner. For example, for the intra prediction residual, the transform type may be selected according to the intra prediction mode. For the inter prediction residual, the transform type of the chrominance block may be selected according to the transform type selection of the co-located luma block. Therefore, for the chrominance component, the transform type is not written into the bitstream.

[0073] In some embodiments, for an intra prediction residual block of the luma color component, before applying quantization at the decoder, a secondary transform method, i.e., intra secondary transform (IST), may be applied to the primary transform coefficient block. Correspondingly, before applying the inverse primary transform at the decoder, a secondary inverse transform may be applied to the dequantized transform coefficient block. IST is not applied to the chrominance color component. A schematic diagram of using IST in the encoding and decoding processes is as Figure 12 shown.

[0074] In some embodiments using IST, a non-separable transform process will be applied. To apply a forward non-separable transform to a specific region of an input transform coefficient block consisting of N samples, first, according to the relative coordinates of each sample in the N input samples, the N samples are organized into an N×1 vector using the coefficient scan order Then, an N×N transform kernel (K) is selected, and the non-separable transform is performed using the following arithmetic operation: where is the output N×1 vector that replaces a specific region of the transform coefficient block using the coefficient scan order.

[0075] In some embodiments, to apply the inverse non-separable transform, first, the dequantized transform coefficient block is given as the input, and then the specific region of the dequantized transform coefficient block is identified based on the transform block size. By reorganizing the transform coefficients using the coefficient scan order, an input vector is formed Given a selected N×N transform kernel and an input vector Perform an inverse non-separable transform using the following arithmetic operations: where (·) T denotes the matrix transpose operation. The output is an N×1 vector that replaces a specific region of the input block according to a specific coefficient scan order.

[0076] In some embodiments, the input to the forward IST is a coefficient vector composed of low-frequency main transform coefficients in a zig-zag scan. Depending on the size of the block, a 16-point or 64-point non-separable quadratic transform can be selected. When the minimum of the main transform width and the main transform height is less than 8, a 16-point IST is used, and the low-frequency main transform coefficients refer to the first 16 main transform coefficients in the zig-zag scan order. When both the main transform width and height are greater than or equal to 8, a 64-point IST is applied, and the low-frequency main transform coefficients refer to the first 64 main transform coefficients in the zig-zag scan order. The 16-point non-separable transform uses an 8×16 transform kernel, and the 64-point non-separable transform uses a 32×64 transform kernel. Additionally, when the IST is applied, the high-frequency transform coefficients that are not processed by the quadratic transform are set to zero.

[0077] In some embodiments, a total of 12 quadratic transform sets (or IST sets) can be defined, each set containing 3 quadratic transform kernels. For each intra-coded transform block, first identify the nominal intra-prediction mode and the main transform type, and then select the IST set based on Table 1. In some embodiments, for the Paeth prediction mode and the recursive intra-prediction mode, neither the IST is applied nor is the IST signaled. Table 1: Mapping of intra-nominal modes and main transform types to IST sets

[0078] In some embodiments, given an IST set, there can be four encoder choices: 1) no secondary transform, 2) perform a secondary transform using the first transform kernel in the given IST set, 3) perform a secondary transform using the second transform kernel in the given IST set, and / or 4) perform a secondary transform using the third transform kernel in the given IST set. In some embodiments, the encoder signals the choice using the syntax element ist_idx. At the decoder, the value of the syntax element ist_idx is first parsed. Then, given the IST set and the value associated with ist_idx, the secondary transform kernel is identified. After signaling the primary transform type, for each intra-coded luminance transform block, the syntax element ist_idx is signaled. Signaling of ist_idx is performed only when all of the following conditions hold: the current block is an intra-coded luminance transform block; the primary transform type is DCT in both dimensions or ADST in both dimensions; the intra-prediction mode is neither the Paeth prediction mode nor the recursive intra-prediction mode; the transform partition depth is 0; and the end-of-block (EOB) position falls within the low-frequency transform coefficient region where the secondary transform is applied. In some embodiments, an entropy coding context for ist_idx is derived based on the transform block size.

[0079] In some embodiments, a secondary transform can be performed on the primary transform coefficients. For example, the low-frequency non-separable transform (LFNST), also known as the reduced secondary transform, can be applied (at the encoder) between the forward primary transform and quantization, and (at the decoder) between dequantization and the inverse primary transform to further decorrelate the primary transform coefficients. In essence, the LFNST can take a portion of the primary transform coefficients, such as the low-frequency portion (hence “reduced” from the complete set of primary transform coefficients of the transform block) for the secondary transform. In an exemplary LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform can be applied depending on the size of the transform block. For example, a 4×4 LFNST can be applied to smaller transform blocks (e.g., min(width, height) < 8), while an 8×8 LFNST can be applied to larger transform blocks (e.g., min(width, height) > 8). For example, if an 8×8 transform block is subjected to a 4×4 LFNST, only the low-frequency 4×4 portion of the 8×8 primary transform coefficients will further undergo the secondary transform.

[0080] In some embodiments, to perform a primary transform on a residual block, a primary transform kernel needs to be specified. The primary transform kernel is a primary transform matrix composed of transform basis vectors, and the transform process essentially involves performing matrix multiplication of the residual block with the given primary transform matrix. The output of the transform process is a block of transform coefficients. In some embodiments, to further perform a secondary transform on the block of transform coefficients, a secondary transform kernel needs to be specified. The secondary transform kernel is a secondary transform coefficient matrix, and the secondary transform process essentially involves performing matrix multiplication of the block of transform coefficients with the given secondary transform matrix. In some embodiments, transforms can be classified as separable transforms and non-separable transforms. A separable transform transforms a 2-D block by first performing a 1-D transform on each row (or column), which derives an intermediate block of coefficients, and then performing a 1-D transform on each column (or row) of the intermediate block of coefficients. Meanwhile, a 2-D non-separable transform defines its transform basis in 2-D format, i.e., each transform basis is 2-D rather than 1-D as in the separable transform. A feasible way to perform a non-separable 2-D transform is to reshape the 2-D block into a 1-D vector with the same number of elements, then apply the 1-D transform basis on the reshaped 1-D input vector, and then reshape the output back into a 2-D block of transform coefficients.

[0081] There are some problems / difficulties associated with how to organize multiple transform kernels. This application describes various embodiments of a transform dictionary, which can have a hierarchical structure to organize multiple transform kernels, thereby improving the efficiency of constructing, signaling, or switching transforms used in transform coding / decoding in various different situations, and improving transform coding / decoding for the field of video coding / encoding.

[0082] In various embodiments, the transform hierarchy may include two or more levels, such as three levels. In one example, the overall hierarchy can be referred to as a transform dictionary (or other appropriate name), the next level is called the transform pile level (or other appropriate name), the second next level is called the transform set level, and / or the third next level is called the transform kernel level. The transform dictionary defines all transform piles available for encoding and decoding. A transform pile contains multiple transform sets, and the selection of the transform set within the transform pile can be signaled or implicitly derived (e.g., based on encoded information such as an intra prediction mode).

[0083] As Figure 13 shown, the transform dictionary 1300 contains multiple transform piles (1310, 1320, 1330,..., and 1390).

[0084] Each transform heap may contain multiple transform sets. In some embodiments, the first transform heap (transform heap 1, 1310) may contain N_1 transform sets, where N_1 is a non-negative integer: for example, the first transform set (transform set 1, 1311), the second transform set (transform set 2, 1312), the third transform set (transform set 3, 1313),..., the N_1th transform set (transform set N_1, 1319). In some embodiments, the second transform heap (transform heap 2, 1320) may contain N_2 transform sets, where N_2 is a non-negative integer: for example, the first transform set (transform set 1, 1321), the second transform set (transform set 2, 1322), the third transform set (transform set 3, 1333),..., the N_2th transform set (transform set N_2, 1329). In some embodiments, the third transform heap (transform heap 3, 1330) may contain N_3 transform sets, where N_3 is a non-negative integer: for example, the first transform set (transform set 1, 1331), the second transform set (transform set 2, 1332), the third transform set (transform set 3, 1333),..., the N_3th transform set (transform set N_3, 1339). In some embodiments, the kth transform heap (transform heap k, 1390) may contain N_k transform sets, where N_k is a non-negative integer: for example, the first transform set (transform set 1, 1391), the second transform set (transform set 2, 1392), the third transform set (transform set 3, 1393),..., the N_kth transform set (transform set N_k, 1399).

[0085] Each transform set may contain multiple transform kernels, and the selection of transform kernels within the transform set may be signaled or implicitly derived (e.g., based on encoded information such as an intra prediction mode or the relative position of the transform block within the coded block).

[0086] In some embodiments, each transform heap may include multiple transform sets selected for better transform coding in certain situations. As a non-limiting example, one transform heap may include transform sets for better transform coding of gaming videos, while another transform heap may contain transform sets for better transform coding of conference videos.

[0087] In some embodiments, only one transform heap may be used for a sequence / group of pictures (GOP) / picture / sub-picture / slice / tile. In some embodiments, the selection of transform sets within the transform heap may be signaled or implicitly derived (e.g., based on encoded information such as an intra prediction mode).

[0088] The various embodiments and / or implementations described in this application can be executed individually or in any order of combination, and are applicable to decoding, encoding, or streaming. In addition, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). The one or more processors execute programs stored in a non-transitory computer-readable medium. In this application, the term block can be interpreted as a prediction block, an encoding block, or a coding unit (CU).

[0089] Figure 14 FIG. 1400 is a flowchart showing an exemplary method that follows the basic principles implemented for the transform dictionary as described above. The exemplary decoding method flow starts at 1401 and may include some or all of the following steps: S1410, receiving an encoded video bitstream; S1420, based on the encoded video bitstream, determining a transform heap for the current block, the transform heap being one of the transform heaps in the transform dictionary, the transform dictionary including a plurality of transform heaps, and the transform heap including a plurality of transform sets; S1430, determining a transform set in the transform heap; S1440, determining a transform kernel in the transform set; S1450, based on the transform kernel, performing an inverse transform on the current block to obtain a residual block of the current block; and / or, S1460, reconstructing the current block based on the residual block. The example method terminates at S1499. Method 1400 can be executed by a device including a memory and a processor, where the memory stores instructions and the processor communicates with the memory.

[0090] In any part or combination of the above embodiments, the determining of the transform heap for the current block includes some or all of the following: the device extracts syntax for indicating the transform heap based on the encoded video bitstream; and / or, the device selects a transform heap from the plurality of transform heaps based on the syntax.

[0091] In any part or combination of the above embodiments, the determining of the transform heap for the current block includes: the device derives an index for selecting a transform heap from the plurality of transform heaps based on at least one of the following: one or more quantization parameters, the spatial resolution of the encoded video bitstream, the temporal resolution of the encoded video bitstream, the bit depth of the encoded video bitstream, or the reconstructed samples from a previous picture group, picture, sub-picture, slice, or tile.

[0092] In any part or combination of the above embodiments, the determining of the transform heap for the current block includes some or all of the following: the device extracts a flag for indicating whether the transform heap is changed based on the encoded video bitstream; and / or, in response to the flag indicating that the transform heap is not changed, the device determines the previous transform heap of the previous block as the transform heap of the current block.

[0093] In any part or combination of the above embodiments, the determination of the transform heap for the current block includes some or all of the following: the device extracts an index group based on the encoded video bitstream, and the index group indicates a set of transform sets in a set of candidate transform sets; and / or, the device constructs a transform heap to include a set of transform sets based on the index group.

[0094] In any part or combination of the above embodiments, the determination of the transform heap for the current block includes some or all of the following: for the transform sets in the transform heap, the device extracts a flag indicating whether the transform set is updated based on the encoded video bitstream; and / or, in response to the flag indicating that the transform set is updated: the device extracts an index indicating the updated transform set from a set of candidate transform sets based on the encoded video bitstream; and / or, the device updates the transform heap by replacing the transform set with the updated transform set.

[0095] In any part or combination of the above embodiments, each transform heap among multiple transform heaps includes a different number of transform sets.

[0096] In any part or combination of the above embodiments, the determination of the transform heap for the current block includes: the device extracts syntax indicating the number of transform sets in the transform heap based on the encoded video bitstream.

[0097] In any part or combination of the above embodiments, a transform set is a transform set in the transform heap.

[0098] In any part or combination of the above embodiments, two or more different transform heaps in the transform dictionary include one or more identical transform sets.

[0099] In any part or combination of the above embodiments, the separable transform kernels in two or more different transform heaps in the transform dictionary are the same.

[0100] In any part or combination of the above embodiments, the separable non - Karhunen - Loève transform (non - KLT) kernels in two or more different transform heaps in the transform dictionary are the same.

[0101] In any part or combination of the above embodiments, different transform kernels in the transform dictionary include one or more identical bases.

[0102] In various embodiments of the present application, there may be at least one transform dictionary to select transform kernels for performing encoding and decoding of blocks. The transform dictionary includes multiple transform heaps.

[0103] In one embodiment of video encoding / decoding, a transform heap is selected for encoding a current sequence, group of pictures (GOP), picture, sub-picture, slice, or tile. The selection of the transform heap can be explicitly signaled at the sequence, GOP, picture, sub-picture, slice, or tile level; or, the selection of the transform heap can be implicitly derived based on encoded information including, but not limited to: quantization parameter, spatial and / or temporal resolution of the video, bit depth of the video, any other value of syntax signaled, reconstructed samples from a previous GOP, picture, sub-picture, slice, or tile.

[0104] In one embodiment of video encoding / decoding, a flag is signaled to indicate whether the transform heap is changed. When the value of the signaled flag indicates that the transform heap does not need to be changed, the same used transform heap is applied. Otherwise, when the value of the signaled flag indicates that the transform heap is changed, any method in any other embodiment / mode described in this application is used to further signal the change of the transform heap.

[0105] In one embodiment, a transform heap is selected for encoding a current sequence, GOP, picture, sub-picture, slice, or tile, and the selection of the transform heap can be specified by signaling the index of the selected transform set, where the selected transform set is selected from a group of transform sets for each slot of the transform sets in the transform heap.

[0106] In one example, for a specific slot of the transform sets within the transform heap, a flag is signaled to indicate whether the transform set needs to be updated. If the value of the signaled flag indicates that the transform set needs to be updated, then the index of the transform set selected from a group of transform sets (repository) is further signaled from a group of candidate transform sets.

[0107] In one embodiment, different numbers of transform sets can be specified for different transform heaps. For example, as Figure 13 shown, the N_1, N_2,... N_k specifying the number of transform sets for heap 1 to heap N can be different from each other. In some embodiments, the N_1, N_2,... N_k specifying the number of transform sets for heap 1 to heap N can be the same. In one embodiment, to signal the transform sets for a specific transform heap, for each slot of the transform sets in the transform heap, the index of the transform set selected from the repository (selected from a group of candidate transform sets) is signaled. In one embodiment, to signal the transform heap, the number of transform sets (i.e., the number of slots of the transform sets) is signaled.

[0108] In one embodiment, given a selected transform heap, at the level of signaling the selection of the transform heap, only the set of transforms included in the selected transform heap can be applied. In one embodiment, different transform heaps may share one or more sets of transforms. In one embodiment, for different transform heaps, the separable transform kernels are the same. In one embodiment, for different transform heaps, the separable non-Karhunen Lòeve Transform (non-KLT) kernels are the same. In one embodiment, different transform kernels may share some bases.

[0109] Various embodiments of the present application may include methods for encoding a current block into a video bitstream, which are performed by an encoder and include the inverse processes of any part or all of the processes described for the decoder.

[0110] Various embodiments of the present application may include methods for encoding a current block of a streamed video, which are performed by one or more electronic devices (e.g., a streaming media player) and include any part or all of the processes described for the decoder, and / or any part or all of the processes described for the encoder.

[0111] As needed, the above operations can be combined or arranged in any number or order. Two or more steps and / or operations can be performed in parallel. The embodiments and implementations in the present application can be performed separately or combined in any order. In addition, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in the present application can be applied to luminance blocks or chrominance blocks. The term "block" here can be interpreted as a prediction block, a coding block, or a coding unit (i.e., a CU). The term "block" can also be used to refer to a transform block. In the following content, when referring to the block size, it can refer to the width or height of the block, or the maximum of the width and height of the block, or the minimum of the width and height of the block, or the area size (width * height) of the block, or the aspect ratio (width: height, or height: width) of the block.

[0112] The above technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 15 A computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0113] Computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can undergo assembly, compilation, linking, or similar mechanisms to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0114] The instructions can be executed on various types of computers or their components, which include, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0115] Figure 15 The components of the computer system (1800) shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (1800).

[0116] The computer system (1800) may include certain human-machine interface input devices. The human-machine interface input devices may include one or more of the following (only one of each type is shown): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).

[0117] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., the tactile feedback of the touch screen (1810), data glove (not shown), or joystick (1805), but can also be a tactile feedback device that is not used as an input device), audio output devices (e.g., speakers (1809), headphones (not labeled)), visual output devices (e.g., screens (1810) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality, some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not labeled), holographic displays, and smoke boxes (not labeled), as well as printers (not labeled)).

[0118] The computer system (1800) may also include a human-accessible storage device and its associated media, such as an optical medium including a CD / DVD ROM / RW (1820) with a medium (1821) such as a CD / DVD, a thumb drive (1822), a removable hard disk drive or a solid-state drive (1823), traditional magnetic media such as tapes and floppy disks (not labeled), a dedicated ROM / ASIC / PLD-based device such as a security dongle (not labeled), etc.

[0119] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover a transmission medium, a carrier wave, or other transient signals.

[0120] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc.

[0121] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the kernel (1840) of the computer system (1800).

[0122] The kernel (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1843), a hardware accelerator (1844) for certain tasks, a graphics adapter (1850), etc. These devices, as well as a read-only memory (ROM) (1845), a random access memory (1846), an internal mass storage such as an internal non-user-accessible hard disk drive, SSD, etc. (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus (1848) of the kernel or attached to the system bus (1848) of the kernel via a peripheral bus (1849). In one example, a touch screen (1810) may be connected to the graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.

[0123] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be those specially designed and constructed for the purposes of this application, or the medium and the computer code can be of the type well-known and available to those having skill in the field of computer software.

[0124] Although this application has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of this application. Accordingly, it is to be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of this application and thus fall within the spirit and scope of this application.

Claims

1. A method for decoding a current block in an encoded video bitstream, the method comprising: receiving, by a device, the encoded video bitstream, the device including a memory storing instructions and a processor communicating with the memory; determining, by the device, based on the encoded video bitstream, a transform heap for the current block, the transform heap being one of a plurality of transform heaps in a transform dictionary, the transform dictionary including a plurality of transform heaps, and the transform heap including a plurality of transform sets; determining, by the device, a transform set in the transform heap; determining, by the device, a transform kernel in the transform set; performing, by the device, an inverse transform on the current block based on the transform kernel to obtain a residual block of the current block; and reconstructing, by the device, the current block based on the residual block.

2. The method according to claim 1, wherein The determining the transform heap for the current block includes: extracting, by the device, based on the encoded video bitstream, syntax for indicating the transform heap; and selecting, by the device, the transform heap from the plurality of transform heaps based on the syntax.

3. The method according to claim 1, wherein The determining the transform heap for the current block includes: the device deriving an index for selecting the transform heap from the plurality of transform heaps based on at least one of: one or more quantization parameters, a spatial resolution of the encoded video bitstream, a temporal resolution of the encoded video bitstream, a bit depth of the encoded video bitstream, or reconstructed samples from a previous picture group, picture, sub - picture, slice, or tile.

4. The method according to claim 1, wherein, The determining the transform heap for the current block includes: extracting, by the device, based on the encoded video bitstream, a flag for indicating whether the transform heap is changed; and in response to the flag indicating that the transform heap is not changed, the device determining the previous transform heap of a previous block as the transform heap of the current block.

5. The method according to claim 1, wherein The determining the transform heap for the current block includes: extracting, by the device, based on the encoded video bitstream, an index group for indicating a set of transform sets in a set of candidate transform sets; and constructing, by the device, the transform heap to include the set of transform sets based on the index group.

6. The method according to claim 1, wherein The determining the transform heap for the current block includes: for the transform sets in the transform heap, extracting, by the device, based on the encoded video bitstream, a flag indicating whether the transform set is updated; and in response to the flag indicating that the transform set is updated: extracting, by the device, based on the encoded video bitstream, an index of the updated transform set, the updated transform set being from a set of candidate transform sets; and updating, by the device, the transform heap by replacing the transform set with the updated transform set.

7. The method according to claim 1, wherein: each of the plurality of transform heaps includes a different number of transform sets.

8. The method according to claim 1, wherein The determining the transform heap for the current block includes: extracting, by the device, based on the encoded video bitstream, syntax indicating the number of transform sets in the transform heap.

9. The method according to claim 1, wherein: the transform set is one of the transform sets in the transform heap.

10. The method according to claim 1, wherein: Two or more different transform stacks in the transform dictionary include one or more identical transform sets.

11. The method according to claim 1, wherein: The separable transform kernels in two or more different transform stacks in the transform dictionary are the same.

12. The method according to claim 1, wherein: The separable non-Karhunen-Loève transform (non-KLT) kernels in two or more different transform stacks in the transform dictionary are the same.

13. The method according to claim 1, wherein: The different transform kernels in the transform dictionary include one or more identical bases.

14. An apparatus for decoding a current block of a current frame in an encoded video bitstream, the apparatus comprising: A memory storing instructions; And A processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1 to 13.

15. A non-transitory computer-readable storage medium stores instructions, wherein, When the instructions are executed by the processor, the instructions are configured to cause the processor to perform the method according to any one of claims 1 to 13.