A method and apparatus for identifying sub-block transform information

By identifying the flag bits that enable sub-block transformation (SBT) and maximum conversion block (TB) sizes in the sequence parameter set (SPS) of the video sequence, optimizing the video encoding unit (CU) size, the problem of insufficient coding performance in the prior art is solved, and more efficient video encoding is achieved.

CN114402547BActive Publication Date: 2025-06-13HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080064235.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-07-24
Publication Date
2025-06-13
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

While improving encoding efficiency, existing video encoding technologies are difficult to effectively use sub-block transformation (SBT) to optimize the encoding unit (CU) size of video sequences, resulting in insufficient coding performance.

Method used

The maximum conversion block (TB) size that allows SBT is indicated by identifying the first flag in the sequence parameter set (SPS) of the video sequence, and the second flag is indicated. In response to SBT activation, the maximum encoding unit (CU) size of the SBT is allowed to be determined directly based on the maximum conversion block (TB) size.

Benefits of technology

By enabling SBT and optimizing CU size, the efficiency of video encoding is significantly improved, the storage and transmission bandwidth requirements are reduced, while maintaining the same subjective quality as HEVC/H.265.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402547B_ABST
    Figure CN114402547B_ABST
Patent Text Reader

Abstract

The present invention provides an apparatus and method for identifying sub-block transform (SBT) information. The SBT information is used for encoding video data. According to some disclosed embodiments, an exemplary method includes: identifying, in a sequence parameter set (SPS) of a video sequence, a first flag bit to indicate whether sub-block transform (SBT) is enabled; and identifying a second flag bit to indicate a maximum transform block (TB) size allowed for SBT. In response to the first flag bit indicating that the SBT is enabled, a maximum coding unit (CU) size allowed for the SBT is directly determined according to the maximum TB size.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This disclosure claims priority to U.S. Provisional Application No. 62 / 900,395, filed on September 13, 2019, which is incorporated herein by reference in its entirety. Background Art

[0003] Video is a set of still pictures (or "images") that capture visual information. To reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission and decompressed for display. The compression process is generally referred to as encoding, and the decompression process is generally referred to as decoding. Currently, there are various video coding formats that employ standardized video coding techniques, and the most common ones are video coding formats based on prediction, transformation, quantization, entropy coding, and in - loop filtering. The video coding standards, such as the HEVC / H.265 (High Efficiency Video Coding) standard, the VVC / H.266 (Versatile Video Coding) standard, and the AVS (AVS Standards) standard, specify specific video coding formats by standardization organizations. With the application of more and more advanced video coding techniques in the video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Invention

[0004] Embodiments of the present invention provide a method and an apparatus for video processing. In an exemplary embodiment, the video processing method includes: identifying a first flag bit in a sequence parameter set (SPS) of a video sequence to indicate whether to enable sub - block transform (SBT); and identifying a second flag bit to indicate the maximum transform block (TB) size allowed for the SBT. In response to the first flag bit indicating that the SBT is enabled, the maximum coding unit (CU) size allowed for the SBT is directly determined according to the maximum transform block TB size.

[0005] In another exemplary embodiment, a video processing apparatus includes: at least one memory for storing an instruction set and at least one processor. The at least one processor executes the instruction set to cause the device to perform: identifying a first flag bit in a sequence parameter set (SPS) of a video sequence to indicate whether to enable sub - block transform (SBT); and identifying a second flag bit to indicate the maximum transform block (TB) size allowed for the SBT. In response to the first flag bit indicating that the SBT is enabled, the maximum coding unit (CU) size allowed for the SBT is directly determined according to the maximum transform block TB size.

[0006] In another exemplary embodiment, a non - volatile computer - readable storage medium stores a set of instruction sets. The instruction sets can be executed by at least one processor to enable a computer to execute a video processing method. The method includes: identifying a first flag bit in a sequence parameter set (SPS) of a video sequence to indicate whether sub - block transform (SBT) is enabled; and identifying a second flag bit to indicate the maximum transform block (TB) size allowed for the SBT. In response to the first flag bit indicating that the SBT is enabled, the maximum coding unit (CU) size allowed for the SBT is directly determined according to the maximum transform block (TB) size. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and aspects of the present disclosure are described in the following detailed implementation methods and the accompanying drawings. The various features shown in the drawings are not drawn to scale.

[0008] Figure 1 According to some embodiments of the present disclosure, a schematic diagram showing the structure of an exemplary video sequence is presented.

[0009] Figure 2 According to some embodiments of the present disclosure, a schematic diagram of an exemplary encoder in a certain hybrid video coding system is presented.

[0010] Figure 3 According to some embodiments of the present disclosure, a schematic diagram of an exemplary decoder in a certain hybrid video coding system is presented.

[0011] Figure 4 According to some embodiments of the present disclosure, a block diagram of an exemplary apparatus for encoding or decoding video is presented.

[0012] Figure 5 According to some embodiments of the present invention, an example of sub - block transform (SBT) types and SBT positions of an inter - prediction coding unit (CU) is presented.

[0013] Figure 6 And Figure 6-1 According to some embodiments of the present disclosure, an exemplary Table 1 showing a part of the SPS syntax table is presented.

[0014] Figure 7 According to some embodiments of the present disclosure, a flowchart of an exemplary video processing method is presented.

[0015] Figure 8 And Figure 8-1 According to some embodiments of the present disclosure, an exemplary Table 2 showing a part of the SPS syntax table is presented.

[0016] Figure 9 According to some embodiments of the present disclosure, an exemplary Table 3 showing a part of the CU syntax table is presented.

[0017] Figure 10 According to some embodiments of the present disclosure, a flowchart of another exemplary video processing method is shown. Detailed implementation manners

[0018] Preferred embodiments will now be described in detail, and the examples are provided with illustrations in the accompanying drawings. Unless otherwise specified, the following description refers to the accompanying drawings, in which the same numerals in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with aspects related to the present invention described in the appended claims. Specific aspects of the present disclosure will be described in more detail below. If there is a conflict with the terms and / or definitions included in the reference citations, the terms and definitions provided herein shall prevail.

[0019] The Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) are currently developing the Versatile Video Coding (VVC / H.266) standard. The goal of the VVC standard is to double the compression efficiency compared to its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 and use half the bandwidth.

[0020] In order to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth, JVET has been leveraging the Joint Exploration Model (JEM) reference software to explore technologies beyond HEVC. Since the coding technologies are incorporated into JEM, the coding performance of JEM is much higher than that of HEVC.

[0021] The VVC standard has only recently been developed and continues to incorporate more coding technologies that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0022] Video is a set of static pictures (or "frames") arranged in chronological order for storing visual information. Video capture devices (such as cameras) can be used to capture and store these pictures in a time sequence, and video playback devices (such as TVs, computers, smartphones, tablets, video players, or any end-user terminal with a display function) can be used to display such pictures in the time sequence. Similarly, in some application fields, video capture devices can transmit the captured video to a video playback device (for example, a computer with a monitor) in real time, such as for surveillance, conferencing, or live streaming.

[0023] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally referred to as an "encoder", and the module for decompression is generally referred to as a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented in any form, such as various suitable hardware, software, or a combination thereof. For example, the hardware implementation of the encoder and decoder can include circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination of the above. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using a variety of algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and then recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder".

[0024] The video encoding process can identify and retain useful information that can be used to reconstruct an image and ignore information that is not important for the reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy as a trade-off to reduce the required storage space and transmission bandwidth.

[0025] The useful information of a coded image (referred to as the "current frame") includes changes relative to a reference frame (e.g., a previously coded and reconstructed frame). These changes can include changes in the position of pixels, photometric changes, or color changes, with the most concerned being the position changes. The position changes of a group of pixels representing an object can reflect the movement of the object between the reference frame and the current frame.

[0026] An encoded frame that does not reference another frame (i.e., it is its own reference frame) is called an "I-frame". An encoded frame that uses a previous frame as a reference frame is called a "P-frame". An encoded frame that uses both a previous frame and a future frame as reference frames (i.e., the reference is "bidirectional") is called a "B-frame".

[0027] Figure 1According to some embodiments of the present disclosure, the structure of an exemplary video sequence 100 is shown. The video sequence 100 can be a live video or a video that has been captured and archived. Video 100 can be a real-life video, a computer-generated video (such as a computer game video), or a combination of both (such as a real-life video with augmented reality effects). The video sequence 100 can be input from a video capture device (such as a camera), a video archive containing previously captured videos (such as a video file stored in a storage device), or from a video feed interface (such as a video broadcast transceiver) to receive video from a video content provider.

[0028] As Figure 1 shown, the video sequence 100 can include a series of frames arranged along a time axis, including frames 102, 104, 106, 108. Frames 102-106 are consecutive, and there are many frames between frame 106 and frame 108. In Figure 1 , frame 102 is an I-frame, and its reference frame is frame 102 itself. Frame 104 is a P-frame, and its reference frame is frame 102, as indicated by the arrow. Frame 106 is a B-frame, and the reference frames are frame 104 and frame 108, as shown by the arrows in the figure. In some embodiments, the reference frame of a frame (e.g., frame 104) is not directly in front of or behind the frame. For example, the reference frame of frame 104 can be a frame before frame 102. It should be noted that the reference frames of frames 102-106 are only examples, and the present disclosure is not limited to the embodiments of the reference frames shown Figure 1 as examples.

[0029] Generally, due to the computational complexity of such tasks, video codecs do not encode or decode an entire frame at once. Instead, they can divide a frame into basic segments and encode or decode the frame segment by segment. Such a basic segment is referred to as a basic processing unit ("BPUs") in this disclosure. For example, Figure 1 the structure 110 in shows an example structure of a certain frame of the video sequence 100 (e.g., any of frames 102-108). In structure 110, a frame is divided into 4×4 basic processing units, and their boundaries are indicated by dashed lines. In some embodiments, the basic processing unit is referred to as a "macroblock" in some video coding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in a certain frame can have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or can be pixels of any shape and size. The size and shape of the basic processing unit can be selected for a frame based on the balance between coding efficiency and the level of detail to be retained in the basic processing unit.

[0030] The basic processing unit can be a logical unit, and the logical unit can include a set of different types of video data, which are stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image can contain a luminance component (Y) representing achromatic luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information and related syntax elements, and the basic processing units of the luminance and chrominance components can have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the said luminance and chrominance components can be referred to as "coding tree blocks" (CTBs). Any operation performed on the basic processing unit can be repeatedly performed on its respective luminance and chrominance components.

[0031] Video coding has multiple operation stages, examples of which are Figure 2 and Figure 3 shown. For each stage, the size of the basic processing unit may still be too large to process, so it can be further divided into segments called "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunit can be called a "block" in some video coding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or a "coding unit" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit can have the same or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which can include a set of different types of video data (e.g., Y, Cb, Cr, and related syntax elements), and these data are stored in a computer memory (e.g., a video frame buffer). Any operation performed on the basic processing subunit can be repeatedly performed on its respective luminance and chrominance components. It should be noted that this division can be further performed according to processing needs. It should also be noted that different stages can use different schemes to divide the said basic processing unit.

[0032] For example, in the mode decision stage (as an example shown in Figure 2 ), the encoder should decide what prediction mode (e.g., intra prediction or inter prediction) to use for a basic processing unit; but the basic processing unit may be too large to make such a decision, and the encoder can split the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC), and determine a prediction type for each individual basic processing subunit.

[0033] For another example, in the prediction stage ( Figure 2(Shown as one example), the encoder may perform prediction operations at the level of basic processing units (such as CUs). However, in some cases, the basic processing units may still be too large to process. The encoder may further divide the basic processing units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and the prediction operations may be performed at this level.

[0034] As another example, in the transform stage ( Figure 2 (Shown as one example), the encoder may perform transform operations on the remaining basic processing units (such as CUs). However, in some cases, the basic processing units may still be too large to process. The encoder may further divide the basic processing units into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transform operations may be performed at this level. It should be noted that the partitioning schemes of the same basic processing unit in the prediction stage and the transform stage may be different. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0035] In Figure 1 the structure 110, the base processing unit 112 is further divided into 3×3 basic processing units, and the boundaries of the units are shown as dotted lines in the figure. In different schemes, different basic processing units of the same frame may be divided into different basic processing units.

[0036] In some embodiments, to provide the ability of parallel processing and error recovery for video encoding and decoding, a frame may be divided into several regions for processing, such that for one region of the frame, the encoding or decoding process may not depend on any information from other regions of the frame. In other words, each region in the frame can be processed independently. In this way, the codec can process different regions of the image frame in parallel, thereby improving the encoding efficiency. In addition, when the data of one region is damaged or lost during network transmission during the processing, the codec can correctly encode or decode other regions of the same image frame without depending on the damaged or lost data, thereby providing fault tolerance. In some video coding standards, a frame can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different frames of the video sequence 100 may have different partitioning schemes for dividing the frame into regions.

[0037] For example, in Figure 1In [the figure], structure 110 is divided into three regions 114, 116, and 118, and the boundaries of the regions are represented by solid lines inside structure 110. Region 114 includes four basic processing units. Each of region 116 and region 118 includes six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and regions of structure 110 in [the figure] are only examples, and the present disclosure does not limit its implementation manners.

[0038] According to some embodiments of the present disclosure, Figure 2 a schematic diagram of an exemplary encoder 200 in a hybrid video coding system is shown. Video encoder 200 can perform intra-coding or inter-coding of blocks within a video frame, including video blocks, or partitions or sub-partitions of video blocks. Intra-coding can rely on spatial prediction to reduce or eliminate video spatial redundancy within a given video frame. Inter-coding can rely on temporal prediction to reduce or remove temporal redundancy in adjacent frames in a video sequence. Intra modes can refer to some space-based compression modes. Inter modes (such as single prediction or dual prediction) can refer to some time-based compression modes.

[0039] Referring to Figure 2 , the input video signal 202 can be processed block by block. For example, a video block unit can be a 16×16 pixel block (e.g., a macroblock (MB)). The size of the video block unit may vary depending on the coding technology used, as well as the required accuracy and efficiency. In HEVC, extended block sizes (such as coding tree units (CTUs)) can be used to compress video signals with a resolution of 1080p or higher. In HEVC, a CTU can include up to 64×64 luminance samples corresponding to chrominance samples, as well as related syntax elements. In VVC, the size of the CTU can be further increased to 128x128 luminance samples corresponding to chrominance samples, as well as related syntax elements. A certain CTU can be further divided into coding units (CUs), for example, using a quadtree, binary tree, or ternary tree. A CU can be further divided into prediction units (PUs), and different prediction methods can be applied to these units. Each input video block can be processed using a spatial prediction unit 260 or a temporal prediction unit 262.

[0040] The spatial prediction unit 260 performs spatial prediction (such as intra prediction) on the current block / CU using information on the same frame / slice that contains the current block. Spatial prediction can utilize pixels in adjacent blocks that have already been encoded in the same video frame / slice to predict the current video block. Spatial prediction can reduce the spatial redundancy inherent in the video signal.

[0041] The temporal prediction unit 262 performs temporal prediction (such as inter prediction) on the current block by using information of other frames / slices outside the frame / slice containing the current block. The temporal prediction of a video block can be identified by one or more motion vectors. In uni-directional temporal prediction, only one motion vector identifying one reference frame is used to generate the prediction of the current block. On the other hand, in bi-directional temporal prediction, two motion vectors, each identifying a reference frame, can be used to generate the prediction of the current block. The motion vector(s) can indicate the amount of motion and the direction of motion between the current block and one or more related blocks in the reference frame. If multiple reference frames are supported, one or more reference frame indices can be sent for a video block. The one or more reference indices are used to identify from which reference frame in the reference picture buffer or decoded picture buffer (DPB) 264 the temporal prediction signal can be generated.

[0042] The mode decision and encoder control unit 280 in the encoder can select the prediction mode, e.g., based on rate-distortion optimization. According to the determined prediction mode, a predicted block can be obtained. The predicted block can be subtracted from the current video block at the adder 216. The prediction residual can be transformed by the transform unit 204 and quantized by the quantization unit 206. The quantized residual coefficients can be dequantized by the dequantization unit 210 and inverse-transformed by the inverse-transform unit 212 to form a reconstructed residual. The reconstructed residual is added to the predicted block at the adder 226 to form a reconstructed video block. The reconstructed video block before loop filtering can be used as a reference sample for intra prediction.

[0043] The reconstructed video block can be loop-filtered by the loop filter 266. For example, loop filtering such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filter (ALF) can be employed. The loop-filtered reconstructed block can be stored in the reference picture buffer 264 and can provide an inter prediction reference sample for encoding other video blocks. To form the output video bitstream 220, before compressing and packing the data into the bitstream 220, the coding mode (such as inter (frame) or intra (frame)), prediction mode information, motion information, quantized residual coefficients, etc. can be sent to the entropy coding unit 208 to further reduce the bit rate.

[0044] According to some embodiments of the present disclosure, Figure 3 shows a schematic diagram of an exemplary decoder 300 in a hybrid video coding system. Referring to Figure 3 , the video bitstream 302 can be unpacked or entropy decoded at the entropy decoding unit 308. The coding mode information can be used to determine whether to select the spatial prediction unit 360 or the temporal prediction unit 362. The prediction mode information can be sent to the corresponding prediction unit to generate a predicted block. For example, the temporal prediction unit 362 can apply motion compensation prediction to form the temporal prediction block.

[0045] Send the residual coefficients to the inverse quantization unit 310 and the inverse transform unit 312 to obtain the reconstructed residual. Add the prediction block and the reconstructed residual at 326 to form a reconstructed block before loop filtering. Then, the reconstructed block can be loop filtered at the loop filter 366. For example, loop filtering such as deblocking filtering, SAO, ALF, etc. can be applied. The reconstructed block after loop filtering can be stored in the reference picture buffer 364. The reconstructed data in the reference picture buffer 364 can be used to obtain the decoded video 320, or to predict subsequent video blocks. The decoded video 320 can be displayed on a display device, such as a TV, PC, smartphone, or tablet, for viewing by the end user.

[0046] According to some embodiments of the present invention, Figure 4 A block diagram showing an exemplary apparatus 400 for encoding or decoding video is shown. As Figure 4 shown, the apparatus 400 may include a processor 402. When the processor 402 executes the instruction set described herein, the apparatus 400 can become a dedicated machine for video encoding or decoding. The processor 402 can be any type of circuitry capable of operating or processing information. For example, the processor 402 may include any number of any combination of central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs), a microcontroller unit (MCU), optical processors, programmable logic controllers, single-chip microcomputers, a microprocessor, digital signal processors, IP cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), system-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 may also be a group of processors grouped as a single logical component. For example, as Figure 4 shown, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0047] The apparatus 400 may also include a memory 404 configured to store data (e.g., a set of instruction sets, computer code, intermediate data, or the like). For example, as Figure 4 shown, the stored data may include program instructions (e.g., for performing Figure 2 or Figure 3The program instructions at each stage) and the data to be processed. The processor 402 can access the program instructions and data to be processed (e.g., via the bus 410), and execute the program instructions to perform operations or controls on the data for processing. The memory 404 can include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 can include any combination of any number of random access memories (RAMs), a read-only memory (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, or other similar components. The memory 404 can also be a set of memories combined into a single logical component ( Figure 4 not shown).

[0048] The bus 410 can be a communication device for transferring data between the internal components of the device 400, such as an internal bus (e.g., a central processor - memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or a similar device.

[0049] For ease of explanation without ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this disclosure. The data processing circuits can be implemented entirely in hardware, or in a combination of software, hardware, or firmware. In addition, the data processing circuits can be an independent module, or can be incorporated in whole or in part into other elements of the device 400.

[0050] The device 400 can also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, the network interface 406 can include a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network card, a wireless network card, a Bluetooth network card, an infrared network card, a near field communication (NFC) adapter, a cellular network chip, or the like.

[0051] In some embodiments, optionally, the device 400 can also include a peripheral interface 408 to provide connections to one or more peripheral devices. As Figure 4 shown, the peripheral devices can include, but are not limited to, cursor control devices (such as a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface connected to a video file), or other similar devices.

[0052] It should be noted that the video codec can be implemented as any combination of any software or hardware modules in the device 400. For example, Figure 2 the encoder 200 ofFigure 3 Some or all stages of the decoder 300 of can be implemented as one or more software modules of the apparatus 400, e.g., program instructions loadable into the memory 404. In another example, Figure 2 Some or all stages of the encoder 200 of or Figure 3 Some or all stages of the decoder 300 of can be implemented as one or more hardware modules of the apparatus 400, e.g., dedicated data processing circuits (e.g., FPGA, ASIC, NPU or similar circuits).

[0053] In the quantization and dequantization functional blocks (such as Figure 2 the quantization unit 206 and the dequantization unit 210 of and Figure 3 the dequantization unit 310 of ), the quantization parameter (QP) is used to determine the quantization amount (and dequantization amount) applied to the prediction residual. The initial QP value for encoding a frame or slice can be identified at a high level, e.g., using the syntax element t_qp_minus26 in the picture parameter set (PPS), and using the syntax element slice_qp_delta in the slice header. In addition, the incremental QP value sent at the quantization group granularity can be used to adjust the QP value for each CU at this level.

[0054] In WC, sub-block transform (SBT) is used for inter prediction coding units (CUs). In this transform mode, only a sub-part of the residual block is encoded and provided to the encoding unit. When the inter prediction unit with the syntax element cu_cbf is equal to 1, the syntax element cu_sbt_flag can be marked to indicate whether to encode the entire residual block or a sub-part of the residual block. In the former case, the MTS (inter multiple transform selected) information is further parsed to determine the transform type of the CU. In the latter case, the partial residual block is encoded using the inferred adaptive transform, and the other parts of the residual block are set to zero.

[0055] When SBT is used for a certain inter prediction CU, the SBT type and SBT position information are identified in the bitstream. There are two SBT types and two SBT positions, such as Figure 5As shown. For SBT-V (or SBT-H), the width (or height) of the transform unit (TU) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), forming a 2:2 split or a 1:3 / 3:1 split. The 2:2 split is similar to a binary tree (BT) split, while the 1:3 / 3:1 split is similar to an asymmetric binary tree (ABT) split. In the ABT split, only its small region contains non-zero residuals. If a coding unit has 8 luma samples in a certain dimension, then a 1:3 / 3:1 split along that dimension is not allowed. A coding unit can have at most 8 SBT modes.

[0056] The sequence parameter set (SPS) level syntax can use the syntax element sps_sbt_enabled_flag to specify enabling or disabling SBT. When the syntax element sps_sbt_enabled_flag is equal to 0, it indicates that SBT for inter-predicted coding units is disabled in the entire video sequence that references this SPS. When the syntax element sps_sbt_enabled_flag is equal to 1, it indicates that SBT for inter-predicted coding units is enabled in the entire video sequence that references this SPS.

[0057] In addition, when sps_sbt_enabled_flag is equal to 1, another SPS syntax element sps_sbt_max_size_64_flag can be used to specify the maximum CU width and height allowed for SBT. When the syntax element sps_sbt_max_size_64_flag is equal to 0, it indicates that the maximum CU width and height allowed for SBT are 32 luma samples. When the syntax element sps_sbt_max_size_64_flag is equal to 1, it indicates that the maximum CU width and height allowed for SBT are 64 luma samples. The variable MaxSbtSize, which can specify the maximum CU size allowed for SBT, is calculated according to the following formula 1:

[0058] MaxSbtSize = Min(MaxTbSizeY, sps_sbt_max_size_64_flag? 64 : 32) (Formula 1)

[0059] where MaxTbSizeY is the maximum transform block (TB) size allowed, which can be derived from another SPS-level syntax element sps_max_luma_transform_size_64_flag according to the following formula 2:

[0060] MaxTbSizeY = sps_max_luma_transform_size_64_flag? 64 : 32 (Formula 2)

[0061] As described above, the MaxSbtSize source depends on two syntax elements, sps_max_luma_transform_size_64_flag and sps_sbt_max_size_64_flag. If the value of the syntax element sps_max_luma_transform_size_64_flag is 0, then regardless of the value of the syntax element sps_sbt_max_size_64_flag, MaxSbtSize is always 32. Therefore, when the syntax element sps_max_luma_transform_size_64_flag is 0, there is no need to identify the syntax element sps_sbt_max_size_64_flag. This syntax redundancy in VVC unnecessarily increases the signaling overhead.

[0062] To improve video coding efficiency, according to some disclosed embodiments, the syntax element sps_sbt_max_size_64_flag is identified only when both the syntax elements sps_max_luma_transform_size_64_flag and sps_sbt_enabled_flag are 1. Figure 6 Exemplary Table 1 according to some embodiments of the present disclosure is shown. Table 1 shows an exemplary SPS syntax table of some embodiments. As shown in Table 1 (the emphasized part is in italics), the syntax element sps_sbt_max_size_64_flag is identified only when both the syntax elements sps_max_luma_transform_size_64_flag and sps_sbt_enabled_flag are 1. If the syntax element sps_max_luma_transform_size_64_flag is 0, then the syntax element sps_sbt_max_size_64_flag can be inferred as 0, which means that the maximum width and height of the CU allowing SBT are 32 (in terms of luma samples).

[0063] Figure 7 A flowchart of an exemplary video processing method 700 according to some embodiments of the present disclosure is shown. In some embodiments, method 700 may be performed by an encoder (e.g., Figure 2 encoder 200), a decoder (e.g., Figure 3 decoder 300) or a device of one or more software or hardware components (e.g., Figure 4 device 400 in Figure 4The processor 402) may execute method 700. In some embodiments, method 700 may be implemented by a computer program product included in a computer-readable medium, the product including computer-executable instructions, such as program code executed by a computer (e.g., Figure 4 the apparatus 400 in

[0064] In step 702, method 700 may include determining whether sub-block transform (SBT) is enabled in a sequence parameter set (SPS) of a certain video sequence. In some embodiments, a flag bit (e.g., a syntax element sps_sbt_enabled_flag as shown in Figure 6 Table 1) may be identified in the SPS indicating whether SBT is enabled. For example, the syntax element sps_sbt_enabled_flag being equal to 0 may specify that SBT of inter prediction coding units is disabled for the entire video sequence referencing the SPS. And the syntax element sps_sbt_enabled_flag = 1 may specify that SBT of inter prediction coding units is enabled for the entire video sequence referencing the SPS.

[0065] In step 704, method 700 may include determining the value of a first flag bit in the SPS that indicates the maximum transform block (TB) size allowed for SBT. The first flag bit may be set to a first value or a second value. For example, the first value is 1 and the second value is 0. The maximum TB size may be 32, 64, or the like. In some embodiments, method 700 may further include setting the first flag bit to the first value corresponding to the maximum TB size being 64, and setting the value of the first flag bit to the second value corresponding to the maximum TB size being 32. In some embodiments, the first flag bit may be Figure 6 the syntax element sps_max_luma_transform_size_64_flag in Table 1.

[0066] In step 706, method 700 may include identifying a second flag bit indicating the maximum coding unit (CU) size allowed for SBT in response to SBT being enabled and the value of the first flag bit being equal to the first value. In response to SBT being disabled or the value of the first flag bit being equal to the second value, the second flag bit is not identified. For example, the second flag bit may be a syntax element sps_sbt_max_size_64_flag as shown in Figure 6 Table 1. The syntax element sps_sbt_max_size_64_flag is only identified when both the syntax elements sps_max_luma_transform_size_64_flag and sps_sbt_enabled_flag are 1.

[0067] In some embodiments, method 700 may further include identifying a third flag bit in the SPS (e.g., the syntax element sps_sbt_enabled_flag as shown in Figure 6 Table 1) to indicate whether SBT is enabled, and identifying the first flag bit in the SPS (e.g., Figure 6 the syntax element sps_max_luma_transform_size_64_flag in Table 1 in

[0068] In certain embodiments, the maximum CU size may be 32 or 64. The maximum CU width or height allowing SBT may be determined according to the smaller one of the maximum TB size and the maximum CU size (e.g., according to Formula 1).

[0069] In some disclosed embodiments, the syntax element sps_sbt_max_size_64_flag is not identified at all. In this case, the maximum width and height of the CU allowing SBT directly depend on the syntax element sps_max_luma_transform_size_64_flag. If the syntax element sps_max_luma_transform_size_64_flag is equal to 0, the maximum CU width and height allowing SBT are 32 luma samples. If the syntax element sps_max_luma_transform_size_64_flag is equal to 1, the maximum CU width and height allowing SBT are 64 luma samples. In other words, MaxSbtSize is set to be equal to MaxTbSizeY. Figure 8 According to some embodiments of the present disclosure, an exemplary Table 2 is shown. Table 2 shows an exemplary SPS syntax for implementing these embodiments. As shown in Table 2, the syntax element sps_sbt_max_size_64_flag is not identified and is removed from the syntax. Figure 9 According to some embodiments of the present disclosure, an exemplary Table 3 is shown. Table 3 (emphasized in italics) shows an exemplary coding unit (CU) syntax table that directly uses MaxTbSizeY to set the maximum width and height of the CU.

[0070] MaxTbSizeY is calculated by Equation 3 as follows:

[0071] MaxTbSizeY = sps_max_luma_transform_size_64_flag? 64 : 32 (Formula 3)

[0072] Figure 10FIG. 0 shows a flowchart of another exemplary video processing method 1000 according to some embodiments of the present disclosure. In some embodiments, method 1000 may be performed by an encoder (e.g., Figure 2 encoder 200), a decoder (e.g., Figure 3 decoder 300), or one or more software or hardware components of a device (e.g., Figure 4 device 400). For example, a processor (e.g., Figure 4 processor 402) may perform method 1000. In certain embodiments, method 1000 may be implemented by a computer program product included in a computer-readable medium, the product including computer-executable instructions, such as program code executed by a computer (e.g., Figure 4 device 400) in

[0073] In step 1002, method 1000 includes identifying a first flag bit in the sequence parameter set (SPS) of a video sequence, the flag bit indicating whether sub-block transform (SBT) is enabled. In some embodiments, the first flag bit may be the syntax element sps_sbt_enabled_flag, as shown in Figure 8 Table 2 of

[0074] For example, the syntax element sps_sbt_enabled_flag being equal to 0 may specify that SBT for inter prediction coding units is disabled for the entire video sequence that references the SPS. Also, the syntax element sps_sbt_enabled_flag being equal to 1 may specify that SBT for inter prediction coding units is enabled for the entire video sequence that references the SPS. Figure 8 In step 1004, method 1000 may include identifying a second flag bit to indicate the maximum transform block (TB) size that allows SBT. The second flag bit may be set to a first value or a second value. For example, the first value is 1 and the second value is 0. The maximum TB size may be 32, 64, or a similar value. In some embodiments, method 1000 may further include setting the value of the second flag bit to 0 when the maximum TB size corresponds to 32; setting the value of the second flag bit to 1 when the maximum TB size corresponds to 64. In some embodiments, the first flag bit may be

[0075] the syntax element sps_max_luma_transform_size_64_flag in Table 2 of

[0076] In some embodiments, a non-volatile computer-readable storage medium including an instruction set is also provided, and the instruction set can be executed by a device (such as the encoder and decoder) for performing the above method. Common forms of non-volatile media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any punched physical medium pattern, RAM, PROM, and EPROM, flash EPROM or other flash memory, NVRAM, cache, registers, any other storage chip or tape, and network versions of the like. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0077] The embodiments can be further described using the following clauses:

[0078] 1. A video processing method, comprising:

[0079] Determining whether sub-block transform (SBT) is enabled in a sequence parameter set (SPS) of a video sequence;

[0080] Determining the value of a first flag bit in the SPS, the flag bit indicating the maximum transform block (TB) size allowed for SBT; and

[0081] In response to SBT being enabled and the value of the first flag bit being equal to a first value, identifying a second flag bit indicating the maximum coding unit (CU) size allowed for SBT.

[0082] 2. The method according to clause 1, wherein in response to SBT not being enabled, or the value of the first flag bit being equal to a second value, the second flag bit is not identified.

[0083] 3. The method according to clause 1 or clause 2, further comprising:

[0084] Identifying a third flag bit in the SPS to indicate whether SBT is enabled; and

[0085] Identifying the first flag bit in the SPS.

[0086] 4. The method according to clause 2, wherein the first value is 1 and the second value is 0.

[0087] 5. The method according to any one of clauses 1 - 4, wherein the maximum TB size is 32 or 64.

[0088] 6. The method according to clause 5, further comprising: in response to the maximum TB size being 64, setting the value of the first flag bit to the first value.

[0089] 7. The method according to item 5 further includes:

[0090] In response to the maximum TB size being 32, set the value of the first flag bit to a second value.

[0091] 8. According to any of the methods in items 1 - 7, wherein the maximum CU size allowing SBT is 32 or 64.

[0092] 9. According to any of the methods in items 1 - 8, wherein the maximum CU width allowing SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowing SBT.

[0093] 10. According to any of the methods in items 1 - 9, wherein the maximum CU height allowing SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowing SBT.

[0094] 11. A video processing device, including:

[0095] At least one memory for storing instruction sets; and

[0096] At least one processor that executes the instruction sets to cause the device to perform:

[0097] Determine whether sub - block transform (SBT) is enabled in the sequence parameter set (SPS) of a video sequence;

[0098] Determine the value of the first flag bit in the SPS, which indicates the maximum transform block (TB) size allowing SBT; and

[0099] In response to the SBT being enabled and the value of the first flag bit being equal to a first value, identify a second flag bit indicating the maximum coding unit (CU) size allowing SBT.

[0100] 12. The device according to item 11, wherein corresponding to the SBT not being enabled or the value of the first flag bit being equal to a second value, the second flag bit is not identified.

[0101] 13. The device according to item 11 or 12, wherein at least one processor further executes the instruction sets to cause the device to perform:

[0102] Identify the third flag bit in the SPS to indicate whether the SBT is enabled; and

[0103] Identify the first flag bit in the SPS.

[0104] 14. The device according to item 12, wherein the first value is 1 and the second value is 0.

[0105] 15. According to any one of Articles 11 - 14, where the maximum TB size is 32 or 64.

[0106] 16. The device according to Article 15, further comprising:

[0107] In response to the maximum TB size being 64, set the value of the first flag bit to a first value.

[0108] 17. The device according to Article 15, further comprising:

[0109] In response to the maximum TB size being 32, set the value of the first flag bit to a second value.

[0110] 18. The device according to any one of Articles 11 - 17, wherein the maximum CU size allowed for the SBT is 32 or 64.

[0111] 19. The device according to any one of Articles 11 - 18, wherein the maximum CU width allowed for the SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowed for the SBT.

[0112] 20. The device according to any one of Articles 11 - 19, wherein the maximum CU height allowed for the SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowed for the SBT.

[0113] 21. A non - volatile computer - readable storage medium storing a set of instructions, which are executed by at least one processor to cause a computer to perform a video processing method, including:

[0114] Determine whether sub - block transform (SBT) is enabled in the sequence parameter set (SPS) of a video sequence;

[0115] Determine the value of the first flag bit in the SPS, which indicates the maximum transform block (TB) size allowed for the SBT; and

[0116] In response to the SBT being enabled and the value of the first flag bit being equal to the first value, identify a second flag bit indicating the maximum coding unit (CU) size allowed for the SBT.

[0117] 22. The non - volatile computer - readable storage medium according to Article 21, wherein corresponding to the SBT not being enabled, or the value of the first flag bit being equal to the second value, the second flag bit is not identified.

[0118] 23. The non - volatile computer - readable storage medium according to Article 21 or 22, wherein the set of instructions executed by at least one processor causes the computer to further perform:

[0119] Identify the third flag bit in the SPS to indicate whether the SBT is enabled; and

[0120] Identify the first flag bit in the SPS.

[0121] 24. The non - volatile computer - readable storage medium according to Article 22, wherein the first value is 1 and the second value is 0.

[0122] 25. A non - volatile computer - readable storage medium according to any one of Articles 21 - 24, wherein the maximum TB size is 32 or 64.

[0123] 26. The non - volatile computer - readable storage medium according to Article 25, wherein the instruction set executed by at least one processor causes the computer to further perform:

[0124] In response to the maximum TB size being 64, set the value of the first flag bit to the first value.

[0125] 27. The non - volatile computer - readable storage medium according to Article 25, wherein the instruction set executed by at least one processor causes the computer to further perform:

[0126] In response to the maximum TB size being 32, set the value of the first flag bit to the second value.

[0127] 28. The non - volatile computer - readable storage medium according to any one of Articles 21 - 27, wherein the maximum CU size allowed for the SBT is 32 or 64.

[0128] 29. The non - volatile computer - readable storage medium according to any one of Articles 21 - 28, wherein the maximum CU width allowed for the SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowed for the SBT.

[0129] 30. The non - volatile computer - readable storage medium according to any one of Articles 21 - 29, wherein the maximum CU height allowed for the SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowed for the SBT.

[0130] 31. A video processing method, comprising:

[0131] In the sequence parameter set SPS of a video sequence, identify the first flag bit to indicate whether to enable sub - block transform SBT; and

[0132] Identify the second flag bit to indicate the maximum transform block TB size allowed for the SBT,

[0133] In response to the first flag bit indicating that the SBT is enabled, allowing the maximum coding unit (CU) size of the SBT to be directly determined according to the maximum transform block (TB) size.

[0134] 32. The method according to Article 31, wherein the maximum CU size allowed for the SBT is the maximum CU width or the maximum CU height.

[0135] 33. The method according to Article 31 or 32, wherein determining the maximum CU size allowed for the SBT is determined to be equal to the maximum TB size.

[0136] 34. The method according to any one of Articles 31 - 33, wherein the maximum TB size is 32 or 64.

[0137] 35. The method according to Article 34, further comprising:

[0138] In response to the maximum TB size being 32, setting the value of the second flag bit to 0.

[0139] 36. The method according to Article 34, further comprising:

[0140] In response to the maximum TB size being 64, setting the value of the second flag bit to 1.

[0141] 37. A video processing apparatus, comprising:

[0142] At least one memory for storing instructions; and

[0143] At least one processor for executing an instruction set to cause the apparatus to perform:

[0144] In a sequence parameter set (SPS) of a video sequence, identifying a first flag bit to indicate whether sub - block transform (SBT) is enabled; and

[0145] In response to the first flag bit indicating that the SBT is enabled, allowing the maximum coding unit (CU) size of the SBT to be directly determined according to the maximum transform block (TB) size.

[0146] 38. The apparatus according to Article 37, wherein the maximum CU size allowed for the SBT is the maximum CU width or the maximum CU height.

[0147] 39. The apparatus according to Article 37 or 38, wherein the maximum CU size allowed for the SBT is determined to be equal to the maximum TB size.

[0148] 40. The apparatus according to any one of Articles 37 - 39, wherein the maximum TB size is 32 or 64.

[0149] 41. The apparatus according to clause 40, wherein at least one processor further executes instructions to cause the apparatus to perform: in response to the maximum TB size being 32, set the value of the second flag bit to 0.

[0150] 42. The apparatus according to clause 40, wherein at least one processor further executes a set of instructions to cause the apparatus to perform: in response to the maximum TB size being 64, set the value of the second flag bit to 1.

[0151] 43. A non - volatile computer - readable storage medium storing a set of instructions executable by at least one processor to cause the computer to perform a video processing method, including:

[0152] In a sequence parameter set SPS of a video sequence, identify a first flag bit to indicate whether sub - block transform (SBT) is enabled; and

[0153] Identify a second flag bit to indicate the maximum transform block (TB) size allowed for SBT,

[0154] In response to the first flag bit indicating that the SBT is enabled, allow the maximum coding unit (CU) size of the SBT to be directly determined according to the maximum transform block (TB) size.

[0155] 44. The non - volatile computer - readable storage medium according to clause 43, wherein the maximum CU size allowed for SBT is the maximum CU width or the maximum CU height.

[0156] 45. The non - volatile computer - readable storage medium according to clause 43 or 44, wherein the maximum CU size allowed for SBT is determined to be equal to the maximum TB size.

[0157] 46. The non - volatile computer - readable storage medium according to any one of clauses 43 - 45, wherein the maximum TB size is 32 or 64.

[0158] 47. The non - volatile computer - readable storage medium according to clause 46, wherein the set of instructions executable by at least one processor causes the computer to further perform:

[0159] In response to the maximum TB size being 32, set the value of the second flag bit to 0.

[0160] 48. The non - volatile computer - readable storage medium according to clause 46, wherein the set of instructions executed by the at least one processor causes the computer to further perform:

[0161] In response to the maximum TB size being 64, set the value of the second flag bit to 1.

[0162] It should be noted that relational terms such as "first", "second", etc. in this text are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, words such as "include", "have", "contain", "comprise" and other similar forms have the same meaning, and any one or more items following any of the above words are open-ended, and none of the above nouns indicate that the one or more items have been enumerated exhaustively, or are limited to these enumerated one or more items.

[0163] As used herein, unless otherwise expressly stated, the term "or" includes all possible combinations, except where infeasible. For example, if it is stated that a database may include A or B, then unless otherwise specifically provided or infeasible, it may include database A, or B, or A and B. As a second example, if it is stated that a certain database may include A, B or C, then unless otherwise specifically provided or infeasible, the database may include database A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0164] It is worth noting that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When the software is executed by a processor, it can execute the methods disclosed above. The computing units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above multiple modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0165] In the above detailed description, the embodiments have been described with reference to many specific details, and these details may vary depending on the implementation. Certain adaptations and modifications can be made to the embodiments. For those skilled in the art, some other embodiments can be obviously obtained from the specific implementation manners disclosed in the present invention. This specification and examples are for illustrative purposes only, and the true scope and essence of the present invention are defined by the claims. The step order shown in the drawings is also for illustrative purposes only and does not mean to be limited to any specific steps or order. Therefore, those skilled in the art will realize that when implementing the same method, these steps can be executed in a different order.

[0166] In the drawings and detailed description of this application, exemplary embodiments are disclosed. However, many variations and modifications can be made to these embodiments. Correspondingly, although specific terms are used, these terms are only general and descriptive, and not for the purpose of limitation.

Claims

1. A video encoding method, comprising: identifying a first flag bit in a sequence parameter set (SPS) of a bitstream associated with a video sequence to indicate whether to enable sub-block transform (SBT); and identifying a second flag bit in the sequence parameter set SPS to indicate the maximum transform size in luminance samples, wherein, based on the value of the second flag bit, the maximum coding unit (CU) size allowing the SBT is determined, wherein the CU size allowing the SBT is determined as: if the value of the second flag bit is 1, the size of the maximum CU allowing the SBT is 64.

2. The method according to claim 1, wherein the maximum CU size allowing the SBT is the maximum CU width or the maximum CU height.

3. The method according to claim 1, wherein, determining that the maximum CU size allowing the SBT is determined to be equal to the maximum transform size.

4. The method according to claim 1, further comprising: in response to the value of the second flag bit being 1, setting the maximum transform size to 64.

5. The method according to claim 1, wherein, the second flag bit is sps_max_luma_transform_size_64_flag.

6. The method according to claim 1, wherein, the first flag bit is sps_sbt_enabled_flag.

7. A video decoding method, comprising: receiving a bitstream associated with a video sequence, the sequence parameter set (SPS) of the bitstream comprising: a first flag bit indicating whether to enable sub-block transform (SBT), and a second flag bit indicating the maximum transform size in luminance samples; and in response to the value of the second flag bit being 1, determining that the size of the maximum coding unit (CU) allowing the SBT is 64.

8. The method according to claim 7, wherein, the maximum CU size allowing the SBT is the maximum CU width or the maximum CU height.

9. The method according to claim 7, further comprising: wherein the maximum CU size allowing the SBT is determined to be equal to the maximum transform size.

10. The method according to claim 7, further comprising: in response to the value of the second flag bit being 1, determining the maximum transform size to be 64.

11. The method according to claim 7, wherein, the second flag bit is sps_max_luma_transform_size_64_flag.

12. The method according to claim 7, wherein, the first flag bit is sps_sbt_enabled_flag.

13. A non-transitory computer-readable storage medium that stores a bitstream associated with a video sequence, the non-transitory computer-readable storage medium being part of a computing device, the bitstream being generated by execution of an instruction set by one or more processors of the computing device, wherein, execution of the instruction set causes the computing device to perform: In a sequence parameter set (SPS) of a bitstream associated with a video sequence, identify a first flag bit to indicate whether sub-block transform (SBT) is enabled; And In the sequence parameter set SPS, identify a second flag bit to indicate the maximum transform size in luma samples, wherein, based on the second flag bit, the decoder determines the maximum coding unit (CU) size that allows the SBT to be: If the value of the second flag bit is 1, the size of the maximum CU that allows the SBT is 64.

14. The non-volatile computer-readable storage medium according to claim 13, wherein the maximum CU size that allows the SBT is the maximum CU width or the maximum CU height.

15. The non-volatile computer-readable storage medium according to claim 13, wherein, If the value of the second flag bit is 1, the second flag bit causes the decoder to determine the maximum transform size to be 64.

16. The non-volatile computer-readable storage medium according to claim 13, wherein, The second flag bit is sps_max_luma_transform_size_64_flag.

17. The non-volatile computer-readable storage medium according to claim 13, wherein, The first flag bit is sps_sbt_enabled_flag.