Method and device for identifying sub-block transformation information
By enabling sub-block transformation (SBT) in the sequence parameter set (SPS) of the video sequence and optimizing the maximum encoding unit (CU) size, the problem of insufficient utilization of encoding efficiency in the prior art is solved, and more efficient video encoding performance is achieved.
Patent Information
- Application Number
- CN202510423454.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-13
- Filing Date
- 2020-07-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-07-24
AI Technical Summary
While improving encoding efficiency, existing video encoding technologies are difficult to effectively use sub-block transformation (SBT) to optimize the encoding unit size of video sequences, resulting in insufficient coding performance.
In the sequence parameter set (SPS) of the video sequence, whether sub-block transformation (SBT) is enabled is indicated by identifying the first flag bit, and a second flag bit is indicated by identifying the maximum conversion block (TB) size that allows the SBT. In response to SBT activation, the maximum encoding unit (CU) size of the SBT is allowed to be determined directly based on the maximum conversion block (TB) size.
By enabling SBT and optimizing the encoding unit size, the efficiency and performance of video encoding are improved, and the storage and transmission bandwidth can be reduced while maintaining subjective quality.
Smart Images

Figure CN120111255A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to U.S. Provisional Application No. 62 / 900,395, filed on September 13, 2019, which is incorporated herein by reference in its entirety. Background Art
[0002] A video is a set of static pictures (or "images") that capture visual information. In order to reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission, and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are currently a variety of video coding formats that use standardized video coding technologies, the most common of which are video coding formats based on prediction, transformation, quantization, entropy coding, and in-loop filtering. The video coding standards, such as the HEVC / H.265 (High Efficiency video coding) standard, the VVC / H.266 (Versatile video coding) standard, and the AVS (AVS standards) standard, are formulated by standardization organizations for specific video coding formats. With the application of more and more advanced video coding technologies in the video standards, the coding efficiency of the new video coding standards is also getting higher and higher. Summary of the invention
[0003] Embodiments of the present invention provide a method and apparatus for video processing. In an exemplary embodiment, the video processing method includes: in a sequence parameter set (SPS) of a video sequence, identifying a first flag bit to indicate whether a sub-block transform (SBT) is enabled; and identifying a second flag bit to indicate a maximum transform block (TB) size allowed for the SBT. In response to the first flag bit indicating that the SBT is enabled, the maximum coding unit (CU) size allowed for the SBT is directly determined based on the maximum transform block (TB) size.
[0004] In another exemplary embodiment, a video processing device includes: at least one memory for storing an instruction set and at least one processor. The at least one processor executes the instruction set to cause the device to perform: in a sequence parameter set (SPS) of a video sequence, identify a first flag bit to indicate whether a sub-block transform (SBT) is enabled; and identify a second flag bit to indicate a maximum transform block (TB) size that allows the SBT. In response to the first flag bit indicating that the SBT is enabled, the maximum coding unit CU size that allows the SBT is directly determined based on the maximum transform block TB size.
[0005] In another exemplary embodiment, a non-volatile computer-readable storage medium stores a set of instruction sets. The instruction set can be executed by at least one processor to enable a computer to perform a video processing method. The method includes: in a sequence parameter set (SPS) of a video sequence, identifying a first flag bit to indicate whether a sub-block transform (SBT) is enabled; and, identifying a second flag bit to indicate a maximum transform block (TB) size allowed for the SBT. In response to the first flag bit indicating that the SBT is enabled, the maximum coding unit CU size allowed for the SBT is directly determined based on the maximum transform block TB size. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Embodiments and aspects of the present disclosure are described in the following detailed description and accompanying drawings. The various features shown in the drawings are not drawn to scale.
[0007] Figure 1 According to some embodiments of the present disclosure, a schematic diagram of the structure of an exemplary video sequence is shown.
[0008] Figure 2 According to some embodiments of the present disclosure, a schematic diagram of an exemplary encoder in a hybrid video coding system is shown.
[0009] Figure 3 According to some embodiments of the present disclosure, a schematic diagram of an exemplary decoder in a hybrid video coding system is shown.
[0010] Figure 4 According to some embodiments of the present disclosure, a block diagram of an exemplary apparatus for encoding or decoding a video is shown.
[0011] Figure 5 According to some embodiments of the present invention, exemplary sub-block transform (SBT) types and SBT positions of an inter-prediction coding unit (CU) are shown.
[0012] Figure 6 (Table 1: Exemplary SPS syntax, sps_sbt_max_size_64_flag flag bit is only identified when the sps_max_luma_transform_size_64_flag flag bit and the sps_sbt_enabled_flag flag bit are both 1 (first part)) and Figure 6(Continued) (Table 1: Exemplary SPS syntax, the sps_sbt_max_size_64_flag flag is identified only when the sps_max_luma_transform_size_64_flag flag and the sps_sbt_enabled_flag flag are both 1 (second part)) According to some embodiments of the present disclosure, an exemplary Table 1 is shown, showing a portion of the SPS syntax table.
[0013] Figure 7 According to some embodiments of the present disclosure, a flowchart of an exemplary video processing method is shown.
[0014] Figure 8 (Table 2: Exemplary SPS syntax table, sps_sbt_max_size_64_flag flag not used (first part)) and Figure 8 (Continued) (Table 2: Exemplary SPS syntax table, not using the sps_sbt_max_size_64_flag flag (second part)) According to some embodiments of the present disclosure, an exemplary Table 2 is shown, showing a portion of the SPS syntax table.
[0015] Fig. 9 (Table 3: Exemplary CU syntax, directly using MaxTbSizeY to set the maximum CU width and height) According to some embodiments of the present disclosure, an exemplary Table 3 is shown, showing a portion of the CU syntax table.
[0016] Fig.10 According to some embodiments of the present disclosure, a flowchart of another exemplary video processing method is shown. DETAILED DESCRIPTION
[0017] Preferred embodiments will now be described in detail, with examples provided as illustrated in the accompanying drawings. Unless otherwise noted, the following description refers to the accompanying drawings, in which the same numbers in different figures represent the same or similar elements. The implementations described in the following exemplary embodiment descriptions do not represent all implementations consistent with the present invention. Instead, they are merely examples of devices and methods consistent with aspects related to the present invention described in the appended claims. Specific aspects of the present disclosure will be described in more detail below. In the event of a conflict with terms and / or definitions contained in the reference citation, the terms and definitions provided herein shall prevail.
[0018] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) are currently developing the Versatile Video Coding (VVC / H.266) standard. The goal of the VVC standard is to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 and use half the bandwidth.
[0019] In order to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. As the coding techniques are incorporated into JEM, the coding performance of JEM is significantly higher than that of HEVC.
[0020] The VVC standard was developed only recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used for modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0021] A video is a set of static pictures (or "frames") arranged in time sequence, used to store visual information. A video acquisition device (such as a camera) can be used to capture and store these pictures in a time sequence, and a video playback device (such as a TV, computer, smartphone, tablet, video player, or any end-user terminal with a display function) can be used to display such pictures in the time sequence. Similarly, in some application areas, a video acquisition device can transmit the acquired video to a video playback device (such as a computer with a monitor) in real time, such as for monitoring, conferencing, or live broadcasting.
[0022] In order to reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software or dedicated hardware executed by a processor (e.g., a processor of a general-purpose computer). The module used for compression is usually called an "encoder" and the module used for decompression is usually called a "decoder". Encoders and decoders can be collectively referred to as "codecs". Encoders and decoders can be implemented in any form such as various suitable hardware, software, or a combination thereof. For example, the hardware implementation of encoders and decoders may include circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination of the above. The software implementation of encoders and decoders may include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using a variety of algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress video from a first coding standard and then recompress the decompressed video using a second coding standard, in which case the codec can be referred to as a "transcoder."
[0023] A video encoding process can identify and retain useful information that can be used to reconstruct an image, and ignore information that is not important to the reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy as a trade-off to reduce the required storage space and transmission bandwidth.
[0024] Useful information for a certain coded image (referred to as the "current frame") includes changes relative to a reference frame (e.g., a previously coded and reconstructed frame). These changes may include positional changes, luminance changes, or color changes of pixels, with positional changes being of greatest interest. The positional changes of a group of pixels representing an object may reflect the motion of the object between the reference frame and the current frame.
[0025] A coded frame that does not reference another frame (i.e., it is its own reference frame) is called an "I-frame." A coded frame that uses a previous frame as a reference frame is called a "P-frame." A coded frame that uses both some previous frame and some future frame as a reference frame (i.e., the reference is "bidirectional") is called a "B-frame."
[0026] Figure 1According to some embodiments of the present disclosure, the structure of an exemplary video sequence 100 is shown. Video sequence 100 can be a real-time video or a video that has been captured and archived. Video 100 can be a real-life video, a computer-generated video (such as a computer game video), or a combination of the two (such as a real-life video with augmented reality effects). Video sequence 100 can be input from a video capture device (such as a camera), a video archive containing previously captured videos (such as a video file stored in a storage device), or input from a video feed interface (such as a video broadcast transceiver) to receive video from a video content provider.
[0027] like Figure 1 As shown, video sequence 100 may include a series of frames arranged along a time axis, including frames 102, 104, 106, and 108. Frames 102-106 are continuous, and there are many frames between frame 106 and frame 108. Figure 1 , frame 102 is an I-frame whose reference frame is frame 102 itself. Frame 104 is a P-frame whose reference frame is frame 102, as indicated by the arrows. Frame 106 is a B-frame whose reference frames are frame 104 and frame 108, as indicated by the arrows in the figure. In some embodiments, the reference frame of a frame (e.g., frame 104) is not directly before or after the frame. For example, the reference frame of frame 104 can be a frame before frame 102. It should be noted that the reference frames of frames 102-106 are only examples and the present disclosure is not limited to such frames. Figure 1 Embodiments of the reference frames are shown as examples.
[0028] Typically, video codecs do not encode or decode an entire frame at once, due to the computational complexity of such tasks. Instead, they may segment a frame into basic segments and encode or decode the frame segment by segment. Such basic segments are referred to as basic processing units ("BPUs") in this disclosure. For example, Figure 1 Structure 110 in shows an example structure of a frame of video sequence 100 (e.g., any frame of frames 102-108). In structure 110, a frame is divided into 4×4 basic processing units, whose boundaries are represented by dashed lines. In some embodiments, the basic processing units are referred to as "macroblocks" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in a frame can have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or can be pixels of any shape and size. The size and shape of the basic processing units can be selected for a frame based on a balance between coding efficiency and the level of detail to be retained in the basic processing units.
[0029] A basic processing unit may be a logical unit that may include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image may include a luma component (Y) representing achromatic brightness information, one or more chroma components (e.g., Cb and Cr) representing color information and associated syntax elements, and the basic processing units for luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" (CTBs). Any operation performed on a basic processing unit may be repeatedly performed on its individual luma and chroma components.
[0030] Video encoding has multiple stages of operation, examples of which are Figure 2 and Figure 3 As shown. For each stage, the size of the basic processing unit may still be too large to process, so it can be further divided into segments referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or a "coding unit" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same or smaller size as the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., a video frame buffer). Any operation performed on the basic processing subunit can be repeatedly performed on its various luminance and chrominance components. It should be noted that this division can be further performed according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing unit.
[0031] For example, in the mode decision stage (such as Figure 2 As an example shown), the encoder should decide what prediction mode to use for a basic processing unit (e.g., intra prediction or inter prediction); but the basic processing unit may be too large to make such a decision, the encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine a prediction type for each individual basic processing sub-unit.
[0032] For example, in the prediction stage ( Figure 2An example is shown), the encoder can perform prediction operations at the level of basic processing sub-units (such as CUs). However, in some cases, the basic processing sub-units may still be too large to be processed. The encoder can further split the basic processing sub-units into smaller fragments (for example, called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations can be performed.
[0033] As another example, in the transformation phase ( Figure 2 An example is shown), the encoder can perform transform operations on the remaining basic processing sub-units (such as CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder can further split the basic processing sub-units into smaller fragments (for example, called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transform operation can be performed at this level. It should be noted that the division scheme of the same basic processing sub-unit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.
[0034] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, and the boundaries of the sub-units are shown as dot-dashed lines in the figure. In different schemes, different basic processing units of the same frame can be divided into different basic processing sub-units.
[0035] In some embodiments, the video encoding and decoding are provided with parallel processing capabilities and error recovery, and a frame can be divided into several regions for processing, so that for a region of the frame, the encoding or decoding process can be independent of any information from other regions of the frame. In other words, each region in the frame can be processed independently. In this way, the codec can process different regions of the image frame in parallel, thereby improving the coding efficiency. In addition, when the data of a region is damaged during processing or lost in network transmission, the codec can correctly encode or decode other regions of the same image frame without relying on the damaged or lost data, thereby providing fault tolerance. In some video coding standards, a frame can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different frames of the video sequence 100 can have different division schemes for dividing the frame into regions.
[0036] For example, in Figure 1In FIG. 1 , the structure 110 is divided into three regions 114, 116 and 118, and the boundaries of the regions are represented by solid lines inside the structure 110. The region 114 includes four basic processing units. The regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and regions of the structure 110 are just examples, and the present disclosure does not limit implementation thereof.
[0037] According to some embodiments of the present disclosure, Figure 2 A schematic diagram of an exemplary encoder 200 in a hybrid video coding system is shown. The video encoder 200 can perform intra-coding or inter-coding of blocks within a video frame, including video blocks, or partitions or sub-partitions of video blocks. Intra-coding can rely on spatial prediction to reduce or eliminate video spatial redundancy within a given video frame. Inter-coding can rely on temporal prediction to reduce or remove temporal redundancy in adjacent frames in a video sequence. Intra-mode can refer to some spatial-based compression modes. Inter-mode (such as uni-prediction or bi-prediction) can refer to some temporal-based compression modes.
[0038] Reference Figure 2 , the input video signal 202 can be processed block by block. For example, the video block unit can be a 16×16 pixel block (e.g., a macroblock (MB)). The size of the video block unit may vary, depending on the coding technology used, and the accuracy and efficiency required. In HEVC, extended block sizes such as coding tree units (CTUs) can be used to compress video signals with a resolution of 1080p or higher. In HEVC, a CTU can include up to 64×64 luma samples corresponding to chroma samples, and related syntax elements. In VVC, the size of the CTU can be further increased to 128x128 luma samples corresponding to chroma samples, and related syntax elements. A CTU can be further divided into coding units (CUs), for example, using a quadtree, a binary tree, or a ternary tree. A CU can be further divided into prediction units (PUs), and different prediction methods can be applied to these units. Each input video block can be processed using a spatial prediction unit 260 or a temporal prediction unit 262.
[0039] The spatial prediction unit 260 uses information on the same frame / slice containing the current block to perform spatial prediction (e.g., intra prediction) on the current block / CU. Spatial prediction can use pixels in neighboring blocks that have been encoded in the same video frame / slice to predict the current video block. Spatial prediction can reduce the inherent spatial redundancy in the video signal.
[0040] The temporal prediction unit 262 performs temporal prediction (such as inter-prediction) on the current block using information from other frames / slices other than the frame / slice containing the current block. The temporal prediction of a video block can be identified by one or more motion vectors. In unidirectional temporal prediction, only one motion vector identifying a reference frame is used to generate a prediction identifier for the current block. On the other hand, in bidirectional temporal prediction, two motion vectors, each identifying a reference frame, can be used to generate a prediction identifier for the current block. The motion vector can show the amount and direction of motion between the current block and one or more related blocks in the reference frame. If multiple reference frames are supported, one or more reference frame indexes can be sent for a video block. The one or more reference indexes are used to identify which reference frame in the reference picture library or decoded picture buffer (DPB) 264 can generate a temporal prediction signal.
[0041] The mode decision and encoder control unit 280 in the encoder can select the prediction mode, for example, based on rate distortion optimization. According to the determined prediction mode, a prediction block can be obtained. The prediction block can be subtracted from the current video block at the adder 216. The prediction residual can be transformed by the transform unit 204 and quantized by the quantization unit 206. The quantized residual coefficients can be dequantized in the dequantization unit 210 and detransformed in the inverse transform unit 212 to form a reconstructed residual. The reconstructed residual is added to the prediction block at the adder 226 to form a reconstructed video block. The reconstructed video block before loop filtering can be used as a reference sample for internal prediction.
[0042] The reconstructed video block may be loop filtered on the loop filter 266. For example, loop filtering such as deblocking filtering, sample adaptive offset (SAO) and adaptive loop filter (ALF) may be used. The reconstructed block after loop filtering may be stored in the reference picture library 264 and may provide inter-prediction reference samples for encoding other video blocks. To form the output video bitstream 220, before the data is compressed and packaged to form the bitstream 220, the coding mode (such as (frame) inter or (frame) intra), prediction mode information, motion information, quantized residual coefficients, etc. may be sent to the entropy coding unit 208 to further reduce the bit rate.
[0043] According to some embodiments of the present disclosure, Figure 3 FIG. 3 is a schematic diagram of an exemplary decoder 300 in a hybrid video coding system. Figure 3 , the video bitstream 302 may be unpacked or entropy decoded at the entropy decoding unit 308. The coding mode information may be used to determine whether to select the spatial prediction unit 360 or the temporal prediction unit 362. The prediction mode information may be sent to the corresponding prediction unit to generate a prediction block. For example, the temporal prediction unit 362 may apply motion compensated prediction to form the temporal prediction block.
[0044] The residual coefficients are sent to the inverse quantization unit 310 and the inverse transform unit 312 to obtain the reconstructed residual. The predicted block and the reconstructed residual are added at 326 to form a reconstructed block before loop filtering. The reconstructed block can then be loop filtered at the loop filter 366. For example, deblocking filtering, SAO, ALF and other loop filtering can be applied. The reconstructed block after loop filtering can be stored in the reference picture library 364. The reconstructed data in the reference picture library 364 can be used to obtain the decoded video 320, or to predict subsequent video blocks. The decoded video 320 can be displayed on a display device, such as a TV, PC, smart phone or tablet computer, for viewing by the end user.
[0045] According to some embodiments of the present invention, Figure 4 4 is a block diagram of an exemplary apparatus 400 for encoding or decoding video. Figure 4 As shown, the device 400 may include a processor 402. When the processor 402 executes the instruction set described herein, the device 400 may become a dedicated machine for video encoding or decoding. The processor 402 may be any type of circuit system capable of operating or processing information. For example, the processor 402 may include any number of any combination of a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a microcontroller unit (MCU), an optical processor, a programmable logic controller, a single-chip microcomputer, a microprocessor, a digital signal processor, an IP core, a programmable logic array (PLA), a programmable array logic (PAL), a general array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), etc. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0046] The apparatus 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). Figure 4 As shown, the stored data may include program instructions (e.g., for executing Figure 2 or Figure 3The processor 402 can access the program instructions and data to be processed (for example, through the bus 410), and execute the program instructions to perform operations or control on the data for processing. The memory 404 can include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 can include any combination of any number of random access memories (RAM), a read-only memory (ROM), an optical disk, a magnetic disk, a hard disk, a solid-state drive, a flash drive, a secure digital (SD) card, a memory stick, a compact flash (CF) card, or other similar elements. The memory 404 can also be a group of memories ( Figure 4 not shown).
[0047] The bus 410 may be a communication device for transmitting data between internal components of the apparatus 400, such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.
[0048] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this disclosure. The data processing circuits may be implemented entirely in hardware, or in a combination of software, hardware, or firmware. In addition, the data processing circuit may be an independent module, or may be fully or partially integrated into other components of the device 400.
[0049] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, the network interface 406 may include a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network card, a wireless network card, a Bluetooth network card, an infrared network card, a near field communication (NFC) adapter, a cellular network chip, or the like.
[0050] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, a cursor control device (such as a mouse, touch pad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface for connecting to a video file), or other similar devices.
[0051] It should be noted that the video codec may be implemented as any combination of any software or hardware modules in the apparatus 400. For example, Figure 2 Encoder 200 or Figure 3 Some or all stages of the decoder 300 may be implemented as one or more software modules of the apparatus 400, for example, program instructions loadable into the memory 404. In another example, Figure 2 Encoder 200 or Figure 3 Some or all stages of decoder 300 may be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuits (eg, FPGA, ASIC, NPU, or similar circuits).
[0052] In the quantization and dequantization blocks (such as Figure 2 The quantization unit 206 and the inverse quantization unit 210, Figure 3 The inverse quantization unit 310 of the quantization group granularity determines the amount of quantization (and inverse quantization) applied to the prediction residual using a quantization parameter (QP). The initial QP value for encoding a frame or slice can be identified at a high level, for example, using the syntax element init_qp_minus26 in the frame parameter set (PPS) and the syntax element slice_qp_delta in the slice header. In addition, the incremental QP value sent at the quantization group granularity can be used to adjust the QP value for each CU at this level.
[0053] In VVC, sub-block transform (SBT) is used for inter-prediction coding units (CUs). In this transform mode, only a sub-part of the residual block is encoded and provided to the coding unit. When the inter-prediction unit with the syntax element cu_cbf is equal to 1, the syntax element cu_sbt_flag can be marked to indicate whether the entire residual block is encoded or a sub-part of the residual block is encoded. In the former case, the MTS (inter multiple transform selected) information is further parsed to determine the transform type of the CU. In the latter case, part of the residual block is encoded using the inferred adaptive transform, and the other parts of the residual block are set to zero.
[0054] When SBT is used for an inter-prediction CU, the SBT type and SBT position information are identified in the bitstream. There are two SBT types and two SBT positions, such as Figure 5As shown. For SBT-V (or SBT-H), the width (or height) of the transform unit (TU) can be equal to half the CU width (or height) or 1 / 4 of the CU width (or height), forming a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to the binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to the asymmetric binary tree (ABT) partition. In the ABT partition, only its small area contains non-zero residuals. If a coding unit has 8 luminance samples in a certain dimension, then 1:3 / 3:1 partition along this dimension is not allowed. A coding unit has a maximum of 8 SBT modes.
[0055] The sequence parameter set (SPS) level syntax can use the syntax element sps_sbt_enabled_flag to specify enabling or disabling of SBT. When the syntax element sps_sbt_enabled_flag is equal to 0, it indicates that SBT for inter-predicted coding units is disabled in the entire video sequence that references this SPS. When the syntax element sps_sbt_enabled_flag is equal to 1, it indicates that SBT for inter-predicted coding units is enabled for the entire video sequence that references this SPS.
[0056] In addition, when sps_sbt_enabled_flag is equal to 1, another SPS syntax element sps_sbt_max_size_64_flag can be used to specify the maximum CU width and height allowed for SBT. When the syntax element sps_sbt_max_size_64_flag is equal to 0, it indicates that the maximum CU width and height allowed for SBT is 32 luma samples. When the syntax element sps_sbt_max_size_64_flag is equal to 1, it indicates that the maximum CU width and height allowed for SBT is 64 luma samples. The variable MaxSbtSize is calculated according to the following formula 1, which can specify the maximum CU size allowed for SBT: MaxSbtSize=Min(MaxTbSizeY,sps_sbt_max_size_64_flag?64:32)(Formula 1) Where MaxTbSizeY is the maximum allowed transform block (TB) size, which can be derived from another SPS level syntax element sps_max_luma_transform_size_64_flag according to the following formula 2: MaxTbSizeY=sps_max_luma_transform_size_64_flag? 64:32(Formula 2)
[0057] As mentioned above, the source of MaxSbtSize depends on two syntax elements sps_max_luma_transform_size_64_flag and sps_sbt_max_size_64_flag. If the value of the syntax element sps_max_luma_transform_size_64_flag is 0, MaxSbtSize is always 32 regardless of the value of the syntax element sps_sbt_max_size_64_flag. Therefore, when the syntax element sps_max_luma_transform_size_64_flag is 0, there is no need to identify the syntax element sps_sbt_max_size_64_flag. This syntax redundancy in VVC unnecessarily increases signal overhead.
[0058] In order to improve video coding efficiency, according to some disclosed embodiments, the syntax element sps_sbt_max_size_64_flag is only signaled when the syntax elements sps_max_luma_transform_size_64_flag and sps_sbt_enabled_flag are both 1. Figure 6 An exemplary Table 1 according to some embodiments of the present disclosure is shown. Table 1 shows an exemplary SPS syntax table for some embodiments. As shown in Table 1 (emphasized parts are shown in italics), the syntax element sps_sbt_max_size_64_flag is only identified when the syntax elements sps_max_luma_transform_size_64_flag and sps_sbt_enabled_flag are both 1. If the syntax element sps_max_luma_transform_size_64_flag is 0, the syntax element sps_sbt_max_size_64_flag can be inferred to be 0, which means that the maximum width and height of the CU that allows SBT is 32 (in units of luma samples).
[0059] Figure 7 FIG. 7 is a flowchart of an exemplary video processing method 700 according to some embodiments of the present disclosure. In some embodiments, the method 700 may be performed by an encoder (e.g., Figure 2 200), a decoder (e.g., Figure 3 decoder 300) or one or more software or hardware components (e.g., Figure 4 For example, a processor (e.g., Figure 4The method 700 may be performed by a processor 402 of the computer. In some embodiments, the method 700 may be implemented by a computer program product contained in a computer-readable medium, the product including computer-executable instructions, such as program codes (e.g., Figure 4 Device 400 in).
[0060] In step 702, method 700 may include determining whether a sub-block transform (SBT) is enabled in a sequence parameter set (SPS) of a video sequence. In some embodiments, a flag bit (e.g., Figure 6 The syntax element sps_sbt_enabled_flag shown in Table 1 of ) can be identified in the SPS indicating whether SBT is enabled. For example, the syntax element sps_sbt_enabled_flag equal to 0 can specify that the SBT of the inter-prediction coding unit is disabled for the entire video sequence referencing the SPS. And the syntax element sps_sbt_enabled_flag = 1 can specify that the SBT of the inter-prediction coding unit is enabled for the entire video sequence referencing the SPS.
[0061] In step 704, method 700 may include determining a value of a first flag bit in the SPS, which indicates the maximum transition block (TB) size allowed for the SBT. The first flag bit may be set to a first value or a second value. For example, the first value is 1 and the second value is 0. The maximum TB size may be 32, 64, or the like. In some embodiments, method 700 may also include setting the first flag bit to a first value corresponding to the maximum TB size being 64, and setting the first flag bit to a second value corresponding to the maximum TB size being 32. In some embodiments, the first flag bit may be Figure 6 The syntax element sps_max_luma_transform_size_64_flag in Table 1.
[0062] In step 706, the method 700 may include, in response to SBT being enabled and the value of the first flag bit being equal to the first value, marking a second flag bit indicating a maximum coding unit (CU) size for which SBT is allowed. In response to SBT being disabled or the value of the first flag bit being equal to the second value, the second flag bit is not marked. For example, the second flag bit may be as follows: Figure 6 The syntax element sps_sbt_max_size_64_flag is shown in Table 1. The syntax element sps_sbt_max_size_64_flag is signaled only when both the syntax element sps_max_luma_transform_size_64_flag and the sps_sbt_enabled_flag are 1.
[0063] In some embodiments, method 700 may further include identifying a third flag bit in the SPS (eg, Figure 6 The syntax element sps_sbt_enabled_flag shown in Table 1 indicates whether SBT is enabled, and identifies the first flag bit in the SPS (eg, Figure 6 The syntax element sps_max_luma_transform_size_64_flag in Table 1).
[0064] In some embodiments, the maximum CU size may be 32 or 64. The maximum CU width or height allowed for the SBT may be determined according to the smaller of the maximum TB size and the maximum CU size (eg, according to Formula 1).
[0065] In some disclosed embodiments, the syntax element sps_sbt_max_size_64_flag is not flagged at all. In this case, the maximum width and height allowed for a CU of the SBT depends directly on the syntax element sps_max_luma_transform_size_64_flag. If the syntax element sps_max_luma_transform_size_64_flag is equal to 0, the maximum CU width and height allowed for the SBT is 32 luma samples. If the syntax element sps_max_luma_transform_size_64_flag is equal to 1, the maximum CU width and height allowed for the SBT is 64 luma samples. In other words, MaxSbtSize is set equal to MaxTbSizeY. Figure 8 According to some embodiments of the present disclosure, an exemplary Table 2 is shown. Table 2 shows an exemplary SPS syntax for implementing these embodiments. As shown in Table 2, the syntax element sps_sbt_max_size_64_flag is not identified and is deleted from the syntax. Fig. 9 According to some embodiments of the present disclosure, an exemplary Table 3 is shown. Table 3 (emphasized in italics) shows an exemplary coding unit (CU) syntax table, which directly uses MaxTbSizeY to set the maximum width and height of the CU. MaxTbSizeY is calculated by the following formula 3: MaxTbSizeY=sps_max_luma_transform_size_64_flag? 64:32(Formula 3)
[0066] Fig.10 FIG. 1 is a flowchart of another exemplary video processing method 1000 according to some embodiments of the present disclosure. In some embodiments, the method 1000 may be performed by an encoder (eg, Figure 2 200), a decoder (e.g., Figure 3 decoder 300) or one or more software or hardware components of a device (e.g., Figure 4 For example, a processor (e.g., Figure 4 The method 1000 may be performed by a processor 402 of the computer system. In some embodiments, the method 1000 may be implemented by a computer program product contained in a computer-readable medium, the product including computer-executable instructions, such as by a computer (e.g., Figure 4 Program code executed by device 400 in the device.
[0067] In step 1002, method 1000 includes identifying a first flag in a sequence parameter set (SPS) of a video sequence, the flag indicating whether sub-block transform (SBT) is enabled. In some embodiments, the first flag may be a syntax element sps_sbt_enabled_flag, such as Figure 8 As shown in Table 2. For example, the syntax element sps_sbt_enabled_flag equal to 0 can specify that the SBT of the inter-prediction coding unit is disabled for the entire video sequence referencing the SPS. And the syntax element sps_sbt_enabled_flag equal to 1 can specify that the SBT of the inter-prediction coding unit is enabled for the entire video sequence referencing the SPS.
[0068] In step 1004, method 1000 may include identifying a second flag bit to indicate the maximum transition block (TB) size allowed for the SBT. The second flag bit may be set to a first value or a second value. For example, the first value is 1 and the second value is 0. The maximum TB size may be 32, 64, or a similar value. In some embodiments, method 1000 may further include setting the value of the second flag bit to 0 when corresponding to a maximum TB size of 32; and setting the value of the second flag bit to 1 when corresponding to a maximum TB size of 64. In some embodiments, the first flag bit may be Figure 8 The syntax element sps_max_luma_transform_size_64_flag in Table 2.
[0069] In response to the first flag indicating that SBT is enabled, the maximum CU size that allows SBT can be directly determined based on the maximum TB size. For example, the maximum CU size is equal to the maximum TB size. The maximum CU size includes the maximum width and maximum height of the CU.
[0070] In some embodiments, a non-volatile computer-readable storage medium including an instruction set is also provided, and the instruction set can be executed by a device (such as the encoder and decoder) for performing the above method. Common forms of non-volatile media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any perforated physical media mode, RAM, PROM, and EPROM, flash EPROM or other flash memory, NVRAM, cache, registers, any other storage chip or tape, and network versions of the same. The device may include one or more processors (CPU), input / output interfaces, network interfaces, and / or memory.
[0071] The embodiments may be further described using the following terms: 1. A video processing method, comprising: Determining whether sub-block transform (SBT) is enabled in a sequence parameter set (SPS) of a video sequence; determining the value of a first flag bit in the SPS, the flag bit indicating the maximum transform block (TB) size allowed for the SBT; and In response to SBT being enabled and the value of the first flag bit being equal to a first value, a second flag bit is identified to indicate a maximum coding unit (CU) size that allows SBT. 2. The method according to clause 1, wherein, in response to the SBT not being enabled or the value of the first flag bit being equal to a second value, the second flag bit is not identified. 3. The method according to clause 1 or clause 2, further comprising: Identify the third flag bit in the SPS to indicate whether SBT is enabled; and Identifies the first flag bit in the SPS. 4. The method according to clause 2, wherein the first value is 1 and the second value is 0. 5. According to any method of clauses 1-4, wherein the maximum TB size is 32 or 64. 6. The method according to clause 5 further comprises: in response to the maximum TB size being 64, setting the value of the first flag bit to a first value. 7. The method according to clause 5, further comprising: In response to the maximum TB size being 32, the value of the first flag bit is set to the second value. 8. A method according to any of clauses 1-7, wherein the maximum CU size allowed for SBT is 32 or 64. 9. The method according to any one of clauses 1 to 8, wherein the maximum CU width allowed for the SBT is determined according to the smaller of the maximum TB size and the maximum CU size allowed for the SBT. 10. A method according to any one of clauses 1-9, wherein the maximum CU height allowed for the SBT is determined based on the smaller of the maximum TB size and the maximum CU size allowed for the SBT. 11. A video processing device, comprising: at least one memory for storing an instruction set; and At least one processor executes the set of instructions to cause the device to: Determining whether sub-block transform (SBT) is enabled in a sequence parameter set (SPS) of a video sequence; determining the value of a first flag bit in the SPS, the flag bit indicating the maximum transform block (TB) size allowed for the SBT; and In response to the SBT being enabled and the value of the first flag bit being equal to the first value, a second flag bit is identified to indicate a maximum coding unit (CU) size that allows the SBT. 12. The apparatus according to clause 11, wherein the second flag bit is not identified corresponding to the SBT not being enabled or the value of the first flag bit being equal to the second value. 13. The apparatus of clause 11 or clause 12, wherein at least one processor further executes the set of instructions to cause the apparatus to perform: Identify the third flag bit in the SPS to indicate whether the SBT is enabled; and Identifies the first flag bit in the SPS. 14. The apparatus of clause 12, wherein the first value is 1 and the second value is 0. 15. According to any of Clauses 11-14, wherein the maximum TB size is 32 or 64. 16. The apparatus according to clause 15, further comprising: In response to the maximum TB size being 64, the value of the first flag bit is set to a first value. 17. The apparatus according to clause 15, further comprising: In response to the maximum TB size being 32, the value of the first flag bit is set to a second value. 18. The apparatus of any one of clauses 11-17, wherein the maximum CU size allowed for the SBT is 32 or 64. 19. The apparatus according to any one of clauses 11-18, wherein the maximum CU width allowed for the SBT is determined according to the smaller of a maximum TB size and a maximum CU size allowed for the SBT. 20. The apparatus of any one of clauses 11-19, wherein the maximum CU height for which the SBT is allowed is determined according to the smaller of the maximum TB size and the maximum CU size for which the SBT is allowed. 21. A non-volatile computer-readable storage medium storing a set of instructions, wherein the instruction set is executed by at least one processor to enable a computer to perform a video processing method, comprising: Determining whether sub-block transform (SBT) is enabled in a sequence parameter set (SPS) of a video sequence; determining the value of a first flag bit in the SPS, the flag bit indicating the maximum transform block (TB) size allowed for the SBT; and In response to the SBT being enabled and the value of the first flag bit being equal to the first value, a second flag bit is identified to indicate a maximum coding unit (CU) size that allows the SBT. 22. The non-volatile computer-readable storage medium according to clause 21, wherein the second flag bit is not identified corresponding to the SBT not being enabled or the value of the first flag bit being equal to a second value. 23. The non-transitory computer-readable storage medium of clause 21 or clause 22, wherein the set of instructions executed by at least one processor causes the computer to further perform: Identify a third flag bit in the SPS to indicate whether the SBT is enabled; and Identifies the first flag bit in the SPS. 24. The non-volatile computer-readable storage medium of clause 22, wherein the first value is 1 and the second value is 0. 25. A non-volatile computer-readable storage medium according to any of clauses 21-24, wherein the maximum TB size is 32 or 64. 26. The non-transitory computer-readable storage medium of clause 25, wherein a set of instructions executed by at least one processor causes the computer to further perform: In response to the maximum TB size being 64, the value of the first flag bit is set to a first value. 27. The non-transitory computer-readable storage medium of clause 25, wherein a set of instructions executed by at least one processor causes the computer to further perform: In response to the maximum TB size being 32, the value of the first flag bit is set to the second value. 28. A non-volatile computer-readable storage medium according to any one of clauses 21-27, wherein the maximum CU size allowed for the SBT is 32 or 64. 29. The non-volatile computer-readable storage medium of any one of clauses 21-28, wherein the maximum CU width allowed for the SBT is determined based on the smaller of the maximum TB size and the maximum CU size allowed for the SBT. 30. The non-volatile computer-readable storage medium of any one of clauses 21-29, wherein the maximum CU height allowed for the SBT is determined based on the smaller of the maximum TB size and the maximum CU size allowed for the SBT. 31. A video processing method, comprising: In a sequence parameter set SPS of a video sequence, a first flag is identified to indicate whether a sub-block transform SBT is enabled; and Identify the second flag to indicate the maximum conversion block TB size allowed for SBT, In response to the first flag indicating that the SBT is enabled, a maximum coding unit CU size of the SBT is allowed to be directly determined according to the maximum transform block TB size. 32. The method of clause 31, wherein the maximum CU size allowed for the SBT is the maximum CU width or the maximum CU height. 33. The method of clause 31 or 32, wherein determining a maximum CU size that allows the SBT is determined to be equal to a maximum TB size. 34. The method according to any one of clauses 31-33, wherein the maximum TB size is 32 or 64. 35. The method according to clause 34, further comprising: In response to the maximum TB size being 32, the value of the second flag bit is set to 0. 36. The method according to clause 34, further comprising: In response to the maximum TB size being 64, the value of the second flag bit is set to 1. 37. A video processing device, comprising: at least one memory for storing instructions; and At least one processor executes a set of instructions to cause the device to: In a sequence parameter set SPS of a video sequence, a first flag is identified to indicate whether a sub-block transform SBT is enabled; and In response to the first flag indicating that the SBT is enabled, a maximum coding unit CU size of the SBT is allowed to be directly determined according to the maximum transform block TB size. 38. An apparatus according to clause 37, wherein the maximum CU size allowed for the SBT is a maximum CU width or a maximum CU height. 39. An apparatus as described in claim 37 or 38, wherein the maximum CU size allowed for the SBT is determined to be equal to the maximum TB size. 40. An apparatus according to any one of clauses 37-39, wherein the maximum TB size is 32 or 64. 41. The apparatus of clause 40, wherein at least one processor further executes instructions to cause the apparatus to perform: in response to a maximum TB size being 32, setting the value of the second flag bit to 0. 42. The apparatus of clause 40, wherein at least one processor further executes an instruction set to cause the apparatus to perform: in response to the maximum TB size being 64, setting the value of the second flag bit to 1. 43. A non-volatile computer-readable storage medium storing a set of instructions executed by at least one processor to cause the computer to perform a video processing method, comprising: In a sequence parameter set SPS of a video sequence, a first flag is identified to indicate whether a sub-block transform SBT is enabled; and Identify the second flag to indicate the maximum conversion block TB size allowed for SBT, In response to the first flag indicating that the SBT is enabled, a maximum coding unit CU size of the SBT is allowed to be directly determined according to the maximum transform block TB size. 44. The non-transitory computer-readable storage medium of clause 43, wherein the maximum CU size allowed for the SBT is a maximum CU width or a maximum CU height. 45. The non-transitory computer-readable storage medium of clause 43 or 44, wherein a maximum CU size allowed for the SBT is determined to be equal to a maximum TB size. 46. The non-volatile computer-readable storage medium of any one of clauses 43-45, wherein the maximum TB size is 32 or 64. 47. The non-transitory computer-readable storage medium of clause 46, wherein a set of instructions executable by at least one processor causes the computer to further perform: In response to the maximum TB size being 32, the value of the second flag bit is set to 0. 48. The non-transitory computer-readable storage medium of clause 46, wherein the set of instructions executed by the at least one processor causes the computer to further perform: In response to the maximum TB size being 64, the value of the second flag bit is set to 1.
[0072] It should be noted that the relational terms such as "first" and "second" in this article are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include", "have", "include" and "includes" and other similar forms are the same in meaning, and the end of any one or more items after any of the above words is open-ended, and any of the above nouns does not mean that the one or more items have been exhaustively listed or are limited to the one or more items listed.
[0073] When used herein, unless otherwise expressly stated, the term "or" includes all possible combinations, except those that are not feasible. For example, if it is expressed that a database may include A or B, then unless otherwise specifically stated or not feasible, it may include database A, or B, or A and B. As a second example, if it is expressed that a database may include A, B, or C, then unless otherwise specifically stated or not feasible, the database may include database A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0074] It is worth noting that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. When the software is executed by a processor, it can execute the above-mentioned disclosed method. The computing unit and other functional units described in the present disclosure can be implemented by hardware or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.
[0075] In the above detailed description, the embodiments have been described with reference to many specific details, which may vary from implementation to implementation. Certain adaptations and modifications may be made to the embodiments. For those skilled in the art, other embodiments may be apparent from the specific embodiments disclosed in the present invention. This specification and examples are for exemplary purposes only, and the true scope and essence of the present invention are described by the claims. The order of steps shown in the diagrams is also for the purpose of explanation only and is not meant to be limited to any particular steps or order. Therefore, those skilled in the art will appreciate that these steps may be performed in different orders when implementing the same method.
[0076] In the drawings and detailed description of the present application, exemplary embodiments are disclosed. However, many variations and modifications may be made to these embodiments. Accordingly, although specific terms are used, these terms are only general and descriptive and not for the purpose of limitation.
Claims
1. A method for decoding a bitstream to output one or more images of a video sequence, the method comprising: include: Decoding the bit stream; Determining, based on the decoded bitstream, a maximum transform size in units of luma samples; as well as decoding a first flag in a sequence parameter set (SPS) associated with the video sequence, the first flag indicating whether a sub-block transform (SBT) is enabled; Determining whether the SBT is allowed for a coding unit (CU) of the video sequence; Wherein, whether the SBT is allowed to be used for the CU is determined based on a comparison between the size of the CU and the maximum transform size in units of luma samples.
2. The method according to claim 1, in: Decoding the bitstream includes decoding a second flag associated with the video sequence; and Determining the maximum transform size in units of luma samples includes determining the maximum transform size in units of luma samples based on a value of the second flag.
3. The method according to claim 2, in: The second flag is signaled in an SPS of the bitstream.
4. The method according to claim 2, in: The second flag is sps_max_luma_transform_size_64_flag.
5. The method according to claim 2, further comprising: include: In response to the value of the second flag being 1, determining that the maximum conversion size in units of luma samples is equal to 64; or In response to the value of the second flag being 0, the maximum transform size in units of luma samples is determined to be equal to 32.
6. The method according to claim 2, further comprising: include: In response to the second flag having a first value, determining the maximum transform size in units of luma samples to be equal to a second value; or In response to the second flag having a third value, the maximum transform size in units of luma samples is determined to be equal to a fourth value.
7. The method according to claim 1, in, The comparison of the size of the CU with the maximum transform size in units of luma samples comprises at least one of the following: a comparison of the width of the CU to the maximum transform size in units of luma samples, or A comparison of the height of the CU to the maximum transform size in units of luma samples.
8. The method according to claim 1, further comprising: include: A size of a maximum CU that allows the SBT to be equal to the maximum transform size in units of luma samples is determined.
9. A method for encoding a video sequence into a bitstream, the method comprising: determining a maximum transform size in units of luma samples for the video sequence; encoding a first flag in a sequence parameter set (SPS) associated with the video sequence, the first flag indicating whether a sub-block transform (SBT) is enabled; determining whether to use a sub-block transform (SBT) for a coding unit of the video sequence, in, Whether to use the SBT for a coding unit (CU) is determined based on a comparison of the size of the CU with a maximum transform size in units of luma samples.
10. The method according to claim 9, further comprising: include: A second flag indicating the maximum transform size in units of luma samples is encoded.
11. The method of claim 10, encoding the second flag in an SPS of a bitstream associated with the video sequence.
12. The method according to claim 10, in: The second flag is sps_max_luma_transform_size_64_flag.
13. The method according to claim 10, further comprising: include: In response to a maximum transform size in units of luma samples being determined to be equal to 64, setting the value of the second flag to 1; or In response to the maximum transform size in units of luma samples being determined to be equal to 32, the value of the second flag is set to 0.
14. The method according to claim 10, further comprising: In response to the maximum transform size in units of luma samples being determined to be equal to the first value, setting the second flag to have a second value; or In response to the maximum transform size in units of luma samples being determined to be equal to the third value, the second flag is set to have a fourth value.
15. The method of claim 9, wherein the comparison of the size of the CU with the maximum transform size in units of luma samples comprises at least one of: a comparison of the width of the CU to the maximum transform size in units of luma samples, or A comparison of the height of the CU to the maximum transform size in units of luma samples.
16. A non-transitory computer-readable storage medium storing a bitstream associated with a video sequence, in, The bitstream can be decoded by: determining a maximum transform size in units of luma samples for the video sequence; and decoding a first flag in a sequence parameter set (SPS) associated with the video sequence, the first flag indicating whether a sub-block transform (SBT) is enabled; Determine whether SBT is allowed for a coding unit (CU) of a video sequence, Wherein whether the SBT is allowed for the CU is determined based on a comparison between the size of the CU and a maximum transform size in units of luma samples.
17. The non-transitory computer-readable storage medium of claim 16, It is characterized in that The bitstream comprises a second flag indicating a maximum transform size in units of luma samples.
18. The non-transitory computer-readable storage medium of claim 17, It is characterized in that The bitstream comprises an SPS associated with the video sequence, the second flag being signaled in the SPS.
19. The non-transitory computer-readable storage medium of claim 17, It is characterized in that The bitstream can be decoded by: In response to the value of the second flag being 1, determining that the maximum conversion size in units of luma samples is equal to 64; or In response to the value of the second flag being 0, the maximum conversion size in units of luma samples is determined to be equal to 32.
20. The non-transitory computer-readable storage medium of claim 17, It is characterized in that The bitstream can be decoded by: In response to the second flag having a first value, determining that the maximum transform size in units of luma samples is equal to a second value; or In response to the second flag having a third value, a maximum transform size in units of luma samples is determined to be equal to a fourth value.
Citation Information
Patent Citations
Method and apparatus of block partition with smallest block size in video coding
CN108293109A
Method and apparatus for encoding / decoding image
CN110063058A
Video coding with large macroblocks
US20110194613A1