Coding and decoding tree splitting

By increasing the depth of multi-type tree hierarchy and specifying the maximum MTT hierarchy depth for each QT level, the problem of inflexibility of codec tree segmentation in the prior art is solved, and the compression efficiency of video encoding is improved.

CN114175642BActive Publication Date: 2025-07-01INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080053335.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2020-09-18
Publication Date
2025-07-01
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

The existing video encoding technology lacks flexibility in codec tree segmentation, resulting in low compression efficiency.

Method used

Provide greater flexibility by increasing the maximum allowable multi-type tree hierarchy depth to double the difference between the codec tree unit size and the minimum allowable coded block size and specifying the maximum MTT hierarchy depth for each QT level.

Benefits of technology

Added the accessible set of codec tree nodes and leaves, improving the compression efficiency of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175642B_ABST
    Figure CN114175642B_ABST
Patent Text Reader

Abstract

To encode a picture, the coding tree units (CTUs) in the picture are partitioned by a quadtree structure, and the quadtree leaf nodes can be further partitioned by a multi-type tree (MTT) structure. To increase the set of reachable coding tree nodes and leaves, we propose to increase the maximum allowed MTT hierarchical depth to twice the difference between the CTU size and the minimum allowed size of a coding unit (CU). The maximum allowed MTT hierarchical depth can be specified for all QT levels to provide greater flexibility in the split tree. Alternatively, only two levels of the maximum allowed MTT depth are signaled: one when QT splitting is allowed and the other when no more QT splitting is allowed. In addition, an upper limit can be set for the minimum allowed coding block size based on the coding tree unit size or the maximum allowed transform size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This embodiment generally relates to methods and apparatuses for codec tree splitting in video encoding or decoding. Background Art

[0002] To achieve high compression efficiency, image and video codec schemes typically employ prediction and perform transforms to take full advantage of spatial and temporal redundancies in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-picture or inter-picture correlations, and then the difference between the original block and the predicted block (often represented as prediction error or prediction residue) is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded through inverse processes corresponding to entropy coding, quantization, transform, and prediction. Summary of the Invention

[0003] To encode a picture, the codec tree units (CTUs) in the picture are split by a quadtree structure, and the quadtree leaf nodes can be further split by a multi-type tree (MTT) structure. To increase the set of reachable codec tree nodes and leaves, we propose increasing the maximum allowed MTT hierarchical depth to twice the difference between the CTU size and the minimum allowed size of a coding unit (CU). The maximum allowed MTT hierarchical depth can be specified for all quadtree (QT) levels to provide greater flexibility in the split tree. Alternatively, only two levels of the maximum allowed MTT depth are signaled: one when QT splitting is allowed and the other when no more QT splitting is allowed. Additionally, an upper limit can be set for the minimum allowed coding block size based on the codec tree unit size or the maximum allowed transform size. Moreover, flags can be used to indicate whether a binary tree (BT) is allowed and whether a ternary tree (TT) is enabled for the MTT. Flags indicating whether to enable BT or TT can be sent separately for intra-frame stripes / inter-frame stripes and luminance / chrominance components. Brief Description of the Drawings

[0004] Figure 1 A block diagram of a system in which aspects of this embodiment can be implemented is illustrated.

[0005] Figure 2 A block diagram of an embodiment of a video encoder is illustrated.

[0006] Figure 3 A block diagram of an embodiment of a video decoder is illustrated.

[0007] Figure 4 Illustrates splitting a coding unit (CU) into a codec tree unit (CTU) according to the High Efficiency Video Coding (HEVC) standard.

[0008] Figure 5 Illustrates splitting a CTU into a CU, a prediction unit (PU), and a transform unit (TU) according to the HEVC standard.

[0009] Figure 6 Illustrates the quadtree plus binary tree (QTBT) CTU representation in VVC.

[0010] Figure 7 Illustrates the set of all coding unit split modes supported in VVC draft 6.

[0011] Figure 8 Illustrates splitting a 32x32 block using only BT splits of depth 5.

[0012] Figure 9 Illustrates splitting a 32x32 block using only BT splits of depth 2.

[0013] Figure 10 Illustrates splitting a 32x32 block using only TT splits of depth 2.

[0014] Figure 11 Illustrates the minimum, maximum, and common test condition (CTC) values for syntax elements related to splitting in VVC draft 6.

[0015] Figure 12 Illustrates the modified maximum allowable value of the minimum decoded block size based on the maximum allowable transform size according to an embodiment.

[0016] Figure 13 Illustrates the modified maximum allowable value of the minimum decoded block size based on the CTU size according to another embodiment. Detailed Description

[0017] Figure 1 Illustrates a block diagram of an example of a system in which various embodiments can be implemented. System 100 can be implemented as a device including various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 can be implemented individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in this application.

[0018] System 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing various aspects described, for example, in the present application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., volatile memory devices and / or non-volatile memory devices). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0019] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents the (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of System 100 or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.

[0020] The program code to be loaded onto the processor 110 or the encoder / decoder 130 to execute the various aspects described in the present application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, during the execution of the processes described in the present application, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items. Such stored items may include but are not limited to input video, decoded video or a portion of the decoded video, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operation logic.

[0021] In several embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as for MPEG-2, HEVC, or VVC.

[0022] As indicated in block 105, input to the elements of the system 100 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal, e.g., transmitted over the air by a broadcaster, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.

[0023] In various embodiments, the input devices of block 105 have associated corresponding input processing elements, as is known in the art. For example, the RF section can be associated with elements adapted to (i) select a desired frequency (also referred to as selecting a signal, or restricting a signal to a frequency band), (ii) down-convert the selected signal, (iii) restrict the frequency band again to a narrower frequency band to select a signal frequency band that can be referred to as a channel, for example, in some embodiments, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, e.g., a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments reorder the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, e.g., inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0024] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 100 to other electronic devices via the USB and / or HDMI connections. It should be understood that various aspects of the input processing (e.g., Reed-Solomon error correction) may be implemented, for example, in a separate input processing IC or within the processor 110 as needed. Similarly, various aspects of the USB or HDMI interface processing may be implemented, for example, in a separate interface IC or within the processor 110 as needed. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, the processor 110, and an encoder / decoder 130 operating in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0025] The various elements of the system 100 may be provided within an integrated housing. Within the integrated housing, a suitable connection arrangement 115 (e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards) may be used to interconnect the various elements and transfer data between them.

[0026] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0027] In various embodiments, data streams are transmitted to the system 100 using a Wi-Fi network such as IEEE 802.11. The wireless signals of these embodiments are received on a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-air communications. Other embodiments use a set-top box to provide streamed data to the system 100, and the set-top box delivers the data via the HDMI connection of the input block 105. Still other embodiments use the RF connection of the input block 105 to provide streamed data to the system 100.

[0028] System 100 can provide an output signal to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various examples of the embodiments, the other peripheral devices 185 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices based on the output providing function of System 100. In various embodiments, control signals are communicated between System 100 and the display 165, the speaker 175, or the other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols enabling device-to-device control, with or without user intervention. The output devices can be communicatively coupled to System 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to System 100 using a communication channel 190 via a communication interface 150. The display 165 and the speaker 175 can be integrated with other components of System 100 in an electronic device (e.g., a television) in a single unit. In various embodiments, the display interface 160 includes a display driver, e.g., a timing controller (T Con) chip.

[0029] For example, if the RF portion of the input terminal 105 is part of a separate set-top box, then the display 165 and the speaker 175 can alternatively be separate from one or more of the other components. In various embodiments where the display 165 and the speaker 175 are external components, the output signal can be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0030] Figure 2 An example video encoder 200 is illustrated, such as a High Efficiency Video Coding (HEVC) encoder. Figure 2 An encoder that improves on the HEVC standard or an encoder that employs techniques similar to HEVC, such as the Versatile Video Coding (VVC) encoder developed by the JVET (Joint Video Exploration Team), can also be illustrated.

[0031] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "encoding" or "codec" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Generally but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0032] Before being encoded, the video sequence may undergo pre-encoding processing (201). For example, a color transformation (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0) may be applied to the input color picture, or remapping may be performed on the input picture components to obtain a signal distribution that is more resilient to compression (e.g., histogram equalization using one of the color components). Metadata may be associated with the preprocessing and appended to the bitstream.

[0033] To encode a video sequence having one or more pictures, for example, the picture is segmented (202) into one or more strips, where each strip may include one or more strip segments. In HEVC, the strip segments are organized into coding tree units, prediction units, and transform units. The HEVC specification differentiates between "blocks" and "units", where a "block" addresses a specific region in the sample array (e.g., luminance, Y), while a "unit" includes the juxtaposed blocks of all encoded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data associated with the block (e.g., motion vectors).

[0034] For encoding and decoding according to HEVC, the picture is segmented into square coding tree blocks (CTBs) with configurable sizes (usually 64x64, 128x128, or 256x256 pixels), and a set of consecutive coding tree blocks is grouped into a strip. A coding tree unit (CTU), also known as a largest coding unit (LCU), contains the CTBs of the encoded color components. The CTB (also known as the largest coding block, LCB) is the root of a quadtree segmented into coding blocks (CBs), as Figure 4 shown, and the coding block can be segmented into one or more prediction blocks (PBs) and form the root of a quadtree segmented into transform blocks (TBs), as Figure 5 shown.

[0035] Corresponding to the coding block, prediction block, and transform block, a coding unit (CU) includes a set of prediction units (PUs) and a tree-structured transform unit (TU), the PU includes prediction information for all color components, and the TU includes the residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luminance component are applicable to the corresponding CU, PU, and TU. In this application, the term "block" may be used to refer to any one of, for example, CTU, CU, PU, TU, CB, PB, and TB. In addition, the term "block" may also be used to refer to the macroblocks and partitions specified in H.264 / AVC or other video coding standards, and more generally refers to an array of data of various sizes.

[0036] In encoder 200, pictures are encoded by encoder elements as described below. The pictures to be encoded are processed in units such as CUs. Each coding unit is encoded using either an intra mode or an inter mode. When a coding unit is encoded in the intra mode, it performs intra prediction (260). In the inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of the intra mode or the inter mode to use for encoding the coding unit, and indicates the intra / inter decision by a prediction mode flag. The prediction residual is calculated by subtracting (210) the predicted block from the original image block.

[0037] Then the prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy coded (245) to output a bitstream. As a non-limiting example, context-based adaptive binary arithmetic coding (CABAC) can be used to code the syntax elements into the bitstream.

[0038] The encoder may also skip the transform and apply quantization directly to the untransformed residual signal, for example, on a 4x4 TU basis. The encoder may also bypass both the transform and quantization, i.e., the residual is directly coded without applying the transform or quantization process. In direct PCM coding, no prediction is applied, and the coding unit samples are directly coded into the bitstream.

[0039] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse-transformed (250) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (255) to reconstruct the image block. For example, a loop filter (265) is applied to the reconstructed picture to perform deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored at the reference picture buffer (280).

[0040] Figure 3 The block diagram of an example video decoder 300 such as an HEVC decoder is illustrated. In decoder 300, the bitstream is decoded by decoder elements as described below. Video decoder 300 generally performs a decoding pass that is the reverse of the encoding pass Figure 2 described above, and it performs video decoding as part of encoding video data. Figure 3 A decoder that improves on the HEVC standard or a decoder that employs techniques similar to HEVC (such as a VVC decoder) can also be illustrated.

[0041] In particular, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, picture segmentation information, and other encoding / decoding information. The picture segmentation information indicates how the picture is segmented, for example, the size of the CTU, and the way the CTU is split into CUs and, when applicable, possibly into PUs. The decoder can thus partition (335) the picture into, for example, CTUs according to the decoded picture segmentation information, and partition each CTU into CUs. The transform coefficients are dequantized (340) and inverse-transformed (350) to decode the prediction residuals.

[0042] The decoded prediction residuals and the predicted blocks are combined (355) to reconstruct the image blocks. The predicted blocks can be obtained (370) from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored at the reference picture buffer (380).

[0043] The decoded picture can be further post-processed (385) after decoding, for example, inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or performing the inverse of the remapping process performed in the pre-coding process (201). The post-processing after decoding can use the metadata derived in the pre-coding process and signaled in the bitstream.

[0044] The new video compression tools in VVC include the codec tree unit representation in the compression domain, which can represent picture data in a more flexible way. In VVC, a quadtree with a nested multi-type tree (MTT) using binary and ternary split segmentation structures replaces the concept of multiple segmentation unit types, i.e., except for a few special cases, VVC removes the separation of the CU, PU, and TU concepts. In the VVC codec tree structure, a CU can have a square or rectangular shape. The coding tree unit (CTU) is first segmented by a quadtree structure. Then the quadtree leaf nodes can be further segmented by a multi-type tree structure.

[0045] In particular, the tree decomposition of the CTU is performed in different stages: first, the CTU is split in a quadtree manner, and then each quadtree leaf can be further divided in a binary or ternary manner. This is shown on the Figure 6 right side, where the solid lines represent the quadtree decomposition stage, and the dashed lines represent the binary decomposition in the spatially embedded quadtree leaves. In the intra slice, when the dual-tree mode is activated, the luminance and chrominance block segmentation structures are separate and are determined independently.

[0046] As Figure 7As shown in the figure, there are four splitting types in the multi-type tree structure: vertical binary splitting (VER), horizontal binary splitting (HOR), vertical triple splitting (VER_TRIPLE), and horizontal triple splitting (HOR_TRIPLE). The HOR_TRIPLE or VER_TRIPLE splitting (horizontal or vertical ternary tree splitting mode) involves dividing a coding unit (CU) into 3 sub-coding units (sub-CUs), whose sizes are respectively equal to 1 / 4, 1 / 2, and 1 / 4 of the size of the parent CU in the considered spatial partitioning direction.

[0047] The leaf nodes of the multi-type tree are called coding units (CUs), and except for a few special cases, this splitting is used for prediction and transformation processing without further splitting. Exceptions occur under the following conditions:

[0048] - If the width or height of the CU is greater than 64, then the CU is tiled into transform units (TUs) with sizes equal to the maximum supported transform size. Usually, the maximum transform size can be equal to 64.

[0049] - If the intra-CU is coded in the ISP (intra-sub-partitioning) mode, then the CU is split into 2 or 4 transform units, depending on the type of the ISP mode used and the shape of the CU.

[0050] - If the inter-CU is coded in the SBT (sub-block transform) mode, then the CU is split into 2 transform units, and one of the resulting TUs must have residual data equal to zero.

[0051] - If the inter-CU is coded in the triangle prediction merge (TPM, Triangle Prediction Merge) mode, then the CU consists of 2 triangle prediction units, and each PU is assigned its own motion data.

[0052] According to VVC Draft 6, the syntax related to splitting is coded in the sequence parameter set (SPS). If the partition_constraints_override_enabled_flag is true, then the syntax related to splitting can be overridden in the slice header (SH). The SPS syntax and SH syntax used in VVC Draft 6 are shown in Tables 1 and 2.

[0053] Table 1. Sequence Parameter Set Syntax in VVC Draft 6

[0054]

[0055]

[0056] The semantic descriptions of some SPS syntax elements are as follows:

[0057] log2_ctu_size_minus5 plus 5 specifies the luma coding tree block size for each CTU. The value of log2_ctu_size_minus5 less than or equal to 2 is a requirement for bitstream consistency.

[0058] log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size.

[0059] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, IbcBufWidthY, IbcBufWidthC, and Vsize are derived as follows:

[0060] CtbLog2SizeY = log2_ctu_size_minus5 + 5 (7-15)

[0061] CtbSizeY = 1 << CtbLog2SizeY (7-16)

[0062] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-17)

[0063] MinCbSizeY = 1 << MinCbLog2SizeY (7-18)

[0064] IbcBufWidthY = 128 * 128 / CtbSizeY (7-19)

[0065] IbcBufWidthC = IbcBufWidthY / SubWidthC (7-20)

[0066] VSize = Min(64, CtbSizeY) (7-21)

[0067] The variables CtbWidthC and CtbHeightC, which specify the width and height of the array for each chroma CTB respectively, are derived as follows:

[0068] - If chroma_format_idc is equal to 0 (monochrome) or separate_colour_plane_flag is equal to 1, then both CtbWidthC and CtbHeightC are equal to 0.

[0069] - Otherwise, CtbWidthC and CtbHeightC are derived as follows:

[0070] CtbWidthC = CtbSizeY / SubWidthC (7-22)

[0071] CtbHeightC = CtbSizeY / SubHeightC (7-23)

[0072] sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size among the luma samples of the luma leaf blocks resulting from the quadtree splitting of the CTU in the reference SPS and the base-2 logarithm of the minimum decoded block size among the luma samples of the luma CUs in a slice with slice_type equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_min_qt_min_cb_luma present in the slice header of the slice. The value of sps_log2_diff_min_qt_min_cb_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY. The base-2 logarithm of the minimum size among the luma samples of the luma leaf blocks resulting from the quadtree splitting of the CTU is derived as follows: MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY (7-24)

[0073] sps_log2_diff_min_qt_min_cb_inter_slice specifies the default difference between the log2 of the minimum size of the luma samples in the luma leaf blocks resulting from the quadtree splitting of a CTU in the reference SPS and the log2 of the minimum luma coding block size of the luma CUs in a slice where slice_type is equal to 0 (B) or 1 (P). When partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_min_qt_min_cb_luma which is present in the slice header of the slice and refers to the SPS. The value of sps_log2_diff_min_qt_min_cb_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY. The log2 of the minimum size of the luma samples in the luma leaf blocks resulting from the quadtree splitting of a CTU is derived as follows: MinQtLog2SizeInterY = sps_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY (7 - 25)

[0074] sps_max_mtt_hierarchy_depth_inter_slice specifies the default maximum hierarchy depth for the coding units resulting from the multi-type tree splitting of the quadtree leaves in a slice where slice_type is equal to 0 (B) or 1 (P) in the reference SPS. When partition_constraints_override_flag is equal to 1, the default maximum hierarchy depth can be overridden by slice_max_mtt_hierarchy_depth_luma which is present in the slice header of the slice and refers to the SPS. The value of sps_max_mtt_hierarchy_depth_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY.

[0075] sps_max_mtt_hierarchy_depth_intra_slice_luma specifies the default maximum hierarchy depth for the coding units generated by the multi-type tree splitting of the quadtree leaves in the slices with slice_type equal to 2 (I) with reference to the SPS. When partition_constraints_override_flag is equal to 1, the default maximum hierarchy depth may be overridden by slice_max_mtt_hierarchy_depth_luma present in the slice header of the slice with reference to the SPS. The value of sps_max_mtt_hierarchy_depth_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY.

[0076] sps_log2_diff_max_bt_min_qt_intra_slice_luma specifies the default difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in the luma coding blocks that can use binary splitting with reference to the SPS and the minimum size (width or height) of the luma samples in the luma leaf blocks generated by the quadtree splitting of the CTUs in the slices with slice_type equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default difference may be overridden by slice_log2_diff_max_bt_min_qt_luma present in the slice header of the slice with reference to the SPS. The value of sps_log2_diff_max_bt_min_qt_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeIntraY. When sps_log2_diff_max_bt_min_qt_intra_slice_luma is not present, the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma is inferred to be equal to 0.

[0077] sps_log2_diff_max_tt_min_qt_intra_slice_luma specifies the default difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in a luma coding block that can use ternary splitting in the reference SPS and the minimum size (width or height) of the luma samples in a luma leaf block resulting from the quadtree splitting of a CTU in a slice with slice_type equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default difference can be overridden with slice_log2_diff_max_tt_min_qt_luma present in the slice header of the slice that references the SPS. The value of sps_log2_diff_max_tt_min_qt_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeIntraY. When sps_log2_diff_max_tt_min_qt_intra_slice_luma is not present, the value of sps_log2_diff_max_tt_min_qt_intra_slice_luma is inferred to be equal to 0.

[0078] sps_log2_diff_max_bt_min_qt_inter_slice specifies the default difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in a luma coding block that can use ternary splitting in the reference SPS and the minimum size (width or height) of the luma samples in a luma leaf block resulting from the quadtree splitting of a CTU in a slice with slice_type equal to 0 (B) or 1 (P). When partition_constraints_override_flag is equal to 1, the default difference can be overridden with slice_log2_diff_max_bt_min_qt_luma present in the slice header of the slice. The value of sps_log2_diff_max_bt_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeInterY, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeInterY. When sps_log2_diff_max_bt_min_qt_inter_slice is not present, the value of sps_log2_diff_max_bt_min_qt_inter_slice is inferred to be equal to 0.

[0079] sps_log2_diff_max_tt_min_qt_inter_slice specifies the default difference between the log2 of the maximum size (width or height) in the luma samples of a luma coded block that can use ternary splitting in the reference SPS and the minimum size (width or height) in the luma samples of a luma leaf block resulting from the quadtree splitting of a CTU in a slice where slice_type is equal to 0 (B) or 1 (P). When partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_max_tt_min_qt_luma present in the slice header of the slice with reference to the SPS. The value of sps_log2_diff_max_tt_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeInterY, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeInterY. When sps_log2_diff_max_tt_min_qt_inter_slice is not present, the value of sps_log2_diff_max_tt_min_qt_inter_slice is inferred to be equal to 0.

[0080] sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the default difference between the log2 of the minimum size in the luma samples of the chroma leaf blocks resulting from the quadtree splitting of the chroma CTUs with treeType equal to DUAL_TREE_CHROMA in the reference SPS and the log2 of the minimum decoded block size in the luma samples of the chroma CUs with treeType equal to DUAL_TREE_CHROMA in a slice with slice_type equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_min_qt_min_cb_chroma present in the slice header of the slice with reference to the SPS. The value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma shall be in the range from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY. When not present, the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to 0. The log2 of the minimum size in the luma samples of the chroma leaf blocks resulting from the quadtree splitting of the CTUs with treeType equal to DUAL_TREE_CHROMA is derived as follows:

[0081] MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY(7 - 26)

[0082] sps_max_mtt_hierarchy_depth_intra_slice_chroma specifies the default maximum hierarchical depth of chroma coding units generated by multi-type tree splitting of chroma quadtree leaves with treeType equal to DUAL_TREE_CHROMA in slices where slice_type is equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default maximum hierarchical depth can be overridden by slice_max_mtt_hierarchy_depth_chroma present in the slice header of the slice as per the reference SPS. The value of sps_max_mtt_hierarchy_depth_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY. When not present, the value of sps_max_mtt_hierarchy_depth_intra_slice_chroma is inferred to be equal to 0.

[0083] sps_log2_diff_max_bt_min_qt_intra_slice_chroma specifies the default difference between the base-2 logarithm of the maximum size (width or height) of a chroma coding block that can be split using binary splitting as per the reference SPS and the minimum size (width or height) in the luma samples of the chroma leaf blocks resulting from the quadtree splitting of chroma CTUs with treeType equal to DUAL_TREE_CHROMA in slices where slice_type is equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_max_bt_min_qt_chroma present in the slice header of the slice as per the reference SPS. The value of sps_log2_diff_max_bt_min_qt_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraC, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeIntraC. When sps_log2_diff_max_bt_min_qt_intra_slice_chroma is not present, the value of sps_log2_diff_max_bt_min_qt_intra_slice_chroma is inferred to be equal to 0.

[0084] sps_log2_diff_max_tt_min_qt_intra_slice_chroma specifies the default difference between the base-2 logarithm of the maximum size (width or height) of the luma samples of the chroma coding blocks that can be split using ternary splitting in the reference SPS and the minimum size (width or height) of the luma samples of the chroma leaf blocks resulting from the quadtree splitting of the chroma CTUs with treeType equal to DUAL_TREE_CHROMA in slices where the slice type is equal to 2 (I). When partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_max_tt_min_qt_chroma present in the slice header of the slice with reference to the SPS. The value of sps_log2_diff_max_tt_min_qt_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraC, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeIntraC. When sps_log2_diff_max_tt_min_qt_intra_slice_chroma is not present, the value of sps_log2_diff_max_tt_min_qt_intra_slice_chroma is inferred to be equal to 0.

[0085] sps_max_luma_transform_size_64_flag being equal to 1 specifies that the maximum transform size in luma samples is equal to 64. sps_max_luma_transform_size_64_flag being equal to 0 specifies that the maximum transform size in luma samples is equal to 32.

[0086] When CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag shall be equal to 0.

[0087] The variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, and MaxTbSizeY are derived as follows:

[0088] MinTbLog2SizeY = 2 (7 - 27)

[0089] MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag? 6 : 5 (7 - 28)

[0090] MinTbSizeY = 1 << MinTbLog2SizeY (7-29)

[0091] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-30)

[0092] Table 2. Syntax of the slice header in VVC Draft 6

[0093]

[0094]

[0095] Next, the maximum allowed hierarchical depth (maxMTT hierarchy depth) of the multi-type tree split from the quadtree leaf is described using the syntax element sps_max_mtt_hierarchy_depth_inter_slice of the luma color component for the inter slice as an example. However, this principle can also be applied to the intra slice or the chroma color component (e.g., the syntax elements sps_max_mtt_hierarchy_depth_intra_slice_luma and sps_log2_diff_min_qt_min_cb_intra_slice_chroma).

[0096] In VVC Draft 6, the value of sps_max_mtt_hierarchy_depth_inter_slice should be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, including 0 and CtbLog2SizeY - MinCbLog2SizeY. Generally, MinCbLog2SizeY is equal to 2, corresponding to a 4x4 block, and CtbLog2SizeY is equal to 7, corresponding to a 128x128 CTU. In this configuration, max_mtt_hierarchy_depth should be in the range of 0 to 5. If the minimum QT size (i.e., 1 << MinQtLog2SizeInterY) is equal to 32 and only BT is used, then the minimum block size that can be achieved is 4x8, and it is 8x4 when the BT split depth is 5, as Figure 8 shown. That is, the encoder lacks the flexibility to support the 4x4 block size in this example. More generally, the encoder will only support a subset of the split patterns specified by VVC.

[0097] To highlight the lack of flexibility, we use a configuration where the CTU size is equal to 32, the minimum Cb size is 8, and the minimum QT size is 32. This means that only BT and TT can be used. In this configuration, max_mtt_hierarchy_depth is set to 2. If only BT is used, the minimum block size that can be achieved is 16x16 or 8x32; if only TT is used, some regions can only be split into 16x16, as Figure 9 and 10 shown in

[0098] To better illustrate the syntax elements related to splitting, Figure 11 Figure shows the possible values of CTU size, minimum decoded block size, maximum transform size, minimum QT size, and maximum BT / TT size. The actual values used for common test conditions (CTC) in VTM6.0 are also shown. We can see that the minimum decoded block size (min_luma_coding_block_size) is only bounded by the CTU size.

[0099] As mentioned above, due to the combined use of the maximum block size allowing BT or TT splitting and the maximum multi-type tree hierarchy depth, the way the codec tree depth is canonically bounded may make the VVC compression scheme suboptimal in terms of coding efficiency given fixed maximum and minimum block sizes. Another issue is that the log2_min_luma_coding_block_size_minus2 syntax element is independent of other syntax elements and has no maximum value. This can cause the encoder to generate a VVC bitstream with a value of log2_min_luma_coding_block_size_minus2 that is higher than the maximum block size, making things inconsistent.

[0100] To address the lack of flexibility of VVC for the maximum split depth, the maximum binary tree (BT) size, maximum ternary tree (TT) size, and maximum hierarchy depth information are defined separately for intra slices and inter slices. In the case of the dual tree, the maximum BT size / maximum TT size and maximum MTT depth are also defined for the chroma tree of intra slices. The proposed method can increase the set of reachable codec tree nodes and leaves under the constraints of pre-fixed maximum and minimum decoded block sizes, and thus improve the compression efficiency by allowing a higher degree of flexibility in the allowed codec tree representation.

[0101] In one embodiment, the maximum values of sps_max_mtt_hierarchy_depth_inter_slice and sps_max_mtt_hierarchy_depth_intra_slice_luma are increased. Hereinafter, for the sake of convenience of representation, max_mtt_hierarchy_depth is used as a general term to refer to the syntax elements related to the maximum MTT hierarchy depth, such as sps_max_mtt_hierarchy_depth_inter_slice and sps_max_mtt_hierarchy_depth_intra_slice_luma. In another embodiment, max_mtt_hierarchy_depth is described for all available QT depths in order to provide more flexibility in describing the split tree. In another embodiment, max_mtt_hierarchy_depth is different when QT splitting is available for a given depth and when QT splitting is not available for a given depth.

[0102] In another embodiment, the syntax element sps_max_luma_transform_size_64_flag is moved before log2_min_luma_coding_block_size_minus2, and the maximum value of the coding block size (log2_min_luma_coding_block_size_minus2) is defined according to the maximum transform size (sps_max_luma_transform_size_64_flag).

[0103] Hereinafter, different embodiments will be described in more detail.

[0104] Maximum hierarchical depth in VVC Draft 6

[0105] In this embodiment, we propose to increase the allowed number of consecutive splits (i.e., split depth or MTT hierarchy depth) to twice the difference between the CTU size and the minimum size of the CU. With this increase, in the worst case where QT (the minimum QT is defined as the CTU size) is not used and BT and TT are the only splits used, we can reach the minimum CU size.

[0106] The changes in the specification text are underlined as follows:

[0107] sps_max_mtt_hierarchy_depth_inter_slice specifies the default maximum hierarchy depth for the coding units generated by the multi-type tree splitting of the quadtree leaves in the slice where slice_type is equal to 0 (B) or 1 (P) with reference to the SPS. When partition_constraints_override_flag is equal to 1, the default maximum hierarchy depth can be overridden by slice_max_mtt_hierarchy_depth_luma which is present in the slice header of the slice with reference to the SPS. The value of sps_max_mtt_hierarchy_depth_inter_slice shall be in the range of 0 to 2* (CtbLog2SizeY - MinCbLog2SizeY), inclusive of 0 and 2* (CtbLog2SizeY - MinCbLog2SizeY).

[0108] slice_max_mtt_hierarchy_depth_luma specifies the maximum hierarchy depth for the coding units generated by the multi-type tree splitting of the quadtree leaves in the current slice. The value of slice_max_mtt_hierarchy_depth_luma shall be in the range of 0 to 2* (CtbLog2SizeY - MinCbLog2SizeY), inclusive of 0 and 2* (CtbLog2SizeY - MinCbLog2SizeY). When it is not present, the value of slice_max_mtt_hierarchy_depth_luma is inferred as follows:

[0109] If slice_type is equal to 2 (I), then the inferred value of slice_max_mtt_hierarchy_depth_luma is equal to sps_max_mtt_hierarchy_depth_intra_slice_luma.

[0110] Otherwise (slice_type is equal to 0 (B) or 1 (P)), then the inferred value of slice_max_mtt_hierarchy_depth_luma is equal to sps_max_mtt_hierarchy_depth_inter_slice.

[0111] The reason for the value 2*(CtbLog2SizeY - MinCbLog2SizeY) instead of (CtbLog2SizeY - MinCbLog2SizeY) specified in VVC Draft 6 is that it allows reaching the minimum allowable block size regardless of which of the QT, BT, or TT split types is used. In particular, it can only be achieved by binary tree (BT) splitting, which is not the case with the current specification constraints in VVC Draft 6. Thus, the advantage of the proposed method is that it maximizes the compression performance, which can be achieved with a VVC encoder under the constraints of the maximum and minimum decoded block sizes.

[0112] Adaptive maximum MTT depth

[0113] In this embodiment, max_mtt_hierarchy_depth is specified canonically for all QT levels to provide more flexibility in the split tree. The advantage of specifying the maximum multi-type tree depth associated with each level that a quadtree leaf can reach is that it enables a fine-grained way of allocating the encoder rate-distortion combination. In fact, the rate-distortion search for the optimal codec tree implies a large combination of the encoder search space. Thus, it makes sense to fine-tune the combination of the multi-type codec tree search in order to achieve a good trade-off between the encoder's search over all combinations and the compression performance. Allocating the maximum mtt hierarchy depth for each quadtree level provides a way to obtain a better trade-off between the RD search combination and the compression performance. Thus, a higher degree of flexibility in the canonical signaling of the maximum mtt codec tree depth for each quadtree level may lead to an encoder complexity / compression efficiency trade-off that cannot be achieved with the current VVC Draft 6 specification.

[0114] Table 3. SPS specification with one maximum BT depth and one maximum TT depth signaled for each quadtree leaf, where binary tree or ternary tree splitting is allowed.

[0115]

[0116]

[0117] Here, we first signal a syntax element that indicates the maximum size relative to the minimum quadtree size at which the mtt tree can start, sps_log2_diff_max_mtt_size_min_qt_size_inter_slice_luma. The maximum size at which the mtt can start is derived as:

[0118] maxMTTsize = 1 << (sps_log2_diff_max_mtt_size_min_qt_size_inter_slice_luma + MinQtLog2SizeInterY).

[0119] The maximum value of sps_log2_diff_max_mtt_size_min_qt_size_inter_slice_luma is CtbLog2SizeY - MinQtLog2SizeInterY.

[0120] If sps_log2_diff_max_mtt_size_min_qt_size_inter_slice_luma is equal to 0, it means that MTT splitting is not allowed. If it is not zero, then for each QT level for which MTT is allowed, the maximum depths for BT and TT are signaled. For each level, the allowed range of sps_max_bt_depth_inter_slice_luma[i] is from 0 to 2x(i + MinQtLog2SizeInterY - MinCbLog2SizeY). The last value corresponds to twice the difference between the log2 of the current considered quadtree leaf node size and the log2 of the minimum decoded block size. It ensures that the minimum decoded block size can be reached by means of binary splitting.

[0121] In one example, we can define the splitting tree in VVC as follows:

[0122] Table 4. Maximum MTT depth for each size

[0123] CU size Allow QT split? Maximum MTT depth 128 Yes 0 64 Yes 0 32 Yes 2 16 Yes 4 8 No 2 4 No 0

[0124] - By signaling the following values:

[0125] - log2_ctu_size_minus5 = 2 (CTU size is 128)

[0126] - log2_min_luma_coding_block_size_minus2 = 0 (minimum CU size is 4)

[0127] - sps_log2_diff_min_qt_min_cb_intra_slice_luma = 1 (the minimum CU size resulting from QT splitting is 8)

[0128] -sps_log2_diff_max_mtt_size_min_qt_size = 2 (The maximum CU size for MTT splitting is 32)

[0129] -sps_max_bt_depth_inter_slice_luma[0] = 2 (2BT splitting allowed for 8x8 CU)

[0130] -sps_max_bt_depth_inter_slice_luma[l] = 4 (4BT splitting allowed for 16x16 CU)

[0131] -sps_max_bt_depth_inter_slice_luma[2] = 2 (2BT splitting allowed for 32x32 CU)

[0132] The same principle applies to intra_slice_luma and intra_slice_chroma.

[0133] In another embodiment, the maximum allowed hierarchical depth is specified canonically for multi-type tree splitting of quadtree leaves, which is associated with each quadtree level that allows BT or TT splitting of quadtree leaves. Basically, the coding unit corresponding to the quadtree leaf for which MTT splitting is performed must be a square CU, the log2 of whose size is included between MinQtLog2SizeIntraY and (MinQtLog2SizeIntraY + log2_diff_max_mtt_min_qt_intra_slice_luma).

[0134] Here, we define log2_diff_max_mtt_min_qt_intra_slice_luma by the maximum value between the signaled values sps_log2_diff_max_bt_min_qt_intra_slice_luma and sps_log2_diff_max_tt_min_qt_intra_slice_luma. For each block size whose log2 is included between MinQtLog2SizeIntraY and (MinQtLog2SizeIntraY + log2_diff_max_mtt_min_qt_intra_slice_luma), the maximum multi-type tree depth is signaled.

[0135] The value sps_max_mtt_hierarchy_depth_intra_slice_luma[i] specifies the maximum hierarchical depth of the multi-type tree for splitting the CU corresponding to the quadtree leaf.

[0136] Table 5

[0137]

[0138]

[0139] In the same way as log2_diff_max_mtt_min_qt_intra_slice_luma, the parameter log2_diff_max_mtt_min_qt_inter_slice is defined as the maximum of the signaled values sps_log2_diff_max_bt_min_qt_inter_slice_luma and sps_log2_diff_max_tt_min_qt_inter_slice_luma. For each block size whose log2 is included between MinQtLog2SizeInterY and (MinQtLog2SizeInterY + log2_diff_max_mtt_min_qt_inter_slice), the maximum multi-type tree depth is signaled.

[0140] In the same way as log2_diff_max_mtt_min_qt_intra_slice_luma, the parameter log2_diff_max_mtt_min_qt_intra_slice_chroma is defined as the maximum of the signaled values sps_log2_diff_max_bt_min_qt_intra_slice_chroma and sps_log2_diff_max_tt_min_qt_intra_slice_chroma. For each block size whose log2 is included between MinQtLog2SizeIntraC and (MinQtLog2SizeIntraC + log2_diff_max_mtt_min_qt_intra_slice), the maximum multi-type tree depth is signaled.

[0141] According to the variant of the embodiment in Table 5, the syntax elements sps_max_mtt_hierarchy_depth_intra_slice_luma_present_flag, sps_max_mtt_hierarchy_depth_inter_slice_luma_present_flag, and sps_max_mtt_hierarchy_depth_intra_slice_chroma_present_flag are not included in the SPS specification. This variant can take the form of Table 6 below.

[0142] Table 6

[0143]

[0144]

[0145] In another variant, the encoding and decoding of the maximum MTT hierarchy depth is indexed by the quadtree depth instead of log2 of the quadtree leaf size. This can be slightly different from the form of Table 7.

[0146] In the variant of Table 7, the quantity strat_qt_depth_inter_slice is defined as:

[0147] start_qt_depth_inter_slice = CtbLog2 Size Y - max(MaxBtLog2 Size Y, MaxTtLog2SizeY)

[0148] where:

[0149] MaxBtLog2SizeY = (MinQtLog2SizeInterY + sps_log2_diff_max_bt_min_qt_inter_slice)

[0150] MaxTtLog2SizeY = (MinQtLog2SizeY + sps_log2_diff_max_bt_min_qt_inter_slice)

[0151] Moreover, max_qt_depth_inter_slice is defined as:

[0152] max_qt_depth_inter_slice = CtbLog2SizeY - MinQtLog2SizeInterY

[0153] The quantities start_qt_depth_intra_slice_luma, max_qt_depth_intra_slice_luma, start_qt_depth_intra_slice_chroma, and max_qt_depth_intra_slice_chroma are defined in a similar manner to start_qt_depth_inter_slice and max_qt_depth_inter_slice, but for the case of intra-slice luma and intra-slice chroma, respectively (in the dual-tree case).

[0154] Table 7

[0155]

[0156]

[0157] In yet another embodiment, any of the aforementioned variants proposed herein is also used for the encoding and decoding of the slice header. In fact, in the VVC draft 6 specification, the codec tree parameters signaled in the SPS can be overwritten in the slice header, for example, according to the syntax table given in Table 2.

[0158] Note that, on the encoder side, an upper limit on the maximum MTT hierarchy depth for encoding and decoding can be determined based on the depth difference between the log2 of the size of the quadtree tree child node under consideration and the log2 of the size of the minimum decoded block size. This can take the following form. Given a depth value i (index in one of the loops of Table 7), the upper limit of the maximum mtt hierarchy depth to be encoded can be the value 2*(CtbLog2SizeY-i-MinCbLog2SizeY), where 2*(CtbLog2SizeY-i-MinCbLog2SizeY) represents the number of splits required to achieve the minimum block size of width and height using only BT splits. In fact, from a given block to the split, exactly 2 symmetrical binary split stages are required to obtain a block of half the size in width and height. Clipping the value by the upper limit 2*(CtbLog2SizeY-i-MinCbLog2SizeY) can be beneficial in terms of bit savings in the encoding and decoding of SPS and slice headers.

[0159] Also, note that the maximum mtt hierarchy depth is usually allowed to range from 0 to a value of 2*(CtbLog2SizeY-i-MinCbLog2SizeY) to ensure that the minimum decoding block size can be achieved with BT splitting.

[0160] Finally, during the CU-level decoder-side parsing process of CU split information, high-level signaling of the maximum MTT hierarchical depth for each quadtree level is considered.

[0161] To this end, when the decoder evaluates whether the current node of the coding tree for a given CTU allows a given binary or ternary split mode, it compares the multi-type tree depth of the current coding tree node with the maximum multi-type tree depth at the quadtree depth associated with the current coding tree node. If the multi-type tree depth is higher than or equal to the maximum allowed multi-type tree depth at the considered quadtree depth, then all binary and ternary split modes are prohibited for the current tree node. Thus, the decoder infers that the split mode of the considered tree node is different from any binary or ternary split mode.

[0162] The difference from the parsing process of the split information in VVC Draft 6 is that in VVC Draft 6, the maximum allowed multi-type tree depth of the current tree node does not depend on the quadtree depth associated with the considered coding tree node. In the case of the intra slice type, it only depends on the slice type and the component type.

[0163] Binary / Ternary split enabled

[0164] In this embodiment, two new flags are introduced in the SPS to signal whether BT and TT splits are used. Then, if at least one of the two splits is used, the sps_max_mtt_hierarchy_depth syntax element is coded as shown in Table 8. In a variant of the embodiment of Table 8, sps_bt_enabled_flag is defined as:

[0165] sps_bt_enabled_flag being equal to 1 specifies that binary split is allowed during the coding block splitting process.

[0166] Table 8

[0167]

[0168]

[0169] In another variant, BT and TT splits are enabled or disabled differently for intra luma, inter chroma, and intra chroma to allow for greater flexibility. In the embodiment shown in Table 9, sps_bt_enabled_flag is defined as:

[0170] sps_bt_enabled_flag being equal to 1 specifies that binary split is allowed during the coding block splitting of the quadtree leaves in slices where slice_type is equal to 2 (I) with reference to the SPS.

[0171] Table 9

[0172]

[0173]

[0174] In another variant, the sps_log2_diff_max_bt_min_qt or sps_log2_diff_max_tt_min_qt syntax elements are used to disable BT or TT. In the current VVC draft 6, BT and TT splitting are enabled together by setting a value greater than 0 to the sps_max_mtt_hierarchy_depth syntax element. In the variant of Table 9, sps_log2_diff_max_bt_min_qt and sps_log2_diff_max_bt_min_qt are first defined, and then sps_max_mtt_hierarchy_depth is conditionally parsed. Changing sps_log2_diff_max_bt_min_qt (sps_log2_diff_max_tt_min_qt) to sps_log2_diff_max_bt_min_qt_plus_one (sps_log2_diff_max_tt_min_qt_plus_one), a value of 0 indicates that BT / TT is disabled. In fact, in this case, the maximum BT size is strictly lower than the minimum QT size, so it is never used. The sps_log2_diff_max_bt_min_qt_plus_one (sps_log2_diff_max_tt_min_qt_plus_one) syntax element is defined as: sps_log2_diff_max_bt_min_qt_plus_one_intra_slice_luma specifies the default difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in the luma coding block that can be split using binary splitting in the reference SPS and the minimum size (width or height) of the luma samples in the luma leaf block resulting from the quadtree splitting of the CTU in the slice with slice_type equal to 2 (I) plus one. When the partition_constraints_override_flag is equal to 1, the default difference can be overridden by slice_log2_diff_max_bt_min_qt_luma present in the slice header of the slice with reference to the SPS. The value of sps_log2_diff_max_bt_min_qt_minus_one_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY + 1, inclusive of 0 and CtbLog2SizeY - MinQtLog2SizeIntraY + 1.When sps_log2_diff_max_bt_min_qt_minus_intra_slice_luma does not exist, the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma is inferred to be equal to 0.

[0175] Table 10

[0176]

[0177]

[0178] Simplified adaptive maximum mtt depth

[0179] In another variant, only two levels of the maximum MTT depth are signaled (instead of 1 for each QT level as in the previous embodiments): the maximum MTT depth when QT splitting is allowed, and the maximum MTT depth when no more QT splitting is allowed.

[0180] Table 11

[0181]

[0182]

[0183] In this embodiment, we first signal the syntax element sps_max_mtt_hierarchy_depth_before_minqt_inter_slice.

[0184] sps_max_mtt_hierarchy_depth_before_minqt_inter_slice specifies the default maximum hierarchical depth of the coding units generated by the multi-type tree splitting of this quadtree leaf when the quadtree leaf size in the slices where slice_type is equal to 0 (B) or 1 (P) of the reference SPS is not equal to (strictly greater than) MinQtLog2SizeInterY. When partition_constraints_override_flag is equal to 1, the default maximum hierarchical depth can be overridden by slice_max_mtt_hierarchy_depth_before_minqt_luma present in the slice header of the slice in the strip with reference to the SPS. The value of sps_max_mtt_hierarchy_depth_before_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY.

[0185] If sps_log2_diff_min_qt_min_cb_inter_slice is not equal to 0, it means that the QT tree will stop before reaching the minimum decoded block size, so we need to use more binary / ternary splits to reach the minimum decoded block size. Therefore, this sps_log2_diff_min_qt_min_cb_inter_slice value is a condition for parsing the sps_log2_diff_max_hierarchy_depth_after_minqt_intra_slice_luma syntax element.

[0186] sps_max_mtt_hierarchy_depth_after_minqt_inter_slice specifies the default maximum hierarchy depth of the coding unit generated by the multi-type tree split of the quadtree leaf when the quadtree leaf size is equal to MinQtLog2SizeInterY in a slice where slice_type is equal to 0 (B) or 1 (P) in the reference SPS. When partition_constraints_override_flag is equal to 1, the default maximum hierarchy depth can be overridden by slice_max_mtt_hierarchy_depth_after_minqt_luma present in the slice header of the slice in the SPS. The value of sps_max_mtt_hierarchy_depth_after_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of 0 and CtbLog2SizeY - MinCbLog2SizeY.

[0187] The same principle applies to intra_slice_luma and intra_slice_chroma.

[0188] The previous embodiments can be used alone or in combination. For example, the embodiment of doubling the hierarchy depth as described above is combined with the embodiment described in Table 11. This generally means that the maximum mtt hierarchy depth is doubled compared to its maximum allowed value in VVC Draft 6. This takes the following form.

[0189] The value range of sps_max_mtt_hierarchy_depth_before_minqt_inter_sliced is from 0 (no splitting allowed) to 2 * (CtbLog2SizeY - MinCbLog2SizeY).

[0190] The value range of sps_max_mtt_hierarchy_depth_before_minqt_intra_slice_luma is from 0 (no splitting allowed) to 2 * (CtbLog2SizeY - MinCbLog2SizeY).

[0191] The value range of sps_max_mtt_hierarchy_depth_after_minqt_inter_slice is changed from 0 (no splitting allowed) to 2 * (MinQtLog2 Size Y - MinCbLog2 Size Y).

[0192] The value range of sps_max_mtt_hierarchy_depth_after_minqt_intra_slice_luma is changed from 0 (no splitting allowed) to 2 * (MinQtLog2SizeY - MinCbLog2SizeY).

[0193] The value range of sps_max_mtt_hierarchy_depth_before_minqt_intra_slice_chroma is changed from 0 (no splitting allowed) to 2 * (MinQtLog2SizeY - MinCbLog2SizeY).

[0194] The value range of sps_max_mtt_hierarchy_depth_after_minqt_intra_slice_chroma is changed from 0 (no splitting allowed) to 2 * (MinQtLog2SizeY - MinCbLog2SizeY).

[0195] In one example, we can define a splitting tree as shown in Table 12.

[0196] Table 12. Maximum MTT depth for each size when QT is enabled and disabled

[0197] CU size Allow QT split? Maximum MTT depth 128,64,32 Yes 2 16,8,4 No 4

[0198] Maximum value of log2_min_luma_coding_block_size_minus2 syntax element

[0199] In this embodiment, the sps_max_luma_transform_size_64_flag syntax element is moved after log2_ctu_size_minus5 and before log2_min_luma_coding_block_size_minus2, as shown in Table 13.

[0200] Table 13. Syntax Table for Signaling the Maximum Value of log2_min_luma_coding_block_size_minus2

[0201]

[0202]

[0203] When sps_max_luma_transform_size_64_flag equals 1, it specifies that the maximum transform size in the luma samples is equal to 64. When sps_max_luma_transform_size_64_flag equals 0, it specifies that the maximum transform size in the luma samples is equal to 32.

[0204] When CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag shall be equal to 0.

[0205] The variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, and MaxTbSizeY are derived as follows:

[0206] MinTbLog2SizeY = 2 (7 - 27)

[0207] MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag? 6 : 5 (7 - 28)

[0208] MinTbSizeY = 1 << MinTbLog2SizeY (7 - 29)

[0209] MaxTbSizeY = 1 << MaxTbLog2SizeY

[0210] log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size.

[0211] In VVC Draft 6, the upper limit of the log2_min_luma_coding_block_size_minus2 syntax element is not specified.

[0212] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, IbcBufWidthY, IbcBufWidthC, and Vsize are derived as follows:

[0213] CtbLog2SizeY = log2_ctu_size_minus5 + 5 (7-15)

[0214] CtbSizeY = 1 << CtbLog2SizeY (7-16)

[0215] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-17)

[0216] MinCbSizeY = 1 << MinCbLog2SizeY (7-18)

[0217] IbcBufWidthY = 128 * 128 / CtbSizeY (7-19)

[0218] IbcBufWidthC = IbcBufWidthY / SubWidthC (7-20)

[0219] VSize = Min(64, CtbSizeY) (7-21)

[0220] Variables CtbWidthC and CtbHeightC, which respectively specify the width and height of the array of each chroma CTB, are derived as follows:

[0221] - If chroma_format_idc is equal to 0 (monochrome) or separate_colour_plane_flag is equal to 1, then both CtbWidthC and CtbHeightC are equal to 0.

[0222] - Otherwise, CtbWidthC and CtbHeightC are derived as follows:

[0223] CtbWidthC = CtbSizeY / SubWidthC (7-22)

[0224] CtbHeightC = CtbSizeY / SubHeightC (7-23)

[0225] The proposed modification to the specification of the allowed range of the syntax element log2_min_luma_coding_block_size_minus2 is described by Figure 12 as can be seen from Figure 12 In this embodiment, the allowed range for the minimum decoded block size changes from 4 to the maximum transform size.

[0226] Based on an alternative way of defining the bounds of the syntax element log2_min_luma_coding_block_size_minus2, the maximum possible value of log2_min_luma_coding_block_size_minus2 is specified based on the CTU size.

[0227] Therefore, the value of log2_min_luma_coding_block_size_minus2 should be in the range from 0 to (CtbLog2SizeY - 2). This ensures that each coded / decoded block has a size at most equal to the CTU size. It can be larger than the maximum transform size. In this case, the VVC specification already mentions that a coded / decoded block with a size larger than the maximum transform size and not split into sub-coded / decoded units should be tiled into transform units in order to code / decodethe residual data for it.

[0228] The proposed modification to the specification of the allowed range of the syntax element log2_min_luma_coding_block_size_minus2 is Figure 13 illustrated. As can be seen, in this embodiment, the allowed range of the minimum coded / decoded block size is from 4 to the CTU size.

[0229] According to another embodiment of defining the bounds of the syntax element log2_min_luma_coding_block_size_minus2, the maximum possible value of log2_min_luma_coding_block_size_minus2 is specified based on the CTU size and the Virtual Pipeline Decoding Unit (VPDU) size. The VPDU represents a decoding unit assumed in the hardware implementation of the VVC decoder. The VVC decoding process is designed in such a way that for each 64x64 picture region, all the luminance and chrominance data in that picture region can be fully decoded and reconstructed before starting to decode and reconstruct the next 64x64 region in the picture being considered.

[0230] In this embodiment, the value of log2_min_luma_coding_block_size_minus2 should be in the range from 0 to (min(CtbLog2SizeY, 6) - 2). In other words, the minimum coded / decoded block size should be in the range from 0 to min(CtbSizeY, 64), which is exactly equal to the variable VSize (VPDU size) specified in VVC draft 6.

[0231] The advantages of this embodiment are as follows. According to the VVC Draft 6 specification, CtbSize can be equal to 128 and the minimum decoded block size (MinCbSizeY) can also be equal to 128. Using the proposed constraint on the size of the minimum decoded block based on the VPDU size, each 128x128 CTU must be split into 4 64x64 luminance CUs in the luminance component. Synchronously, the 64x64 chroma blocks corresponding to the 128x128 luminance CTU must be split into 4 32x32 chroma CUs. Therefore, the codec block size will comply with the VPDU constraint.

[0232] This embodiment solves the above problem in an alternative way to the foregoing embodiment that aligns the upper limit of the minimum block size with the maximum transform size.

[0233] Compressed strip-level segmentation constraint coverage

[0234] Another embodiment of the present disclosure lies in making the codec of the slice header partition information more compact than in VVC Draft 6.

[0235] In VVC Draft 6, the slice header flag partition_constraints_override_flag is coded and decoded to indicate that the codec tree configuration signaled in the sequence parameter set is overridden in the considered slice. If this override flag is true, then the parameters related to the minimum quadtree node size, the maximum BT size, the maximum TT size, and the maximum MTT level depth level are signaled in the slice header. They are coded and decoded respectively for the luminance component of the considered slice, and also for the chroma component in the case where the dual-tree codec is active.

[0236] However, in the VTM6 encoder strategy, some of these codec tree parameters are changed in some slices, while others never change. Therefore, for some specific coding strategies, the VVC Draft 6 slice header syntax specification can lead to the duplication of redundant data. Generally, the maximum MTT level depth information never changes.

[0237] In this embodiment, a signaling flag max_mtt_hierarchy_depth_override_flag is proposed, which indicates whether the (one or more) maximum hierarchy depth parameters are overridden in the slice header. If so, then the slice-level maximum hierarchy depth information is coded / decoded in the slice header. Otherwise, the slice-level maximum hierarchy depth value is set to be equal to the value of the SPS, for the luma component and the chroma component (in the case of the dual-tree) respectively. Additionally, the slice-level coding / decoding of the maximum BT size and the maximum MTT size, in the form of the syntax elements slice_log2_diff_max_bt_min_qt_luma and slice_log2_diff_max_tt_min_qt_luma, no longer depends on the value of the maximum MTT hierarchy depth as in VVC Draft 6. In the case of dual-tree coding / decoding, this dependency is also removed for the chroma component. The following table illustrates the proposed slice header syntax modification. The advantage of this embodiment is a more compact slice header syntax, resulting in a bitrate saving of up to 0.1% for video sequences with non-negligible overhead related to the high-level syntax.

[0238] Table 14: Slice header syntax modification proposed in this embodiment

[0239]

[0240]

[0241] According to another variant of this embodiment, the minimum quadtree node size information equivalent to the maximum quadtree depth is also coded / decoded at the slice level based on the flag maximum_hierarchy_depth_override signaled before the slice header. Thus, this flag maximum_hierarchy_depth_override controls the signaling of the minimum QT size and the maximum MTT hierarchy depth parameters. Compared with VVC Draft 6, the advantage of this variant is a further compressed slice header.

[0242] Table 15: Proposed further compressed slice header syntax

[0243]

[0244]

[0245] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or overlapping the period of the second decoding.

[0246] The various methods and other aspects described in this application can be used to modify the modules of video encoder 200 and decoder 300, such as, for example, the segmentation, entropy encoding, and decoding modules (202, 335, 245, 330), as Figure 2 and Figure 3 shown. Moreover, the aspects given are not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, as well as extensions of any such standards and recommendations. Unless otherwise stated or technically precluded, the aspects described in this application can be used alone or in combination.

[0247] Various numerical values are used in this application. The specific values are for illustrative purposes, and the aspects described are not limited to these specific values.

[0248] Various embodiments relate to decoding. As used in this application, "decoding" can cover, for example, all or part of the process of performing on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as, for example, entropy decoding, inverse quantization, inverse transform, and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to the broader decoding process, and it is believed that those skilled in the art will well understand.

[0249] Various embodiments relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can cover, for example, all or part of the process of performing on an input video sequence to produce an encoded bitstream.

[0250] Note that the grammatical elements used herein are descriptive terms. Thus, they do not preclude the use of other grammatical element names. Above, the grammatical elements for SPS and SH are mainly used to illustrate various embodiments. It should be noted that these grammatical elements can be placed in other grammatical structures.

[0251] The embodiments and aspects described herein can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a program). An apparatus can be implemented, for example, with appropriate hardware, software, and firmware. A method can be implemented, for example, in an apparatus (e.g., a processor), which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes a communication device, such as a computer, a mobile phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication among end users.

[0252] References to “one embodiment” or “an embodiment” or “one implementation” or “an implementation” and other variations thereof refer to the specific features, structures, characteristics, etc. described in connection with that embodiment being included in at least one embodiment. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation” and any other variations thereof throughout this application do not necessarily all refer to the same embodiment.

[0253] In addition, this application may refer to “determining” various information. Determining information can include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0254] Furthermore, this application may refer to “accessing” various information. Accessing information can include, for example, one or more of the following: receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0255] In addition, this application may refer to “receiving” various information. Like “accessing”, receiving is a broad term. Receiving information can include one or more of the following: for example, accessing information or retrieving information (e.g., from a memory). Additionally, in one way or another, during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, “receiving” is generally involved.

[0256] It should be recognized that in cases such as "A / B", "A and / or B", and "at least one of A and B", the use of any of the following, namely " / ", "and / or", and "at least one of...", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of the first and second-listed options (A and B), or the selection of the first and third-listed options (A and C), or the selection of the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related fields, this can be extended to multiple listed items.

[0257] As will be apparent to those of ordinary skill in the art, the embodiments can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the described embodiment. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over various different wired or wireless links. The signal can be stored on a processor-readable medium.

Claims

1. A method, comprising: Encoding or decoding a value indicating a maximum allowable depth of a tree structure, wherein the maximum allowable depth is limited by twice the difference between a first value and a second value, the first value indicating the size of a maximum allowable encoding / decoding block, and the second value indicating the size of a minimum allowable encoding / decoding block; and Dividing a node representing the maximum allowable encoding / decoding block or a split of the maximum allowable encoding / decoding block into encoding / decoding blocks through the tree structure, wherein the tree structure uses at least one of horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.

2. The method according to claim 1, wherein the tree structure does not include quadtree splitting.

3. The method according to claim 1, further comprising: Dividing the maximum allowable encoding / decoding block into quadtree leaf nodes through a quadtree structure, the quadtree leaf nodes including the nodes representing the split of the maximum allowable encoding / decoding block.

4. The method according to claim 1, wherein the first value corresponds to the base-2 logarithm of the size of the maximum allowable encoding / decoding block.

5. The method according to claim 1, wherein the second value corresponds to the base-2 logarithm of the size of the minimum allowable encoding / decoding block.

6. An apparatus, comprising at least one memory and one or more processors, wherein the one or more processors are configured to: Encode or decode a value indicating a maximum allowable depth of a tree structure, wherein the maximum allowable depth is limited by twice the difference between a first value and a second value, the first value indicating the size of a maximum allowable encoding / decoding block, and the second value indicating the size of a minimum allowable encoding / decoding block; and Divide a node representing the maximum allowable encoding / decoding block or a split of the maximum allowable encoding / decoding block into encoding / decoding blocks through the tree structure, wherein the tree structure uses at least one of horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.

7. The apparatus according to claim 6, wherein the tree structure does not include quadtree splitting.

8. The apparatus according to claim 6, wherein the one or more processors are further configured to: Divide the maximum allowable encoding / decoding block into quadtree leaf nodes through a quadtree structure, the quadtree leaf nodes including the nodes representing the split of the maximum allowable encoding / decoding block.

9. The apparatus according to claim 6, wherein the first value corresponds to the base-2 logarithm of the size of the maximum allowable encoding / decoding block.

10. The apparatus according to claim 6, wherein the second value corresponds to the base-2 logarithm of the size of the minimum allowable encoding / decoding block.

11. A non-transitory computer-readable storage medium storing instructions that, when executed, implement an encoding or decoding method, the encoding or decoding method comprising: Encoding or decoding a value indicating a maximum allowable depth of a tree structure, wherein the maximum allowable depth is limited by twice the difference between a first value and a second value, the first value indicating the size of a maximum allowable encoding / decoding block, and the second value indicating the size of a minimum allowable encoding / decoding block; and The nodes representing the maximum allowable coding / decoding block or the splitting of the maximum allowable coding / decoding block are split into coding / decoding blocks through the tree structure, where the tree structure uses at least one of horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.

12. The medium according to claim 11, wherein the tree structure does not include quadtree splitting.

13. The medium according to claim 11, the encoding or decoding method further comprising: The maximum allowable coding / decoding block is split into quadtree leaf nodes through a quadtree structure, and the quadtree leaf nodes include the nodes representing the splitting of the maximum allowable coding / decoding block.

14. The medium according to claim 11, wherein the first value corresponds to the base-2 logarithm of the size of the maximum allowable coding / decoding block.

15. The medium according to claim 11, wherein the second value corresponds to the base-2 logarithm of the minimum allowable coding / decoding block size.