Interdependence of transform size and coding tree unit size in video coding

By allowing CTU-level signaling at different video unit levels in the video codec standard, relying on TT/BT partitioning based on VPDU and CTU dimensions, disabling large CU/PU tools, optimizing block segmentation when the CTU size is greater than 128, resolving the inconsistency between CTU and transform size in the VVC draft, and improving codec efficiency and quality.

CN114175650BActive Publication Date: 2026-05-01DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2020-07-27
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In the existing VVC draft video codec standard, the definitions of CTU size and transform size are inconsistent, resulting in low efficiency of block segmentation processing. Furthermore, when the CTU size is greater than 128, the block segmentation structure and signaling notification need to be modified.

Method used

Explicit signaling of CTU dimensions is allowed at different video unit levels such as layers, images, strips, slices, and bricks. CTU dimensions can span different layers. TT or BT partitioning depends on VPDU dimensions. Recursive QT partitioning is performed when the CTU dimension is greater than 128. TT/BT partitioning flags are disabled or signaled. Affine model parameter calculation and IBC buffering depend on CTU dimensions. Encoding and decoding tools for large CU/PU are disabled. The maximum TU size depends on the CTU dimension.

Benefits of technology

It improves video encoding and decoding efficiency, adapts to image encoding and decoding at different resolutions, optimizes block segmentation processing, and enhances encoding and decoding quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175650B_ABST
    Figure CN114175650B_ABST
Patent Text Reader

Abstract

A method, system, and apparatus for video coding or decoding including a configurable coding tree unit (CTU) are described. An example method of video processing includes performing a conversion between a video comprising one or more video regions and a bitstream of the video, the video regions comprising one or more video blocks, wherein the conversion complies with a rule that allows the conversion to be performed using different sizes for one or more video blocks among different video regions of the one or more video regions. Another example method of video processing includes determining, based on a dimension of a video block of a video region of a video exceeding a threshold, whether to signal an indication of a binary tree (BT) partitioning of the video block in a bitstream of the video, and performing a conversion between the video and the bitstream based on the determination.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] Pursuant to applicable patent law and / or the provisions of the Paris Convention, this application promptly claims priority and benefit to International Patent Application No. PCT / CN2019 / 097926, filed on July 26, 2019. For all purposes under that law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This document covers video and image encoding and decoding technologies. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] The disclosed techniques can be used by a video or image decoder or encoder to perform encoding or decoding of video, wherein a configurable codec tree unit size is used.

[0006] In one exemplary aspect, a video processing method is disclosed. The method includes: performing a conversion between a video comprising one or more video regions and a bitstream representation of the video, the video regions comprising one or more video blocks, wherein the conversion conforms to a rule that allows different sizes to be used for one or more video blocks within different video regions in order to perform the conversion.

[0007] In another exemplary aspect, a method for video processing is disclosed. The method includes: determining, based on the fact that the size of a video block representing a video region exceeds a threshold, to partition the video block using a quadtree-based partitioning until a size condition is met and an indication to partition using a quadtree-based partitioning is excluded from the bitstream representation of the video; and performing a conversion between the video and the bitstream representation based on the determination.

[0008] In yet another exemplary aspect, a method for video processing is disclosed. The method includes: determining, based on a video block of a video region of the video having a dimension exceeding a threshold, whether to signal an indication of a ternary-tree (TT) partition of the video block in the bitstream representation of the video; and performing a conversion between the video and the bitstream representation based on the determination.

[0009] In yet another exemplary aspect, a method for video processing is disclosed. The method includes: determining, based on a video block of a video region of the video having a dimension exceeding a threshold, whether to signal an indication of a binary-tree (BT) partition of the video block in the bitstream representation of the video; and performing a conversion between the video and the bitstream representation based on the determination.

[0010] In yet another exemplary aspect, a method for video processing is disclosed. The method includes: performing a conversion between a video comprising a video region and a bitstream representation of the video, the video region comprising video blocks, wherein the conversion includes affine model parameter calculation, and wherein the affine model parameter calculation is based on the dimension of the video blocks.

[0011] In yet another exemplary aspect, a method for video processing is disclosed. The method includes: performing a conversion between a video comprising a video region and a bitstream representation of the video, the video region comprising video blocks, wherein the conversion includes the application of an intra-block copy (IBC) tool, and wherein the size of the IBC buffer is based on the maximum configurable and / or permissible dimension of the video block.

[0012] In yet another exemplary aspect, a method for video processing is disclosed. The method includes: performing a conversion between a video comprising one or more video regions and a bitstream representation of the video, the video regions comprising one or more video blocks, wherein the conversion is performed according to a rule specifying a relationship between an indication of the size of the video block of the one or more video blocks and an indication of the maximum size of a transform block (TB) used for the video block.

[0013] In another example, the above method can be implemented by a video encoder device that includes a processor.

[0014] In yet another exemplary aspect, these methods may be embodied in the form of processor-executable instructions and stored on a computer-readable program medium.

[0015] This document further describes these and other aspects. Attached Figure Description

[0016] Figure 1 This is a block diagram of an example hardware platform used to implement the technologies described in this document.

[0017] Figure 2 This is a block diagram of an exemplary video processing system that can implement the disclosed technology.

[0018] Figure 3 This is a flowchart of an exemplary method for video processing.

[0019] Figure 4 This is a flowchart of another exemplary method for video processing.

[0020] Figure 5 This is a flowchart of yet another exemplary method for video processing.

[0021] Figure 6 This is a flowchart of yet another exemplary method for video processing.

[0022] Figure 7 This is a flowchart of yet another exemplary method for video processing.

[0023] Figure 8 This is a flowchart of yet another exemplary method for video processing.

[0024] Figure 9 This is a flowchart of yet another exemplary method for video processing. Detailed Implementation

[0025] This document provides a variety of techniques that decoders of image or video bitstreams can use to improve the quality of decompressing or decoding digital video or images. For the sake of brevity, the term "video" is used herein to include both sequences of images (conventionally referred to as video) and individual images. Furthermore, video encoders can implement these techniques during encoding processing to reconstruct decoded frames for further encoding.

[0026] The use of chapter headings in this document is for ease of understanding and not to limit the embodiments and techniques to the corresponding chapters. Thus, embodiments from one chapter can be combined with embodiments from other chapters.

[0027] 1. Overview

[0028] This document relates to video codec technology. Specifically, it covers configurable codec tree units (CTUs) in video codecs and decoders. It can be applied to existing video codec standards, such as HEVC, or pending standards (Multi-Functional Video Codec). It can also be applied to future video codec standards or codecs.

[0029] 2. Preliminary Discussion

[0030] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). JVET meetings are currently held quarterly, and new codec standards aim for a 50% bitrate reduction compared to HEVC. This new codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC test model (VTM) was released at that time. With ongoing efforts to promote VVC standardization, new codec technologies have been adopted by the VVC standard at each JVET meeting. Therefore, the VVC working draft and test model VTM are updated after each meeting. Currently, the VVC project aims to achieve Technical Completion and Standardization Indication (FDIS) at the meeting in July 2020.

[0031] 2.1 CTU Dimensions in VVC

[0032] The VTM-5.0 software allows four different CTU sizes: 16×16, 32×32, 64×64, and 128×128. However, at the JVET meeting in July 2019, the minimum CTU size was redefined as 32×32 due to the adoption of JVET-O0526. Furthermore, the CTU size in VVC Working Draft 6 was encoded into the UE encoding syntax element called log2_ctu_size_minus_5 in the SPS (sequence parameter set) header.

[0033] The following is a specification modification in VVC draft 6 that defines the Virtual Pipeline Data Unit (VPDU) and adopts JVET-O0526.

[0034] 7.3.2.3. Sequence Parameter Set (RBSP) Syntax

[0035] seq_parameter_set_rbsp(){ descriptor … log2_ctu_size_minus5 u(2) …

[0036] 7.4.3.3. Sequence Parameter Set (RBSP) Semantics

[0037]

[0038] The value of log2_ctu_size_minus5 plus 5 specifies the luminance codec tree block size for each CTU. Bitstream consistency requires that the value of log2_ctu_size_minus5 be less than or equal to 2.

[0039] log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma encoding / decoding block size.

[0040] The following derivations are used to derive the variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, IbcBufWidthY, IbcBufWidthC, and Vsize:

[0041] CtbLog2SizeY=log2_ctu_size_minus5+5 (7-15)

[0042] CtbSizeY = 1 <CtbLog2SizeY (7-16)

[0043] MinCbLog2SizeY=log2_min_luma_coding_block_size_minus2+2 (7-17)

[0044] MinCbSizeY=1< <MinCbLog2SizeY (7-18)

[0045] IbcBufWidthY=128*128 / CtbSizeY (7-19)

[0046] IbcBufWidthC=IbcBufWidthY / SubWidthC (7-20)

[0047] VSize=Min(64,CtbSizeY) (7-21)

[0048] The following derivation specifies the variables CtbWidthC and CtbHeightC for the width and height of the array for each chromaticity CTB:

[0049] – If chroma_format_idc is equal to 0 (monochrome) or separate_colour_plane_flag is equal to 1, then both CtbWidthC and CtbHeightC are equal to 0.

[0050] – Otherwise, derive CtbWidthC and CtbHeightC as follows:

[0051] CtbWidthC = CtbSizeY / SubWidthC (7-22)

[0052] CtbHeightC = CtbSizeY / SubHeightC (7-23)

[0053] For log2BlockWidth in the range 0 to 4 (including the endpoints) and for log2BlockHeight in the range 0 to 4 (including the endpoints), call the upper-right diagonal and raster scan order array initialization process as specified in Clause 6.5.2 with 1 << log2BlockWidth and 1 << log2BlockHeight as inputs, and assign the output to DiagScanOrder[log2BlockWidth][log2BlockHeight] and RasterScanOrder[log2BlockWidth][log2BlockHeight].

[0054]

[0055] slice_log2_diff_max_bt_min_qt_luma specifies the difference between the base-2 logarithm of the maximum size (width or height) in luminance samples of the luma coding tree blocks that can be partitioned using binary tree partitioning in the current slice and the base-2 logarithm of the minimum size (width or height) in luminance samples of the luma leaf blocks obtained by quadtree partitioning of the CTU. The value of slice_log2_diff_max_bt_min_qt_luma shall be in the range 0 to CtbLog2SizeY - MinQtLog2SizeY (including the endpoints). When not present, infer the value of slice_log2_diff_max_bt_min_qt_luma as follows:

[0056] – If slice_type equals 2(I), then the value of slice_log2_diff_max_bt_min_qt_luma is inferred to be equal to sps_log2_diff_max_bt_min_qt_intra_slice_luma.

[0057] Otherwise (slice_type equals 0 (B) or 1 (P)), the value of slice_log2_diff_max_bt_min_qt_luma is inferred to be equal to sps_log2_diff_max_bt_min_qt_inter_slice.

[0058] `slice_log2_diff_max_tt_min_qt_luma` specifies the base-2 logarithm of the maximum size (width or height) of the luma codec block that can be partitioned using a ternary tree in the current slice, and the base-2 logarithm of the minimum size (width or height) of the luma leaf block obtained by quadtree partitioning of the CTU. The value of `slice_log2_diff_max_tt_min_qt_luma` should be in the range of 0 to `CtbLog2SizeY - MinQtLog2SizeY` (inclusive). When it does not exist, the value of `slice_log2_diff_max_tt_min_qt_luma` is inferred as follows:

[0059] – If slice_type equals 2(I), then the value of slice_log2_diff_max_tt_min_qt_luma is inferred to be equal to sps_log2_diff_max_tt_min_qt_intra_slice_luma.

[0060] Otherwise (slice_type equals 0 (B) or 1 (P)), the value of slice_log2_diff_max_tt_min_qt_luma is inferred to be equal to sps_log2_diff_max_tt_min_qt_inter_slice.

[0061] `slice_log2_diff_min_qt_min_cb_chroma` specifies the base-2 logarithm of the smallest size (in luma samples) of the chroma leaf blocks obtained by quadtree partitioning of the chroma CTU when `treeType` equals `DUAL_TREE_CHROMA`, and the base-2 logarithm of the smallest size (in luma samples) of the chroma CU when `treeType` equals `DUAL_TREE_CHROMA`. The value of `slice_log2_diff_min_qt_min_cb_chroma` should be within the range of 0 to `CtbLog2SizeY - MinCbLog2SizeY` (inclusive). If it does not exist, the value of `slice_log2_diff_min_qt_min_cb_chroma` is inferred to be equal to `sps_log2_diff_min_qt_min_cb_intra_slice_chroma`.

[0062] `slice_max_mtt_hierarchy_depth_chroma` specifies the maximum hierarchical depth of the codec units obtained by partitioning quadtree leaves into multiple tree types within the current slice, provided that `treeType` equals `DUAL_TREE_CHROMA`. The value of `slice_max_mtt_hierarchy_depth_chroma` should be within the range of 0 to `CtbLog2SizeY - MinCbLog2SizeY` (inclusive). If it does not exist, the value of `slice_log2_diff_min_qt_min_cb_chroma` is inferred to be equal to `sps_max_mtt_hierarchy_depth_intra_slices_chroma`.

[0063] `slice_log2_diff_max_bt_min_qt_chroma` specifies the base-2 logarithm of the maximum size (width or height) of the chroma codec block that can be partitioned using binary tree partitioning when `treeType` equals `DUAL_TREE_CHROMA` in the current slice, and the base-2 logarithm of the minimum size (width or height) of the chroma leaf block obtained by quadtree partitioning of the chroma CTU. The value of `slice_log2_diff_max_bt_min_qt_chroma` should be in the range of 0 to `CtbLog2SizeY - MinQtLog2SizeC` (inclusive). If it does not exist, the value of `slice_log2_diff_max_bt_min_qt_chroma` is inferred to be equal to `sps_log2_diff_max_bt_min_qt_intra_slice_chroma`.

[0064] `slice_log2_ddiff_max_tt_min_qt_chroma` specifies the base-2 logarithm of the maximum size (width or height) of the chroma codec block that can be partitioned using a ternary tree when `treeType` equals `DUALTREE CHROMA` in the current slice, and the base-2 logarithm of the minimum size (width or height) of the chroma leaf block obtained by quadtree partitioning the chroma CTU. The value of `slice_log2_diff_max_tt_min_qt_chroma` should be in the range of 0 to `CtbLog2SizeY - MinQtLog2SizeC` (inclusive). If it does not exist, the value of `slice_log2_diff_max_tt_min_qt_chroma` is inferred to be equal to `sps_log2_diff_max_tt_min_qt_intra_slice_chroma`.

[0065] The variables MinQtLog2SizeY, MinQtLog2SizeC, MinQtSizeY, MinQtSizeC, MaxBtSizeY, MaxBtSizeC, MinBtSizeY, MaxTtSizeY, MaxTtSizeC, MinTtSizeY, MaxMttDepthY, and MaxMttDepthC are derived as follows:

[0066] MinQtLog2SizeY=MinCbLog2SizeY+slice_log2_diff_min_qt_min_cb_luma(7-86)

[0067] MinQtLog2SizeC=MinnCbLog2SizeY+slice_log2_diff_min_qt_min_cb_chroma(7-87)

[0068] MinQtSizeY=1<MinQtLog2SizeY

[0069] (7-88)

[0070] MinQtSizeC=1<<MinQtLog2SizeC /

[0071] (7-89)

[0072] MaxBtSizeY=1<<(MiinQtLog2SizeY+slice_log2_diff_max_bt_min_qt_luma)(7-90)

[0073] MaxBtSizeC=1<<(MinQtLog2SizeC+slice_log2_diff_max_bt_min_qt_chroma)(7-91)

[0074] MinBtSizeY=1<MinCbLog2SizeY

[0075] (7-92)

[0076] MaxTtSizeY=1<<(MinQtLog2SizeY+slice_log2_diff_max_tt_min_qt_luma)(7-93)

[0077] MaxTtSizeC=1<<(MinQtLog2SizeC+slice_log2_ddiff_max_tt_min_qt_chroma)(7-94)MinTtSizeY=1<<MinCbLog2SizeY

[0078] (7-95)

[0079] MaxMttDepthY=slice_max_mtt_hierarchy_depth_luuma (7-96)

[0080] MaxMtttDepthC=slice_max_mtt_hierarchy_depth_chroma (7-97)

[0081] 2.2 Maximum transformation size in VVC

[0082] In VVC draft 5, the maximum luma transform size was signaled in the SPS, but it was fixed at a length of 64 and was not configurable. However, at the JVET meeting in July 2019, it was decided that the maximum luma transform size could be either 64 or 32 using a flag at the SPS level. The maximum chroma transform size is derived from the chroma sampling ratio relative to the maximum luma transform size.

[0083] The following are the corresponding specification modifications adopted from VVC Draft 6 of JVET-O05xxx.

[0084] 7.3.2.3. Sequence Parameter Set (RBSP) Syntax

[0085] seq_parameter_set_rbsp(){ descriptor … sps_max_luma_transform_size_64__flag u(1) …

[0086] 7.4.3.3. Sequence Parameter Set (RBSP) Semantics

[0087]

[0088] A value of 1 for `sps_max_luma_transform_size_64_flag` specifies that the maximum transformation size, measured in luminance samples, is 64. A value of 0 for `sps_max_luma_transform_size_64__flag` specifies that the maximum transformation size, measured in luminance samples, is 32.

[0089] When CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag should be equal to 0.

[0090] The variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, and MaxTbSizeY are derived as follows:

[0091] MinTbLog2SizeY=2 (7-27)

[0092] MaxTbLog2SizeY=sps_max_luma_transform_size64_flag? 6:5 (7-28)

[0093] MinTbSizeY=1<<MinTbLog2SizeY (7-29)

[0094] MaxTbSizeY=1<<MaxTbLog2SizeY

[0095] (7-30)

[0096]

[0097] A value of 0 for `sps_sbt_max_size_64_flag` specifies that the maximum CU width and height allowed for subblock transformations is 32 samples. A value of 1 for `sps_sbt_max_size_64_flag` specifies that the maximum CU width and height allowed for subblock transformations is 64 samples.

[0098] MaxSbtSize=Min(MaxTbSizeY, sps_sbt_max_size_64_flag? 64:32)(7-31)

[0099]

[0100] 3. Examples of technical problems solved by the disclosed technical solutions

[0101] Several issues exist in the recent VVC working draft JVET-O2001-v11, which will be described below.

[0102] 1) In the current VVC draft 6, the maximum transform size and CTU size are defined separately. For example, the CTU size can be 32, while the transform size can be 64. It is desirable that the maximum transform size be equal to or less than the CTU size.

[0103] 2) In the current VVC draft 6, block partitioning depends on the maximum transform block size, not the VPDU size. Therefore, if the maximum transform block size is 32×32, then in addition to prohibiting 128×128 TT partitioning, 64×128 vertical BT partitioning, and 128×64 horizontal BT partitioning to comply with VPDU rules, TT partitioning for 64×64 blocks is also prohibited, as are vertical BT partitioning for 32×64 / 16×64 / 8×64 codec blocks, and horizontal BT partitioning for 64×8 / 64×16 / 64×32 codec blocks, which are inefficient for encoding / decoding.

[0104] 3) The current VVC draft 6 allows CTU sizes of 32, 64, and 128. However, it is also possible for the CTU size to be larger than 128. Therefore, some syntax elements need to be modified.

[0105] a) If a larger CTU size is allowed, then the block partitioning structure and the signaling notification for the block partitioning flag can be redesigned.

[0106] b) If a larger CTU size is permissible, then some aspects of the current design can be redesigned (e.g., affine parameter derivation, IBC prediction, IBC cache size, Merge triangle prediction, CIIP, regular Merge pattern, etc.).

[0107] 4) In the current VVC draft 6, the CTU size is signaled at the SPS level. However, since adopting reference image resampling (aka adaptive resolution variation) allows images to be encoded and decoded into a single bitstream at different resolutions, the CTU size can vary across multiple layers.

[0108] 4. Exemplary embodiments and technologies

[0109] The solutions listed below should be considered as examples to illustrate some of the ideas. These projects should not be interpreted narrowly. Furthermore, these projects can be combined in any way possible.

[0110] In this document, C = min(a,b) indicates that C is equal to the minimum of a and b.

[0111] In this document, the video unit size / dimension can be the height or width of the video unit (e.g., the width or height of an image / sub-image / strip / brick / piece / CTU / CU / CB / TU / TB). If a video unit is represented by M×N, then M represents the width of the video unit and N represents the height of the video unit.

[0112] In this document, a "codec block" can be a luma codec block and / or a chroma codec block. In this invention, the size / dimension of the codec block, measured in luma samples, can be used to represent the size / dimension measured in luma samples. For example, a 128×128 codec block (or a codec block size of 128×128 in luma samples) can indicate a 128×128 luma codec block and / or a 64×64 chroma codec block for 4:2:0. Similarly, for a 4:2:2 chroma format, it can refer to a 128×128 luma codec block and / or a 64×128 chroma codec block. For a 4:4:4 chroma format, it can refer to a 128×128 luma codec block and / or a 128×128 chroma codec block.

[0113] Related configurable CTU size

[0114] 1. It proposes that different CTU dimensions (such as width and / or height) can be allowed for different video units (such as layers / pictures / subpictures / strips / pieces / bricks).

[0115] a) In one example, one or more CTU dimensions can be explicitly signaled at the video unit level (such as VPS / DPS / SPS / PPS / APS / picture / subpicture / strip / strip header / piece / brick level).

[0116] b) In one example, this allows for reference image resampling (aka adaptive resolution variation).

[0117] In this case, the CTU dimension can be different across different layers.

[0118] i. For example, the CTU dimension of the interlayer image can be explicitly derived from the downsampling / upsampling scaling factor.

[0119] 1. For example, if the signaling notification CTU dimension of the base layer is M×N (e.g., M=128 and N=128) and the interlayer codec image is resampled by a scaling factor S for width and a scaling factor T for height, where S and T can be greater than or less than 1 (e.g., S=1 / 4 and T=1 / 2 means downsampling the interlayer codec image by a factor of 4 in width and a factor of 2 in height), then the CTU dimension in the interlayer codec image can be derived as (M×S)×(N×T) or (M / S)×(N / T).

[0120] ii. For example, at the video unit level, different CTU dimensions can be explicitly signaled for multiple layers. For example, for inter-layer resampled images / sub-images, signaling can be used at the VPS / DPS / SPS / PPS / APS / image / sub-image / strip / strip header / piece / brick level to indicate CTU dimensions different from the basic layer CTU size.

[0121] 2. It is proposed that whether TT or BT partitioning is allowed can depend on the VPDU dimensions (such as width and / or height). Assume that the VPDU has a dimension VSize measured in luminance samples, and the codec tree block has a dimension CtbSizeY measured in luminance samples.

[0122] a) In one example, VSize = min(M, CtbSizeY). M is an integer value, for example, 64.

[0123] b) In one example, whether TT or BT partitioning is allowed can be independent of the maximum transform size.

[0124] c) In one example, TT partitioning can be disabled when the width or height of the codec block, measured in luminance samples, is greater than min(VSize, maxTtSize).

[0125] i. In one example, when the maximum transform size is 32×32 but the VSize is 64×64, TT partitioning can be disabled for 128×128 / 128×64 / 64×128 codec blocks.

[0126] ii. In one example, when the maximum transform size is 32×32 but the VSize is 64×64, TT partitioning is allowed for the 64×64 codec block.

[0127] d) In one example, vertical BT partitioning can be disabled when the width of the codec block, measured in luminance samples, is less than or equal to VSize, but its height, measured in luminance samples, is greater than VSize.

[0128] i. In one example, when the maximum transform size is 32×32, but the VPDU size is 64×64, vertical BT partitioning can be disabled for 64×128 codec blocks.

[0129] ii. In one example, when the maximum transform size is 32×32, but the VPDU size is equal to 64×64, vertical BT partitioning is allowed for 32×64 / 16×64 / 8×64 codec blocks.

[0130] e) In one example, vertical BT partitioning can be disabled when the codec block exceeds the width of the image / sub-image as measured in luminance samples, but its height as measured in luminance samples is greater than VSize.

[0131] i. Alternatively, horizontal BT partitioning may be allowed when the codec block exceeds the width of the picture / subpicture as measured in luminance samples.

[0132] f) In one example, horizontal BT partitioning can be disabled when the width of the codec block, measured in luminance samples, is greater than VSize, but its height, measured in luminance samples, is less than or equal to VSize.

[0133] i. In one example, when the maximum transform size is 32×32, but the VPDU size is 64×64, vertical BT partitioning can be disabled for a 128×64 codec block.

[0134] ii. In one example, when the maximum transform size is 32×32, but the VPDU size is equal to 64×64, horizontal BT partitioning is allowed for 64×8 / 64×16 / 64×32 codec blocks.

[0135] g) In one example, horizontal BT partitioning can be disabled when the codec block exceeds the height of the image / sub-image as measured in luminance samples, but its width as measured in luminance samples is greater than VSize.

[0136] i. Alternatively, vertical BT partitioning may be allowed when the codec block exceeds the height of the picture / subpicture as measured in luminance samples.

[0137] h) In one example, when TT or BT partitioning is disabled, the TT or BT partitioning flag can be signaled without being explicitly set to zero.

[0138] i. Alternatively, when TT and / or BT partitioning is enabled, the TT and / or BT partitioning flags can be explicitly signaled in the bitstream.

[0139] ii. Alternatively, when TT or BT partitioning is disabled, the TT or BT partitioning flag can be signaled, but the decoder can ignore it.

[0140] iii. Alternatively, when disabling TT or BT partitioning, the TT or BT partitioning flag can be signaled, but it must be zero in the consistent bitstream.

[0141] 3. It was proposed that CTU dimensions (such as width and / or height) can be greater than 128.

[0142] a) In one example, the CTU dimension of the signaling notification can be 256 or even larger (e.g., log2_ctu_size_minus5 can be equal to 3 or larger).

[0143] b) In one example, the derived CTU size could be 256 or even larger.

[0144] i. For example, the derived CTU dimension used for resampling images / sub-images can be greater than 128.

[0145] 4. It is proposed that when a large CTU dimension is allowed (e.g., CTU width and / or height greater than 128), then the QT partitioning flag can be inferred to be true and QT partitioning can be applied recursively until the dimension of the partitioned codec block reaches a specified value (e.g., this specified value can be set to the maximum transform block size, or 128, or 64, or 32).

[0146] a) In one example, recursive QT partitioning can be performed implicitly without signaling notification until the partitioned codec block size reaches the maximum transform block size.

[0147] b) In one example, when CTU 256×256 is applied to a dual tree, for codec blocks larger than the maximum transform block size, the QT partitioning flag can be notified without signaling, and QT partitioning can be forced for codec blocks until the partitioned codec block size reaches the maximum transform block size.

[0148] 5. A conditional signaling notification of TT partition flags is proposed for CU / PU dimensions (width and / or height) greater than 128.

[0149] a) In one example, for a 256×256CU, signaling can be used to notify both the horizontal and vertical TT partition flags.

[0150] b) In one example, for 256×128 / 256×64CU / PU, signaling can be used to notify the vertical TT partition, but not the horizontal TT partition.

[0151] c) In one example, for 128×256 / 64×256CU / PU, the horizontal TT partition can be signaled, but the vertical TT partition cannot be signaled.

[0152] d) In one example, when the TT partitioning flag is disabled for CU dimensions greater than 128, then signaling notification for it can be omitted and implicitly deduced to be zero.

[0153] i. In one example, for 256×128 / 256×64CU / PU, horizontal TT partitioning can be disabled.

[0154] ii. In one example, for 128×256 / 64×256CU / PU, vertical TT partitioning can be disabled.

[0155] 6. A proposal was made to conditionally signal BT partition flags for CU / PU dimensions (width and / or height) greater than 128.

[0156] a) In one example, both horizontal and vertical BT division flags can be used for 256×256 / 256×128 / 128×256CU / PU signaling notifications.

[0157] b) In one example, the BT flag can be used to define the 64×256CU / PU signaling notification level.

[0158] c) In one example, the vertical BT partition flag can be signaled to the 256×64CU / PU.

[0159] d) In one example, when the BT partitioning flag is disabled for CU dimensions greater than 128, then it is possible to not signal it and implicitly infer it to be zero.

[0160] i. In one example, vertical BT segmentation can be disabled for K×256CU / PU (e.g., K equals or is less than 64 when measured in luminance samples), and the vertical BT segmentation flag can be signaled without signaling and deduced to be zero.

[0161] 1. For example, in the above case, vertical BT partitioning can be disabled for 64×256CU / PU.

[0162] 2. For example, in the above case, vertical BT partitioning can be disabled to avoid 32×256CU / PU at the image / sub-image boundary.

[0163] ii. In one example, vertical BT partitioning can be disabled when the codec block exceeds the width of the image / sub-image as measured in luminance samples, but its height as measured in luminance samples is greater than M (e.g., M = 64 when measured in luminance samples).

[0164] iii. In one example, horizontal BT partitioning can be disabled for 256×K codec blocks (e.g., K equals or is less than 64 when measured in luminance samples), and the horizontal BT partitioning flag can be signaled without signaling and derived to zero.

[0165] 1. For example, in the above case, horizontal BT partitioning can be disabled for 256×64 codec blocks.

[0166] 2. For example, in the above case, horizontal BT partitioning can be disabled to avoid 256×32 codec blocks at the boundaries of images / sub-images.

[0167] iv. In one example, horizontal BT partitioning can be disabled when the codec block exceeds the height of the picture / subpicture as measured in luminance samples, but its width as measured in luminance samples is greater than M (e.g., M = 64 when measured in luminance samples).

[0168] 7. It was proposed that the calculation of affine model parameters can depend on the CTU dimension.

[0169] a) In one example, the derivation of scaled motion vectors and / or the control point motion vectors in affine predictions may depend on the CTU dimension.

[0170] 8. It was proposed that the intra block copy (IBC) buffer can depend on the maximum configurable / allowable CTU dimension.

[0171] a) For example, the IBC cache width, measured in luminance samples, can be equal to N×N divided by the CTU width (or height), measured in luminance samples, where N can be the maximum configurable CTU size measured in luminance samples, for example, N = 1 << (log2_ctu_size_minus5 + 5).

[0172] 9. A method is proposed to disable a set (or more) of specified codec tools for larger CUs / PUs, where a larger CU / PU is defined as a CU / PU with a width or height greater than N (e.g., N = 64 or 128).

[0173] a) In one example, the specified codec tools mentioned above could be a palette and / or Intra-Block Copy (IBC) and / or Intra-Skip mode and / or Triangle Prediction mode and / or CIIP mode and / or Regular Merge mode and / or Decoder-Side Motion Derivation and / or Bidirectional Optical Flow and / or Optical Flow-Based Prediction Refinement and / or Affine Prediction and / or Sub-Block-Based TMVP, etc.

[0174] i. Alternatively, multiple screen content codecs, such as palette and intra-block copy (IBC) modes, can be applied to larger CUs / PUs.

[0175] b) In one example, syntax constraints can be explicitly used to disable (multiple) specified codec tools for larger CUs / PUs.

[0176] i. For example, for CU / PUs that are not large CU / PUs, the color palette / IBC flag can be explicitly signaled.

[0177] c) In one example, it can use bitstream constraints to disable (multiple) specified codec tools for larger CUs / PUs.

[0178] Related configurable maximum transformation size

[0179] 10. It was proposed that the maximum TU size can depend on the CTU dimensions (width and / or height), or the CTU dimensions can depend on the maximum TU size.

[0180] a) In one example, the bitstream constraint that the maximum TU size should be less than or equal to the CTU dimension can be used.

[0181] b) In one example, signaling notification for the largest TU size can depend on the TU dimension.

[0182] i. For example, when the CTU dimension is less than N (e.g., N = 64), the maximum TU size for signaling notification must be less than N.

[0183] ii. For example, when the CTU dimension is less than N (e.g., N = 64), it is possible to indicate whether the maximum luminance transformation size is 64 or 32 without signaling (e.g.,

[0184] sps_max_luma_transform_size_64flag), and the maximum luminance transform size can be implicitly deduced to be 32.

[0185] 5. Examples

[0186] Newly added parts are enclosed in bold double parentheses; for example, {{a}} indicates the addition of "a". Parts deleted from the VVC working draft are enclosed in bold double brackets; for example, [[b]] ​​indicates the deletion of "b". These modifications are based on the latest VVC working draft (JVET-O2001-v11).

[0187] 5.1 Exemplary Example #1

[0188] The following embodiments are for the inventive method, which makes the maximum TU size depend on the CTU size.

[0189] 7.4.3.3. Sequence Parameter Set (RBSP) Semantics

[0190]

[0191] A value of 1 for sps_max_luma_transform_size64_flag specifies that the maximum transformation size, measured in luminance samples, is 64. A value of 0 for sps_max_luma_transform_size64_flag specifies that the maximum transformation size, measured in luminance samples, is 32.

[0192] When CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag should be equal to 0.

[0193] The variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, and MaxTbSizeY are derived as follows:

[0194] MinTbLog2SizeY=2 (7-27)

[0195] MaxTbLog2SizeYsps_max_luma_transform_size_64_flag? 6:5 (7-28)

[0196] MinTbSizeY=1<<MinTbLog2SizeY (7-29)

[0197] MaxTbSizeY={{min(CtbSizeY,1<<MaxTbLog2SizeY)}} (7-30)

[0198]

[0199] 5.2 Exemplary Example #2

[0200] The following embodiments are for the inventive method, which makes the TT and BT partitioning process dependent on the VPDU size.

[0201] 6.4.2 Allowed Binary Tree Partitioning Processing

[0202] The following is the derivation of the variable allowBtSplit:

[0203]

[0204] Otherwise, if all of the following conditions are true, then set allowBtSplit to equal FALSE.

[0205] –btSplit equals Split_BT_VER

[0206] –cbHeight is greater than [[MaxTbSizeY]]{{VSize}}

[0207] –x0+cbWidth is greater than pic_width_in_luma_samples

[0208] Otherwise, if all of the following conditions are true, then set allowBtSplit to equal FALSE.

[0209] –btSplit equals Split_BT_HOR

[0210] –cbWidth is greater than [[MaxTbSizeY]]{{VSize}}

[0211] –y0+cbHeight is greater than pic_height_in_luma_samples

[0212]

[0213] Otherwise, if all of the following conditions are true, then set allowBtSplit to equal FALSE.

[0214] –btSplit equals Split_BT_VER

[0215] –cbWidth is less than or equal to [[MaxTbSizeY]]{{VSize}}

[0216] –cbHeight is greater than [[MaxTbSizeY]]{{VSize}}

[0217] Otherwise, if all of the following conditions are true, then set allowBtSplit to equal FALSE.

[0218] –btSplit equals Split_BT_HOR

[0219] –cbWidth is greater than [[MaxTbSizeY]]{{VSize}}

[0220] –cbHeight is less than or equal to [[MaxTbSizeY]]{{VSize}}

[0221] 6.4.3 Permissible ternary tree partitioning processes

[0222]

[0223] The variable allowTtSplit is derived as follows:

[0224] – Set allowTtSplit to FALSE if one or more of the following conditions are true:

[0225] –cbSize is less than or equal to 2*MinTtSizeY

[0226] –cbWidth is greater than Min([[MaxTbSizeY]]{{VSize}},maxTtSize)

[0227] –cbHeight is greater than Min([[MaxTbSizeY]]{{VSize}},maxTtSize)

[0228] -mttDepth is greater than or equal to maxMttDepth

[0229] -x0+cbWidth is greater than pic_width_in_luma__samples

[0230] -y0+cbHeight is greater than pic_height_in_luma_samples

[0231] -treeType equals DUAL_TREE_CHROMA, and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 32.

[0232] -TreeType equals DUAL_TREE_CHROMA and modeType equals INTRA -Otherwise, set allowTtSplit to equal TRUE.

[0233] 5.3 Exemplary Model #3

[0234] The following embodiments are for the inventive method, which makes the calculation of affine model parameters dependent on the CTU size.

[0235] 7.4.3.3. Sequence Parameter Set (RBSP) Semantics

[0236]

[0237] log2_ctu_size_minus5 plus 5 specifies the luminance codec tree block size for each CTU. One requirement for bitstream consistency is that the value of log2_ctu__size_minus5 is less than or equal to [[2]]{{3 (which can be larger depending on the specification)}}.

[0238]

[0239] CtbLog2SizeY=log2_ctu_size_minus5+5

[0240] {{CtbLog2SizeY is used to indicate the CTU size of the current video unit, measured in luminance samples. When using a single CTU size for the current video unit, CtbLog2SizeY is calculated using the equation above. Otherwise, CtbLog2SizeY may depend on the actual CTU size, and can be explicitly signaled or implicitly deduced for the current video unit's sideline. (An example)}}

[0241]

[0242] 8.5.5.5 Derivation of the motion vector of the affine control point of the brightness from adjacent blocks

[0243]

[0244] The variables mvScaleHor, mvScaleVer, dHorX, and dVerX are derived as follows:

[0245] - If isCTUboundary equals TRUE, then the following applies:

[0246] mvScaleHor=MvLX[xNb][yNb+nNbH-1][0]<<[[7]]{{CtbLog2SizeY}} (8-533)

[0247] mvScaleVer=MvLX[xNb][yNb+nNbH-1][1]<<[[7]]{{CtbLog2SizeY}} (8-534)

[0248]

[0249] - Otherwise (isCTUboundary equals FALSE), then the following applies:

[0250] mvScaleHor=CpMvLX[xNb][yNb][0][0]<<[[7]]{{CtbLog2SizeY}} (8-537)

[0251] mvScaleVer=CpMvLX[xNb][yNb][0][1]<<[[7]]{{CtbLog2SizeY}}

[0252] (8-538)

[0253]

[0254] 8.5.5.6 Derivation and processing of merging candidates for the constructed affine control point motion vectors

[0255]

[0256] The following applies when availableFlagCorner[0] equals TRUE and availableFlagCorner[2] equals TRUE:

[0257] - For X replaced by 0 or 1, the following applies:

[0258] -The following is the derivation of the variable availableFlagLX:

[0259] - If all of the following conditions are TRUE, then set availableFlagLX to equal TRUE:

[0260] -predFlagLXCorner[0] equals 1

[0261] -predFlagLXCorner[2] equals 1

[0262] -refIdxLXCorner[0] is equal to refIdxLXCorner[2].

[0263] - Otherwise, set availableFlagLX to equal FALSE.

[0264] - When availableFlagLX equals TRUE, the following applies:

[0265] -The motion vector cpMvLXCorner[1] of the second control point is derived as follows:

[0266] cpMvLXCorner[1][0]=(cpMvLXCorner[0][0]<<[[7]]

[0267] {{CtbLog2SizeY}})+

[0268] ((cpMvLXCorner[2][1]-cpMvLXCorner[0][1])

[0269] (8-606)

[0270] <<([[7]]{{CtbLog2SizeY}}+Log2(cbHeight / cbWidth)))

[0271] cpMvLXCorner[1][1]=(cpMvLXCorner[0][1]<<[[7]]

[0272] {{CtbLog2SizeY}})+

[0273] ((cpMvLXCorner[2][0]-cpMvLXCorner[0][0])

[0274] (8-607)

[0275] <<([[7]]{{CtbLog2SizeY}}+Log2(cbHeight / cbWidth)))

[0276] 8.5.5.9 Derivation of the motion vector array from the motion vectors of the affine control points

[0277] The variables mvScaleHor, mvScaleVer, dHorX, and dVerX are derived as follows:

[0278] mvScaleHor=cpMvLX[0][0]<<[[7]]{{CtbLog2SizeY}}

[0279] (8-665)

[0280] mvScaleVer=cpMvLX[0][1]<<[[7]]{{CtbLog2SizeY}}

[0281] (8-666)

[0282] Figure 1This is a block diagram of a video processing apparatus 1300. Apparatus 1300 can be used to implement one or more of the methods described herein. Apparatus 1300 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 1300 may include one or more processors 1302, one or more memories 1304, and video processing hardware 1306. The processors(multiple) 1302 may be configured to implement one or more methods described herein. The memories(multiple) 1304 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1306 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, hardware 1306 may be at least partially included within processor 1302 (e.g., a graphics coprocessor).

[0283] In some embodiments, it can be used in combination Figure 1 The device described is implemented on the hardware platform to implement these video encoding and decoding methods.

[0284] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream representation of the video will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to a video block will be performed using the video processing tool or mode enabled based on a decision or determination.

[0285] Some embodiments of the disclosed technology include making a decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion from video blocks to a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that no modification has been made to the bitstream using a video processing tool or mode enabled based on the decision or mode.

[0286] Figure 2This is a block diagram illustrating an exemplary video processing system 200 in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 200. System 200 may include an input 202 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or it may be received in a compressed or encoded format. Input 202 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0287] System 200 may include an encoding / decoding unit 204, which may implement the various encoding / decoding or coding methods described in this document. Encoding / decoding unit 204 may reduce the average bit rate of the video from input 202 to the output of encoding / decoding unit 204 to produce an encoded / decoded representation of the video. Therefore, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding unit 204 may be stored or transmitted via connected communication, as shown in unit 206. The stored or transmitted bitstream (or encoded / decoded) representation of the video received at input 202 may be used by unit 208 to generate pixel values ​​or to send displayable video to display interface 210. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding / decoding results will be performed by the decoder.

[0288] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0289] Figure 3This is a flowchart of video processing method 300. Method 300 includes, in operation 310, performing a conversion between a video comprising one or more video regions and a bitstream representation of the video, the video regions comprising one or more video blocks, the conversion following a rule that allows the conversion to be performed using different sizes for one or more video blocks in different video regions within the one or more video regions.

[0290] Figure 4 This is a flowchart of video processing method 400. Method 400 includes: in operation 410, if the size of a video block based on a video region of a video exceeds a threshold, determining to use a quadtree-based partitioning to partition the video block until a size condition is met and excluding indications for quadtree-based partitioning from the bitstream representation of the video.

[0291] Method 400 includes performing a conversion between the video and the bitstream representation in operation 420 based on the determination.

[0292] Figure 5 This is a flowchart of video processing method 500. Method 500 includes: in operation 510, if the dimension of a video block in a video region of a video exceeds a threshold, determining whether to signal an indication of ternary-tree (TT) partitioning of the video block in the bitstream representation of the video.

[0293] Method 500 includes performing a conversion between the video and the bitstream representation based on the determination in operation 520.

[0294] Figure 6 This is a flowchart of video processing method 600. Method 600 includes: in operation 610, if the dimension of a video block in a video region of a video exceeds a threshold, determining whether to signal an indication of binary-tree (BT) partitioning of the video block in the bitstream representation of the video.

[0295] Method 600 includes performing a conversion between the video and the bitstream representation in operation 620 based on the determination.

[0296] Figure 7 This is a flowchart of video processing method 700. Method 700 includes: in operation 710, performing a conversion between a video including a video region and a bitstream representation of the video, the video region including video blocks, the conversion including affine model parameter calculation, and the affine model parameter calculation being based on the dimension of the video block.

[0297] Figure 8This is a flowchart of video processing method 800. Method 800 includes: in operation 810, performing a conversion between a video comprising a video region and a bitstream representation of the video, the video region comprising video blocks, the conversion including the application of an intra-block copy (IBC) tool, and the size of the IBC buffer being based on the maximum configurable and / or permissible dimension of the video block.

[0298] Figure 9 This is a flowchart of a video processing method 900. Method 900 includes, in operation 910, performing a conversion between a video comprising one or more video regions and a bitstream representation of the video, the video regions comprising one or more video blocks, performing the conversion according to a rule specifying a relationship between an indication of the video block size of the one or more video blocks and an indication of the maximum size of a transform block (TB) used for the video block.

[0299] In some embodiments, the following technical solutions may be implemented.

[0300] A1. A method for video processing, comprising: performing a conversion between a video comprising one or more video regions and a bitstream representation of the video, the video regions comprising one or more video blocks, wherein the conversion conforms to a rule that allows the conversion to be performed using different sizes for one or more video blocks in different video regions within the one or more video regions.

[0301] A2. The method according to solution A1, wherein the rule further specifies that a syntax element indicating one or more sizes of video blocks allowed in the bitstream representation is included in the bitstream representation.

[0302] A3. The method described in solution A2, wherein the syntax element is included in the sequence parameter set (SPS).

[0303] A4. According to the method described in solution A2, the syntax element is included in the picture parameter set (PPS).

[0304] A5. As described in solution A2, wherein the syntax element is included in the video parameter set (VPS), decoding parameter set (DPS), adaptation parameter set (APS), image header, sub-image header, strip header, slice header, or brick header.

[0305] A6. The method according to solution A1, wherein the one or more video regions correspond to a video layer, and wherein the one or more video blocks correspond to a codec tree unit (CTU), the codec tree unit representing a logical segmentation for encoding and decoding the video into the bitstream representation.

[0306] A7. The method according to solution A6, wherein when the reference image resampling tool is enabled for at least one video region of the one or more video regions, different sizes are used for the one or more video blocks in the video layer.

[0307] A8. The method according to solution A6, wherein at least one of the one or more video regions comprises an inter-layer image or an intra-layer image, and wherein the dimensions of the one or more video blocks used for inter-layer or intra-layer referencing are implicitly based on a scaling factor.

[0308] A9. The method according to solution A8, wherein the scaling factor includes an upsampling scaling factor or an downsampling scaling factor.

[0309] A10. The method according to solution A8, wherein the scaling factor is derived from the size of the current image comprising one or more blocks and the size of a reference image associated with the current image.

[0310] A11. The method according to solution A8, wherein the scaling factor is derived from one or more syntax elements in the bitstream representation.

[0311] A12. The method according to solution A8, wherein the video block size of one or more video blocks of inter-layer or intra-layer images is M×N, wherein the inter-layer or intra-layer images are resampled in the width dimension according to a first scaling factor (S) and a second scaling factor (T), wherein the dimension of the video block used for inter-layer or intra-layer reference is (M×S)×(N×T) or (M / S)×(N / T), and wherein M, N, S and T are positive integers.

[0312] A13. The method according to solution A8, wherein the signaling in the bitstream representation notifies the different sizes of one or more video blocks used within the video layer.

[0313] A14. The method described in solution A13, wherein these different sizes are signaled in the Sequence Parameter Set (SPS) or Picture Parameter Set (PPS).

[0314] A15. The method described in solution A14, wherein each of these different sizes is different from the size of the base layer CTU.

[0315] A16. The method according to solution A6, wherein the dimensions of the CTU include height and width, and wherein the height and / or width is greater than 128.

[0316] A17. The method described in solution A6, wherein the dimensions of the CTU include height and width, and wherein the height and / or width is greater than or equal to 256.

[0317] A18. A method for video processing, comprising: determining, based on the fact that the size of a video block of a video region of a video exceeds a threshold, to partition the video block using a quadtree-based partitioning until a size condition is met and an indication to partition the video using a quadtree-based partitioning is excluded from the bitstream representation of the video; and performing a conversion between the video and the bitstream representation based on the determination.

[0318] A19. The method described in solution A18, wherein the threshold is 128.

[0319] A20. The method according to solution A18 or A19, wherein the size condition corresponds to a maximum transform block size of 64 or 32.

[0320] A21. The method according to any of the solutions A18 to A20, wherein the video block corresponds to a codec tree unit (CTU) that represents a logical segmentation for encoding and decoding the video into the bitstream representation.

[0321] A22. A video processing method, comprising: determining, based on a video block of a video region of a video exceeding a threshold, whether to signal an instruction for a ternary tree (TT) partition of the video block in a bitstream representation of the video; and performing a conversion between the video and the bitstream representation based on the determination.

[0322] A23. The method described in solution A22, wherein the threshold is 128.

[0323] A24. The method described in solution A22 or A23, wherein when the dimension of the video block is 256×256, the indication includes a horizontal TT flag and a vertical TT flag.

[0324] A25. The method according to solution A22 or A23, wherein when the dimensions of the video block are 256×128 or 256×64, the indication is constituted by a vertical TT flag.

[0325] A26. The method according to solution A22 or A23, wherein when the dimension of the video block is 128×256 or 64×256, the indication is composed of a horizontal TT flag.

[0326] A27. The method according to any of the solutions A22 to A26, wherein the video block is a coding unit (CU) or a prediction unit (PU).

[0327] A28. A method for video processing, comprising: determining, based on a video block of a video region of a video exceeding a threshold dimension, whether to signal an indication of a binary tree (BT) partition of the video block in a bitstream representation of the video; and performing a conversion between the video and the bitstream representation based on the determination.

[0328] A29. The method described in solution A28, wherein the threshold is 128.

[0329] A30. The method according to solution A28 or A29, wherein when the dimensions of the video block are 256×256, 256×128 or 128×256, the indication includes a horizontal TT flag and a vertical TT flag.

[0330] A31. The method according to solution A28 or A29, wherein when the dimension of the video block is 64×256, the indication is composed of a horizontal TT flag.

[0331] A32. The method according to solution A28 or A29, wherein when the dimension of the video block is 256×64, the indication is composed of a vertical TT flag.

[0332] A33. The method described according to any of the solutions A28 to A32, wherein the video block is a codec unit (CU) or a prediction unit (PU).

[0333] A34. A method for video processing, comprising: performing a conversion between a video comprising a video region and a bitstream representation of the video, the video region comprising video blocks, wherein the conversion includes affine model parameter calculation, and wherein the affine model parameter calculation is based on the dimension of the video blocks.

[0334] A35. The method according to solution A34, wherein the affine model parameter calculation is part of the affine prediction process, which further includes the derivation of scaled motion vectors and / or control point motion vectors, and wherein the derivation is based on the dimension of the video block.

[0335] A36. The method according to solution A34 or A35, wherein the video block corresponds to a codec tree unit (CTU) that represents a logical segmentation for encoding and decoding the video into the bitstream representation.

[0336] A37. A method of video processing, comprising: performing a conversion between a video comprising a video region and a bitstream representation of the video, the video region comprising video blocks, wherein the conversion comprises the application of an intra-block copy (IBC) tool, and wherein the size of the IBC buffer is based on the maximum configurable and / or permissible dimension of the video block.

[0337] A38. The method according to solution A37, wherein the width of the IBC buffer, measured in luminance samples, is equal to N×N divided by the width or height of the video block, wherein N×N is the maximum configurable dimension of the video block, measured in luminance samples, and wherein N is an integer.

[0338] A39. The method described in solution A38, wherein N = 1 << (log2_ctu_size_minus5 + 5), where log2_ctu_size_minus5 represents an indication of the codec tree unit (CTU) size.

[0339] A40. The method according to any of the solutions A37 to A39, wherein the video block corresponds to a codec tree unit (CTU), which represents a logical segmentation used to encode and decode the video into the bitstream representation.

[0340] A41. The method according to any of the solutions A1 to A40, wherein performing the conversion includes generating the bitstream representation from the video.

[0341] A42. The method according to any of the solutions A1 to A40, wherein performing the conversion includes generating the video from the bitstream representation.

[0342] A43. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any of the solutions A1 to A42.

[0343] A44. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method described according to any of the solutions A1 to A42.

[0344] In some embodiments, the following technical solutions may be implemented.

[0345] B1. A method of video processing, comprising: performing a conversion between a video comprising one or more video regions and a bitstream representation of the video, the video regions comprising one or more video blocks, wherein the conversion is performed according to a rule specifying a relationship between an indication of the video block size of the one or more video blocks and an indication of the maximum size of a transform block (TB) for the video block.

[0346] B2. The method described in solution B1, wherein the relationship specifies that the maximum size of TB is based on the size of the video block.

[0347] B3. The method described in solution B1, wherein the relationship specifies that the size of the video block is based on the maximum size of the TB.

[0348] B4. The method described in solution B2 or B3, wherein the maximum size of the TB is less than or equal to the dimension of the video block.

[0349] B5. The method according to solution B2 or B3, wherein the inclusion of the bitstream indicating the maximum size of the TB is based on the dimension of the video block.

[0350] B6. The method according to solution B5, wherein at least one dimension of the video block is less than N, wherein the indication of the maximum size of the TB indicates that the maximum size of the TB is less than N, and wherein N is a positive integer.

[0351] B7. The method described in solution B6, where N = 64.

[0352] B8. The method described in solution B5, wherein the maximum size of the luminance transformation block associated with the video region is 64 or 32.

[0353] B9. The method according to solution B8, wherein when at least one dimension of the video block is less than N, the bitstream indicates an exclusion of the maximum size of the brightness transformation block, wherein the maximum size of the brightness transformation block is implicitly derived to be 32, and wherein N is a positive integer.

[0354] B10. The method described in solution B9, wherein N = 64.

[0355] B11. The method according to any of the solutions B1 to B10, wherein the video block corresponds to a coding tree block (CTB) that represents a logical segmentation used to encode and decode the video into the bitstream representation.

[0356] B12. The method according to any of the solutions B1 to B10, wherein the video block corresponds to a luminance codec tree block (CTB) that represents a logical segmentation for encoding and decoding the luminance components of the video into the bitstream representation.

[0357] B13. The method described in any of the solutions B1 to B10, wherein the indication of the size of the video block corresponds to a syntax element or variable indicating whether the size of the Luminance Codec Tree Block (CTB) is greater than 32.

[0358] B14. The method described in any of the solutions B1 to B10, wherein the indication of the size of the video block corresponds to a syntax element or variable indicating whether the size of the Luminance Codec Tree Block (CTB) is greater than or equal to 64.

[0359] B15. The method according to any of the solutions B1 to B10, wherein the maximum size of the transformation block corresponds to the maximum size of the brightness transformation block.

[0360] B16. The method described in any of the solutions B1 to B10, wherein the indication of the maximum size of the transform block corresponds to a syntax element or variable indicating whether the maximum size of the luminance transform block is equal to 64.

[0361] B17. The method described in solution B16, wherein the syntax element is a flag.

[0362] B18. The method according to any of the solutions B1 to B17, wherein performing the conversion includes generating the bitstream representation from the video region.

[0363] B19. The method according to any of the solutions B1 to B17, wherein performing the conversion includes generating the video region from the bitstream representation.

[0364] B20. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any of solutions B1 to B19.

[0365] B21. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for implementing the method described according to any of solutions B1 to B19.

[0366] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuit systems or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition affecting machine-readable propagation signals, or combinations thereof. The term "data processing apparatus" covers all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. The transmitted signal is a man-made signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.

[0367] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0368] The processes and logic flows described in this specification can be executed by one or more programmable processors executing one or more computer programs, thereby performing functions by manipulating input data and generating outputs. These processes and logic flows can also be executed by dedicated logic circuit systems, and the device can also be implemented as a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0369] For example, processors suitable for executing computer programs include general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into a dedicated logic circuit system.

[0370] While this patent document contains numerous details, it should not be construed as limiting any subject matter or scope of the claims, but rather as a description of specific features of particular embodiments of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features from the claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.

[0371] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequential execution of such operations to obtain the desired result, or requiring the execution of all illustrated operations. Furthermore, the division of various system components in the embodiments described in this patent document should not be construed as requiring such division in all embodiments.

[0372] Only a few implementation methods and examples have been described. Other implementation methods, enhancements and variations can be made based on the content described and illustrated in this patent document.

Claims

1. A video processing method, comprising: Perform a conversion between a video comprising one or more video regions and the bitstream of said video, wherein said video regions comprise one or more video blocks. The transformation is performed according to rules that define the relationship between an indication of the video block size of the one or more video blocks and an indication of the maximum size of the transform block (TB) used for the video blocks, wherein the inclusion of the indication of the maximum size of the TB in the bitstream is based on the dimensions of the video blocks.

2. The method according to claim 1, wherein, The relationship specifies that the maximum size of the TB is based on the size of the video block.

3. The method according to claim 1, wherein, The relationship specifies that the size of the video block is based on the maximum size of the TB.

4. The method according to claim 1, wherein, The maximum size of the TB is less than or equal to the dimension of the video block.

5. The method according to claim 1, wherein, At least one dimension of the video block is less than N, wherein the indication of the maximum size of the TB indicates that the maximum size of the TB is less than N, and wherein N is a positive integer.

6. The method according to claim 5, wherein, N = 64。 7. The method according to claim 1, wherein, The maximum size of the brightness transformation block associated with the video region is 64 or 32.

8. The method according to claim 7, wherein, When at least one dimension of the video block is less than N, the bitstream does not contain an indication of the maximum size of the luminance transform block, wherein the maximum size of the luminance transform block is implicitly derived to be 32, and wherein N is a positive integer.

9. The method according to claim 8, wherein, N = 64。 10. The method according to claim 1, wherein, The video block corresponds to a codec tree block (CTB), which represents a logical segmentation used to encode and decode the video into the bitstream.

11. The method according to claim 1, wherein, The video block corresponds to a luminance codec tree block (CTB), which represents a logical segmentation used to encode and decode the luminance components of the video into the bitstream.

12. The method according to claim 1, wherein, The indication of the size of the video block corresponds to a syntax element or variable indicating whether the size of the Luminance Codec Tree Block (CTB) is greater than 32.

13. The method according to claim 1, wherein, The indication of the size of the video block corresponds to a syntax element or variable indicating whether the size of the Luminance Codec Tree Block (CTB) is greater than or equal to 64.

14. The method according to claim 1, wherein, The maximum size of the transformation block corresponds to the maximum size of the brightness transformation block.

15. The method according to claim 1, wherein, The indication of the maximum size of the transform block corresponds to a syntax element or variable indicating whether the maximum size of the brightness transform block is equal to 64.

16. The method according to claim 15, wherein, The syntax element is a flag.

17. The method according to any one of claims 1 to 16, wherein, Performing the conversion includes generating the bitstream from the video region.

18. The method according to any one of claims 1 to 16, wherein, Performing the conversion includes generating the video region from the bitstream.

19. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, When executed by the processor, the instructions cause the processor to perform the method according to any one of claims 1 to 18.

20. A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing program code of the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Image coding method, image decoding method, image coding device, image decoding device, and image coding / decoding device

    CN104604225A

  • Method and apparatus of video data processing with restricted block size in video coding

    CN108713320A