Video processing method and apparatus, and non-transitory computer-readable storage medium

By forcing quadtree segmentation at image boundaries, the problem of insufficient coding block partitioning efficiency in existing video coding standards is solved, coding efficiency is improved, and the high-performance requirements of efficient video coding standards are met.

CN120201190BActive Publication Date: 2026-03-31ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-24
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing video coding standards fail to effectively utilize quadtree segmentation in their block partitioning methods at image boundaries, resulting in insufficient coding efficiency. This is especially true in high-efficiency video coding standards such as VVC/H.266, where higher coding efficiency is required to achieve the same subjective quality in video transmission and storage.

Method used

When a sample outside the image boundary is detected in the coded block, quadtree segmentation is forced, regardless of the setting of the first parameter, to ensure that quadtree segmentation is allowed, thereby improving the segmentation efficiency of the coded block.

Benefits of technology

By enforcing quadtree partitioning, the coding efficiency of video coding is improved, meeting the higher coding performance requirements of high-efficiency video coding standards and achieving better subjective quality under the same bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201190B_ABST
    Figure CN120201190B_ABST
Patent Text Reader

Abstract

A video processing method and apparatus are provided. An example method includes determining whether a coding block includes samples outside of a picture boundary, and responsive to the coding block being determined to include samples outside of a picture boundary, performing quad tree partitioning of the coding block regardless of a value of a first parameter, wherein the first parameter indicates whether use of the quad tree to partition the coding block is allowed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This disclosure claims priority to U.S. Provisional Application No. 62 / 948,856, filed on December 17, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to video processing, and more specifically, to methods and apparatus for dividing images into blocks at boundaries. Background Technology

[0004] Video is a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, the most common being prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards, such as High Efficiency Video Coding (HEVC / H.265), Universal Video Coding (VVC / H.266), and AVS, specify particular video coding formats and are developed by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention

[0005] In some embodiments, an exemplary video processing method includes: determining whether a coded block includes samples outside an image boundary; and in response to the coded block being determined to include samples outside an image boundary, performing quadtree segmentation of the coded block regardless of the value of a first parameter, wherein the first parameter indicates whether the quadtree is allowed to be used to segment the coded block.

[0006] In some embodiments, an exemplary video processing apparatus includes at least one memory for storing instructions and at least one processor. The at least one processor is configured to execute the instructions to cause the apparatus to: determine whether a coded block includes samples outside an image boundary; and, in response to the coded block being determined to include samples outside an image boundary, perform quadtree segmentation of the coded block regardless of the value of a first parameter, wherein the first parameter indicates whether the quadtree is permitted for segmenting the coded block.

[0007] In some embodiments, an exemplary non-transitory computer-readable storage medium stores a set of instructions. The set of instructions can be executed by one or more processing devices to cause a video processing device to: determine whether a coded block includes samples outside an image boundary, and, in response to the coded block being determined to include samples outside an image boundary, perform a quadtree segmentation of the coded block regardless of the value of a first parameter, wherein the first parameter indicates whether the quadtree is permitted for segmenting the coded block. Attached Figure Description

[0008] Embodiments and aspects of this disclosure are illustrated in the following detailed description and accompanying drawings. The various features shown in the figures are not drawn to scale.

[0009] Figure 1 This is a schematic diagram of the structure of an exemplary video sequence according to some embodiments of the present disclosure.

[0010] Figure 2A This is a schematic diagram illustrating an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure.

[0011] Figure 2B This is a schematic diagram illustrating another exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure.

[0012] Figure 3A This is a schematic diagram illustrating an exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure.

[0013] Figure 3B This is a schematic diagram illustrating another exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure.

[0014] Figure 4 This is a block diagram of an exemplary apparatus for encoding or decoding video according to some embodiments of the present disclosure.

[0015] Figure 5 This is a schematic diagram illustrating an example of a multi-type tree splitting pattern according to some embodiments of the present disclosure.

[0016] Figure 6 This is a schematic diagram illustrating an exemplary signaling mechanism for partitioning information in a quadtree (QT) with a nested multi-type tree-coded tree structure, according to some embodiments of the present disclosure.

[0017] Figure 7 Exemplary Table 1 is shown according to some embodiments of the present disclosure, illustrating the derivation of an exemplary multi-type tree splitting mode (MttSplitMode) based on multi-type tree syntax elements.

[0018] Figure 8This is a schematic diagram illustrating examples of disallowed ternary tree (TT) and binary tree (BT) partitions according to some embodiments of the present disclosure.

[0019] Figure 9 This is a schematic diagram illustrating exemplary blocks divided on an image boundary according to some embodiments of the present disclosure.

[0020] Figure 10 Exemplary Table 2 is shown according to some embodiments of the present disclosure, illustrating exemplary specifications for parallel ternary tree partitioning (parallelTtSplit) and code block size (cbSize) based on binary partitioning mode (btSplit).

[0021] Figure 11 Exemplary Table 3 is shown according to some embodiments of the present disclosure, illustrating an exemplary specification of cbSize based on a ternary tree partitioning pattern (ttSplit).

[0022] Figure 12 Exemplary Table 4 is shown as an example of some embodiments according to this disclosure, illustrating an exemplary encoding tree syntax.

[0023] Figure 13 This is a schematic diagram illustrating an example of a multi-type tree splitting mode indicated by MttSplitMode according to some embodiments of the present disclosure.

[0024] Figure 14 Exemplary Table 5 is shown according to some embodiments of the present disclosure, illustrating an exemplary specification of MttSplitMode.

[0025] Figure 15 Exemplary Table 6 is shown according to some embodiments of the present disclosure, illustrating an exemplary sequence parameter set RBSP syntax.

[0026] Figure 16 Exemplary Table 7 is shown according to some embodiments of the present disclosure, illustrating exemplary image header RBSP syntax.

[0027] Figure 17 This is a schematic diagram illustrating exemplary blocks where QT, TT, or BT partitioning is not allowed at image boundaries according to some embodiments of this disclosure.

[0028] Figure 18 This is a schematic diagram illustrating exemplary blocks where QT, TT, or BT partitioning is not allowed at image boundaries according to some embodiments of this disclosure.

[0029] Figure 19 This is a schematic diagram illustrating examples of using BT and TT partitioning according to some embodiments of the present disclosure.

[0030] Figure 20 A flowchart of an exemplary video processing method according to some embodiments of the present disclosure is shown. Detailed Implementation

[0031] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, wherein, unless otherwise stated, the same numerals in different drawings denote the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with aspects of this disclosure as described in the appended claims. Specific aspects of this disclosure are described below in more detail. In the event of any conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0032] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Universal Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0033] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing techniques other than HEVC using the Joint Exploratory Model (JEM) reference software. With the incorporation of coding techniques into JEM, JEM achieves higher coding performance than HEVC.

[0034] The VVC standard is a recent development and continues to include more coding techniques that provide better compression performance. VVC is based on a hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0035] Video is a set of still images (or “frames”) arranged in chronological order to store visual information. These images can be captured and stored in chronological order using video capture devices (e.g., cameras), and displayed in a time-series using video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display capability). Furthermore, in some applications, video capture devices can transmit captured video in real time to video playback devices (e.g., computers with monitors), such as for surveillance, conferencing, or live broadcasting.

[0036] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. The module used for compression is typically called an "encoder," and the module used for decompression is typically called a "decoder." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code embedded in a computer-readable medium, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, a codec can decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard; in this case, the codec can be referred to as a "transcoder."

[0037] Video encoding processes identify and retain useful information that can be used to reconstruct the image, while ignoring unimportant reconstruction information. If ignoring unimportant information prevents complete reconstruction, such an encoding process can be called "lossy." Otherwise, it can be called "lossless." Most encoding processes are lossy, a trade-off to reduce required storage space and transmission bandwidth.

[0038] Useful information about an encoded image (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include variations in pixel position, brightness, or color, with positional changes being the most important. The positional changes of a set of pixels representing an object can reflect the object's movement between the reference and current images.

[0039] An image encoded without referencing another image (i.e., it is its own reference image) is called an "I-image". An image encoded using a previous image as a reference image is called a "P-image", and an image encoded using both a previous image and a future image as reference images is called a "B-image" (the reference is "bidirectional").

[0040] Figure 1The structure of an example video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 may be live video or video that has been captured and archived. The video 100 may be real-life video, computer-generated video (e.g., computer game video), or a combination of both (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored on a storage device), or a video feed interface (e.g., a video broadcast transceiver) receiving video from a video content provider.

[0041] like Figure 1 As shown, video sequence 100 may include a series of images arranged temporally along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, with more images between images 106 and 108. Figure 1 In this diagram, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrow. In some embodiments, the reference image of an image (e.g., image 104) may not immediately precede or follow the image. For example, the reference image of image 104 may be an image preceding image 102. It should be noted that the reference images of images 102-106 are merely examples, and this disclosure does not limit the scope to such cases. Figure 1 An example of the reference image shown.

[0042] Typically, due to the computational complexity of encoding and decoding tasks, video codecs do not encode or decode the entire image at once. Instead, they can segment the image into basic segments and encode or decode each segment sequentially. In this disclosure, such basic segments are referred to as basic processing units (“BPUs”). For example, Figure 1Structure 110 illustrates an example structure of an image (e.g., any of images 102-108) from video sequence 100. In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a “macroblock” in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (“CTU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit may have a variable size in the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing units for the image can be chosen based on a balance between coding efficiency and the level of detail to be maintained within the basic processing units.

[0043] A basic processing unit can be a logical unit that may include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image may include a luminance component (Y) representing achromatic luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components may be referred to as “code tree blocks” (“CTBs”). Any operation performed on a basic processing unit may be performed repeatedly on each of its luminance and chrominance components.

[0044] Video encoding involves multiple operational stages, examples of which are as follows: Figure 2A-2B and Figures 3A-3BAs shown. For each stage, the size of the basic processing unit may still be too large for the processing, and therefore can be further divided into segments referred to herein as "basic processing subunits". In some embodiments, the basic processing subunit may be referred to as a "block" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size. Similar to the basic processing unit, the basic processing subunit is also a logical unit that may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit may be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to further levels as needed for processing. It should also be noted that different schemes can be used to divide the basic processing units for different stages.

[0045] For example, in the pattern decision-making stage (examples of which are in...) Figure 2B As shown, the encoder can decide which prediction mode (e.g., intra-frame prediction or inter-frame prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing subunit.

[0046] For another example, in the prediction phase (the example is in...) Figure 2A-2B As shown in the diagram, the encoder can perform prediction operations at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to handle. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) at which prediction operations can be performed.

[0047] For another example, in the transformation phase (the example of which is in...) Figure 2A-2BAs shown in the diagram, the encoder can perform transformation operations on residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level transformation operations can be performed. It is important to note that the partitioning scheme of the same basic processing subunit can differ between the prediction and transformation phases. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0048] exist Figure 1 In structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, the boundaries of which are shown by dashed lines. Different basic processing units of the same image can be divided into basic processing sub-units in different schemes.

[0049] In some implementations, to provide the ability to perform parallel processing of video encoding and decoding, as well as fault tolerance, an image can be divided into regions for processing, such that for a given region of the image, the encoding or decoding process can be independent of information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving encoding efficiency. Furthermore, when data in one region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thus providing fault tolerance. In some video coding standards, images can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: “slices” and “tiles.” It should also be noted that different images in the video sequence 100 may have different partitioning schemes for dividing the image into regions.

[0050] For example, in Figure 1 In the diagram, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 comprises four basic processing units. Regions 116 and 118 each comprise six basic processing units. It should be noted that... Figure 1 The areas of the basic processing unit, basic processing subunit, and structure 110 are merely examples, and this disclosure does not limit its embodiments.

[0051] Figure 2A A schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure is shown. For example, the encoding process 200A may be performed by an encoder. Figure 2AAs shown, the encoder can encode the video sequence 202 into a video bitstream 228 according to process 200A. Similar to... Figure 1 Video sequence 100 and video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to... Figure 1 In structure 110, each raw image of video sequence 202 can be divided into basic processing units, basic processing subunits, or regions by an encoder for processing. In some embodiments, the encoder can perform process 200A at the level of basic processing units for each raw image of video sequence 202. For example, the encoder can perform process 200A iteratively, wherein the encoder can encode basic processing units in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for regions (e.g., regions 114-118) of each raw image of video sequence 202.

[0052] refer to Figure 2A The encoder feeds the basic processing unit (referred to as the "raw BPU") of the original image of video sequence 202 to prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder can subtract the predicted BPU 208 from the raw BPU to generate residual BPU 210. The encoder can feed residual BPU 210 to transform stage 212 and quantization stage 214 to generate quantization transform coefficients 216. The encoder can feed prediction data 206 and quantization transform coefficients 216 to binary encoding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder can feed quantization transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as the "reconstruction path". The reconstruction path can be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0053] The encoder can iteratively execute process 200A to encode each raw BPU (in the forward path) of the original image and generate a prediction reference 224 for encoding the next raw BPU (in the reconstruction path) of the original image. After encoding all raw BPUs of the original image, the encoder can continue to encode the next image in the video sequence 202.

[0054] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term "receive" can refer to any action that receives, inputs, acquires, retrieves, obtains, reads, accesses, or is used for inputting data in any manner.

[0055] In prediction phase 204, during the current iteration, the encoder can receive the original BPU and prediction reference 224, and perform prediction operations to generate prediction data 206 and prediction BPU 208. Prediction reference 224 can be generated from the reconstruction path of previous iterations of process 200A. The purpose of prediction phase 204 is to reduce information redundancy by extracting prediction data 206 from prediction data 206 and prediction reference 224 that can be used to reconstruct the original BPU into prediction BPU 208.

[0056] Ideally, the predicted BPU 208 should be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To record these differences, the encoder can subtract the predicted BPU 208 from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the value of the corresponding pixel in the predicted BPU 208 (e.g., grayscale or RGB value) from the pixel value of the original BPU. Each pixel in the residual BPU 210 can have a residual value as the result of this subtraction between the corresponding pixel in the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without a significant quality degradation. Thus, the original BPU is compressed.

[0057] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional “base patterns”. Each base pattern is associated with “transform coefficients”. The base patterns can have the same size (e.g., the size of the residual BPU 210), and each base pattern can represent the frequency (e.g., the frequency of brightness variation) component of the residual BPU 210. None of the base patterns can be reproduced from any combination (e.g., a linear combination) of any other base patterns. In other words, the decomposition decomposes the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to the discrete Fourier transform of a function, where the base patterns are analogous to the base functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are analogous to the coefficients associated with the base functions.

[0058] Different transform algorithms can use different base patterns. Various transform algorithms can be used at transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform at transform stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transform (called the "inverse transform"). For example, to recover the pixels of the residual BPU 210, the inverse transform can be to multiply the values ​​of the corresponding pixels of the base pattern by the corresponding correlation coefficients and sum the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore have the same base pattern). Therefore, the encoder can only record the transform coefficients, from which the decoder can reconstruct the residual BPU 210 without receiving the base pattern from the encoder. The transform coefficients can have fewer bits than the residual BPU 210, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0059] The encoder can further compress the transform coefficients during the quantization stage 214. During the transform process, different fundamental patterns can represent different frequencies of change (e.g., brightness change frequencies). Because the human eye is generally better at recognizing low-frequency changes, the encoder can ignore information about high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called the “quantization parameter”) and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of the high-frequency fundamental patterns can be converted to zero, and the transform coefficients of the low-frequency fundamental patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thus further compressing the transform coefficients. This quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called “inverse quantization”).

[0060] Because the encoder ignores the remainder of the division during rounding, quantization stage 214 can be lossy. Typically, quantization stage 214 contributes the most to information loss in process 200A. The greater the information loss, the fewer bits are required for the quantization transform coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values ​​or any other parameter of the quantization process.

[0061] In the binary encoding stage 226, the encoder can encode the prediction data 206 and the quantization transform coefficients 216 using binary encoding techniques, such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantization transform coefficients 216, the encoder can encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type at the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder can use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 can be further packaged for network transmission.

[0062] Following the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate a reconstruction residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 that will be used in the next iteration of process 200A.

[0063] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in different orders. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A One or more stages in the process.

[0064] Figure 2B A schematic diagram of another example encoding process 200B according to an embodiment of the present disclosure is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230, and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B also additionally includes a loop filtering stage 232 and a buffer 234.

[0065] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-frame image prediction or "intra-prediction") uses pixels from one or more already encoded neighboring BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of images. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") uses regions from one or more already encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the inherent temporal redundancy of images.

[0066] In reference process 200B, during the forward path, the encoder performs prediction operations in spatial prediction phase 2042 and temporal prediction phase 2044. For example, in spatial prediction phase 2042, the encoder may perform intra-frame prediction. For the original BPU of the encoded image, prediction reference 224 may include one or more adjacent BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate the predicted BPU 208 by interpolating adjacent BPUs. Interpolation techniques may include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder may perform interpolation at the pixel level, for example, by interpolating to predict the value of the corresponding pixel for each pixel of BPU 208. The adjacent BPUs used for interpolation may be located in various directions relative to the original BPU, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, the interpolation parameters, the orientation of the neighboring BPUs used relative to the original BPU, etc.

[0067] In another example, during the temporal prediction phase 2044, the encoder can perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 can include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images can be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same image have been generated, the encoder can generate a reconstructed image as the reference image. The encoder can perform a "motion estimation" operation to search for matching regions within the range of the reference image (referred to as a "search window"). The position of the search window in the reference image can be determined based on the position of the original BPU in the current image. For example, the search window can be centered at a location in the reference image that has the same coordinates as the original BPU in the current image and can extend outward by a predetermined distance. When the encoder identifies a region in the search window that resembles the original BPU (e.g., by using a PEL recursive algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region can have a different size than the original BPU (e.g., less than, equal to, greater than, or with a different shape). This is because the reference image and the current image are temporally separated on the timeline (e.g., as...). Figure 1 As shown in the image, the matching region can be considered to have "moved" to the original BPU's location over time. The encoder can record the direction and distance of this movement as a "motion vector." When using multiple reference images (e.g., such as...), Figure 1 In image 106, the encoder can search for matching regions and determine the associated motion vector for each reference image. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching regions of each matching reference image.

[0068] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0069] To generate the predicted BPU 208, the encoder can perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the predicted data 206 (e.g., motion vectors) and the predicted reference 224. For example, the encoder can move a matching region of the reference image according to the motion vectors, where the encoder can predict the original BPU of the current image. When using multiple reference images (e.g., such as...), Figure 1In image 106), the encoder can move the matching region of the reference image based on the individual motion vectors and average pixel values ​​of the matching region. In some embodiments, if the encoder has already assigned weights to the pixel values ​​of the matching regions of the respective matching reference images, the encoder can add the weighted sums of the pixel values ​​of the moved matching regions.

[0070] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in the diagram is a one-way inter-frame prediction image, where the reference image (i.e., image 102) precedes image 104. Two-way inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 Image 106 in the image is a bidirectional inter-frame prediction image, in which the reference image (i.e., images 104 and 08) is relative to image 104 in two temporal directions.

[0071] Referring again to the forward path of process 200B, after spatial prediction 2042 and temporal prediction stages 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform a rate distortion optimization technique, whereby the encoder selects a prediction mode based on the bit rate of the candidate prediction modes and the distortion of the reconstructed reference image under the candidate prediction modes to minimize the value of the cost function. Based on the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.

[0072] In the reconstruction path of process 200B, if intra-frame prediction mode has been selected in the forward path, the encoder can directly feed prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image) to spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image) after generating prediction reference 224. If inter-frame prediction mode has been selected in the forward path, the encoder can feed prediction reference 224 (e.g., the current image where all BPUs have been encoded and reconstructed) to loop filter stage 232 after generating prediction reference 224. In this stage, the encoder can apply loop filters to prediction reference 224 to reduce or eliminate distortions introduced by inter-frame prediction (e.g., block artifacts). The encoder can apply various loop filter techniques at loop filter stage 232, such as deblocking, adaptive sampling compensation, adaptive loop filtering, etc. The loop-filtered reference image can be stored in buffer 234 (or "decoded image buffer") for later use (e.g., as an inter-frame prediction reference image for future images of video sequence 202). The encoder can store one or more reference images in buffer 234 for use at temporal prediction stage 2044. In some embodiments, the encoder can encode parameters of the loop filter (e.g., loop filter strength) as well as quantization transform coefficients 216, prediction data 206, and other information at binary encoding stage 226.

[0073] Figure 3A A schematic diagram of an exemplary decoding process 300A according to an embodiment of the present invention is shown. Process 300A may be corresponding to Figure 2A The compression process 200A in the video stream is followed by the decompression process. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into video stream 304 according to process 300A. Video stream 304 can be very similar to video sequence 202. However, due to information loss during compression and decompression (e.g., Figure 2A-2B In the quantization stage 214), typically, video stream 304 differs from video sequence 202. Similar to... Figure 2A-2B In processes 200A and 200B, the decoder can perform process 300A at the basic processing unit (BPU) level for each image encoded in the video bitstream 228. For example, the decoder can perform process 300A iteratively, where the decoder can decode the basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for a region (e.g., region 114-118) of each image encoded in the video bitstream 228.

[0074] like Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit (referred to as the "encoded BPU") of the encoded image to a binary decoding stage 302, where the decoder can decode this portion into prediction data 206 and quantization transform coefficients 216. The decoder can feed the quantization transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstruction residual BPU 222. The decoder can feed the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder can feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.

[0075] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for the next encoded BPU of the encoded image. After decoding all encoded BPUs of the encoded image, the decoder can output the image to video stream 304 for display and continue decoding the next encoded image in video bit stream 228.

[0076] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary encoding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder can unpack it before feeding the video bitstream 228 to the binary decoding stage 302.

[0077] Figure 3B A schematic diagram of another example decoding process 300B according to an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.

[0078] In process 300B, for the encoding basic processing unit (referred to as the "current BPU") of the decoded encoded image (referred to as the "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder uses intra-frame prediction to encode the current BPU, the prediction data 206 can include prediction mode indicators (e.g., flag values) that indicate intra-frame prediction, parameters of the intra-frame prediction operation, etc. Parameters of the intra-frame prediction operation can include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, the sizes of neighboring BPUs, interpolation parameters, the orientation of neighboring BPUs relative to the original BPU, etc. For another example, if the current BPU is encoded by inter-frame prediction used by the encoder, the prediction data 206 can include prediction mode indicators (e.g., flag values) that indicate inter-frame prediction, parameters of the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference images, and one or more motion vectors associated with the matching regions respectively.

[0079] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra-frame prediction) in the spatial prediction phase 2042 or temporal prediction (e.g., inter-frame prediction) in the temporal prediction phase 2044. The details of performing this spatial or temporal prediction are... Figure 2B As described herein, it will not be repeated below. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208, which can be added to the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown below. Figure 3A As described in [the text].

[0080] In process 300B, the decoder can feed prediction reference 224 to either spatial prediction stage 2042 or temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if intra-frame prediction is used to decode the current BPU in spatial prediction stage 2042, the decoder can feed prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image) after generating prediction reference 224 (e.g., the decoded current BPU). If inter-frame prediction is used to decode the current BPU in temporal prediction stage 2044, the encoder can feed prediction reference 224 to loop filter stage 232 to reduce or eliminate distortion (e.g., block artifacts) after generating prediction reference 224 (e.g., a reference image where all BPUs are decoded). The decoder can, as follows: Figure 2BThe loop filter is applied to prediction reference 224 in the manner shown. The loop-filtered reference image can be stored in buffer 234 (e.g., a decoded image buffer in computer memory) for later use (e.g., as an inter-prediction reference image for future encoded images of video bitstream 228). The decoder can store one or more reference images in buffer 234 for use at temporal prediction stage 2044. In some embodiments, the prediction data can further include parameters of the loop filter (e.g., loop filter strength) when the prediction mode indicator of prediction data 206 indicates that inter-frame prediction is used to encode the current BPU.

[0081] Figure 4 This is a block diagram of an example apparatus 400 for encoding or decoding video according to embodiments of the present disclosure. Figure 4 As shown, device 400 may include processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for video encoding or decoding. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or “CPU”), graphics processing units (or “GPU”), neural processing units (“NPU”), microcontroller units (“MCU”), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-a-chip (SoCs), application-specific integrated circuits (ASICs), and any combination thereof. In some embodiments, processor 402 may also be a group of processors grouped into individual logic components. For example, such as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b and processor 402n.

[0082] The device 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, such as Figure 4As shown, the stored data may include program instructions (e.g., for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any combination of any number of random access memories (RAM), read-only memories (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital cards (SD cards), memory sticks, compact flash memory (CF cards), etc. Memory 404 may also be a group of memories grouped into single logical components. Figure 4 (Not shown in the image).

[0083] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), or the like.

[0084] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or it may be wholly or partially integrated into any other component of the device 400.

[0085] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transceivers, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0086] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connectivity to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video files), etc.

[0087] It should be noted that the video codec (e.g., the codec for executing processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instances that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0088] In quantization and dequantization function blocks (e.g., Figure 2A or Figure 2B Quantization 214 and inverse quantization 218, Figure 3A or Figure 3B In inverse quantization (218), the quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used to encode an image or slice can be signaled at a higher level, for example, using the `init_qp_minus26` syntax element in the Image Parameter Set (PPS) and the `slic_qp_delta` syntax element in the slice header. Furthermore, incremental (delta) QP values ​​sent at the granularity of quantization groups can be used to adapt QP values ​​at the local level for each CU.

[0089] According to some embodiments, an image can be divided into multiple coding tree units (CTUs). Then, the CTUs are further divided into one or more coding units (CUs) using a quadtree (SPLIT_QT) with a nested multi-type tree having a segmentation structure using binary and ternary partitioning. Figure 5 This is a schematic diagram illustrating examples of multi-type tree partitioning patterns according to some embodiments of the present disclosure. For example... Figure 5 As shown, the partitioning types in a multi-type tree structure can include quadtree partitioning (SPLIT_QT) 501, vertical binary tree partitioning (SPLIT BT VER) 502, horizontal binary tree partitioning (splitting SPLIT_BT_HOR) 503, vertical ternary tree partitioning (SPLIT_TT_VER) 504, and horizontal ternary tree partitioning (SPLIT_TT_HQR) 505. The leaf nodes of the multi-type tree are called coding units (CUs), which can have square or rectangular shapes.

[0090] Figure 6 This is a schematic diagram of an exemplary signaling mechanism for partitioning information in a quadtree with a nested multi-type tree-coded tree structure, according to some embodiments of this disclosure. Figure 6In this approach, the CTU is considered the root of a quadtree and is initially partitioned by the quadtree structure. The leaf nodes of each quadtree (when large enough to allow it) are further partitioned by a multi-type tree structure. In the multi-type tree structure, a first flag (e.g., `mtt_split_cu_flag`) is signaled to indicate whether a node is to be further partitioned. When a node is further partitioned, a second flag (e.g., `mtt_split_cu_vertical_flag`) is signaled to indicate the partition direction, and then a third flag (e.g., `mtt_split_cu_binary_flag`) is signaled to indicate whether the partition is a binary or ternary tree partition. Based on the values ​​of the second and third flags, `mtt_split_cu_vertical_flag` and `mtt_split_cu_binary_flag`, the multi-type tree partitioning mode (`MttSplitMode`) of a CU can be derived. Figure 7 Exemplary Table 1 is shown according to some embodiments of the present disclosure, illustrating an exemplary MTTSplitMode derivation based on multi-type tree syntax elements.

[0091] Virtual Pipeline Data Units (VPDUs) are defined as non-overlapping units in an image. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so maintaining the VPDU size is important. In most hardware decoders, the VPDU size can be set to 64×64 luma samples. However, in some embodiments, ternary tree (TT) and binary tree (BT) partitioning may lead to an increase in the VPDU size. Consistent with this disclosure, certain canonical partitioning constraints can be applied to maintain the VPDU size at 64×64 luma samples. Figure 8 Examples of disallowed TT and BT partitioning according to some embodiments of this disclosure are shown. Figure 8 As shown, TT splits are not allowed for blocks with a width or height, or both width and height equal to 128. For CUs of 128×N, where N≤64 (e.g., width equal to 128 and height less than 128), horizontal BT splits are not allowed. For CUs of N×128, where N≤64 (e.g., height equal to 128 and width less than 128), vertical BT splits are not allowed.

[0092] According to HEVC requirements, when a portion of a tree node block extends beyond the bottom or right edge of the image, the tree node block must be forcibly split until all samples of each encoded CU are within the image boundaries. The following splitting rules were applied in VVC draft 7:

[0093] -If part of a tree node block extends beyond the bottom and right boundaries of the image.

[0094] - If the block is a QT node and its size is greater than the minimum QT size, then force the block to be split using the QT splitting mode.

[0095] Otherwise, the block will be forced to be split using the SPLIT_BT_HOR mode.

[0096] Otherwise, if part of a tree node block extends beyond the bottom boundary of the image,

[0097] - If the block is a QT node, and the block size is greater than the minimum QT size and the block size is greater than the maximum BT size, then force the block to be split using the QT splitting mode.

[0098] Otherwise, if the block is a QT node and the block size is greater than the minimum QT size and the block size is less than or equal to the maximum BT size, then force the block to be split using either the QT split mode or the SPLIT_BT_HQR mode.

[0099] Otherwise (if the block is a BT node or the block size is less than or equal to the minimum QT size), the block will be split using the SPLIT_BT_HQR mode.

[0100] Otherwise, if a portion of a tree node block extends beyond the right boundary of the image edge,

[0101] - If the block is a QT node, and the block size is greater than the minimum QT size and the maximum BT size, then force the block to be split using QT splitting mode.

[0102] Otherwise, if the block is a QT node and the block size is greater than the minimum QT size and the block size is less than or equal to the maximum BT size, then force the block to be split using either the QT split mode or the SPLIT_BT_VER mode.

[0103] Otherwise (e.g., if the block is a BT node or the block size is less than or equal to the minimum QT size), the block is forced to be split using the SPLIT_BT_VER mode.

[0104] Figure 9 Exemplary block divisions on image boundaries are shown according to some embodiments of this disclosure. For example... Figure 9 As shown, for CTU 911, either SPLIT_QT or SPLIT_BT_VER can be executed. For CTU 913, if SPLIT_QT is enabled, SPLIT_QT can be executed; if SPLIT_QT is not enabled, SPLIT_BT_HOR can be executed. For CTU 915, either SPLIT_QT or SPLIT_BT_HOR can be executed.

[0105] In VVC Draft 7, there are two sections related to block partitioning. The first is Section 6.4, which defines whether quadtrees, binary trees, or ternary trees can be used for block partitioning. The output of Section 6.4 is the variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer. These variables are used in... Figure 12 The table shown in Section 7.3.9.4 is used to determine whether to signal the corresponding CU level division flag (by...). Figure 12 (Boxes 1201-1204 in Table 4).

[0106] Section 6.4.1 of VVC Draft 7 describes:

[0107] 6.4 Availability Process

[0108] 6.4.1 Permissible Quadrilateral Tree Splitting Process

[0109] The inputs to this process include:

[0110] - The size of the encoded block in the luminance sample, cbSize.

[0111] -Multi-type tree depth mttDepth,

[0112] - The variable `treeType` specifies whether to use a single tree (SINGLE_TREE) or a dual tree to divide the coding tree nodes, and if using a dual tree, whether the current component being processed is luma (DUAL_TREE_LUMA) or chroma (DUAL_TREE_CHROMA).

[0113] - A variable modeType that specifies whether intra-frame (MODE_INTRA), IBC (MODE_IBC), and inter-frame coding modes (MODE_TYPE_ALL) can be used, or whether only intra-frame and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter-frame coding modes (MODE_TYPE_INTER) can be used for coding units within coding tree nodes.

[0114] The output of this process is the variable allowSplitQt. The variable allowSplitQt is derived as follows:

[0115] - Set allowSplitQt to FALSE if one or more of the following conditions are true:

[0116] -treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY

[0117] -treeType is equal to DUAL_TREE_CHROMA, and cbSize / SubWidthC is less than or equal to MinQtSizeC.

[0118] -mttDepth is not equal to 0

[0119] -treeType equals DUAL_TREE_CHROMA, and (cbSize / SubWidthC) is less than or equal to 4.

[0120] -treeType equals DUAL_TREE_CHROMA, modeType equals MODE_TYPE_INTER

[0121] Otherwise, allowSplitQt is set to TRUE.

[0122] Section 6.4.2 of VVC Draft 7 describes:

[0123] 6.4.2 Permissible Binary Tree Partitioning Process

[0124] The input to this process is:

[0125] - Binary tree partitioning mode btSplit

[0126] - The width of the coded block in the luminance sample, cbWidth.

[0127] - The height of the coded block in the luminance sample, cbHeight

[0128] - The position (x0, y0) of the top-left corner brightness sample of the considered coded block relative to the top-left corner brightness sample of the image.

[0129] -Multi-type tree depth mttDepth,

[0130] - Maximum multi-type tree depth with offset maxMttDepth

[0131] -Maximum binary tree size maxBtSize

[0132] -Minimum quadtree size minQtSize

[0133] -Partition index partIdx

[0134] - The variable `treeType` specifies whether to use a single tree (SINGLE_TREE) or a dual tree to divide the coding tree nodes, and if using a dual tree, whether the current component being processed is the luma component (DUAL_TREE_LUMA) or the chroma component (DUAL_TREE_CHROMA).

[0135] - A variable modeType specifies whether intra-frame (MODE_INTRA), IBC (MODE_JBC), and inter-frame coding modes (MODE_TYPE_ALL) can be used, or whether only intra-frame and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter-frame coding modes (MODE_TYPE_INTER) can be used for coding units within coding tree nodes.

[0136] The output of this process is the variable allowBtSplit. Figure 10 Exemplary Table 2 is shown according to some embodiments of the present disclosure, illustrating exemplary specifications of variables parallelTtSplit and cbSize based on btSplit.

[0137] The derivation of the variable allowBtSplit is as follows:

[0138] - AllowBtSplit is set to FALSE if one or more of the following conditions are true:

[0139] -cbSize is less than or equal to MinBtSizeY

[0140] -cbWidth is greater than maxBtSize

[0141] -cbHeight is greater than maxBtSize

[0142] -mttDepth is greater than or equal to maxMttDepth

[0143] -treeType equals DUAL_TREE_CHROMA, and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 16.

[0144] -treeType equals DUAL_TREE_CHROMA, (cbWidth / SubWidthC) equals 4, and btSplit equals Split_BT_VER

[0145] -treeType equals DUAL_TREE_CHROMA, modeType equals MODE_TYPE_INTRA

[0146] -cbWidth*cbHeight equals 32, modeType equals MODE_TYPE_INTER

[0147] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0148] -btSplit equals Split_BT_YER

[0149] -y0+cbHeight is greater than pic_height_in_luma_samples

[0150] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0151] -btSplit equals Split_BT_VER

[0152] -cbHeight is greater than 64

[0153] -x0+cbWidth is greater than pic_width_in_luma_samples

[0154] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0155] -btSplit equals Split_BT_HOR

[0156] -cbWidth is greater than 64

[0157] -y0+cbHeight is greater than pic_height_in_luma_samples

[0158] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0159] -x0+cbWidth is greater than pic_width_in_luma_samples

[0160] -y0+cbHeight is greater than pic_height_in_luma_samples

[0161] -cbWidth is greater than minQtSize

[0162] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0163] -btSplit equals Split_BT_HOR

[0164] -x0+cbWidth is greater than pic_width_in_luma_samples

[0165] -y0+cbHeight is less than or equal to pic_height_in_luma_samples

[0166] - Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true:

[0167] -mttDepth is greater than 0

[0168] -partldx equals 1

[0169] -MttsplitMode[x0][y0][mttDepth-1] equals parallelTtSplit

[0170] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0171] -btSplit equals Split_BT_VER

[0172] -cbWidth is less than or equal to 64

[0173] -cbHeight is greater than 64

[0174] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0175] -btSplit equals Split_BT_HOR

[0176] -cbWidth is greater than 64

[0177] -cbHeight is less than or equal to 64

[0178] Otherwise, allowBtSplit is set to TRUE.

[0179] Section 6.4.3 of VVC Draft 7 describes:

[0180] 6.4.3 Permissible ternary tree partitioning process

[0181] The input to this process is:

[0182] - Tritree partitioning ttSplit,

[0183] - The width of the coded block in the luminance sample, cbWidth.

[0184] - The height of the coded block in the luminance sample, cbHeight

[0185] - The position (x0, y0) of the top-left corner luminance sample of the coded block relative to the top-left corner luminance sample of the image.

[0186] -Multi-type tree depth mttDepth

[0187] - Maximum multi-type tree depth with offset maxMttDepth

[0188] -Maximum ternary tree size maxTtSize

[0189] - The variable `treeType` specifies whether to use a single tree (SINGLE_TREE) or a dual tree to partition the coding tree nodes, and if using a dual tree, whether the current processing is the luma component (DUAL_TREE_LUMA) or the chroma component (DUAL_TREE_CHROMA).

[0190] - A variable modeType specifies whether intra-frame (MODE_INTRA), IBC (MODE_IBC), and inter-frame coding modes (MODE_TYPE_ALL) can be used, or whether only intra-frame and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter-frame coding modes (MODE_TYPE_INTER) can be used for coding units within coding tree nodes.

[0191] The output of this process is the variable allowTtSplit. Figure 11 Exemplary Table 3 is shown according to some embodiments of the present disclosure, illustrating an exemplary specification of the variable cbSize based on ttSplit.

[0192] The derivation of the variable allowTtSplit is as follows:

[0193] - AllowTtSplit is set to FALSE if one or more of the following conditions are true:

[0194] -cbSize is less than or equal to 2 * MinTtSizeY

[0195] -cbWidth is greater than Min(64, maxTtSize).

[0196] -cbHeight is greater than Min(64, maxTtSize)

[0197] -mttDepth is greater than or equal to maxMttDeptb

[0198] -x0+cbWidth is greater than pic_width_in_luma_samples

[0199] -y0+cbHeight is greater than pic_height_in_luma_samples

[0200] -treeTpye equals DUAL_TREE_CHROMA, and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 32.

[0201] -treeType equals DUAL_TREE_CHROMA, (cbWidth / SubWidthC) equals 8, and ttSplit equals Split_TT_VER.

[0202] -treeType equals DUAL_TREE_CHROMA, and modeType equals MODE_TYPE_INTRA

[0203] -cbWidth*cbHeight equals 64, modeType equals MODE_TYPE_INTER

[0204] Otherwise, allowTtSplit is set to TRUE.

[0205] Figure 12 Exemplary Table 4 is shown according to some embodiments of this disclosure, which illustrates the exemplary section 7.3.9.4 of VVC draft 7, the coding tree syntax (emphasized with italics and shading).

[0206] The derivation of the variables allowSplitQt, allowSplitBtVer, allowSplitBtHor, allowSplitTtVer, and allowSplitTtHor is as follows:

[0207] - Call the allowed quadtree partitioning process specified in section 6.4.1, using the coded block size cbSize set to equal cbWidth, the current multitype tree depth mttDepth, treeTypeCurr, and modeTypeCurr as input, and assign the output to allowSplitQt.

[0208] The derivation of variables minQtSize, maxBtSize, maxTtSize, and maxMttDepth is as follows:

[0209] - If treeType equals DUAL_TREE_CHROMA, then set minQtSize, maxBtSize, maxTtSize, and maxMttdepth to equal MinQtSizeC, MaxBtSizeC, MaxTtSizeC, and MaxMttDepthc+depthOffset, respectively.

[0210] Otherwise, minQtSize, maxBtSize, maxTtSize, and maxMttDepth are set to equal to MinQtSizeY, MaxBtSizeY, MaxTtSizeY, and MaxMttDepthY + depthOffset, respectively.

[0211] - Call the allowed binary tree partitioning procedure specified in Section 6.4.2, using the binary tree partitioning mode SPLIT_BT_VER, the block width cbWith, the block height cbHeight, the position (x0, y0), the current multi-type tree depth mttDepth, the maximum multi-type tree depth with offset maxMttdepth, the maximum binary tree size maxBtSize, the minimum quadtree size minQtSize, the current partition index partIdx, treeTypeCurr, and modeTypeCurr as input, and assign the output to allowSplitBtVer.

[0212] - Calls the allowed binary tree splitting procedure specified in clause 6.4.2, taking the binary tree splitting mode SPLIT_BT_HOR, the block height cbHeight, the block width cbWidth, the position (x0, y0), the current multi-type tree depth mttDepth, the maximum multi-type tree depth with offset maxMttDepth, the maximum binary tree size maxBtSize, the minimum quadtree size minQtSize, the current partition index partIdx, treeTypeCurr, and modeTypeCurr as input, and assigns the output to allowSplitBtHor.

[0213] - Call the allowed ternary tree partitioning procedure specified in Section 6.4.3, using the ternary tree partitioning mode SPLIT_TT_VER, the block width cbWidth, the block height cbHeight, the position (x0, y0), the current multi-type tree depth mttDepth, the maximum multi-type tree depth with offset maxMttDepth, the maximum ternary tree size maxTtSize, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitTtVer.

[0214] - Call the allowed ternary tree partitioning procedure specified in Section 6.4.3, taking the ternary tree partitioning mode SPLIT_TT_HOR, the block height cbHeight, the block width cbWidth, the position (x0, y0), the current multitype tree depth mttDepth, the maximum multitype tree depth with offset maxMttDepth, the maximum ternary tree size maxTtSize, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitTtHor.

[0215] The syntax element `split_cu_flag` equal to 0 specifies that the coding unit is not split. The syntax element `split_cu_flag` equal to 1 specifies that the coding unit is divided into four coding units using a quadtree partition as indicated by the syntax element `split_qt_flag`, or into two coding units using a binary tree partition as indicated by the syntax element `mtt_split_cu_binary_flag`, or into three coding units using a ternary tree partition as indicated by the syntax element `mtt_split_cu_vertical_flag`. As indicated by the syntax element `mtt_split_cu_vertical_flag`, the binary or ternary tree partition can be vertical or horizontal.

[0216] When the syntax element split_cu_flag does not exist, the value of split_cu_flag is inferred as follows:

[0217] - If one or more of the following conditions are true, then the value of split_cu_flag will be inferred to be equal to 1:

[0218] -x0+cbWidth is greater than pic_width_in_luma_samples.

[0219] -y0+cbHeight is greater than pic_height_in_luma_sampless.

[0220] Otherwise, it is assumed that the value of split_cu_flag is equal to 0.

[0221] The syntax element split_qt_flag specifies whether the coding unit is divided into coding units that are half the size horizontally and vertically.

[0222] The following applies when the syntax element split_qt_flag is not present:

[0223] - If allowSplitQt equals TRUE, then it is inferred that the value of split_qt_flag is equal to 1.

[0224] Otherwise, it is inferred that the value of split_qt_flag is equal to 0.

[0225] The syntax element `mtt_split_cu_vertical_flag` equal to 0 specifies that the coding unit is horizontally divided. The syntax element `mtt_split_cu_vertical_flag` equal to 1 specifies that the coding unit is vertically divided.

[0226] When the syntax element mtt_split_cu_vertical_flag is not present, the following inference is made:

[0227] - If allowSplitBtHor equals TRUE or allowSplitTtHor equals TRUE, then it is inferred that the value of mtt_split_cu_vertical_flag is equal to 0.

[0228] Otherwise, it is assumed that the value of mtt_split_cu_vertical_flag is equal to 1.

[0229] The syntax element `mtt_split_cu_binary_flag` equal to 0 specifies that the coding unit is split into three coding units using a ternary tree. The syntax element `mtt_split_cu_binary_flag` equal to 1 specifies that the coding unit is split into two coding units using a binary tree.

[0230] When the syntax element mtt_split_cu_binary_flag is not present, the following inference is made:

[0231] - If allowSplitBtVer equals FALSE and allowSplitBtHor equals FALSE, then it is inferred that the value of mtt_split_cu_binary_flag is 0.

[0232] Otherwise, if allowSplitTtVer equals FALSE and allowSplitTtHor equals FALSE, then it is inferred that the value of mtt_split_cu_binary_flag is equal to 1.

[0233] Otherwise, if allowSplitBtHor equals TRUE and allowSplitTtVer equals TRUE, then it is inferred that the value of mtt_split_cu_binary_flag is equal to mtt_split_cu_vertical_flag.

[0234] - Otherwise (allowSplitBtVer equals TRUE, allowSplitTtHor equals TRUE), infer that the value of mtt_split_cu_binary_flag is equal to mtt_split_cu_vertical_flag.

[0235] Figure 14 Exemplary Table 5 is shown according to some embodiments of this disclosure, illustrating an exemplary specification of MttSplitMode. The variable MttSplitMode[x][y][mttDepth] is derived from the values ​​of the syntax element mtt_split_cu_vertical_flag and the syntax element mtt_split_cu_binary_flag when x = x0..x0+cbWidth-1 and y = y0...y0+cbHeight-1 are defined in Table 4.

[0236] MttSplitMode[x][y][mttDepth] represents the horizontal binary tree, vertical binary tree, horizontal ternary tree, and vertical ternary tree partitioning of the coding unit within the multi-type tree. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the considered coding block relative to the top-left luminance sample of the image. Figure 13 Examples of multi-type tree partitioning modes indicated by MttSplitMode, according to some embodiments of this disclosure, are shown. Figure 13 As shown, the multi-type tree partitioning modes can include vertical binary tree partitioning (SPLIT_BT_VER) 1301, horizontal binary tree partitioning (SPLIT_BT_HOR) 1302, vertical branch tree partitioning (SPLIT_TT_VER) 1303 and horizontal branch tree partitioning (SPLIT_TT_HOR) 1304.

[0237] It should be noted that the CTU size, minimum block size, and block size restrictions for quadtree, binary tree, and ternary tree partitions are signaled in the sequence parameter set or image header.

[0238] Figure 15 Exemplary Table 6 is shown according to some embodiments of the present disclosure, illustrating the exemplary section 7.3.2.3 of the VVC draft 7, the Sequence Parameter Set RBSP syntax. Figure 16 Exemplary Table 7 is shown according to some embodiments of the present disclosure, illustrating the exemplary portion 7.3.2.6 of the VVC draft 7, the image header RBSP syntax.

[0239] In some implementations, when a portion of a tree node block extends beyond the bottom or right side of the image boundary, the tree node block is forced to split until all samples of each coded block are within the image boundary. However, in some cases, for blocks located at the image boundary but containing samples extending beyond the image boundary, all tree partitioning patterns are not permitted. Figure 17 Exemplary blocks are shown that disallow QT, TT, or BT partitioning at image boundaries according to some embodiments of this disclosure. For example, Figure 17 The CTU 1701, CTU 1703, and CTU 1705 at the image boundaries do not allow QT, TT, or BT partitioning.

[0240] As a first exemplary case where all tree partitioning modes are not allowed, when both the CTU size and the minimum QT size are set to 128, all variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer are set to false.

[0241] Because the following condition variable, allowSplitQt, is set to false (italicized for emphasis):

[0242] 6.4.1 Permissible Quadtree Partitioning Process ...

[0243] - AllowSplitQt is set to FALSE if one or more of the following conditions are true:

[0244] -treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, and chSize is less than or equal to MinQtSizeY

[0245] -treeType equals DUAL_TREE_CHROMA, and cbSize / SubWidthC is less than or equal to MinQtSizeC ...

[0246] The variable allowSplitBtHor (in italics for emphasis) is set to false because of the following condition:

[0247] 6.4.2 Permissible Binary Partitioning Processes ...

[0248] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0249] -btSplit equals Split_BT_HOR

[0250] -cbWidth is greater than 64

[0251] -y0+cbHeight is greater than pic_height_luma_samples

[0252] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0253] -x0+cbWidth is greater than pic_width_in_luma_samples

[0254] -y0+cbHeight is greater than pic_height_in_luma_samples

[0255] -cbWidth is greater than minQtSize

[0256] Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true.

[0257] -btSplit equals Split_BT_HOR

[0258] -x0+cbWidth is greater than pic_width_in_luma_samples

[0259] -y0+cbHeight is less than or equal to pic_height_in_luma_samples ...

[0260] The variable allowSplitBtVer is set to false due to the following condition (italicized for emphasis):

[0261] 6.4.2 Permissible Binary Partitioning Processes ...

[0262] Otherwise, set allowBiSplit to FALSE if all of the following conditions are true.

[0263] -btSplit equals Split_BT_VER

[0264] -y0+cbHeight is greater than pic_height_in_luma_samples

[0265] Otherwise, set allowBtSplit to FALSE if all of the following conditions are true.

[0266] -btSplit equals Split_BT_VER

[0267] -chHeight is greater than 64

[0268] -x0+cbWidth is greater than pic_width_in_luma_samples

[0269] The variables allowSplitTtHor and allowSplilTtVer (in italics for emphasis) are set to false for the following reason:

[0270] 6.4.3 Permissible ternary tree partitioning process ...

[0271] - AllowTtSplit is set to FALSE if one or more of the following conditions are true:

[0272] -cbSize is less than or equal to 2 * MinTtSizeY

[0273] -cbWidth is greater than min(64, maxTtSize).

[0274] -cbHeight is greater than min(64, maxTtSize)

[0275] -mttDepth is greater than or equal to maxMttDepth

[0276] -x0+cbWidth is greater than pic_width_in_luma_samples

[0277] -y0+cbHeight is greater than pic_height_in_luma_samples ...

[0278] When all these variables are set to false, CU-level partition flags can be sent without signaling. The inferred syntax elements `split_cu_flag` are set to 1, `split_qt_flag` to 0, `mtt_split_cu_vertical_flag` to 1, and `mtt_split_cu_binary_flag` to 0. In this case, `SPLIT_TT_VER` can be used for partitioning, which may violate VPDU limitations.

[0279] As a second example of a partitioning pattern that disallows all trees, partitioning of all trees is not allowed for blocks located at image boundaries and containing samples that extend beyond the image boundaries. When the minimum QT size is greater than the minimum CU size (syntax element log2_min_luma_codign_block_size_minus2 in the previous table) and the maximum BT / TT depth (syntax elements sps_max_mtt_hierarchy_depth_inter_slice, sps_max_mtt_hierarchy_depth_intra_slice_luma, pic_max_mtt_hierarchy_depth_inter_slice, pic_max_mtt_hierarchy_depth_intra_slice_luma, and pic_max_mtt_hierarchy_depth_intra_slice_chroma) equals 0, all variables allowSplitQt, allowSplitBtHor, allowSplitBtHor, allowSplitTtHor, and allowSplitTtVer are set to false.

[0280] Figure 18 Exemplary blocks where QT, BT, or BT partitioning is not allowed at image boundaries are shown according to some embodiments of this disclosure. First, a CTU (e.g., CTU 1801, CTU 1803, or CTU 1805) is partitioned into four 64x64 blocks using a quadtree. Then, each 64x64 block cannot be further partitioned. However, portions of the block extending beyond the right and / or bottom image boundaries are marked in gray, which is not allowed in VVC designs. For blocks marked in gray, the variable allowSpiltQt is set to false due to the following condition (emphasis in italics):

[0281] 6.4.1 Permissible Quadtree Partitioning Process

[0282]

[0283] - AllowSplitQt is set to FALSE if one or more of the following conditions are true:

[0284] -treeTpye is equal to SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY

[0285] -treeTpye equals DUAL_TREE_CHROMS, and cbSize / SubWidthC is less than or equal to MinQtSizeC.

[0286]

[0287] For blocks marked in gray, the variables allowSplitBtHor and allowSplitBtVer are set to false due to the following (italicized for emphasis):

[0288] 6.4.2 Permissible Binary Splitting Procedures

[0289]

[0290] The derivation of the variable allowBtSplit is as follows:

[0291] - AllowBtSplit is set to FALSE if one or more of the following conditions are true:

[0292] -cbSize is less than or equal to MinBtSizeY

[0293] -cbWidth is greater than maxBtSize

[0294] -cbHeight is greater than maxBtSize

[0295] Mttdepth is greater than or equal to maxMttDeptih

[0296]

[0297] For blocks marked in gray, the variables allowSplitTtHor and allowSplitTtVer are set to false due to the following (italicized for emphasis):

[0298] 6.4.3 Permissible ternary splitting processes

[0299]

[0300] The derivation of the variable allowTtSplit is as follows:

[0301] - allowTtSplit is set to FALSE if one or more of the following conditions are true:

[0302] -cbSize is less than or equal to 2 * MinTtSizeY

[0303] -cbWidth is greater than Min(64, maxTtSize).

[0304] -cbHeight is greater than Min(64, maxTtSize)

[0305] -mttDepth is greater than or equal to maxMttDepth ...

[0306] As described in the first exemplary case where all tree partitioning patterns are not allowed, in the current VVC draft 7, a CU may contain samples outside the image boundaries, but under certain conditions, the CU cannot be partitioned. In some embodiments of this disclosure, the QT partitioning conditions in VVC can be modified. In one aspect, for a block containing samples outside the image boundaries and whose width or height is equal to N (e.g., N = 128), QT partitioning is used when the minimum QT size is less than N (e.g., 128). Furthermore, in some embodiments, QT partitioning can also be used when the minimum QT size is equal to N (e.g., 128). In another aspect, using QT partitioning can be more straightforward than using BT or TT partitioning. Figure 19 This is a schematic diagram illustrating examples of BT and TT partitioning according to some embodiments of the present disclosure. Partitioning blocks may require multiple steps. Furthermore, the partitioning may differ for blocks located in different positions, which can be complex. For example, for block 1903, SPLIT_BT_HQR, SPLIT_BT_HOR, SPLIT_TT_VER, and SPLIT_TT_VER are executed sequentially.

[0307] In some embodiments, when block partitioning is not allowed, quadtree partitioning can be used, and it can be inferred that the syntax element split_qt_flag is 1. The syntax element split_qt_flag specifies whether to divide the coding unit into coding units of half size horizontally and vertically.

[0308] When the syntax element split_qt_fiag is not present, the following applies (italicized for emphasis):

[0309] - If all of the following conditions are true, then it is inferred that split_qt_flag equals 1:

[0310] -split_cu_flag equals I

[0311] -allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor and allowSplitTtVer equal FALSE.

[0312] Otherwise, if allowSplitQt equals TRUE, then the value of split_qt_flag is inferred to be 1.

[0313] Otherwise, the value of split_qt_flag will be inferred to be 0.

[0314] In some embodiments, the minimum QT size constraint cannot be applied to blocks located at image boundaries. When a portion of a block extends beyond the bottom or right edge of the image, a quadtree can be used to partition the block. The permitted quadtree partitioning process is described below:

[0315] 6.4.1 Permissible Four-Way Split Process

[0316] The input to this process is:

[0317] - Size of the encoded block in the luminance sample, cbSize

[0318] -Multi-type tree depth mttDepth,

[0319] - Variable tree type, used to specify whether to use a single tree (SINGLE_TREE) or a dual tree to divide the coding tree nodes, and when using a dual tree, whether the current component being processed is luma (DUAL_TREE_LUMA) or chroma (DUAL_TREE_CHROMA).

[0320] - The variable modeType specifies whether intra-frame (MODE_INTRA), IBC (MODE_IBC), and inter-frame coding modes (MODE_TYPE_ALL) can be used, or whether only intra-frame and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter-frame coding modes (MODE_TYPE_INTER) can be used for coding units within coding tree nodes.

[0321] The output of this process is the variable allowSplitQt.

[0322] The derivation of the variable allowSplitQt is as follows (italicized for emphasis):

[0323] - AllowSplitQt is set to true if all of the following conditions are true:

[0324] -treeTpye is equivalent to SINGLE_TREE or DUAL_TREE_LUMA

[0325] -cbSize equals 128

[0326] -MinQtSizeY equals 128

[0327] -x0+cbWidth is greater than pic_width_in_luma_samples or y0+cbHeight is greater than pic_height_inluma_samples

[0328] Otherwise, allowSplitOt is set to true if all of the following conditions are true:

[0329] -treeTpye equals DUAL_TREE_CHROMA

[0330] -CbSize / SubWidthC equals 128

[0331] -MinQtSizeC equals 128

[0332] -x0+cbWidth is greater than pic_width_in_luma_sample, or y0+cbHeight is greater than pic_height_in_luma_samples

[0333] Otherwise, allowSplitQt is set to FALSE if one or more of the following conditions are true:

[0334] -treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY

[0335] -treeType equals DUAL_TREE_CHROMA, and cbSize / SubWidthC is less than or equal to MinQtSizeC

[0336] -mttDepth is not equal to 0

[0337] -treeType equals DUAL_TREE_CHROMA, and (ebSize / SubWidthC) is less than or equal to 4.

[0338] -treeType equals DUAL_TREE_CHROMA, modeType equals MODE_TYPE_INTRA

[0339] - Otherwise, allowSplitQt is set equal to TRUE.

[0340] In some embodiments, bitstream conformance may be added to the syntax of the minimum QT size. It may be required that the minimum QT size be less than or equal to 64.

[0341] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size in the luma samples of the luma leaf blocks resulting from the quadtree partitioning of the CTU and the base-2 logarithm of the minimum coding block size in the luma samples of the luma CUs in slices where the slice_type of the SPS is equal to 2 (I). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference may be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma in the PH that is present in the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. The base-2 logarithm of the minimum size in the luma samples of the luma leaf blocks resulting from the quadtree partitioning of the CTU is derived as follows (italics are used for emphasis):

[0342] MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY

[0343] VSize = Min(64, CtbSizeY)

[0344] In some embodiments, bitstream conformance d requires that the value of (1 << MinQtLog2SizeIntraY) be less than or equal to VSize.

[0345] The syntax element sps_log2_diff_min_qt_min_cb_inter_slice_luma specifies the default difference between the log2 of the minimum size in the luma samples of a luma leaf block produced by the quadtree partitioning of a CTU, and the log2 of the minimum luma coding block size in the luma CU luma samples in a slice where the slice_type in the SPS is equal to 0 (B) or 1 (P). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH in the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_inter_slice is from 0 to the value of CtbLog2SizeY MinCbLog2SizeY, including the end values. The log2 of the minimum size in the luma samples of a luma leaf block produced by the quadtree partitioning of a CTU is derived as follows (italic is used for emphasis):

[0346] MinQtLog2SizeInterY = sps_log2_diff_min_qt_min_cb_inter_slice MinCbLog2SizeY

[0347] VSize = Min(64, CihSizeY)

[0348] In some embodiments, the consistency requirement of the bitstream is that the value of (1 << MinQtLog2SizeInterY) is less than or equal to VSize.

[0349] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the default difference between the log2 of the minimum size in luma samples in chroma leaf blocks resulting from the quadtree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA, and the log2 of the minimum coded block size in luma samples of a chroma CU with treeType equal to DUAL_TREE_CHROMA in a slice with slice_type equal to 2 (I) in the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_chroma present in the PH of the SPS. The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma has a value range from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to 0, and the log2 of the minimum size in luma samples in chroma leaf blocks resulting from the quadtree partitioning of a CTU with treeType equal to DUAL_TREE_CHROMA is derived as follows (italic emphasis):

[0350] MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_chroma + MinCbLog2SizeY

[0351] VSize = Min(64, CthSizeY)

[0352] In some embodiments, the bitstream conformance requirement is that the value of (1 << MinQtLog2SizeIntraC) is less than or equal to VSize.

[0353] The syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_luma` specifies the base-2 logarithm of the smallest size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU, and the base-2 logarithm of the smallest code block size in the luminance CU luminance samples in the slices associated with `p` and `slice_type` equal to 2(I). The value of the syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_luma` is in the range of 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. When it does not exist, the value of the syntax element `pic_log2_diff_min_qt_min_cb_luma` is inferred to be equal to the syntax element `sps_log2_diff_min_qt_min_cb_intra_slice_luma`. In some embodiments, the bitstream consistency requirement (1 << (pic_log2_diff_min_qt_min_cb_intra_slice_luma + MinCblog2SizeY)) is less than or equal to Min(64, CtbSizeY).

[0354] The syntax element `pic_log2_diff_min_qt_min_cb_inter_slice` specifies the difference between the base-2 logarithm of the smallest size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU, and the base-2 logarithm of the smallest luminance coding block size of the luminance CU in the tune points where the slice_type associated with PH is 0 (B) or 1 (P). The value of the syntax element `pic_log2_diff_min_qt_min_cb_inter_slice` ranges from 0 to `CtbLog2SizeY-MinCbLog2SizeY`, inclusive. When it does not exist, the value of the syntax element `pic_log2_diff_min_qt_min_cb_luma` is inferred to be equal to the syntax element `sps_log2_diff_min_qt_min_cb_inter_slice`. In some embodiments, the value of the bitstream consistency requirement (1 << (pic_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY)) is less than or equal to Min(64, CtbSizeY).

[0355] The syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_chroma` specifies the difference, in base 2, between the smallest logarithm of the luminance sample size of the chrominance leaf block generated by partitioning a quadtree of chrominance CTU with `treeType` equal to `DUAL_TREE_CHROMA`, and the smallest logarithm of the smallest coding block size of the luminance sample of chrominance CU with `treeType` equal to `DUAL_TREE_CHROMA` in a slice with `slice_type` equal to 2(I) associated with `PH`. The value of the syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_chroma` is in the range from 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. When it does not exist, the value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma. In some embodiments, the bitstream consistency requirement (1 << (pic_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY)) is less than or equal to Min(64, CtbSizeY).

[0356] In some embodiments, when block partitioning is not allowed in the second example above (where all tree partitioning modes are not allowed), quadtree partitioning can be used, and the syntax element split_qt_flag is inferred to be 1.

[0357] The syntax element `split_qt_flag` specifies whether a coding unit is divided into coding units with half the horizontal and half the vertical dimensions. When the syntax element `split_qt_flag` is not present, the following applies (italicized emphasis):

[0358] - If all of the following conditions are true, then it is inferred that split_qt_flag equals 1:

[0359] -split_cu_flag equals I

[0360] -allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor and allowSplitTtVer equal FALSE.

[0361] Otherwise, if allowSplitQt equals TRUE, then it is inferred that the value of split_qt_flag is equal to 1.

[0362] Otherwise, it is assumed that the value of split_qt_flag is equal to 0.

[0363] In some embodiments, in the second exemplary case described above where all tree partitioning modes are not allowed, bitstream consistency can be added to the syntax for minimum QT size and maximum BT / TT depth.

[0364] The syntax element `sps_log2_diff_min_qt_min_cb_intra_slice_luma` specifies the default difference, based on base 2, between the minimum logarithm of the size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU and the minimum logarithm of the size of the luminance block of the luminance samples in the luminance CUs of the slices involving SPS with slice_type equal to 2(I). When the syntax element `partition_constraints_override_enabled_flag` is equal to 1, the default difference can be overridden by the syntax element `pic_log2_diff_min_qt_min_cb_luma` present in the PH involving SPS. The value of the syntax element `sps_log2_diff_min_qt_min_cb_intra_slice_luma` ranges from 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. The minimum logarithm of the size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU is as follows:

[0365] MinQtLog2SizeIntraY=sps_log2_diff_min_qt_min_cb_intra_slice_luma+MinCbLog2SizeY

[0366] The syntax element `sps_max_mtt_hierarchy_depth_intra_slice_luma` specifies the default maximum hierarchical depth of coding units generated by multi-type tree partitioning of quad-leaf trees in strips with `slice_type` equal to 2(I) involving the SPS. This default maximum hierarchical depth can be overridden by the syntax element `pic_max_mtt_hierarchy_depth_intra_slice_luma` present in the PH involving the SPS when the syntax element `partition_constraints_override_enabled_flag` is equal to 1. The value of the syntax element `sps_max_mtt_hierarchy_depth_intra_slice_luma` ranges from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. In some embodiments, the value of the bitstream consistency requirement (MmQtLog2SizeIntraY - sps_max_mtt_hierachy_depth_intra_slice_luma / 2) is less than or equal to MinCbLog2SizeY.

[0367] The syntax element `sps_log2_diff_min_qt_min_cb_inter_slice` specifies the base-2 logarithm of the minimum size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU, and the base-2 logarithm of the minimum luminance coding block size of the luminance samples in the luminance CUs of the slices involving SPS with `slice_type` equal to 0 (B) or 1 (P). This default difference can be overridden by the syntax element `pic_log2_diff_min_qt_min_cb_luma` present in the PH involving SPS when the syntax element `partition_constraints_override_enabled_flag` is equal to 1. The value of the syntax element `sps_log2_diff_min_qt_min_cb_inter_slice` ranges from 0 to `CtbLog2SizeY-MinCbLog2SizeY`, inclusive. The base-2 logarithm of the minimum size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU is derived as follows:

[0368] MinQtLog2SizeInterY=sps_log2_diff_min_qt_min_cb_inter_slice+MinCbLog2SizeY

[0369] The syntax element `sps_max_mtt_hierarchy_depth_inter_slice` specifies the default maximum hierarchical depth of coding units generated by multi-type tree partitioning of quad-leaf trees in stripes involving SPS with slice_type equal to 0 (B) or 1 (P). When the syntax element `partition_constraints_override_enabled_flag` is equal to 1, the default maximum hierarchical depth can be overridden by the syntax element `pic_max_mtt_hierarchy_depth_inter_slice` present in the PH involving SPS. The value of the syntax element `sps_max_mtt_hierarchy_depth_inter_slice` ranges from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), including end values. In some embodiments, the bitstream consistency requirement (MinQtLog2SizeInterY - sps_max_mtt_hierarchy_depth_inter_slice / 2) is less than or equal to MinCbLog2SizeY.

[0370] The syntax element `sps_log2_diff_min_qt_min_cb_intra_slice_chroma` specifies the base-2 logarithm of the smallest size of the luminance samples from the chrominance CTUs partitioned by a quadtree with `treeType` equal to `DUAL_TREE_CHROMA`, and the base-2 logarithm of the smallest code block size of the luminance samples from the chrominance CUs with `treeType` equal to `DUAL_TREE_CHROMA` in a slice with `slice_type` equal to 2(I) involving SPS, to the default difference. This default difference can be overridden by the syntax element `pic_log2_diff_min_qt_min_cb_chroma` present in the PH involving SPS when the syntax element `partition_constraints_override_enabled_flag` is equal to 1. The syntax element `sps_log2_diff_min_qt_min_cb_intra_slice_chroma` takes values ​​from 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. If it does not exist, the value of `sps_log2_diff_min_qt_min_cb_intra_slice_chroma` is inferred to be 0. The base-2 logarithm derivation of the smallest size of the luminance samples in the chroma leaf blocks generated by the quadtree partitioning of a CTU with `treeType` equal to `DUAL_TREE_CHROMA` is as follows:

[0371] MinQtLog2SizeIntraC=sps_log2_diff_min_qt_min_cb_intra_slice_chroma+MinCbLog2SizeY

[0372] The syntax element `sps_max_mtt_hierarchy_depth_intra_slice_chroma` is the default maximum hierarchical depth of the chroma coding unit generated by multi-type tree partitioning of chroma quadtree leaves with `treeType` equal to `DUAL_TREE_CHROMA` in strips involving `slice_type` equal to 2(I) of the SPS. When the syntax element `partition_constraints_override_enabled_flag` is equal to 1, this default maximum hierarchical depth can be overridden by the syntax element `pic_max_mtt_hierarchy_depth_chroma` present in the `PH` involving the SPS. The value of the syntax element `sps_max_mtt_hierarchy_depth_intra_slice_chroma` is between 0 and 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When it does not exist, the value of the syntax element `sps_max_mtt_hierarchy_depth_intra_slice_chroma` is inferred to be 0. In some embodiments, the bitstream consistency requirement (MinQtLog2SizeTntraC-sps_max_mtt_hierarchy_depth_intra_slice_chroma / 2) is less than or equal to MinCbLog2SizeY.

[0373] The syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_luma` specifies the base-2 logarithm of the minimum size of the luminance samples of the luminance leaf blocks generated by the quadtree partitioning of the CTU, and the base-2 logarithm of the minimum coding block size of the luminance samples of the luminance CUs in the slices associated with `p` and `slice_type` equal to 2(I). The value of the syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_luma` ranges from 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. When it does not exist, it is inferred that the value of the syntax element `pic_log2_diff_min_qt_min_cb_luma` is equal to the syntax element `sps_log2_diff_min_qt_min_cb_intra_slice_luma`.

[0374] The syntax element `pic_max_mtt_hierarchy_depth_intra_slice_luma` specifies the maximum hierarchical depth of the coding unit resulting from a multi-type tree partition of quadrangular leaves in a stripe with a slice_type of 2(I) associated with PH. The value of the syntax element `pic_max_mtt_hierarchy_depth_intra_slice_luma` ranges from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When it does not exist, it is inferred that the value of the syntax element `pic_max_mtt_hierarchy_depth_intra_slice_luma` is equal to that of the syntax element `sps_max_mtt_hierarchy_depth_intra_slice_luma`. In some embodiments, the bitstream consistency requirement (pic_log2_diff_min_qt_min_cb_intra_slice_luma+MinCbLog2SizeY-pic_max_mtt_hierarchy_depth_intra_slice_luma / 2) is less than or equal to MinCbLog2SizeY.

[0375] The syntax element `pic_log2_diff_min_qt_min_cb_inter_slice` specifies the base-2 logarithm of the minimum size of the luminance samples in the luminance leaf blocks generated by the quadtree partitioning of the CTU, and the base-2 logarithm of the minimum luminance coding block size in the luminance samples of the luminance CUs in the slices with a slice_type equal to 0 (B) or 1 (P) associated with PH. The value of the syntax element `pic_log2_diff_min_qt_min_cb_inter_slice` ranges from 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. When it does not exist, the value of the syntax element `pic_log2_diff_min_qt_min_cb_luma` is inferred to be equal to the syntax element `sp_log2_diff_min_qt_min_cb_inter_slice`.

[0376] The syntax element `pic_max_mtt_hierarchy_depth_inter_slice` specifies the maximum hierarchical depth of a coding unit resulting from a multi-type tree partition of quadrangular leaves in a stripe with a slice_type of 0 (B) or 1 (P) associated with PH. The value of the syntax element `pic_max_mtt_hierarchy_depth_inter_slice` ranges from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When it does not exist, the value of the syntax element `pic_max_mtt_hierarchy_depth_inter_slice` is inferred to be equal to the syntax element `sps_max_mtt_hierarchy_depth_inter_slice`. In some embodiments, the bitstream consistency requirement (pic_log2_diff_min_qt_min_cb_inter_slice+MinCbLog2SizeY-pic_max_mtt_hierarchy_depth_inter_slice / 2) is less than or equal to MinCbLog28izeY.

[0377] The syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_chroma` specifies the base-2 logarithm of the minimum size of the luminance samples of the chrominance leaf blocks generated by partitioning a chrominance CTU with `treeType` equal to `DUAL_TREE_CHROMA`, and the base-2 logarithm of the minimum code block size of the luminance samples of the chrominance CU with `treeType` equal to `DUAL_TREE_CHROMA` in the slice with `slice_type` equal to 2(I) associated with `PH`. The value of the syntax element `pic_log2_diff_min_qt_min_cb_intra_slice_chroma` is in the range from 0 to `CtbLog2SizeY - MinCbLog2SizeY`, inclusive. When it does not exist, the value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma.

[0378] The syntax element `pic_max_mtt_hierarchy_depth_intra_slice_chroma` specifies the maximum hierarchical depth of a chroma coding unit, which is generated by a multi-type tree partitioning of chroma quadrilateral leaves with `treeType` equal to `DUAL_TREE_CHROMA` from a slice with `slice_type` equal to 2(I) associated with `PH`. The value of the syntax element `pic_max_mtt_hierarchy_depth_intra_slice_chroma` ranges from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When it does not exist, the value of the syntax element `pic_max_mtt_hierarchy_depth_intra_slice_chroma` is inferred to be equal to the syntax element `sps_max_mtt_hierarchy_depth_intra_slice_chroma`. In some embodiments, the consistency requirement of the bit stream (pic_log2_diff_min_qt_min_cb_intra_slice_chroma+MinCbLog2SizeY-pic_max_mtt_hierarchy_depth_intra_slice_chroma / 2) is less than or equal to MinCbLog2SizeY.

[0379] Figure 20 A flowchart of an exemplary video processing method 2000 according to some embodiments of the present disclosure is shown. Method 2000 may be generated by an encoder (e.g., via...) Figure 2A Process 200A or Figure 2B 200B), decoder (e.g., via Figure 3A Process 300A or Figure 3B Process 300B) or by means of a device (e.g., Figure 4 The device 400) is executed by one or more software or hardware components. For example, a processor (e.g., Figure 4 The processor 402) can execute method 2000. In some embodiments, method 2000 can be implemented by a computer program product contained in a computer-readable medium, the computer program product comprising components implemented by a computer (e.g., a processor 402). Figure 4 The device 400 executes computer-executable instructions, such as program code.

[0380] In step 2001, it can be determined whether the coded block includes samples outside the image boundary. In some embodiments, the image boundary may be the bottom image boundary or the right image boundary. As an example of a coded block that includes samples outside the image boundary, in Figure 17In the image 1700, the coding block 1701 exceeds the right image boundary of the image 1700, the coding block 1705 exceeds the bottom image boundary of the image 1700, and the coding block 1703 exceeds the bottom and right image boundaries of the image 1700.

[0381] In step 2003, in response to the coding block being determined to include samples outside the image boundary, the coding block can be segmented using the QT mode. In some embodiments, in response to the coding block being determined to include samples outside the image boundary, method 2000 can determine that the coding block cannot be segmented using the BT mode and TT mode. For example, the variables allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer can be determined to be equal to FALSE.

[0382] In some embodiments, in response to a coding block being determined to include samples outside image boundaries, method 2000 may determine that the coding block will be segmented using QT mode, regardless of whether a QT flag is present in the bitstream including the coding block. The QT flag indicates whether the coding block is segmented using QT mode. For example, when the syntax element `split_qt_flag` is not present, if the coding block includes samples outside image boundaries and the variables `allowSplitQt`, `allowSplitBtHor`, `allowSplitBtVer`, `allowSplitTtHor`, and `allowSplitTtVer` are equal to `FALSE` or `allowSplitQt` is equal to `TRUE`, then the value of `split_qt_flag` is inferred to be 1.

[0383] In some embodiments, in response to determining that the coding block includes samples outside the image boundary, method 2000 may determine whether the coding block is allowed to be segmented using a QT mode, regardless of a preset constraint on the minimum block size for which the QT mode is allowed to be applied. For example, a minimum QT size constraint cannot be applied to coding blocks located at the image boundary.

[0384] In some embodiments, preset constraints may include bitstream consistency of coded blocks. Bitstream consistency can set the minimum block size that allows QT mode to be applied. For example, the minimum block size can be set to less than or equal to 64. Bitstream consistency can also set a maximum BT depth or a maximum TT depth.

[0385] In some embodiments, method 2000 may include determining that a coded block will be split. For example, the syntax element split_cu_flag may be used to indicate whether a coded block should be split. When the syntax element split_cu_flag is not present, it can be inferred that it is equal to 1, which indicates that the coded block should be split.

[0386] It should be understood that the embodiments of this disclosure may be combined with another embodiment or some other embodiments.

[0387] The embodiments may be further described using the following terms:

[0388] 1. A video processing method, comprising:

[0389] Determine whether the coded block includes samples outside the image boundary; and

[0390] In response to determining that the coded block includes samples outside the image boundary, regardless of the value of the first parameter, quadtree segmentation of the coded block is performed, wherein the first parameter indicates whether the quadtree is allowed to be used to segment the coded block.

[0391] 2. The method described in Clause 1 further includes:

[0392] Determine the value of a first flag of the coded block, the first flag indicating whether the coded block is divided into multiple sub-blocks; and

[0393] The values ​​of the second, third, fourth, and fifth parameters of the encoded block are determined, and the second, third, fourth, and fifth parameters respectively indicate whether it is allowed to use a binary horizontal tree, a binary vertical tree, a ternary horizontal tree, and a ternary vertical tree to split the encoded block.

[0394] 3. The method described in Clause 2 further includes:

[0395] In response to the first flag being equal to 1 and the first, second, third, fourth, and fifth parameters being equal to 0, the second flag of the coded block is set to 1, the second flag indicating whether the coded block is segmented using the quadtree.

[0396] 4. The method described in Clause 2 further includes:

[0397] In response to the first parameter being equal to 1, the value of the second flag of the coded block is set to 1, the second flag indicating whether the coded block is segmented using the quadtree.

[0398] 5. The method according to any one of clauses 2-4 further includes:

[0399] When it is determined that the coded block includes samples outside the image boundary, the value of the first flag is set to 1.

[0400] 6. A video processing apparatus, comprising:

[0401] At least one memory for storing instructions, and

[0402] At least one processor is configured to execute the instructions to cause the device to perform the following operations:

[0403] Determine whether the coded block includes samples outside the image boundary; and

[0404] In response to determining that the coded block includes samples outside the image boundary, a quadtree segmentation of the coded block is performed regardless of the value of the first parameter, wherein the first parameter indicates whether the quadtree is allowed to be used to segment the coded block.

[0405] 7. The apparatus according to Clause 6, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:

[0406] Determine the value of a first flag of the coded block, the first flag indicating whether the coded block is divided into multiple sub-blocks; and

[0407] The values ​​of the second, third, fourth, and fifth parameters of the coded block are determined, and the second, third, fourth, and fifth parameters respectively indicate whether the binary horizontal tree, binary vertical tree, ternary horizontal tree, and ternary vertical tree are allowed to be used to segment the coded block.

[0408] 8. The apparatus according to Clause 7, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:

[0409] In response to the first flag being equal to 1, and the first, second, third, fourth, and fifth parameters being equal to 0, the second flag of the coded block is set to 1. The second flag indicates whether the coded block is segmented using a quadtree.

[0410] 9. The apparatus according to Clause 7, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:

[0411] In response to the first parameter being equal to 1, the value of the second flag of the coded block is set to 1, the second flag indicating whether the coded block is segmented using the quadtree.

[0412] 10. The apparatus according to any one of clauses 7-9, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:

[0413] When the coded block is determined to include samples outside the image boundary, the value of the first flag is set to 1.

[0414] 11. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processing devices to cause a video processing apparatus to perform the following operations:

[0415] Determine whether the coded block includes samples outside the image boundary, and

[0416] In response to the coded block being determined to be a sample included outside the image boundary, a quadtree segmentation of the coded block is performed regardless of the value of the first parameter, wherein the first parameter indicates whether the quadtree is allowed to be used to segment the coded block.

[0417] 12. The non-transitory computer-readable storage medium as described in Clause 11, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform:

[0418] Determine the value of a first flag of the coded block, the first flag indicating whether the coded block is divided into multiple sub-blocks; and

[0419] The values ​​of the second, third, fourth, and fifth parameters of the encoded block are determined, wherein the second, third, fourth, and fifth parameters respectively indicate whether binary horizontal trees, binary vertical trees, ternary horizontal trees, and ternary vertical trees are allowed to be used to segment the encoded block.

[0420] 13. The non-transitory computer-readable storage medium as described in Clause 12, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform:

[0421] In response to the first flag being equal to 1 and the first, second, third, fourth, and fifth parameters being equal to 0, the second flag of the coded block is set to 1, the second flag indicating whether the coded block is segmented using a quadtree.

[0422] 14. The non-transitory computer-readable storage medium as described in Clause 12, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform:

[0423] In response to the first parameter being equal to 1, the value of the second flag of the coded block is set to 1, the second flag indicating whether the coded block is segmented using the quadtree.

[0424] 15. A non-transitory computer-readable storage medium according to any one of clauses 12-14, wherein segmenting the coded block using the QT mode in response to the coded block being determined to include samples outside an image boundary comprises:

[0425] When the coded block is determined to include samples outside the image boundary, the value of the first flag is set to 1.

[0426] 16. A video processing method, comprising:

[0427] Determine whether the coded block includes samples outside the image boundary;

[0428] In response to the coding block being determined to include samples outside the image boundary, the coding block is segmented using a quadtree (QT) pattern.

[0429] 17. The method described in Clause 16 further includes:

[0430] In response to the coded block being determined to be a sample included outside the image boundary, it is determined that binary tree (BT) mode and ternary tree (TT) mode are not allowed to be used to segment the coded block.

[0431] 18. The method according to any one of clauses 16 and 17, wherein segmenting the coded block using a quadtree (QT) pattern in response to the coded block being determined to include samples outside the image boundary comprises:

[0432] Regardless of whether a QT flag is present in the bitstream including the encoded block, the encoded block is segmented using the QT mode, whereby the QT flag indicates whether the encoded block should be segmented using the QT mode.

[0433] 19. The method according to Clause 16, wherein segmenting the coded block using a quadtree (QT) pattern in response to the coded block being determined to include samples outside the image boundary comprises:

[0434] Regardless of any preset constraints on the minimum block size that allow the application of the QT mode, the encoded block is split using the QT mode.

[0435] 20. The method as described in Clause 19, wherein:

[0436] The preset constraints include bitstream consistency associated with the encoded block, and the bitstream consistency setting allows the minimum block size for applying the QT mode.

[0437] 21. The method according to Clause 20, wherein the minimum block size is set to be less than or equal to 64.

[0438] 22. The method according to any one of Clauses 20 and 21, wherein the bitstream consistency setting is a maximum BT depth or a maximum TT depth.

[0439] 23. The method according to any one of Clauses 16-22, wherein the image boundary is the bottom image boundary or the right image boundary.

[0440] 24. A video processing apparatus, comprising:

[0441] At least one memory for storing instructions; and

[0442] At least one processor is configured to execute the instructions to cause the device to perform:

[0443] Determine whether the coded block includes samples outside the image boundary;

[0444] In response to the coding block being determined to include samples outside the image boundary, the coding block is segmented using a quadtree (QT) pattern.

[0445] 25. The apparatus according to clause 24, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:

[0446] In response to the coded block being determined to be a sample included outside the image boundary, it is determined that binary tree (BT) mode and ternary tree (TT) mode are not allowed to be used to segment the coded block.

[0447] 26. The apparatus according to any one of clauses 24 and 25, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform;

[0448] Regardless of whether a QT flag is present in the bitstream including the encoded block, the encoded block is segmented using the QT mode, whereby the QT flag indicates whether the encoded block should be segmented using the QT mode.

[0449] 27. The apparatus according to clause 24, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:

[0450] Regardless of any preset constraints on the minimum block size that allow the application of the QT mode, the QT mode is used to segment the coded blocks.

[0451] 28. The equipment as described in Clause 27, wherein

[0452] The preset constraints include bitstream consistency associated with the encoded block, and the bitstream consistency setting allows the minimum block size for applying the QT mode.

[0453] 29. The apparatus according to Clause 28, wherein the minimum block size is set to be less than or equal to 64.

[0454] 30. The apparatus according to any one of clauses 28 and 29, wherein the bit stream consistency setting is a maximum BT depth or a maximum TT depth.

[0455] 31. The apparatus according to any one of clauses 24 to 30, wherein the image boundary is a bottom image boundary or a right image boundary.

[0456] 32. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processing devices to cause a video processing apparatus to perform the following operations:

[0457] Determine whether the coded block includes samples outside the image boundary;

[0458] In response to the coding block being determined to include samples outside the image boundary, the coding block is segmented using a quadtree (QT) pattern.

[0459] 33. The non-transitory computer-readable storage medium as described in Clause 32, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform:

[0460] In response to the coded block being determined to include samples outside the image boundary, it is determined that binary tree (BT) mode and ternary tree (TT) mode are not allowed to be used to segment the coded block.

[0461] 34. A non-transitory computer-readable storage medium according to any one of clauses 32 and 33, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform:

[0462] Regardless of whether a QT flag is present in the bitstream including the coded block, the coded block is segmented using the QT mode, whereby the QT flag indicates whether the coded block should be segmented using the QT mode.

[0463] 35. The non-transitory computer-readable storage medium as described in Clause 32, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform:

[0464] Regardless of any preset constraints on the minimum block size that allow the application of the QT mode, the coded blocks are segmented using the QT mode.

[0465] 36. The non-transitory computer-readable storage medium as described in Clause 35, wherein: the preset constraint includes bitstream consistency associated with the coded block, the bitstream consistency setting allowing the minimum block size for applying the QT mode.

[0466] 37. The non-lateral computer-readable storage medium as described in Clause 36, wherein the minimum block size is set to be less than or equal to 64.

[0467] 38. A non-transitory computer-readable storage medium according to any one of clauses 36 and 37, wherein the bit stream consistency setting is a maximum BT depth or a maximum TT depth.

[0468] 39. The non-transitory computer-readable storage medium according to any one of clauses 32 to 38, wherein the image boundary is the bottom image boundary or the right image boundary.

[0469] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and these instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the methods described above. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, caches, registers, any other memory chips or cassettes, and their networking versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0470] It should be noted that relational terms in this document, such as “first” and “second”, are used only to distinguish an entity or operation from another entity or operation and do not require or imply any actual relationship or order between these entities or operations. Furthermore, the words “contains,” “has,” “includes,” and “includes,” as well as other similar forms, are semantically equivalent and open-ended, because one or more entries following any of these words do not imply an exhaustive list of those entries or entries, or are limited to only the listed entries or entries.

[0471] As used herein, unless otherwise specified, the term "or" includes all possible combinations unless impractical. For example, if a database is declared to include A or B, then unless otherwise expressly stated or impractical, the database may include A, or B, or A and B. As a second example, if a database is declared to include A, B, or C, then unless otherwise expressly stated or impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0472] It should be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that the above-described multiple modules / units can be combined into one module / unit, and each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0473] In the foregoing description, numerous specific details have been described with reference to embodiments, which may vary as implementation progresses. Certain modifications and changes may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims herein. The sequence of steps shown in the accompanying drawings is for illustrative purposes only and is not intended to limit one to any particular sequence of steps. Therefore, those skilled in the art will understand that these steps may be performed in a different order while implementing the same method.

[0474] Exemplary embodiments have been disclosed in the accompanying drawings and description. However, many variations and modifications can be made to these embodiments. Therefore, although specific terminology has been used, it is used only in a general and descriptive sense and not for limiting purposes.

Claims

1. A method of video processing applied to an encoder, comprising: determining a minimum size of a plurality of luma samples of a luma leaf block resulting from a quad-tree partitioning of a coding tree unit for an inter slice; determining a minimum luma coding block size; determining whether a coding block includes samples outside of a picture boundary; determining a size of the coding block; in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quad-tree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, the minimum luma coding block size being equal to the first value, and the coding block including samples outside of a picture boundary, splitting the coding block into coding blocks having a horizontal size and a vertical size that are half.

2. The method of claim 1, wherein, the first value is 128.

3. The method of claim 1, wherein, in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quad-tree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, the minimum luma coding block size being equal to the first value, and the coding block including samples outside of a picture boundary, determining a first flag signaled in a bitstream has a second value, wherein the first flag equal to the second value indicates that each of a plurality of coding blocks is split into coding blocks having a horizontal size and a vertical size that are half.

4. The method of claim 3, wherein, the second value is equal to 1.

5. The method of claim 3, wherein, the first flag comprises split_qt_flag.

6. The method of claim 1, further comprising: determining values of first, second, third, fourth, and fifth parameters associated with the coding block, the parameters respectively indicating whether a quad split, a binary horizontal split, a binary vertical split, a ternary horizontal split, and a ternary vertical split is allowed to be used to split the coding block.

7. The method of claim 6, further comprising: in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quad-tree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, determining that the values of the first, second, third, fourth, and fifth parameters are all equal to a third value.

8. The method of claim 6, wherein, the minimum size of the plurality of luma samples of the luma leaf block resulting from the quad-tree partitioning of the coding tree unit being equal to a first value, the values of the first, second, third, fourth, and fifth parameters all being equal to a third value, indicating that none of the quad split, the binary horizontal split, the binary vertical split, the ternary horizontal split, or the ternary vertical split is allowed to be used to split the coding block.

9. The method of claim 1, further comprising: determining that the coding block is split into a plurality of coding blocks based on a value of a first flag signaled in a bitstream.

10. The method of claim 9, wherein, the first flag is equal to 1.

11. The method of claim 9, wherein, the first flag comprises split_qt_flag.

12. A system for video processing, comprising: a memory storing a set of instructions; at least one processor configured to execute the set of instructions to cause the system to perform: determining a minimum size of a plurality of luma samples of a luma leaf block resulting from a quad-tree partitioning of a coding tree unit for an inter slice; determining a minimum luma coding block size; determining whether the coding block includes samples outside of the picture boundary; determining a size of the coding block; in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quadtree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, the minimum luma coding block size being equal to the first value, and the coding block including samples outside of the picture boundary, partitioning the coding block into coding blocks having a horizontal size and a vertical size that is half.

13. The video processing system of claim 12, wherein, the first value is equal to 128.

14. A non-transitory computer-readable storage medium storing a set of instructions and a video bitstream, the set of instructions executable by one or more processors to perform a method to generate the video bitstream, the method comprising: determining a minimum size of a plurality of luma samples of a luma leaf block resulting from the quadtree partitioning of the coding tree unit for an inter slice; determining a minimum luma coding block size; determining whether the coding block includes samples outside of the picture boundary; determining a size of the coding block; in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quadtree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, the minimum luma coding block size being equal to the first value, and the coding block including samples outside of the picture boundary, partitioning the coding block into coding blocks having a horizontal size and a vertical size that is half.

15. The non-transitory computer-readable storage medium of claim 14, wherein, the first value is 128.

16. The non-transitory computer-readable storage medium of claim 14, wherein, the video bitstream includes the first value, the method further comprising: in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quadtree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, the minimum luma coding block size being equal to the first value, and the coding block including samples outside of the picture boundary, determining that a first flag transmitted in the video bitstream has a second value, wherein the first flag equal to the second value indicates that each of the plurality of coding blocks is partitioned into coding blocks having a horizontal size and a vertical size that is half.

17. The non-transitory computer-readable storage medium of claim 16, wherein the second value is equal to 1.

18. The non-transitory computer-readable storage medium of claim 16, wherein the first flag comprises split_qt_flag.

19. The non-transitory computer-readable storage medium of claim 14, the method further comprising: determining values of first, second, third, fourth, and fifth parameters associated with the coding block, the parameters respectively indicating whether use of quad partitioning, binary horizontal partitioning, binary vertical partitioning, ternary horizontal partitioning, and ternary vertical partitioning is allowed to partition the coding block.

20. The non-transitory computer-readable storage medium of claim 14, the method further comprising: in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quadtree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, determining that the values of the first, second, third, fourth, and fifth parameters are all equal to a third value.

21. A method of video processing, applied to a decoder, comprising: determining a minimum size of a plurality of luma samples of a luma leaf block resulting from a quad-tree partitioning of a coding tree unit for an inter slice; determining a minimum luma coding block size; determining whether a coding block includes samples outside of a picture boundary; determining a size of the coding block; in response to the minimum size of the plurality of luma samples of the luma leaf block resulting from the quad-tree partitioning of the coding tree unit being equal to a first value, the size of the coding block being equal to the first value, the minimum luma coding block size being equal to the first value, and the coding block including samples outside of a picture boundary, partitioning the coding block into coding blocks having a horizontal size and a vertical size that are half of the coding block.

Citation Information

Patent Citations

  • METHOD of DECODING coding units from bitstream of video data

    CN108810540A

  • Signaling of quantization information in non-quadtree-only partitioned video coding

    CN109479140A