Method and apparatus for dividing blocks at image boundaries
By forcibly applying quadtree segmentation technology at image boundaries, the problem of low block partitioning efficiency in existing video coding standards is solved, coding efficiency is improved, and the high-efficiency compression requirements of the VVC/H.266 standard are met.
Patent Information
- Application Number
- CN202510449475.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-17
- Filing Date
- 2020-11-24
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2040-11-24
AI Technical Summary
Existing video coding standards struggle to efficiently divide blocks when dealing with image boundaries, resulting in low coding efficiency. This is especially true in emerging, high-efficiency video coding standards such as VVC/H.266, where block division at image boundaries still presents an efficiency bottleneck.
The quadtree segmentation technique is used to force segmentation of coding blocks that contain samples outside the image boundary, regardless of the first parameter, ensuring the application of quadtree segmentation and improving the segmentation efficiency of coding blocks.
By forcing quadtree segmentation, the coding efficiency at image boundaries is improved, the overall compression performance of video coding is enhanced, and the high-efficiency compression requirements of emerging video coding standards are met.
Smart Images

Figure CN120223886B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This disclosure claims priority to U.S. Provisional Application No. 62 / 948,856, filed on December 17, 2019, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to video processing, and more specifically, to methods and apparatus for dividing images into blocks at boundaries. Background Technology
[0004] Video is a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, the most common being prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards, such as High Efficiency Video Coding (HEVC / H.265), Universal Video Coding (VVC / H.266), and AVS, specify particular video coding formats and are developed by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention
[0005] In some embodiments, an exemplary video processing method includes: determining whether a coded block includes samples outside an image boundary; and in response to the coded block being determined to include samples outside an image boundary, performing quadtree segmentation of the coded block regardless of the value of a first parameter, wherein the first parameter indicates whether the quadtree is allowed to be used to segment the coded block.
[0006] In some embodiments, an exemplary video processing apparatus includes at least one memory for storing instructions and at least one processor. The at least one processor is configured to execute the instructions to cause the apparatus to: determine whether a coded block includes samples outside an image boundary; and, in response to the coded block being determined to include samples outside an image boundary, perform quadtree segmentation of the coded block regardless of the value of a first parameter, wherein the first parameter indicates whether the quadtree is permitted for segmenting the coded block.
[0007] In some embodiments, an example non-transitory computer-readable storage medium stores a set of instructions. The set of instructions can be executable by one or more processing devices to cause a video processing device to perform determining whether a coding block includes samples outside of an image boundary, and responsive to the coding block being determined to include samples outside of an image boundary, performing quad-tree partitioning of the coding block regardless of a value of a first parameter, wherein the first parameter indicates whether the quad-tree is allowed to be used for partitioning the coding block. BRIEF DESCRIPTION OF DRAWINGS
[0008] Embodiments and aspects of the disclosure are illustrated by way of example in the following detailed description and in conjunction with the figures. Various features shown in the figures were not necessarily drawn to scale.
[0009] Figure 1 is a structural diagram of an example video sequence in accordance with some embodiments of the disclosure.
[0010] Figure 2A is a diagram illustrating an example encoding process of a hybrid video coding system in accordance with embodiments of the disclosure.
[0011] Figure 2B is a diagram illustrating another example encoding process of a hybrid video coding system in accordance with embodiments of the disclosure.
[0012] Figure 3A is a diagram illustrating an example decoding process of a hybrid video coding system in accordance with embodiments of the disclosure.
[0013] Figure 3B is a diagram illustrating another example decoding process of a hybrid video coding system in accordance with embodiments of the disclosure.
[0014] Figure 4 is a block diagram of an example apparatus for encoding or decoding a video in accordance with some embodiments of the disclosure.
[0015] Figure 5 is a diagram illustrating an example of multi-type tree split modes in accordance with some embodiments of the disclosure.
[0016] Figure 6 is a diagram illustrating an example signaling mechanism for partitioning information in a quad-tree (QT) with nested multi-type tree coding tree structure in accordance with some embodiments of the disclosure.
[0017] Figure 7 is an example Table 1 illustrating example multi-type tree split mode (MttSplitMode) derivation based on multi-type tree syntax elements in accordance with some embodiments of the disclosure.
[0018] Figure 8FIG. 1 is a diagram illustrating an example of a not allowed triple tree (TT) and binary tree (BT) split according to some embodiments of the disclosure.
[0019] Figure 9 FIG. 2 is a diagram illustrating an example of a block split at an image boundary according to some embodiments of the disclosure.
[0020] Figure 10 FIG. 3 illustrates an example Table 2 showing example specifications of parallel triple tree split (parallelTtSplit) and coding block size (cbSize) based on binary split mode (btSplit) according to some embodiments of the disclosure.
[0021] Figure 11 FIG. 4 illustrates an example Table 3 showing example specifications of cbSize based on triple tree split mode (ttSplit) according to some embodiments of the disclosure.
[0022] Figure 12 FIG. 5 illustrates an example Table 4 showing example coding tree syntax according to some embodiments of the disclosure.
[0023] Figure 13 FIG. 6 is a diagram illustrating an example of multi-type tree split mode indicated by MttSplitMode according to some embodiments of the disclosure.
[0024] Figure 14 FIG. 7 illustrates an example Table 5 showing example specifications of MttSplitMode according to some embodiments of the disclosure.
[0025] Figure 15 FIG. 8 illustrates an example Table 6 showing example sequence parameter set RBSP syntax according to some embodiments of the disclosure.
[0026] Figure 16 FIG. 9 illustrates an example Table 7 showing example specifications of picture header RBSP syntax according to some embodiments of the disclosure.
[0027] Figure 17 FIG. 10 is a diagram illustrating an example of a block for which QT, TT, or BT split is not allowed at an image boundary according to some embodiments of the disclosure.
[0028] Figure 18 FIG. 11 is a diagram illustrating an example of a block for which QT, TT, or BT split is not allowed at an image boundary according to some embodiments of the disclosure.
[0029] Figure 19 FIG. 12 is a diagram illustrating an example of using BT and TT split according to some embodiments of the disclosure.
[0030] Figure 20 A flowchart illustrating an exemplary video processing method according to some embodiments of the disclosure is shown. DETAILED DESCRIPTION
[0031] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers represent the same or similar elements unless otherwise represented. The implementation set forth in the following description of exemplary embodiments does not represent all of the implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present disclosure as described in the appended claims. Certain aspects of the present disclosure are described in more detail below. If there is a contradiction between the claims and the
[0032] The Joint Video Expert Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality using half the bandwidth of HEVC / H.265.
[0033] To achieve the same subjective quality using half the bandwidth of HEVC / H.265, the JVET has been developing techniques beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM achieves higher coding performance than HEVC.
[0034] The VVC standard is recently developed and continues to include more coding techniques that provide better compression performance. VVC is based on the hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0035] A video is a set of still images (or “frames”) arranged in a time sequence to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images in a time sequence, and video playback devices (e.g., televisions, computers, smartphones, tablet computers, video players, or any end-user terminal with display functionality) can be used to display such images in a time sequence. Furthermore, in some applications, video capture devices can transmit captured videos to video playback devices (e.g., computers with monitors) in real time, for example, for surveillance, conferencing, or live broadcasting.
[0036] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission, and decompressed before display. The compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general purpose computer) or by specialized hardware. The module for compression is generally referred to as an “encoder”, and the module for decompression is generally referred to as a “decoder”. The encoder and the decoder can be collectively referred to as a “codec”. The encoder and the decoder can be implemented in any of a variety of suitable hardware, software, or combination thereof. For example, the hardware implementation of the encoder and the decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic or any combination thereof. The software implementation of the encoder and the decoder can include program code, computer executable instructions, firmware or any suitable computer-implemented algorithm or process fixed in a computer readable medium. The video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, the codec can decompress the video from a first encoding standard, and recompress the decompressed video using a second encoding standard, in which case the codec can be referred to as a “transcoder”.
[0037] The video encoding process can identify and retain useful information that can be used to reconstruct the image, and ignore unimportant reconstruction information. If the unimportant information cannot be completely reconstructed, such an encoding process can be referred to as “lossy”. Otherwise, it can be referred to as “lossless”. Most encoding processes are lossy, which is a tradeoff for reducing the required storage space and transmission bandwidth.
[0038] The useful information of an encoded image (referred to as “current image”) includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include position changes, luminance changes, or color changes of pixels, among which the position changes are the most concerned. The position changes of a group of pixels representing an object can reflect the motion of the object between the reference image and the current image.
[0039] An image that is encoded without reference to another image (i.e., it is its own reference image) is referred to as an “I-image”. An image that is encoded using a previous image as the reference image is referred to as a “P-image”, and an image that is encoded using a previous image and a future image as the reference images is referred to as a “B-image” (the reference is “bidirectional”).
[0040] Figure 1The structure of an example video sequence 100 is shown in accordance with some embodiments of the present disclosure. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination of both (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), an archive containing previously captured videos (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives videos from a video content provider.
[0041] As shown in Figure 1 , the video sequence 100 can include a series of images arranged in time along a time line, including images 102, 104, 106, and 108. The images 102-106 are consecutive, with more images in between image 106 and 108. In Figure 1 , image 102 is an I-image, whose reference image is image 102 itself. Image 104 is a P-image, whose reference image is image 102, as shown by the arrow. Image 106 is a B-image, whose reference images are images 104 and 108, as shown by the arrows. In some embodiments, a reference image of an image (e.g., image 104) can not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It is noted that the reference images of images 102-106 are merely examples, and the present disclosure is not limited to the embodiments of reference images as shown in Figure 1 .
[0042] Generally, due to the computational complexity of the coding task, a video codec does not encode or decode an entire image at once. Instead, they can partition an image into basic segments and encode or decode the image segments one by one. In the present disclosure, such a basic segment is referred to as a basic processing unit (“BPU”). For example, Figure 1Structure 110 in FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, a basic processing unit can be referred to as a “macroblock” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (“CTU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have a variable size in pixels, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size. The size and shape of a basic processing unit can be selected for a picture based on a balance of coding efficiency and level of detail to be maintained in the basic processing unit.
[0043] A basic processing unit can be a logical unit that can include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components can be referred to as “coding tree blocks” (“CTBs”). Any operation performed on a basic processing unit can be repeated on each of its luma and chroma components.
[0044] Video coding has multiple stages of operations, examples of which are shown in Figures 2A-2B and Figures 3A-3BThe basic processing unit can be too large for processing, so it can be further divided into segments, referred to in this disclosure as“basic processing subunits.” In some embodiments, the basic processing subunits can be referred to as“blocks” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as“coding units” (“CUs”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits can have the same size as the basic processing unit or have a smaller size than the basic processing unit. Like the basic processing unit, the basic processing subunit is also a logical unit, which can include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operations performed on the basic processing subunit can be repeated for each of its luma and chroma components. It should be noted that such division can be performed to further levels as needed for processing. It should also be noted that different stages can use different schemes to divide the basic processing unit.
[0045] For example, at the mode decision stage (examples of which are shown in Figure 2B The encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for the basic processing unit, which can be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., as CUs in H.265 / HEVC or H.266 / VVC), and decide the prediction type for each individual basic processing subunit.
[0046] For another example, at the prediction stage (examples of which are shown in Figures 2A-2B The encoder can perform the prediction operation at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit can still be too large for processing. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as“prediction blocks” or“PBs” in H.265 / HEVC or H.266 / VVC), at which level the prediction operation can be performed.
[0047] For another example, at the transform stage (examples of which are shown in Figures 2A-2BAs shown in FIG. 1, the encoder can perform a transform operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit can still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as “transform blocks” or “TBs” in H.265 / HEVC or H.266 / VVC), on which the transform operation can be performed. It is noted that the division scheme of the same basic processing subunit can be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and the transform blocks of the same CU can have different sizes and numbers.
[0048] In Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3x3 basic processing subunits, the boundaries of which are shown in dashed lines. Different basic processing units of the same picture can be divided into basic processing subunits in different schemes.
[0049] In some embodiments, to provide the capability of parallel processing of video encoding and decoding and the capability of fault tolerance, a picture can be divided into regions for processing such that, for a region of the picture, the encoding or decoding process can not depend on information from any other region of the picture. In other words, each region of the picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, thereby improving the encoding efficiency. Furthermore, when the data of a region is corrupted in processing or lost in network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing the capability of fault tolerance. In some video coding standards, a picture can be divided into regions of different types. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: “slices” and “tiles”. It is also noted that different pictures of the video sequence 100 can have different division schemes for dividing the pictures into regions.
[0050] For example, in Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3x3 basic processing subunits, the boundaries of which are shown in dashed lines. Different basic processing units of the same picture can be divided into basic processing subunits in different schemes. Figure 1 It is noted that the basic processing units, the basic processing subunits, and the regions of the structure 110 in
[0051] Figure 2A A schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure is shown. For example, the encoding process 200A can be performed by an encoder. As shown in FIG. 2A, the encoding process 200A can include the following steps. Figure 2AAs shown, the encoder can encode the video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 in Figure 1 , the video sequence 202 can include a set of pictures (referred to as “original pictures”) arranged in a temporal order. Similar to the structure 110 in Figure 1 , each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform the process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder can perform the process 200A in an iterative manner, where the encoder can encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder can perform the process 200A in parallel for regions (e.g., regions 114-118) of each original picture of the video sequence 202.
[0052] Referring to Figure 2A , the encoder can feed a basic processing unit of an original picture of the video sequence 202 (referred to as an “original BPU”) to the prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder can feed the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder can feed the prediction data 206 and the quantized transform coefficients 216 to the binary encoding stage 226 to generate the video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the “forward path.” During the process 200A, after the quantization stage 214, the encoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of the process 200A. The components 218, 220, 222, and 224 of the process 200A can be referred to as the “reconstruction path.” The reconstruction path can be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0053] The encoder can iteratively perform the process 200A to encode each original BPU of an original picture (in the forward path) and generate a prediction reference 224 for the next original BPU of the original picture (in the reconstruction path) to be encoded. After all original BPUs of an original picture are encoded, the encoder can proceed to encode the next picture in the video sequence 202.
[0054] Referring to process 200A, an encoder can receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term “receive” can refer to any action that gets, obtains, retrieves, acquires, reads, accesses, or otherwise inputs data in any manner.
[0055] At a prediction stage 204, at a current iteration, the encoder can receive an original BPU and a prediction reference 224, and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting, from the prediction data 206 and the prediction reference 224, prediction data 206 that can be used to reconstruct the original BPU into the predicted BPU 208.
[0056] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To account for these differences, upon generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the values (e.g., grayscale or RGB values) of the corresponding pixels of the predicted BPU 208 from the values of the pixels of the original BPU. Each pixel of the residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. The prediction data 206 and the residual BPU 210 can have a smaller number of bits compared to the original BPU, but they can be used to reconstruct the original BPU without a noticeable quality degradation. Thus, the original BPU is compressed,
[0057] To further compress the residual BPU 210, at a transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional “base patterns.” Each base pattern is associated with a “transform coefficient.” The base patterns can have the same size (e.g., the size of the residual BPU 210), and each base pattern can represent a component of the residual BPU 210 in terms of frequency of variation (e.g., frequency of luminance variation). None of the base patterns can be reproduced from any combination (e.g., linear combination) of any of the other base patterns. In other words, the decomposition can decompose the variations of the residual BPU 210 into the frequency domain. This decomposition is similar to a discrete Fourier transform of a function, where the base images are analogous to the base functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the base functions.
[0058] Different transform algorithms can use different basis patterns. Various transform algorithms can be used at transform stage 212, e.g., discrete cosine transform, discrete sine transform, etc. The transform at transform stage 212 is invertible. That is, the encoder can recover the residual BPU 210 through an inverse operation of the transform, referred to as an “inverse transform.” For example, to recover a pixel of the residual BPU 210, the inverse transform can be multiplying the values of the corresponding pixels of the basis pattern by the respective correlation coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transform algorithm (and thus have the same basis pattern). Thus, the encoder can record only the transform coefficients from which the decoder can reconstruct the residual BPU 210 without receiving the basis pattern from the encoder. The transform coefficients can have fewer bits than the residual BPU 210, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.
[0059] The encoder can further compress the transform coefficients at quantization stage 214. During the transform process, different basis patterns can represent different frequencies of variation (e.g., frequencies of luminance variation). Because the human eye is generally better at recognizing low-frequency variations, the encoder can ignore information of high-frequency variations without causing noticeable quality degradation in decoding. For example, at quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value, referred to as a “quantization parameter,” and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the quantized transform coefficients 216 of zero value, whereby the transform coefficients are further compressed. This quantization process is also invertible, where the quantized transform coefficients 216 can be reconstructed to transform coefficients in an inverse operation of quantization, referred to as “inverse quantization.”
[0060] Because the encoder ignores the remainder of the division in the rounding operation, quantization stage 214 can be lossy. Generally, quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer the number of bits required for the quantized transform coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.
[0061] At the binarization stage 226, the encoder can binarize the prediction data 206 and the quantized transform coefficients 216 using a binarization technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can binarize other information at the binarization stage 226, such as the prediction modes used at the prediction stage 204, parameters of the prediction operations, the type of transform at the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. The encoder can use the output data of the binarization stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 can be further packetized for network transmission.
[0062] Referring to the reconstruction path of the process 200A, at the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of the process 200A.
[0063] It should be noted that other variants of the process 200A can be used to encode the video sequence 202. In some embodiments, the stages of the process 200A can be performed by the encoder in a different order. In some embodiments, one or more stages of the process 200A can be combined into a single stage. In some embodiments, a single stage of the process 200A can be split into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, the process 200A can include additional stages. In some embodiments, the process 200A can omit one or more stages of the process 200A. Figure 2A
[0064] Figure 2B A schematic diagram illustrating another example encoding process 200B is shown, in accordance with an embodiment of the disclosure. The process 200B can be a modification of the process 200A. For example, the process 200B can be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x family of standards). In comparison to the process 200A, the forward path of the process 200B further includes a mode decision stage 230 and splits the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and the reconstruction path of the process 200B further additionally includes a loop filtering stage 232 and a buffer 234.
[0065] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or “intra-prediction”) can use pixels from one or more already encoded neighboring BPU in the same image to predict a current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPU. Spatial prediction can reduce spatial redundancy inherent in images. Temporal prediction (e.g., inter-image prediction or “inter-prediction”) can use regions from one or more already encoded images to predict a current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce temporal redundancy inherent in images.
[0066] Referring to the process 200B, in the forward path, the encoder performs prediction operations at a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the encoder can perform intra-prediction. For an original BPU of an image being encoded, the prediction reference 224 can include one or more neighboring BPU in the same image that have been encoded (in the forward path) and reconstructed (in the reconstruction path). The encoder can generate a predicted BPU 208 by interpolating the neighboring BPU. Interpolation techniques can include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at a pixel level, e.g., by interpolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPU used for interpolation can be located in various directions relative to the original BPU, e.g., in a vertical direction (e.g., at the top of the original BPU), a horizontal direction (e.g., at the left of the original BPU), a diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the neighboring BPU used, the size of the neighboring BPU used, parameters for interpolation, the direction of the neighboring BPU used relative to the original BPU, etc.
[0067] For another example, at the temporal prediction stage 2044, the encoder can perform inter prediction. For the original BPU of the current image, the prediction reference 224 can include one or more images (referred to as “reference images”) that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images can be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same image are generated, the encoder can generate a reconstructed image as a reference image. The encoder can perform an operation of “motion estimation” to search for a matching region in a range (referred to as a “search window”) of the reference image. The location of the search window in the reference image can be determined based on the location of the original BPU in the current image. For example, the search window can be centered at a location in the reference image that has the same coordinates as the original BPU in the current image, and can extend outward by a predetermined distance. When the encoder identifies (e.g., by using a pel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder can determine such a region as a matching region. The matching region can have a different size (e.g., smaller, equal, larger, or have a different shape) than the original BPU. Because the reference image and the current image are separated in time on a timeline (e.g., as shown in FIG. 1), the matching region can be considered to “move” to the location of the original BPU over time. The encoder can record the direction and distance of such motion as a “motion vector.” When multiple reference images are used (e.g., as in FIG. 1), the encoder can search for matching regions and determine their associated motion vectors for each reference image. In some embodiments, the encoder can assign weights to the pixel values of the matching regions of the respective reference images. Figure 1 Figure 1
[0068] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0069] To generate the predicted BPU 208, the encoder can perform an operation of “motion compensation.” Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference image according to the motion vector, where the encoder can predict the original BPU of the current image. When multiple reference images are used (e.g., as in FIG. 1), the encoder can move the matching region of each reference image according to its associated motion vector, where the encoder can predict the original BPU of the current image. Figure 1 In some embodiments, the encoder can move the matching region of the reference image according to the individual motion vectors and average pixel values of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of the individual matching reference images, the encoder can add the weighted sum of the pixel values of the moved matching region.
[0070] In some embodiments, inter prediction can be uni-directional or bi-directional. Uni-directional inter prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in FIG. 1 is a uni-directional inter prediction image, where the reference image (i.e., image 102) precedes image 104. Bi-directional inter prediction can use one or more reference images in both temporal directions relative to the current image. For example, Figure 1 Image 106 in FIG. 1 is a bi-directional inter prediction image, where the reference images (i.e., images 104 and 108) precede image 106 in both temporal directions.
[0071] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stage 2044, at a mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, where the encoder can select the prediction mode to minimize the value of a cost function according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate the corresponding predicted BPU 208 and prediction data 206.
[0072] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all the BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the in-loop filter stage 232. At this stage, the encoder can apply in-loop filters to the prediction reference 224 to reduce or eliminate the distortion (e.g., blockiness artifacts) introduced by inter prediction. The encoder can apply various in-loop filter techniques at the in-loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The in-loop filtered reference picture can be stored in the buffer 234 (or “decoded picture buffer”) for later use (e.g., as an inter prediction reference picture for future pictures of the video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode parameters of the in-loop filters (e.g., in-loop filter strength) at the binary encoding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.
[0073] Figure 3A A schematic diagram illustrating an example decoding process 300A in accordance with an embodiment of the present application is shown. The process 300A can be a decompression process corresponding to the compression process 200A in Figure 2A some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in Figures 2A-2B some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in Figures 2A-2B some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in
[0074] As described above with respect to the process 200A in Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded image (referred to as an "encoded BPU") to a binarization decoding stage 302, where the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder can feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of the process 300A.
[0075] The decoder can iteratively perform the process 300A to decode each encoded BPU of an encoded image and generate a prediction reference 224 for a next encoded BPU of the encoded image. After decoding all encoded BPUs of an encoded image, the decoder can output the image to a video stream 304 for display and continue decoding a next encoded image in the video bitstream 228.
[0076] At the binarization decoding stage 302, the decoder can perform an inverse operation of the binarization encoding technique used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information at the binarization decoding stage 302, such as prediction modes, parameters of prediction operations, transform types, parameters of quantization processes (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder can depacketize the video bitstream 228 before feeding it to the binarization decoding stage 302.
[0077] Figure 3B A schematic diagram illustrating another example decoding process 300B according to embodiments of the disclosure is shown. The process 300B can be a modification of the process 300A. For example, the process 300B can be used by a decoder conforming to a hybrid video coding standard (e.g., the H.26x family of standards). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.
[0078] In process 300B, for a coded base processing unit (referred to as "current BPU") of a decoded coded picture (referred to as "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on what prediction mode is used by the encoder to code the current BPU. For example, if the current BPU is coded using intra prediction by the encoder, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the intra prediction, parameters of the intra prediction operation, and the like. The parameters of the intra prediction operation can include, for example, locations (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of interpolation, directions of the neighboring BPUs relative to the original BPU, and the like. For another example, if the current BPU is coded using inter prediction by the encoder, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the inter prediction, parameters of the inter prediction operation, and the like. The parameters of the inter prediction operation can include, for example, a number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions in the respective reference pictures, one or more motion vectors respectively associated with the matching regions, and the like.
[0079] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) at the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) at the temporal prediction stage 2044, details of performing such spatial or temporal prediction are described in Figure 2B , which will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208, to which the decoder can add the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A .
[0080] In process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction at the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If the current BPU is decoded using inter prediction at the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can feed the prediction reference 224 to the loop filter stage 232 as described in Figure 2BThe illustrated manner applies the loop filter to the prediction reference 224. The loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data can further include parameters of the loop filter (e.g., loop filter strength).
[0081] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding a video according to embodiments of the present disclosure. As Figure 4 illustrated, the apparatus 400 can include a processor 402. When the processor 402 executes instructions as described herein, the apparatus 400 can become a special purpose machine for video encoding or decoding. The processor 402 can be any type of circuitry capable of manipulating or processing information. For example, the processor 402 can include any combination of central processing units (or “CPUs”), graphics processing units (or “GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), a field-programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), and the like. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as Figure 4 illustrated, the processor 402 can include multiple processors, including a processor 402a, a processor 402b, and a processor 402n.
[0082] The apparatus 400 can also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, and the like). For example, as Figure 4As shown, the stored data can include program instructions (e.g., for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 can include a high-speed random access memory or a nonvolatile memory device. In some embodiments, memory 404 can include any combination of any number of random access memories (RAM), read only memories (ROM), optical disk drives, magnetic disk drives, hard drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, and the like. Memory 404 can also be a group of memories grouped as a single logical component (not shown in FIG. 4). Figure 4
[0083] Bus 410 can be a communication device that transfers data between components within device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or the like.
[0084] For ease of explanation and without causing ambiguity, in this disclosure, processor 402 and other data processing circuitry are collectively referred to as “data processing circuitry.” The data processing circuitry can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Moreover, the data processing circuitry can be a single standalone module, or can be combined, in whole or in part, into any other component of device 400.
[0085] Device 400 can also include network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, network interface 406 can include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, and the like.
[0086] In some embodiments, optionally, device 400 can further include peripheral interface 408 to provide connection to one or more peripheral devices. As Figure 4 shown, the peripheral devices can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), and the like.
[0087] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) can be implemented as any combination of any of the software or hardware modules in the apparatus 400. For example, some or all stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of the apparatus 400, such as a program instance that can be loaded into memory 404. For another example, some or all stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of the apparatus 400, such as a special-purpose data processing circuit (e.g., FPGA, ASIC, NPU, etc.).
[0088] In the quantization and inverse quantization function blocks (e.g., quantization 214 and inverse quantization 218 of Figure 2A or Figure 2B , a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. An initial QP value for encoding an image or slice can be signaled at a higher level, e.g., using an init_qp_minus26 syntax element in a picture parameter set (PPS) and using a slice_qp_delta syntax element in a slice header. In addition, a delta QP value sent at the granularity of a quantization group can be used to adapt the QP value locally for each CU. Figure 3A Figure 3B According to some embodiments, an image can be divided into a plurality of coding tree units (CTU). The CTU is then further divided into one or more coding units (CU) using a quadtree (SPLIT QT) with a nested multi-type tree partitioning structure with binary and ternary splitting.
[0089] According to some embodiments of the present disclosure, a picture can be divided into a plurality of coding tree units (CTU). The CTU is then further divided into one or more coding units (CU) using a quadtree (SPLIT QT) with a nested multi-type tree partitioning structure with binary and ternary splitting. Figure 5 is a schematic diagram showing examples of multi-type tree partitioning modes according to some embodiments of the present disclosure. As shown in Figure 5 the partitioning types in the multi-type tree structure can include a quadtree split (SPLIT QT) 501, a vertical binary tree split (SPLIT BT VER) 502, a horizontal binary tree split (SPLIT BT HOR) 503, a vertical ternary tree split (SPLIT TT VER) 504, and a horizontal ternary tree split (SPLIT TT HOR) 505. The leaf nodes of the multi-type tree are called coding units (CU), which can have a square or rectangular shape.
[0090] Figure 6 is a schematic diagram of an exemplary signaling mechanism for partition split information in a quadtree with nested multi-type tree coding tree structure according to some embodiments of the present disclosure. In Figure 6 In the multi-type tree (MTT) structure, a CTU is treated as the root of a quad-tree, and is first partitioned by a quad-tree structure. Each quad-tree leaf node (when large enough to allow it) is further partitioned by a multi-type tree structure. In the multi-type tree structure, a first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned. When the node is further partitioned, a second flag (e.g., mtt_split_cu_vertical_flag) is signaled to indicate the partition direction, and then a third flag (e.g., mtt_split_cu_binary_flag) is signaled to indicate whether the partition is a binary tree partition or a ternary tree partition. From the values of the second and third flags, mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree partition mode (MttSplitMode) of a CU can be derived. Figure 7 An exemplary Table 1 is shown, which illustrates exemplary MTTSplitMode derivation based on multi-type tree syntax elements, according to some embodiments of the disclosure.
[0091] A virtual pipeline data unit (VPDU) is defined as a non-overlapping unit in an image. In a hardware decoder, consecutive VPDU is processed by multiple pipeline stages simultaneously. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is important to maintain the VPDU size. In most hardware decoders, the VPDU size can be set to 64x64 luma samples. However, in some embodiments, ternary tree (TT) and binary tree (BT) partitioning can cause an increase in VPDU size. In line with the present disclosure, to maintain the VPDU size to 64x64 luma samples, certain normative partitioning restrictions can be applied. Figure 8 An example of disallowed TT and BT partitioning is shown, according to some embodiments of the disclosure. As shown, Figure 8 For blocks with width or height, or both width and height equal to 128, TT splitting is disallowed. For CUs with 128xN, N≤64 (e.g., width equal to 128 and height less than 128), horizontal BT is disallowed. For CUs with Nx128, N≤64 (e.g., height equal to 128 and width less than 128), vertical BT is disallowed.
[0092] According to the requirement in HEVC, when a portion of a tree node block exceeds the bottom or right side boundary of the picture, the tree node block is forced to be split until all samples of each coded CU are within the picture boundary. The following splitting rules are applied in VVC Draft 7:
[0093] - If a portion of a tree node block exceeds the bottom and right side boundary of the picture,
[0094] - If the block is a QT node and the size of the block is greater than the minimum QT size, then the block is forced to be split using the QT split mode.
[0095] - Otherwise, the block is forced to be split using the SPLIT_BT_HOR mode.
[0096] - Otherwise, if a portion of the tree node block exceeds the image bottom boundary,
[0097] - If the block is a QT node and the size of the block is greater than the minimum QT size and the size of the block is greater than the maximum BT size, then the block is forced to be split using the QT split mode.
[0098] - Otherwise, if the block is a QT node and the size of the block is greater than the minimum QT size and the size of the block is less than or equal to the maximum BT size, then the block is forced to be split using the QT split mode or the SPLIT_BT_HOR mode.
[0099] - Otherwise (the block is a BT node or the size of the block is less than or equal to the minimum QT size), then the block is forced to be split using the SPLIT_BT_HOR mode.
[0100] - Otherwise, if a portion of the tree node block exceeds the image right boundary,
[0101] - If the block is a QT node and the size of the block is greater than the minimum QT size and the size of the block is greater than the maximum BT size, then the block is forced to be split using the QT split mode,
[0102] - Otherwise, if the block is a QT node and the size of the block is greater than the minimum QT size and the size of the block is less than or equal to the maximum BT size, then the block is forced to be split using the QT split mode or the SPLIT_BT_VER mode.
[0103] - Otherwise (e.g., the block is a BT node or the size of the block is less than or equal to the minimum QT size), then the block is forced to be split using the SPLIT_BT_VER mode.
[0104] Figure 9 Example block partitioning on image boundaries is shown in accordance with some embodiments of the disclosure. As shown in Figure 9 For CTU 911, SPLIT_QT or SPLIT_BT_VER can be performed. For CTU 913, SPLIT_QT can be performed if SPLIT_QT is allowed, and SPLIT_BT_HOR can be performed if SPLIT_QT is not allowed. For CTU 915, SPLIT_QT or SPLIT_BT_HOR can be performed.
[0105] In VVC Draft 7, there are two parts related to block partitioning. The first one is Section 6.4, which defines whether quad-tree, binary-tree or ternary-tree can be used to split a block. The output of Section 6.4 is variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor and allowSplitTtVer. These variables are used in Section 7.3.9.4 as shown in Table 4 in Figure 12 to determine whether the corresponding CU-level partitioning flag (marked by boxes 1201-1204 in Table 4 in Figure 12 ) is signaled.
[0106] Section 6.4.1 of VVC Draft 7 describes:
[0107] 6.4 Availability process
[0108] 6.4.1 Allowed quad-tree splitting process
[0109] The inputs of this process include:
[0110] - the coding block size cbSize in luma samples,
[0111] - the multi-type tree depth mttDepth,
[0112] - the variable treeType specifies whether single tree (SINGLE_TREE) or dual tree is used to split the coding tree node, and when dual tree is used, whether the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) is currently being processed,
[0113] - a variable modeType specifies whether intra (MODE_INTRA), IBC (MODE_IBC) and inter coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter coding modes (MODE_TYPE_INTER) can be used for coding units inside the coding tree node.
[0114] The output of this process is the variable allowSplitQt. The variable allowSplitQt is derived as follows:
[0115] - If one or more of the following conditions are true, set allowSplitQt to false (FALSE):
[0116] - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY
[0117] - treeType is equal to DUAL_TREE_CHROMA, cbSize / SubWidthC is less than or equal to MinQtSizeC
[0118] - mttDepth is not equal to 0
[0119] - treeType is equal to DUAL_TREE_CHROMA, and (cbSize / SubWidthC) is less than or equal to 4
[0120] - treeType is equal to DUAL_TREE_CHROMA, modeType is equal to MODE_TYPE_INTER
[0121] - Otherwise, allowSplitQt is set equal to TRUE.
[0122] Section 6.4.2 of VVC Draft 7 describes:
[0123] 6.4.2 Allowed binary tree split process
[0124] The inputs of this process are:
[0125] - a binary tree split mode btSplit,
[0126] - a coding block width in luma samples cbWidth,
[0127] - a coding block height in luma samples cbHeight,
[0128] - the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture,
[0129] - a multi-type tree depth mttDepth,
[0130] - a maximum multi-type tree depth with offset maxMttDepth,
[0131] - a maximum binary tree size maxBtSize,
[0132] - a minimum quad tree size minQtSize,
[0133] - a partition index partIdx
[0134] - a variable treeType specifying whether a single tree (SINGLE_TREE) or a dual tree is used to split the coding tree nodes, and when a dual tree is used, whether a luma component (DUAL_TREE_LUMA) or a chroma component (DUAL_TREE_CHROMA) is currently being processed,
[0135] - a variable modeType specifying whether intra (MODE_INTRA), IBC (MODE_IBC) and inter coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter coding modes (MODE_TYPE_INTER) can be used for coding units inside the coding tree node.
[0136] The output of this process is the variable allowBtSplit. Figure 10 An exemplary Table 2 is shown, which shows exemplary specifications for the btSplit based variables parallelTtSplit and cbSize, according to some embodiments of the disclosure.
[0137] The variable allowBtSplit is derived as follows:
[0138] - If one or more of the following conditions are true, allowBtSplit is set equal to FALSE:
[0139] - cbSize is less than or equal to MinBtSizeY
[0140] - cbWidth is greater than maxBtSize
[0141] - cbHeight is greater than maxBtSize
[0142] - mttDepth is greater than or equal to maxMttDepth
[0143] - treeType is equal to DUAL_TREE_CHROMA and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 16
[0144] - treeType is equal to DUAL_TREE_CHROMA and (cbWidth / SubWidthC) is equal to 4 and btSplit is equal to SPLIT_BT_VER
[0145] - treeType is equal to DUAL_TREE_CHROMA and modeType is equal to MODE_TYPE_INTRA
[0146] - cbWidth * cbHeight is equal to 32, modeType is equal to MODE TYPE INTER
[0147] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0148] - btSplit is equal to SPLIT BT YER
[0149] - yO + cbHeight is greater than pic height in luma samples
[0150] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0151] - btSplit is equal to SPLIT BT VER
[0152] - cbHeight is greater than 64
[0153] - xO + cbWidth is greater than pic width in luma samples
[0154] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0155] - btSplit is equal to SPLIT BT HOR
[0156] - cbWidth is greater than 64
[0157] - yO + cbHeight is greater than pic height in luma samples
[0158] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0159] - xO + cbWidth is greater than pic width in luma samples
[0160] - yO + cbHeight is greater than pic height in luma samples
[0161] - cbWidth is greater than minQtSize
[0162] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0163] - btSplit is equal to SPLIT_BT_HOR
[0164] - x0 + cbWidth is greater than pic width in luma samples
[0165] - y0 + cbHeight is less than or equal to pic height in luma samples
[0166] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE:
[0167] - mttDepth is greater than 0
[0168] - partIdx is equal to 1
[0169] - MttsplitMode[ x0 ][ y0 ][ mttDepth - 1 ] is equal to parallelTtSplit
[0170] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0171] - btSplit is equal to SPLIT_BT_VER
[0172] - cbWidth is less than or equal to 64
[0173] - cbHeight is greater than 64
[0174] - Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE
[0175] - btSplit is equal to SPLIT_BT_HOR
[0176] - cbWidth is greater than 64
[0177] - cbHeight is less than or equal to 64
[0178] - Otherwise, allowBtSplit is set equal to TRUE.
[0179] VVC Draft 7 Section 6.4.3 describes:
[0180] 6.4.3 Allowed triple tree splitting processes
[0181] The inputs to this process are:
[0182] - triple tree splitting ttSplit,
[0183] - the coding block width in luma samples cbWidth,
[0184] - the coding block height in luma samples cbHeight,
[0185] - the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture,
[0186] - the multi-type tree depth mttDepth
[0187] - the maximum multi-type tree depth with offset maxMttDepth,
[0188] - the maximum ternary tree size maxTtSize,
[0189] - the variable treeType specifying whether a single tree (SINGLE_TREE) or a dual tree is used to split the coding tree node, and when a dual tree is used, whether the current processing is for the luma component (DUAL_TREE_LUMA) or the chroma component (DUAL_TREE_CHROMA),
[0190] - a variable modeType specifying whether intra (MODE_INTRA), IBC (MODE_IBC) and inter coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter coding modes (MODE_TYPE_INTER) can be used for coding units inside the coding tree node.
[0191] The output of this process is the variable allowTtSplit. Figure 11 An exemplary Table 3 is shown, which shows exemplary specifications for the variable cbSize based on ttSplit, according to some embodiments of the disclosure.
[0192] The variable allowTtSplit is derived as follows:
[0193] - If one or more of the following conditions are true, allowTtSplit is set equal to FALSE:
[0194] - cbSize is less than or equal to 2 * MinTtSizeY
[0195] - cbWidth is greater than Min(64, maxTtSize)
[0196] - cbHeight is greater than Min(64, maxTtSize)
[0197] - mttDepth is greater than or equal to maxMttDeptb
[0198] - x0 + cbWidth is greater than pic_width_in_luma_samples
[0199] - y0 + cbHeight is greater than pic_height_in_luma_samples
[0200] - treeType is equal to DUAL_TREE_CHROMA, and
[0201] - (cbWidth / SubWidthC) * (cbHeight / SubHeightC) is less than or equal to 32
[0202] - treeType is equal to DUAL_TREE_CHROMA, and (cbWidth / SubWidthC) is equal to 8, ttSplit is equal to SPLIT_TT_VER
[0203] - treeType is equal to DUAL_TREE_CHROMA, and modeType is equal to MODE_TYPE_INTRA
[0204] - cbWidth * cbHeight is equal to 64, modeType is equal to MODE_TYPE_INTER
[0205] - Otherwise, allowTtSplit is set equal to TRUE.
[0206] Figure 12 An example Table 4 is shown, which shows an example portion 7.3.9.4 Coding Tree Syntax (with added emphasis in italics and shading) of VVC Draft 7, according to some embodiments of the present disclosure.
[0207] The derivation of the variables allowSplitQt, allowSplitBtVer, allowSplitBtHor, allowSplitTtVer, and allowSplitTtHor is as follows:
[0208] - The allowed quadtree splitting process specified in clause 6.4.1 is invoked with the coding block size cbSize set equal to cbWidth, the current multi-type tree depth mttDepth, treeTypeCurr, and modeTypeCurr as inputs, and the output is assigned to allowSplitQt.
[0209] - The variables minQtSize, maxBtSize, maxTtSize and maxMttDepth are derived as follows:
[0210] - If treeType is equal to DUAL_TREE_CHROMA, minQtSize, maxBtSize, maxTtSize and maxMttDepth are set equal to MinQtSizeC, MaxBtSizeC, MaxTtSizeC and MaxMttDepthC + depthOffset, respectively.
[0211] - Otherwise, minQtSize, maxBtSize, maxTtSize and maxMttDepth are set equal to MinQtSizeY, MaxBtSizeY, MaxTtSizeY and MaxMttDepthY + depthOffset, respectively.
[0212] - The allowed binary tree splitting process specified in clause 6.4.2 is invoked with binary tree splitting mode SPLIT_BT_VER, coded block width cbWidth, coded block height cbHeight, position (x0, y0), current multi-type tree depth mttDepth, maximum multi-type tree depth with offset maxMttDepth, maximum binary tree size maxBtSize, minimum quad tree size minQtSize, current partition index partldx, treeTypeCurr and modeTypeCurr as inputs, and the output is assigned to allowSplitBtVer.
[0213] - The allowed binary tree splitting process specified in clause 6.4.2 is invoked with binary tree splitting mode SPLIT_BT_HOR, coded block height cbHeight, coded block width cbWidth, position (x0, y0), current multi-type tree depth mttDepth, maximum multi-type tree depth with offset maxMttDepth, maximum binary tree size maxBtSize, minimum quad tree size minQtSize, current partition index partldx, treeTypeCurr and modeTypeCurr as inputs, and the output is assigned to allowSplitBtHor.
[0214] - The allowed ternary tree split process specified in clause 6.4.3 is invoked with the ternary tree split mode SPLIT_TT_VER, the coded block width cbWidth, the coded block height cbHeight, the position (x0, y0), the current multi-type tree depth mttDepth, the maximum multi-type tree depth with offset maxMttDepth, the maximum ternary tree size maxTtSize, treeTypeCurr, and modeTypeCurr as inputs, and the output is assigned to allowSplitTtVer.
[0215] - The allowed ternary tree split process specified in clause 6.4.3 is invoked with the ternary tree split mode SPLIT_TT_HOR, the coded block height cbHeight, the coded block width cbWidth, the position (x0, y0), the current multi-type tree depth mttDepth, the maximum multi-type tree depth with offset maxMttDepth, the maximum ternary tree size maxTtSize, treeTypeCurr, and modeTypeCurr as inputs, and the output is assigned to allowSplitTtHor.
[0216] The syntax element split_cu_flag equal to 0 specifies that the coding unit is not split. The syntax element split_cu_flag equal to 1 specifies that the coding unit is split into four coding units using quad-tree splitting as indicated by the syntax element split_qt_flag, or into two coding units using binary tree splitting, or into three coding units using ternary tree splitting as indicated by the syntax element mtt_split_cu_binary_flag. The binary or ternary tree splitting can be vertical or horizontal as indicated by the syntax element mtt_split_cu_vertical_flag.
[0217] When the syntax element split_cu_flag is not present, the value of split_cu_flag is inferred as follows:
[0218] - If one or more of the following conditions are true, the value of split_cu_flag is inferred to be equal to 1:
[0219] - x0 + cbWidth is greater than pic_width_in_luma_samples.
[0220] - y0 + cbHeight is greater than pic_height_in_luma_samples.
[0221] - Otherwise, the value of split_cu_flag is inferred to be equal to 0.
[0222] The syntax element split_qt_flag specifies whether the coding unit is split into coding units with half horizontal and vertical size.
[0223] When the syntax element split_qt_flag is not present, the following applies:
[0224] - If allowSplitQt is equal to TRUE, the value of split_qt_flag is inferred to be equal to 1.
[0225] - Otherwise, the value of split_qt_flag is inferred to be equal to 0.
[0226] The syntax element mtt_split_cu_vertical_flag equal to 0 specifies that the coding unit is split horizontally. The syntax element mtt_split_cu_vertical_flag equal to 1 specifies that the coding unit is split vertically.
[0227] When the syntax element mtt_split_cu_vertical_flag is not present, the following is inferred:
[0228] - If allowSplitBtHor is equal to TRUE or allowSplitTtHor is equal to TRUE, the value of mtt_split_cu_vertical_flag is inferred to be equal to 0.
[0229] - Otherwise, the value of mtt_split_cu_vertical_flag is inferred to be equal to 1.
[0230] The syntax element mtt_split_cu_binary_flag equal to 0 specifies that the coding unit is split into three coding units using a ternary tree. The syntax element mtt_split_cu_binary_flag equal to 1 specifies that the coding unit is split into two coding units using a binary tree split.
[0231] When the syntax element mtt_split_cu_binary_flag is not present, the following is inferred:
[0232] - If allowSplitBtVer is equal to FALSE and allowSplitBtHor is equal to FALSE, the value of mtt_split_cu_binary_flag is inferred to be equal to 0.
[0233] - Otherwise, if allowSplitTtVer is equal to FALSE and allowSplitTtHor is equal to FALSE, the value of mtt_split_cu_binary_flag is inferred to be equal to 1.
[0234] - Otherwise, if allowSplitBtHor is equal to TRUE and allowSplitTtVer is equal to TRUE, the value of mtt_split_cu_binary_flag is equal to mtt_split_cu_vertical_flag.
[0235] - Otherwise (allowSplitBtVer is equal to TRUE, allowSplitTtHor is equal to TRUE), the value of mtt_split_cu_binary_flag is inferred to be equal to mtt_split_cu_vertical_flag.
[0236] Figure 14 An exemplary Table 5 is shown, which illustrates exemplary specifications of MttSplitMode according to some embodiments of the disclosure. The variable MttSplitMode[x][y][mttDepth] is derived from the value of the syntax element mtt_split_cu_vertical_flag and the value of the syntax element mtt_split_cu_binary_flag when x = x0..x0+cbWidth-1 and y = y0..y0+cbHeight-1 are defined in Table 4.
[0237] MttSplitMode[x][y][mttDepth] represents the horizontal binary tree, vertical binary tree, horizontal ternary tree and vertical ternary tree split of the multi-type tree coding unit. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. Figure 13 An example of multi-type tree split modes indicated by MttSplitMode according to some embodiments of the disclosure is shown. As shown in Figure 13 The multi-type tree split modes can include a vertical binary tree split (SPLIT_BT_VER) 1301, a horizontal binary tree split (SPLIT_BT_HOR) 1302, a vertical ternary tree split (SPLIT_TT_VER) 1303, and a horizontal ternary tree split (SPLIT_TT_HOR) 1304, as shown.
[0238] It should be noted that the CTU size, the minimum block size and the block size limit for quad-tree, binary-tree and ternary-tree splits are signaled in the sequence parameter set or the picture header.
[0239] Figure 15 An exemplary Table 6 is shown, which illustrates an exemplary portion 7.3.2.3 Sequence Parameter Set RBSP syntax of VVC Draft 7, according to some embodiments of the present disclosure. Figure 16 An exemplary Table 7 is shown, which illustrates an exemplary portion 7.3.2.6 Picture Header RBSP syntax of VVC Draft 7, according to some embodiments of the present disclosure.
[0240] According to some embodiments, when a portion of a tree node block exceeds the picture boundary bottom or right side, the tree node block is forced to be split until all samples of each coding block are within the picture boundary. However, in certain cases, all tree split modes are not allowed for a block that is located at a picture boundary and contains samples that exceed the picture boundary. Figure 17 An exemplary block where QT, TT or BT split is not allowed at a picture boundary is shown, according to some embodiments of the present disclosure. For example, Figure 17 CTU 1701, CTU 1703 and CTU 1705 at a picture boundary are shown not to allow QT, TT or BT split.
[0241] As a first exemplary case where all tree split modes are not allowed, when both CTU size and MinQtSize are set to 128, all variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor and allowSplitTtVer are set to false.
[0242] Since the following condition variable allowSplitQt is set to false (emphasis in italics):
[0243] 6.4.1 Allowed quadtree splitting process ...
[0244] - If one or more of the following conditions are true, allowSplitQt is set equal to FALSE:
[0245] - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA and chSize is less than or equal to MinQtSizeY
[0246] - treeType is equal to DUAL_TREE_CHROMA and cbSize / SubWidthC is less than or equal to MinQtSizeC ...
[0247] The variable allowSplitBtHor (emphasized in italics) is set to false due to the following conditions:
[0248] 6.4.2 Allowed binary split process ...
[0249] - Otherwise, if all of the following conditions are true, allowBtSplit is set to FALSE
[0250] - btSplit is equal to SPLIT_BT_HOR
[0251] - cbWidth is greater than 64
[0252] - y0 + cbHeight is greater than pic_height_luma_samples
[0253] - Otherwise, if all of the following conditions are true, allowBtSplit is set to FALSE
[0254] - x0 + cbWidth is greater than pic_width_in_luma_samples
[0255] - y0 + cbHeight is greater than pic_height_in_luma_samples
[0256] - cbWidth is greater than minQtSize
[0257] - Otherwise, if all of the following conditions are true, allowBtSplit is set to FALSE
[0258] - btSplit is equal to SPLIT_BT_HOR
[0259] - x0 + cbWidth is greater than pic_width_in_luma_samples
[0260] - y0 + cbHeight is less than or equal to pic_height_in_luma_samples ...
[0261] The variable allowSplitBtVer is set to false (emphasized in italics) due to the following conditions:
[0262] 6.4.2 Allowed binary split process ...
[0263] - Otherwise, if all of the following are true, set allowBiSplit to FALSE
[0264] - btSplit is equal to SPLIT_BT_VER
[0265] - y0 + cbHeight is greater than pic_height_in_luma_samples
[0266] - Otherwise, if all of the following are true, set allowBtSplit to FALSE
[0267] - btSplit is equal to SPLIT_BT_VER
[0268] - chHeight is greater than 64
[0269] - x0 + cbWidth is greater than pic_width_in_luma_samples
[0270] The variables allowSplitTtHor and allowSplilTtVer (emphasized in italics) are set to false due to the following:
[0271] 6.4.3 Allowed ternary tree split process ...
[0272] - If one or more of the following conditions are true, set allowTtSplit to equal FALSE:
[0273] - cbSize is less than or equal to 2 * MinTtSizeY
[0274] - cbWidth is greater than min(64, maxTtSize)
[0275] - cbHeight is greater than min(64, maxTtSize)
[0276] - mttDepth is greater than or equal to maxMttDepth
[0277] - x0 + cbWidth is greater than pic_width_in_luma_samples
[0278] - y0 + cbHeight is greater than pic_height_in_luma_samples ...
[0279] When all these variables are set to false, the CU-level split flag can not be signaled. The syntax element split_cu_flag is inferred to be 1, the syntax element split_qt_flag is inferred to be 0, the syntax element mtt_split_cu_vertical_flag is inferred to be 1, and the syntax element mtt_split_cu_binary_flag is inferred to be 0. In this case, a SPLIT_TT_VER split block can be used, which can violate the constraint of VPDU.
[0280] As a second example of a split mode that is not allowed for all trees, the split of all trees is not allowed for a block that is located at the picture boundary and contains samples that are outside the picture boundary. When the minimum QT size (syntax element log2_min_luma_coding_block_size_minus2 in the previous table) is greater than the minimum CU size and the maximum BT / TT depth (syntax elements sps_max_mtt_hierarchy_depth_inter_slice, sps_max_mtt_hierarchy_depth_intra_slice_luma, pic_max_mtt_hierarchy_depth_inter_slice, pic_max_mtt_hierarchy_depth_intra_slice_luma, and pic_max_mtt_hierarchy_depth_intra_slice_chroma) are equal to 0, all variables allowSplitQt, allowSplitBtHor, allowSplitBtHor, allowSplitTtHor, and allowSplitTtVer are set to false.
[0281] Figure 18 An exemplary block where QT, BT, or BT split is not allowed at the picture boundary is shown according to some embodiments of the disclosure. A CTU (e.g., CTU 1801, CTU 1803, or CTU 1805) is first split into four 64x64 blocks using quad-tree. Then, each 64x64 block cannot be further split. However, the portion of the block that is outside the right and / or bottom picture boundary is marked in gray, which is not allowed in the VVC design. For the block marked in gray, the variable allowSpiltQt is set to false due to the following condition (emphasis in italics):
[0282] 6.4.1 Allowed quad-tree split process
[0283] …
[0284] - If one or more of the following conditions are true, allowSplitQt is set equal to FALSE:
[0285] - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA and cbSize is less than or equal to MinQtSizeY
[0286] - treeType is equal to DUAL_TREE_CHROMS and cbSize / SubWidthC is less than or equal to MinQtSizeC
[0287] …
[0288] For the blocks marked in grey, the variables allowSplitBtHor and allowSplitBtVer are set to false due to the following (italics are used to emphasize):
[0289] 6.4.2 Allowed binary splitting process
[0290] …
[0291] The variable allowBtSplit is derived as follows:
[0292] - If one or more of the following conditions are true, allowBtSplit is set to FALSE:
[0293] - cbSize is less than or equal to MinBtSizeY
[0294] - cbWidth is greater than maxBtSize
[0295] - cbHeight is greater than maxBtSize
[0296] - MttDepth is greater than or equal to maxMttDepth
[0297] …
[0298] For the blocks marked in grey, the variables allowSplitTtHor and allowSplitTtVer are set to false due to the following (italics are used to emphasize):
[0299] 6.4.3 Allowed ternary splitting process
[0300] …
[0301] The variable allowTtSplit is derived as follows:
[0302] - If one or more of the following conditions are true, allowTtSplit is set to FALSE:
[0303] - cbSize is less than or equal to 2 * MinTtSizeY
[0304] - cbWidth is greater than Min(64, maxTtSize)
[0305] - cbHeight is greater than Min(64, maxTtSize)
[0306] - mttDepth is greater than or equal to maxMttDepth ...
[0307] As described in the first exemplary case where all tree partition modes are not allowed, in the current VVC Draft 7, a CU can contain samples outside the picture boundary, but the CU cannot be further partitioned under certain conditions. In some embodiments of the present disclosure, the QT partitioning condition in VVC can be changed. In one aspect, for a block containing samples outside the picture boundary and whose width or height is equal to N (e.g., N = 128), QT partitioning is used when the minimum QT size is less than N (e.g., 128). In addition, in some embodiments, QT partitioning can also be used when the minimum QT size is equal to N (e.g., 128). In another aspect, using QT partitioning can be more straightforward than using BT or TT partitioning. Figure 19 is a schematic diagram showing an example of using BT and TT partitioning according to some embodiments of the present disclosure. Partitioning a block can require multiple steps. In addition, the partitioning can be different for blocks located in different positions, which can be complex. For example, for block 1903, SPLIT_BT_HQR, SPLIT_BT_HOR, SPLIT_TT_VER, and SPLIT_TT_VER are executed in sequence.
[0308] In some embodiments, when block partitioning is not allowed, quad-tree partitioning can be used, and the syntax element split_qt_flag can be inferred to be 1. The syntax element split_qt_flag specifies whether to partition a coding unit into coding units with half the horizontal and vertical size.
[0309] When the syntax element split_qt_fiag is not present, the following applies (emphasis in italics):
[0310] - If all of the following conditions are true, split_qt_flag is inferred to be equal to 1:
[0311] - split_cu_flag is equal to I
[0312] - allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor and allowSplitTtVer are equal to FALSE.
[0313] - Otherwise, if allowSplitQt is equal to TRUE, infer the value of split_qt_flag to be 1.
[0314] - Otherwise, infer the value of split_qt_flag to be 0.
[0315] In some embodiments, the minimum QT size constraint cannot be applied to blocks located at the image boundary. When a part of a block exceeds the image bottom or right boundary, the block can be split using a quad-tree. The allowed quad-tree splitting process is described as follows:
[0316] 6.4.1 Allowed quad-split process
[0317] The input of this process is:
[0318] - cbSize, the size of the coding block in luma samples,
[0319] - mttDepth, the multi-type tree depth,
[0320] - the variable treeType used to specify whether a single tree (SINGLE_TREE) or dual tree is used to split the coding tree node, and when dual tree is used, whether the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) is being processed.
[0321] - the variable modeType specifies whether intra (MODE_INTRA), IBC (MODE_IBC) and inter coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter coding modes (MODE_TYPE_INTER) can be used for coding units inside the coding tree node.
[0322] The output of this process is the variable allowSplitQt.
[0323] The derivation of the variable allowSplitQt is as follows (italics are used to emphasize):
[0324] - If all the following conditions are true, allowSplitQt is set equal to true:
[0325] - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA
[0326] - cbSize is equal to 128
[0327] - MinQtSizeY is equal to 128
[0328] - x0 + cbWidth is greater than pic width in luma samples or y0 + cbHeight is greater than pic height in luma samples
[0329] - Otherwise, if all of the following conditions are true, allowSplitQt is set equal to true:
[0330] - treeType is equal to DUAL_TREE_CHROMA
[0331] - cbSize / SubWidthC is equal to 128
[0332] - MinQtSizeC is equal to 128
[0333] - x0 + cbWidth is greater than pic width in luma samples or y0 + cbHeight is greater than pic height in luma samples
[0334] - Otherwise, if one or more of the following conditions are true, allowSplitQt is set equal to FALSE:
[0335] - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA and cbSize is less than or equal to MinQtSizeY
[0336] - treeType is equal to DUAL_TREE_CHROMA and cbSize / SubWidthC is less than or equal to MinQtSizeC
[0337] - mttDepth is not equal to 0
[0338] - treeType is equal to DUAL_TREE_CHROMA and (ebSize / SubWidthC) is less than or equal to 4
[0339] - treeType is equal to DUAL_TREE_CHROMA and modeType is equal to MODE_TYPE_INTRA
[0340] - Otherwise, allowSplitQt is set equal to TRUE.
[0341] In some embodiments, bitstream conformance can be added to the syntax of the minimum QT size. It can be required that the minimum QT size is less than or equal to 64.
[0342] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size of luma samples of a luma leaf block resulting from the quadtree partitioning of a CTU and the base-2 logarithm of the minimum coding block size in luma samples of a luma CU in a slice with slice_type equal to 2 (I) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH referring to the SPS. The value of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. The base-2 logarithm of the minimum size of luma samples of a luma leaf block resulting from the quadtree partitioning of a CTU is derived as follows (italics are used to emphasize):
[0343] MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY
[0344] VSize = Min(64, CtbSizeY)
[0345] In some embodiments, bitstream conformance d requires that the value of (1 « MinQtLog2SizeIntraY) is less than or equal to VSize.
[0346] The syntax element sps_log2_diff_min_qt_min_cb_inter_slice_luma specifies a default difference, in base-2 logarithm, of the minimum size in luma samples of a luma leaf block resulting from the quad-tree partitioning of a CTU from the minimum luma coding block size in luma samples of a luma CU in a slice with slice_type equal to 0 (B) or 1 (P) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH referring to the SPS. The syntax element sps_log2_diff_min_qt_min_cb_inter_slice has a value in the range of 0 to CtbLog2SizeY value MinCbLog2SizeY, inclusive. The minimum size in luma samples of a luma leaf block resulting from the quad-tree partitioning of a CTU from the minimum luma coding block size in luma samples of a luma CU in a slice with slice_type equal to 0 (B) or 1 (P) referring to the SPS is derived as follows (italics are used to emphasize):
[0347] MinQtLog2SizeInterY = sps_log2_diff_min_qt_min_cb_inter_slice
[0348] MinCbLog2SizeY
[0349] VSize = Min(64, CihSizeY)
[0350] In some embodiments, a conformance requirement of the bitstream is that the value of (1 « MinQtLog2SizeInterY) is less than or equal to VSize.
[0351] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies a default difference, in base 2 logarithm, between the minimum size of luma samples in chroma leaf blocks resulting from quad-tree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA and the minimum coding block size in luma samples of chroma CUs with treeType equal to DUAL_TREE_CHROMA in slices with slice_type equal to 2 (I) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_chroma present in the PH referring to the SPS. The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma has a range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to 0, the minimum size of luma samples in chroma leaf blocks resulting from quad-tree partitioning of a CTU with treeType equal to DUAL_TREE_CHROMA is derived as follows (italics emphasis):
[0352] MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_chroma + MinCbLog2SizeY
[0353] VSize = Min(64, CtbSizeY)
[0354] In some embodiments, bitstream conformance requires that the value of (1 « MinQtLog2SizeIntraC) is less than or equal to VSize.
[0355] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma specifies a default difference, in base 2 logarithm, between the minimum size in luma samples of a luma leaf block resulting from the quad-tree partitioning of a CTU and the base 2 logarithm of the minimum coding block size in luma samples of a luma CU in a slice with slice type equal to 2 (I) associated with the PH. The value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma. In some embodiments, the bitstream conformance requirement that (1 « (pic_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY)) is less than or equal to Min(64, CtbSizeY).
[0356] The syntax element pic_log2_diff_min_qt_min_cb_inter_slice specifies a difference, in base 2 logarithm, between the minimum size in luma samples of a luma leaf block resulting from the quad-tree partitioning of a CTU and the base 2 logarithm of the minimum luma coding block size of a luma CU in an intra slice with slice type equal to 0 (B) or 1 (P) associated with the PH. The value of the syntax element pic_log2_diff_min_qt_min_cb_inter_slice is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_inter_slice. In some embodiments, the bitstream conformance requirement that (1 « (pic_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY)) is less than or equal to Min(64, CtbSizeY).
[0357] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the difference, in base 2 logarithm, of the minimum size of luma samples of a chroma leaf block resulting from quad-tree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA and the base 2 logarithm of the minimum coding block size in luma samples of a chroma CU with treeType equal to DUAL_TREE_CHROMA in the slice with slice_type equal to 2 (I) associated with PH. The value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma. In some embodiments, the bitstream conformance requirement (1 « (pic_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY)) has a value less than or equal to Min(64, CtbSizeY).
[0358] In some embodiments, when the block partitioning case is not allowed in the second example above (where all tree partitioning modes are not allowed), quad-tree partitioning can be used and the syntax element split_qt_flag is inferred to be equal to 1.
[0359] The syntax element split_qt_flag specifies whether a coding unit is partitioned into a coding unit with half the horizontal and vertical size. When the syntax element split_qt_flag is not present, the following applies (in italics emphasis):
[0360] - If all of the following conditions are true, split_qt_flag is inferred to be equal to 1:
[0361] - split_cu_flag is equal to I
[0362] - allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer are equal to FALSE.
[0363] - Otherwise, if allowSplitQt is equal to TRUE, infer the value of split_qt_flag to be equal to 1.
[0364] - Otherwise, infer the value of split_qt_flag to be equal to 0.
[0365] In some embodiments, bitstream conformance can be added to the syntax of the minimum QT size and the maximum BT / TT depth in the above second exemplary case where all tree partition modes are not allowed.
[0366] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference, in base 2 logarithm, of the minimum size of luma samples of a luma leaf block resulting from a quad-tree partition of a CTU from the minimum coding block size of luma samples of luma CUs in slices with slice_type equal to 2 (I) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, the default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH referring to the SPS. The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma has a value range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. The minimum size of luma samples of a luma leaf block resulting from a quad-tree partition of a CTU, in base 2 logarithm, is as follows:
[0367] MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY
[0368] The syntax element sps_max_mtt_hierarchy_depth_intra_slice_luma specifies the default maximum hierarchy depth of coding units resulting from multi-type tree partitioning of quad-tree leaves in slices with slice_type equal to 2 (I) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default maximum hierarchy depth can be overridden by the syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma present in the PH referring to the SPS. The value of the syntax element sps_max_mtt_hierarchy_depth_intra_slice_luma is in the range of 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. In some embodiments, the bitstream conformance requires that the value of (MmQtLog2SizeIntraY - sps_max_mtt_hierachy_depth_intra_slice_luma / 2) is less than or equal to MinCbLog2SizeY.
[0369] The syntax element sps_log2_diff_min_qt_min_cb_inter_slice specifies the default difference, in base 2 logarithm, of the minimum size of luma samples of luma leaf blocks resulting from quad-tree partitioning of CTU, from the minimum luma coding block size of luma CUs in slices with slice_type equal to 0 (B) or 1 (P) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH referring to the SPS. The value of the syntax element sps_log2_diff_min_qt_min_cb_inter_slice is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. The minimum size of luma samples of luma leaf blocks resulting from quad-tree partitioning of CTU, in base 2 logarithm, is derived as follows:
[0370] MinQtLog2SizeInterY = sps_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY
[0371] The syntax element sps_max_mtt_hierarchy_depth_inter_slice specifies the default maximum level depth of coding units resulting from quad-tree leaf multi-type tree partitioning of slices with slice_type equal to 0 (B) or 1 (P) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, the default maximum level depth can be overridden by the syntax element pic_max_mtt_hierarchy_depth_inter_slice present in the PH referring to the SPS. The syntax element sps_max_mtt_hierarchy_depth_inter_slice has a value range of 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. In some embodiments, the bitstream conformance requires that the value of (MinQtLog2SizeInterY - sps_max_mtt_hierachy_depth_inter_slice / 2) is less than or equal to MinCbLog2SizeY.
[0372] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the default difference, in base 2 logarithm, of the minimum size in luma samples of a chroma leaf block resulting from quad-tree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA, and the minimum coding block size in luma samples of a chroma CU with treeType equal to DUAL_TREE_CHROMA in slices with slice_type equal to 2 (I) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_chroma present in the PH referring to the SPS. The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma has a value range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be 0. The minimum size in luma samples of a chroma leaf block resulting from quad-tree partitioning of a CTU with treeType equal to DUAL_TREE_CHROMA is derived as follows:
[0373] MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY
[0374] The syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma specifies the default maximum level depth of chroma coding units that are produced by multi-type tree partitioning of color quad-tree leaves with treeType equal to DUAL_TREE_CHROMA in slices of slices_type equal to 2 (I) referring to the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default maximum level depth can be overridden by the syntax element pic_max_mtt_hierarchy_depth_chroma present in the PH referring to the SPS. The value of the syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma is in the range of 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When not present, the value of the syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma is inferred to be 0. In some embodiments, the bitstream conformance requires that the value of (MinQtLog2SizeIntraC - sps_max_mtt_hierarchy_depth_intra_slice_chroma / 2) is less than or equal to MinCbLog2SizeY.
[0375] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma specifies the difference between the base-2 logarithm of the minimum size of luma samples of luma leaf blocks produced by quad-tree partitioning of CTUs and the base-2 logarithm of the minimum coding block size of luma samples of luma CUs in slices of slices_type equal to 2 (I) associated with the PH. The value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma.
[0376] The syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma specifies the maximum hierarchy depth of coding units resulting from multi-type tree partitioning of quad-tree leaves in slices of a slice type equal to 2 (I) associated with a PH. The value of the syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma is in the range of 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When not present, the value of the syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma is inferred to be equal to the syntax element sps_max_mtt_hierarchy_depth_intra_slice_luma. In some embodiments, the requirement of bitstream conformance (pic_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY - pic_max_mtt_hierarchy_depth_intra_slice_luma / 2) is less than or equal to MinCbLog2SizeY.
[0377] The syntax element pic_log2_diff_min_qt_min_cb_inter_slice specifies the difference between the base-2 logarithm of the minimum size of luma samples of luma leaf blocks resulting from quad-tree partitioning of CTUs and the base-2 logarithm of the minimum luma coding block size in luma samples of luma CUs in slices of a slice type equal to 0 (B) or 1 (P) associated with a PH. The value of the syntax element pic_log2_diff_min_qt_min_cb_inter_slice is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sp_log2_diff_min_qt_min_cb_inter_slice.
[0378] The syntax element pic_max_mtt_hierarchy_depth_inter_slice specifies the maximum level depth of coding units resulting from multi-type tree partitioning of quad-tree leaves in slices with slice type equal to 0 (B) or 1 (P) associated with a PH. The syntax element pic_max_mtt_hierarchy_depth_inter_slice has a value range of 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When not present, the value of the syntax element pic_max_mtt_hierarchy_depth_inter_slice is inferred to be equal to the syntax element sps_max_mtt_hierarchy_depth_inter_slice. In some embodiments, the value of the requirement for bitstream conformance (pic_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY - pic_max_mtt_hierarchy_depth_inter_slice / 2) is less than or equal to MinCbLog2SizeY.
[0379] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the difference, in base 2 logarithm, of the minimum size of luma samples of chroma leaf blocks resulting from quad-tree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA and the base 2 logarithm of the minimum coding block size in luma samples of a chroma CU with treeType equal to DUAL_TREE_CHROMA in slices with slice type equal to 2 (I) associated with a PH. The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma has a value range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma.
[0380] The syntax element pic max mtt hierarchy depth intra slice chroma specifies the maximum hierarchy depth of a chroma coding unit that is produced by multi-type tree partitioning of a chroma quad-tree leaf in a slice with slice type equal to 2 (I) associated with a PH having treeType equal to DUAL_TREE_CHROMA. The syntax element pic max mtt hierarchy depth intra slice chroma has a value range of 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When not present, the value of the syntax element pic max mtt hierarchy depth intra slice chroma is inferred to be equal to the syntax element sps max mtt hierarchy depth intra slice chroma. In some embodiments, a conformance requirement of the bitstream is that (pic_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY - pic_max_mtt_hierarchy_depth_intra_slice_chroma / 2) is less than or equal to MinCbLog2SizeY.
[0381] Figure 20 A flowchart of an example video processing method 2000 according to some embodiments of the disclosure is shown. The method 2000 can be performed, for example, by an encoder (e.g., by a process 200A of Figure 2A or a process 200B of Figure 2B , a decoder (e.g., by a process 300A of Figure 3A or a process 300B of Figure 3B ), or by one or more software or hardware components of an apparatus (e.g., the apparatus 400 of Figure 4 ). For example, a processor (e.g., the processor 402 of Figure 4 ) can perform the method 2000. In some embodiments, the method 2000 can be implemented by a computer program product embodied in a computer readable medium, including computer executable instructions, such as program code, executed by a computer (e.g., the apparatus 400). Figure 4
[0382] At step 2001, it can be determined whether a coding block includes samples outside of an image boundary. In some embodiments, the image boundary can be a bottom image boundary or a right image boundary. As an example of a coding block that includes samples outside of an image boundary, in Figure 17 In some embodiments, the coding block 1701 exceeds the right image boundary of the picture 1700, the coding block 1705 exceeds the bottom image boundary of the picture 1700, and the coding block 1703 exceeds the bottom and right image boundaries of the picture 1700.
[0383] At step 2003, in response to the coding block being determined to include samples outside of the image boundary, the coding block can be split using QT mode. In some embodiments, in response to the coding block being determined to include samples outside of the image boundary, the method 2000 can determine that the coding block is not allowed to be split using BT mode and TT mode. For example, the variables allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer can be determined to equal FALSE.
[0384] In some embodiments, in response to the coding block being determined to include samples outside of the image boundary, the method 2000 can determine that the coding block is to be split using QT mode regardless of whether a QT flag is present in the bitstream that includes the coding block. The QT flag indicates whether the coding block is split using QT mode. For example, when the syntax element split_qt_flag is not present, if the coding block includes samples outside of the image boundary and the variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer equal FALSE or allowSplitQt equals TRUE, then the value of split_qt_flag is inferred to equal 1.
[0385] In some embodiments, in response to determining that the coding block includes samples outside of the image boundary, the method 2000 can determine that the coding block is allowed to be split using QT mode regardless of a preset constraint on a minimum block size for which the QT mode is allowed to be applied. For example, a minimum QT size constraint can not be applied to a coding block that is located at an image boundary.
[0386] In some embodiments, the preset constraint can include bitstream conformance for the coding block. The bitstream conformance can set a minimum block size for which the QT mode is allowed to be applied. For example, the minimum block size can be set to be less than or equal to 64. The bitstream conformance can also set a maximum BT depth or a maximum TT depth.
[0387] In some embodiments, the method 2000 can include determining that the coding block is to be split. For example, a syntax element split_cu_flag can be used to indicate whether the coding block is to be split. When the syntax element split_cu_flag is not present, it can be inferred to equal 1, which indicates that the coding block is to be split.
[0388] It should be understood that embodiments of the disclosure can be combined with another embodiment or some other embodiments.
[0389] Embodiments can be further described using the following clauses:
[0390] 1. A method of video processing, comprising:
[0391] determining whether a coding block includes samples outside of an image boundary; and
[0392] in response to determining that the coding block includes samples outside of an image boundary, performing quad-tree partitioning of the coding block regardless of a value of a first parameter, wherein the first parameter indicates whether the quad-tree is allowed to be used for partitioning the coding block.
[0393] 2. The method of clause 1, further comprising:
[0394] determining a value of a first flag of the coding block, the first flag indicating whether the coding block is partitioned into sub-blocks; and
[0395] determining values of a second parameter, a third parameter, a fourth parameter, and a fifth parameter of the coding block, the second parameter, the third parameter, the fourth parameter, and the fifth parameter respectively indicating whether a binary-horizonal tree, a binary-vertical tree, a ternary-horizonal tree, a ternary-vertical tree is allowed to be used to split the coding block.
[0396] 3. The method of clause 2, further comprising:
[0397] in response to the value of the first flag being equal to 1 and the values of the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter being equal to 0, setting a value of a second flag of the coding block to 1, the second flag indicating whether the quad-tree is used to partition the coding block.
[0398] 4. The method of clause 2, further comprising:
[0399] in response to the value of the first parameter being equal to 1, setting a value of a second flag of the coding block to 1, the second flag indicating whether the quad-tree is used to partition the coding block.
[0400] 5. The method of any of clauses 2-4, further comprising:
[0401] setting the value of the first flag to 1 when it is determined that the coding block includes samples outside of an image boundary.
[0402] 6. A video processing apparatus comprising:
[0403] at least one memory configured to store instructions, and
[0404] at least one processor configured to execute the instructions to cause the apparatus to perform operations comprising:
[0405] determining whether the coding block includes samples outside of an image boundary; and
[0406] in response to determining that the coding block includes samples outside of an image boundary, performing quad-tree partitioning of the coding block regardless of a value of a first parameter, wherein the first parameter indicates whether the quad-tree is allowed to be used to partition the coding block.
[0407] 7. The apparatus of clause 6, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform operations comprising:
[0408] determining a value of a first flag of the coding block, the first flag indicating whether the coding block is partitioned into a plurality of sub-blocks; and
[0409] determining values of a second parameter, a third parameter, a fourth parameter, and a fifth parameter of the coding block, the second parameter, the third parameter, the fourth parameter, and the fifth parameter respectively indicating whether a binary-horizonal tree, a binary-vertical tree, a ternary-horizonal tree, a ternary-vertical tree is allowed to be used to partition the coding block.
[0410] 8. The apparatus of clause 7, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform operations comprising:
[0411] in response to the value of the first flag being equal to 1 and the values of the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter being equal to 0, setting a value of a second flag of the coding block to 1, the second flag indicating whether the coding block is partitioned using a quad-tree,
[0412] 9. The apparatus of clause 7, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform operations comprising:
[0413] in response to the value of the first parameter being equal to 1, setting a value of a second flag of the coding block to 1, the second flag indicating whether the coding block is partitioned using the quad-tree.
[0414] 10. The apparatus of any of clauses 7-9, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform operations comprising:
[0415] setting the value of the first flag to 1 when the coding block is determined to include samples outside of an image boundary.
[0416] 11. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processing devices to cause a video processing apparatus to perform operations comprising:
[0417] determining whether the coding block includes samples outside of a picture boundary, and
[0418] in response to the coding block being determined to include samples outside of a picture boundary, performing quad-tree partitioning of the coding block regardless of a value of a first parameter, wherein the first parameter indicates whether the quad-tree is allowed to be used to partition the coding block.
[0419] 12. The non-transitory computer-readable storage medium of clause 11, wherein the set of instructions is executable by the one or more processing apparatuses to cause the video processing device to perform:
[0420] determining a value of a first flag of the coding block, the first flag indicating whether the coding block is partitioned into a plurality of sub-blocks; and
[0421] determining values of a second parameter, a third parameter, a fourth parameter, and a fifth parameter of the coding block, the second parameter, the third parameter, the fourth parameter, and the fifth parameter respectively indicating whether a binary-horizonal tree, a binary-vertical tree, a ternary-horizonal tree, a ternary-vertical tree is allowed to be used to partition the coding block.
[0422] 13. The non-transitory computer-readable storage medium of clause 12, wherein the set of instructions is executable by the one or more processing apparatuses to cause the video processing device to perform:
[0423] in response to the value of the first flag being equal to 1 and the values of the first parameter, the second parameter, the third parameter, the fourth, and the fifth parameter being equal to 0, setting a value of a second flag of the coding block to 1, the second flag indicating whether the coding block is partitioned using a quad-tree.
[0424] 14. The non-transitory computer-readable storage medium of clause 12, wherein the set of instructions is executable by the one or more processing apparatuses to cause the video processing device to perform:
[0425] in response to the value of the first parameter being equal to 1, setting a value of a second flag of the coding block to 1, the second flag indicating whether the coding block is partitioned using the quad-tree.
[0426] 15. The non-transitory computer-readable storage medium of any of clauses 12-14, wherein in response to the coding block being determined to include samples outside of a picture boundary, partitioning the coding block using the QT mode comprises:
[0427] setting the value of the first flag to 1 when the coding block is determined to include samples outside of a picture boundary.
[0428] 16. A video processing method comprising:
[0429] determining whether the coding block includes samples outside of an image boundary;
[0430] responsive to the coding block being determined to include samples outside of an image boundary, partitioning the coding block using a quad-tree (QT) mode.
[0431] 17. The method of clause 16, further comprising:
[0432] responsive to the coding block being determined to include samples outside of an image boundary, determining that a binary tree (BT) mode and a ternary tree (TT) mode are not allowed to partition the coding block.
[0433] 18. The method of any of clauses 16 and 17, wherein responsive to the coding block being determined to include samples outside of an image boundary, partitioning the coding block using a quad-tree (QT) mode comprises:
[0434] partitioning the coding block using the QT mode regardless of whether a QT flag is present in a bitstream that includes the coding block, the QT flag indicating whether the coding block is to be partitioned using the QT mode.
[0435] 19. The method of clause 16, wherein responsive to the coding block being determined to include samples outside of an image boundary, partitioning the coding block using a quad-tree (QT) mode comprises:
[0436] partitioning the coding block using the QT mode regardless of a preset constraint on a minimum block size at which the QT mode is allowed to be applied.
[0437] 20. The method of clause 19, wherein:
[0438] the preset constraint comprises a bitstream conformance associated with the coding block that sets a minimum block size at which the QT mode is allowed to be applied.
[0439] 21. The method of clause 20, wherein the minimum block size is set to be less than or equal to 64.
[0440] 22. The method of any of clauses 20 and 21, wherein the bitstream conformance sets a maximum BT depth or a maximum TT depth.
[0441] 23. The method of any of clauses 16-22, wherein the image boundary is a bottom image boundary or a right image boundary.
[0442] 24. A video processing apparatus comprising:
[0443] at least one memory configured to store instructions; and
[0444] at least one processor configured to execute the instructions to cause the apparatus to perform:
[0445] determining whether the coding block includes samples outside of an image boundary;
[0446] responsive to the coding block being determined to include samples outside of an image boundary, partitioning the coding block using a quad-tree (QT) mode.
[0447] 25. The apparatus of clause 24, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:
[0448] responsive to the coding block being determined to include samples outside of an image boundary, determining that a binary-tree (BT) mode and a ternary-tree (TT) mode are not allowed to partition the coding block.
[0449] 26. The apparatus of any of clauses 24 and 25, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:
[0450] partitioning the coding block using the QT mode regardless of whether a QT flag is present in a bitstream that includes the coding block, the QT flag indicating whether the coding block is to be partitioned using the QT mode.
[0451] 27. The apparatus of clause 24, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform:
[0452] partitioning the coding block using the QT mode regardless of a preset constraint on a minimum block size at which the QT mode is allowed to be applied.
[0453] 28. The apparatus of clause 27, wherein
[0454] the preset constraint comprises a bitstream conformance associated with the coding block that sets a minimum block size at which the QT mode is allowed to be applied.
[0455] 29. The apparatus of clause 28, wherein the minimum block size is set to be less than or equal to 64.
[0456] 30. The apparatus of any of clauses 28 and 29, wherein the bitstream conformance sets a maximum BT depth or a maximum TT depth.
[0457] 31. The apparatus of any of clauses 24 to 30, wherein the image boundary is a bottom image boundary or a right image boundary.
[0458] 32. A non-transitory computer-readable storage medium storing a set of instructions capable of being executed by one or more processing devices to cause a video processing apparatus to perform the following operations:
[0459] determining whether the coding block includes samples outside of an image boundary;
[0460] responsive to the coding block being determined to include samples outside of an image boundary, partitioning the coding block using a quad-tree (QT) mode.
[0461] 33. The non-transitory computer-readable storage medium of clause 32, wherein the set of instructions is capable of being executed by the one or more processing devices to cause the video processing apparatus to perform:
[0462] responsive to the coding block being determined to include samples outside of an image boundary, determining that a binary-tree (BT) mode and a ternary-tree (TT) mode are not allowed to partition the coding block.
[0463] 34. The non-transitory computer-readable storage medium of any of clauses 32 and 33, wherein the set of instructions is capable of being executed by the one or more processing devices to cause the video processing apparatus to perform:
[0464] partitioning the coding block using the QT mode regardless of whether a QT flag is present in a bitstream that includes the coding block, the QT flag indicating whether the coding block is to be partitioned using the QT mode.
[0465] 35. The non-transitory computer-readable storage medium of clause 32, wherein the set of instructions is capable of being executed by the one or more processing devices to cause the video processing apparatus to perform:
[0466] partitioning a coding block using a QT mode regardless of a preset constraint on a minimum block size to which the QT mode is allowed to be applied.
[0467] 36. The non-transitory computer-readable storage medium of clause 35, wherein: the preset constraint comprises a bitstream conformance associated with the coding block that sets a minimum block size to which the QT mode is allowed to be applied.
[0468] 37. The non-transitory computer-readable storage medium of clause 36, wherein the minimum block size is set to be less than or equal to 64.
[0469] 38. The non-transitory computer-readable storage medium of any of clauses 36 and 37, wherein the bitstream conformance sets a maximum BT depth or a maximum TT depth.
[0470] 39. The non-transitory computer-readable storage medium of any of clauses 32 to 38, wherein the picture boundary is a bottom picture boundary or a right picture boundary.
[0471] In some embodiments, non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by a device, such as the disclosed encoders and decoders, for performing the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and a networked version of any of the foregoing. The device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0472] It should be noted that relational terms, such as“first” and“second,” are used herein solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the words“comprises,”“has,”“includes,” and“containing” and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items.
[0473] As used herein, unless specifically stated otherwise, the term“or” includes all possible combinations, unless impractical. For example, if a database is said to include A or B, then unless specifically stated otherwise or impractical, the database can include A, or B, or A and B. As a second example, if a database is said to include A, B, or C, then unless specifically stated otherwise or impractical, the database can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0474] It should be understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computing units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.
[0475] In the foregoing specification, embodiments have been described with reference to numerous specific details that can vary with implementation. Certain modifications and changes can be made thereto by those of ordinary skill in the art. Other embodiments will be apparent to those of ordinary skill in the art from consideration of the specification and practice of the application disclosed herein. The specification and examples are to be considered exemplary only, with the true scope and spirit of the application indicated by the claims. The sequence of steps shown in the drawings is for illustrative purposes only and is not intended to be limited to any particular sequence of steps. Thus, those skilled in the art will appreciate that the steps can be performed in different orders while still implementing the same method.
[0476] In the drawings and specification, there have been disclosed exemplary embodiments. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A method of video processing, comprising: determining a value of a first parameter for an intra slice, wherein the first parameter indicates a minimum size in a plurality of luma samples of a chroma leaf block resulting from a quad-tree partitioning of a coding tree unit; determining a value of a second parameter, wherein the second parameter indicates a minimum luma coding block size; determining whether a coding block includes samples outside of a picture boundary; determining a size of the coding block; in response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside of the picture boundary, partitioning the coding block into coding blocks having a horizontal size and a vertical size that are each half of the coding block.
2. The video processing method of claim 1, wherein, the first value is 128.
3. The method of claim 1, wherein, in response to the first parameter being equal to the first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside of the picture boundary, determining that a first flag signaled in a bitstream has a second value, wherein the first flag equal to the second value indicates that each of a plurality of coding blocks is partitioned into coding blocks having a horizontal size and a vertical size that are each half of the coding block.
4. The method of claim 3, wherein, the second value is equal to 1.
5. The method of claim 3, wherein, the first flag comprises split_qt_flag.
6. The method of claim 1, further comprising: determining values of second, third, fourth, fifth, and sixth parameters associated with the coding block, the parameters respectively indicating whether use of quad partitioning, binary horizontal partitioning, binary vertical partitioning, ternary horizontal partitioning, and ternary vertical partitioning is allowed to partition the coding block.
7. The method of claim 6, further comprising: in response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, determining that the values of the second, third, fourth, fifth, and sixth parameters are each equal to a third value.
8. The method of claim 6, wherein, the values of the second, third, fourth, fifth, and sixth parameters each equal to the third value indicating that none of quad partitioning, binary horizontal partitioning, binary vertical partitioning, ternary horizontal partitioning, or ternary vertical partitioning is allowed to partition the coding block.
9. The method of claim 1, further comprising: determining that the coding block is partitioned into a plurality of coding blocks based on a value of a first flag signaled in a bitstream.
10. The method of claim 9, wherein, the first flag is equal to 1.
11. The method of claim 9, wherein, the first flag comprises split_qt_flag.
12. A system for video processing, comprising: a memory storing a set of instructions; at least one processor configured to execute the set of instructions to cause the system to perform: determining a value of a first parameter for an intra slice, wherein the first parameter indicates a minimum size in a plurality of luma samples of a chroma leaf block resulting from a quad-tree partitioning of a coding tree unit; determining a value of a second parameter, wherein the second parameter indicates a minimum luma coding block size; determining whether a coding block includes samples outside of a picture boundary; determining a size of the coding block; in response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside of the picture boundary, partitioning the coding block into coding blocks having half the horizontal and vertical size.
13. The video processing system of claim 12, wherein, the first value is equal to 0.
14. A non-transitory computer-readable storage medium storing a video bitstream generated by a method executed by a video processing apparatus, the method comprising: determining a value for a first parameter for an intra slice, wherein the first parameter indicates a minimum size in a plurality of luma samples of a chroma leaf block resulting from a quad-tree partitioning of a coding tree unit; determining a value for a second parameter, wherein the second parameter indicates a minimum luma coding block size; determining whether a coding block includes samples outside of a picture boundary; determining a size of the coding block; in response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside of the picture boundary, partitioning the coding block into coding blocks having half the horizontal and vertical size.
15. The non-transitory computer-readable storage medium of claim 14, wherein, the first value is equal to 128.
16. The non-transitory computer-readable storage medium of claim 14, wherein, the bitstream includes the first value, the method further comprising: in response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside of the picture boundary, determining that a first flag sent in the bitstream has a second value, wherein the first flag equal to the second value indicates that each of the plurality of coding blocks is partitioned into coding blocks having half the horizontal size and half the vertical size.
17. The non-transitory computer-readable storage medium of claim 16, wherein the second value is equal to 1.
18. The non-transitory computer-readable storage medium of claim 16, wherein the first flag comprises split_qt_flag.
19. The non-transitory computer-readable storage medium of claim 14, the method further comprising: determining values for second, third, fourth, fifth, and sixth parameters associated with the coding block, the parameters respectively indicating whether use of quad partitioning, binary horizontal partitioning, binary vertical partitioning, ternary horizontal partitioning, and ternary vertical partitioning is allowed to partition the coding block.
20. The non-transitory computer-readable storage medium of claim 14, the method further comprising: in response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, determining that the values for the second, third, fourth, fifth, and sixth parameters are all equal to a third value.
Citation Information
Patent Citations
Decoder and method for generating high dynamic range image, and storage medium
CN107896332A
Video decoding method and video decoder
US20150350682A1