Video processing method and device, and non-temporary computer readable storage medium
By performing quad-tree segmentation at the image boundary, the problem of inefficient video encoding block division in the prior art is solved, and more efficient video encoding performance is achieved.
Patent Information
- Application Number
- CN202510449480.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-17
- Filing Date
- 2020-11-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-24
AI Technical Summary
The block division method of existing video encoding technology at image boundaries is inefficient, especially in efficient video encoding standards such as VVC/H.266. How to effectively divide encoding blocks to improve encoding performance is a challenge.
By determining whether the encoding block includes samples outside the image boundary and performing quad-tree segmentation in this case, the effective partitioning of the encoding blocks is ensured, thus adapting to the needs of different encoding standards.
More efficient block division at image boundaries is achieved, and the performance and efficiency of video encoding are improved, especially when using efficient video encoding standards.
Smart Images

Figure CN120201190A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This disclosure claims priority to U.S. Provisional Application No. 62 / 948,856, filed on December 17, 2019, which is hereby incorporated by reference in its entirety. Technical Field
[0002] This disclosure generally relates to video processing, and more particularly, to methods and apparatuses for block partitioning at image boundaries. Background Art
[0003] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, and the most common ones are based on prediction, transformation, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, specify specific video coding formats and are developed by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is getting higher and higher. Summary of the Invention
[0004] In some embodiments, an exemplary video processing method includes: determining whether an encoding block includes samples outside an image boundary; and in response to the encoding block being determined to include samples outside the image boundary, performing quadtree partitioning of the encoding block regardless of the value of a first parameter, where the first parameter indicates whether the quadtree is allowed to be used to partition the encoding block.
[0005] In some embodiments, an exemplary video processing apparatus includes at least one memory for storing instructions and at least one processor. The at least one processor is configured to execute the instructions to cause the apparatus to perform: determining whether an encoding block includes samples outside an image boundary; and in response to the encoding block being determined to include samples outside the image boundary, performing quadtree partitioning of the encoding block regardless of the value of a first parameter, where the first parameter indicates whether the quadtree is allowed to be used to partition the encoding block.
[0006] In some embodiments, an exemplary non - transitory computer - readable storage medium stores a set of instructions. The set of instructions can be executed by one or more processing devices to cause a video processing device to perform: determining whether an encoding block includes samples outside an image boundary, and in response to the encoding block being determined to include samples outside the image boundary, performing quadtree splitting of the encoding block regardless of the value of a first parameter, where the first parameter indicates whether the quadtree is allowed to be used for splitting the encoding block. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments and aspects of the present disclosure are shown in the following detailed description and the drawings. The various features shown in the figures are not drawn to scale.
[0008] Figure 1 is a schematic structural diagram of an exemplary video sequence according to some embodiments of the present disclosure.
[0009] Figure 2A is a schematic diagram showing an exemplary encoding process of a hybrid video coding system according to an embodiment of the present disclosure.
[0010] Figure 2B is a schematic diagram showing another exemplary encoding process of a hybrid video coding system according to an embodiment of the present disclosure.
[0011] Figure 3A is a schematic diagram showing an exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure.
[0012] Figure 3B is a schematic diagram showing another exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure.
[0013] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding a video according to some embodiments of the present disclosure.
[0014] Figure 5 is a schematic diagram showing an example of a multi - type tree splitting pattern according to some embodiments of the present disclosure.
[0015] Figure 6 is a schematic diagram showing an exemplary signaling mechanism for partitioning division information in a quadtree (QT) having a nested multi - type tree coding tree structure according to some embodiments of the present disclosure.
[0016] Figure 7 shows an exemplary Table 1 according to some embodiments of the present disclosure, which shows an exemplary derivation of a multi - type tree partitioning pattern (MttSplitMode) based on multi - type tree syntax elements.
[0017] Figure 8is a schematic diagram showing an example of an unacceptable ternary tree (TT) and binary tree (BT) partitioning according to some embodiments of the present disclosure.
[0018] Figure 9 is a schematic diagram showing an exemplary block partitioned at an image boundary according to some embodiments of the present disclosure.
[0019] Figure 10 Shows an exemplary Table 2 according to some embodiments of the present disclosure, which shows exemplary specifications of parallel ternary tree partitioning (parallelTtSplit) and coded block size (cbSize) based on a binary partitioning pattern (btSplit).
[0020] Figure 11 Shows an exemplary Table 3 according to some embodiments of the present disclosure, which shows exemplary specifications of cbSize based on a ternary tree partitioning pattern (ttSplit).
[0021] Figure 12 Shows an exemplary Table 4 according to some embodiments of the present disclosure, which shows an exemplary coding tree syntax.
[0022] Figure 13 is a schematic diagram showing an example of a multi-type tree splitting mode indicated by MttSplitMode according to some embodiments of the present disclosure.
[0023] Figure 14 Shows an exemplary Table 5 according to some embodiments of the present disclosure, which shows exemplary specifications of MttSplitMode.
[0024] Figure 15 Shows an exemplary Table 6 according to some embodiments of the present disclosure, which shows an exemplary sequence parameter set RBSP syntax.
[0025] Figure 16 Shows an exemplary Table 7 according to some embodiments of the present disclosure, which shows an exemplary picture header RBSP syntax.
[0026] Figure 17 is a schematic diagram showing an exemplary block where QT, TT, or BT partitioning is not allowed at an image boundary according to some embodiments of the present disclosure.
[0027] Figure 18 is a schematic diagram showing an exemplary block where QT, TT, or BT partitioning is not allowed at an image boundary according to some embodiments of the present disclosure.
[0028] Figure 19 is a schematic diagram showing an example of using BT and TT partitioning according to some embodiments of the present disclosure.
[0029] Figure 20 The flowchart of an exemplary video processing method according to some embodiments of the present disclosure is shown. Detailed implementation
[0030] Reference will now be made in detail to exemplary embodiments, which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, unless otherwise specified, where the same numerals in different drawings represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present disclosure as set forth in the appended claims. Specific aspects of the present disclosure are described in more detail below. If there is a conflict with the terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0031] The Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0032] In order to achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been using the Joint Exploration Model (JEM) reference software to explore technologies other than HEVC. As coding technologies are incorporated into JEM, JEM has achieved higher coding performance than HEVC.
[0033] The VVC standard has been recently developed and continues to include more coding technologies that provide better compression performance. VVC is based on the hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0034] Video is a set of static images (or "frames") arranged in chronological order to store visual information. These images can be acquired and stored in chronological order using a video acquisition device (e.g., a camera), and such images in the time series can be displayed using a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display function). In addition, in some applications, the video acquisition device can send the acquired video to the video playback device (e.g., a computer with a monitor) in real time, such as for surveillance, conferencing, or live broadcasting.
[0035] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally referred to as an "encoder", and the module for decompression is generally referred to as a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code fixed in a computer-readable medium, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be referred to as a "transcoder".
[0036] The video encoding process can identify and retain useful information that can be used to reconstruct the image and ignore unimportant reconstruction information. If the unimportant information cannot be fully reconstructed by ignoring it, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off to reduce the required storage space and transmission bandwidth.
[0037] The useful information of the encoded image (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in the position of pixels, brightness changes, or color changes, with position changes being the most concerned. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.
[0038] An image encoded without referring to another image (i.e., it is its own reference image) is called an "I-image". An image encoded using a previous image as a reference image is called a "P-image", and an image encoded using a previous image and a future image as reference images is called a "B-image" (the reference is "bidirectional").
[0039] Figure 1Shows the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a video that has been captured and archived. Video 100 can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination of both (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider.
[0040] As Figure 1 shown, the video sequence 100 can include a series of images arranged in time along a timeline, including images 102, 104, 106, and 108. Images 102 - 106 are consecutive, with more images between images 106 and 108. In Figure 1 this example, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrows. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It should be noted that the reference images of images 102 - 106 are merely examples, and the present disclosure does not limit the embodiments of the reference images as Figure 1 shown.
[0041] Generally, due to the computational complexity of the encoding and decoding tasks, video codecs do not encode or decode an entire image at once. Instead, they can divide an image into basic segments and encode or decode the image segments segment by segment. In the present disclosure, such a basic segment is referred to as a basic processing unit (“BPU”). For example, Figure 1The structure 110 in [the context] shows an example structure of an image (e.g., any of the images 102 - 108) in the video sequence 100. In the structure 110, the image is divided into 4×4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit may have a variable size in the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing unit can be selected for the image based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit.
[0042] The basic processing unit can be a logical unit, which may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image may include a luminance component (Y) representing achromatic luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components may be referred to as "coding tree blocks" ("CTB"). Any operation performed on the basic processing unit can be repeated for each of its luminance and chrominance components.
[0043] Video coding has multiple operation stages, examples of which are Figures 2A - 2B and Figures 3A - 3BAs shown. For each stage, the size of the basic processing unit may still be too large for processing, so it can be further divided into segments called "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or an "encoding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to further levels according to processing needs. It should also be noted that different stages may use different schemes to divide the basic processing unit.
[0044] For example, in the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for the basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC), and decide the prediction type for each individual basic processing subunit.
[0045] For another example, in the prediction stage (an example of which is shown in Figures 2A - 2B ), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations can be performed.
[0046] For another example, in the transform stage (an example of which is shown in Figures 2A - 2BAs shown (in [reference], etc.), the encoder may perform a transformation operation on a residual basic processing unit (e.g., a CU). However, in some cases, the basic processing unit may still be too large to process. The encoder may further divide the basic processing unit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level the transformation operation can be performed. It should be noted that the partitioning scheme for the same basic processing unit may be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU may have different sizes and numbers.
[0047] In Figure 1 the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, the boundaries of which are shown as dashed lines. Different basic processing units of the same image may be divided into basic processing sub-units in different schemes.
[0048] In some embodiments, to provide the ability for parallel processing of video encoding and decoding and fault tolerance, an image may be divided into regions for processing such that for a region of the image, the encoding or decoding process may not depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving the encoding efficiency. Additionally, when the data of a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thereby providing fault tolerance. In some video coding standards, an image may be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images of the video sequence 100 may have different partitioning schemes for dividing the image into regions.
[0049] For example, in Figure 1 the structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing sub-units, and regions of the structure 110 in [example] are only examples, and the present disclosure does not limit its embodiments.
[0050] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As Figure 2AAs shown, the encoder may encode video sequence 202 into video bitstream 228 according to process 200A. Similar to Figure 1 the video sequence 100 in Figure 1 , the video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to
[0051] the structure 110 in Figure 2A , each original image of the video sequence 202 may be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original image of the video sequence 202. For example, the encoder may perform process 200A in an iterative manner, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions (e.g., regions 114 - 118) of each original image of the video sequence 202.
[0052] Referring to Figure 2A , the encoder may feed a basic processing unit of an original image of the video sequence 202 (referred to as "original BPU") to prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. The encoder may feed the residual BPU 210 to transform stage 212 and quantization stage 214 to 216 generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder may feed the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0052] The encoder may iteratively perform process 200A to encode each original BPU (in the forward path) of the encoded original image and generate prediction reference 224 for the next original BPU (in the reconstruction path) of the encoded original image. After encoding all the original BPUs of the original image, the encoder may continue to encode the next image in the video sequence 202.
[0053] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, fetching, reading, accessing, or using for inputting data in any way.
[0054] In the prediction stage 204, at the current iteration, the encoder may receive an original BPU and a prediction reference 224, and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and the prediction reference 224 that can be used to reconstruct the original BPU into the predicted BPU 208.
[0055] Ideally, the predicted BPU 208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To record these differences, when generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the value of the corresponding pixel of the predicted BPU 208 (e.g., grayscale value or RGB value) from the value of the pixel of the original BPU. Each pixel of the residual BPU 210 may have a residual value as the result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared with the original BPU, the prediction data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0056] To further compress the residual BPU 210, in the transform stage 212, the encoder may reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "base patterns". Each base pattern is associated with a "transformation coefficient". The base patterns may have the same size (e.g., the size of the residual BPU 210), and each base pattern may represent a frequency component of the variation of the residual BPU 210 (e.g., the frequency of brightness variation). None of the base patterns can be reproduced from any combination (e.g., linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the base images are similar to the basic functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basic functions.
[0057] Different transformation algorithms can use different basic patterns. Various transformation algorithms can be used at transformation stage 212, for example, discrete cosine transform, discrete sine transform, etc. The transformation at transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, to recover the pixels of the residual BPU 210, the inverse transformation can be multiplying the values of the corresponding pixels of the basic pattern by the corresponding correlation coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus have the same basic pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from them without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.
[0058] The encoder can further compress the transformation coefficients at quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without causing significant quality degradation in decoding. For example, at quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, and the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the quantized transformation coefficients 216 with zero values, whereby the transformation coefficients are further compressed. This quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0059] Since the encoder ignores the remainder of this division in the rounding operation, quantization stage 214 can be lossy. Generally, quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer bits required for the quantized transformation coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.
[0060] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using binary coding techniques, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context - adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform at the transform stage 212, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameter), etc. The encoder may use the output data of the binary coding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 may be further packed for network transmission.
[0061] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the encoder may generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction reference 224 that will be used in the next iteration of process 200A.
[0062] It should be noted that other variants of process 200A may be used to encode the video sequence 202. In some embodiments, the stages of process 200A may be executed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit Figure 2A one or more of the stages.
[0063] Figure 2B FIG. shows a schematic diagram of another example encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and the reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0064] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-prediction") can use pixels from one or more already-encoded neighboring BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the spatial redundancy inherent in an image. Temporal prediction (e.g., inter-image prediction or "inter-prediction") can use regions from one or more already-encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in an image.
[0065] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-prediction. For the original BPU of the encoded image, the prediction reference 224 can include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate a predicted BPU 208 by interpolating the neighboring BPUs. The interpolation technique can include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at the pixel level, for example, by interpolating the values of the corresponding pixels of each pixel of the predicted BPU 208. The neighboring BPUs used for interpolation can be located in various directions relative to the original BPU, such as in the vertical direction (e.g., on top of the original BPU), horizontal direction (e.g., to the left of the original BPU), diagonal direction (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard being used. For intra-prediction, the prediction data 206 can include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, the parameters of the interpolation, the direction of the neighboring BPUs relative to the original BPU, etc.
[0066] For another example, at the time prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a per-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same image have been generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform an operation of "motion estimation" to search for a matching region within a range of the reference images (referred to as the "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a position in the reference image that has the same coordinates as the original BPU in the current image and may extend outward a predetermined distance. When the encoder identifies (e.g., by using a pel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller, equal to, larger, or a different shape) than the original BPU. Since the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 shown), the matching region may be considered to "move" over time to the position of the original BPU. The encoder may record the direction and distance of this motion as a "motion vector". When multiple reference images are used (e.g., as in the image 106 in Figure 1 ), the encoder may search for the matching region and determine its associated motion vector for each reference image. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0067] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0068] To generate the predicted BPU 208, the encoder may perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., as in Figure 1For the image 106) therein, the encoder can move the matching region of the reference image according to the respective motion vectors and average pixel values of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder can sum up the weighted sums of the pixel values of the moved matching regions.
[0069] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same time direction relative to the current image. For example, Figure 1 the image 104 therein is an unidirectional inter-frame prediction image, where the reference image (i.e., image 102) is before image 04. Bidirectional inter-frame prediction can use one or more reference images in two time directions relative to the current image. For example, Figure 1 the image 106 therein is a bidirectional inter-frame prediction image, where the reference images (i.e., images 104 and 08) are in two time directions relative to image 104.
[0070] Still referring to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of the intra-frame prediction or the inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode to minimize the value of the cost function according to the bit rate of the candidate prediction modes and the distortion of the reconstructed reference images under the candidate prediction modes. According to the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0071] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232. At this stage, the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion introduced by inter prediction (e.g., blocking artifacts). The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference image can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter prediction reference image for future images of the video sequence 202). The encoder can store one or more reference images in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., loop filter strength) as well as the quantized transform coefficients 216, prediction data 206, and other information at the binary coding stage 226.
[0072] Figure 3A FIG. shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present invention. Process 300A can be a decompression process corresponding to Figure 2A the compression process 200A therein. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss during the compression and decompression processes (e.g., Figures 2A - 2B the quantization stage 214 therein), generally, the video stream 304 is different from the video sequence 202. Similar to Figures 2A - 2B processes 200A and 200B therein, the decoder can perform process 300A on each image encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform process 300A in an iterative manner, where the decoder can decode the basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region (e.g., regions 114-118) of each image encoded in the video bitstream 228.
[0073] As Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with the basic processing unit of the encoded image (referred to as the "encoded BPU") into the binary decoding stage 302, where the decoder can decode this portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 into the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. The decoder can feed the prediction data 206 into the prediction stage 204 to generate the predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in a computer memory). The decoder can feed the prediction reference 224 into the prediction stage 204 for performing prediction operations in the next iteration of process 300A.
[0074] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate the prediction reference 224 for the next encoded BPU of the encoded image. After decoding all the encoded BPUs of the encoded image, the decoder can output the image to the video stream 304 for display and continue to decode the next encoded image in the video bitstream 228.
[0075] In the binary decoding stage 302, the decoder can perform the inverse operations of the binary coding techniques used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context - adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, the parameters of the prediction operation, the transform type, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameter), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder can unpack the video bitstream 228 before feeding it into the binary decoding stage 302.
[0076] Figure 3B A schematic diagram of another example decoding process 300B according to an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0077] In process 300B, for an encoded basic processing unit (referred to as the "current BPU") of a decoded encoded image (referred to as the "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder uses to encode the current BPU. For example, if the encoder uses intra prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, interpolation parameters, the directions of the adjacent BPUs relative to the original BPU, etc. For another example, if the encoder uses inter prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation can include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in the respective reference images, one or more motion vectors respectively associated with the matching regions, etc.
[0078] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B and will not be repeated hereinafter. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A
[0079] In process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference image in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can, as Figure 2B Apply the loop filter to the prediction reference 224 in the manner shown. The reference image for loop filtering can be stored in buffer 234 (e.g., the decoded picture buffer in a computer memory) for later use (e.g., as an inter-prediction reference picture for future encoded pictures of video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).
[0080] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video according to an embodiment of the present disclosure. As Figure 4 shown, apparatus 400 may include a processor 402. When processor 402 executes the instructions described herein, apparatus 400 may become a dedicated machine for video encoding or decoding. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), a field programmable gate array (FPGA), system on a chip (SoC), application specific integrated circuit (ASIC), etc. in any combination. In some embodiments, processor 402 may also be a group of processors grouped as a single logical component. For example, as Figure 4 shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0081] Apparatus 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, as Figure 4As shown, the stored data may include program instructions (e.g., for implementing the stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410), and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a group of memories grouped as a single logical component ( Figure 4 not shown).
[0082] The bus 410 may be a communication device for transferring data between components inside the device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., universal serial bus port, peripheral component interconnect express port), or the like.
[0083] For ease of explanation without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single independent module, or may be fully or partially combined into any other component of the device 400.
[0084] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0085] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As Figure 4 shown, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touch screen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video archives), etc.
[0086] Note that a video codec (e.g., the codec performing processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instances that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0087] In the quantization and inverse quantization functional blocks (e.g., Figure 2A or Figure 2B quantization 214 and inverse quantization 218 of Figure 3A or Figure 3B inverse quantization 218 of , the quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value for encoding an image or slice can be signaled at a higher level, e.g., using the init_qp_minus26 syntax element in the picture parameter set (PPS) and using the slic_qp_delta syntax element in the slice header. Additionally, delta QP values sent at the granularity of quantization groups can be used to adapt the QP value locally for each CU.
[0088] According to some embodiments, an image can be divided into multiple coding tree units (CTUs). Then, the CTUs are further divided into one or more coding units (CUs) using a quadtree (SPLIT_QT) with a nested multi-type tree having a binary and ternary partition splitting structure. Figure 5 is a schematic diagram showing an example of a multi-type tree partitioning pattern according to some embodiments of the present disclosure. As Figure 5 shown, the partitioning types in the multi-type tree structure can include quadtree split (SPLIT_QT) 501, vertical binary tree partitioning (SPLIT BT VER) 502, horizontal binary tree partitioning (split SPLIT_BT_HOR) 503, vertical ternary tree partitioning (SPLIT_TT_VER) 504, and horizontal ternary tree partitioning (SPLIT_TT_HQR) 505. The leaf nodes of the multi-type are called coding units (CUs), which can have a square or rectangular shape.
[0089] Figure 6 is a schematic diagram of an exemplary signaling mechanism for partition partitioning information in a quadtree with a nested multi-type tree coding tree structure according to some embodiments of the present disclosure. In Figure 6In it, the CTU is regarded as the root of the quadtree and is first divided by the quadtree structure. Each leaf node of the quadtree (when large enough to allow it) is further divided by the multi-type tree structure. In the multi-type tree structure, the first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the node is further divided. When the node is further divided, the second flag (e.g., mtt_split_cu_vertical_flag) is signaled to indicate the division direction, and then the third flag (e.g., mtt_split_cu_binary_flag) is signaled to indicate whether the division is a binary tree division or a ternary tree division. According to the values of the second and third flags, mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree division mode (MttSplitMode) of a CU can be derived. Figure 7 Exemplary Table 1 according to some embodiments of the present disclosure is shown, which shows an exemplary MTTSplitMode derivation based on multi-type tree syntax elements.
[0090] The virtual pipeline data unit (VPDU) is defined as a non-overlapping unit in the image. In the hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is important to keep the VPDU size. In most hardware decoders, the VPDU size can be set to 64×64 luminance samples. However, in some embodiments, the ternary tree (TT) and binary tree (BT) divisions may cause an increase in the VPDU size. Consistent with the present disclosure, in order to keep the VPDU size as 64×64 luminance samples, certain regular partition restrictions can be applied. Figure 8 Examples of disallowed TT and BT divisions according to some embodiments of the present disclosure are shown. As Figure 8 shown, for a block with a width or height, or both width and height equal to 128, TT splitting is not allowed. For a CU of 128×N, N≤64 (e.g., width equal to 128 and height less than 128), horizontal BT is not allowed. For a CU of N×128, N≤64 (e.g., height equal to 128 and width less than 128), vertical BT is not allowed.
[0091] According to the requirements in HEVC, when a part of the tree node block exceeds the bottom or right boundary of the image, the tree node block needs to be forced to split until all samples of each encoded CU are within the image boundary. The following splitting rules are applied in VVC draft 7: - If a part of the tree node block exceeds the bottom and right boundaries of the image, - If the block is a QT node and the size of the block is greater than the minimum QT size, split the block using the QT split mode. - Otherwise, split the block using the SPLIT_BT_HOR mode. - Otherwise, if a part of the tree node block exceeds the bottom boundary of the image, - If the block is a QT node, and the size of the block is greater than the minimum QT size and the size of the block is greater than the maximum BT size, split the block using the QT split mode. - Otherwise, if the block is a QT node, and the size of the block is greater than the minimum QT size and the size of the block is less than or equal to the maximum BT size, split the block using the QT split mode or the SPLIT_BT_HQR mode. - Otherwise (the block is a BT node or the size of the block is less than or equal to the minimum QT size), split the block using the SPLIT_BT_HQR mode. - Otherwise, if a part of the tree node block extends beyond the right boundary of the image, - If the block is a QT node, and the size of the block is greater than the minimum QT size and the size of the block is greater than the maximum BT size, split the block using the QT split mode, - Otherwise, if the block is a QT node, and the size of the block is greater than the minimum QT size and the size of the block is less than or equal to the maximum BT size, split the block using the QT split mode or the SPLIT_BT_VER mode. - Otherwise (e.g., the block is a BT node or the size of the block is less than or equal to the minimum QT size), split the block using the SPLIT_BT_VER mode.
[0092] Figure 9 Shows an exemplary block partitioning on the image boundary according to some embodiments of the present disclosure. As Figure 9 shown, for CTU 911, SPLIT_QT or SPLIT_BT_VER can be performed. For CTU 913, if SPLIT_QT is allowed, SPLIT_QT can be performed, and if SPLIT_QT is not allowed, SPLIT_BT_HOR can be performed. For CTU 915, SPLIT_QT or SPLIT_BT_HOR can be performed.
[0093] In VVC Draft 7, there are two parts related to block partitioning. The first is Section 6.4, which defines whether a quadtree, binary tree, or ternary tree can be used to partition a block. The output of Section 6.4 are the variables allowSplitQt, allowsplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer. These variables are used in Section 7.3.9.4 as shown in Table 4 in Figure 12 to determine whether to signal the corresponding CU-level partitioning flags (marked by boxes 1201 - 1204 in Table 4 of Figure 12 ).
[0094] Section 6.4.1 of VVC Draft 7 describes: 6.4 Availability Process 6.4.1 Allowed Quadtree Split Process The inputs to this process include: - The coded block size cbSize in the luma samples, - The multi-type tree depth mttDepth, - The variable treeType that specifies whether to use a single tree (SINGLE_TREE) or a dual tree to partition the coded tree node, and when using a dual tree, whether it is the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) being processed currently, - A variable modeType that specifies whether intra (MODE_INTRA), IBC (MODE_IBC), and inter coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter coding mode (MODE_TYPE_INTER) can be used for the coding units inside the coded tree node.
[0095] The output of this process is the variable allowSplitQt. The variable allowSplitQt is derived as follows: - If one or more of the following conditions are true, then set allowSplitQt to false (FALSE): - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY - treeType is equal to DUAL_TREE_CHROMA, and cbSize / SubWidthC is less than or equal to MinQtSizeC -mttDepth is not equal to 0 -treeType is equal to DUAL_TREE_CHROMA, and (cbSize / SubWidthC) is less than or equal to 4 -treeType is equal to DUAL_TREE_CHROMA and modeType is equal to MODE_TYPE_INTER -Otherwise, allowSplitQt is set to be equal to TRUE.
[0096] Section 6.4.2 of VVC Draft 7 describes: 6.4.2 Allowed Binary Tree Partitioning Process The inputs to this process are: - Binary tree partitioning mode btSplit, - Encoding block width cbWidth in luma samples, - Encoding block height cbHeight in luma samples, - Position (x0, y0) of the top - left luma sample of the encoding block being considered, relative to the top - left luma sample of the image, - Multi - type tree depth mttDepth, - Maximum multi - type tree depth with offset maxMttDepth, - Maximum binary tree size maxBtSize, - Minimum quadtree size minQtSize, - Partition index partIdx - Variable treeType, which specifies whether to use a single tree (SINGLE_TREE) or a dual tree to partition the coding tree node, and when using a dual tree, whether the current component being processed is the luma component (DUAL_TREE_LUMA) or the chroma component (DUAL_TREE_CHROMA), - A variable modeType, which specifies whether intra (MODE_INTRA), IBC (MODE_JBC), and inter - coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter - coding modes (MODE_TYPE_INTER) can be used for coding units inside the coding tree node.
[0097] The output of this process is the variable allowBtSplit. Figure 10 Exemplary Table 2 according to some embodiments of the present disclosure is shown, which shows exemplary specifications of variables parallelTtSplit and cbSize based on btSplit.
[0098] The variable allowBtSplit is derived as follows: - If one or more of the following conditions are true, then allowBtSplit is set equal to FALSE: - cbSize is less than or equal to MinBtSizeY - cbWidth is greater than maxBtSize - cbHeight is greater than maxBtSize - mttDepth is greater than or equal to maxMttDepth - treeType is equal to DUAL_TREE_CHROMA, and (cbWidth / SubWidthC) * (cbHeight / SubHeightC) is less than or equal to 16 - treeType is equal to DUAL_TREE_CHROMA, and (cbWidth / SubWidthC) is equal to 4, and btSplit is equal to SPLIT_BT_VER - treeType is equal to DUAL_TREE_CHROMA, and modeType is equal to MODE_TYPE_INTRA - cbWidth * cbHeight is equal to 32, and modeType is equal to MODE_TYPE_INTER - Otherwise, if all of the following conditions are true, then allowBtSplit is set to FALSE - btSplit is equal to SPLIT_BT_YER - y0 + cbHeight is greater than pic_height_in_luma_samples - Otherwise, if all of the following conditions are true, then allowBtSplit is set equal to FALSE - btSplit is equal to SPLIT_BT_VER - cbHeight is greater than 64 - x0 + cbWidth is greater than pic_width_in_luma_samples - Otherwise, if all of the following conditions are true, then allowBtSplit is set equal to FALSE - btSplit is equal to SPLIT_BT_HOR - cbWidth is greater than 64 -y0 + cbHeight is greater than pic_height_in_luma_samples - Otherwise, if all of the following conditions are true, then allowBtSplit is set to FALSE -x0 + cbWidth is greater than pic_width_in_luma_samples -y0 + cbHeight is greater than pic_height_in_luma_samples -cbWidth is greater than minQtSize - Otherwise, if all of the following conditions are true, then allowBtSplit is set to FALSE -btSplit is equal to SPLIT_BT_HOR -x0 + cbWidth is greater than pic_width_in_luma_samples -y0 + cbHeight is less than or equal to pic_height_in_luma_samples - Otherwise, if all of the following conditions are true, then allowBtSplit is set to FALSE: -mttDepth is greater than 0 -partldx is equal to 1 -MttsplitMode[x0][y0][mttDepth - 1] is equal to parallelTtSplit - Otherwise, if all of the following conditions are true, then allowBtSplit is set to be equal to FALSE -btSplit is equal to SPLIT_BT_VER -cbWidth is less than or equal to 64 -cbHeight is greater than 64 - Otherwise, if all of the following conditions are true, then allowBtSplit is set to be equal to FALSE -btSplit is equal to SPLIT_BT_HOR -cbWidth is greater than 64 -cbHeight is less than or equal to 64 - Otherwise, allowBtSplit is set to be equal to TRUE.
[0099] Described in Section 6.4.3 of VVC Draft 7: 6.4.3 Allowed ternary tree partitioning process The inputs to this process are: - The ternary tree partition ttSplit, - The coded block width cbWidth in the luma samples, - The coded block height cbHeight in the luma samples, - The position (x0, y0) of the top - left luma sample of the coded block under consideration relative to the top - left luma sample of the image, - The multi - type tree depth mttDepth - The maximum multi - type tree depth with offset maxMttDepth, - The maximum ternary tree size maxTtSize, - A variable treeType that specifies whether to use a single tree (SINGLE_TREE) or a dual tree to partition the coding tree nodes, and when using a dual tree, whether the current component being processed is the luma component (DUAL_TREE_LUMA) or the chroma component (DUAL_TREE_CHROMA), - A variable modeType that specifies whether intra (MODE_INTRA), IBC (MODE_IBC), and inter - coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter - coding modes (MODE_TYPE_INTER) can be used for coding units inside the coding tree nodes.
[0100] The output of this process is the variable allowTtSplit. Figure 11 Exemplary Table 3 according to some embodiments of the present disclosure is shown, which shows an exemplary specification of the variable cbSize based on ttSplit.
[0101] The variable allowTtSplit is derived as follows: - If one or more of the following conditions are true, then allowTtSplit is set to be equal to FALSE: - cbSize is less than or equal to 2 * MinTtSizeY - cbWidth is greater than Min(64, maxTtSize) - cbHeight is greater than Min(64, maxTtSize) - mttDepth is greater than or equal to maxMttDeptb - x0 + cbWidth is greater than pic_width_in_luma_samples -y0 + cbHeight is greater than pic_height_in_luma_samples -treeTpye is equal to DUAL_TREE_CHROMA, and (cbWidth / SubWidthC) * (cbHeight / SubHeightC) is less than or equal to 32 -treeType is equal to DUAL_TREE_CHROMA, and (cbWidth / SubWidthC) is equal to 8, ttSplit is equal to SPLIT_TT_VER -treeType is equal to DUAL_TREE_CHROMA, and modeType is equal to MODE_TYPE_INTRA -cbWidth * cbHeight is equal to 64, modeType is equal to MODE_TYPE_INTER -Otherwise, allowTtSplit is set to be equal to TRUE.
[0102] Figure 12 Exemplary Table 4 according to some embodiments of the present disclosure is shown, which shows an exemplary part 7.3.9.4 of the VVC Draft 7 coding tree syntax (with emphasis symbols added in italics and shading).
[0103] The derivation of the variables allowSplitQt, allowSplitBtVer, allowSplitBtHor, allowSplitTtVer, and allowSplitTtHor is as follows: - Call the allowed quadtree partitioning process specified in Article 6.4.1, with the coding block size cbSize set to be equal to cbWidth, the current multi-type tree depth mttDepth, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitQt. - The derivation of the variables minQtSize, maxBtSize, maxTtSize, and maxMttDepth is as follows: - If treeType is equal to DUAL_TREE_CHROMA, then set minQtSize, maxBtSize, maxTtSize, and maxMttdepth to be equal to MinQtSizeC, MaxBtSizeC, MaxTtSizeC, and MaxMttDepthc + depthOffset, respectively. - Otherwise, minQtSize, maxBtSize, maxTtSize, and maxMttDepth are set to be equal to MinQtSizeY, MaxBtSizeY, MaxTtSizeY, and MaxMttDepthY + depthOffset, respectively. - Call the allowed binary tree partitioning process specified in Clause 6.4.2, using the binary tree partitioning mode SPLIT_BT_VER, coding block width cbWith, coding block height cbHeight, position (x0, y0), current multi-type tree depth mttDepth, maximum multi-type tree depth with offset maxMttdepth, maximum binary tree size maxBtSize, minimum quadtree size minQtSize, current partition index partIdx, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitBtVer. - Call the allowed binary tree splitting process specified in Clause 6.4.2, using the binary tree splitting mode SPLIT_BT_HOR, coding block height cbHeight, coding block width cbWidth, position (x0, y0), current multi-type tree depth mttDepth, maximum multi-type tree depth with offset maxMttDepth, maximum binary tree size maxBtSize, minimum quadtree size minQtSize, current partition index partIdx, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitBtHor. - Call the allowed ternary tree partitioning process specified in Clause 6.4.3, using the ternary tree partitioning mode SPLIT_TT_VER, coding block width cbWidth, coding block height cbHeight, position (x0, y0), current multi-type tree depth mttDepth, maximum multi-type tree depth with offset maxMttDepth, maximum ternary tree size maxTtSize, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitTtVer. - Call the allowed ternary tree partitioning process specified in Clause 6.4.3, using the ternary tree partitioning mode SPLIT_TT_HOR, coding block height cbHeight, coding block width cbWidth, position (x0, y0), current multi-type tree depth mttDepth, maximum multi-type tree depth with offset maxMttDepth, maximum ternary tree size maxTtSize, treeTypeCurr, and modeTypeCurr as inputs, and assign the output to allowSplitTtHor.
[0104] The syntax element split_cu_flag being equal to 0 specifies not to split the coding unit. The syntax element split_cu_flag being equal to 1 specifies to split the coding unit into four coding units using the quadtree partitioning indicated by the syntax element split_qt_flag, or into two coding units using binary tree partitioning, or into three coding units using the ternary tree partitioning indicated by the syntax element mtt_split_cu_binary_flag. As indicated by the syntax element mtt_split_cu_vertical_flag, the binary tree or ternary tree partitioning can be vertical or horizontal.
[0105] When the syntax element split_cu_flag does not exist, the value of split_cu_flag is inferred as follows: - If one or more of the following conditions are true, the value of split_cu_flag is inferred to be equal to 1: - x0 + cbWidth is greater than pic_width_in_luma_samples. - y0 + cbHeight is greater than pic_height_in_luma_sampless. - Otherwise, the value of split_cu_flag is inferred to be equal to 0.
[0106] The syntax element split_qt_flag specifies whether the coding unit is split into coding units with half the horizontal and vertical sizes.
[0107] When the syntax element split_qt_flag does not exist, the following applies: - If allowSplitQt is equal to TRUE, the value of split_qt_flag is inferred to be equal to 1. - Otherwise, the value of split_qt_flag is inferred to be equal to 0.
[0108] The syntax element mtt_split_cu_vertical_flag being equal to 0 specifies that the coding unit is horizontally partitioned. The syntax element mtt_split_cu_vertical_flag being equal to 1 specifies that the coding unit is vertically partitioned.
[0109] When the syntax element mtt_split_cu_vertical_flag is absent, the following inferences are made: - If allowSplitBtHor is equal to TRUE or allowSplitTtHor is equal to TRUE, then the value of mtt_split_cu_vertical_flag is inferred to be equal to 0. - Otherwise, the value of mtt_split_cu_vertical_flag is inferred to be equal to 1.
[0110] The syntax element mtt_split_cu_binary_flag being equal to 0 specifies that the coding unit is split into three coding units using a ternary tree. The syntax element mtt_split_cu_binary_flag being equal to 1 specifies that the coding unit is split into two coding units using a binary tree partition.
[0111] When the syntax element mtt_split_cu_binary_flag is absent, the following inferences are made: - If allowSplitBtVer is equal to FALSE and allowSplitBtHor is equal to FALSE, then the value of mtt_split_cu_binary_flag is inferred to be equal to 0. - Otherwise, if allowSplitTtVer is equal to FALSE and allowSplitTtHor is equal to FALSE, then the value of mtt_split_cu_binary_flag is inferred to be equal to 1. - Otherwise, if allowSplitBtHor is equal to TRUE and allowSplitTtVer is equal to TRUE, then the value of mtt_split_cu_binary_flag is inferred to be equal to mtt_split_cu_vertical_flag. - Otherwise (allowSplitBtVer is equal to TRUE, allowSplitTtHor is equal to TRUE), the value of mtt_split_cu_binary_flag is inferred to be equal to mtt_split_cu_vertical_flag.
[0112] Figure 14Exemplary Table 5 according to some embodiments of the present disclosure is shown, which shows exemplary specifications of MttSplitMode. The variable MttSplitMode[x][y][mttDepth] is derived from the value of the syntax element mtt_split_cu_vertical_flag and the value of the syntax element mtt_split_cu_binary_flag when x = x0..x0+cbWidth-1 and y = y0...y0+cbHeight-1 are defined as in Table 4.
[0113] MttSplitMode[x][y][mttDepth] represents the horizontal binary tree, vertical binary tree, horizontal ternary tree, and vertical ternary tree partitioning of coding units within a multi-type tree. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration relative to the top-left luminance sample of the image. Figure 13 An example of the multi-type tree partitioning mode indicated by MttSplitMode according to some embodiments of the present disclosure is shown. As Figure 13 shown, the multi-type tree partitioning mode may include vertical binary tree partitioning (SPLIT_BT_VER) 1301, horizontal binary tree partitioning (SPLIT_BT_HOR) 1302, vertical ternary tree partitioning (SPLIT_TT_VER) 1303, and horizontal ternary tree partitioning (SPLIT_TT_HOR) 1304.
[0114] It should be noted that the CTU size, minimum block size, and block size limits for quadtree, binary tree, and ternary tree partitioning are signaled in the sequence parameter set or the picture header.
[0115] Figure 15 Exemplary Table 6 according to some embodiments of the present disclosure is shown, which shows an exemplary part 7.3.2.3 of the VVC draft 7 sequence parameter set RBSP syntax. Figure 16 Exemplary Table 7 according to some embodiments of the present disclosure is shown, which shows an exemplary part 7.3.2.6 of the VVC draft 7 picture header RBSP syntax.
[0116] According to some embodiments, when a part of a tree node block exceeds the bottom or right side of the image boundary, the tree node block is forced to be split until all samples of each coding block are within the image boundary. However, in some cases, for blocks located at the image boundary and containing samples that exceed the image boundary, not all tree partitioning modes are allowed. Figure 17 An exemplary block where QT, TT, or BT partitioning is not allowed at the image boundary according to some embodiments of the present disclosure is shown. For example, Figure 17CTUs 1701, 1703, and 1705 at the image boundary are not allowed to have QT, TT, or BT partitions.
[0117] As a first exemplary case where all tree partition patterns are not allowed, when both the CTU size and the minimum QT size are set to 128, all variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer are set to false.
[0118] Due to the following condition, variable allowSplitQt is set to false (italicized for emphasis): 6.4.1 Allowed Quad - tree Partition Process ... - If one or more of the following conditions are true, allowSplitQt is set to equal FALSE: - treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and chSize is less than or equal to MinQtSizeY - treeType equals DUAL_TREE_CHROMA, and cbSize / SubWidthC is less than or equal to MinQtSizeC ...
[0119] Due to the following condition, variable allowSplitBtHor (italicized for emphasis) is set to false: 6.4.2 Allowed Binary Partition Process ... - Otherwise, if all of the following conditions are true, allowBtSplit is set to FALSE - btSplit equals SPLIT_BT_HOR - cbWidth is greater than 64 - y0 + cbHeight is greater than pic_height_luma_samples - Otherwise, if all of the following conditions are true, allowBtSplit is set to FALSE - x0 + cbWidth is greater than pic_width_in_luma_samples - y0 + cbHeight is greater than pic_height_in_luma_samples - cbWidth is greater than minQtSize - Otherwise, if all of the following conditions are true, then allowBtSplit is set to FALSE - btSplit is equal to SPLIT_BT_HOR - x0 + cbWidth is greater than - y0 + cbHeight is less than or equal to pic_height_in_luma_samples ...
[0120] Due to the following conditions, the variable allowSplitBtVer is set to false (italicized for emphasis): 6.4.2 Allowed binary partitioning process ... - Otherwise, if all of the following conditions are true, then allowBiSplit is set to FALSE - btSplit is equal to SPLIT_BT_VER - y0 + cbHeight is greater than pic_height_in_luma_samples - Otherwise, if all of the following conditions are true, then allowBtSplit is set to FALSE - btSplit is equal to SPLIT_BT_VER - chHeight is greater than 64 - x0 + cbWidth is greater than pic_width_in_luma_samples
[0121] Due to the following, the variables allowSplitTtHor and allowSplilTtVer (italicized for emphasis) are set to false: 6.4.3 Allowed ternary tree partitioning process ... - If one or more of the following conditions are true, then allowTtSplit is set to equal FALSE: - cbSize is less than or equal to 2 * MinTtSizeY - cbWidth is greater than min(64, maxTtSize) - cbHeight is greater than min(64, maxTtSize) - mttDepth is greater than or equal to maxMttDepth -x0 + cbWidth is greater than pic_width_in_luma_samples -y0 + cbHeight is greater than pic_height_in_luma_samples ...
[0122] When all these variables are set to false, the CU-level partition flag may not be signaled. Infer the syntax element split_cu_flag as 1, the syntax element split_qt_flag as 0, the syntax element mtt_split_cu_vertical_flag as 1, and the syntax element mtt_split_cu_binary_flag as 0. In this case, the SPLIT_TT_VER partition block may be used, which may violate the VPDU limit.
[0123] As a second example where all tree partitioning patterns are not allowed, for blocks located at the image boundary and containing samples that extend beyond the image boundary, all tree partitioning is not allowed. When the minimum QT size is greater than the minimum CU size (syntax element log2_min_luma_codign_block_size_minus2 in the previous table) and the maximum BT / TT depth (syntax elements sps_max_mtt_hierarcyh_depth_inter_slice, sps_max_mtt_hierarchy_depth_intra_slice_luma, pic_max_mtt_hierarchy_depth_inter_slice, _pic_max_mtt_hierarchy_depth_intra_slice_luma, and pic_max_mtt_hierarchy_depth_intra_slice_chroma) is equal to 0, all variables allowSplitQt, allowSplitBtHor, allowSplitBtHor, allowSplitTtHor, and allowSplitTtVer are set to false.
[0124] Figure 18Shows exemplary blocks that do not allow QT, BT, or BT partitioning at the image boundary. First, a CTU (e.g., CTU 1801, CTU 1803, or CTU 1805) is divided into four 64x64 blocks using a quadtree. Then, each 64x64 block cannot be further divided. However, the portions of the blocks that extend beyond the right and / or bottom image boundaries are marked in gray, which is not allowed in the VVC design. For the blocks marked in gray, the variable allowSpiltQt is set to false due to the following conditions (emphasis in italics): 6.4.1 Allowed quadtree partitioning process … - If one or more of the following conditions are true, allowSplitQt is set to equal FALSE: - treeTpye equals SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY - treeTpye equals DUAL_TREE_CHROMS, and cbSize / SubWidthC is less than or equal to MinQtSizeC …
[0125] For the blocks marked in gray, the variables allowSplitBtHor and allowSplitBtVer are set to false due to the following (emphasis in italics): 6.4.2 Allowed binary splitting process …
[0126] The variable allowBtSplit is derived as follows: - If one or more of the following conditions are true, allowBtSplit is set to FALSE: - cbSize is less than or equal to MinBtSizeY - cbWidth is greater than maxBtSize - cbHeight is greater than maxBtSize Mttdepth is greater than or equal to maxMttDeptih …
[0127] For the blocks marked in gray, the variables allowSplitTtHor and allowSplitTtVer are set to false due to the following (emphasis in italics): 6.4.3 Allowed ternary splitting process …
[0128] The variable allowTtSplit is derived as follows: - If one or more of the following conditions are true, allowTtSplit is set to FALSE: - cbSize is less than or equal to 2 * MinTtSizeY - cbWidth is greater than Min(64, maxTtSize) - cbHeight is greater than Min(64, maxTtSize) - mttDepth is greater than or equal to maxMttDepth ...
[0129] As described in the first exemplary case where not all tree partitioning modes are allowed, in the current VVC draft 7, a CU may contain samples outside the image boundary, but under certain conditions, the CU cannot be partitioned any further. In some embodiments of the present disclosure, the QT partitioning conditions in VVC can be changed. In one aspect, for a block that contains samples outside the image boundary and whose width or height is equal to N (e.g., N = 128), QT partitioning is used when the minimum QT size is less than N (e.g., 128). Additionally, in some embodiments, QT partitioning can also be used when the minimum QT size is equal to N (e.g., 128). In another aspect, using QT partitioning can be more straightforward than using BT or TT partitioning. Figure 19 FIG. is a schematic diagram showing an example of using BT and TT partitioning according to some embodiments of the present disclosure. Partitioning a block may require multiple steps. Additionally, for blocks located at different positions, the partitioning may be different, which can be complex. For example, for block 1903, SPLIT_BT_HQR, SPLIT_BT_HOR, SPLIT_TT_VER, and SPLIT_TT_VER are executed in sequence.
[0130] In some embodiments, when block partitioning is not allowed, quadtree partitioning can be used, and the syntax element split_qt_flag can be inferred to be 1. The syntax element split_qt_flag specifies whether to partition the coding unit into coding units with horizontal and vertical sizes halved.
[0131] When the syntax element split_qt_fiag does not exist, the following applies (italicized for emphasis): - If all of the following conditions are met, split_qt_flag is inferred to be equal to 1: - split_cu_flag is equal to I - allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer are equal to FALSE. - Otherwise, if allowSplitQt is equal to TRUE, then the value of split_qt_flag is inferred to be 1. - Otherwise, the value of split_qt_flag is inferred to be 0.
[0132] In some embodiments, the minimum QT size constraint cannot be applied to blocks located at the image boundary. When a part of a block exceeds the bottom or right boundary of the image, the block can be divided using quadtree partitioning. The allowed quadtree partitioning process is described as follows: 6.4.1 Allowed Four-Way Split Process The inputs to this process are: - The coded block size cbSize in the luma samples, - The multi-type tree depth mttDepth, - A variable tree type that specifies whether to use a single tree (SINGLE_TREE) or a dual tree to partition the coded tree nodes, and when using a dual tree, whether the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) is currently being processed. - A variable modeType that specifies whether intra (MODE_INTRA), IBC (MODE_IBC), and inter-coding modes (MODE_TYPE_ALL) can be used, or whether only intra and IBC coding modes (MODE_TYPE_INTRA) can be used, or whether only inter-coding modes (MODE_TYPE_INTER) can be used for coding units inside the coded tree nodes. The output of this process is the variable allowSplitQt.
[0133] The derivation of the variable allowSplitQt is as follows (italic is used for emphasis): - If all of the following conditions hold, then allowSplitQt is set to be equal to true: - treeTpye is equal to SINGLE_TREE or DUAL_TREE_LUMA - cbSize is equal to 128 - MinQtSizeY is equal to 128 -x0 + cbWidth is greater than pic_width_in_luma_samples or y0 + cbHeight is greater than pic_height_in_luma_samples - Otherwise, if all of the following conditions are true, then allowSplitOt is set equal to true: - treeTpye is equal to DUAL_TREE_CHROMA - CbSize / SubWidthC is equal to 128 - MinQtSizeC is equal to 128 - x0 + cbWidth is greater than pic_width_in_luma_sample, or y0 + cbHeight is greater than pic_height_in_luma_samples - Otherwise, if one or more of the following conditions are true, then allowSplitQt is set equal to FALSE: - treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, and cbSize is less than or equal to MinQtSizeY - treeType is equal to DUAL_TREE_CHROMA, and cbSize / SubWidthC is less than or equal to MinQtSizeC - mttDepth is not equal to 0 - treeType is equal to DUAL_TREE_CHROMA, and (ebSize / SubWidthC) is less than or equal to 4 - treeType is equal to DUAL_TREE_CHROMA, modeType is equal to MODE_TYPE_INTRA - Otherwise, allowSplitQt is set equal to TRUE.
[0134] In some embodiments, bitstream conformance can be added to the syntax of the minimum QT size. It may be required that the minimum QT size be less than or equal to 64.
[0135] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size among the luma samples of the luma leaf blocks resulting from the quadtree partitioning of a CTU, and the base-2 logarithm of the minimum coding block size among the luma samples of the luma CUs in a slice where the slice_type of the SPS is equal to 2 (I). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH and related to the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. The base-2 logarithm of the minimum size among the luma samples of the luma leaf blocks resulting from the quadtree partitioning of a CTU is derived as follows (italic is used for emphasis): MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY VSize = Min(64, CtbSizeY)
[0136] In some embodiments, the bitstream conformance d requires that the value of (1 << MinQtLog2SizeIntraY) be less than or equal to VSize.
[0137] The syntax element sps_log2_diff_min_qt_min_cb_inter_slice_luma specifies the default difference between the log2 of the minimum size in the luma samples of luma leaf blocks resulting from the quadtree partitioning of a CTU and the log2 of the minimum luma coding block size in the luma CU luma samples in slices where the slice_type in the SPS is equal to 0 (B) or 1 (P). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH in the SPS. The syntax element sps_log2_diff_min_qt_min_cb_inter_slice has a value range of 0 to the value CtbLog2SizeY MinCbLog2SizeY, inclusive of the end values. The log2 of the minimum size in the luma samples of luma leaf blocks resulting from the quadtree partitioning of a CTU is derived as follows (italic is used for emphasis): MinQtLog2SizeInterY = sps_log2_diff_min_qt_min_cb_inter_slice MinCbLog2SizeY VSize = Min(64, CihSizeY)
[0138] In some embodiments, the consistency requirement of the bitstream is that the value of (1 << MinQtLog2SizeInterY) is less than or equal to VSize.
[0139] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the default difference between the log2 of the minimum size in luma samples in chroma leaf blocks resulting from the quadtree partitioning of chroma CTUs with treeType equal to DUAL_TREE_CHROMA, and the log2 of the minimum coded block size in luma samples of chroma CUs with treeType equal to DUAL_TREE_CHROMA in slices with slice_type equal to 2 (I) in the SPS. When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_chroma present in the PH of the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to 0. The log2 of the minimum size in luma samples in chroma leaf blocks resulting from the quadtree partitioning of CTUs with treeType equal to DUAL_TREE_CHROMA is derived as follows (italic emphasis): MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_chroma + MinCbLog2SizeY VSize = Min(64, CthSizeY)
[0140] In some embodiments, the bitstream conformance requirement is that the value of (1 << MinQtLog2SizeIntraC) is less than or equal to VSize.
[0141] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size in the luma samples of a luma leaf block resulting from the quadtree partitioning of a CTU, and the base-2 logarithm of the minimum coded block size in the luma samples of a luma CU in a slice where the slice_type associated with PH is equal to 2 (I). The value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When absent, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma. In some embodiments, the bitstream conformance requirement is that the value of (1 << (pic_log2_diff_min_qt_min_cb_intra_slice_luma + MinCblog2SizeY)) is less than or equal to Min(64, CtbSizeY).
[0142] The syntax element pic_log2_diff_min_qt_min_cb_inter_slice specifies the difference between the base-2 logarithm of the minimum size in the luma samples of a luma leaf block resulting from the quadtree partitioning of a CTU, and the base-2 logarithm of the minimum luma coded block size of a luma CU in a slice where the slice_type associated with PH is equal to 0 (B) or 1 (P). The value range of the syntax element pic_log2_diff_min_qt_min_cb_inter_slice is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When absent, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_inter_slice. In some embodiments, the bitstream conformance requirement is that the value of (1 << (pic_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY)) is less than or equal to Min(64, CtbSizeY).
[0143] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the base-2 logarithm of the minimum size of the luma samples of a chroma leaf block resulting from the quadtree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA, minus the base-2 logarithm of the minimum coded block size of the luma samples of a chroma CU with treeTpye equal to DUAL_TREE_CHROMA in a slice where slice_type associated with PH is equal to 2 (I). The value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When absent, the value of the syntax element pic_1og2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma. In some embodiments, the bitstream conformance requirement is that the value of (1 << (pic_1og2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY)) is less than or equal to Min(64, CtbSizeY).
[0144] In some embodiments, when block partitioning is not allowed in the second example above (where all tree partitioning modes are not allowed), quadtree partitioning may be used and the syntax element split_qt_flag is inferred to be 1.
[0145] The syntax element split_qt_flag specifies whether a coding unit is partitioned into coding units with half the horizontal size and vertical size. When the syntax element split_qt_flag is absent, the following applies (italic emphasis): - If all of the following conditions hold, split_qt_flag is inferred to be equal to 1: - split_cu_flag is equal to I - allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer are equal to FALSE. - Otherwise, if allowSplitQt is equal to TRUE, the value of split_qt_flag is inferred to be equal to 1. - Otherwise, the value of split_qt_flag is inferred to be equal to 0.
[0146] In some embodiments, in the above second exemplary case where all tree partitioning patterns are not allowed, bitstream consistency may be added to the syntax of the minimum QT size and the maximum BT / TT depth.
[0147] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size of the luma samples of the luma leaf blocks resulting from the quadtree partitioning of the CTU and the base-2 logarithm of the minimum coding block size of the luma samples of the luma CUs in a slice where the slice_type of the SPS is equal to 2 (I). When the syntax element partition_constraints_override_enabled_flag is equal to 1, the default difference may be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH of the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. The base-2 logarithm of the minimum size of the luma samples of the luma leaf blocks resulting from the quadtree partitioning of the CTU is as follows: MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY
[0148] The syntax element sps_max_mtt_hierarchy_depth_intra_slice_luma specifies the default maximum hierarchy depth of coding units generated by the multi-type tree partitioning of quadtree leaves in a slice where slice_type in the SPS is equal to 2 (I). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default maximum hierarchy depth can be overridden by the syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma present in the PH in the SPS. The value range of the syntax element sps_max_mtt_hierarchy_depth_intra_slice_luma is from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive of the end values. In some embodiments, the bitstream conformance requirement is that the value of (MmQtLog2SizeIntraY - sps_max_mtt_hierachy_depth_intra_slice_luma / 2) is less than or equal to MinCbLog2SizeY.
[0149] The syntax element sps_log2_diff_min_qt_min_cb_inter_slice specifies the default difference between the base-2 logarithm of the minimum size of luma samples in a luma leaf block generated by the quadtree partitioning of a CTU and the base-2 logarithm of the minimum luma coding block size of luma samples in a luma CU in a slice where slice_type in the SPS is equal to 0 (B) or 1 (P). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_luma present in the PH in the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_inter_slice is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. The base-2 logarithm of the minimum size of luma samples in a luma leaf block generated by the quadtree partitioning of a CTU is derived as follows: MinQtLog2SizeInterY = sps_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY
[0150] The syntax element sps_max_mtt_hierarchy_depth_inter_slice specifies the default maximum hierarchical depth of coding units generated by multi-type tree partitioning of quadtree leaves in slices where the slice_type associated with the SPS is equal to 0 (B) or 1 (P). When the syntax element partition_constraints_override_enabled_flag is equal to 1, the default maximum hierarchical depth can be overridden by the syntax element pic_max_mtt_hierarchy_depth_inter_slice present in the PH associated with the SPS. The value range of the syntax element sps_max_mtt_hierarchy_depth_inter_slice is from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive of the end values. In some embodiments, the bitstream conformance requirement that the value of (MinQtLog2SizeInterY - sps_max_mtt_hierachy_depth_inter_slice / 2) is less than or equal to MinCbLog2SizeY is imposed.
[0151] The syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the default difference in the base-2 logarithm of the minimum size in the luma samples of chroma leaf blocks generated by quadtree partitioning of chroma CTUs where treeType is equal to DUAL_TREE_CHROMA, from the base-2 logarithm of the minimum coded block size in the luma samples of chroma CUs where treeTpye is equal to DUAL_TREE_CHROMA in slices where the slice_type associated with the SPS is equal to 2 (I). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default difference can be overridden by the syntax element pic_log2_diff_min_qt_min_cb_chroma present in the PH associated with the SPS. The value range of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When not present, the inferred value of the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma is 0. The base-2 logarithm of the minimum size in the luma samples of chroma leaf blocks generated by quadtree partitioning of CTUs where treeType is equal to DUAL_TREE_CHROMA is derived as follows: MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY
[0152] The syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma is the default maximum hierarchical depth of chroma coding units generated by the multi-type tree partition of chroma quad-tree leaves with treeType equal to DUAL_TREE_CHROMA in slices where slice_type in the SPS is equal to 2 (I). When the syntax element partition_constraints_override_enabled_flag is equal to 1, this default maximum hierarchical depth can be overridden by the syntax element pic_max_mtt_hierarchy_depth_chroma present in the PH of the SPS. The value of the syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma ranges from 0 to 2 * (CtbLog2SizeY - MinCbLog2SizeY), inclusive of the end values. When it is not present, the inferred value of the syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma is 0. In some embodiments, the consistency requirement of the bitstream is that the value of (MinQtLog2SizeTntraC - sps_max_mtt_hierarchy_depth_intra_slice_chroma / 2) is less than or equal to MinCbLog2SizeY.
[0153] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma specifies the difference between the base-2 logarithm of the minimum size of the luma samples of the luma leaf blocks generated by the quadtree partition of the CTU and the base-2 logarithm of the minimum coded block size of the luma samples of the luma CUs in slices where slice_type associated with the PH is equal to 2 (I). The value range of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_luma is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When it is not present, the inferred value of the syntax element pic_log2_diff_min_qt_min_cb_luma is equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma.
[0154] The syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma specifies the maximum hierarchical depth of coding units generated by the multi-type tree partitioning of quadtree leaves in a slice where the slice_type associated with PH is equal to 2 (I). The value of the syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma is in the range of 0 to 2 * (CtbLog2SizeY - MinCbLog2SizeY), inclusive of the end values. When absent, the value of the syntax element pic_max_mtt_hierarchy_depth_intra_slice_luma is inferred to be equal to the syntax element sps_max_mtt_hierarchy_depth_intra_slice_luma. In some embodiments, the value of the bitstream consistency requirement (pic_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY - pic_max_mtt_hierarchy_depth_intra_slice_luma / 2) is less than or equal to MinCbLog2SizeY.
[0155] The syntax element pic_log2_diff_min_qt_min_cb_inter_slice specifies the difference between the base-2 logarithm of the minimum size of the luma samples of a luma leaf block generated by the quadtree partitioning of a CTU and the base-2 logarithm of the minimum luma coding block size among the luma samples of a luma CU in a slice where the slice_type associated with PH is equal to 0 (B) or 1 (P). The value range of the syntax element pic_log2_diff_min_qt_min_cb_mter_slice is from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When absent, the value of the syntax element pic_log2_diff_min_qt_min_cb_luma is inferred to be equal to the syntax element sp_log2_diff_min_qt_min_cb_inter_slice.
[0156] The syntax element pic_max_mtt_hierarchy_depth_inter_slice specifies the maximum hierarchical depth of coding units generated by multi-type tree partitioning of quadtree leaves in a slice where the slice_type associated with PH is equal to 0 (B) or 1 (P). The value range of the syntax element pic_max_mtt_hierarchy_depth_inter_slice is from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive of the end values. When absent, the value of the syntax element pic_max_mtt_hierarchy_depth_inter_slice is inferred to be equal to the syntax element sps_max_mtt_hierarchy_depth_inter_slice. In some embodiments, the value of (pic_log2_diff_min_qt_min_cb_inter_slice + MinCbLog2SizeY - pic_max_mtt_hierarchy_depth_inter_slice / 2) for bitstream conformance requirements is less than or equal to MinCbLog28izeY.
[0157] The syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the difference in the base-2 logarithm of the minimum size of the luminance samples of a chroma leaf block generated by quadtree partitioning of a chroma CTU with treeType equal to DUAL_TREE_CHROMA, and the base-2 logarithm of the smallest coding block size of the luminance samples in a chroma CU with treeType equal to DUAL_TREE_CHROMA in a slice where the slice_type associated with PH is equal to 2 (I). The value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma ranges from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of the end values. When absent, the value of the syntax element pic_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_chroma.
[0158] The syntax element pic_max_mtt_hierarchy_depth_intra_slice_chroma specifies the maximum hierarchical depth of a chroma coding unit that is generated by a multi-type tree partition of a chroma quad tree leaf having treeType equal to DUAL_TREE_CHROMA in a slice with slice_type equal to 2 (I) associated with PH. The value range of the syntax element pic_max_mtt_hierarchy_depth_intra_slice_chroma is from 0 to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive of the end values. When not present, the value of the syntax element pic_max_mtt_hierarchy_depth_intra_slice_chroma is inferred to be equal to the syntax element sps_max_mtt_hierarchy_depth_intra_slice_chroma. In some embodiments, the bitstream consistency requirement (pic_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY - pic_max_mtt_hierarchy_depth_intra_slice_chroma / 2) is less than or equal to MinCbLog2SizeY.
[0159] Figure 20 FIG. 2000 is a flowchart showing an exemplary video processing method 2000 according to some embodiments of the present disclosure. Method 2000 may be performed by an encoder (e.g., by Figure 2A process 200A or Figure 2B 200B), a decoder (e.g., by Figure 3A process 300A or Figure 3B process 300B), or one or more software or hardware components of a device (e.g., Figure 4 device 400). For example, a processor (e.g., Figure 4 processor 402) may execute method 2000. In some embodiments, method 2000 may be implemented by a computer program product included in a computer-readable medium, the computer program product including computer-executable instructions, such as program code, executed by a computer (e.g., Figure 4 device 400).
[0160] At step 2001, it may be determined whether an encoding block includes samples outside an image boundary. In some embodiments, the image boundary may be a bottom image boundary or a right image boundary. As an example of an encoding block that includes samples outside an image boundary, in Figure 17Among them, coding block 1701 exceeds the right image boundary of image 1700, coding block 1705 exceeds the bottom image boundary of image 1700, and coding block 1703 exceeds the bottom and right image boundaries of image 1700.
[0161] In step 2003, in response to a coding block being determined to include samples outside the image boundary, the coding block may be split using the QT mode. In some embodiments, in response to a coding block being determined to include samples outside the image boundary, method 2000 may determine that splitting the coding block using the BT mode and the TT mode is not allowed. For example, the variables allowSplitBtHor, allowSplitBtVer, allowSplitTtHor, and allowSplitTtVer may be determined to be equal to FALSE.
[0162] In some embodiments, in response to a coding block being determined to include samples outside the image boundary, method 2000 may determine to split the coding block using the QT mode regardless of whether a QT flag exists in the bitstream including the coding block. The QT flag indicates whether to split the coding block using the QT mode. For example, when the syntax element split_qt_flag does not exist, if the coding block includes samples outside the image boundary and the variables allowSplitQt, allowSplitBtHor, allowSplitBtVer, allowsplitTtHor, and allowSplitTtVer are equal to FALSE or allowSplitQt is equal to TRUE, it is inferred that the value of split_qt_flag is equal to 1.
[0163] In some embodiments, in response to determining that a coding block includes samples outside the image boundary, method 2000 may determine to allow splitting the coding block using the QT mode regardless of a preset constraint on the minimum block size allowing the application of the QT mode. For example, the minimum QT size constraint cannot be applied to coding blocks located at the image boundary.
[0164] In some embodiments, the preset constraint may include bitstream consistency of the coding block. The bitstream consistency may set the minimum block size allowing the application of the QT mode. For example, the minimum block size may be set to be less than or equal to 64. The bitstream consistency may also set the maximum BT depth or the maximum TT depth.
[0165] In some embodiments, method 2000 may include determining that a coding block is to be split. For example, the syntax element split_cu_flag may be used to indicate whether a coding block is to be split. When the syntax element split_cu_flag does not exist, it may be inferred to be equal to 1, which indicates that the coding block is to be split.
[0166] It should be understood that the embodiments of the present disclosure may be combined with another embodiment or some other embodiments.
[0167] The embodiments may be further described using the following clauses: 1. A video processing method, comprising: determining whether an encoding block includes samples outside an image boundary; and in response to determining that the encoding block includes samples outside the image boundary, performing quadtree splitting of the encoding block regardless of the value of a first parameter, where the first parameter indicates whether the quadtree is allowed to be used for splitting the encoding block. 2. The method according to clause 1, further comprising: determining the value of a first flag of the encoding block, where the first flag indicates whether the encoding block is split into multiple sub-blocks; and determining the values of a second parameter, a third parameter, a fourth parameter, and a fifth parameter of the encoding block, where the second parameter, the third parameter, the fourth parameter, and the fifth parameter respectively indicate whether a binary horizontal tree, a binary vertical tree, a ternary horizontal tree, and a ternary vertical tree are allowed to be used to split the encoding block. 3. The method according to clause 2, further comprising: in response to the value of the first flag being equal to 1, and the values of the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter being equal to 0, setting the value of a second flag of the encoding block to 1, where the second flag indicates whether the quadtree is used to split the encoding block. 4. The method according to clause 2, further comprising: in response to the value of the first parameter being equal to 1, setting the value of a second flag of the encoding block to 1, where the second flag indicates whether the quadtree is used to split the encoding block. 5. The method according to any one of clauses 2-4, further comprising: when determining that the encoding block includes samples outside the image boundary, setting the value of the first flag to 1. 6. A video processing apparatus, comprising: at least one memory for storing instructions, and at least one processor configured to execute the instructions to cause the apparatus to perform the following operations: determining whether an encoding block includes samples outside an image boundary; and in response to determining that the encoding block includes samples outside the image boundary, performing quadtree splitting of the encoding block regardless of the value of a first parameter, where the first parameter indicates whether the quadtree is allowed to be used for splitting the encoding block. 7. The apparatus according to clause 6, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform: Determine the value of a first flag of the coding block, the first flag indicating whether the coding block is divided into a plurality of sub-blocks; and Determine the values of a second parameter, a third parameter, a fourth parameter, and a fifth parameter of the coding block, the second parameter, the third parameter, the fourth parameter, and the fifth parameter respectively indicating whether a binary horizontal tree, a binary vertical tree, a ternary horizontal tree, and a ternary vertical tree are allowed to be used to divide the coding block. 8. The apparatus according to clause 7, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform: In response to the value of the first flag being equal to 1 and the values of the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter being equal to 0, set the value of a second flag of the coding block to 1, the second flag indicating whether a quadtree is used to divide the coding block, 9. The apparatus according to clause 7, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform: In response to the value of the first parameter being equal to 1, set the value of a second flag of the coding block to 1, the second flag indicating whether the quadtree is used to divide the coding block. 10. The apparatus according to any one of clauses 7-9, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform: When the coding block is determined to include samples outside the image boundary, set the value of the first flag to 1. 11. A non-transitory computer-readable storage medium storing an instruction set that can be executed by one or more processing devices to cause a video processing apparatus to perform the following operations: Determine whether a coding block includes samples outside the image boundary, and In response to the coding block being determined to include samples outside the image boundary, perform quadtree division of the coding block regardless of the value of a first parameter, where the first parameter indicates whether the quadtree is allowed to be used to divide the coding block. 12. The non-transitory computer-readable storage medium according to clause 11, wherein the instruction set can be executed by the one or more processing devices to cause the video processing apparatus to perform: Determine the value of a first flag of the coding block, the first flag indicating whether the coding block is divided into a plurality of sub-blocks; and Determine the values of the second parameter, third parameter, fourth parameter, and fifth parameter of the coding block, where the second parameter, third parameter, fourth parameter, and fifth parameter respectively indicate whether a binary horizontal tree, a binary vertical tree, a ternary horizontal tree, and a ternary vertical tree are allowed to be used to split the coding block. 13. The non-transitory computer-readable storage medium according to clause 12, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform: In response to the value of the first flag being equal to 1, and the values of the first parameter, second parameter, third parameter, fourth, and fifth parameters being equal to 0, set the value of the second flag of the coding block to 1, where the second flag indicates whether to use a quadtree to split the coding block. 14. The non-transitory computer-readable storage medium according to clause 12, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform: In response to the value of the first parameter being equal to 1, set the value of the second flag of the coding block to 1, where the second flag indicates whether to use the quadtree to split the coding block. 15. The non-transitory computer-readable storage medium according to any one of clauses 12-14, wherein in response to the coding block being determined to include samples outside the image boundary, using the QT mode to split the coding block includes: When the coding block is determined to include samples outside the image boundary, set the value of the first flag to 1. 16. A video processing method, comprising: Determine whether the coding block includes samples outside the image boundary; In response to the coding block being determined to include samples outside the image boundary, split the coding block using a quadtree (QT) mode. 17. The method according to clause 16, further comprising: In response to the coding block being determined to include samples outside the image boundary, determine that the binary tree (BT) mode and the ternary tree (TT) mode are not allowed to be used to split the coding block. 18. The method according to any one of clauses 16 and 17, wherein in response to the coding block being determined to include samples outside the image boundary, using the quadtree (QT) mode to split the coding block includes: Regardless of whether there is a QT flag in the bitstream including the coding block, split the coding block using the QT mode, where the QT flag indicates whether to use the QT mode to split the coding block. 19. The method according to clause 16, wherein in response to the coded block being determined to include samples outside the image boundary, using a quadtree (QT) mode to partition the coded block comprises: Partitioning the coded block using the QT mode regardless of a preset constraint on a minimum block size allowing application of the QT mode. 20. The method according to clause 19, wherein: The preset constraint includes a bitstream consistency associated with the coded block, and the bitstream consistency sets a minimum block size allowing application of the QT mode. 21. The method according to clause 20, wherein the minimum block size is set to be less than or equal to 64. 22. The method according to any one of clauses 20 and 21, wherein the bitstream consistency sets a maximum BT depth or a maximum TT depth. 23. The method according to any one of clauses 16 - 22, wherein the image boundary is a bottom image boundary or a right - hand image boundary. 24. A video processing apparatus, comprising: At least one memory for storing instructions; and At least one processor configured to execute the instructions to cause the apparatus to perform: Determine whether a coded block includes samples outside an image boundary; In response to the coded block being determined to include samples outside the image boundary, partition the coded block using a quadtree (QT) mode. 25. The apparatus according to clause 24, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform: In response to the coded block being determined to include samples outside the image boundary, determine that a binary tree (BT) mode and a ternary tree (TT) mode are not allowed to be used to partition the coded block. 26. The apparatus according to any one of clauses 24 and 25, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform; Partition the coded block using the QT mode regardless of whether a QT flag exists in a bitstream including the coded block, the QT flag indicating whether to use the QT mode to partition the coded block. 27. The apparatus according to clause 24, wherein the at least one processor is configured to execute the instructions to cause the apparatus to perform: Partition the coded block using the QT mode regardless of a preset constraint on a minimum block size allowing application of the QT mode. 28. The apparatus according to clause 27, wherein The preset constraint includes bitstream consistency associated with the coding block, and the bitstream consistency setting allows the minimum block size for applying the QT mode. 29. The apparatus according to clause 28, wherein the minimum block size is set to be less than or equal to 64. 30. The apparatus according to any one of clauses 28 and 29, wherein the bitstream consistency setting sets the maximum BT depth or the maximum TT depth. 31. The apparatus according to any one of clauses 24 to 30, wherein the image boundary is the bottom image boundary or the right image boundary. 32. A non-transitory computer-readable storage medium storing an instruction set executable by one or more processing devices to cause a video processing apparatus to perform the following operations: Determine whether a coding block includes samples outside an image boundary; In response to the coding block being determined to include samples outside the image boundary, segment the coding block using a quadtree (QT) mode. 33. The non-transitory computer-readable storage medium according to clause 32, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform: In response to the coding block being determined to include samples outside the image boundary, determine that it is not allowed to use a binary tree (BT) mode and a ternary tree (TT) mode to segment the coding block. 34. The non-transitory computer-readable storage medium according to any one of clauses 32 and 33, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform: Regardless of whether a QT flag exists in a bitstream including the coding block, segment the coding block using the QT mode, where the QT flag indicates whether to use the QT mode to segment the coding block. 35. The non-transitory computer-readable storage medium according to clause 32, wherein the instruction set is executable by the one or more processing devices to cause the video processing apparatus to perform: Regardless of the preset constraint on the minimum block size allowing the application of the QT mode, segment the coding block using the QT mode. 36. The non-transitory computer-readable storage medium according to clause 35, wherein: the preset constraint includes bitstream consistency associated with the coding block, and the bitstream consistency setting allows the minimum block size for applying the QT mode. 37. The non-horizontal computer-readable storage medium according to clause 36, wherein the minimum block size is set to be less than or equal to 64. 38. The non-transitory computer-readable storage medium according to any one of clauses 36 and 37, wherein the bitstream consistency setting is a maximum BT depth or a maximum TT depth. 39. The non-transitory computer-readable storage medium according to any one of clauses 32 to 38, wherein the image boundary is a bottom image boundary or a right image boundary.
[0168] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical medium with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM or any other flash memory, NVRAM, caches, registers, any other storage chip or cartridge storage, and their networked versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0169] It should be noted that relational terms herein, such as "first" and "second", are only used to distinguish an entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. Additionally, the words "comprise", "have", "include", and "contain" and other similar forms are equivalent in meaning and are open-ended, because one or more items after any of these words do not mean an exhaustive list of that item or items, or are limited to the listed item or items.
[0170] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, unless infeasible. For example, if it is stated that a database may include A or B, then unless otherwise explicitly stated or infeasible, the database may include A, or B, or A and B. As a second example, if it is stated that a database may include A, B, or C, then unless otherwise explicitly stated or infeasible, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0171] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When executed by a processor, the software can execute the disclosed methods. The computing units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.
[0172] In the foregoing specification, embodiments have been described with reference to numerous specific details, which may vary with the implementation. Certain modifications and changes can be made to the described embodiments. By considering the specification and practice of the invention disclosed herein, other embodiments will be apparent to those skilled in the art. The specification and embodiments are to be considered as exemplary only, and the true scope and spirit of the invention are indicated by the claims herein. The step sequences shown in the drawings are for illustrative purposes only and are not intended to be limited to any particular step sequence. Thus, those skilled in the art can understand that these steps can be executed in a different order while implementing the same method.
[0173] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are employed, they are used in a general and descriptive sense only and not for purposes of limitation.
Claims
1. A video processing method, comprising: Determining a value of a first parameter for an inter-frame strip, wherein the first parameter indicates a minimum size among a plurality of luminance samples of a luminance leaf block generated by quadtree partitioning of a coding tree unit; Determining a value of a second parameter, wherein the second parameter indicates a minimum luminance coding block size; Determining whether a coding block includes samples outside an image boundary; Determining the size of the coding block; In response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside the image boundary, dividing the coding block into coding blocks having a horizontal size and a vertical size that are half of the original; 2. The method according to claim 1, wherein, The first value is 128.
3. The method according to claim 1, wherein, In response to the first parameter being equal to the first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside the image boundary, Determining that a first flag sent in a bitstream has a second value, wherein the first flag equal to the second value indicates that each of a plurality of coding blocks is divided into coding blocks having a horizontal size and a vertical size that are half of the original; 4. The method according to claim 3, characterized in that, The second value is equal to 1.
5. The method according to claim 3, characterized in that The first flag includes split_qt_flag.
6. The method according to claim 1, further comprising: Determining values of second, third, fourth, fifth, and sixth parameters associated with the coding block, the parameters indicating whether quadtree partitioning, binary horizontal partitioning, binary vertical partitioning, ternary horizontal partitioning, and ternary vertical partitioning are allowed to partition the coding block, respectively.
7. The method according to claim 6, further comprising: In response to the first parameter being equal to the first value and the size of the coding block being equal to the first value, determining that the values of the second, third, fourth, fifth, and sixth parameters are all equal to a third value.
8. The method according to claim 6, wherein, The values of the second, third, fourth, fifth, and sixth parameters all being equal to the third value indicates that none of quadtree partitioning, binary horizontal partitioning, binary vertical partitioning, ternary horizontal partitioning, or ternary vertical partitioning is allowed to partition the coding block.
9. The method according to claim 1, further comprising: Determining that the coding block is divided into a plurality of coding blocks based on a value of a first flag sent in a bitstream.
10. The method according to claim 9, wherein, The first flag is equal to 1.
11. The method according to claim 9, wherein, The first flag includes split_qt_flag.
12. A video processing system, comprising: A memory that stores an instruction set. At least one processor configured to execute the instruction set to cause the system to perform: Determining a value of a first parameter for an inter-frame strip, wherein the first parameter indicates a minimum size among a plurality of luminance samples of a luminance leaf block generated by quadtree partitioning of a coding tree unit; Determining a value of a second parameter, wherein the second parameter indicates a minimum luminance coding block size; Determining whether a coding block includes samples outside an image boundary; Determining the size of the coding block; In response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside the image boundary, the coding block is divided into coding blocks having half the horizontal size and half the vertical size.
13. The video processing system according to claim 12, characterized in that, The first value is equal to 0.
14. A non-transitory computer-readable storage medium storing a video bitstream generated by a method executed by a video processing device, the method comprising: Determining a value of a first parameter for an inter-frame stripe, where the first parameter indicates a minimum size among a plurality of luminance samples of a luminance leaf block generated by quadtree splitting of a coding tree unit; Determining a value of a second parameter, where the second parameter indicates a minimum luminance coding block size; Determining whether a coding block includes samples outside the image boundary; Determining the size of the coding block; In response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside the image boundary, the coding block is divided into coding blocks having half the horizontal size and half the vertical size.
15. The non-transitory computer-readable storage medium according to claim 14, wherein, The first value is 128.
16. The non-transitory computer-readable storage medium according to claim 14, wherein, The bitstream includes a first value, and the method further comprises: In response to the first parameter being equal to a first value, the size of the coding block being equal to the first value, the second parameter being equal to the first value, and the coding block including samples outside the image boundary, Determining that a first flag sent in the bitstream has a second value, where the first flag equal to the second value indicates that each of a plurality of coding blocks is divided into coding blocks having half the horizontal size and half the vertical size.
17. The non-transitory computer-readable storage medium according to claim 16, wherein, The second value is equal to 1.
18. The non-transitory computer-readable storage medium according to claim 16, wherein, The first flag includes split_qt_flag.
19. The non-transitory computer-readable storage medium according to claim 14, the method further comprising: Determining values of second, third, fourth, fifth, and sixth parameters associated with the coding block, the parameters indicating whether quadtree splitting, binary horizontal splitting, binary vertical splitting, ternary horizontal splitting, and ternary vertical splitting are allowed to split the coding block, respectively.
20. The non-transitory computer-readable storage medium according to claim 14, the method further comprising: In response to the first parameter being equal to a first value and the size of the coding block being equal to the first value, determining that the values of the second, third, fourth, fifth, and sixth parameters are all equal to a third value.
Citation Information
Patent Citations
METHOD of DECODING coding units from bitstream of video data
CN108810540A
Signaling of quantization information in non-quadtree-only partitioned video coding
CN109479140A
Method and apparatus for binary-tree split mode coding
CN109792539A
Method and apparatus for processing a video signal based on adaptive block patitioning
KR1020180033030A
Methods and apparatuses for encoding and decoding adaptive quantization parameter based on quadtree structure
US20140161177A1