Block division method for video coding.

By dividing video blocks into partitions and blending prediction signals, the method enhances video coding efficiency, addressing bandwidth and storage challenges in high-definition video applications.

JP2026041814APending Publication Date: 2026-03-10ALIBABA GROUP HOLDING LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing high-definition video data, particularly in applications like video surveillance, due to high bandwidth and storage demands, with limited improvements from reducing the use of I-pictures.

Method used

A method and system for video processing that involves dividing blocks along a division edge into first and second partitions, performing inter-prediction on these partitions, and blending the generated prediction signals to enhance compression efficiency.

Benefits of technology

This approach improves video coding efficiency by reducing the bit rate and storage requirements for high-definition video, enabling more effective utilization of bandwidth and storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041814000001_ABST
    Figure 2026041814000001_ABST
Patent Text Reader

Abstract

A system and method for processing video content is provided. [Solution] The method may include dividing a plurality of blocks associated with a picture into a first partition and a second partition along a division edge, performing inter prediction on the plurality of blocks to generate a first prediction signal for the first partition and a second prediction signal for the second partition, and mixing the first prediction signal and the second prediction signal for an edge block associated with the division edge.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 887,039, filed August 15, 2019, and U.S. Provisional Patent Application No. 62 / 903,970, filed September 23, 2019, both of which are incorporated by reference in their entireties.

[0002] Technical Field FIELD OF THE DISCLOSURE

[0002] The present disclosure relates generally to video processing, and more particularly to methods and systems for motion estimation using triangulation or geometric partitioning. [Background technology]

[0003] background

[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]

[0004] Disclosure Overview

[0004] Embodiments of the present disclosure provide a method for processing video content, which may include dividing a plurality of blocks associated with a picture along a division edge into first and second partitions, performing inter prediction on the plurality of blocks to generate a first predicted signal for the first partition and a second predicted signal for the second partition, and blending the first predicted signal and the second predicted signal for an edge block associated with the division edge.

[0005]

[0005] An embodiment of the present disclosure provides a system for processing video content, including a memory that stores a set of instructions and at least one processor, wherein the at least one processor is configured to execute the set of instructions to cause the system to divide a plurality of blocks associated with a picture along a division edge into first and second partitions, perform inter prediction on the plurality of blocks to generate a first prediction signal for the first partition and a second prediction signal for the second partition, and mix the first prediction signal and the second prediction signal for an edge block associated with the division edge.

[0006]

[0006] An embodiment of the present disclosure provides a non-transitory computer-readable medium storing instructions executable by at least one processor of a computer system, wherein execution of the instructions causes the computer system to perform a method including dividing a plurality of blocks associated with a picture along a dividing edge into a first partition and a second partition, performing inter-prediction on the plurality of blocks to generate a first predicted signal for the first partition and a second predicted signal for the second partition, and mixing the first predicted signal and the second predicted signal for an edge block associated with the dividing edge.

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]

[0008] [Figure 1] 8 illustrates the structure of an exemplary video sequence consistent with embodiments of the present disclosure. [Figure 2A]

[0009] 1 shows a schematic diagram of an exemplary encoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 2B]

[0010] 1 shows a schematic diagram of another exemplary encoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3A]

[0011] 1 shows a schematic diagram of an exemplary decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3B]

[0012] 1 shows a schematic diagram of another exemplary decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 4]

[0013] 1 is a block diagram of an exemplary device for encoding or decoding video consistent with embodiments of the present disclosure. [Figure 5]

[0014] 1 illustrates an exemplary inter-prediction based on triangulation, consistent with embodiments of this disclosure. [Figure 6]

[0015] 10 shows an example table for associating merge indexes with motion vectors, consistent with embodiments of this disclosure. [Figure 7]

[0016] 1 illustrates an example chroma weight map and example luma weight samples consistent with embodiments of the present disclosure. [Figure 8]

[0017] 1 illustrates an example of a 4x4 sub-block for storing motion vectors located in a uni-predictive or bi-predictive region, consistent with embodiments of the present disclosure. [Figure 9]

[0018] 1 illustrates an example syntax structure for merge mode, consistent with embodiments of the present disclosure. [Figure 10]

[0019] 10 illustrates another example syntax structure for merge mode, consistent with embodiments of the present disclosure. [Figure 11]

[0020] 1 illustrates an exemplary geometric division consistent with embodiments of the present disclosure. [Figure 12]

[0021] 10 shows an exemplary lookup table for dis[], consistent with embodiments of the present disclosure. [Figure 13]

[0022] 10 shows an exemplary lookup table for GeoFilter[ ], consistent with embodiments of the present disclosure. [Figure 14A]

[0023] 10 shows an exemplary lookup table for angleIdx and distanceIdx consistent with embodiments of the present disclosure. [Figure 14B]

[0023] An exemplary lookup table for angleIdx and distanceIdx is shown, consistent with embodiments of the present disclosure. [Figure 14C]

[0023] An exemplary lookup table for angleIdx and distanceIdx is shown, consistent with embodiments of the present disclosure. [Figure 14D]

[0023] An exemplary lookup table for angleIdx and distanceIdx is shown, consistent with embodiments of the present disclosure. [Figure 15A1]

[0024] 10 shows an exemplary lookup table for stepDis, consistent with embodiments of the present disclosure. [Figure 15A2] 10 shows an exemplary lookup table for stepDis, consistent with an embodiment of the present disclosure. [Figure 15B1]

[0025] 10 shows another exemplary lookup table for stepDis, consistent with embodiments of the present disclosure. [Figure 15B2] 10 shows another exemplary lookup table for stepDis, consistent with embodiments of the present disclosure. [Figure 15C1]

[0026] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with embodiments of the present disclosure. [Figure 15C2]

[0026] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with an embodiment of the present disclosure, is shown. [Figure 15C3]

[0026] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with an embodiment of the present disclosure, is shown. [Figure 15C4]

[0026] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with an embodiment of the present disclosure, is shown. [Figure 15D1]

[0027] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 108, consistent with embodiments of the present disclosure. [Figure 15D2]

[0027] Figure 1 shows an exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 108, consistent with an embodiment of the present disclosure. [Figure 15D3]

[0027] Figure 1 shows an exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 108, consistent with an embodiment of the present disclosure. [Figure 15E1]

[0028] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 80, consistent with embodiments of the present disclosure. [Figure 15E2]

[0028] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 80, consistent with an embodiment of the present disclosure, is shown. [Figure 15F1]

[0029] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 64, consistent with embodiments of the present disclosure. [Figure 15F2]

[0029] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 64, consistent with an embodiment of the present disclosure, is shown. [Figure 16]

[0030] 10 illustrates an example syntax structure for a geometric partitioning mode, consistent with embodiments of the present disclosure. [Figure 17A]

[0031] 10 illustrates another example syntax structure for a geometric partitioning mode, consistent with embodiments of the present disclosure. [Figure 17B]

[0032] 10 illustrates yet another example syntax structure for a geometric partitioning mode, consistent with embodiments of the present disclosure. [Figure 18]

[0033] 10 illustrates an example sub-block transform for an inter-predicted block consistent with embodiments of this disclosure. [Figure 19]

[0034] 1 illustrates an example of a unified syntax structure consistent with embodiments of the present disclosure. [Figure 20]

[0035] 10 illustrates another example of a unified syntax structure consistent with embodiments of the present disclosure. [Figure 21]

[0036] 10 illustrates yet another example of a unified syntax structure consistent with embodiments of the present disclosure. [Figure 22A1]

[0037] 10 shows an example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 22A2]

[0037] Figure 10 shows an exemplary lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 22A3]

[0037] Figure 10 shows an exemplary lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 22A4]

[0037] Figure 10 shows an exemplary lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 22B]

[0038] 10 shows another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 23]

[0039] 10 shows an example that allows angles that divide only larger block dimensions, consistent with embodiments of the present disclosure. [Figure 24A]

[0040] 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 24B]

[0040] Figure 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 24C]

[0040] Figure 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 24D]

[0040] Figure 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 25A]

[0041] 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 25B]

[0041] Figure 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 25C]

[0041] Figure 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 25D]

[0041] Figure 10 shows yet another example lookup table for angleIdx and distanceIdx including triangulation and geometric division, consistent with embodiments of the present disclosure. [Figure 26]

[0042] 10 shows an exemplary lookup table for Dis[ ] consistent with embodiments of the present disclosure. [Figure 27]

[0043] 1 illustrates the syntax structure of an exemplary coding unit consistent with embodiments of the present disclosure. [Figure 28]

[0044] 1 illustrates examples of SBT partitioning and GEO partitioning consistent with embodiments of the present disclosure. [Figure 29]

[0045] 1 illustrates the syntax structure of another exemplary coding unit consistent with embodiments of the present disclosure. [Figure 30A]

[0046] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 80, consistent with embodiments of the present disclosure. [Figure 30B]

[0046] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 80, consistent with an embodiment of the present disclosure, is shown. [Figure 31A]

[0047] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 64, consistent with embodiments of the present disclosure. [Figure 31B]

[0047] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 64, consistent with an embodiment of the present disclosure, is shown. [Figure 32]

[0048] 1 shows an exemplary lookup table for Rho[ ] consistent with embodiments of the present disclosure. [Figure 33]

[0049] 10 shows a table of respective numbers of operations per block consistent with embodiments of the present disclosure. [Figure 34]

[0050] 10 shows an exemplary lookup table for Rhosubblk[ ] consistent with embodiments of the present disclosure. [Figure 35A]

[0051] 13 shows an exemplary mask with a 135° angle consistent with embodiments of the present disclosure. [Figure 35B]

[0052] 1 illustrates an exemplary mask with a 45° angle consistent with embodiments of the present disclosure. [Figure 36A]

[0053] 13 shows an exemplary mask with a 135° angle consistent with embodiments of the present disclosure. [Figure 36B]

[0054] 1 illustrates an exemplary mask with a 45° angle consistent with embodiments of the present disclosure. [Figure 37]

[0055] 10 illustrates example angles of triangulation modes for various block shapes consistent with embodiments of the present disclosure. [Figure 38A]

[0056] 10 shows an example lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with embodiments of the present disclosure. [Figure 38B]

[0056] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with an embodiment of the present disclosure, is shown. [Figure 38C]

[0056] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with an embodiment of the present disclosure, is shown. [Figure 38D]

[0056] An exemplary lookup table for angleIdx and distanceIdx when the total number of geometric division submodes is set to 140, consistent with an embodiment of the present disclosure, is shown. [Figure 39]

[0057] 1 is a flowchart of an exemplary method for processing video content consistent with embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0009] Detailed Description

[0058] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise indicated, like numerals in different figures represent the same or similar elements. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present invention as recited in the appended claims. Unless otherwise specified, the term "or" encompasses all possible combinations, unless impracticable. For example, if a component is stated to include A or B, that component can include A or B, or A and B, unless otherwise specified or impracticable. As a second example, if a component is stated to include A, B, or C, that component can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise specified or impracticable.

[0010]

[0059] Video coding systems are often used to compress digital video signals, for example to reduce the storage space consumed or to reduce the transmission bandwidth consumption associated with such signals. With the increasing popularity of high-definition (HD) video (e.g., having a resolution of 1920x1080 pixels) in various applications of video compression, such as online video streaming, video conferencing, or video surveillance, there is a continuing need to develop video coding tools that can increase the efficiency of compressing video data.

[0011]

[0060] For example, video surveillance applications are becoming more and more widely used in many application scenarios (e.g., security, traffic, environmental monitoring, etc.), and the number and resolution of surveillance devices are increasing rapidly. Many video surveillance application scenarios choose to provide users with HD video to capture more information, and HD video has more pixels per frame to capture such information. However, HD video bitstreams may have high bitrates that require high bandwidth for transmission and large storage space. For example, a surveillance video stream with an average resolution of 1920 x 1080 may require as much as 4 Mbps of bandwidth for real-time transmission. Furthermore, video surveillance generally involves continuous monitoring, which can pose great challenges to storage systems when storing video data. Therefore, the high bandwidth and large storage space demands of HD video have become major limitations to the large-scale deployment of HD video in video surveillance.

[0012]

[0061] A video is a set of still pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display these pictures in chronological order. Furthermore, in some applications, a video capture device can transmit captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conferencing, or live broadcasting.

[0013]

[0062] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. A module for compression is generally referred to as an "encoder," and a module for decompression is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as various suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be called a "transcoder."

[0014]

[0063] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, then such a coding process can be called "lossy." Otherwise, such a coding process can be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.

[0015]

[0064] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are the most relevant. Position changes of pixels representing an object can reflect the object's movement between the reference picture and the current picture.

[0016]

[0065] A picture that is coded without reference to another picture (i.e., such a picture is its own reference picture) is called an "I-picture." A picture that is coded using a past picture as a reference picture is called a "P-picture." A picture that is coded using both past and future pictures as reference pictures (i.e., the referencing is "bidirectional") is called a "B-picture."

[0017]

[0066] As mentioned above, video surveillance using HD video faces the challenges of high bandwidth and large storage demands. To address this challenge, the bit rate of the encoded video can be reduced. Among I-, P-, and B-pictures, I-pictures have the highest bit rate. Because the background of most surveillance video is nearly static, one way to reduce the overall bit rate of the encoded video can be to use fewer I-pictures for video encoding.

[0018]

[0067] However, because I-pictures are generally undominant in coded video, the improvement of using fewer I-pictures may be trivial. For example, in a typical video bitstream, the ratio of I-pictures, B-pictures, and P-pictures may be 1:20:9, with I-pictures accounting for less than 10% of the total bitrate. In other words, in such an example, removing all I-pictures may only result in a 10% reduction in bitrate.

[0019]

[0068] 1 illustrates the structure of an example video sequence 100 consistent with embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.

[0020]

[0069] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the examples shown in FIG.

[0021]

[0070] Typically, video codecs do not encode or decode an entire picture at once because such a task is computationally complex. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. In this disclosure, such elementary segments are referred to as basic processing units ("BPUs"). For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Basic processing units can have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail one wishes to preserve within the basic processing unit.

[0022]

[0071] A basic processing unit may be a logical unit that may include various types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit, where the luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated for each of its luma and chroma components.

[0023]

[0072] Video coding involves multiple stages of operation, examples of which are detailed in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process and therefore may be further divided into segments referred to in this disclosure as "basic processing sub-units." In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size than the basic processing units. Similar to basic processing units, basic processing sub-units are also logical units that may contain various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that such division can be performed to further levels depending on the processing needs. It should also be noted that different stages can divide the basic processing unit using different schemes.

[0024]

[0073] For example, in a mode decision stage (one example of which is detailed in FIG. 2B ), the encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large for such a decision to be made. The encoder may divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.

[0025]

[0074] In another example, in the prediction stage (one example of which is detailed in FIG. 2A), the encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.

[0026]

[0075] In another example, in the transform stage (one example of which is detailed in FIG. 2A ), the encoder can perform a transform operation on a residual elementary processing sub-unit (e.g., a CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder can further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0027]

[0076] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.

[0028]

[0077] In some implementations, to provide parallel processing and error resilience for video encoding and decoding, a picture can be divided into regions for processing, thereby allowing the encoding or decoding process for a region of a picture to not depend on information from any other region of the picture. In other words, each region of a picture can be processed independently. This allows a codec to process different regions of a picture in parallel, thus increasing coding efficiency. Furthermore, if data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that various pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.

[0029]

[0078] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0030]

[0079] FIG. 2A shows a schematic diagram of an example encoding process 200A consistent with embodiments of the present disclosure. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 of FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to structure 110 of FIG. 1, each original picture of video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of a basic processing unit for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, and the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for each original picture region of video sequence 202 (eg, regions 114-118).

[0031]

[0080] In FIG. 2A , an encoder may feed a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may feed quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to predicted BPU 208 to generate a prediction reference 224 used in prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0032]

[0081] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a predicted reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0033]

[0082] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any manner of inputting data.

[0034]

[0083] In the prediction stage 204, in the current iteration, the encoder receives the original BPU and a prediction reference 224 and can perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and the prediction reference 224 that can be used to reconstruct the original BPU as a predicted BPU 208.

[0035]

[0084] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values ​​(e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values ​​of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208, each pixel of residual BPU 210 may have a residual value. Compared to the original BPU, prediction data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant loss of quality.

[0036]

[0085] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, the basis patterns are similar to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0037]

[0086] Different transform algorithms can use different basis patterns. For example, various transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, etc. The transform in transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inversely operating the transform (referred to as an "inverse transform"). For example, to reconstruct a pixel of residual BPU 210, the inverse transform can multiply the value of the corresponding pixel in the basis pattern by the associated respective coefficient and add the products to obtain a weighted sum. In video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from these transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they can be used to reconstruct residual BPU 210 without significant loss of quality. Therefore, the residual BPU 210 is further compressed.

[0038]

[0087] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest neighbor. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").

[0039]

[0088] Quantization stage 214 may be lossy because the encoder ignores the remainder of such a division in a rounding operation. Typically, quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0040]

[0089] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may generate a video bitstream 228 using the output data of the binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0041]

[0090] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0042]

[0091] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, an encoder can perform the stages of process 200A in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be separated into multiple stages. For example, transform stage 212 and quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit one or more stages in FIG. 2A .

[0043]

[0092] 2B shows a schematic diagram of another example encoding process 200B consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0044]

[0093] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels of one or more already coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions of one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0045]

[0094] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being coded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction specified within the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.

[0046]

[0095] In another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. With respect to the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, a reference picture may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a location having the same coordinates in the reference picture as the original BPU in the current picture, and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine that region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or different shape) than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., picture 106 in FIG. 1), the encoder can find the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching region in each matching reference picture.

[0047]

[0096] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, scaling, etc. In inter prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.

[0048]

[0097] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift matching regions of a reference picture according to a motion vector, within which the encoder may predict the original BPU of the current picture. If multiple reference pictures are used (e.g., picture 106 of FIG. 1), the encoder may shift matching regions of the reference pictures according to their respective motion vectors and average pixel values ​​of the matching regions. In some embodiments, if the encoder assigns weights to pixel values ​​of the matching regions of the respective matching reference pictures, the encoder may add a weighted sum of pixel values ​​of the shifted matching regions.

[0049]

[0098] In some embodiments, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter-prediction can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which the reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0050]

[0099] Continuing with reference to the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrates of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0051]

[0100] In the reconstruction path of process 200B, if an intra-prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., a current BPU that has been coded and reconstructed in a current picture), the encoder can directly feed the prediction reference 224 to a spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If an inter-prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., a current picture in which all BPUs have been coded and reconstructed), the encoder can feed the prediction reference 224 to a loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused by inter-prediction. For example, the encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as inter-predicted reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in a binary coding stage 226.

[0052]

[0101] FIG. 3A shows a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0053]

[0102] In FIG. 3A , a decoder may feed a portion of a video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a “coded BPU”) to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the predicted reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0054]

[0103] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a predicted reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0055]

[0104] In binary decoding stage 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before feeding it to binary decoding stage 302.

[0056]

[0105] 3B shows a schematic diagram of another example decoding process 300B consistent with embodiments of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.

[0057]

[0106] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, directions of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.

[0058]

[0107] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (e.g., intra prediction) in a spatial prediction step 2042 or temporal prediction (e.g., inter prediction) in a temporal prediction step 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.

[0059]

[0108] In process 300B, the decoder may feed the predicted reference 224 to a spatial prediction stage 2042 or a temporal prediction stage 2044 for performing a prediction operation within the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may feed the prediction reference 224 to a loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as inter-prediction reference pictures for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).

[0060]

[0109] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video consistent with embodiments of the present disclosure. As shown in FIG. 4, device 400 may include a processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for encoding or decoding video. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in Figure 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0061]

[0110] Device 400 may also include memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or process the data for processing. Memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, memory 404 may include any combination of any number of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical entity.

[0062]

[0111] Bus 410, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like, may be a communication device that transfers data between components within device 400.

[0063]

[0112] For ease of explanation and to avoid ambiguity, this disclosure will collectively refer to the processor 402 and other data processing circuitry as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module or may be fully or partially combined within any other component of the device 400.

[0064]

[0113] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0065]

[0114] In some embodiments, device 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), etc.

[0066]

[0115] It should be noted that a video codec (e.g., a codec that executes process 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0067]

[0116] The present disclosure provides block partitioning methods for use in motion estimation. It is anticipated that the disclosed methods may be performed by an encoder or a decoder.

[0068]

[0117] For inter prediction, triangulation modes are supported. Triangulation modes can be applied to blocks that are 8x8 or larger and coded by triangle skip or merge modes. Triangle skip / merge modes are signaled in parallel with normal merge mode, MMVD mode, combined inter and intra prediction (CIIP) mode, or sub-block merge mode.

[0069]

[0118] When the triangular partitioning mode is used, a block can be evenly divided into two triangular partitions using diagonal or anti-diagonal partitioning (FIG. 5). Each triangular partition in a block is inter-predicted using its own motion. Only uni-prediction is allowed for each partition. In other words, each partition has one motion vector and one reference index. As with traditional bi-prediction, a uni-prediction motion constraint is applied to ensure that only two motion-compensated predictions are required for each block. The uni-prediction motion for each partition is derived directly from the merge candidate list constructed for enhanced merge prediction, and the selection of the uni-prediction motion from a given merge candidate in the list is according to the procedure described below.

[0070]

[0119] If triangulation mode is used for the current block, a flag indicating the triangulation direction (diagonal or anti-diagonal) and two merge indices (one per partition) are further signaled. After predicting each triangular partition, a blending process with adaptive weights is used to adjust the sample values ​​along the diagonal or anti-diagonal edges. This is a prediction signal for the entire block, and the transform and quantization processes can be applied to the entire block as with other prediction modes. Note that sub-block transform (SBT) mode cannot be applied to blocks coded using triangulation mode. The motion field of a block predicted using triangulation mode can be stored in 4x4 sub-blocks.

[0071]

[0120] The construction of the uni-prediction candidate list for triangulation mode is described below.

[0072]

[0121] Given a merge candidate index, a uni-predictive motion vector is derived from the merge candidate list constructed for extended merge prediction, as illustrated in Figure 6. For a candidate in the list, its LX (L0 or L1) motion vector is used as the uni-predictive motion vector for triangulation mode, where X is equal to the parity of the merge candidate index value (i.e., X = 0 or 1). In Figure 6, these motion vectors are marked with "x". If there is no corresponding LX motion vector, the L(1-X) motion vector of the same candidate in the extended merge prediction candidate list is used as the uni-predictive motion vector for triangulation mode.

[0073]

[0122] Blending along triangle section edges is described below.

[0074]

[0123] For each triangular partition, after predicting using the respective motion, we apply a blending of the two prediction signals to derive samples around the diagonal or anti-diagonal edges. The following weights are used in the blending process:

[0075]

[0124] As shown in FIG. 7, the values ​​are {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} for luma and {6 / 8, 4 / 8, 2 / 8} for chroma.

[0076]

[0125] The weights for each luma and chroma sample in a block predicted using the triangulation mode are calculated using the following formula: - Ratio

number

number

[0077]

[0126] Next, motion field storage in triangulation mode is described below.

[0078]

[0127] The motion vectors of blocks coded by triangulation mode are stored in 4x4 sub-blocks. Depending on the location of each 4x4 sub-block, a uni-predictive motion vector or a bi-predictive motion vector is stored. Mv1 and Mv2 represent the uni-predictive motion vectors of partition 1 and partition 2 in Figure 5, respectively. If a 4x4 sub-block is located in a uni-predictive region, Mv1 or Mv2 is stored for that 4x4 sub-block. Otherwise, if the 4x4 sub-block is located in a bi-predictive region, a bi-predictive motion vector is stored. The bi-predictive motion vector is derived from Mv1 and Mv2 according to the following process:

[0079]

[0128] 1. If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), simply combine Mv1 and Mv2 to form a bi-predictive motion vector.

[0080]

[0129] 2. Otherwise, if Mv1 and Mv2 are from the same list, store only the uni-predictive motion Mv2 instead of the bi-predictive motion.

[0081]

[0130] Note that if all samples in a 4x4 sub-block are weighted, the 4x4 sub-block is considered to be in a bi-predictive region. Otherwise, the 4x4 sub-block is considered to be in a uni-predictive region. Examples of bi-predictive and uni-predictive regions (which are shaded regions) are shown in Figure 8.

[0082]

[0131] To determine whether a 4x4 sub-block is located within a bi-predictive region, the following equation may be used: - Ratio

number

number

number

number

number

number

[0083]

[0132] An exemplary syntax structure for the triangulation mode is described below.

[0084]

[0133] Exemplary syntax structures for merge modes are shown in Figures 9 and 10, respectively. The CIIP flag shown is used to indicate whether a block is predicted using triangulation mode.

[0085]

[0134] A geometric division mode consistent with this disclosure is described as follows.

[0086]

[0135] The disclosed embodiments also allow for the use of a geometric partitioning mode to code video content. In the geometric partitioning mode, a block is divided into two partitions, which may be rectangular or non-rectangular, as shown in FIG. 11 . The two partitions are then inter-predicted using their unique motion vectors. The uni-predictive motion is derived using the same process described above with reference to FIG. 6 . After each geometric partition is predicted, a blending process with adaptive weights is used to adjust the sample values ​​along the partition edges, similar to the process used in the triangulation mode. This is a prediction signal for the entire block, and the transformation and quantization processes can be applied to the entire block, as with other prediction modes. Note that the SBT mode can be applied to blocks coded using the geometric partitioning mode. Finally, the motion field of a block predicted using the geometric partitioning mode can be stored in 4×4 sub-blocks. The advantage of the geometric partitioning mode is that it provides a more flexible partitioning method for motion compensation.

[0087]

[0136] The geometric partition mode applies only to blocks whose width and height are both 8 or greater and whose max(width, height) / min(width, height) ratio is 4 or less, and is coded by the geometric skip or merge mode. The geometric partition mode is signaled per block in parallel with the normal merge mode, MMVD mode, CIIP mode, sub-block merge mode, or triangulation mode. When this mode is used for the current block, a geometric partition mode index and two merge indices are also signaled, indicating which of 140 partitioning methods (32 quantized angles + 5 quantized distances) will be used to divide the current block. Note that depending on various settings, the total number of geometric partition submodes can be one of 140 (16 quantized angles + 9 quantized distances), 108 (16 quantized angles + 7 quantized distances), 80 (12 quantized angles + 7 quantized distances), and 64 (10 quantized angles + 7 quantized distances).

[0088]

[0137] Blending along geometric division edges is described as follows: After predicting each geometric partition using its intrinsic motion, a blending process is applied to the two predicted signals to derive samples around the partition edge. In some embodiments, the weight of each luma sample is calculated using the following formula: distFromLine=((x<<1)+1)×Dis[displacementX]+((y<<1)+1)×Dis[displacementY]-rho distScaled=Min((abs(distFromLine))>>4,14) sampleWeight L [x][y]=distFromLine≦0?GeoFilter[distScaled]:8-GeoFilter[distScaled] Here, (x, y) represents the position of each luma sample, and Dis[ ] and GeoFilter[ ] are two lookup tables shown in Table 12 and Tables 13A and 13B in FIGS. 12 and 13, respectively.

[0089]

[0138] The parameters displacementX, displacementY and rho are calculated as follows: displacementX=angleIdx displacementY=(disprancementX+NumAngles>>2)%NumAngles rho=distanceIdx×stepSize×scaleStep+CuW×Dis[displacementX]+CuH×Dis[displacementY] stepSize=stepDis+64 scaleStep=(wIdx≧hIdx)?(1< <hIdx):(1<<wIdx) wIdx=log2(CuW)-3 hIdx=log2(CuH)-3 whRatio=(wIdx≧hIdx)?(wIdx-hIdx):(hIdx-wIdx)

number

[0090]

[0139] In some embodiments, the parameters displacementX, displacementY and rho may be calculated as follows: displacementX=angleIdx displacementY=(disprancementX+NumAngles>>2)%NumAngles rho=distanceIdx×(stepSize< <scaleStep)+Dis[displacementX]<<wIdx +Dis[displacementY]< <hIdx stepSize=stepDis+77 scaleStep=(wIdx≧hIdx)?hIdx-3:wIdx-3 wIdx=log2(CuW) hIdx=log2(CuH) whRatio=(wIdx≧hIdx)?(wIdx-hIdx):(hIdx-wIdx)

number

[0091]

[0140] For example, in the YUV4:2:0 video format, the weights of the chroma samples are subsampled from the weight of the top-left luma sample of each 2x2 luma sub-block.

[0092]

[0141] Motion field storage in geometric partitioning mode is described below.

[0093]

[0142] The motion vectors of blocks coded by geometric partitioning mode are stored in 4x4 sub-blocks. A uni-predictive motion vector or a bi-predictive motion vector is stored for each 4x4 sub-block. The derivation process of bi-predictive motion is the same as the above process. Two methods are proposed to determine whether a uni-predictive motion vector or a bi-predictive motion vector is stored for a 4x4 sub-block.

[0094]

[0143] In the first method, for a 4x4 sub-block, the sample weight values ​​of its four corners are summed. If the sum is less than threshold2 and greater than threshold1, a bi-predictive motion vector is stored for this 4x4 sub-block. Otherwise, a uni-predictive motion vector is stored. Threshold1 and threshold2 are set to 32>>(((log2(CuW)+log2(CuH))>>1)-1) and 32-threshold1, respectively.

[0095]

[0144] In the second method, the following equation is used to determine which motion vector is stored for a 4x4 sub-block depending on the position of the 4x4 sub-block: rho subblk =3×Dis[displacementX]+3×Dis[displacementY] distFromLine subblk =((x subblk <<3)+1)×Dis[displacementX] +((y subblk <<3)+1)×Dis[displacementY]-rho+rho subblk motionMask[x subblk ][y subblk ]=abs(distFromLine subblk )<256?2:(distFromLine subblk ≦0?0:1) where (x subblk , y subblk ) represents the position of each 4x4 sub-block. The variables Dis[], displacementX, displacementY and rho are the same as above. motionMask[x subblk ][y subblk ] is equal to 2, then store a bi-predictive motion vector for this 4x4 sub-block. Otherwise, store a uni-predictive motion vector for this 4x4 sub-block.

[0096]

[0145] 16-17B show three exemplary syntax structures for the geometric partition mode, respectively.

[0097]

[0146] In some embodiments, a sub-block transform may be used, which divides the residual block into two residual sub-blocks as shown in Figure 18. Only one of the two residual sub-blocks is coded, and for the other residual sub-block, the residual is set equal to 0.

[0098]

[0147] For inter-predicted blocks with residual, a CU level flag is signaled to indicate whether sub-block transform is applied. If sub-block transform mode is used, a parameter is signaled to indicate horizontal or vertical symmetric or asymmetric division of the residual block into two sub-blocks.

[0099]

[0148] Triangulation mode and geometric partitioning mode are two partitioning methods for improving the coding efficiency of motion compensation. Triangulation can be considered as a part of geometric partitioning. However, in this implementation, the syntax structure, blending process, and motion field storage of the geometric partitioning mode are different from those of the triangulation mode. For example, the following processes are different between these two modes:

[0100]

[0149] 1. In a block coded in merge mode, two flags (including a triangulation mode flag and a geometric subdivision mode flag) are signaled. Furthermore, the triangulation mode can be applied to blocks whose width or height is equal to 4. However, the geometric subdivision mode cannot be applied to those blocks.

[0101]

[0150] 2. The formula used to calculate the weights of luma samples coded by triangulation mode is different from that of luma samples in geometric partitioning mode. Furthermore, the weights of chroma samples coded by triangulation mode are calculated separately, whereas the weights of chroma samples coded by geometric partitioning mode are subsampled from the corresponding luma sample.

[0102]

[0151] 3. The motion vectors of blocks coded by triangulation mode or geometric partitioning mode are both stored in 4x4 sub-blocks. Depending on the location of each 4x4 sub-block, either uni-predictive or bi-predictive motion vectors are stored. However, the process of selecting whether to store a uni-predictive or bi-predictive motion vector for a 4x4 sub-block is different between triangulation mode and geometric partitioning mode.

[0103]

[0152] 4. SBT mode is not allowed in triangulation mode, but can be applied to geometric division mode.

[0104]

[0153] The triangulation mode can be considered as a part of the geometric subdivision mode, so that it can unify all the processes used in the triangulation mode and the geometric subdivision mode.

[0105]

[0154] To unify the syntax of the triangulation mode and the geometric division mode, only one flag can be used to indicate whether a block is divided into two partitions. The flag is signaled if the size of the block is equal to or greater than 64 luma samples. If the flag is true, a division mode index is further signaled to indicate which division method is used to divide the block.

[0106]

[0155] In one embodiment, a flag (eg, the CIIP flag in FIGS. 19-20) is signaled if a block is not coded using sub-block merging mode, normal merging mode, and MMVD mode.

[0107]

[0156] In another embodiment, a flag (eg, the triangle / geometry flag in FIG. 21) is signaled at the beginning of the merge syntax structure.

[0108]

[0157] If the block is divided into two partitions, a partitioning mode index is further signaled to indicate which partitioning method is used.

[0109]

[0158] In one embodiment, two triangulation modes (e.g., dividing a block from the upper-left corner to the lower-right corner or from the upper-right corner to the lower-left corner) are placed at the front of the division mode list, followed by the geometric division mode. In other words, if the division mode index is equal to 0 or 1, the block is divided using the triangulation mode. Otherwise, the block is divided using the geometric division mode.

[0110]

[0159] In another embodiment, the triangulation mode is treated as one of the geometric division modes, as shown in Table 22A of Figure 22A. For example, if the division mode index is equal to 19, the block is divided from the upper left corner to the lower right corner. As another example, the division mode index 58 indicates that the block is divided from the upper right corner to the lower left corner.

[0111]

[0160] In yet another embodiment, the triangulation mode is also treated as one of the geometric division modes, as shown in Table 22B of Figure 22B. When the division mode index is equal to 10, the block is divided from the upper left corner to the lower right corner. Furthermore, a division mode index of 24 indicates that the block is divided from the upper right corner to the lower left corner.

[0112]

[0161] It should be noted that the number of division modes for a block may depend on the block size and / or the block shape.

[0113]

[0162] In one embodiment, only two partition modes are allowed for a block if the width or height of the block is equal to 4 or the ratio max(width, height) / min(width, height) is greater than 4. Otherwise, 142 partition modes are allowed.

[0114]

[0163] In another embodiment, if the size of the block is above a threshold, the number of partition modes is reduced, for example, if the size of the block is above 1024 luma samples, only 24 quantized angles and 4 quantized distances are allowed.

[0115]

[0164] In yet another embodiment, the number of division modes is reduced when the block shape is portrait or landscape. For example, if the ratio of max(width, height) / min(width, height) is greater than 2, only 24 quantized angles and 4 quantized distances are allowed. Furthermore, only angles that divide the block along the larger dimension may be allowed, as shown in Figure 23. As shown in Figure 23, for landscape blocks, the three angles shown in dashed lines are not allowed, and only the angles shown in solid lines are allowed.

[0116]

[0165] In yet another embodiment, if the block size exceeds a threshold and the block shape is portrait or landscape, the number of partition modes may be reduced. For example, if the block size exceeds 1024 luma samples and the ratio max(width, height) / min(width, height) is greater than 2, only 24 quantized angles and 4 quantized distances are allowed. This can be further combined with the restrictions shown in Figure 23.

[0117]

[0166] The look-up tables for the angle index and distance index can be changed.

[0118]

[0167] In one embodiment, a split mode index is used, which is a first order distance index instead of a first order angle index, as shown in Table 24 in FIG.

[0119]

[0168] In another embodiment, the order of the split mode index is related to the occurrence of the split methods. More frequently occurring split methods are placed at the front of the lookup table. An example is shown in Table 25 of Figure 25. Angle indices 0, 4, 8 and distance 0 = 12 are more likely to be used in splitting a block.

[0120]

[0169] As mentioned above, the triangular division mode can be applied to blocks whose size is equal to or greater than 64 luma samples. However, the geometric division mode can be applied to blocks whose width and height are both equal to or greater than 8 and whose max(width, height) / min(width, height) ratio is equal to or less than 4. This allows the triangular division mode and the geometric division mode to unify the restrictions on block size and block shape.

[0121]

[0170] In one embodiment, both triangulation mode and geometric division mode can be applied to blocks whose width and height are both 8 or greater and whose max(width, height) / min(width, height) ratio is 4 or less.

[0122]

[0171] In another embodiment, both the triangulation mode and the geometric division mode can be applied to blocks whose size is equal to or greater than 64 luma samples.

[0123]

[0172] In yet another embodiment, both triangulation and geometric division modes can be applied to blocks whose size is 64 luma samples or more and whose ratio max(width, height) / min(width, height) is 4 or less.

[0124]

[0173] Embodiments of the present disclosure also provide a method for unifying the weight calculation process of the triangulation mode and the geometric division mode.

[0125]

[0174] In one embodiment, the weight calculation process for the luma samples of the triangulation is replaced with the process used for the geometric decomposition mode (described in the blending process above) with the following two modifications:

[0126]

[0175] 1. Replace the values ​​in Dis[] with the values ​​in Table 26 in Figure 26.

[0127]

[0176] 2. wIdx = log2(CuW)-2 and hIdx = log2(CuH)-2.

[0128]

[0177] Additionally, for blocks coded by triangulation mode, if the block is divided from the upper-left corner to the lower-right corner, angleIdx and distanceIdx are set to 4 and 0, respectively. Otherwise (e.g., the block is divided from the upper-right corner to the lower-left corner), angleIdx and distanceIdx are set to 12 and 0, respectively.

[0129]

[0178] In another embodiment, for both triangulation and geometric division modes, the chroma sample weights are subsampled from the weight of the top-left luma sample of each 2x2 luma sub-block.

[0130]

[0179] In yet another embodiment, for both triangular and geometric partitioning modes, the weights for chroma samples are calculated using the same process as used for luma samples coded by geometric partitioning mode.

[0131]

[0180] This disclosure also allows for unifying the process of motion field storage used in the triangulation mode and the geometric division mode.

[0132]

[0181] In one embodiment, the motion field storage for triangulation mode is replaced by that for geometric partitioning mode. That is, the weights of the four luma samples located at the four corners of a 4x4 sub-block are summed. If the sum is less than threshold2 and greater than threshold1, a bi-predictive motion vector may be stored for this 4x4 sub-block. Otherwise, uni-predictive motion is stored. Threshold1 and threshold2 are set to 32>>(((log2(CuW)+log2(CuH))>>1)-1) and 32-threshold1, respectively.

[0133]

[0182] In another embodiment, for blocks coded using triangulation mode or geometric partitioning mode, the weight of each luma sample is checked. If the weight of a luma sample is not equal to 0 or 8, the luma sample is considered a weighted sample. If all luma samples in a 4x4 sub-block are weighted, bi-predictive motion is stored for the 4x4 sub-block. Otherwise, uni-predictive motion is stored.

[0134]

[0183] To harmonize the interaction between SBT and geometric partitioning mode with the interaction between SBT and triangulation mode, the present disclosure allows disabling SBT in geometric partitioning mode. The combination of SBT and geometric partitioning mode may create two intersecting boundaries within a block, which may cause subjective quality issues.

[0135]

[0184] In one embodiment, cu_sbt_flag is not signaled when geometric partitioning is used, as shown in Table 27 of Figure 27, where the relevant syntax is shown in italics and highlighted in gray.

[0136]

[0185] In another embodiment, some SBT partitioning modes are disabled depending on the GEO partitioning mode. If an SBT partitioning edge intersects with a GEO partitioning edge, this SBT partitioning mode is not allowed. Otherwise, the SBT partitioning mode is allowed. Figure 28 shows an example of SBT partitioning modes and GEO partitioning modes. Furthermore, angle index and distance index can be used to determine whether there is an intersection of a GEO partitioning edge with an SBT partitioning edge. In one example, if the angle index of the current block is 0 (i.e., a vertical partitioning edge), horizontal SBT partitioning cannot be applied to the current block. Furthermore, the SBT syntax can be modified as shown in Table 29 in Figure 29, where the changes are italicized and highlighted in gray.

[0137]

[0186] The geometric partitioning mode divides a block into two geometric partitions, and each geometric partition performs motion compensation using its own motion vector. The geometric partitioning mode improves the prediction accuracy of inter prediction. However, it may be complicated in the following aspects:

[0138]

[0187] First, the total number of geometric partitioning submodes is huge. Therefore, in a practical implementation, it is impossible to store all the masks for blending weights and motion field storage. The total number of bits required to store the masks for 140 submodes is as follows: - For mixed weights, (8x8 + 8x16 + 8x32 + 8x64 + 16x8 + 16x16 + 16x32 + 16x64 + 32x8 + 32x16 + 32x32 + 32x64 + 64x8 + 64x16 + 64x32 + 64x64 + 64x128 + 128x64 + 128x128) x 140 x 4 = 26,414,080 bits = 3,301,760 bytes ≒ 3.3 MB - For motion field memory, (2 x 2 + 2 x 4 + 2 x 8 + 2 x 16 + 4 x 2 + 4 x 4 + 4 x 8 + 4 x 16 + 8 x 2 + 8 x 4 + 8 x 8 + 8 x 16 + 16 x 2 + 16 x 4 + 16 x 8 + 16 x 16 + 16 x 32 + 32 x 16 + 32 x 32) x 140 x 2 = 825,440 bits = 103,180 bytes ≒ 103 kilobytes

[0139] [Table 1]

[0140]

[0188] Second, the computational complexity increases when the masks are computed on the fly instead of being stored. The formulas for computing the masks for blending weights and motion field storage are complex. More specifically, the number of multiplication (×), shift (<<), addition (+), and comparison operations is enormous. Assuming blocks of size W×H, the number of respective operations per block is as follows: - Multiplication: 5 + 2 × W × H + 2 × (W × H / 16) - Shift: 4 + 3 × W × H + 2 × (W × H / 16) - Addition: 8 + 6 × W × H + 5 × (W × H / 16) - Comparison: 4+2×W×H+2×(W×H / 16)

[0141]

[0189] The details are listed in the table below. [Table 2]

[0142]

[0190] Additionally, additional memory is required to store four pre-calculated tables: Dis[], GeoFilter[], stepDis[], and lookup tables for angleIdx and distanceIdx. The size of each table is listed below:

[0143] [Table 3]

[0144]

[0191] In the third aspect, in the current design of the geometric division mode, the combinations of 45° / 135° and distanceIdx0 are always disallowed because it is assumed that the triangulation mode in VVC supports these division choices. However, as shown in the table below, for non-square blocks, the division angles in the triangulation mode are neither 45° nor 135°. Therefore, it is meaningless to exclude these two division angles for non-square blocks.

[0145] [Table 4]

[0146]

[0192] In a fourth aspect, because some combinations of angles and distances are not supported in geometric division modes, such as horizontal division by distanceIdx0 or vertical division by distanceIdx0 (to avoid redundancy with binary tree division), a lookup table is used to derive angles and distances for each geometric division submode. If the restriction on angle and distance combinations is removed, the lookup table may not be necessary.

[0147]

[0193] In a fifth aspect, the blending process, motion field storage, and syntax structure used in the triangulation mode and the geometric division mode are not unified, which means that two types of logic are required for the two modes in both software and hardware implementations. In addition, the total number of bits required to store the mask for the triangle mode is: - For mixed weights, (4x16 + 4x32 + 4x64 + 8x8 + 8x16 + 8x32 + 8x64 + 16x4 + 16x8 + 16x16 + 16x32 + 16x64 + 32x4 + 32x8 + 32x16 + 32x32 + 32x64 + 64x4 + 64x8 + 64x16 + 64x32 + 64x64 + 64x128 + 128x64 + 128x128) x 2 x 4 = 384,512 bits = 48,064 bytes ≒ 48 kilobytes - For motion field storage, (1 x 4 + 1 x 8 + 1 x 16 + 2 x 2 + 2 x 4 + 2 x 8 + 2 x 16 + 4 x 1 + 4 x 2 + 4 x 4 + 4 x 8 + 4 x 16 + 8 x 1 + 8 x 2 + 8 x 4 + 8 x 8 + 8 x 16 + 16 x 1 + 16 x 2 + 16 x 4 + 16 x 8 + 16 x 16 + 16 x 32 + 32 x 16 + 32 x 32) x 2 x 2 = 12,016 bits = 1,502 bytes ≒ 1.5 kilobytes

[0148] [Table 5]

[0149]

[0194] To solve the above problems, several solutions are proposed.

[0150]

[0195] The first solution is directed to simplifying the geometric division mode.

[0151]

[0196] To avoid on-the-fly calculation of masks for blending weights and motion field storage, we propose to derive the mask for each block from several pre-computed masks whose size is 256 × 256 or 64 × 64. The proposed method can reduce the memory required to store the masks.

[0152]

[0197] In the proposed cropping method, a first mask set and a second mask set are predefined. L The first set of masks, g_motionMask[ ], can contain several masks, each of which has a size of 256x256, and is used to derive blending weights for each block. The second set of masks, g_motionMask[ ], can contain several masks, each of which has a size of 64x64, and is used to derive a mask for motion field storage for each block. The number of masks in the first and second sets depends on the number of geometric partitioning submodes. For blocks of various sizes, the masks for those blocks are cropped from one of the masks in the first and second sets.

[0153]

[0198] In one embodiment, the predefined masks in the first and second sets may be calculated using the formulas described with respect to Figures 14 and 15A-15F and with respect to motion field storage. The number of masks in both the first and second sets is N, where N is set to the number of angles supported in the geometric partitioning mode. The nth mask in the first and second sets with index n represents the mask for angle n, where n is in the range of 0 to N-1.

[0154]

[0199] In one example, if the number of geometric division submodes is set to 140, i.e., 16 angles and 9 distances, the variable N is set to 16. In another example, if the number of geometric division submodes is set to 108, i.e., 16 angles and 7 distances, the variable N is set to 16. In other examples, if the number of geometric division submodes is set to 80 (12 angles and 7 distances) and 64 (10 angles and 7 distances), the variable N is set to 12 and 10, respectively.

[0155]

[0200] For a block whose geometric partitioning index is set to K and whose size is W×H, the mask for the blending weights of the luma samples is derived as follows: The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table, examples of which are shown in Table 15C in FIG. 15C, Table 15D in FIG. 15D, Table 22B in FIG. 22B, Table 30 in FIG. 30 and Table 31 in FIG. - Calculate the variables offsetX and offsetY as follows:

number

[0156]

[0201] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0157]

[0202] Furthermore, the mask for the motion field storage is derived as follows. - Calculate the variables offsetXmotion and offsetYmotion as follows:

number

[0158]

[0203] The number of bits required to store the defined masks is listed below. - For mixed weights, (256 x 256) x 16 x 4 = 4,193,304 bits = 524,288 bytes ≒ 524 kilobytes - For motion field storage, (64 x 64) x 16 x 2 = 131,072 bits = 16,384 bytes ≒ 16 kilobytes

[0159] [Table 6]

[0160]

[0204] Additionally, the mask can be calculated on the fly using the following simplified formula: distFromLine=(((x+offsetX)<<1)+1)×Dis[displacementX]+(((y+offsetY)<<1)+1)×Dis[displacementY]-Rho[displacementX] distScaled=Min((abs(distFromLine)+4)>>3,26) sampleWeight L [x][y]=distFromLine≦0?GeoFilter[distScaled]:8-GeoFilter[distScaled] where (x, y) represents the position of each luma sample, Dis[] and GeoFilter[] are two lookup tables shown in Table 12 and Table 13, respectively, and Rho[] is the lookup table shown in Table 32 of Figure 32.

[0161]

[0205] The parameters displacementX and displacementY are calculated as follows: displacementX=angleIdx%16 displacementY=(disprancementX+NumAngles>>2)%NumAngles

[0162]

[0206] Here, NumAngles is set to 32. When the total number of geometric division submodes is set to 140, 108, 80, and 64, respectively, angleIdx is derived from Figure 15C, Table 15D, Table 15E, and Table 15F. angleIdx can also be derived from the lookup tables shown in Table 22B, Table 30, and Table 31. distFromLine subblk =(((x subblk +offsetXmotion)<<3)+1)×Dis[displacementX]+(((y subblk +offsetYmotion)<<3)+1)×Dis[displacementY]-Rho subblk [displacementX] motioMask[x subblk ][y subblk ]=abs(distFromLine subblk )<256?2:(distFromLine subblk ≦0?0:1)

[0163]

[0207] Assuming blocks of size W×H, the respective number of operations per block are: - Multiplication: 4 + 2 × W × H + 2 × (W × H / 16) - Shift: 9 + 3 × W × H + 2 × (W × H / 16) - Addition: 7 + 8 × W × H + 6 × (W × H / 16) - Comparison: 6+2×W×H+2×(W×H / 16)

[0164]

[0208] Details regarding the number of operations per block are listed in Table 33 of FIG.

[0165]

[0209] Dis[], GeoFilter[], Rho[], Rho subblk Memory is required to store five pre-calculated tables: [] and lookup tables for angleIdx and distanceIdx. subblk The reference table for [ ] is shown in Figure 34. The size of each table is listed below.

[0166] [Table 7]

[0167]

[0210] The computational complexity of the proposed method is similar to that of the original geometric design. More specifically, we compare it with the original geometric design. - The number of multiplication operations increases by 1 for blocks of WxH - The number of shift operations increases by 5 for blocks of W x H - The number of comparison operations increases by 2 for blocks of W x H - The number of addition operations increases by 2*W*H+(W*H / 16)-1 for a block of W*H. - Memory usage increases by 17 bits

[0168]

[0211] The formula used to calculate the mask can be further simplified as follows: The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31. If angleIdx and distanceIdx are derived using Tables 15C, 15D and 22B, then the variable N (number of masks in the first and second sets) is set to 16. On the other hand, if angleIdx and distanceIdx are derived using Tables 30 and 31, then the variable N is set to 12 and 10, respectively. - Calculate the variables offsetX and offsetY as follows:

number

[0169] [Table 8]

[0170]

[0212] In another embodiment, the predefined masks in the first and second sets may be calculated using the formulas described with respect to Figures 14 and 15A-15F and with respect to the motion field storage. The number of masks in both the first and second sets may be N reduced and N reduced = (N>>1)+1, where N is the number of angles supported within the geometric division mode. In one example, if the number of geometric division submodes is set to 140, then the variable N reduced is set to 9, i.e., 16 angles and a distance of 9. In another example, if the number of geometric division submodes is set to 108 (16 angles and a distance of 7), 80 (12 angles and a distance of 7), and 64 (10 angles and a distance of 7), the variable N reduced are set to 9, 7 and 6 respectively.

[0171]

[0213] 0~N reduced At an angle of −1, the masks are directly cropped from the masks in the first and second sets. reduced The mask at angle ∼N-1 is cropped from the masks in the first and second sets and flipped horizontally. Consistent with this embodiment, Figure 35A shows an example of a mask at an angle of 135°, and Figure 35B shows an example of a mask at an angle of 45°.

[0172]

[0214] For a block whose geometric partitioning index is set to K and whose size is W×H, the mask for the blending weights of the luma samples is derived as follows: The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31. The variables offsetX and offsetY can be calculated as follows:

number

number

[0173]

[0215] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0174]

[0216] On the other hand, the mask for the motion field storage is derived as follows.

[0175]

[0217] The variables offsetXmotion and offsetYmotion are calculated as follows:

number

[0176]

[0218] The number of bits required to store the defined masks is listed below. - For mixed weights, (256 x 256) x 9 x 4 = 2,359,296 bits = 294,912 bytes ≒ 295 kilobytes - For motion field storage, (64 x 64) x 9 x 2 = 131,072 bits = 16,384 bytes ≒ 16 kilobytes

[0177] [Table 9]

[0178]

[0219] In the third embodiment, the predefined masks in the first set and the second set can be calculated using the above formula. The number of masks in the first set and the second set is both N reduced and N reduced = (N>>1)+1, where N is the number of angles supported in the geometric division mode. reducedAt an angle of −1, the masks are directly cropped from the masks in the first and second sets. reduced The mask at angle ∼N-1 is cropped from the masks in the first and second sets and flipped vertically. Consistent with this embodiment, Figure 36 shows an example of a mask at an angle of 135°, and Figure 36B shows an example of a mask at an angle of 45°.

[0179]

[0220] For a block whose geometric partitioning index is set to K and whose size is W×H, the mask for the blending weights of the luma samples is derived as follows:

[0180]

[0221] The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31.

[0181]

[0222] The variables offsetX and offsetY can be calculated as follows:

number

number

[0182]

[0223] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0183]

[0224] On the other hand, the mask for the motion field storage is derived as follows.

[0184]

[0225] The variables offsetXmotion and offsetYmotion can be calculated as follows:

number

[0185]

[0226] The number of bits required to store the defined masks is listed below. - For mixed weights, (256 x 256) x 9 x 4 = 2,359,296 bits = 294,912 bytes ≒ 295 kilobytes - For motion field storage, (64 x 64) x 9 x 2 = 131,072 bits = 16,384 bytes ≒ 16 kilobytes

[0186] [Table 10]

[0187]

[0227] It should be noted that in the above embodiment, the methods for offset derivation and chroma weight derivation can be modified.

[0188]

[0228] The different methods for offset derivation are shown below, with the differences highlighted in italics and bold compared to the above embodiment.

[0189]

[0229] The formula for deriving the offset can be modified to ensure that the offset is not equal to 0 when distanceIdx is not 0. In one example, the variables offsetX and offsetY are calculated as follows:

number

[0190]

[0230] The offset derivation may be based on the number of distances supported in the geometric division mode. In one example, if a distance of 7 is supported, the variables offsetX and offsetY are calculated as follows (same method as shown in the first embodiment):

number

[0191]

[0231] In another example, if a distance of 9 is supported, then the variables offsetX and offsetY are calculated as follows:

number

[0192]

[0232] The offset derivation may be based on the divided angle. In one example, for angles between 135° and 225° and between 315° and 45°, the offset is added vertically. Otherwise, for angles between 45° and 135° and between 225° and 315°, the offset is added horizontally. The variables offsetX and offsetY are calculated as follows:

number

[0193]

[0233] The chroma sample weights are the first mask set g_sampleWeight L For a block whose geometric partitioning index is set to K and whose size is W × H, the mask for the blending weights of the chroma samples is derived as follows: - The size of the chroma block is W' x H'. The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31. - Calculate the variables offsetXchroma and offsetYchroma as follows:

number

[0194]

[0234] The size of the defined masks in the first set may not be 256x256. This size may depend on the maximum block size and the maximum shift offset. Assuming the maximum block size is S, the number of supported distances is N d and the shift offset for each distance is defined as offset=(D×S)>>0. Then the width and height of the defined mask can be calculated as follows: S+((((N d -1)>>1)×S)>>O)<<1

[0195]

[0235] In one example, the variables S and N d and O are set to 128, 9, and 4, respectively. The size of the defined mask is set to 192x192. In another example, the variables S, N d and O are set to 128, 7 and 3, respectively. The size of the defined mask is set to 224x224.

[0196]

[0236] In some embodiments, similar to the cropping method, a first mask set and a second mask set are predefined. L The set of masks g_motionMask[ ] may contain several masks of size 256x256, which are used to derive blending weights for each block. The second set of masks g_motionMask[ ] may contain several masks of size 64x64, which are used to derive masks for motion field storage for each block. For square blocks, a mask is cropped from one of the masks in the first set and the second set, similar to the cropping method. For non-square blocks, a mask is cropped from one of the masks in the first set and the second set, followed by the upsampling process.

[0197]

[0237] In one embodiment, the predefined masks in the first set and the second set may be calculated using the above formula. The number of masks in both the first set and the second set is N, where N is set to the number of angles supported in the geometric partitioning mode. The nth mask with index n in the first set and the second set represents the mask for angle n, where n is in the range of 0 to N-1. For a block whose geometric partitioning index is set to K and whose size is W×H, the masks for the blending weights of luma samples are derived as follows: The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31. - Set the variable minSize to min(W,H). - Set the variables ratioWH and ratioHW to log2(max(W / H,1)) and log2(max(H / W,1)), respectively. The variables offsetX and offsetY can be calculated as follows:

number

[0198]

[0238] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0199]

[0239] On the other hand, the mask for the motion field storage is derived as follows. - Set the variable minSubblk to min(W,H)>>2. The variables offsetXmotion and offsetYmotion can be calculated as follows:

number

[0200]

[0240] In another embodiment, the predefined masks in the first set and the second set may be calculated using the above formula. The number of masks in the first set and the second set may both be N reduced and N reduced = (N>>1)+1, where N is the number of angles supported in the geometric partitioning mode. For a block whose geometric partitioning index is set to K and whose size is W × H, the mask for the blending weights of the luma samples is derived as follows: The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31. - Set the variable minSize to min(W,H). - Set the variables ratioWH and ratioHW to log2(max(W / H,1)) and log2(max(H / W,1)), respectively. The variables offsetX and offsetY can be calculated as follows:

number

[0201]

[0241] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0202]

[0242] On the other hand, the mask for the motion field storage is derived as follows. - Set the variable minSubblk to min(W,H)>>2. The variables offsetXmotion and offsetYmotion can be calculated as follows:

number

[0203]

[0243] In the third embodiment, the predefined masks in the first set and the second set can be calculated using the above formula. The number of masks in the first set and the second set is both N reduced and N reduced = (N>>1)+1, where N is the number of angles supported in the geometric partitioning mode. For a block whose geometric partitioning index is set to K and whose size is W × H, the mask for the blending weights of the luma samples is derived as follows: The geometric division index K is used to obtain the variables angleIdx A and distanceIdx D from a look-up table. Examples of look-up tables are shown in Tables 15C, 15D, 22B, 30 and 31. - Set the variable minSize to min(W,H). - Set the variables ratioWH and ratioHW to log2(max(W / H,1)) and log2(max(H / W,1)), respectively. The variables offsetX and offsetY can be calculated as follows:

number

[0204]

[0244] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0205]

[0245] On the other hand, the mask for the motion field storage is derived as follows. - Set the variable minSubblk to min(W,H)>>2. The variables offsetXmotion and offsetYmotion can be calculated as follows:

number

[0206]

[0246] It should be noted that the various methods for offset derivation and chroma weight derivation described above may be applied herein.

[0207]

[0247] In the original design of the geometric partitioning mode, the combination of dividing a block across its center at 135° or 45° is always excluded. The main purpose is to remove redundant partitioning options using the triangulation mode from the geometric partitioning mode. However, for non-square blocks coded using the triangulation mode, the partitioning angle is neither 135° nor 45°. Therefore, the two angles can be adaptively excluded based on the shape of the block.

[0208]

[0248] In some embodiments of the present disclosure, the two excluded angles are changed based on the shape of the block. For square blocks, 135° and 45° are excluded, which is the same as the original geometric division design. For other block shapes, the excluded angles are listed in Table 37 of FIG. 37. For example, for a block whose size is 8x16 (i.e., a width-to-height ratio of 1:2), the angles of 112.5° and 67.5°, i.e., angleIdx10 and angleIdx6, are excluded. The lookup table for the geometric division indexes is then modified as shown in Table 38 of FIG. 38. In FIG. 38, the portion of Table 38 related to the excluded angles is highlighted in gray.

[0209]

[0249] Note that this embodiment can be combined with other embodiments of the present disclosure. For example, the blending weights and motion field storage masks using the excluded angles can be calculated using the triangulation mode method. For other angles, the crop method is used to derive the masks.

[0210]

[0250] As mentioned earlier, the blending process, motion field storage and syntax structures used in the triangulation and geometric modes are not unified. This disclosure proposes to unify all processes.

[0211]

[0251] In one embodiment, the proposed cropping method described above is used in the blending and motion field storage process for both triangulation mode and geometric division mode. In addition, the angle and distance lookup tables for each geometric division submode are eliminated. The first set of masks and the second set of masks are predefined and can be calculated using the above formulas, respectively. The number of masks in both the first set and the second set is N reduced and N reduced = (N>>1)+1, where N is the number of angles supported in the geometric division mode. D Let denote the number of distances supported within a geometric partitioning mode. Thus, the total number of geometric partitioning submodes is N × N DIn one example, N and N D Set N and N to 8 and 7, respectively. D Set N and N to 12 and 7, respectively. D are set to even and odd numbers, respectively.

[0212]

[0252] For a block whose geometric partitioning index is set to K and whose size is W×H, the mask for the blending weights of the luma samples is derived as follows: - Variable N halfD N D >>Set to 1. - Set the variables angleIdx A and distanceIdx D to K%N and K / N respectively. - Calculate the variables offsetX and offsetY as follows:

number

number

[0213]

[0253] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0214]

[0254] On the other hand, the mask for the motion field storage is derived as follows. - Calculate the variables offsetXmotion and offsetYmotion as follows:

number

[0215]

[0255] The number of bits required to store the defined mask is (256 x 256) x ((N>>1)+1) x 4 + (64 x 64) x ((N>>1)+1) x 2.

[0216] [Table 11]

[0217]

[0256] In another embodiment, the proposed cropping method described above is used in the blending and motion field storage process for both triangulation mode and geometric division mode. In addition, the angle and distance lookup tables for each geometric division submode are eliminated. The first set of masks and the second set of masks are predefined and can be calculated using the above formulas, respectively. The number of masks in both the first set and the second set is N reduced and N reduced = (N>>1)+1, where N is the number of angles supported in the geometric division mode. D Let denote the number of distances supported within a geometric partitioning mode. Thus, the total number of geometric partitioning submodes is N × N D is.

[0218]

[0257] For a block whose geometric partitioning index is set to K and whose size is W×H, the mask for the blending weights of the luma samples is derived as follows: - Variable N halfD N D >>Set to 1. - Set the variables angleIdx A and distanceIdx D to K%N and K / N respectively. - Calculate the variables offsetX and offsetY as follows:

number

number

[0219]

[0258] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0220]

[0259] On the other hand, the mask for the motion field storage is derived as follows. - Calculate the variables offsetXmotion and offsetYmotion as follows:

number

[0221]

[0260] In the third embodiment, the proposed upsampling method described above is used in the blending and motion field storage process for both the triangulation mode and the geometric division mode. In addition, the angle and distance lookup tables for each geometric division submode are eliminated. The first set of masks and the second set of masks are predefined and can be calculated using the above formulas, respectively. The number of masks in both the first set and the second set is N reduced and N reduced = (N>>1)+1, where N is the number of angles supported in the geometric division mode. D Let denote the number of distances supported within a geometric partitioning mode. Thus, the total number of geometric partitioning submodes is N × N D For a block whose geometric partitioning index is set to K and whose size is W×H, the mask for the blending weights of luma samples is derived as follows: - Variable N halfD N D >>Set to 1. - Set the variables angleIdx A and distanceIdx D to K%N and K / N respectively. - Set the variable minSize to min(W,H). - Set the variables ratioWH and ratioHW to log2(max(W / H,1)) and log2(max(H / W,1)), respectively. - Calculate the variable offset as follows: offset=((256-minSize)>>1)+D>N halfD ?((DN halfD )×minSize)>>3:-((D×minSize)>>3)

number

[0222]

[0261] The blending weights for the chroma samples are subsampled from the weights of the luma samples, i.e., the weight of the top-left luma sample of each corresponding 2x2 luma sub-block is used as the weight of the chroma sample in YUV4:2:0 video format.

[0223]

[0262] On the other hand, the mask for the motion field storage is derived as follows. - Set the variable minSubblk to min(W,H)>>2. - Calculate the variable offsetmotion as follows: offsetmotion=((64-minSubblk)>>1)+D>N halfD ?((DN halfD )×minSize)>>5:-((D×minSize)>>5)

number

[0224]

[0263] FIG. 39 is a flowchart of an example method 3900 for processing video content according to some embodiments of the present disclosure. In some embodiments, method 3900 may be performed by a codec (e.g., an encoder using encoding process 200A or 200B of FIGS. 2A-2B or a decoder using decoding process 300A or 300B of FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or converting a video sequence into another code. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence that may be captured by a surveillance device (e.g., video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence may include multiple pictures. The device may perform method 3900 at the picture level. For example, the device may process pictures one at a time in the method 3900. In another example, the device may process multiple pictures at a time in the method 3900. The method 3900 may include the following steps.

[0225]

[0264] In step 3902, the plurality of blocks may be divided into first and second partitions along dividing edges.

[0226]

[0265] The multiple blocks are sub-blocks of a first block associated with the picture. It will be understood that the picture may be associated with multiple blocks (including the first block), and each of the blocks may be divided into multiple sub-blocks. The first block may be associated with a chroma block and a luma block. Thus, each of the multiple blocks (e.g., sub-blocks) may be associated with a chroma sub-block and a luma sub-block. A partition mode for the multiple blocks (e.g., sub-blocks) may be determined, and the multiple blocks (e.g., sub-blocks) may be partitioned based on the partition mode. The partitioning may result in improved inter-prediction on the first block. Exemplary partition modes may include a triangular partition mode or a geometric partition mode.

[0227]

[0266] As discussed above, the partition mode may be determined according to at least one indication signal. For example, with reference to FIG. 19, a first indication signal (e.g., a sub-block merge flag), a second indication signal (e.g., a normal merge / MMVD flag), and a third indication signal (e.g., a CIIP flag) are provided to determine the partition mode of a first block. As shown in FIG. 19, it may be determined whether the first block is coded using the sub-block merge mode according to the first indication signal (e.g., the sub-block merge flag). In response to determining that the first block is not coded using the sub-block merge mode, it may be determined whether the first block is coded using one of the normal mode or MMVD (merge mode with motion vector differences) according to the second indication signal (e.g., the normal merge / MMVD flag). In response to determining that the first block is not coded using the normal mode or MMVD, it may be determined whether the first block is coded using CIIP using the CIIP flag.

[0228]

[0267] In some embodiments, before determining whether the first block is to be coded using CIIP in accordance with the third indication signal, step 3902 may further include determining whether a size of the first block satisfies a given condition, and generating the third indication signal in response to determining that the size of the first block satisfies the given condition, or determining that the first block is to be coded using CIIP mode in response to determining that the size of the first block does not satisfy the given condition. The given conditions may include that the width and height of the first block are both equal to or greater than 8, and that the ratio of the larger value between the width and height to the smaller value between the width and height is equal to or less than 4.

[0229]

[0268] Then, in response to determining that the first block is not coded using the CIIP mode, it may be determined that the partitioning mode of the first block is one of a triangulation mode or a geometric partitioning mode.

[0230]

[0269] If the division mode of the first block is determined to be one of the triangular division mode and the geometric division mode, the target division scheme can be further determined according to the division mode index, the angle index, or the distance index, and then the division edge corresponding to the target division scheme can be determined.

[0231]

[0270] Generally, a division mode (triangular division mode or geometric division mode) may be associated with multiple division schemes, and a division mode index may indicate the number of division schemes among the multiple division schemes. The division mode index may be associated with an angle index and a distance index. The angle index may indicate the angle of a division edge of a given division scheme corresponding to the division mode index, and the distance index may indicate the distance between the division edge and the center of the first block.

[0232]

[0271] In some embodiments, a lookup table (e.g., Table 22A of FIG. 22A or Table 22B of FIG. 22B) may include multiple division mode indexes, multiple angle indexes, and multiple distance indexes associated with multiple division schemes. Given a division mode index (e.g., K), the angle index and distance index associated with the division scheme can be determined. Thus, the target division scheme can be determined in the lookup table according to the division mode index.

[0233]

[0272] Among the multiple division schemes, the lookup table can include a first division scheme associated with a first division mode index and a second division scheme associated with a second division mode index, where the first division scheme and the second division scheme are for a triangular division mode. For example, with respect to Table 22B in FIG. 22B, a division mode index of "10" is associated with dividing the block from the upper left corner to the lower right corner. A division mode index of "24" is associated with dividing the block from the upper right corner to the lower left corner.

[0234]

[0273] In step 3904, inter prediction may be performed on the plurality of blocks to generate a first prediction signal for the first partition and a second prediction signal for the second partition. A first motion vector may be applied to the first partition to generate the first prediction signal (e.g., the motion vector for the first partition), and a second motion vector may be applied to the second partition to generate the second prediction signal (e.g., the motion vector for the second partition). The first prediction signal and the second prediction signal may be stored in 4x4 sub-blocks.

[0235]

[0274] In 3906, the first prediction signal and the second prediction signal may be mixed for an edge block associated with a split edge. Because the first block is split before applying inter prediction to the first block, partitions of the first block may be mixed to complete coding of the block. In some embodiments, edge blocks associated with a split edge may be mixed. It will be understood that the prediction signal for each partition block is the same. To mix the first prediction signal and the second prediction signal for an edge block associated with a split edge, a weight for each edge block may be determined.

[0236]

[0275] In some embodiments, a set of masks may be generated to determine the blending weights. For example, a first set of masks (e.g., g_sampleWeight L [ ]) can include several masks, each of which has a size of 256x256, and the masks are used to derive the blending weights for each block. The number of mask sets (i.e., N) can be set to the number of angles supported in the geometric partitioning mode. For example, with respect to Table 22B in FIG. 22B, the number of angles supported in the geometric partitioning mode is 64, and therefore the number of masks is 16. A first offset (e.g., offsetX) and a second offset (e.g., offsetY) can be determined based on the size of the block, the angle value of the target partitioning scheme, and the distance value of the target partitioning scheme. For example, the following equations can be used to determine offsetX and offsetY for a block having a size of WxH:

number

[0237]

[0276] The angle index A is used to determine whether the division is horizontal or vertical. In the case where the angle index A is derived from Table 22B, which is a lookup table in FIG. 22B, if the angle index A is equal to 8 or 24, the block is a horizontal division. On the other hand, if the angle index A is equal to 0 or 16, the block is a vertical division. Therefore, the condition of horizontal division or (non-vertical division and H≧W) is equivalent to A%16==8 or (A%16!=0 and H≧W).

[0238]

[0277] As another example, the following equations can be used to determine offsetX and offsetY for a block having size W×H:

number

[0239]

[0278] A set of masks (e.g., g_sampleWeight L Based on [ ]), a first offset and a second offset (offsetX and offsetY) can be used to generate and determine a plurality of blending weights for the edge sub-blocks.

[0240]

[0279] In some embodiments, the first offset (e.g., offsetX) and the second offset (e.g., offsetY) can be determined without using a set of masks, and the blending weights can be determined on the fly. For example, the following formula can be used to calculate the weights of each of the edge blocks:

number

[0241]

[0280] The weight of the luma sample at position (x,y) (e.g., sampleWeight L [x][y]) can be calculated as follows: weightIdx=(((x+offsetX)<<1)+1)*disLut[displacementX]+(((y+offsetY)<<1)+1))*disLut[displacementY] partFlip=(angleIdx>=13&&angleIdx<=27)?0:1 weightIdxL=partFlip?32+weightIdx:32-weightIdx sampleWeight L [x][y]=Clip3(0,8,(weightIdxL+4)>>3)

[0242]

[0281] Because the first block includes a chroma block and a luma block, the multiple blending weights may include multiple luma weights for edge blocks and multiple chroma weights for edge blocks. A chroma weight among the multiple chroma weights is determined based on the luma weight of the upper left corner of the 2x2 sub-block corresponding to the chroma weight. For example, with reference to Figure 7, the luma weight of the upper left corner of the 2x2 luma sub-block may be used as the chroma weight for the chroma block.

[0243]

[0282] Therefore, mixing the first prediction signal and the second prediction signal for an edge block associated with a split edge may further include mixing the first prediction signal and the second prediction signal to determine a luma value of the edge block according to a plurality of luma weights for the edge block, and mixing the first prediction signal and the second prediction signal to determine a chroma value of the edge block according to a plurality of chroma weights for the edge block.

[0244]

[0283] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by an apparatus (such as the disclosed encoders and decoders) to perform the above-described methods. Common non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, flash EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. An apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0245]

[0284] The embodiments can be further described using the following clauses. 1. A method for processing video content, comprising: Dividing a plurality of blocks associated with a picture into a first partition and a second partition along a dividing edge; performing inter prediction on a plurality of blocks to generate a first prediction signal for a first partition and a second prediction signal for a second partition; blending the first predicted signal and the second predicted signal for an edge block associated with a split edge; A method comprising: 2. Splitting multiple blocks along a dividing edge is determining a partition mode for a plurality of blocks; Dividing the plurality of blocks based on a division mode; 2. The method of clause 1, further comprising: 3. The plurality of blocks are sub-blocks of the first block, and determining a partition mode for the plurality of blocks includes: determining whether the first block is coded using a sub-block merging mode according to a first indication signal; In response to determining that the first block is not coded using the sub-block merge mode, determining whether the first block is coded using one of a normal mode or a merge mode with motion vector differences (MMVD) according to a second indication signal; In response to determining that the first block is not coded using one of the normal mode or the MMVD mode, determining whether the first block is coded using a combined inter- and intra-prediction (CIIP) mode according to a third indication signal; and determining, in response to determining that the first block is not coded using the CIIP mode, that a partitioning mode of the plurality of blocks is one of a triangulation mode or a geometric partitioning mode; 3. The method of clause 2, further comprising: 4. Splitting multiple blocks along a dividing edge is determining a target division manner according to a division mode index, an angle index, or a distance index; determining split edges corresponding to a target splitting scheme; 4. The method of clause 2 or 3, further comprising: 5. Generating a set of masks; determining a first offset and a second offset based on a size of the first block, a target division scheme angle value, and a target division scheme distance value; generating a plurality of blending weights based on the set of masks using the first offset and the second offset; 5. The method of clause 4, further comprising: 6. The method of clause 5, wherein the number of mask sets is determined based on the number of angles supported in the geometric division mode, the angle values ​​of the target division scheme are determined based on the division mode index and the number of angles supported in the geometric division mode, and the distance values ​​of the target division scheme are determined based on the division mode index and the number of angles supported in the geometric division mode. 7. Determining the target division method according to the division mode index, angle index or distance index; determining a target division scheme according to a lookup table, the lookup table including a plurality of division mode indexes, a plurality of angle indexes, and a plurality of distance indexes associated with a plurality of division schemes; 7. The method of clause 5 or 6, further comprising: 8. The method described in clause 7, wherein the number of mask sets is determined based on the number of angle indexes in the lookup table, the angle values ​​of the target division scheme are determined based on the angle indexes in the lookup table corresponding to the target division scheme, and the distance values ​​of the target division scheme are determined based on the distance indexes in the lookup table corresponding to the target division scheme. 9. The method of clause 7 or 8, wherein the lookup table includes, among a plurality of division schemes, a first division scheme associated with a first division mode index and a second division scheme associated with a second division mode index, and the first division scheme and the second division scheme correspond to triangulation modes. 10. The method of clause 9, wherein the first division mode index is equal to 10 and the first division scheme associated with the first division mode index corresponds to dividing the block from the upper left corner to the lower right corner of the block, and the second division mode index is equal to 24 and the second division scheme associated with the second division mode index corresponds to dividing the block from the upper right corner to the lower left corner of the block. 11. The plurality of mixing weights includes a plurality of luma weights for the edge block and a plurality of chroma weights for the edge block, and mixing the first predicted signal and the second predicted signal for the edge block associated with the split edge includes: determining a luma value of the edge block according to a plurality of luma weights for the edge sub-block; determining a chroma value of the edge sub-block according to a plurality of chroma weights for the edge sub-block; 11. The method of any one of clauses 5 to 10, further comprising: 12. The method of clause 11, wherein, among the plurality of chroma weights, a chroma weight is determined based on a luma weight for a top-left corner of a 2x2 block corresponding to the chroma weight. 13. Determining a first offset and a second offset based on the size of the first block, the angle value of the target division scheme, and the distance value of the target division scheme; generating a plurality of blending weights for the luma samples in the first block using the first offset and the second offset; 5. The method of clause 4, further comprising: 14. The first offset and the second offset are determined by the following formula: First Offset

number

number

[0246]

[0285] It should be noted that relational terms such as "first" and "second" herein are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, terms such as "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and are open-ended in that the items following any one of these terms are not intended to be an exhaustive list of such items or to be limited only to the items they list.

[0247]

[0286] As used herein, unless otherwise specified, the word "or" includes all possible combinations unless impracticable. For example, if a database is stated to include A or B, the database can include A or B, or A and B, unless otherwise specified or impracticable. As a second example, if a database is stated to include A, B, or C, the database can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise specified or impracticable.

[0248]

[0287] It will be understood that the above-described embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0249]

[0288] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications to the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims. The order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. As such, one skilled in the art will recognize that steps can be performed in different orders while implementing the same method.

[0250]

[0289] Although illustrative embodiments have been disclosed in the drawings and herein, many variations and modifications to those embodiments may be made. Accordingly, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. 1. A method for processing video content, comprising: Dividing a plurality of blocks associated with a picture into a first partition and a second partition along a dividing edge; performing inter prediction on the plurality of blocks to generate a first prediction signal for the first partition and a second prediction signal for the second partition; blending the first prediction signal and the second prediction signal for an edge block associated with the split edge; A method comprising:

2. Dividing the plurality of blocks along the dividing edges includes: determining a division mode for the plurality of blocks; Dividing the plurality of blocks based on the division mode; The method of claim 1 further comprising:

3. the plurality of blocks are sub-blocks of a first block, and determining the partition mode for the plurality of blocks includes: determining whether the first block is coded using a sub-block merging mode according to a first indication signal; in response to determining that the first block is not coded using the sub-block merge mode, determining whether the first block is coded using one of a normal mode or a merge mode with motion vector differences (MMVD) according to a second indication signal; In response to determining that the first block is not coded using one of the normal mode or the MMVD, determining whether the first block is coded using a combined inter- and intra-prediction (CIIP) mode according to a third indication signal; and determining, in response to determining that the first block is not coded using the CIIP mode, that the partitioning mode of the plurality of blocks is one of a triangulation mode or a geometric partitioning mode; The method of claim 2 further comprising:

4. Dividing the plurality of blocks along the dividing edges includes: determining a target division manner according to a division mode index, an angle index, or a distance index; determining the splitting edges corresponding to the target splitting scheme; The method of claim 2 further comprising:

5. generating a set of masks; determining a first offset and a second offset based on a size of the first block, an angle value of the target division scheme, and a distance value of the target division scheme; generating a plurality of blending weights based on the set of masks using the first offset and the second offset; The method of claim 4 further comprising:

6. 6. The method of claim 5, wherein the number of mask sets is determined based on a number of angles supported in a geometric partitioning mode, the angle values ​​of the target partitioning scheme are determined based on the partitioning mode index and the number of angles supported in the geometric partitioning mode, and the distance values ​​of the target partitioning scheme are determined based on the partitioning mode index and the number of angles supported in the geometric partitioning mode.

7. Determining the target division manner according to the division mode index, the angle index, or the distance index includes: determining the target division scheme according to a lookup table, the lookup table including a plurality of division mode indexes, a plurality of angle indexes, and a plurality of distance indexes associated with a plurality of division schemes; The method of claim 5 further comprising:

8. 8. The method of claim 7, wherein the number of mask sets is determined based on a number of angle indices in the lookup table, the angle values ​​of the target division scheme are determined based on angle indices in the lookup table corresponding to the target division scheme, and the distance values ​​of the target division scheme are determined based on distance indices in the lookup table corresponding to the target division scheme.

9. 8. The method of claim 7, wherein the lookup table includes a first decomposition scheme associated with a first decomposition mode index and a second decomposition scheme associated with a second decomposition mode index among the plurality of decomposition schemes, the first decomposition scheme and the second decomposition scheme corresponding to triangulation modes.

10. 10. The method of claim 9, wherein the first partition mode index is equal to 10 and the first partitioning scheme associated with the first partition mode index corresponds to dividing the block from the upper left corner to the lower right corner of the block, and the second partition mode index is equal to 24 and the second partitioning scheme associated with the second partition mode index corresponds to dividing the block from the upper right corner to the lower left corner of the block.

11. the plurality of mixing weights include a plurality of luma weights for the edge block and a plurality of chroma weights for the edge block, and mixing the first prediction signal and the second prediction signal for the edge block associated with the split edge includes: determining a luma value of the edge block according to the plurality of luma weights for the edge sub-block; determining a chroma value for the edge sub-block according to the plurality of chroma weights for the edge sub-block; The method of claim 5 further comprising:

12. The method of claim 11 , wherein a chroma weight of the plurality of chroma weights is determined based on a luma weight for a top-left corner of a 2×2 block corresponding to the chroma weight.

13. determining a first offset and a second offset based on a size of the first block, an angle value of the target division scheme, and a distance value of the target division scheme; generating a plurality of blending weights for luma samples in the first block using the first offset and the second offset; and The method of claim 4 further comprising:

14. The first offset and the second offset are determined by the following formula: The first offset [Equation 1] The second offset [Equation 2] 14. The method of claim 13, wherein "W" represents a width of the first block, "H" represents a height of the first block, "A" represents the angle value of the target division scheme, and "D" represents the distance value of the target division scheme.

15. before determining whether the first block is coded using a combined inter- and intra-prediction (CIIP) mode according to the third indication signal; determining whether the size of the first block satisfies a given condition; generating the third indication signal in response to the determining that the size of the first block satisfies the given condition; or determining that the block is to be coded using the CIIP mode in response to determining that the size of the block does not satisfy the given condition; and The method of claim 3 further comprising:

16. The given condition is: The width and height of the first block are each 8 or more; the ratio of the larger value between the width and the height to the smaller value between the width and the height is 4 or less; 16. The method of claim 15, comprising:

17. 1. A system for processing video content, comprising: a memory for storing a set of instructions; at least one processor; The at least one processor may include: Dividing a plurality of blocks associated with a picture into a first partition and a second partition along a dividing edge; performing inter prediction on the plurality of blocks to generate a first prediction signal for the first partition and a second prediction signal for the second partition; blending the first prediction signal and the second prediction signal for an edge block associated with the split edge; a system configured to execute the set of instructions to cause

18. In dividing the plurality of blocks along the dividing edges, the at least one processor may cause the system to: determining a division mode for the plurality of blocks; Dividing the plurality of blocks based on the division mode; 20. The system of claim 17, configured to execute the set of instructions to further:

19. The plurality of blocks are sub-blocks of a first block, and in determining the partitioning mode for the plurality of blocks, the at least one processor may cause the system to: determining whether the first block is coded using a sub-block merging mode according to a first indication signal; in response to determining that the first block is not coded using the sub-block merge mode, determining whether the first block is coded using one of a normal mode or a merge mode with motion vector differences (MMVD) according to a second indication signal; In response to determining that the first block is not coded using one of the normal mode or the MMVD, determining whether the first block is coded using a combined inter- and intra-prediction (CIIP) mode according to a third indication signal; and determining, in response to determining that the first block is not coded using the CIIP mode, that the partitioning mode of the plurality of blocks is one of a triangulation mode or a geometric partitioning mode; 20. The system of claim 18, configured to execute the set of instructions to further:

20. A non-transitory computer-readable medium storing instructions executable by at least one processor of a computer system, execution of said instructions comprising: Dividing a plurality of blocks associated with a picture into a first partition and a second partition along a dividing edge; performing inter prediction on the plurality of blocks to generate a first prediction signal for the first partition and a second prediction signal for the second partition; blending the first prediction signal and the second prediction signal for an edge block associated with the split edge; A non-transitory computer-readable medium that causes the computer system to perform a method including: