Geometric segmentation mode
By introducing the Geometric Partitioning (GEO) mode, which optimizes the division of video blocks using angle and distance offsets, the problem of insufficient encoding and decoding efficiency in high-resolution video is solved, achieving higher encoding and decoding efficiency and compression rate, and is applicable to existing and future video encoding and decoding standards.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2021-02-07
- Publication Date
- 2026-05-26
AI Technical Summary
Existing video encoding and decoding technologies are insufficient in terms of encoding and decoding efficiency and compression rate when processing high-resolution video, making it difficult to meet the ever-increasing bandwidth and quality requirements.
The video blocks are segmented using a geometric segmentation mode (GEO). The boundary is defined by angle and distance offsets. Combined with multi-parameter set configuration and signaling, the encoding and decoding process of the video blocks is optimized, including the use of angle, displacement and distance parameters, to improve encoding and decoding efficiency.
It improves the performance of video encoding and decoding, enhances encoding and decoding efficiency and compression ratio, and is applicable to existing standards such as HEVC and future video encoding and decoding standards.
Smart Images

Figure CN115136601B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] In accordance with applicable patent law and / or the rules of the Paris Convention, this application timely claims priority and benefit to International Patent Application No. PCT / CN2020 / 074499, filed on February 7, 2020. The entire disclosure of International Patent Application No. PCT / CN2020 / 074499 is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This document covers video encoding and decoding technologies, systems, and equipment. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This invention relates to video codec technology. Specifically, it concerns inter-frame prediction and correlation techniques in video codec. This invention can be applied to existing video codec standards, such as HEVC, and can also be applied to pending standards (Multi-Functional Video Codec). This invention is also applicable to future video codec standards or video codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of video and a bitstream representation of the video, wherein a codec mode of the current block divides the current block into two or more sub-regions comprising at least one non-rectangular or non-square sub-region, wherein the bitstream representation includes signaling associated with the codec mode, and wherein the signaling corresponds to a set of parameters having a first set of values for the current block of video and a second set of values for subsequent blocks.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of video and a bitstream representation of the video, wherein a codec mode of the current block divides the current block into two or more sub-regions comprising at least one non-rectangular or non-square sub-region, wherein the codec mode can be configured using multiple parameter sets, and wherein the bitstream representation includes signaling for a subset of the multiple parameter sets, and wherein the parameter sets include angles, displacements, and distances associated with at least one non-rectangular or non-square sub-region.
[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for a conversion between a current block of video and a bitstream representation of the video, determining the enabling of a first codec mode and a second codec mode different from the first codec mode, wherein the first codec mode divides the current block into two or more sub-regions comprising at least one non-rectangular or non-square sub-region; and performing the conversion based on the determination.
[0009] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded and decoded using a geometric segmentation pattern, wherein the geometric segmentation pattern includes an entire set of segmentation patterns; determining one or more segmentation patterns for the current video block based on segmentation pattern indices included in the bitstream, wherein the segmentation pattern indices correspond to different segmentation patterns from one video block to another; and performing a conversion based on the one or more segmentation patterns.
[0010] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for a conversion between a current video block and a bitstream of that current video block, determining that the video block is encoded and decoded using geometric segmentation patterns, wherein the geometric segmentation patterns comprise an entire set of segmentation patterns, and each segmentation pattern is associated with a set of parameters including at least one of angle, distance, and / or displacement; deriving a subset of the segmentation patterns or parameters from the entire set of segmentation patterns or parameters; and performing a conversion based on the subset of segmentation patterns or parameters.
[0011] In another representative aspect, the disclosed technology can be used to provide a method for video processing. This method includes: for a video block and a bitstream of that video block, determining the enabling of a geometric segmentation mode and a second encoding / decoding mode different from the geometric segmentation mode for the video block; and performing the transformation based on that determination.
[0012] In another representative aspect, the disclosed technology can be used to provide a method for video processing. This method includes: for a video block and its bitstream, determining a deblocking process associated with the video block based on whether the current video block is encoded / decoded in the current video block's geometric segmentation mode and / or color format; and performing the conversion based on the deblocking process.
[0013] In another representative aspect, a method for storing a bitstream of video is disclosed. The method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded and decoded using a geometric segmentation pattern, wherein the geometric segmentation pattern includes an entire set of segmentation patterns; determining one or more segmentation patterns of the current video block based on segmentation pattern indices included in the bitstream, wherein the segmentation pattern indices correspond to different segmentation patterns from one video block to another; generating a bitstream from the video block using the one or more segmentation patterns; and storing the bitstream in a non-transitory computer-readable recording medium.
[0014] In yet another example, a video encoder apparatus including a processor configured to implement the methods described above is disclosed.
[0015] In yet another example, a video decoder apparatus including a processor configured to implement the methods described above is disclosed.
[0016] In yet another example, a computer-readable medium is disclosed. Code is stored on this computer-readable medium. When executed by a processor, this code causes the processor to perform the methods described above.
[0017] These and other aspects are described in this document. Attached Figure Description
[0018] Figure 1 An example of the Triangle Prediction Pattern (TPM) is shown.
[0019] Figure 2 An example of a boundary description for a geometric codec mode (GEO) is shown.
[0020] Figure 3A An example of edges supported in GEO is shown.
[0021] Figure 3B It shows the geometric relationship between a given pixel location and two edges.
[0022] Figure 4 Examples of different angles of the GEO and their corresponding aspect ratios are shown.
[0023] Figure 5 An example of the angular distribution of 64 GEO patterns is shown.
[0024] Figure 6 An example of GEO / TPM boundary division with angleIdx 0 to 31 in a counterclockwise direction is shown.
[0025] Figure 7 Another example of GEO / TPM boundary division with angleIdx 0 to 31 in a clockwise direction is shown.
[0026] Figure 8 A flowchart of an example method for video processing is shown.
[0027] Figure 9 This is a block diagram of an example video processing device.
[0028] Figure 10 This is a block diagram illustrating an example video codec system.
[0029] Figure 11 This is a block diagram showing an example encoder.
[0030] Figure 12 This is a block diagram showing an example decoder.
[0031] Figure 13 This is a block diagram of an example video processing system in which the disclosed technology can be implemented.
[0032] Figure 14 A flowchart of an example method for video processing is shown.
[0033] Figure 15 A flowchart of an example method for video processing is shown.
[0034] Figure 16 A flowchart of an example method for video processing is shown.
[0035] Figure 17 A flowchart of an example method for video processing is shown.
[0036] Figure 18 A flowchart of an example method for video processing is shown. Detailed Implementation
[0037] Due to the ever-increasing demand for higher resolution video, video encoding and decoding methods and technologies are ubiquitous in modern technology. Video codecs typically consist of electronic circuitry or software that compresses or decompresses digital video and are constantly being improved to provide higher encoding and decoding efficiency. Video codecs convert uncompressed video into a compressed format and vice versa. There are complex relationships between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end latency (delay). Compression formats typically conform to standard video compression specifications, such as the High Efficiency Video Codec (HEVC) standard (also known as H.265 or MPEG-H Part 2), pending multi-functional video codec standards, or other current and / or future video codec standards.
[0038] Embodiments of the disclosed techniques can be applied to existing video codec standards (e.g., HEVC, H.265) and future standards to improve runtime performance. Specifically, it relates to merge modes in video codecs. Section headings are used in this document to improve readability and do not in any way limit the discussion or embodiments (and / or implementations) to the respective sections.
[0039] 1. Background
[0040] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 video codec, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new codec standard is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.
[0041] 1.1. Geometric Segmentation (GEO) for Inter-Frame Prediction
[0042] The following descriptions are excerpted from JVET-P0884, JVET-P0107, JVET-P0304, JVET-P0264, JVET-Q0079, JVETQ0059, JEVT-Q0077 and JVET-Q0309.
[0043] The Geometric Merge Model (GEO) was proposed at the 15th JVET conference in Gothenburg and is an extension of the existing Triangular Prediction Model (TPM). At the 16th JVET conference in Geneva, a simpler design of the GEO model from JVET-P0884 was selected as the CE anchor for further research. At the 17th JVET conference in Brussels, GEO was adopted into VTM8 to replace the TPM model in VTM7, and the GEO model was renamed the GPM model in VVC WD8.
[0044] Figure 1 The diagram illustrates the TPM in VTM-6.0 and the additional shapes proposed for GEO inter-frame blocks.
[0045] The boundary of the geometric merge pattern is determined by angle. and distance offset ρ i Description, such as Figure 2 As shown. Angle This represents the quantized angle between 0 and 360 degrees, and the distance offset ρ. i Represents the maximum distance ρ max The quantization offset was calculated. Furthermore, partitioning directions overlapping with binary tree partitioning and TPM partitioning were excluded.
[0046] In JVET-P0884, GEO is applied to block sizes of not less than 8×8, and for each block size, there are 82 different partitioning methods, distinguished by 24 angles and 4 edges relative to the center of the CU. Figure 3A The diagram shows four edges uniformly distributed along the normal vector direction within the CU, starting from Edge0 which passes through the center of the CU. Each segmentation pattern in the GEO (i.e., a pair of angular and edge indices) is assigned a pixel-adaptive weight table to blend samples from the two segmentation parts. The weight values of the samples range from 0 to 8 and are determined by the L2 distance from the center of the pixel to the edge. Essentially, when assigning weight values, a cell gain constraint is followed; that is, when a small weight value is assigned to a GEO partition, a large complementary weight value is assigned to the other partition, totaling 8.
[0047] The calculation of the weight value for each pixel is done in two steps: (a) calculating the displacement from the pixel position to a given edge, and (c) mapping the calculated displacement to a weight value using a predefined lookup table. The way the displacement from pixel position (x,y) to a given edge is calculated is essentially the same as calculating the displacement from (x,y) to Edge0 and subtracting the distance ρ between Edge0 and Edgei from that displacement. Figure 3B The geometric relationship between (x,y) and the edge is shown. Specifically, the displacement from (x,y) to the edge can be expressed by the following formula:
[0048]
[0049] The value of ρ is a function of the maximum length of the normal vector (denoted by ρmax) and the edge index i, that is:
[0050]
[0051] Where N is the number of edges supported by GEO, and "1" is to prevent the last edge, EdgeN-1, from getting too close to the CU angle at some angular index. Substituting equation (8) into (6), we can calculate the displacement from each pixel (x,y) to a given Edgei. In short, we will This is represented as wIdx(x,y). ρ needs to be calculated once for each CU, and wIdx(x,y) needs to be calculated once for each sample point, which involves multiplication.
[0052] 1.1.1.JVET-P0884
[0053] Based on CE4-1.14 of the 16th Geneva JVET Conference, JVET-P0884 incorporates the simplifications of the proposed slope-based version 2 of JVET-P0107, JVET-P0304, and JVET-P0264 test 1.
[0054] a) In the joint contribution, the geo angle is the same slope (an entangled power of 2) defined in JVET-P0107 and JVET-P0264. The slope used in this proposal is (1, 1 / 2, 1 / 4, 4, 2). In this case, if the blending mask is computed on the fly, the multiplication will be replaced by a shift operation.
[0055] b) The rho calculation is replaced by offsets X and Y, as described in JVET-P304. In this case, only 24 mixing masks need to be stored without calculating the mixing mask on the fly.
[0056] 1.1.2.JVET-P0107
[0057] Based on this slope-based GEO version 2, Table 1 illustrates the Dis[.] lookup table.
[0058] Table 1. 2-bit Dis[.] lookup table for slope-based GEO
[0059]
[0060]
[0061] Using slope-based GEO version 2, the computational complexity of GEO blending mask derivation is considered to be multiplication (up to 2-bit shift) and addition. There are no different partitions compared to TPM. Furthermore, the rounding operation of distFromLine has been removed for easier storage of blending masks. This bug fix guarantees that sample weights are repeated in each row or column by shifting.
[0062] 1.1.3.JVET-P0264
[0063] In JVET-P0264, the angles in GEO are replaced with angles whose tangent is a power of 2. Because the tangent of the proposed angles is a power of 2, most multiplications can be replaced by shifting. Furthermore, the weight values of these angles can be implemented by repeated phase shifting row-wise or column-wise. Using the proposed angles, each block size and each partitioning pattern requires one row or one column to store.
[0064] 1.1.4.JVET-P0304
[0065] In JVET-P0304, weights and mask masks for motion field storage of all block and segmentation patterns are derived from two predefined mask sets: one set for mixed weight derivation and the other set for motion field storage masks. Each set contains a total of 16 masks. Each mask for each angle is computed using the same equations from GEO, with block width and block height set to 256 and displacement set to 0. For blocks of size W×H with distance ρ, the blending weights of the luminance samples are directly clipped from the predefined mask, and the offset is calculated as follows:
[0066] - Variables offsetX and offsetY are calculated as follows:
[0067]
[0068]
[0069] -
[0070] Where g_sampleWeight L [] represents a predefined mask for the mixed weights.
[0071] 1.1.5.JVET-Q0079
[0072] At the JVET-P conference, a simplified GEO mode was proposed in JVET-P0884 / JVET-P0885 and suggested as a common basis for CE4 core experiments. In this common basis, the GEO mode is applied to merge blocks with a width and height greater than or equal to 8. When encoding and decoding a block using the GEO mode, a signaling index indicates which of the 82 segmentation modes should be used to divide the block into two partitions. Each partition uses its own motion vector for inter-frame prediction. After predicting each partition, a weighted mixing process is used to adjust the sample values along the segmentation edges. This is the predicted signal for the entire block, and the transform and quantization processes are applied to the entire block as in other prediction modes.
[0073] The weights of the brightness samples used in the mixing process are calculated as follows:
[0074] weightIdx=(((x+offsetX)<<1)+1)*Dis[displacementX]+(((y+offsetY)<<1)+1))*Dis[displacementY]-rho.
[0075] weightIdxAbs=Clip3(0,26,abs(weightIdx)).
[0076] sampleWeight=weightIdx<=0? GeoFilter[weightIdxAbs]:8-GeoFilter[weightIdxAbs]
[0077] Dis[.] is a lookup table with 24 entries and possible output values {0, 1, 2, 4}. GeoFilter[.] is a lookup table with 27 entries. The variables rho, offsetX, and offsetY are pre-computed based on the following:
[0078] rho=(Dis[displacementX]<<8)+(Dis[displacementY]<<8)
[0079]
[0080]
[0081] ρ and ρ represent the angle and distance derived from the lookup table using the index notified by signaling.
[0082] The weights of the chroma samples are subsampled from the luminance weights. The motion storage mask is derived independently using the same weight derivation method. More detailed information about this common foundation can be found in JVET-P0884 / JVET-P0885.
[0083] 1.1.6.JVET-Q0059
[0084] In the JVET-Q conference, geometric inter-frame prediction of 64 modes was adopted (i.e., JVET-Q0059).
[0085] Because moving objects in natural video sequences are mostly vertically positioned, near-horizontal partitioning patterns are not frequently used. In the proposed 64-pattern GEO, we removed angles {5, 7, 17, 19}. Figure 5 The angular distribution of GEO for 64 modes is shown.
[0086] Furthermore, in the 82-pattern GEO, the distance index of 2 for horizontal angles {0,12} and vertical angles {6,18} overlaps with the ternary tree partition boundary. These were also removed in the proposed 64-pattern GEO.
[0087] The total number of patterns proposed by the method can be calculated as 10*4 + 10*3 – 2 – 4 = 64 patterns.
[0088] Furthermore, since the geo partitioning mode is signaled by using truncated binary, the 64 modes will be signaled most efficiently in 6-bit TB.
[0089] 1.1.7. JVET-Q0077 and JVET-Q0309
[0090] Disable GEO for blocks larger than 64x64, and disable GEO for 64x8 and 8x64 blocks.
[0091] 1.1.8. GEO / GPM Angle Index and Angle Dimensions
[0092] In the latest VVC working draft 8, the GEO mode has been renamed the Geometric Partition Mode (GPM). The GEO / GPM angular index is used to represent the boundary that divides a GEO / GPM block into two sub-regions, such as... Figure 6 As shown. As described in JVET-P0264, the tangent of the GEO / GPM angle is a power of 2, depending on the block's aspect ratio. In JVET-Q2001-vB, the size range of the GEO / GPM angle is from 0° to 352.87°, and the associated GEO / GPM angle index ranges from 0 to 31, as shown. Figure 6As shown. The angle between the vertical direction (e.g., overlapping with the partition boundary where angleIdx equals 0) and the specified GEO partition boundary is defined as the size of the GEO / GPM angle in the GEO / GPM mode. The correspondence between the size of the GEO / GPM angle (angle size) and the GEO / GPM angle index (angleIdx) can be found in Table 2.
[0093] Table 2: Example of the relationship between angle index and angle dimension (counterclockwise)
[0094]
[0095] 1.2. GEO / GPM Specifications in JVET-Q2001-vB
[0096] The following specifications are excerpted from the working draft provided in JVET-Q2001-vB.
[0097] 7.3.10.5 Encoding / Decoding Unit Syntax
[0098]
[0099]
[0100]
[0101]
[0102] 7.3.10.7 Merge Data Syntax
[0103]
[0104] `sps_gpm_enabled_flag` specifies whether geometry-partition-based motion compensation can be used for inter-frame prediction. `sps_gpm_enabled_flag` equal to 0 indicates that the syntax should be constrained so that geometry-partition-based motion compensation is not used in CLVS, and `merge_gpm_partition_idx`, `merge_gpm_idx0`, and `merge_gpm_idx1` do not exist in the CLVS codec unit syntax. `sps_gpm_enabled_flag` equal to 1 indicates that geometry-partition-based motion compensation can be used in CLVS. When it does not exist, the value of `sps_gpm_enabled_flag` is inferred to be 0.
[0105] max_num_merge_cand_minus_max_num_gpm_cand specifies the maximum number of geometric segmentation merge pattern candidates supported in the SPS, subtracted from MaxNumMergeCand.
[0106] If sps_gpm_enabled_flag equals 1 and MaxNumMergeCand is greater than or equal to 3, then the maximum number of geometric segmentation merge pattern candidates, MaxNumGeoMergeCand, is derived as follows:
[0107] if(sps_gpm_enabled_flag&&MaxNumMergeCand>=3)
[0108] MaxNumGpmMergeCand=MaxNumMergeCand-max_num_merge_cand_minus_max_num_gpm_cand
[0109] else if(sps_gpm_enabled_flag&&MaxNumMergeCand==2)
[0110] MaxNumMergeCand=2
[0111] else MaxNumGeoMergeCand=0
[0112] The value of MaxNumGeoMergeCand should be in the range of 2 to MaxNumMergeCand, inclusive.
[0113] The variable MergeGpmFlag[x0][y0] specifies whether to use geometry-based motion compensation to generate prediction samples for the current codec unit when decoding B stripes. Its derivation is as follows:
[0114] - MergeGpmFlag[x0][y0] is set to 1 if all of the following conditions are true:
[0115] -sps_gpm_enabled_flag equals 1.
[0116] -slice_type equals B.
[0117] -general_merge_flag[x0][y0] equals 1.
[0118] -cbWidth is greater than or equal to 8.
[0119] -cbHeight is greater than or equal to 8.
[0120] -cbWidth is less than 8*cbHeight.
[0121] -cbHeight is less than 8*cbWidth.
[0122] -regular_merge_flag[x0][y0] equals 0.
[0123] -merge_subblock_flag[x0][y0] equals 0.
[0124] -ciip_flag[x0][y0] equals 0.
[0125] Otherwise, MergeGpmFlag[x0][y0] is set to 0.
[0126] `merge_gpm_partition_idx[x0][y0]` specifies the segmentation shape of the geometric segmentation merge pattern. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0127] When merge_gpm_partition_idx[x0][y0] does not exist, it is inferred to be equal to 0.
[0128] merge_gpm_idx0[x0][y0] specifies the first Merge candidate index of the motion compensation candidate list based on geometric segmentation, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0129] When merge_gpm_idx0[x0][y0] does not exist, it is inferred to be equal to 0.
[0130] merge_gpm_idx1[x0][y0] specifies the second Merge candidate index of the motion compensation candidate list based on geometric segmentation, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0131] When merge_gpm_idx1[x0][y0] does not exist, it is inferred to be equal to 0.
[0132] 8.5.7 Decoding process of inter-frame blocks in geometric segmentation mode
[0133] 8.5.7.1 Overview
[0134] This procedure is invoked when decoding a codec unit where MergeGpmFlag[xCb][yCb] equals 1.
[0135] The input to this process is:
[0136] -Luminance position (xCb, yCb), specifies the top-left luminance sample of the current codec block relative to the top-left luminance sample of the current image.
[0137] - The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0138] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0139] Luminance motion vectors mvA and mvB with a precision of -1 / 16 fractional sample points
[0140] - Chromaticity motion vectors mvCA and mvCB
[0141] -Refer to indices refIdxA and refIdxB,
[0142] - Prediction list flags predListFlagA and predListFlagB.
[0143] The output of this process is:
[0144] - An array of (cbWidth) x (cbHeight) brightness prediction samples, named predSamples L ,
[0145] - When ChromaArrayType is not equal to 0, the array predSamples of the chromaticity prediction samples of component Cb is (cbWidth / SubWidthC)x(cbHeight / SubHeightC). Cb ,
[0146] - When ChromaArrayType is not equal to 0, the array predSamples of (cbWidth / SubWidthC)x(cbHeight / SubHeightC) of the chromaticity prediction samples of component Cr. Cr .
[0147] Let predSamplesLA L and predSamplesLB L Let predSamplesLA be an array of (cbWidth) x (cbHeight) values for predicting luminance samples, and let predSamplesLA be a non-zero array.Cb ,predSamplesLB Cb ,predSamplesLA Cr and predSamplesLB Cr This is an array of (cbWidth / SubWidthC)x(cbHeight / SubHeightC) for predicting chromaticity sample values.
[0148] predSamples L ,predSamples Cb and predSamples Cr The following sequential steps will be used to derive the following:
[0149] 1. For N to be each of A and B, the following applies:
[0150] - composed of an ordered two-dimensional array of brightness samples, refPicLN L The two ordered two-dimensional arrays refPicLN for chromaticity samples Cb and refPicLN Cr The resulting reference image is derived by invoking the procedure specified in Clause 8.5.6.2, with X set to be equal to predListFlagN and refIdxX set to be equal to refIdxN as input.
[0151] - array predSamplesLN L The derivation is performed by calling the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xCb, yCb), the luminance codec block width sbWidth is set to equal cbWidth, the luminance codec block height sbHeight is set to equal cbHeight, the motion vector offset mvOffset is set to equal (0,0), the motion vector mvLX is set to equal mvN, and the refPicLN is set to equal refPicLN. L The reference array refPicLX L The variables bdofFlag (set to FALSE), cIdx (set to 0), RprConstraintsActive[X][refIdxLX], and RefPicScale[predListFlagN][refIdxN] are used as inputs.
[0152] - When ChromaArrayType is not equal to 0, the array predSamplesLN CbThe derivation is performed by invoking the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xCb, yCb), the codec block width sbWidth set to equal cbWidth / SubWidthC, the codec block height sbHeight set to equal cbHeight / SubHeightC, the motion vector offset mvOffset set to equal (0,0), the motion vector mvLX set to equal mvCN, and the refPicLN set to equal... Cb The reference array refPicLX Cb The variables bdofFlag (set to FALSE), cIdx (set to 1), RprConstraintsActive[X][refIdxLX], and RefPicScale[predListFlagN][refIdxN] are used as inputs.
[0153] - When ChromaArrayType is not equal to 0, the array predSamplesLN Cr The derivation is performed by invoking the fractional sample interpolation procedure specified in Clause 8.5.6.3, where the luminance position (xCb, yCb), the codec block width sbWidth set to equal cbWidth / SubWidthC, the codec block height sbHeight set to equal cbHeight / SubHeightC, the motion vector offset mvOffset set to equal (0,0), the motion vector mvLX set to equal mvCN, and the refPicLN set to equal... Cr The reference array refPicLX Cr The variables bdofFlag (set to FALSE), cIdx (set to 2), RprConstraintsActive[X][refIdxLX], and RefPicScale[predListFlagN][refIdxN] are used as inputs.
[0154] 2. The segmentation angle variable angleIdx and distance variable distanceIdx in the geometric segmentation mode are set according to the value of merge_gpm_partition_idx[xCb][yCb], as specified in Table 36.
[0155] 3. PredSamples within the current luminance codec block L [x L ][y L (where x) L =0..cbWidth-1, and y L=0..cbHeight-1) is derived by calling the weighted sample prediction procedure of the geometric segmentation mode specified in Clause 8.5.7.2, where the codec block width nCbW is set to equal cbWidth, the codec block height nCbH is set to equal cbHeight, and the sample array predSamplesLA is... L and predSamplesLB L The variables angleIdx, distanceIdx, and cIdx (which is equal to 0) are used as inputs.
[0156] 4. When ChromaArrayType is not equal to 0, the predicted samples (predSamples) within the current chroma component Cb codec block. Cb [x C ][y C (where x) C =0..cbWidth / SubWidthC-1, and y C =0..cbHeight / SubHeightC-1) is derived by calling the weighted sample prediction procedure of the geometric segmentation mode specified in Clause 8.5.7.2, where the codec block width nCbW is set to be equal to cbWidth / SubWidthC, the codec block height nCbH is set to be equal to cbHeight / SubHeightC, and the sample array predSamplesLA is... Cb and predSamplesLB Cb The variables angleIdx, distanceIdx, and cIdx equal to 1 are used as inputs.
[0157] 5. When ChromaArrayType is not equal to 0, the predicted samples predSamplesCr[x] within the current chroma component Cr codec block. C ][y C (where x) C =0..cbWidth / SubWidthC-1 and y C =0..cbHeight / SubHeightC-1) is derived by calling the weighted sample prediction procedure of the geometric segmentation mode specified in Clause 8.5.7.2, where the codec block width nCbW is set to be equal to cbWidth / SubWidthC, the codec block height nCbH is set to be equal to cbHeight / SubHeightC, and the sample array predSamplesLA is... Cr and predSamplesLB CrThe variables angleIdx, distanceIdx, and cIdx equal to 2 are used as inputs.
[0158] 6. Call the motion vector storage procedure for the merge geometry segmentation pattern specified in Clause 8.5.7.3, with the luma codec block position (xCb, yCb), luma codec block width cbWidth, luma codec block height cbHeight, segmentation angle angleIdx and distanceIdx, luma motion vectors mvA and mvB, reference indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB as input.
[0159] Table 36 – Specifications of angleIdx and distanceIdx based on merge_gpm_partition_idx
[0160] merge_gpm_partition_idx 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 angleIdx 0 0 2 2 2 2 3 3 3 3 4 4 4 4 5 5 distanceIdx 1 3 0 1 2 3 0 1 2 3 0 1 2 3 0 1 merge_gpm_partition_idx 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 angleIdx 5 5 8 8 11 11 11 11 12 12 12 12 13 13 13 13 distanceIdx 2 3 1 3 0 1 2 3 0 1 2 3 0 1 2 3 merge_gpm_partition_idx 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 angleIdx 14 14 14 14 16 16 18 18 18 19 19 19 20 20 20 21 distanceIdx 0 1 2 3 1 3 1 2 3 1 2 3 1 2 3 1 merge_gpm_partition_idx 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 angleIdx 21 21 24 24 27 27 27 28 28 28 29 29 29 30 30 30 distanceIdx 2 3 1 3 1 2 3 1 2 3 1 2 3 1 2 3
[0161] 8.5.7.2 Weighted Sample Prediction Process for Geometric Segmentation Pattern
[0162] The input to this process is:
[0163] - Two variables, nCbW and nCbH, specify the width and height of the current codec block.
[0164] - Two arrays, predSamplesLA and predSamplesLB, of type (nCbW)x(nCbH).
[0165] - The variable angleIdx specifies the angle index of the geometric segmentation.
[0166] - The variable distanceIdx specifies the distance index of the geometric segment.
[0167] - The variable cIdx specifies the color component index.
[0168] The output of this process is an array pbSamples of (nCbW)x(nCbH) predicted sample values.
[0169] The variables nW, nH, shift1, offset1, hwRatio, displacementX, displacementY, partFlip, and shiftHor are derived as follows:
[0170] nW=(cIdx==0)? nCbW:nCbW*SubWidthC (1030)
[0171] nH=(cIdx==0)? nCbH:nCbH*SubHeightC (1031)
[0172] shift1=Max(5,17-BitDepth) (1032)
[0173] offset1=1<<(shift1-1) (1033)
[0174] hwRatio=nH / nW (1034)
[0175] displacementX=angleIdx (1035)
[0176] displacementY=(angleIdx+8)%32 (1036)
[0177] partFlip=(angleIdx>=13&&angleIdx<=27)? 0:1 (1037)
[0178] shiftHor=(angleIdx%16==8||(angleIdx%16!=0&&hwRatio>0))? 0:1(1038)
[0179] The variables offsetX and offsetY are derived as follows:
[0180] - If shiftHor equals 0, the following applies:
[0181] offsetX = (-nW) >> 1 (1039)
[0182] offsetY=((-nH)>>1)+(angleIdx<16?(distanceIdx*nH)>>3:-((distanceIdx*nH)>>3)) (1040)
[0183] - Otherwise (shiftHor equals 1), the following applies:
[0184] offsetX=((-nW)>>1)+(angleIdx<16?(distanceIdx*nW)>>3:-((distanceIdx*nW)>>3)) (1041)
[0185] offsetY = (-nH) >> 1 (1042)
[0186] The predicted sample points pbSamples[x][y] (where x = 0..nCbW-1 and y = 0..nCbH-1) are derived as follows:
[0187] - The variables xL and yL are derived as follows:
[0188] xL=(cIdx==0)? x:x*SubWidthC (1043)
[0189] yL=(cIdx==0)? y:y*SubHeightC (1044)
[0190] - The variable wValue, which specifies the weights of the predicted samples, is derived based on the array disLut specified in Table 37 as follows:
[0191] weightIdx=(((xL+offsetX)<<1)+1)*disLut[displacementX]+(((yL+offsetY)<<1)+1))*disLut[displacementY] (1045)
[0192] weightIdxL=partFlip? 32+weightIdx:32-weightIdx (1046)
[0193] wValue=Clip3(0,8,(weightIdxL+4)>>3) (1047)
[0194] -The predicted sample values are derived as follows:
[0195] pbSamples[x][y] = Clip3(0, (1< <BitDepth)-1,(predSamplesLA[x][y]*wValue+(1048)
[0196] predSamplesLB[x][y]*(8-wValue)+offset1)>>shift1)
[0197] Table 37 - Specification of the geometric segmentation distance array disLut
[0198] idx 0 2 3 4 5 6 8 10 11 12 13 14 disLut[idx] 8 8 8 4 4 2 0 -2 -4 -4 -8 -8 idx 16 18 19 20 21 22 24 26 27 28 29 30 disLut[idx] -8 -8 -8 -4 -4 -2 0 2 4 4 8 8
[0199] 8.5.7.3 Motion Vector Storage Procedure for Geometric Segmentation Pattern
[0200] This procedure is invoked when decoding a codec unit where MergeGpmFlag[xCb][yCb] equals 1.
[0201] The input to this process is:
[0202] -Luminance position (xCb, yCb), specifies the top-left luminance sample of the current codec block relative to the top-left luminance sample of the current image.
[0203] - The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0204] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0205] - The variable angleIdx specifies the angle index of the geometric segmentation.
[0206] - The variable distanceIdx specifies the distance index of the geometric segment.
[0207] Luminance motion vectors mvA and mvB with a precision of -1 / 16 fractional sample points
[0208] -Refer to indices refIdxA and refIdxB,
[0209] - Prediction list flags predListFlagA and predListFlagB.
[0210] The variables numSbX and numSbY, which specify the number of 4×4 blocks in the horizontal and vertical directions of the current codec block, are set to cbWidth>>2 and cbHeight>>2, respectively.
[0211] The variables hwRatio, displacementX, displacementY, partIdx, and shiftHor are derived as follows:
[0212] hwRatio=cbHeight / cbWidth (1049)
[0213] displacementX=angleIdx (1050)
[0214] displacementY=(angleIdx+8)%32 (1051)
[0215] partIdx=(angleIdx>=13&&angleIdx<=27)? 0:1 (1052)
[0216] shiftHor=(angleIdx%16==8||(angleIdx%16!=0&&hwRatio>0)?0:1(1053)
[0217] The variables offsetX and offsetY are derived as follows:
[0218] - If shiftHor equals 0, the following applies:
[0219] offsetX=(-cbWidth)>>1 (1054)
[0220] offsetY=((-cbHeight)>>1)+(angleIdx<16?(distanceIdx*cbHeight)>>3:-((distanceIdx*cbHeight)>>3)) (1055)
[0221] - Otherwise (shiftHor equals 1), the following applies:
[0222] offsetX=((-cbWidth)>>1)+(angleIdx<16?(distanceIdx*cbWidth)>>3:-((distanceIdx*cbWidth)>>3)) (1056)
[0223] offsetY=(-cbHeight)>>1 (1057)
[0224] For each 4×4 sub-block at sub-block index (xSbIdx, ySbIdx) (where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1), the following applies:
[0225] - The variable motionIdx is calculated based on the array disLut specified in Table 37 as follows:
[0226] motionIdx=(((4*xSbIdx+offsetX)<<1)+5)*disLut[displacementX]+(((4*ySbIdx+offsetY<<1)+5))*disLut[displacementY] (1058)
[0227] The variable sType is deduced as follows:
[0228] sType=abs(motionIdx)<32?2:(motionIdx<=0?(1-partIdx):partIdx) (1059)
[0229] - Depending on the value of sType, the following assignments are made:
[0230] - If sType is equal to 0, the following applies:
[0231] predFlagL0=(predListFlagA==0)?1:0 (1060)
[0232] predFlagL1=(predListFlagA==0)?0:1 (1061)
[0233] refIdxL0=(predListFlagA==0)?refIdxA:-1 (1062)
[0234] refIdxL1=(predListFlagA==0)?-1:refIdxA (1063)
[0235] mvL0[0]=(predListFlagA==0)?mvA[0]:0 (1064)
[0236] mvL0[1]=(predListFlagA==0)?mvA[1]:0 (1065)
[0237] mvL1[0]=(predListFlagA==0)?0:mvA[0] (1066)
[0238] mvL1[1]=(predListFlagA==0)?0:mvA[1] (1067)
[0239] - Otherwise, if sType is equal to 1 or (sType is equal to 2, predListFlagA+predListFlagB is not equal to 1), then the following applies:
[0240] predFlagL0=(predListFlagB==0)?1:0 (1068)
[0241] predFlagL1=(predListFlagB==0)?0:1 (1069)
[0242] refIdxL0=(predListFlagB==0)?refIdxB:-1 (1070)
[0243] refIdxL1=(predListFlagB==0)?-1:refIdxB (1071)
[0244] mvL0[0]=(predListFlagB==0)?mvB[0]:0 (1072)
[0245] mvL0[1] = (predListFlagB == 0)? mvB[1] : 0 (1073)
[0246] mvL1[0] = (predListFlagB == 0)? 0 : mvB[0] (1074)
[0247] mvL1[1] = (predListFlagB == 0)? 0 : mvB[1] (1075)
[0248] - Otherwise (sType equals 2 and predListFlagA + predListFlagB equals 1), the following applies:
[0249] predFlagL0 = 1 (1076)
[0250] predFlagL1 = 1 (1077)
[0251] refIdxL0 = (predListFlagA == 0)? refIdxA : refIdxB (1078)
[0252] refIdxL1 = (predListFlagA == 0)? refIdxB : refIdxA (1079)
[0253] mvL0[0] = (predListFlagA == 0)? mvA[0] : mvB[0] (1080)
[0254] mvL0[1] = (predListFlagA == 0)? mvA[1] : mvB[1] (1081)
[0255] mvL1[0] = (predListFlagA == 0)? mvB[0] : mvA[0] (1082)
[0256] mvL1[1] = (predListFlagA == 0)? mvB[1] : mvA[1] (1083)
[0257] - For x = 0..3 and y = 0..3, perform the following assignments:
[0258] MvL0[(xSbIdx << 2) + x][(ySbIdx << 2) + y] = mvL0 (1084)
[0259] MvL1[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=mvL1 (1085)
[0260] MvDmvrL0[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=mvL0 (1086)
[0261] MvDmvrL1[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=mvL1 (1087)
[0262] RefIdxL0[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=refIdxL0 (1088)
[0263] RedIdxL1[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=refIdxL1 (1089)
[0264] PredFlagL0[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=predFlagL0 (1090)
[0265] PredFlagL1[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=predFlagL1 (1091)
[0266] BcwIdx[(xSbIdx<<2)+x][(ySbIdx<<2)+y]=0 (1092)
[0267] 2. Disadvantages of existing solutions
[0268] There are several potential problems in the current design of GEO, as described below.
[0269] (1) The current GEO mode is distributed as symmetrical GEO angle and symmetrical GEO distance / displacement, which may be ineffective for natural video encoding and decoding.
[0270] (2) The current GEO model can be presented using weighted prediction, which may lead to visual artifacts.
[0271] 3. Embodiments of the disclosed technology
[0272] The detailed inventions described below should be considered as examples for explaining general concepts. These inventions should not be interpreted in a narrow sense. Furthermore, these inventions can be combined in any way.
[0273] The term "GEO" can refer to an encoding / decoding method that divides a block into two or more sub-regions, where at least one sub-region is non-rectangular, or it cannot be generated by any existing segmentation structure (e.g., QT / BT / TT) that divides a block into multiple rectangular sub-regions. In one example, for a GEO-encoded block, one or more weighted masks are derived for the codec block based on how the sub-regions are divided, and the final prediction signal of the codec block is generated by a weighted sum of two or more auxiliary prediction signals associated with the sub-regions. The term "GEO" can also indicate Geometric Merge Mode (GEO), and / or Geometric Partition Mode (GPM), and / or Wedge Prediction Mode, and / or Triangle Prediction Mode (TPM).
[0274] The term "block" can refer to codec block (CB), CU, PU, TU, PB, TB.
[0275] 1. The interpretation of the GEO mode index of signaling notification can adaptively change from one video unit to another. That is, for the same signaling notification value, it can be interpreted as different angles and / or different distances.
[0276] a) In one example, the mapping between the GEO pattern index of the signaling notification and its corresponding GEO angle / distance can depend on the block dimension (e.g., the ratio of block width to height).
[0277] b) Alternatively, binarization for GEO mode index encoding and decoding can be non-fixed-length encoding and decoding, where the number of binary bits to be encoded and decoded can be different for two different modes.
[0278] c) In one example, unequal numbers of GEO modes / angles / displacements / distances can be used in different video units.
[0279] i. In one example, the amount of GEO mode / angle / displacement / distance used in a video unit can depend on the block dimensions (e.g., width, height, aspect ratio, etc.).
[0280] ii. In one example, more GEO angles can be used in block A than in block B, where A and B can have different dimensions (e.g., A can indicate a block with a height greater than its width, and B can indicate a block with a height less than or equal to its width).
[0281] iii. In one example, the allowed GEO angle of the video unit can be asymmetrical.
[0282] iv. In one example, the allowed GEO angle of the video unit may not be rotationally symmetric.
[0283] v. In one example, the allowed GEO angle of the video unit may not be bilaterally symmetrical.
[0284] vi. In one example, the allowed GEO angle of the video unit may not be quarter-symmetric.
[0285] d) In one example, a video unit can be a block, VPDU, slice / strip / picture / subpicture / tile / video.
[0286] 2. How GEO modes are represented in a bitstream can depend on the mode priority, for example, priority is determined by the associated GEO angle and / or distance.
[0287] a) In one example, a GEO mode oriented towards a smaller GEO angle size (i.e., a GEO angle size less than X degrees, such as X = 90 or 45, as described in Table 2 of Section 2.1.8) has a higher priority than a GEO mode oriented towards a larger GEO angle size (i.e., a GEO angle size greater than X degrees).
[0288] b) In one example, a GEO pattern associated with a smaller GEO angle index (i.e., a GEO angle index smaller than Y (such as Y = 4, 8, or 16)) has a higher priority than a GEO pattern associated with a larger GEO angle index (i.e., a GEO angle index greater than Y).
[0289] c) In one example, during signaling notification, the GEO mode with higher priority in the above claims may require fewer binary bits or bits than the GEO mode with lower priority.
[0290] 3. GEO patterns can be classified into two or more categories. Signaling can be used to indicate which category a GEO belongs to before other information related to the GEO pattern.
[0291] a) For example, the signaling notification for GEO codec blocks can be indexed first by whether the signaling notification angle is clockwise or counterclockwise.
[0292] i. Definition of clockwise GEO angle index signaling:
[0293] 1. For example, a GEO angle index notified by counterclockwise signaling can mean that a smaller GEO angle index represents a smaller GEO angle size.
[0294] a) In one example, such as Figure 6As shown in Table 2, the GEO angle index is signaled counterclockwise. If angleIdx = 0 means that the GEO angle size is 0 degrees, then angleIdx = 8 means that the GEO angle size is 90 degrees; angleIdx = 16 means that the GEO angle size is 180 degrees, and angleIdx = 24 means that the GEO angle size is 270 degrees.
[0295] ii. Definition of clockwise GEO angle index signaling:
[0296] 1. For example, a GEO angle index notified by clockwise signaling can mean that a smaller GEO angle index represents a larger GEO angle size.
[0297] a) In one example, with Figure 6 Contrary to Table 2, the GEO angle index is signaled counterclockwise. Figure 7 The clockwise GEO angle signaling is illustrated in Table 3. Assuming that angleIdx = 0 means that the GEO angle size is equal to 0 degrees (equivalent to 360 degrees), then angleIdx = 8 can mean that the GEO angle size is equal to 270 degrees; angleIdx = 16 can mean that the GEO angle size is equal to 180 degrees, and angleIdx = 24 can mean that the GEO angle size is equal to 90 degrees.
[0298] iii. In one example, it can be signaled at the video unit level (such as SPS / VPS / PPS / picture header / subpicture / strip / strip header / piece / tile / CTU / VPDU / CU / block level).
[0299] 1. In one example, a high-level signaling flag (above the block level) can be used to indicate whether the angle associated with the GEO mode of the signaling notification is clockwise or counterclockwise.
[0300] 2. In one example, a signaling notification block-level flag can be used to indicate whether the angle associated with the GEO mode of the signaling notification is clockwise or counterclockwise.
[0301] Table 3: Example of the relationship between angle index and angle dimension (clockwise)
[0302]
[0303]
[0304] 4. A subset of GEO modes / angles / displacements / distances can be derived from the entire set of GEO modes / angles / displacements / distances.
[0305] a) In one example, only a subset of the GEO pattern was used for the block.
[0306] b) In one example, the GEO pattern associated only with a subset of GEO angles / displacements / distances was used for the block.
[0307] c) Alternatively, the selected mode / angle / displacement / distance can be further signaled in the bitstream to indicate whether it is within a subset.
[0308] d) In one example, whether a subset of GEO modes / angles / displacements / distances is used for a video unit or the full set of GEO modes / angles / displacements / distances may depend on the decoding information (e.g., syntax elements and / or block dimensions) of the current video unit or (multiple) previously decoded video units.
[0309] e) In one example, what GEO mode in a subset may depend on the corresponding GEO angle.
[0310] i. In one example, the subset could contain only GEO patterns associated with a distance / displacement equal to 0.
[0311] ii. In one example, the subset may contain only the GEO patterns associated with a specified GEO angle (e.g., GEO patterns associated with a predefined subset of GEO angles, which may be combined with all displacements corresponding to those predefined GEO angles).
[0312] f) In one example, what GEO mode is in a subset may depend on whether LDB (i.e., low-latency B-frame) encoding and decoding are checked.
[0313] i. In one example, different subsets of the GEO mode can be used for LDB (i.e., low-latency B-frames) and RA (i.e., random access) encoding and decoding.
[0314] g) In one example, which GEO mode is in the subset can depend on the reference images in the reference image list of the current image. For example, the state of the reference images in the reference image list can be identified as two cases: Case 1: All reference images are before the current image in the display order; Case 2: At least one reference image is after the current image in the display order.
[0315] i. In one example, different subsets can be used for case 1 and case 2.
[0316] h) In one example, what GEO modes are in a subset may depend on how motion candidates are derived.
[0317] i. In one example, different subsets of the GEO pattern can be used, depending on whether the motion candidates are derived from temporal motion candidates (e.g., TMVP), spatial motion candidates, history-based motion vector predictions (HMVP), or which spatial motion candidates (e.g., left, or top, or top right).
[0318] i) In one example, a subset of GEOs may contain only GEO patterns that divide blocks in the same way as the TPM patterns.
[0319] i. In one example, a subset of GEOs may contain only GEO patterns that divide blocks by lines connecting the top left and bottom right corners of the block or by lines connecting the top right and bottom left corners of the block.
[0320] j) In one example, a subset of GEOs may contain only GEO patterns corresponding to diagonal angles with one or more distance / displacement indices.
[0321] i. In one example, the diagonal angle can indicate the GEO pattern corresponding to the partition boundary, which divides the block by a line connecting the top left and bottom right corners of the block or by a line connecting the top right and bottom left corners of the block.
[0322] ii. In one example, a subset of GEOs may contain only GEO patterns associated with a distance / displacement equal to 0.
[0323] 1. In one example, a subset of GEOs may contain only GEO patterns corresponding to any angle associated with a distance / displacement equal to 0.
[0324] 2. In one example, a subset of GEOs may contain only GEO patterns that correspond to the diagonal angles associated with a distance / displacement equal to 0 (i.e., the block's partition boundaries from top left to bottom right and / or from top right to bottom left).
[0325] iii. In one example, a subset of GEOs may contain only GEO patterns corresponding to the diagonal angles (i.e., the block's partition boundaries from top left to bottom right and / or from top right to bottom left) associated with all distance / displacement indices corresponding to these GEO angles.
[0326] 1. For example, for a block with an aspect ratio (i.e., the ratio of 2 to the power of log2(width) - log2(height)) equal to X, a subset of the GEO may contain only the GEO angular dimensions equal to arctan(X) and / or π-arctan(X), and / or π+arctan(X) and / or 2π-arctan(X), and all distance indices corresponding to these GEO angles (e.g., distanceIdx from 0 to 3 as defined in JVET-Q2001-vB).
[0327] k) For example, for a block with an aspect ratio of 1, a subset of GEOs may contain only GEO patterns corresponding to GEO angle dimensions equal to 45° and / or 135° and / or 225° and / or 315° (e.g., angle indices = 4 and / or 12 and / or 20 and / or 28 as defined in JVET-Q2001-vB), and all distance indices corresponding to these GEO angles (e.g., distance indices from 0 to 3 as defined in JVET-Q2001-vB). In one example, horizontal and / or vertical angles may be included in a subset of GEO angles.
[0328] i. For example, a horizontal angle can mean an angle index corresponding to 90° and / or 270° as described in Table 2 of Section 2.1.8 (i.e., a GEO angle index equal to 8 and / or 24 in Table 36 of JVET-Q2001-vB).
[0329] ii. For example, a vertical angle can mean an angle index corresponding to 0° and / or 180° as described in Table 2 of Section 2.1.8 (i.e., GEO angle indices equal to 0 and / or 6 in Table 36 of JVET-Q2001-vB).
[0330] iii. In one example, GEO patterns associated with horizontal and / or vertical angles that sum to 0 can be included in a subset of the allowed GEO angles.
[0331] 1. Alternatively, GEO patterns associated with horizontal and / or vertical angles that sum to zero may not be included in a subset of the permissible GEO angles.
[0332] iv. In one example, the GEO pattern associated with the horizontal and / or vertical angles of all distance / displacement index combinations can be included in a subset of the allowed GEO angles.
[0333] 1. Alternatively, the GEO patterns associated with the horizontal and / or vertical angles of all distance / displacement index combinations may not be included in a subset of the allowed GEO angles.
[0334] l) It can be signaled (such as SPS / VPS / PPS / picture header / subpicture / strip / strip header / piece / block / CTU / VPDU / CU / block level) to indicate whether a subset of GEO mode / angle / displacement / distance is used.
[0335] i. It can be further signaled (such as SPS / VPS / PPS / picture header / subpicture / strip / strip header / piece / block / CTU / VPDU / CU / block level) to indicate which subset of GEO mode / angle / displacement / distance to use.
[0336] 5. GEO mode can coexist with X (e.g., X is a different codec tool than GEO).
[0337] a) In one example, X can indicate a weighted prediction.
[0338] i. In one example, when weighted prediction is enabled (e.g. at the strip level), GEO can be disabled at the video unit level (such as strip / PPS / SPS / film / subpicture / CU / PU / TU level).
[0339] ii. In one example, whether GEO is used with weighted prediction can depend on the weighting factor.
[0340] 1. In one example, GEO can be disabled if the weighting factor of the weighted prediction is greater than T (e.g., T is a constant value).
[0341] b) In one example, X can indicate BCW.
[0342] c) In one example, X can indicate PROF
[0343] d) In one example, X can indicate BDOF.
[0344] e) In one example, X can indicate DMVR.
[0345] f) In one example, X can indicate SBT.
[0346] g) In one example, when GEO is enabled, the codec tool X can be disabled.
[0347] i. In one example, if GEO is enabled, the instructions of the codec tool X can be communicated without signaling.
[0348] h) In another example, GEO can be disabled when codec tool X is enabled.
[0349] i. In one example, if codec tool X is enabled, the instruction from GEO can be notified without signaling.
[0350] i) In another example, when weighted prediction is enabled (e.g., at the stripe level), the codec tool X can be disabled.
[0351] j) Alternatively, the deblocking process (such as deblocking strength, deblocking edge detection, deblocking edge type, etc.) may depend on whether GEO coexists with the codec tool X.
[0352] 6. The deblocking process (such as deblocking intensity, deblocking edge detection, deblocking edge type, etc.) can depend on whether GEO is applied. For codec units using GEO encoding and decoding, the weights generated for the first component (such as the luma component) can be used to derive weights for the second component (such as the Cb or Cr component).
[0353] a) This derivation can depend on the color format (such as 4:2:0 or 4:2:2 or 4:4:4).
[0354] b) The weighted value of the second component can be derived by applying upsampling or downsampling to the weighted value of the first component.
[0355] 7. The weighting values generated for components (such as Cb or Cr components) can depend on the color format (such as 4:2:0, 4:2:2, or 4:4:4).
[0356] a) For example, when the color format is 4:2:2, the GEO angle / displacement / distance associated with the GEO mode can be adjusted to generate weighted values for components (such as Cb or Cr components).
[0357] The examples described above can be incorporated into the context of the methods described below (e.g., method 800), which can be implemented at a video decoder or video encoder.
[0358] Figure 8 A flowchart of an example method 800 for video processing is shown. The method includes, in operation 810, determining, for the conversion between a current block of video and a bitstream representation of the video, the enabling of a first codec mode and a second codec mode different from the first codec mode, wherein the first codec mode divides the current block into two or more sub-regions comprising at least one non-rectangular or non-square sub-region.
[0359] The method includes performing a transformation based on the determination in operation 820.
[0360] In some embodiments, the following technical solutions may be implemented:
[0361] A1. A method for video processing, comprising: performing a conversion between a current block of video and a bitstream representation of the video, wherein a codec mode of the current block divides the current block into two or more sub-regions including at least one non-rectangular or non-square sub-region, wherein the bitstream representation includes signaling associated with the codec mode, and wherein the signaling corresponds to a set of parameters having a first set of values for the current block of video and a second set of values for subsequent blocks.
[0362] A2. The method according to solution A1, wherein the signaling includes an index, and wherein binarization of the index includes variable-length encoding and decoding using a first number of binary bits for a first value of the index and a second number of binary bits for a second value of the index.
[0363] A3. The method according to solution A1, wherein the signaling includes an index, and wherein the index is based on the height or width of the current block.
[0364] A4. The method according to solution A1, wherein the parameter set includes multiple angles and multiple distances of at least one non-rectangular or non-square sub-region.
[0365] A5. The method described in solution A4, wherein multiple angles are asymmetrical.
[0366] A6. The method described in solution A4, wherein multiple angles are not rotationally symmetric.
[0367] A7. The method described in solution A4, wherein multiple angles are not bilaterally symmetrical.
[0368] A8. The method described in solution A1, wherein the position of signaling in the bitstream representation is based on the priority of the encoding / decoding mode.
[0369] A9. The method according to solution A1, wherein the signaling includes an indication of the type of encoding / decoding mode and other information related to the encoding / decoding mode.
[0370] A10. The method according to solution A9, wherein the other information includes an angle, and wherein the indication includes an indication of the clockwise or counterclockwise direction of the angle.
[0371] A11. A method for video processing, comprising: performing a conversion between a current block of video and a bitstream representation of the video, wherein a coding / decoding mode of the current block divides the current block into two or more sub-regions including at least one non-rectangular or non-square sub-region, wherein the coding / decoding mode can be configured using a plurality of parameter sets, and wherein the bitstream representation includes signaling for a subset of the plurality of parameter sets, and wherein the parameter sets include angles, displacements, and distances associated with at least one non-rectangular or non-square sub-region.
[0372] A12. The method described in solution A11, wherein the encoding / decoding mode uses only a subset of multiple parameter sets.
[0373] A13. The method according to solution A11, wherein the selection of a parameter set from a subset of multiple parameter sets is based on syntax elements in the bitstream representation or one or more dimensions of the current block.
[0374] A14. The method according to solution A11, wherein the selection of a parameter set from a subset of multiple parameter sets is based on the enabling of low-latency B (LDB) frame encoding / decoding tools.
[0375] A15. The method according to solution A11, wherein the selection of a parameter set from a subset of multiple parameter sets is based on the derivation of motion vector candidates.
[0376] A16. The method according to solution A15, wherein the motion vector candidate is derived from a temporal motion vector prediction (TMVP) candidate, a spatial motion candidate, or a history-based motion vector prediction (HMVP) candidate.
[0377] A17. The method according to solution A11, wherein, when determining that the ratio of the height to the width of the current block is 1, a subset of the plurality of parameter sets includes encoding / decoding modes corresponding to angles of 45°, 135°, 225° and / or 315°.
[0378] A18. The method according to solution A11, wherein a subset of the plurality of parameter sets includes the encoding / decoding mode corresponding to the current block as divided by (a) a line connecting the upper left and lower right corners of the current block or (b) a line connecting the upper right and lower left corners of the current block.
[0379] A19. A method for video processing, comprising: for a conversion between a current block of a video and a bitstream representation of the video, determining the enabling of a first codec mode and a second codec mode different from the first codec mode, wherein the first codec mode divides the current block into two or more sub-regions including at least one non-rectangular or non-square sub-region; and performing the conversion based on the determination.
[0380] A20. The method according to solution A19, wherein the second codec mode includes weighted prediction, and wherein the first codec mode is disabled at the video unit level and the second codec mode is enabled at the stripe level.
[0381] A21. The method described according to solution A20, wherein the video unit level is a strip level, a picture parameter set (PPS) level, a sequence parameter set (SPS) level, a slice level, a subpicture level, a codec unit (CU) level, a prediction unit (PU) level, or a transform unit (TU) level.
[0382] A22. The method according to solution A19, wherein the second encoding / decoding mode is bidirectional prediction (BCW) using codec unit (CU) weights, prediction refinement (PROF) using optical flow, bidirectional optical flow (BDOF) mode, decoder-side motion vector refinement (DMVR) mode, or subblock transform (SBT) mode.
[0383] A23. The method according to solution A19, wherein the second codec mode includes a deblocking process, and wherein enabling the second codec mode is based on enabling the first codec mode.
[0384] A24. The method according to solution A19, wherein the second encoding / decoding mode includes weighted prediction, and wherein the weighted value of the weighted prediction of the component is based on the color format of the current block.
[0385] A25. The method according to solution A24, wherein the component is a Cb component or a Cr component, and wherein the color format is 4:2:0, 4:2:2, or 4:4:4:4.
[0386] A26. The method according to any one of solutions A1 to A25, wherein the transformation generates the current block from the bitstream representation.
[0387] A27. The method according to any one of solutions A1 to A25, wherein the transformation generates a bitstream representation from the current block.
[0388] A28. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, when executed by the processor, the instructions cause the processor to perform the method according to any one of solutions A1 to A27.
[0389] A29. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method according to any one of solutions A1 to A27.
[0390] Figure 9 This is a block diagram of a video processing apparatus 900. Apparatus 900 can be used to implement one or more methods described herein. Apparatus 900 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 900 may include one or more processors 902, one or more memories 904, and video processing hardware 906. Processor 902 can be configured to implement one or more methods described in this document. Memory 904 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 906 can be used to implement some of the techniques described in this document in a hardware circuit system.
[0391] Figure 10 This is a block diagram illustrating an example video codec system 300 that can utilize the techniques disclosed herein.
[0392] like Figure 10 As shown, the video encoding / decoding system 300 may include a source device 310 and a destination device 320. The source device 310 generates encoded video data, and this source device 310 may be referred to as a video encoding device. The destination device 320 can decode the encoded video data generated by the source device 310, and this destination device 320 may be referred to as a video decoding device.
[0393] The source device 310 may include a video source 312, a video encoder 314, and an input / output (I / O) interface 316.
[0394] Video source 312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 314 encodes the video data from video source 312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 320 via network 330a through I / O interface 316. Encoded video data may also be stored on storage medium / server 330b for access by destination device 320.
[0395] The target device 320 may include an I / O interface 326, a video decoder 324, and a display device 322.
[0396] I / O interface 326 may include a receiver and / or a modem. I / O interface 326 may acquire encoded video data from source device 310 or storage medium / server 330b. Video decoder 324 may decode the encoded video data. Display device 322 may display the decoded video data to a user. Display device 322 may be integrated with destination device 320, or it may be external to destination device 320, which is configured to interface with an external display device.
[0397] The video encoder 314 and the video decoder 324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.
[0398] Figure 11 This is a block diagram illustrating an example of a video encoder 400, which can be... Figure 10 The video encoder 314 in the system 300 shown.
[0399] The video encoder 400 can be configured to perform any or all of the technologies disclosed herein. Figure 11 In the example, the video encoder 400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 400. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0400] The functional components of the video encoder 400 may include a segmentation unit 401, a prediction unit 402 (which may include a mode selection unit 403, a motion estimation unit 404, a motion compensation unit 405, and an intra-frame prediction unit 406), a residual generation unit 407, a transform unit 408, a quantization unit 409, an inverse quantization unit 410, an inverse transform unit 411, a reconstruction unit 412, a buffer 413, and an entropy coding unit 414.
[0401] In other examples, the video encoder 400 may include more, fewer, or different functional components. In one example, the prediction unit 402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0402] Furthermore, some components, such as the motion estimation unit 404 and the motion compensation unit 405, can be highly integrated, but for explanatory purposes, in Figure 11 The examples are shown separately.
[0403] The segmentation unit 401 can segment an image into one or more video blocks. The video encoder 400 and the video decoder 500 can support various video block sizes.
[0404] The mode selection unit 403 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 407 to generate residual block data, and to the reconstruction unit 412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 403 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 403 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).
[0405] To perform inter-frame prediction on the current video block, motion estimation unit 404 can generate motion information for the current video block by comparing one or more reference frames from buffer 413 with the current video block. Motion compensation unit 405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 413 other than the image associated with the current video block.
[0406] The motion estimation unit 404 and the motion compensation unit 405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0407] In some examples, motion estimation unit 404 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 404 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0408] In other examples, motion estimation unit 404 can perform bidirectional prediction on the current video block. Motion estimation unit 404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 404 can then generate a reference index indicating the reference images in lists 0 and 1 containing the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0409] In some examples, the motion estimation unit 404 can output a complete set of motion information for use in the decoder's decoding process.
[0410] In some examples, the motion estimation unit 404 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 404 may refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0411] In one example, the motion estimation unit 404 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 500 that the current video block has the same motion information as another video block.
[0412] In another example, motion estimation unit 404 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0413] As described above, the video encoder 400 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 400 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling notification.
[0414] Intra-prediction unit 406 can perform intra-prediction on the current video block. When intra-prediction unit 406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0415] The residual generation unit 407 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0416] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 407 may not perform the subtraction operation.
[0417] The transform processing unit 408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0418] After the transform processing unit 408 generates a transform coefficient video block associated with the current video block, the quantization unit 409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0419] The inverse quantization unit 410 and the inverse transform unit 411 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 402 to produce a reconstructed video block associated with the current block, which is stored in the buffer 413.
[0420] After the video block is reconstructed by reconstruction unit 412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0421] The entropy encoding unit 414 can receive data from other functional components of the video encoder 400. When the entropy encoding unit 414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0422] Figure 12 This is a block diagram illustrating an example of a video decoder 500, which can be... Figure 10 The video decoder 314 in the system 300 shown.
[0423] The video decoder 500 can be configured to perform any or all of the technologies disclosed herein. Figure 12 In the example, the video decoder 500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 500. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0424] exist Figure 12 In the example, the video decoder 500 includes an entropy decoding unit 501, a motion compensation unit 502, an intra-frame prediction unit 503, an inverse quantization unit 504, an inverse transform unit 505, a reconstruction unit 506, and a buffer 507. In some examples, the video decoder 500 can perform functions typically associated with the video encoder 400. Figure 11 The encoding process described is the opposite of the decoding process.
[0425] Entropy decoding unit 501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 501 can decode the entropy-coded video data, and based on the entropy-coded video data, motion compensation unit 502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 502 can determine such information, for example, by executing AMVP and merge modes.
[0426] The motion compensation unit 502 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.
[0427] The motion compensation unit 502 can use an interpolation filter, such as that used by the video encoder 400 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 502 can determine the interpolation filter used by the video encoder 400 based on the received syntax information, and use the interpolation filter to generate the prediction block.
[0428] The motion compensation unit 502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0429] Intra-prediction unit 503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 503 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 501. Inverse transform unit 503 applies an inverse transform.
[0430] The reconstruction unit 506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 502 or the intra-frame prediction unit 503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 507 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates the decoded video for presentation on the display device.
[0431] Figure 13 This is a block diagram illustrating an example video processing system 1300 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1300. System 1300 may include an input 1302 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1302 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0432] System 1300 may include a codec component 1304 capable of implementing the various codec or encoding methods described in this document. Codec component 1304 can reduce the average bit rate of the video from input 1302 to the output of codec component 1304 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 1304 may be stored or transmitted via a communication connection, as represented by component 1306. The bitstream (or codec) representation of the video received at input 1302, whether stored or communicated, can be used by component 1308 to generate pixel values or transmit as displayable video to display interface 1310. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.
[0433] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0434] Figure 14A flowchart of an example method for video processing is shown. The method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded / decoded in a geometric segmentation mode (1402), wherein the geometric segmentation mode includes the entire set of segmentation modes; determining one or more segmentation modes of the current video block based on segmentation mode indices included in the bitstream (1404), wherein the segmentation mode indices correspond to different segmentation modes from one video block to another; and performing a conversion based on the one or more segmentation modes (1406).
[0435] In some examples, each segmentation pattern in the entire set of segmentation patterns divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are non-square and non-rectangular.
[0436] In some examples, each segmentation pattern index is associated with a set of parameters including at least one of angle, distance, and / or displacement.
[0437] In some examples, the correspondence between the segmentation pattern index and the segmentation pattern depends on the dimensions of the video block, where the dimensions of the video block include at least one of the video block's height, width, and the ratio of its width to its height.
[0438] In some examples, the binarization of the segmentation pattern index includes a variable-length encoding / decoding that uses a first number of bits for a first segmentation pattern and a second number of bits for a second segmentation pattern.
[0439] In some examples, different numbers of segmentation patterns or different numbers of parameters are used in different video blocks.
[0440] In some examples, the number of segmentation patterns or parameters used in a video block depends on the dimensions of the video block.
[0441] In some examples, the number of angles used in the current video block is greater than the number of angles used in the second video block, where the current video block is a video block with a height greater than its width, and the second video block is a video block with a height less than its width.
[0442] In some examples, the angles are asymmetrical.
[0443] In some examples, the angles are not rotationally symmetric.
[0444] In some examples, the angles are not bilaterally symmetrical.
[0445] In some examples, the angles are not symmetrical.
[0446] In some examples, a video block includes at least one of a codec block, a codec unit, a virtual pipeline data unit (VPDU), a slice, a strip, a picture, a subpicture, a tile, or a video.
[0447] In some examples, the representation of the segmentation pattern in the bitstream depends on the priority of the segmentation pattern.
[0448] In some examples, priority is determined by the angle and / or distance associated with the segmentation pattern.
[0449] In some examples, segmentation patterns oriented towards angles smaller than a predetermined value have higher priority than segmentation patterns oriented towards angles larger than a predetermined value.
[0450] In some examples, the predetermined value is 45 or 90.
[0451] In some examples, segmentation patterns associated with angle indices less than a predetermined value have higher priority than segmentation patterns associated with angle indices greater than a predetermined value.
[0452] In some examples, the predetermined value is 4, 8, or 16.
[0453] In some examples, when signaling notifications are made, a higher-priority segmentation pattern requires fewer bits or bytes than a lower-priority segmentation pattern.
[0454] In some examples, the segmentation pattern is classified into two or more categories, and the indication of which category the segmentation pattern belongs to is signaled before other information related to the segmentation pattern.
[0455] In some examples, the signaling notification for video blocks is first indexed by the angle of whether it is clockwise or counterclockwise.
[0456] In some examples, when the angle index is signaled in a counter-clockwise direction, the counter-clockwise signaled angle index means a smaller angle index, and a smaller angle index means a smaller angle size.
[0457] In some examples, if angleIdx = 0 means the angle size is 0 degrees, then angleIdx = 8 means the angle size is 90 degrees; angleIdx = 16 means the angle size is 180 degrees, and angleIdx = 24 means the angle size is 270 degrees.
[0458] In some examples, when the angle index is signaled in a clockwise direction, the clockwise signaled angle index implies a larger angle index, which in turn indicates a larger angle size.
[0459] In some examples, if angleIdx = 0 means the angle size is 0 degrees, then angleIdx = 8 means the angle size is 270 degrees; angleIdx = 16 means the angle size is 180 degrees, and angleIdx = 24 means the angle size is 90 degrees.
[0460] In some examples, this instruction is signaled at the video block level.
[0461] In some examples, the instruction is signaled at at least one of the following levels: SPS, VPS, PPS, picture header, subpicture, strip, strip header, slice, tile, CTU, VPDU, CU, and block level.
[0462] In some examples, higher-level flags above the block level are signaled to indicate whether the angle associated with the segmentation mode of the signaling notification is clockwise or counterclockwise.
[0463] In some examples, block-level flags are signaled to indicate whether the angle associated with the segmentation mode of the signaling notification is clockwise or counterclockwise.
[0464] Figure 15 A flowchart of an example method for video processing is shown. The method includes: for a conversion between a current video block and a bitstream of the current video block, determining that the video block is encoded and decoded in a geometric segmentation pattern (1502), wherein the geometric segmentation pattern includes an entire set of segmentation patterns, and each segmentation pattern is associated with a set of parameters including at least one of angle, distance, and / or displacement; deriving a subset of the segmentation patterns or parameters from the entire set of segmentation patterns or parameters (1504); and performing a conversion based on the subset of segmentation patterns or parameters (1506).
[0465] In some examples, each segmentation pattern in the entire set of segmentation patterns divides the current video block into two or more partitions, at least one of which is neither square nor rectangular.
[0466] In some examples, only a subset of the segmentation pattern is used for video blocks.
[0467] In some examples, segmentation patterns associated only with a subset of parameters are used for video blocks.
[0468] In some examples, an indication of whether the selected segmentation mode or parameter is within a subset is also included in the bitstream.
[0469] In some examples, whether a subset or the entire set of segmentation patterns or parameters is used for a video block depends on the decoding information of the current video block or one or more previously decoded video blocks.
[0470] In some examples, the decoding information includes syntax elements and / or the block dimension of video blocks.
[0471] In some examples, the choice of segmentation pattern within a subset of segmentation patterns depends on the corresponding angle.
[0472] In some examples, a subset of the segmentation patterns contains only those associated with a distance or displacement equal to 0.
[0473] In some examples, a subset of the segmentation patterns contains only the segmentation patterns associated with a specified angle.
[0474] In some examples, the segmentation pattern associated with a specified angle includes segmentation patterns associated with a predefined subset of angles and all displacement combinations corresponding to those predefined angles.
[0475] In some examples, the selection of a segmentation mode within a subset of segmentation modes depends on whether low-latency B-frame (LDB) encoding / decoding is checked.
[0476] In some examples, different subsets of the segmentation pattern are used for LDB encoding / decoding and random access (RA) encoding / decoding.
[0477] In some examples, the selection of a segmentation pattern from a subset of segmentation patterns depends on reference images in the list of reference images for the current image.
[0478] In some examples, different subsets of the segmentation pattern are used in the first case where all reference images are displayed before the current image, and in the second case where at least one reference image is displayed after the current image.
[0479] In some examples, the selection of a segmentation pattern from a subset of segmentation patterns depends on the derivation of motion vector candidates.
[0480] In some examples, different subsets of the segmentation pattern are used when deriving motion vector candidates from temporal motion vector prediction (TMVP) candidates, spatial motion candidates, or history-based motion vector prediction (HMVP) candidates.
[0481] In some examples, a subset of the splitting patterns contains only splitting patterns that divide blocks in the same way as the TPM pattern.
[0482] In some examples, a subset of the segmentation patterns contains only segmentation patterns that divide blocks by lines connecting the top left and bottom right corners of the block or by lines connecting the top right and bottom left corners of the block.
[0483] In some examples, a subset of the segmentation patterns contains only segmentation patterns corresponding to diagonal angles with one or more distance or displacement indices.
[0484] In some examples, the diagonal angle indication corresponds to a segmentation pattern in which the block's boundary is defined by either the line connecting the top-left and bottom-right corners of the block or the line connecting the top-right and bottom-left corners of the block.
[0485] In some examples, a subset of the segmentation patterns contains only those associated with a distance or displacement equal to 0.
[0486] In some examples, a subset of the segmentation patterns contains only the segmentation patterns corresponding to any angle associated with a distance or displacement equal to 0.
[0487] In some examples, a subset of the segmentation patterns contains only the segmentation patterns corresponding to the diagonal angles associated with a distance or displacement equal to 0.
[0488] In some examples, a subset of the segmentation patterns contains only the segmentation patterns corresponding to diagonal angles, which are associated with all distance or displacement indices corresponding to those angles.
[0489] In some examples, for blocks with an aspect ratio equal to X, a subset of the segmentation patterns contains only the segmentation patterns corresponding to angular dimensions equal to arctan(X) and / or π-arctan(X) and / or π+arctan(X) and / or 2π-arctan(X) and all distance indices corresponding to these angles, where X is an integer.
[0490] In some examples, X = 1.
[0491] In some examples, horizontal and / or vertical angles are included in a subset of angles.
[0492] In some examples, the horizontal angle indication corresponds to an angle index of 90° and / or 270°.
[0493] In some examples, the vertical angle indication corresponds to the angle index of 0° and / or 180°.
[0494] In some examples, the segmentation patterns associated with horizontal and / or vertical angles that sum to 0 are included in a subset of the allowed angles.
[0495] In some examples, the segmentation patterns associated with horizontal and / or vertical angles that sum to zero are not included in the subset of allowed angles.
[0496] In some examples, the segmentation patterns associated with horizontal and / or vertical angles combined with all distance or displacement indices are included in a subset of the allowed angles.
[0497] In some examples, the segmentation patterns associated with horizontal and / or vertical angles combined with all distance or displacement indices are not included in the subset of allowed angles.
[0498] In some examples, the indication of whether to use a segmentation mode or a subset of parameters is signaled at at least one of the SPS, VPS, PPS, picture header, subpicture, strip, strip header, slice, tile, CTU, VPDU, CU, and block levels.
[0499] In some examples, the indication of which subset of the segmentation mode or parameters is used is further signaled at at least one of the SPS, VPS, PPS, picture header, subpicture, strip, strip header, slice, tile, CTU, VPDU, CU, and block levels.
[0500] Figure 16 A flowchart of an example method for video processing is shown. The method includes: for a conversion between a video block and a bitstream of the video block, determining the enabling of a geometric segmentation mode and a second codec mode different from the geometric segmentation mode for the video block (1602); and performing the conversion based on the determination (1604).
[0501] In some examples, the geometric segmentation pattern includes an entire set of segmentation patterns, each of which divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are non-square and non-rectangular.
[0502] In some examples, the second codec mode includes weighted prediction.
[0503] In some examples, when weighted prediction is enabled, geometric segmentation mode is disabled at the video tile level.
[0504] In some examples, when weighted prediction is enabled at the strip level, the geometric segmentation mode is disabled at at least one of the strip, PPS, SPS, slice, sub-picture, CU, PU, or TU levels.
[0505] In some examples, whether a geometric segmentation pattern is used with weighted prediction depends on the weighting factor of the weighted prediction.
[0506] In some examples, the geometric segmentation mode is disabled if the weighting factor of the weighted prediction is greater than T, where T is a constant value.
[0507] In some examples, the second codec mode includes bidirectional prediction (BCW) using codec unit (CU) weights, prediction refinement (PROF) mode using optical flow, bidirectional optical flow (BDOF) mode, decoder-side motion vector refinement (DMVR) mode, or subblock transform (SBT) mode.
[0508] In some examples, the second codec mode is disabled when the geometry segmentation mode is enabled.
[0509] In some examples, when the geometry segmentation mode is enabled, no signaling is sent to indicate the second codec mode.
[0510] In some examples, the geometric segmentation mode is disabled when the second codec mode is enabled.
[0511] In some examples, when the second codec mode is enabled, no signaling is given to indicate the geometric segmentation mode.
[0512] In some examples, the second codec mode is disabled when weighted prediction is enabled.
[0513] In some examples, the deblocking process associated with a video block depends on whether both the geometric segmentation mode and the second codec mode are enabled.
[0514] Figure 17 A flowchart of an example method for video processing is shown. The method includes: for a video block and a bitstream of the video block, determining a deblocking process associated with the video block (1702) based on whether the current video block is encoded or decoded in the current video block's geometric segmentation mode and / or color format; and performing the conversion based on the deblocking parameters (1704).
[0515] In some examples, the geometric segmentation pattern includes an entire set of segmentation patterns, each of which divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are non-square and non-rectangular.
[0516] In some examples, when encoding and decoding video blocks in a geometric segmentation mode, the weighted predictions generated for the first component of the video block are used to derive weighted values for the second component of the video block.
[0517] In some examples, the first component is the luminance component, and the second component is the Cb or Cr component.
[0518] In some examples, the derivation depends on the color format of the video block.
[0519] In some examples, the color format is 4:2:0, 4:2:2, or 4:4:4.
[0520] In some examples, the weighted value of the second component is derived by applying upsampling or downsampling to the weighted value of the first component.
[0521] In some examples, the weights for the weighted prediction of the components are based on the color format of the video blocks.
[0522] In some examples, the component is a Cb component or a Cr component, and the color format is 4:2:0, 4:2:2, or 4:4:4.
[0523] In some examples, when the color format is 4:2:2, the parameters associated with the segmentation mode are adjusted to generate weighted values for the components.
[0524] In some examples, the geometric segmentation patterns include one or more of the following: geometric merge pattern, geometric partitioning pattern, wedge prediction pattern, and triangle prediction pattern.
[0525] In some examples, the conversion involves encoding video blocks into a bitstream.
[0526] In some examples, the conversion involves decoding video blocks from a bitstream.
[0527] In some examples, the conversion includes generating a bitstream from video blocks; the method also includes storing the bitstream in a non-transitory computer-readable recording medium.
[0528] Figure 17 A flowchart of an example method for video processing is shown. The method includes: for a conversion between a current video block and a bitstream of the video, determining that the current video block is encoded and decoded in a geometric segmentation mode (1702), wherein the geometric segmentation mode includes the entire set of segmentation modes; determining one or more segmentation modes of the current video block based on segmentation mode indices included in the bitstream (1704), wherein the segmentation mode indices correspond to different segmentation modes from one video block to another; generating a bitstream from the video block in one or more segmentation modes (1706); and storing the bitstream in a non-transitory computer-readable recording medium (1708).
[0529] As can be understood from the foregoing, specific embodiments of the currently disclosed technology have been described herein for illustrative purposes; however, various modifications may be made without departing from the scope of the invention. Therefore, the currently disclosed technology is not limited except for the appended claims.
[0530] The embodiments of this subject matter and the functional operations described in this patent document can be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-volatile computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition that influences machine-readable propagated signals, or a combination thereof. The terms "data processing unit" or "data processing apparatus" include all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors or computer groups. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0531] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed and executed on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.
[0532] The processes and logic flows described in this application can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuits, and the apparatus can also be implemented as special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0533] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as one or more of any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0534] While this patent document contains numerous details, it should not be construed as limiting the scope of any invention or claim, but rather as a description of features of specific embodiments of a particular invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in certain circumstances, one or more features from a combination of claims may be removed from the combination, and a combination of claims may refer to a sub-combination or a variation of a sub-combination.
[0535] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequence shown to perform such operations, or all the described operations, in order to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0536] Only some implementations and examples are described. Other implementations, enhancements and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, it is determined that the current video block is encoded and decoded in a geometric segmentation mode, wherein the geometric segmentation mode includes the entire set of segmentation modes; One or more segmentation modes for the current video block are determined based on segmentation mode indices included in the bitstream, wherein the segmentation mode indices correspond to different segmentation modes from one video block to another; and Perform the transformation based on one or more segmentation patterns; The number of segmentation modes or parameters used in a video block depends on the dimensions of the video block.
2. The method according to claim 1, wherein, Each segmentation pattern in the entire set of segmentation patterns divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are neither square nor rectangular.
3. The method according to claim 1 or 2, wherein, Each segmentation pattern index is associated with a set of parameters including at least one of angle, distance, and / or displacement.
4. The method according to claim 3, wherein, The correspondence between the segmentation pattern index and the segmentation pattern depends on the dimensions of the video block, where the dimensions of the video block include at least one of the video block's height, width, and the ratio of width to height.
5. The method according to any one of claims 1-4, wherein, Binarization of the segmentation pattern index includes variable-length encoding / decoding, which uses a first number of bits for a first segmentation pattern and a second number of bits for a second segmentation pattern.
6. The method according to claim 4, wherein, Different numbers of segmentation patterns or different numbers of parameters are used in different video blocks.
7. The method according to claim 1, wherein, The number of angles used in the current video block is greater than the number of angles used in the second video block, where the current video block is a video block with a height greater than its width, and the second video block is a video block with a height less than its width.
8. The method according to claim 7, wherein, The angles mentioned are asymmetrical.
9. The method according to claim 7, wherein, The angle is not rotationally symmetric.
10. The method according to claim 7, wherein, The angle is not bilaterally symmetrical.
11. The method according to claim 7, wherein, The angle is not symmetrical.
12. The method according to claim 11, wherein, The video block includes at least one of the following: codec block, codec unit, virtual pipeline data unit (VPDU), slice, strip, picture, sub-picture, tile, or video.
13. The method according to claim 1, wherein, The representation of a segmentation pattern in a bitstream depends on the priority of the segmentation pattern.
14. The method according to claim 13, wherein, The priority is determined by the angle and / or distance associated with the segmentation pattern.
15. The method according to claim 14, wherein, Segmentation patterns oriented towards angles smaller than a predetermined value have higher priority than segmentation patterns oriented towards angles larger than a predetermined value.
16. The method according to claim 15, wherein, The predetermined value is 45 degrees or 90 degrees.
17. The method of claim 14, wherein, Segmentation patterns associated with angle indices less than a predetermined value have higher priority than segmentation patterns associated with angle indices greater than a predetermined value.
18. The method according to claim 17, wherein, The predetermined value is 4, 8, or 16.
19. The method according to claim 13, wherein, When signaling, higher priority segmentation patterns require fewer bits or bytes than lower priority segmentation patterns.
20. The method according to claim 1, wherein, The segmentation pattern is classified into two or more categories, and the indication of which category the segmentation pattern belongs to is signaled before other information related to the segmentation pattern.
21. The method according to claim 20, wherein, First, the signaling notification for video blocks is indexed by whether it is clockwise or counterclockwise.
22. The method according to claim 20, wherein, When the angle index is signaled in a counter-clockwise direction, a smaller angle index indicates a smaller angle size.
23. The method according to claim 22, wherein, If angleIdx = 0 means the angle size is 0 degrees, then angleIdx = 8 means the angle size is 90 degrees; angleIdx = 16 means the angle size is 180 degrees, and angleIdx = 24 means the angle size is 270 degrees.
24. The method of claim 20, wherein, If angleIdx = 0 means the angle size is 0 degrees, then angleIdx = 8 means the angle size is 270 degrees; angleIdx = 16 means the angle size is 180 degrees, and angleIdx = 24 means the angle size is 90 degrees. Here, angleIdx is used to represent the angle index.
25. The method according to claim 20, wherein, The instruction is signaled at the video block level.
26. The method of claim 25, wherein, The instruction is signaled at at least one of the following levels: SPS, VPS, PPS, image header, sub-image, strip, strip header, slice, tile, CTU, VPDU, CU, and block level.
27. The method according to claim 26, wherein, Higher-level flags above the block level are signaled to indicate whether the angle associated with the segmentation mode of the signaling notification is clockwise or counterclockwise.
28. The method according to claim 26, wherein, Block-level flags are signaled to indicate whether the angle associated with the segmentation mode of the signaling notification is clockwise or counterclockwise.
29. The method according to claim 1, further comprising: For the conversion between the current video block and the bitstream of the current video block, it is determined that the current video block is encoded and decoded in a geometric segmentation mode, wherein the geometric segmentation mode includes a whole set of segmentation modes comprising multiple subsets of segmentation modes, and each segmentation mode is associated with a set of parameters including at least one of angle, distance and / or displacement. Select a subset of segmentation modes or parameters from the entire set of segmentation modes or parameters for the current video block; and Transformation is performed based on a subset of the segmentation pattern or parameters.
30. The method according to claim 29, wherein, Each segmentation pattern in the entire set of segmentation patterns divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are neither square nor rectangular.
31. The method according to claim 29, wherein, Only a subset of the segmentation mode is used for video blocks.
32. The method according to claim 29, wherein, Segmentation patterns associated only with a subset of parameters are used for video blocks.
33. The method according to claim 29, wherein, An indication of whether the selected segmentation mode or parameter is within a subset is also included in the bitstream.
34. The method according to claim 29, wherein, Whether the segmentation mode is a subset of the segmentation mode or a subset of the parameters, or the entire set of segmentation modes or the entire set of parameters, is used for a video block depends on the decoding information of the current video block or one or more previously decoded video blocks.
35. The method according to claim 34, wherein, The decoding information includes syntax elements and / or the block dimension of video blocks.
36. The method according to claim 29, wherein, The selection of a segmentation pattern from a subset of segmentation patterns depends on the angle corresponding to the segmentation pattern in the subset of segmentation patterns.
37. The method of claim 35, wherein, A subset of the segmentation patterns contains only those segmentation patterns associated with a distance or displacement equal to 0.
38. The method according to claim 36, wherein, A subset of segmentation patterns contains only those segmentation patterns associated with a specified angle.
39. The method according to claim 38, wherein, The segmentation patterns associated with a specified angle include segmentation patterns associated with a predefined subset of angles and all displacement combinations corresponding to these predefined angles.
40. The method according to claim 29, wherein, The selection of a segmentation mode within a subset of segmentation modes depends on whether low-latency B-frame (LDB) encoding / decoding is checked.
41. The method according to claim 40, wherein, Different subsets of the segmentation pattern are used for LDB encoding / decoding and random access (RA) encoding / decoding.
42. The method according to claim 29, wherein, The selection of a segmentation pattern from a subset of segmentation patterns depends on the reference images in the reference image list for the current image.
43. The method according to claim 42, wherein, In the first case where all reference images are displayed before the current image, and in the second case where at least one reference image is displayed after the current image, different subsets of the segmentation pattern are used.
44. The method according to claim 29, wherein, The selection of a segmentation pattern from a subset of segmentation patterns depends on the derivation of motion vector candidates.
45. The method according to claim 44, wherein, When deriving motion vector candidates from temporal motion vector prediction (TMVP) candidates, spatial motion candidates, or history-based motion vector prediction (HMVP) candidates, different subsets of the segmentation pattern are used respectively.
46. The method of claim 36, wherein, A subset of the partitioning patterns contains only partitioning patterns that divide blocks in the same way as the TPM pattern.
47. The method according to claim 46, wherein, A subset of the segmentation patterns contains only segmentation patterns that divide blocks by connecting the top left and bottom right corners of the block or by connecting the top right and bottom left corners of the block.
48. The method according to claim 36, wherein, A subset of the segmentation patterns contains only segmentation patterns that correspond to diagonal angles with one or more distance or displacement indices.
49. The method according to claim 48, wherein, The diagonal angle indicator corresponds to the segmentation mode that divides the block boundary by lines connecting the top left and bottom right corners of the block, or by lines connecting the top right and bottom left corners of the block.
50. The method according to claim 48, wherein, A subset of the segmentation patterns contains only those segmentation patterns associated with a distance or displacement equal to 0.
51. The method according to claim 50, wherein, A subset of the segmentation patterns contains only the segmentation patterns corresponding to any angle associated with a distance or displacement equal to 0.
52. The method according to claim 50, wherein, A subset of the segmentation patterns contains only the segmentation patterns corresponding to the diagonal angles associated with a distance or displacement equal to 0.
53. The method according to claim 48, wherein, A subset of the segmentation patterns contains only the segmentation patterns corresponding to diagonal angles, which are associated with all distance or displacement indices corresponding to those angles.
54. The method according to claim 53, wherein, For a block with an aspect ratio equal to X, a subset of the segmentation patterns contains only the segmentation patterns corresponding to angular dimensions equal to arctan(X) and / or π-arctan(X) and / or π+arctan(X) and / or 2π-arctan(X) and all distance indices corresponding to these angles, where X is an integer.
55. The method according to claim 54, wherein, X=1。 56. The method according to claim 55, wherein, Horizontal and / or vertical angles are included in a subset of angles.
57. The method according to claim 56, wherein, The horizontal angle indicator corresponds to an angle index of 90° and / or 270°.
58. The method according to claim 56, wherein, The vertical angle indicator corresponds to the angle index of 0° and / or 180°.
59. The method according to claim 56, wherein, Segmentation patterns associated with horizontal and / or vertical angles and combined with a distance / displacement of 0 are included in a subset of allowed angles.
60. The method of claim 56, wherein, Segmentation patterns associated with horizontal and / or vertical angles and combined with a distance or displacement equal to 0 are not included in the subset of allowed angles.
61. The method according to claim 56, wherein, Segmentation patterns associated with horizontal and / or vertical angles, and combined with all distance or displacement indices, are included in a subset of allowed angles.
62. The method according to claim 56, wherein, Segmentation patterns associated with horizontal and / or vertical angles, and combined with all distance or displacement indices, are not included in the subset of allowed angles.
63. The method according to claim 29, wherein, Whether to use a segmentation mode or a subset of parameters is indicated by signaling at at least one of the following levels: SPS, VPS, PPS, image header, sub-image, strip, strip header, slice, tile, CTU, VPDU, CU, and block level.
64. The method according to claim 63, wherein, The indication of which subset of the segmentation mode or parameters is used is further signaled at at least one of the following levels: SPS, VPS, PPS, image header, sub-image, strip, strip header, slice, tile, CTU, VPDU, CU, and block level.
65. The method according to claim 1, further comprising: The conversion between video blocks and the bitstream of the video blocks determines the enabling of a geometric segmentation mode and a second encoding / decoding mode different from the geometric segmentation mode for the video blocks; as well as The conversion is performed based on the determination.
66. The method according to claim 65, wherein, The geometric segmentation pattern includes the entire set of segmentation patterns. Each segmentation pattern in the entire set of segmentation patterns divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are non-square and non-rectangular.
67. The method according to claim 66, wherein, The second encoding / decoding mode includes weighted prediction.
68. The method according to claim 67, wherein, When weighted prediction is enabled, geometric segmentation mode is disabled at the video tile level.
69. The method according to claim 68, wherein, When weighted prediction is enabled at the strip level, the geometric segmentation mode is disabled at at least one of the strip, PPS, SPS, slice, sub-image, CU, PU, or TU levels.
70. The method of claim 67, wherein, Whether a geometric segmentation pattern can be used in conjunction with weighted prediction depends on the weighting factor of the weighted prediction.
71. The method according to claim 70, wherein, If the weighting factor of the weighted prediction is greater than T, the geometric segmentation mode is disabled, where T is a constant value.
72. The method according to claim 66, wherein, The second codec mode includes Bidirectional Prediction (BCW) using codec unit (CU) weights, Prediction Enhancement (PROF) mode using optical flow, Bidirectional Optical Flow (BDOF) mode, Decoder-Side Motion Vector Enhancement (DMVR) mode, or Subblock Transform (SBT) mode.
73. The method according to claim 68, wherein, When the geometric segmentation mode is enabled, the second encoding / decoding mode is disabled.
74. The method according to claim 73, wherein, When the geometric segmentation mode is enabled, no signaling is sent to indicate the second codec mode.
75. The method according to claim 66, wherein, When the second codec mode is enabled, the geometric segmentation mode is disabled.
76. The method according to claim 75, wherein, When the second codec mode is enabled, no signaling is given to indicate the geometric segmentation mode.
77. The method of claim 66, wherein, When weighted prediction is enabled, the second codec mode is disabled.
78. The method according to claim 66, wherein, The deblocking process associated with video blocks depends on whether both the geometric segmentation mode and the second encoding / decoding mode are enabled.
79. The method according to claim 1, further comprising: For the conversion between video blocks and the bitstream of the video blocks, the deblocking process associated with the video block is determined based on whether the current video block is encoded and decoded in the current video block's geometric segmentation mode and / or color format; as well as The conversion is performed based on the block removal process.
80. The method according to claim 79, wherein, The geometric segmentation pattern includes the entire set of segmentation patterns. Each segmentation pattern in the entire set of segmentation patterns divides the current video block into two or more partitions, and the two or more partitions corresponding to at least one segmentation pattern are non-square and non-rectangular.
81. The method according to claim 80, wherein, When encoding and decoding a video block in a geometric segmentation mode, the weighted prediction weights generated for the first component of the video block are used to derive weights for the second component of the video block.
82. The method according to claim 81, wherein, The first component is the luminance component, and the second component is either the Cb or Cr component.
83. The method according to claim 82, wherein, The derivation depends on the color format of the video blocks.
84. The method according to claim 83, wherein, The color format is 4:2:0, 4:2:2, or 4:4:
4.
85. The method according to claim 81, wherein, The weighted value of the second component is derived by applying upsampling or downsampling to the weighted value of the first component.
86. The method according to claim 79, wherein, The weights for the weighted prediction of the components are based on the color format of the video blocks.
87. The method according to claim 86, wherein, The component is either a Cb component or a Cr component, and the color format is 4:2:0, 4:2:2, or 4:4:
4.
88. The method according to claim 87, wherein, When the color format is 4:2:2, the parameters associated with the segmentation mode are adjusted to generate weighted values for the components.
89. The method according to any one of claims 1-88, wherein, Geometric segmentation patterns include one or more of the following: geometric merge pattern, geometric segmentation pattern, wedge prediction pattern, and triangle prediction pattern.
90. The method according to any one of claims 1-88, wherein, The conversion involves encoding video blocks into a bitstream.
91. The method according to any one of claims 1-88, wherein, The conversion involves decoding video blocks from a bitstream.
92. The method according to any one of claims 1-88, wherein, The conversion includes generating a bitstream from video blocks; The method further includes: The bit stream is stored in a non-transitory computer-readable recording medium.
93. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When executed by the processor, the instructions cause the processor to: For the conversion between the current video block and the bitstream of the video, it is determined that the current video block is encoded and decoded in a geometric segmentation mode, wherein the geometric segmentation mode includes the entire set of segmentation modes; One or more segmentation modes for the current video block are determined based on the segmentation mode index included in the bitstream, wherein the segmentation mode index corresponds to different segmentation modes from one video block to another; and Perform the transformation based on one or more segmentation patterns; The number of segmentation modes or parameters used in a video block depends on the dimensions of the video block.
94. A non-transitory computer-readable medium storing instructions that cause a processor to: For the conversion between the current video block and the video bitstream, it is determined that the current video block is encoded and decoded in a geometric segmentation mode, wherein, Geometric segmentation patterns include the entire set of segmentation patterns; One or more segmentation modes for the current video block are determined based on the segmentation mode index included in the bitstream, wherein the segmentation mode index corresponds to different segmentation modes from one video block to another; and Perform the transformation based on one or more segmentation patterns; The number of segmentation modes or parameters used in a video block depends on the dimensions of the video block.
95. A non-transitory computer-readable medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: Determine that the current video block is encoded and decoded using a geometric segmentation mode, where the geometric segmentation mode includes the entire set of segmentation modes; One or more segmentation modes for the current video block are determined based on the segmentation mode index included in the bitstream, wherein the segmentation mode index corresponds to different segmentation modes from one video block to another; and Generate bitstreams from video blocks based on one or more segmentation patterns; The number of segmentation modes or parameters used in a video block depends on the dimensions of the video block.
96. A method for storing a bitstream of video, comprising: Determine that the current video block is encoded and decoded using a geometric segmentation mode, where the geometric segmentation mode includes the entire set of segmentation modes; One or more segmentation modes for the current video block are determined based on the segmentation mode index included in the bitstream, wherein the segmentation mode index corresponds to different segmentation modes from one video block to another; Generate bitstreams from video blocks using one or more segmentation modes; and Storing bitstreams in a non-transitory computer-readable recording medium; The number of segmentation modes or parameters used in a video block depends on the dimensions of the video block.