Segmentation derivation of geometric partitions

By introducing geometric segmentation mode and nonlinear polynomial model in video encoding and decoding technology, and combining template matching to determine coefficients, the problem of insufficient utilization of nonlinear characteristics of images in the prior art is solved, and the compression efficiency and prediction accuracy of video encoding are improved.

CN120188474APending Publication Date: 2025-06-20TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004693.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-18
Filing Date
2024-05-31
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When processing image/video data, existing video encoding and decoding technologies are difficult to effectively utilize the nonlinear characteristics of the image, resulting in low compression efficiency.

Method used

Geometric Partitioning Mode (GPM) is used in combination with nonlinear polynomial model, and the coefficients of the nonlinear polynomial model are determined through template matching, which is used to calculate weight w0, thereby achieving the weighted average prediction of the current block.

Benefits of technology

The compression efficiency and prediction accuracy of video encoding are improved, and data quantization errors are reduced by better utilizing the nonlinear characteristics of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120188474A_ABST
    Figure CN120188474A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure include an apparatus for video decoding that includes processing circuitry configured to receive encoding information indicating that a current block is encoded in a geometric partition mode (GPM) using a first prediction mode and a second prediction mode. Coefficient of a non-linear polynomial model is determined based on the current template and the reference template, the non-linear polynomial model indicating a weight w0 of a first prediction obtained from the first prediction mode. Each reference template is obtained based on a first prediction mode, a second prediction mode, and a respective candidate nonlinear polynomial model. The nonlinear polynomial model is dependent on at least one of x and y. (x, y) indicates a sample position in the current block. The current block is reconstructed based on a weighted average of the first prediction and a second prediction of the current block obtained using the second prediction mode according to the weight w0.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 544,765, filed on October 18, 2023, entitled "On Partition Derivation of Geometric Partition", which is hereby incorporated by reference in its entirety. Technical Field

[0002] Aspects of the present disclosure generally relate to video coding and decoding. Background Art

[0003] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background art section, the work of the presently named inventors, and aspects of the description that may not otherwise be considered prior art at the time of filing are neither expressly nor implicitly admitted to be prior art against the present disclosure.

[0004] Image / video compression can help to transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial redundancy and temporal redundancy. In an example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the current picture in reconstruction to perform sample prediction. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can utilize motion compensation to predict samples in the current picture based on a previously reconstructed picture. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding.

[0006] In one aspect, a method for processing visual media data includes processing a bitstream of visual media data according to formatting rules. The bitstream includes a syntax element indicating that a current block is encoded in a geometric partitioning mode (GPM) using a first prediction mode and a second prediction mode. The formatting rules specify determining coefficients of a second-degree non-linear polynomial model based on a current template and a reference template of the current block. The non-linear polynomial model indicates a weight w0 of a first prediction of the current block obtained from the first prediction mode. The current template includes neighboring reconstructed samples of the current block. The formatting rules specify that each reference template is obtained based on the first prediction mode, the second prediction mode, and a corresponding candidate second-degree non-linear polynomial model with corresponding candidate coefficients. The formatting rules specify that the non-linear polynomial model depends on x and y, and (x, y) indicates a sample position in the current block. The formatting rules specify processing the current block based on a weighted average of the first prediction and a second prediction of the current block obtained using the second prediction mode according to the weight w0.

[0007] In an example, the formatting rules specify that w0 is ax 2 +bx+cy 2 +dy+e, and a weight w1 of the second prediction is (1 - w0).

[0008] In one aspect, a method for video coding includes determining whether to apply a geometric partitioning mode (GPM) to a current block, the geometric partitioning mode (GPM) including a GPM partitioning method using a weight w0. The weight w0 is indicated by a non-linear polynomial model and is a weight of a first prediction of the current block obtained from the first prediction mode. The non-linear polynomial model depends on x and y. (x, y) indicates a sample position in the current block. When the GPM using the weight w0 indicated by the non-linear polynomial model is applied to the current block, the method for video coding includes determining coefficients of the non-linear polynomial model based on a current template and a reference template of the current block. The current template includes neighboring samples of the current block. Each reference template is obtained based on the first prediction mode, the second prediction mode, and a corresponding candidate non-linear polynomial model with corresponding candidate coefficients of the GPM. The method for video coding includes encoding the current block based on a weighted average of the first prediction and a second prediction of the current block obtained using the second prediction mode according to the weight w0.

[0009] In an example, w0 is ax 2 +bx+cy 2 +dy+e and ax 3 +bx 2 +cx+d y 3+ey 2 +fy+g, and a weight w1 of the second prediction is (1 - w0).

[0010] According to one aspect of the present disclosure, an apparatus for video decoding includes processing circuitry. The processing circuitry is configured to receive coding information that indicates that a current block is encoded in a geometric partitioning mode (GPM) using a first prediction mode and a second prediction mode. The processing circuitry is configured to determine coefficients of a non-linear polynomial model that indicates a weight w0 of a first prediction of the current block obtained from the first prediction mode. The coefficients are determined based on a current template and a reference template of the current block. The current template includes neighboring reconstructed samples of the current block. Each reference template is obtained based on the first prediction mode, the second prediction mode, and a respective candidate non-linear polynomial model having corresponding candidate coefficients. The non-linear polynomial model depends on at least one of x and y, where (x, y) indicates a sample position in the current block. The processing circuitry is configured to reconstruct the current block based on a weighted average of the first prediction and a second prediction of the current block obtained using the second prediction mode according to the weight w0.

[0011] In an example, the degree of the non-linear polynomial model is 2, w0 = ax 2 + bx + cy 2 + dy + e, the determined coefficients include a, b, c, d, and e, and the weight w1 of the second prediction is (1 - w0).

[0012] In an example, the degree of the non-linear polynomial model is 3, w0 = ax 3 + bx 2 + cx + dy 3 + ey 2 + fy + g, the determined coefficients include a, b, c, d, e, f, and g, and the weight w1 of the second prediction is (1 - w0).

[0013] In an example, the processing circuitry is configured to use a regression method to determine the coefficients of the non-linear polynomial model.

[0014] In an example, the processing circuitry is configured to: downsample an initial current template of the current block having a first resolution to obtain a current template having a second resolution before determining the coefficients of the non-linear polynomial model. The second resolution is lower than the first resolution.

[0015] In an example, the coding information includes a syntax element indicating whether the weight w0 is applied to the current block, and the processing circuitry is configured to determine whether to apply the weight w0 to the current block based on the syntax element.

[0016] In an example, the coding information includes a syntax element indicating one type of multiple types of GPM, and the multiple types of GPM include the type of GPM to which the weight w0 is applied.

[0017] In an example, the current template is one of multiple templates, and the multiple templates include an upper template directly above the current block, a left template directly to the left of the current block, and an L-shaped template that includes the upper template and the left template.

[0018] In an example, the encoded information includes a syntax element indicating the current template for the current block, and the processing circuitry is configured to determine the current template as one of the multiple templates based on the syntax element.

[0019] In an example, the encoded information includes a syntax element indicating the current template, and the processing circuitry is configured to determine the current template as one of the following: the L-shaped template when the syntax element is 0, the left template when the syntax element is 1, and the upper template when the syntax element is 2.

[0020] In an example, the encoded information includes a syntax element indicating whether the current template is selected from the multiple templates, and when the syntax element is determined to indicate that the current template is selected from the multiple templates, the processing circuitry is configured to determine the current template as one of the multiple templates based on another syntax element in the encoded information.

[0021] Aspects of the present disclosure also provide an apparatus for video coding. The apparatus for video coding includes processing circuitry configured to implement any of the methods described for video coding.

[0022] Aspects of the present disclosure also provide a method for video decoding. The method includes any of the methods implemented by an apparatus for video decoding.

[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to execute any of the methods described for video decoding / encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0025] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).

[0026] Figure 2 is a schematic illustration of an example of a block diagram of a decoder.

[0027] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.

[0028] Figure 4Shows an example of intra-picture block compensation such as an Intra Block Copy (IBC) mode according to an aspect of the present disclosure.

[0029] Figure 5 Shows an example of intra-picture block compensation with a Coding Tree Unit (CTU) size search range according to an aspect of the present disclosure and, in some examples, reusing memory to search some parts of the left CTU.

[0030] Figure 6 Shows an example of an Intra Template Matching Prediction (IntraTMP) mode according to an aspect of the present disclosure.

[0031] Figure 7 Shows an example of a first method according to an aspect of the present disclosure in which two geometric partitions are obtained by using a specified splitting pattern.

[0032] Figure 8 Shows an example of a dividing line of a geometric splitting pattern in some examples.

[0033] Figure 9 Shows an example of a second method of a geometric splitting pattern such as using template matching according to an aspect of the present disclosure.

[0034] Figure 10 Shows an example of a geometric splitting pattern for predicting a current block, including a non-linear weight indicated by a non-linear polynomial model, according to an aspect of the present disclosure.

[0035] Figure 11 Shows an example of determining coefficients of a non-linear polynomial model using template matching in a geometric splitting pattern according to an aspect of the present disclosure.

[0036] Figure 12 Shows an example of a template type according to an aspect of the present disclosure.

[0037] Figure 13 Shows a flowchart outlining a decoding process according to some aspects of the present disclosure.

[0038] Figure 14 Shows a flowchart outlining an encoding process according to some aspects of the present disclosure.

[0039] Figure 15 Is a schematic illustration of a computer system according to an aspect. Detailed Description

[0040] Figure 1 FIG. 1 shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example for the disclosed subject matter of the application of video encoders and video decoders in a streaming environment. The disclosed subject matter may equally apply to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including, for example, CD (Compact Disc), DVD (Digital Versatile Disc), memory sticks, etc.

[0041] The video processing system (100) includes a capture subsystem (113), which may include a video source (101), such as a digital imaging device, that creates, for example, an uncompressed video picture stream (102). In an example, the video picture stream (102) includes samples taken by the digital imaging device. The video picture stream (102) is depicted as a thick line to emphasize the high data volume when compared to the encoded video data (104) (or encoded video bitstream), and the video picture stream (102) may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter described in more detail below. The encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (102), and the encoded video data (104) (or encoded video bitstream) may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1The client subsystems (106) and (108) therein can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an outgoing stream (111) of video pictures that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to certain video coding / compression standards. Examples of such standards include the ITU-T (International Telecommunication Union - Telecommunication Standardization Sector, ITU-T) H.265 Recommendation. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.

[0042] Note that the electronic devices (120) and (130) can include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown), and the electronic device (130) can also include a video encoder (not shown).

[0043] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., receiving circuitry). The video decoder (210) can be used to replace Figure 1 the video decoder (110) in the example.

[0044] A receiver (231) may receive one or more encoded video sequences to be decoded by a video decoder (210), such as included in a bitstream. In one aspect, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequences from the other data. To guard against network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be external to the video decoder (210) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (210) to, for example, guard against network jitter, and additionally there may be another buffer memory (215) internal to the video decoder (210) to, for example, handle playout timing. When the receiver (231) is receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be needed, or the buffer memory (215) may be small. To make as much use as possible of packet networks such as the Internet, a buffer memory (215) may be needed, which may be relatively large and may advantageously have an adaptive size and may be implemented at least in part in an operating system or a similar element (not depicted) external to the video decoder (210).

[0045] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequences. The categories of these symbols include: information for managing the operation of the video decoder (210), and potential information for controlling a rendering device such as a rendering device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to the electronic device (230), as Figure 2As shown. The control information for the rendering device can be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set segment (not depicted). The parser (220) can perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup can include a Group Of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0046] The parser (220) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (215) to create symbols (221).

[0047] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser (220) from the encoded video sequence. For clarity, this subgroup control information flow between the parser (220) and the multiple units below is not depicted.

[0048] In addition to the functional blocks already mentioned, the video decoder (210) can conceptually be divided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the following functional units.

[0049] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives, from the parser (220), quantized transform coefficients as symbols (221) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) may output a block including sample values, which may be input into the aggregator (255).

[0050] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses the surrounding reconstructed information extracted from the current picture buffer (258) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds, on a per-sample basis, the predictive information already generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0051] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbols (221) belonging to the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or a residual signal) to generate output sample information. The address within the reference picture memory (257) from which the motion compensation prediction unit (253) extracts the prediction samples may be controlled by a motion vector, which may be in the form of symbols (221) available for use by the motion compensation prediction unit (253), which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0052] The output samples of the aggregator (255) may be subjected to various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, as well as to previously reconstructed and loop filtered sample values.

[0053] The output of the loop filter unit (256) may be a sample stream that may be output to a rendering device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.

[0054] Once fully reconstructed, certain coded pictures can be used as reference pictures for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257) and a new current picture buffer can be reallocated before starting to reconstruct a subsequent coded picture.

[0055] The video decoder (210) may perform decoding operations according to a standard or predetermined video compression technology such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows both the syntax of the video compression technology or standard and the profile as recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as tools that can only be used under the profile. For compliance, the complexity of the encoded video sequence is also required to be within a range as defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.

[0056] In one aspect, the receiver (231) may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the encoded video sequence. The video decoder (210) may use the additional data to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0057] Figure 3 FIG. shows an example of a block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) may be used in place of Figure 1 the video encoder (103) in the example.

[0058] The video encoder (303) may receive video samples from a video source (301) (which is not part of the electronic device (320) in the Figure 3 example), and the video source (301) may capture video images to be encoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0059] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on samples.

[0060] According to one aspect, a video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraint as desired. Implementing an appropriate encoding speed is a function of a controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to other functional units. For the sake of brevity, the couplings are not depicted. Parameters set by the controller (350) can include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions belonging to the video encoder (303) optimized for a specific system design.

[0061] In some aspects, the video encoder (303) is configured to operate in an encoding / decoding loop. As an overly simplified description, in an example, the encoding / decoding loop can include a source encoder (330) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder will also create. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder “sees” the same sample values as the sample values that the decoder will “see” when using prediction during decoding as reference picture samples. The basic principle of reference picture synchronization (and the drift that occurs in cases where synchronization cannot be maintained due to, for example, channel errors) is also used in some related technologies.

[0062] The operation of the “local” decoder (333) can be the same as the operation of a “remote” decoder such as the video decoder (210) that has been described in detail above in connection with Figure 2 However, also briefly referring to Figure 2 , since the encoding of the symbols into the encoded video sequence by the entropy encoder (345) and the decoding of the symbols by the parser (220) can be lossless, the entropy decoding part of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0063] In one aspect, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. Since encoder techniques are inverse to the decoder techniques described in detail, the description of encoder techniques can be simplified. More detailed descriptions are provided in some sections below.

[0064] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding that predicts an input picture by referring to one or more previously encoded pictures designated as "reference pictures" from a video sequence. In this way, the encoding engine (332) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.

[0065] The local video decoder (333) may decode the encoded video data of a picture that can be designated as a reference picture based on symbols created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 3 (not shown), the reconstructed video sequence is generally a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that the video decoder may perform on the reference picture and may store the reconstructed reference picture in the reference picture memory (334). In this way, the video encoder (303) may locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by the remote video decoder.

[0066] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata such as reference picture motion vectors, block shapes, etc. that can be used as an appropriate prediction reference for the new picture. The predictor (335) may operate on a per-pixel basis for sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).

[0067] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0068] The outputs of all the above-mentioned functional units can undergo entropy coding in an entropy encoder (345). The entropy encoder (345) converts the symbols into an encoded video sequence by applying lossless compression to the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0069] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can merge the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0070] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can generally be assigned to one of the following picture types:

[0071] An intra picture (I picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh (“IDR (Independent Decoder Refresh)”) pictures.

[0072] A predictive picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses motion vectors and reference indices to predict the sample values of each block.

[0073] A bi-predictive picture (B picture) can be encoded and decoded using intra prediction or inter prediction that uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0074] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded on a block-by-block basis. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment applied to the corresponding pictures of the blocks. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded with reference to a previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be predictively encoded with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.

[0075] The video encoder (303) can perform encoding operations according to standards such as the ITU-T H.265 recommendation or a predetermined video coding technology. In the operation of the video encoder (303), the video encoder (303) can perform various compression operations, including predictive coding operations that utilize the temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technology or standard being used.

[0076] In one aspect, the transmitter (340) can transmit additional data together with the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0077] Video can be captured as multiple source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture in encoding / decoding, which is called the current picture, is segmented into blocks. In the case where a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0078] In some aspects, bidirectional prediction techniques can be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0079] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0080] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC (High-Efficiency Video Coding) standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) in a quadtree manner. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an example, each CU is analyzed to determine the prediction type for that CU, such as an inter-prediction type or an intra-prediction type. According to temporal and / or spatial predictability, a CU is divided into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one aspect, prediction operations in encoding / decoding (encoding / decoding) are performed on a prediction block basis. Using a luminance prediction block as an example of a prediction block, a prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0081] Note that any suitable technique can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In one aspect, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another aspect, one or more processors that execute software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).

[0082] Video coding is widely used in many applications such as broadcasting, video recording, video streaming, etc. Many emerging video coding standards such as H.264, H.265 / HEVC, H.266 / VVC, and AV1 (AOMedia Video 1, AV1) have been released and widely adopted in video applications. In one aspect, a hybrid video codec may include the following codec modules, such as intra prediction, inter prediction, transform coding, quantization, entropy coding, post-loop in-loop filters, etc.

[0083] In various examples, the current picture can be used as a reference region for block-based compensation, such as in the intra block copy (IBC) mode, intra template matching prediction (IntraTMP) mode, etc.

[0084] In one aspect, block compensation can be performed from a previously reconstructed region within the same picture, which may include intra picture block compensation (also known as current picture referencing (CPR) or IBC mode). Figure 4 An example of intra picture block compensation such as the IBC mode according to one aspect of the present disclosure is shown. A displacement vector indicating the offset between the current block (430) and the reference block (440) may be referred to as a block vector (BV) (450). The current block (430) and the reference block (440) are in the current picture (400).

[0085] Different from the motion vector (MV) in motion compensation - which can be any value (positive or negative in the x or y direction), the BV can have some constraints such that the pointed-to reference block is available and has been reconstructed. In an example, referring to Figure 4 , the current picture (400) may include a region to be decoded (420) and a reconstructed region (410). In an example, the BV can be constrained to point to a reference block in the reconstructed region (410). In some examples, for the consideration of parallel processing, some reference regions that are tile boundaries or wavefront trapezoid boundaries can be excluded.

[0086] The coding of the BV can be explicit or implicit. In the explicit mode (referred to as the AMVP (Advanced Motion Vector Prediction, AMVP) mode in inter coding), the difference between the BV and the BV predictor can be signaled; in the implicit mode, the BV can be recovered only from the BV predictor in a manner similar to the MV in the merge mode. In some implementations, the resolution of the BV can be limited to integer positions; in other systems, the resolution of the BV can be allowed to point to fractional positions.

[0087] The block-level flag known as the IBC flag can be used to signal the use of intra-block copy at the block level. In one aspect, the IBC flag is signaled when the current block is not coded in the merge mode. In an example, the IBC flag can be signaled by a reference index method, such as by treating the current decoded picture as a reference picture. In an example, in HEVC SCC (Screen Content Coding, SCC), such a reference picture (e.g., the current decoded picture) is placed at the last position in a list (e.g., the reference picture list). The special reference picture (e.g., the current decoded picture) can be managed together with other temporal reference pictures in the Decoded Picture Buffer (DPB).

[0088] In some examples, the IBC mode can be regarded as an inter prediction mode or an intra prediction mode.

[0089] There may be some variations for intra-block copy. For example, intra-block copy is regarded as a third mode, which is different from the intra prediction mode or the inter prediction mode. By regarding intra-block copy as a third mode, the block vector prediction in the merge mode and the AMVP mode can be separated from the conventional inter mode. In an example, the explicit mode described above can be called the IBC AMVP mode, and the implicit mode described above can be called the IBC merge mode. For example, a separate merge candidate list is defined for the IBC mode (e.g., the IBC merge mode), where all entries in the list are BVs. Similarly, in an example, the block vector prediction list in the IBC AMVP mode only includes BVs. In some examples, the general rules applied to the two lists include: in terms of candidate derivation processing, the two lists can follow the same logic as the inter merge candidate list used in the inter mode or the AMVP predictor list used in the inter mode. For example, for the IBC mode, 5 spatial neighboring positions in the inter merge mode such as the HEVC or VVC inter merge mode can be accessed to derive the merge candidate list for the IBC mode itself.

[0090] Figure 5 An example of intra-picture block compensation with a search range of one CTU size according to an aspect of the present disclosure and reusing memory to search some parts of the left CTU in some examples is shown.

[0091] In some examples, such as in VVC, the search range of the IBC mode is constrained within the current CTU. In the example, the effective memory requirement for storing the reference samples of the IBC mode is one CTU-sized sample. Considering that the existing reference sample memory stores the reconstructed samples in the current 64×64 region, three additional 64×64-sized reference sample memories can be used. Thus, the following method can be used: extend the effective search range of the IBC mode to some parts of the left CTU, while the total memory requirement for storing the reference pixels can remain unchanged. For example, the total memory requirement is one CTU-sized, e.g., a total of four 64×64 reference sample memories. Figure 5 An example of this memory reuse mechanism is shown. Each vertical stripe block is the current coding region (Curr), and the samples in each gray area are the encoded samples. The crossed-out areas (marked with "X") are not available for reference because the crossed-out areas in the reference sample memory can be replaced by the coding regions in the current CTU.

[0092] Figure 6 An example of the IntraTMP mode according to an aspect of the present disclosure is shown. In the example, the IntraTMP mode is a special intra prediction mode. In the example, the IntraTMP mode is different from the intra prediction mode. Referring to Figure 6 , in the example of the IntraTMP mode, the prediction block (621) can be copied, e.g., the best prediction block from the reconstructed part of the current frame. A template (620) such as an L-shaped template of the prediction block (621) can match the current template (630) of the current block (631). For a predefined search range, the encoder can search for the template most similar to the current template (630) in the reconstructed part of the current frame and can use the corresponding block (621) as the prediction block. In the example, the encoder then signals the use of the IntraTMP mode, and the same prediction operation is performed on the decoder side.

[0093] Referring to Figure 6 , a prediction signal can be generated by matching the L-shaped causal neighbor of the current block (631) with another block in the predefined search region. In the example, the predefined search region includes or consists of the following: R1 as the current CTU, R2 as the upper left CTU of the current CTU, R3 as the upper CTU of the current CTU, and R4 as the left CTU of the current CTU. In the example, the sum of absolute differences (SAD) is used as the cost function.

[0094] Within each region, the decoder can search for a template with the minimum cost (e.g., minimum SAD) relative to the current template, and can use the block corresponding to the minimum cost as the prediction block.

[0095] The dimensions (SearchRange_w, SearchRange_h) of all regions can be set to be proportional to the block dimensions (BlkW, BlkH) so that there is a fixed number of SAD comparisons per pixel. In an example, SearchRange_w = a × BlkW, and SearchRange_h = a × BlkH, where "a" is a constant that controls the trade-off between gain and complexity. In an example, "a" is equal to 5.

[0096] To accelerate the template matching process, in some examples, the search ranges of all search regions are subsampled, for example, subsampled by a factor of 2, and thus the template matching search is reduced by 4. After finding the best match, a refinement process can be performed. The refinement process can be performed via a second template matching search in a reduced range around the best match. In an example, the reduced range is defined as min(BlkW, BlkH) / 2.

[0097] The intra-frame template matching tool can be enabled for CUs with a size less than or equal to 64 in width and height. The maximum CU size in IntraTMP mode can be configurable. For example, when the decoder-side Intra Mode Derivation (DIMD) is not used for the current CU, the IntraTMP mode can be signaled at the CU level via a dedicated flag.

[0098] The term "IBC mode" can refer to the IBC mode or variants described in the present disclosure. The term "IntraTMP mode" can refer to the IntraTMP mode or variants described in the present disclosure.

[0099] The present disclosure includes split derivation for geometric splitting - also known as wedge splitting - to enhance signaling and prediction under geometric splitting.

[0100] According to one aspect of the present disclosure, a first method can be designed to divide a block (e.g., an encoded block) into two different partitions in a specified splitting mode. Figure 7 An example of a first method is shown in which two geometric partitions are obtained by using a specified splitting mode. In Figure 7In the example shown, the CU (700) is divided into two parts or two partitions (701) to (702) by an edge (also referred to as a dividing line or partitioning line) (721). The prediction data for the two partitions (701) to (702) can be predicted by using different prediction modes, including but not limited to intra prediction, inter prediction, and intra block copy (e.g., IBC mode), IntraTMP mode, or any suitable mode. For example, the weights of the predictions from two different prediction modes for each sample can be derived based on the distance between the sample position and the edge (721). An example of the first method is the geometric partitioning mode or variant. In some examples, the geometric partitioning mode is referred to as GPM.

[0101] In some examples of GPM, any suitable mode can be used to predict the two partitions (701) to (702), such as those described above in the first method. Any suitable mode (e.g., intra prediction, inter prediction, and IBC mode, IntraTMP mode, etc.) can be used to predict each of the two partitions (701) to (702). In some examples (e.g., VVC), for inter prediction, GPM is supported. A CU-level flag can be used to signal, together with other merge modes, the geometric partitioning mode as one of the merge modes, and the other merge modes are, for example, the regular merge mode, MMVD (Merge Motion Vector Difference, MMVD) mode, CIIP (Combined Inter and Intra Prediction, CIIP) mode, sub-block merge mode, etc. In some examples, for each possible CU size w×h = 2 m ×2 n , where m, n ∈ {3…6}, excluding 8×64 and 64×8, the geometric partitioning mode supports a total of 64 partitions.

[0102] Referring to Figure 7 , when using GPM, the CU (700) is divided into two parts (701) to (702) by the dividing line (721). The dividing line (721) can be a geometrically positioned straight line.

[0103] Figure 8 Examples of the dividing lines of the geometric partitioning mode in some examples are shown. The dividing lines are grouped by the same angle. Specifically, Figure 8 each rectangle (801) in Figure 8 represents a CU. Multiple parallel lines are shown in each rectangle. The multiple parallel lines correspond to the dividing lines of the same angle. In

[0104] The position of the dividing line can be mathematically derived based on the angle and offset parameters of a specific partition. In an example, each of the two geometric partitions formed by the dividing line in a CU uses its own motion for inter-frame prediction. For example, when only unidirectional prediction is allowed for each partition, each part has one motion vector and one reference index. Unidirectional prediction constraints (e.g., only unidirectional prediction is allowed for each partition) can be applied to ensure that a CU in GPM mode can be encoded as in conventional bi-directional prediction. For example, for each CU in GPM mode, two motion compensation predictions are performed (e.g., each of the two motion compensation predictions corresponds to one unidirectional prediction).

[0105] In some examples, when GPM is used for the current CU, a geometric segmentation index indicating the segmentation mode of the geometric partition (e.g., indicating the angle and offset) and two merge indices (one merge index for each partition) are further signaled. In some examples, the value of the maximum GPM candidate size is explicitly signaled in the SPS (Sequence Parameter Set, SPS), and syntax binarization is specified for the GPM merge indices. After predicting each part of the geometric partition, the sample values along the edge of the geometric partition are adjusted using hybrid processing with adaptive weights to obtain the prediction signal for the entire CU. Then, as in other prediction modes, transform and quantization processing can be applied to the entire CU. Finally, the motion field of the CU predicted using the geometric segmentation mode is stored.

[0106] A second method can be designed to reorder the segmentation modes by using Template Matching (TM), as Figure 9 shown. Figure 9 An example of a second method according to an aspect of the present disclosure is shown. An example of the second method is a variant of GPM, such as GPM with TM. According to the second method, a current block or current CU (900) is divided into two partitions (901) to (902) by an edge or dividing line (921). The dividing line (921) can extend from the current block (900) to the current template (910). The current template (910) can be constructed by using the neighboring reconstructed samples of the current block (900). A reference template can be constructed by using the neighboring samples of a reference block in a reference picture. Template matching for each segmentation mode can be applied to the reference template and the current template (910) of the current block (900) to calculate the corresponding TM cost. Thus, the TM cost can be calculated for each segmentation mode. The segmentation mode indices can be reordered based on the TM cost of the segmentation modes, for example, in ascending order.

[0107] As Figure 9As shown, template matching can be applied to GPM. When the GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to two geometric partitions. TM can be used to refine the motion information for each geometric partition. When TM is selected, a template can be constructed using left neighboring samples, upper neighboring samples, or upper-left neighboring samples according to the splitting angle. Motion can be refined by minimizing the difference between the current template (910) and the template in the reference picture - for example, using the same search pattern in the merge mode and disabling the half-pixel interpolation filter. In Figure 9 In the example shown, the current template (910) can include an upper template (911) and a left template (912). The upper template (911) includes the upper neighboring samples of the current block (900), and the left template (912) includes the left neighboring samples of the current block (900). In the example, the current template can consist of the upper template (911). In the example, the current template can consist of the left template (912).

[0108] In the present disclosure, examples of geometric segmentation modes include a first method in which a block is divided into two different partitions. The weight w i (also referred to as the prediction weight or partition weight) can refer to the weight in the first method (e.g., geometric segmentation mode), and i can be 0 or 1. In one aspect, the weight w0 can refer to the weight of the first prediction of the current block obtained from the first prediction mode, and the weight w1 can refer to the weight of the second prediction of the current block obtained from the second prediction mode. In one aspect, the sum of w0 and w1 is 1.

[0109] In some examples, the template can refer to the template in a TM method such as the second method. The current template adjacent to the current block can be constructed using the adjacent reconstructed samples of the current block. In the example, the reference template can be constructed using the adjacent reconstructed samples of the reference block (or prediction block).

[0110] According to one aspect of the present disclosure, a geometric segmentation mode can be applied to the current block. Two different prediction modes including a first prediction mode and a second prediction mode can be applied to the current block to generate two predictions respectively including a first prediction P0 and a second prediction P1. One of the first prediction mode and the second prediction mode can be one of the following: intra prediction mode, inter prediction mode, IBC mode, IntraTMP mode, etc. The current block can be predicted based on the weighted average of the first prediction P0 and the second prediction P1 according to the weights w0 and w1 respectively. When the sum of w0 and w1 is 1, the weighted average of the first prediction P0 and the second prediction P1 is based on w0 or w1.

[0111] Figure 10An example for predicting the geometric partitioning mode of a current block (1000) is shown. A first prediction mode is applied to the current block (1000) to generate a first prediction P0, e.g., a first prediction block (1001). A second prediction mode is applied to the current block (1000) to generate a second prediction P1, e.g., a second prediction block (1002). A prediction (1011) can be generated based on a weighted average of the first prediction P0 and the second prediction P1 according to weights w0 and w1, respectively, to predict the current block (1000).

[0112] According to one aspect of the present disclosure, the weights w0 and w1 can have a non - linear relationship with at least one of x and y, where (x, y) indicates a sample position in the current block (1000). Thus, the weights w0 and w1 can depend on at least one of x and y in a non - linear manner, which may be different from some variants of the GPM.

[0113] Therefore, different samples in the current block (1000) can have different weights w0 (and different weights w1). According to one aspect of the present disclosure, a polynomial model, e.g., a non - linear polynomial model, can be used to indicate the weights w0 (and w1), and the weights w0 and w1 can be referred to as non - linear weights. In one aspect, the degree of the non - linear polynomial model is at least 2, e.g., 2, 3, etc.

[0114] The polynomial model can be specified according to the degree of the polynomial model, the variables in the polynomial model (e.g., x and / or y), and the coefficients of each term in the polynomial model. The degree of the polynomial model can be an integer greater than 1, e.g., 2, 3, etc. The variables in the polynomial model can include: (i) one variable, e.g., x or y, or (ii) multiple variables, e.g., x and y. In an example, the polynomial model can be determined by template matching, e.g., by matching the current template of the current block with one of the reference templates in the reference templates, and thus Figure 10 the GPM described in can be referred to as a regression - based GPM. The reference template can be determined using a candidate polynomial model.

[0115] According to one aspect of the present disclosure, the coefficients of the non - linear polynomial model indicating the weight w0 can be determined by template matching based on the current template of the current block and the reference template. The current template can include neighboring reconstructed samples of the current block. An example of the current template is as Figure 10The current template (1010) of the current block (1000) shown. Each reference template can be obtained based on a first prediction mode, a second prediction mode, and a corresponding candidate non - linear polynomial model with corresponding candidate coefficients (e.g., the i - th candidate non - linear polynomial model). In an example, each candidate non - linear polynomial model and the non - linear polynomial model have the same type of dependence on at least x and y (e.g., the same degree and the same variables) and have different coefficients. For example, in the non - linear polynomial model, w0 = ax 2 +bx + cy 2 +dy + e, and in the i - th candidate non - linear polynomial model, w 0i = aix 2 +b i x + c i +d i y + e i . The i - th candidate non - linear polynomial model and the non - linear polynomial model have the same degree 2 and depend on x and y. The coefficients of the i - th candidate non - linear polynomial model include a i 、b i 、c i 、d i and e i . The coefficients of the non - linear polynomial model include a, b, c, d, and e.

[0116] Figure 11 illustrates an example of determining the coefficients of a non - linear polynomial model using template matching in a geometric segmentation mode according to an aspect of the present disclosure. The geometric segmentation mode as described in Figure 10 is used to predict the current block (1000). In an example, before predicting the current block (1000), the coefficients of the non - linear polynomial model indicating the weight w0 are determined, as described below for example. A plurality of candidate non - linear polynomial models with corresponding candidate coefficients can be determined. As described above, each candidate non - linear polynomial model and the non - linear polynomial model can have the same type of dependence on x and y and can have different coefficients. For each candidate non - linear polynomial model, a first reference template (1021) of the current template (1010) can be predicted based on the first prediction mode. A second reference template (1022) of the current template (1010) can be predicted based on the second prediction mode. For the i - th pair of weights w 0i and w 1i (e.g., w 1i = 1 - w 0i ), where w 0i is indicated by the i - th candidate non - linear polynomial model, based on the first reference template (1021) and the second reference template (1022) according to w 0i and w 1iThe reference template (1031) of the current template (1010) is determined by the weighted sum. The TM cost associated with the i-th candidate non-linear polynomial model can be determined based on the comparison between the current template (1010) and the reference template (1031) of the current template (1010). The above process can be performed for each candidate non-linear polynomial model to obtain the TM cost of the candidate non-linear polynomial model. In an example, the candidate non-linear polynomial model with the minimum TM cost can be selected as the non-linear polynomial model for predicting the current block (1000), and thus the coefficients of the candidate non-linear polynomial model with the minimum TM cost can be selected as the coefficients of the non-linear polynomial model.

[0117] Such as Figures 10 to 11 The geometric segmentation pattern using non-linear weights of template matching described in can be different from some variations of the geometric segmentation pattern such as Figures 7 to 9 described in, and thus may be advantageous in some applications. For example, such as Figures 10 to 11 shown, using TM to determine the weights w0 and w1 is content-adaptive and takes into account the local content adjacent to the current block, which can provide more accurate weights w0 and w1 and a more accurate prediction of the current block. In addition, in some variations of the GPM, the weights have a linear relationship with at least one of x and y, and thus the dividing line is a straight line. In some examples, the straight-line segmentation (e.g., Figures 7 to 9 shown in) may not be as accurate as the curve segmentation, which is indicated by the non-linear relationship between at least one of x and y and the weights w0 and w1 (e.g., Figures 10 to 11 described in). Therefore, the GPM using non-linear weights of template matching described in the present disclosure can have a more accurate prediction of the segmentation (e.g., curve segmentation), and thus improve the prediction accuracy.

[0118] In one aspect, the segmentation weights (e.g., w0 and w1) can be derived by using a template (e.g., the current template (1010)). In an example, a non-linear polynomial model, such as a polynomial model of degree 2 or higher, can be derived based on the current template (1010) and the reference template, such as Figure 11 described in. The derived non-linear polynomial model, such as a polynomial model of degree 2 or higher, can be applied to the prediction weights (e.g., w0 and w1) of each sample in the weighted average of two prediction modes (e.g., the first prediction mode and the second prediction mode), such as Figure 10 shown.

[0119] In one aspect, a square term (e.g., x 2 or y 2 ) or a high-degree term higher than the square term (e.g., x 3 or y3 ) high - order equations to indicate polynomial models of degree 2 or higher. Two examples of polynomial models are shown below. The following polynomial models have derived coefficients (or parameters), a, b, c, d, e, f, and / or g, and x and y are the coordinates of samples in the current block. In the example, the degree of the non - linear polynomial model is 2, and the polynomial equation is w0 = ax 2 + bx + cy 2 + dy + e, the determined coefficients include a, b, c, d, and e, and the weight w1 of the second prediction is (1 - w0). In the example, the degree of the non - linear polynomial model is 3, and the polynomial equation is w0 = ax 3 + bx 2 + cx + dy 3 + ey 2 + fy + g, the determined coefficients include a, b, c, d, e, f, and g, and the weight w1 is (1 - w0).

[0120] In one aspect, the coefficients in the non - linear polynomial model are derived by using any regression method such as the least - squares method, the maximum - likelihood method, etc.

[0121] In one aspect, with reference to Figure 11 , the derived coefficients in the polynomial model are derived by using the data in the current template (1010) and the reference template (1031). The data in the current template (1010) may include the reconstructed samples in the current template (1010). The data in the reference template (1031) may be derived as described in Figure 11 .

[0122] In the example, a subsampling method (e.g., downsampling method) can be applied to coefficient derivation to reduce the number of samples used in the calculation, and thus improve the calculation efficiency. For example, the current template (1010) with the first resolution can be downsampled before deriving the reference template (1031), and thus the current template and the reference template used in determining the coefficients of the non - linear polynomial model can have a second resolution lower than the first resolution.

[0123] In one aspect, a flag at the block level (e.g., the coded block level) is signaled to indicate whether a method (e.g., the GPM of non - linear weights using template matching as described in Figures 10 to 11 ) is applied to the current block. In the example, the coding information in the bitstream includes a syntax element (e.g., a flag) indicating whether the weight w0 indicated by the non - linear polynomial model is applied to the current block. It can be determined whether to apply the weight w0 to the current block based on this syntax element.

[0124] In one aspect, signaling syntax is used to indicate which GPM splitting method to select (e.g., a GPM such as the one with non-linear weights using template matching as described in Figures 10 to 11 ), and this GPM splitting method is one of multiple GPM splitting methods. In an example, the coded information in the bitstream includes a syntax element indicating one type of multiple types of GPMs, and the multiple types of GPMs can include a type of GPM in which non-linear weights w0 of template matching are applied. In an example, TM is used to determine the coefficients of the non-linear weights w0, such as Figure 11 described in.

[0125] In one aspect, templates can have different types. In an example, the current template is one of multiple templates. Referring to Figure 12 , the multiple templates can include an upper template (1211) directly above the current block (1200), a left template (1212) directly to the left of the current block, and an L-shaped template (1210) including the upper template (1211) and the left template (1212). In some examples, the L-shaped template further includes the upper left corner between the upper template (1211) and the left template (1212). For example, the current template can include only the left template (1212), only the upper template (1211), or include the L-shaped template (1210), etc., to derive the coefficients of a non-linear polynomial model (e.g., a polynomial equation).

[0126] In one aspect, a flag at the block level (e.g., the coded block level) is signaled to indicate which template type (e.g., only the left template, only the upper template, or the L-shaped template) to select to derive the non-linear polynomial model. For example, the coded information includes a syntax element indicating the current template of the current block. Based on the syntax element, the current template is determined as one of multiple templates (e.g., the upper template (1211), the left template (1212), and the L-shaped template (1210)).

[0127] In one aspect, a flag is signaled to indicate whether a partial template (e.g., the left template (1212), the upper template (1211), etc.) can be applied. If the flag is true, another syntax is signaled to indicate which template type to use. Otherwise, if the flag is false, a predefined template type (e.g., the L-shaped template) can be used. In an example, the coded information includes a syntax element indicating whether the current template is selected from multiple templates. When the syntax element is determined to indicate that the current template is selected from multiple templates, the current template is determined as one of multiple templates based on another syntax element in the coded information.

[0128] In one aspect, a syntax element is signaled to indicate which template type to use. Table 1 shows examples of syntax elements. In the example, the encoded information in the bitstream includes a syntax element that indicates the template type of the current template. Referring to Figure 11 , the current template can be determined as one of the following: an L-shaped template (1210) when the syntax element is 0; a left-only template (1212) when the syntax element is 1; and an upper-only template (1211) when the syntax element is 2. Table 1 shows examples of syntax elements Syntax value Template type 0 Left template and upper template 1 Only left template 2 Only upper template

[0129] In one aspect, the syntax for implicitly signaling template type selection indicates the template type among the available templates to be selected when the current block is at one of the following boundaries, including but not limited to: picture boundary, slice boundary, tile boundary, coding tree unit (CTU) boundary, and superblock boundary, etc. In the example, which one of the upper template (1211), left template (1212), and L-shaped template (1210) is the current template depends on the position of the current block, where the position of the current block can be one of the picture boundary, slice boundary, tile boundary, CTU boundary, and superblock boundary.

[0130] In one aspect, the syntax for implicitly signaling template type selection indicates the template type among the available templates to be selected when a part of the template is not available for the current block. For example, if the upper template (e.g., the upper part (1211) of template (1210)) is not available, the only available template is the left template (1212). Thus, when the upper template (1211) is not available, the current template is the left template (1212). In the example, when the left template is not available, the current template is the upper template.

[0131] In one aspect, when only the left template (1212) or only the upper template (1211) is used to derive the non-linear polynomial model, the degree of the non-linear polynomial model is one of 2 or 3, w0 depends on x and not on y, and the weight w1 is (1 - w0). In the example, the polynomial equation is w0 = ax 2 + bx + c and w1 = 1 - w0. In the example, w0 = ax 3 + bx 2 + cx + d and w1 = 1 - w0.

[0132] In one aspect, when deriving a non-linear polynomial model using only the left template (1212) or only the upper template (1211), the degree of the non-linear polynomial model can be constrained to one or more predefined degrees, such as one of 2 and 3, w0 depends on y and not on x, and the weight w1 of the second prediction is (1 - w0). In an example, the polynomial equation is w0 = ay 2 + by + c and w1 = 1 - w0. In an example, w0 = ay 3 + by 2 + cy + d and w1 = 1 - w0.

[0133] Figure 13 FIG. shows a flowchart outlining a process (1300) according to one aspect of the present disclosure. The process (1300) can be used in a device such as a video decoder. In various aspects, the process (1300) is performed by processing circuitry such as processing circuitry that performs the functions of a video decoder (110), processing circuitry that performs the functions of a video decoder (210), etc. In some aspects, the process (1300) is implemented as software instructions, so when the processing circuitry executes the software instructions, the processing circuitry performs the process (1300). The process begins at (S1301) and proceeds to (S1310).

[0134] At (S1310), encoded information is received that indicates that the current block is encoded in a geometric partitioning mode (GPM) using a first prediction mode and a second prediction mode.

[0135] At (S1320), a non-linear polynomial model (e.g., the coefficients of the non-linear polynomial model) can be determined by template matching based on the current template of the current block and a reference template. The non-linear polynomial model indicates the weight w0 of the first prediction of the current block obtained from the first prediction mode. In an example, the weight w1 of the second prediction is (1 - w0). The current template includes neighboring reconstructed samples of the current block. Each reference template can be obtained based on the first prediction mode, the second prediction mode, and a corresponding candidate non-linear polynomial model with corresponding candidate coefficients. The non-linear polynomial model can depend on at least one of x and y. (x, y) indicates the sample position in the current block. Under GPM, template matching can be used to determine the non-linear weights.

[0136] In an example, the degree of the non-linear polynomial model is 2, w0 = ax 2 + bx + cy 2 + dy + e, and the determined coefficients include a, b, c, d, and e.

[0137] In an example, the degree of the non-linear polynomial model is 3, w0 = ax 3 + bx2 +cx + dy 3 +ey 2 + fy + g, where the determined coefficients include a, b, c, d, e, f, and g.

[0138] In an example, a regression method is used to determine the coefficients of the non - linear polynomial model.

[0139] In an example, before determining the coefficients of the non - linear polynomial model, the initial current template with the first resolution of the current block is downsampled to obtain a current template with the second resolution. The second resolution is lower than the first resolution. The downsampled current template and the downsampled reference template can be used to perform template matching.

[0140] At (S1330), the current block is reconstructed based on the weighted average of the first prediction and the second prediction of the current block obtained using the second prediction mode according to the weight w0.

[0141] Then, the process proceeds to (S1399) and terminates.

[0142] The process (1300) can be appropriately adjusted. The steps in the process (1300) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0143] In an example, the encoded information includes a syntax element indicating whether the weight w0 is applied to the current block, and based on this syntax element, it is determined whether the weight w0 is applied to the current block.

[0144] In an example, the encoded information includes a syntax element indicating one of multiple types of GPM, and the multiple types of GPM include the type of GPM to which the weight w0 (e.g., the non - linear weight w0 indicated by the non - linear polynomial model) is applied.

[0145] In one aspect, the current template is one of multiple templates, and the multiple templates include an upper template directly above the current block, a left template directly to the left of the current block, and an L - shaped template including the upper template and the left template. In an example, the encoded information includes a syntax element of the current block indicating the current template, and based on this syntax element, the current template is determined as one of the multiple templates. In an example, the encoded information includes a syntax element indicating the current template, and the current template is determined as one of the following: an L - shaped template when the syntax element is 0, a left template when the syntax element is 1, and an upper template when the syntax element is 2.

[0146] In an example, the coding information includes a syntax element indicating whether the current template is selected from multiple templates. When the syntax element is determined to indicate that the current template is selected from multiple templates, the current template is determined as one of the multiple templates based on another syntax element in the coding information.

[0147] In an example, which one of the upper template directly above the current block, the left template directly to the left of the current block, and the L-shaped template including the upper template and the left template is the current template depends on the position of the current block, and the position of the current block is one of a picture boundary, a slice boundary, a tile boundary, a coding tree unit (CTU) boundary, and a superblock boundary.

[0148] In an example, when the left template directly to the left of the current block is not available, the current template is the upper template directly above the current block, and when the upper template is not available, the current template is the left template.

[0149] In an example, when the current template is the upper template directly above the current block or the left template directly to the left of the current block, the degree of the non-linear polynomial model is one of 2 and 3, w0 depends on x and not on y, and the weight w1 of the second prediction is (1 - w0).

[0150] In an example, when the current template is the upper template directly above the current block or the left template directly to the left of the current block, the degree of the non-linear polynomial model is one of 2 and 3, w0 depends on y and not on x, and the weight w1 of the second prediction is (1 - w0).

[0151] In an example, one of the first prediction mode and the second prediction mode is one of the following: an intra prediction mode, an inter prediction mode, an intra block copy (IBC) mode, and an intra template matching (IntraTMP) mode.

[0152] Figure 14 A flowchart showing an overview of a process (1400) according to an aspect of the present disclosure is shown. The process (1400) can be used in a video encoder. In various aspects, the process (1400) is performed by a processing circuitry such as a processing circuitry that performs the functions of the video encoder (103), a processing circuitry that performs the functions of the video encoder (303), etc. In some aspects, the process (1400) is implemented by software instructions, so when the processing circuitry executes the software instructions, the processing circuitry performs the process (1400). The process starts at (S1401) and proceeds to (S1410).

[0153] At (S1410), it is determined whether to apply a geometric partitioning mode (GPM) to the current block, and the geometric partitioning mode (GPM) includes a GPM partitioning method using a weight w0. The weight w0 can be indicated by a non-linear polynomial model and can be the weight of the first prediction of the current block obtained from the first prediction mode. The non-linear polynomial model can depend on x and y, where (x, y) indicates the sample position in the current block.

[0154] At (S1420), when applying the GPM using the weight w0 indicated by the non-linear polynomial model to the current block, the coefficients of the non-linear polynomial model can be determined by template matching based on the current template of the current block and a reference template. The current template includes neighboring samples of the current block. Each reference template can be obtained based on the first prediction mode, the second prediction mode of the GPM, and a corresponding candidate non-linear polynomial model with corresponding candidate coefficients, for example Figure 11 as shown.

[0155] In an example, w0 is one of ax 2 + bx + cy 2 + dy + e and ax 3 + bx 2 + cx + dy 3 + ey 2 + fy + g, and the weight w1 of the second prediction is (1 - w0).

[0156] At (S1430), the current block can be predicted based on a weighted average of the first prediction according to the weight w0 and the second prediction of the current block obtained using the second prediction mode, for example Figure 10 as shown.

[0157] Then, the process proceeds to (S1499) and terminates.

[0158] The process (1400) can be appropriately adjusted. The steps in the process (1400) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0159] In one aspect, a method for processing visual media data includes processing a bitstream of visual media data according to formatting rules. For example, the bitstream can be a bitstream decoded / encoded by any of the decoding and / or encoding methods described herein. The formatting rules can specify one or more constraints of the bitstream and / or one or more processes to be performed by a decoder and / or an encoder.

[0160] The bitstream includes a syntax element indicating that a current block is to be encoded using a geometric partitioning mode (GPM) with a first prediction mode and a second prediction mode. Format rules may specify coefficients of a second-degree non-linear polynomial model based on a current template and a reference template of the current block. The non-linear polynomial model indicates a weight w0 of a first prediction of the current block obtained from the first prediction mode. The current template includes neighboring reconstructed samples of the current block. Format rules may specify obtaining each reference template based on the first prediction mode, the second prediction mode, and a respective candidate second-degree non-linear polynomial model with corresponding candidate coefficients, as described, for example Figure 11 in. Format rules may specify that the non-linear polynomial model depends on x and y, and (x, y) indicates a sample position in the current block. Format rules may specify processing the current block based on a weighted average of the first prediction and a second prediction of the current block obtained using the second prediction mode according to the weight w0.

[0161] In an example, the format rules specify that w0 is ax 2 + bx + cy 2 + dy + e, and a weight w1 of the second prediction is (1 - w0).

[0162] The methods, aspects, and examples described in this disclosure may be used alone or in any combination. For example, some methods, aspects, and / or examples performed by a decoder may be performed by an encoder, and some methods, aspects, and / or examples performed by an encoder may be performed by a decoder. Each of the method (or aspect), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0163] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 15 FIG. shows a computer system (1500) suitable for implementing certain aspects of the disclosed subject matter.

[0164] The computer software may be encoded using any suitable machine code or computer language, which may be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0165] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0166] Figure 15 The components shown for the computer system (1500) are examples and are not intended to place any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components shown in the example aspects of the computer system (1500).

[0167] The computer system (1500) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices can also be used to capture certain media not necessarily directly related to conscious input by humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from still image cameras), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0168] Input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1501), mouse (1502), touchpad (1503), touch screen (1510), data glove (not shown), joystick (1505), microphone (1506), scanner (1507), camera device (1508).

[0169] The computer system (1500) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include: haptic output devices (e.g., haptic feedback through a touch screen (1510), a data glove (not shown), or a joystick (1505), but there may also be haptic feedback devices that do not serve as input devices); audio output devices (e.g., speakers (1509), headphones (not depicted)); visual output devices (e.g., a screen (1510), including a CRT (Cathode Ray Tube) screen, an LCD (Liquid Crystal Display) screen, a plasma screen, an OLED (Organic Light-Emitting Diode) screen, each screen having or not having touch screen input capabilities, each screen having or not having haptic feedback capabilities - some of the screens may be able to output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted); holographic displays, and smoke cans (not depicted)); and printers (not depicted).

[0170] The computer system (1500) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory) / RW (1520) with media such as CD / DVD (1521), thumb drives (1522), removable hard disk drives or solid state drives (1523), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit) / PLD (Programable Logic Device) such as security dongles (not depicted), etc.

[0171] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0172] The computer system (1500) may also include an interface (1554) to one or more communication networks (1555). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN (Local Area Network, LAN); cellular networks including GSM (Global System for Mobile Communications, GSM), 3G (the Third Generation, 3G), 4G (the Fourth Generation, 4G), 5G (the Fifth Generation, 5G), LTE (Long Term Evolution, LTE), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network - BUS, CANBus), etc. Certain networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1549) (such as, for example, the USB (Universal Serial Bus, USB) port of the computer system (1500)); other networks are typically integrated into the core of the computer system (1500) by attaching to a system bus as described below (for example, an Ethernet interface in a PC (Personal Computer, PC) computer system or a cellular network interface in a smartphone computer system). Using any of these networks, the computer system (1500) can communicate with other entities. Such communication can be one-way reception only (e.g., broadcast TV), one-way transmission only (e.g., CAN bus to certain CAN bus devices), or two-way (e.g., to other computer systems using local or wide area digital networks). Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0173] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1540) of the computer system (1500).

[0174] The core (1540) may include one or more central processing units (CPUs) (1541), a graphics processing unit (GPU) (1542), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1543), a hardware accelerator (1544) for certain tasks, a graphics adapter (1550), etc. These devices together with the read-only memory (ROM) (1545), a random access memory (1546), an internal mass storage device (1547) such as an internal non-user-accessible hard disk drive, a solid-state drive (SSD), etc. may be connected via a system bus (1548). In some computer systems, the system bus (1548) may be accessible in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus (1548) of the core or may be attached to the system bus (1548) of the core via a peripheral bus (1549). In an example, a screen (1510) may be connected to the graphics adapter (1550). The architecture of the peripheral bus includes PCI (Peripheral Component Interconnect / Interface), USB, etc.

[0175] The CPU (1541), GPU (1542), FPGA (1543), and accelerator (1544) may execute certain instructions, which when combined may constitute the computer code mentioned above. The computer code may be stored in the ROM (1545) or a random access memory (RAM) (1546). Transient data may also be stored in the RAM (1546), while permanent data may be stored in, for example, the internal mass storage device (1547). Fast storage and retrieval of any memory device in the memory devices may be achieved by using a cache memory, which may be closely associated with one or more CPUs (1541), GPUs (1542), mass storage devices (1547), ROM (1545), RAM (1546), etc.

[0176] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the art of computer software.

[0177] By way of example and not limitation, a computer system (1500) having an architecture and in particular a core (1540) can provide functionality due to a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core (1540) having a non-transitory nature, such as a mass storage device (1547) inside the core or a ROM (1545). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the core (1540). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (1540) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in the RAM (1546) and modifying such data structures according to the processes defined by the software. Additionally or alternatively, the computer system can provide functionality due to being logically hardwired or otherwise embodied in circuitry (e.g., an accelerator (1544)), which can operate in place of or in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In appropriate cases, references to software can include logic, and references to logic can also include software. In appropriate cases, references to computer-readable media can include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.

[0178] The use of "at least one of..." or "one of..." in the present disclosure is intended to include any one of the recited elements or combinations thereof. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include A or B or (A and B). The use of "one of..." does not exclude any combination of the recited elements in cases where, for example, the elements are not mutually exclusive.

[0179] Although several examples of aspects of the present disclosure have been described, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that those skilled in the art will be able to envision many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for processing visual media data, the method comprising: The bit stream of the visual media data is processed according to format rules, wherein: The bitstream includes a syntax element indicating that a current block is encoded in a geometric partition mode (GPM) using a first prediction mode and a second prediction mode; and The format rules specify: determining coefficients of a nonlinear polynomial model of degree 2 by template matching based on a current template of the current block and a reference template, the nonlinear polynomial model indicating a weight w0 of a first prediction of the current block obtained from the first prediction mode, the current template comprising neighboring reconstructed samples of the current block, each reference template being obtained based on the first prediction mode, the second prediction mode and a corresponding candidate nonlinear polynomial model of degree 2 having corresponding candidate coefficients, the nonlinear polynomial model being dependent on x and y, (x, y) indicating a sample position in the current block; and The current block is processed based on a weighted average of the first prediction and a second prediction of the current block obtained using the second prediction mode according to the weight w0.

2. The method according to claim 1, wherein: The format rule specifies that w0 is ax 2 +bx+cy 2 +dy+e, and the weight w1 of the second prediction is (1-w0).

3. A method for video encoding, comprising: determining whether to apply a geometric partition mode (GPM) to a current block, the geometric partition mode (GPM) comprising a GPM partitioning method using a weight w0, the weight w0 being indicated by a non-linear polynomial model and being a weight of a first prediction of the current block obtained from a first prediction mode, the non-linear polynomial model being dependent on x and y, (x, y) indicating a sample position in the current block; When applying the GPM using the weight w0 indicated by the nonlinear polynomial model to the current block, determining coefficients of the nonlinear polynomial model by template matching based on a current template of the current block and reference templates, the current template including neighboring samples of the current block, each reference template being obtained based on the first prediction mode, the second prediction mode of the GPM, and a corresponding candidate nonlinear polynomial model having corresponding candidate coefficients; as well as The current block is encoded based on a weighted average of the first prediction according to the weight w0 and a second prediction of the current block obtained using the second prediction mode.

4. The method according to claim 3, wherein: w0 is ax 2 +bx+cy 2 +dy+e and ax 3 +bx 2 +cx+dy 3 +ey 2 +fy+g; and The weight w1 of the second prediction is (1-w0).

5. The method according to claim 3 or 4, wherein: The current template is one of a plurality of templates; and The plurality of templates include an upper template located right above the current block, a left template located right to the left of the current block, and an L-shaped template including the upper template and the left template.

6. A device for video decoding, the device comprising: processing circuitry configured to: receiving encoding information indicating that a current block is encoded in a geometric partition mode (GPM) using a first prediction mode and a second prediction mode; and determining coefficients of a nonlinear polynomial model, the nonlinear polynomial model indicating a weight w0 of a first prediction of the current block obtained from the first prediction mode, the coefficients being determined by template matching based on a current template of the current block and a reference template, the current template comprising neighboring reconstructed samples of the current block, each reference template being obtained based on the first prediction mode, the second prediction mode, and a corresponding candidate nonlinear polynomial model with corresponding candidate coefficients, the nonlinear polynomial model being dependent on at least one of x and y, (x, y) indicating a sample position in the current block; as well as The current block is reconstructed based on a weighted average of the first prediction according to the weight w0 and a second prediction of the current block obtained using the second prediction mode.

7. The device according to claim 6, wherein: The degree of the nonlinear polynomial model is 2, w0 = ax 2 +bx+cy 2 +dy+e, the determined coefficients include a, b, c, d and e, and the weight w1 of the second prediction is (1-w0).

8. The device according to claim 6, wherein: The degree of the nonlinear polynomial model is 3, w0=ax 3 +bx 2 +cx+dy 3 +ey 2 +fy+g, the determined coefficients include a, b, c, d, e, f and g, and the weight w1 of the second prediction is (1-w0).

9. The device according to any one of claims 6 to 8, wherein: The processing circuitry is configured to determine the coefficients of the nonlinear polynomial model using a regression method.

10. The device according to any one of claims 6 to 9, wherein: The processing circuit system is configured to: Before determining the coefficients of the nonlinear polynomial model, an initial current template of the current block having a first resolution is downsampled to obtain the current template having a second resolution, the second resolution being lower than the first resolution.

11. The device according to any one of claims 6 to 10, wherein: The encoding information includes a syntax element indicating whether the weight w0 is applied to the current block; and The processing circuitry is configured to determine whether to apply the weight w0 to the current block based on the syntax element.

12. The device according to claim 6, wherein: The encoding information includes a syntax element indicating one type among a plurality of types of GPM, and the plurality of types of GPM include a type of GPM in which the weight w0 is applied.

13. The device according to any one of claims 6 to 11, wherein: The current template is one of a plurality of templates; and The plurality of templates include an upper template located right above the current block, a left template located right to the left of the current block, and an L-shaped template including the upper template and the left template.

14. The device according to claim 13, wherein: The encoding information includes a syntax element of the current block indicating the current template; and The processing circuitry is configured to determine the current template as one of: When the syntax element is 0, it is the L-shaped template; When the syntax element is 1, it is the left template; and When the syntax element is 2, it is the upper template.

15. The device according to any one of claims 6 to 11, wherein: The encoding information includes a syntax element indicating whether the current template is selected from a plurality of templates; and When the syntax element is determined to indicate that the current template is selected from the plurality of templates, the processing circuit system is configured to determine the current template to be one of the plurality of templates based on another syntax element in the encoding information.