Adaptive sample-based bi-directional optical flow for GPM with bi-predictive motion vector

By employing geometric segmentation and bidirectional optical flow motion refinement techniques in video coding, the problem of low video coding efficiency in existing technologies is solved, achieving more efficient video data compression and quality improvement.

CN121569478APending Publication Date: 2026-02-24TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480027225.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-24
Filing Date
2024-04-22
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively utilize spatial and temporal redundancy for efficient compression when processing video data. In particular, the accuracy of motion vector indication and prediction is insufficient in inter-frame prediction, resulting in low coding efficiency.

Method used

The video blocks are segmented and motion vectors are refined using geometric segmentation mode (GPM) and sample-based bidirectional optical flow (S-BDOF) motion refinement techniques. By determining the spatial relationship of the partitions and the sample intersection, S-BDOF motion refinement is applied to improve coding accuracy.

Benefits of technology

It improves the efficiency and quality of video encoding, reduces the amount of data, enhances the accuracy of motion vector prediction, and improves the quality of the encoded video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569478A_ABST
    Figure CN121569478A_ABST
Patent Text Reader

Abstract

Some aspects of the present disclosure provide an apparatus for video decoding. The apparatus comprises processing circuitry configured to: receive an encoded video bitstream comprising encoding information for one or more pictures; determining that a current block in the current picture is in a geometric partitioning mode (GPM) according to the encoding information; and determining whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement on at least a first GPM partition of the current block. The first GPM partition has a bidirectional prediction motion vector. The processing circuitry is further configured to: when it is determined to apply S-BDOF motion refinement on the first GPM partition, apply S-BDOF motion refinement on one or more samples in the first GPM partition; and reconstructing the current block using the one or more samples reconstructed based on the S-BDOF motion refinement.
Need to check novelty before this filing date? Find Prior Art

Description

Related applications

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 461,583, filed April 24, 2023, entitled “Adaptive Bi-Directional Sample Based Optical Flow on GPM with Bi-Predictive Motion Vector,” which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure generally describes implementation methods related to video encoding and decoding. Background Technology

[0003] The background description provided herein is for the purpose of presenting the overall context of this disclosure. To the extent that the work described in this background section is intended, neither the work of the currently identified inventors nor any aspect of the description that would not otherwise be considered prior art at the time of filing is expressly or implicitly acknowledged as prior art to this disclosure.

[0004] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. For instance, a video codec can use a technique called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] This disclosure includes methods and apparatus for video encoding / decoding.

[0006] Some aspects of this disclosure provide a method for processing video media data. The method includes processing a bitstream of visual media data according to format rules. The bitstream includes encoded information for one or more images. The format rules specify: encoding a current block in the current image using a geometric segmentation mode (GPM); and whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least a first GPM partition in the current block. The first GPM partition has bidirectional predicted motion vectors. The format rules further specify: when applying S-BDOF motion refinement to the first GPM partition, determining the spatial relationship between sub-blocks in the first GPM partition and the segmentation boundary of the GPM; when a sub-block is not adjacent to the segmentation boundary, applying sub-block-based BDOF motion refinement to the sub-block; and when the segmentation boundary intersects with the sub-block, applying S-BDOF motion refinement to samples in the sub-block.

[0007] Some aspects of this disclosure provide an apparatus for video decoding. The apparatus includes a processing circuitry configured to: receive an encoded video bitstream comprising encoded information of one or more images; determine, based on the encoded information, that a current block in the current image is in a geometric segmentation mode (GPM); and determine whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least a first GPM partition of the current block. The first GPM partition has bidirectional predicted motion vectors. The processing circuitry is further configured to: when it is determined that S-BDOF motion refinement should be applied to the first GPM partition, apply S-BDOF motion refinement to one or more samples in the first GPM partition; and reconstruct the current block using one or more samples reconstructed based on the S-BDOF motion refinement.

[0008] In some examples, the processing circuitry is configured to determine the application of S-BDOF motion refinement to each GPM partition in the current block that has a bidirectional predicted motion vector.

[0009] In some examples, the processing circuitry is configured to: divide the first GPM partition into sub-blocks; determine whether the first sub-block is adjacent to the GPM partition boundary; and when the first sub-block is not adjacent to the partition boundary, apply S-BDOF motion refinement to the first sample in the first sub-block.

[0010] In some examples, the processing circuitry is configured to: divide a first GPM partition into sub-blocks; determine whether the partition boundary of the GPM intersects with the first sub-block within the sub-blocks; when the partition boundary does not intersect with the first sub-block, apply sub-block-based BDOF motion refinement to the first sub-block; and when the partition boundary intersects with the first sub-block, apply S-BDOF motion refinement to the first sub-block.

[0011] In some examples, the processing circuitry is configured to decode a flag from the encoded video bitstream indicating whether S-BDOF motion refinement is applied to each partition with bidirectional predictive motion vectors, the flag being signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0012] In some examples, the processing circuitry is configured to: compare the number of samples in the first GPM partition with a threshold to obtain a comparison result; and determine the application of S-BDOF motion refinement based on the comparison result.

[0013] In some examples, the processing circuitry is configured to: compare the number of sub-blocks in the first GPM partition with a threshold to obtain a comparison result; and determine the application of S-BDOF motion refinement based on the comparison result.

[0014] In some examples, the processing circuitry is configured to compare weight values ​​in a portion of the GPM blending mask with a threshold to obtain a comparison result, this portion corresponding to a first GPM partition. Based on the comparison result, the processing circuitry determines to apply S-BDOF motion refinement to one or more samples in the first GPM partition. In one example, the processing circuitry is configured to determine to apply S-BDOF motion refinement to one or more samples in the first GPM partition when each weight value is greater than or equal to the threshold. In another example, the processing circuitry is configured to determine to apply S-BDOF motion refinement to one or more samples in the first GPM partition when each weight value is less than or equal to the threshold.

[0015] In some examples, the threshold is predefined, or is signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0016] In some examples, the processing circuitry is configured to decode the block-level syntax element associated with the current block, which indicates one of the following: applying S-BDOF motion refinement to both GPM partitions of the current block; not applying S-BDOF motion refinement to any GPM partition of the current block; applying S-BDOF motion refinement only to the first GPM partition; and applying S-BDOF motion refinement only to the second GPM partition of the current block.

[0017] In some examples, the processing circuitry is configured to decode the block-level syntax element associated with the current block, which indicates one of the following: applying S-BDOF motion refinement to two GPM partitions of the current block; not applying S-BDOF motion refinement to any GPM partition of the current block; and applying S-BDOF motion refinement to the GPM partition in the current block that uses the first GPM merge index, but not to another GPM partition of the current block.

[0018] Some aspects of this disclosure provide a method for video coding. The method includes: determining to encode a current block in a current image in a geometric segmentation mode (GPM); determining whether to apply sample-based bidirectional optical flow (S-BDOF) motion thinning to at least a first GPM partition of the current block, the first GPM partition having bidirectional predicted motion vectors; when it is determined that S-BDOF motion thinning will be applied to the first GPM partition, applying S-BDOF motion thinning to one or more samples in the first GPM partition; and reconstructing the current block using one or more samples reconstructed based on the S-BDOF motion thinning.

[0019] In some examples, to determine whether to apply S-BDOF motion refinement, the method includes applying S-BDOF motion refinement to each GPM partition in the current block that has a bidirectional predicted motion vector.

[0020] In some examples, to apply S-BDOF motion refinement, the method includes: dividing a first GPM partition into sub-blocks; determining whether the first sub-block is adjacent to a GPM segmentation boundary; and when the first sub-block is not adjacent to a segmentation boundary, applying S-BDOF motion refinement to a first sample in the first sub-block.

[0021] In some examples, to apply S-BDOF motion refinement, the method includes: dividing a first GPM partition into sub-blocks; determining whether the partition boundary of the GPM intersects with a first sub-block within the sub-blocks; when the partition boundary does not intersect with the first sub-block, applying sub-block-based BDOF to the first sub-block; and when the partition boundary intersects with the first sub-block, applying S-BDOF motion refinement to the first sub-block.

[0022] In some examples, the method includes: encoding a flag in an encoded video bitstream that includes encoding information of the current block in the current picture, the flag indicating whether S-BDOF motion refinement is applied to each partition having a bidirectional predictive motion vector, the flag being signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0023] In some examples, to determine whether to apply S-BDOF motion refinement, the method includes: comparing weight values ​​in a portion of a GPM blending mask with a threshold to obtain a comparison result, the portion corresponding to a first GPM partition; and determining, based on the comparison result, to apply S-BDOF motion refinement to one or more samples in the first GPM partition.

[0024] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes a processing circuitry system. This processing circuitry system can be configured to perform any of the methods described for video decoding / encoding.

[0025] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods described for video decoding / encoding. Attached Figure Description

[0026] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0027] Figure 1 It is a schematic diagram of an exemplary block diagram of a communication system.

[0028] Figure 2 This is a schematic illustration of an exemplary block diagram of a decoder.

[0029] Figure 3 This is a schematic illustration of an exemplary block diagram of an encoder.

[0030] Figure 4 The locations of spatial merging candidates according to an embodiment of this disclosure are shown.

[0031] Figure 5 Candidate pairs for redundancy checking of spatial merging candidates are shown according to an embodiment of this disclosure.

[0032] Figure 6 An exemplary motion vector scaling for time merging candidates is shown.

[0033] Figure 7 An exemplary candidate position for the time merging candidate of the current block is shown.

[0034] Figure 8 The diagram shows some examples of angles used in Geometric Partition Mode (GPM).

[0035] Figure 9 A graph showing possible partition edges in the example is provided.

[0036] Figure 10 A diagram illustrating the hybrid processing in some examples is shown.

[0037] Figure 11 A plot of the ramp function for the weights of GPM mixing based on the displacement from the predicted sample location to the GPM partition boundary and the size of the mixing region is shown.

[0038] Figures 12A to 12D Plots of GPM with inter-frame prediction and intra-frame prediction are shown in some examples.

[0039] Figure 13 The diagram illustrates some examples of expanding GPM-split edges to obtain edges on a template.

[0040] Figure 14 An exemplary schematic view of decoder-side motion vector refinement based on bilateral matching is shown in some examples.

[0041] Figure 15 A graph showing some calculations in the bi-directional optical flow (BDOF) example is presented.

[0042] Figure 16 The search areas are shown in some examples.

[0043] Figure 17 A flowchart outlining some embodiments of the decoding process according to this disclosure is shown.

[0044] Figure 18 A flowchart outlining some embodiments of the coding process according to this disclosure is shown.

[0045] Figure 19 It is a schematic diagram of a computer system according to an implementation method. Detailed Implementation

[0046] Figure 1 Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an example of the application of the disclosed subject matter—video encoders and video decoders—in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs (Compact Discs), DVDs (Digital Video Disks), memory sticks, etc.

[0047] The video processing system (100) includes a capture subsystem (113) that may include a video source (101), such as a digital camera device, which creates, for example, an uncompressed video picture stream (102). In the example, the video picture stream (102) includes samples captured by the digital camera device. The video picture stream (102) is depicted as a thick line to emphasize the high data volume when compared with encoded video data (104) (or encoded video bitstream), which may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize its lower data volume when compared to the video picture stream (102). The encoded video data (104) (or encoded video bitstream) can be stored on a streaming server (105) for future use. One or more streaming client subsystems, for example... Figure 1 Client subsystems (106) and (108) can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an outgoing stream of video pictures (111) that can be displayed on a display (112) (e.g., a screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., video bitstreams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T (International Telecommunication Union-Telecommunication Standardization Sector, ITU-T) Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The topics that are exposed can be used in the context of VVC.

[0048] Note that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may also include a video encoder (not shown).

[0049] Figure 2An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuitry system). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the example.

[0050] The receiver (231) can receive, for example, one or more encoded video sequences to be decoded by the video decoder (210) included in a bitstream. In an implementation, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective user entities (not depicted). The receiver (231) can separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory can be external to the video decoder (210) (not depicted). In some other applications, a buffer memory (not depicted) may exist outside the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may exist inside the video decoder (210) to handle broadcast timing, for example. The buffer memory (215) may not be necessary when the receiver (231) is receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, or the buffer memory (215) may be small. For the purpose of utilizing packet networks such as the Internet as much as possible, a buffer memory (215) may be required, which may be relatively large and advantageously have an adaptive size, and may be implemented at least partially in the operating system or in a similar element (not depicted) outside the video decoder (210).

[0051] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols include: information for managing the operation of the video decoder (210), and potential information for controlling a presentation device, such as a presentation device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to it, such as... Figure 2As shown. The control information used for the presentation device can be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of the encoded video sequence for use in the video decoder based on at least one parameter corresponding to the group. The subgroup can include Group of Picture (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0052] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0053] Depending on the type of the encoded video picture or a portion thereof (e.g., inter-frame picture and intra-frame picture, inter-frame block and intra-frame block) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (220). For clarity, the flow of this subgroup control information between the parser (220) and the multiple units below is not depicted.

[0054] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0055] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) from the parser (220) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks containing sample values ​​that can be input into the aggregator (255).

[0056] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to intra-coded blocks. Intra-coded blocks are blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from the current picture buffer (258) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current picture buffer (258) buffers the partially reconstructed current image and / or the fully reconstructed current image. In some cases, the aggregator (255) adds the predictive information already generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.

[0057] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-frame coded block and potentially to a motion-compensated block. In such cases, the motion-compensated prediction unit (253) can access the reference image memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (221) belonging to the block, these samples can be added to the output of the scaler / inverse transform unit (251) by the aggregator (255) (in this case, referred to as residual samples or residual signals) to generate output sample information. The address in the reference image memory (257) from which the motion-compensated prediction unit (253) extracts its predicted samples can be controlled by motion vectors, which are provided to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0058] The output samples of the aggregator (255) can undergo various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters available to the loop filter unit (256) as included in the encoded video sequence (also referred to as the encoded video bitstream) and as symbols (221) from the parser (220). Video compression may also respond to metadata acquired during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0059] The output of the loop filter unit (256) can be a sample stream, which can be output to the presentation device (212) and stored in the reference image memory (257) for use in future inter-frame image prediction.

[0060] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and that encoded image (by, for example, the parser (220)) is identified as the reference image, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0061] The video decoder (210) can perform decoding operations according to standards such as ITU-T H.265 Recommendation or predetermined video compression technologies. An encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, the profile may select certain tools from all available tools in the video compression technology or standard as tools usable only under said profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is signaled in the encoded video sequence.

[0062] In this implementation, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0063] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuitry system). The video encoder (303) can be used in place of... Figure 1 The video encoder (103) in the example.

[0064] The video encoder (303) can obtain data from the video source (301) (which is not...). Figure 3 In one example, an electronic device (320) receives a video sample, and the video source (301) can capture a video image to be encoded by a video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0065] A video source (301) can provide a sequence of source video samples in the form of a digital video sample stream to be encoded by a video encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera device capturing local image information as a video sequence. Video data can be provided as multiple individual pictures that are given motion when viewed sequentially. The pictures themselves can be organized as spatial pixel arrays, where each pixel can include one or more samples, depending on the sampling structure, color space, etc., used. The following description focuses on samples.

[0066] According to the implementation, the video encoder (303) can encode and compress images of the source video sequence into an encoded video sequence (343) in real time or under any other time constraints as required. Implementing an appropriate encoding rate is a function of the controller (350). In some implementations, the controller (350) controls and is functionally coupled to other functional units as described below. For clarity, the coupling is not depicted. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions belonging to the video encoder (303) optimized for a specific system design.

[0067] In some implementations, the video encoder (303) is configured to operate within an encoding loop. As an oversimplification, in this example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols, creating sample data in a manner similar to how the (remote) decoder would also create them. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding of the symbol stream results in bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-accurate between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values ​​the decoder would "see" when using prediction during decoding. This basic principle of reference image synchronicity (and the drift that occurs, for example, due to channel errors) is also used in some related techniques.

[0068] The operation of the "local" decoder (333) can be combined with what has already been done above. Figure 2 The operation of a "remote" decoder, such as a video decoder (210), is the same as described in the detailed description. However, a brief reference is also made to... Figure 2 Since symbols are available and the encoding of symbols into an encoded video sequence by the entropy encoder (345) and the decoding of symbols by the parser (220) can be lossless, the entropy decoding portion of the video decoder (210), which includes the buffer (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0069] In the implementation, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. Since the encoder technique is the inverse of the fully described decoder technique, the description of the encoder technique can be simplified. A more detailed description is provided below in certain sections.

[0070] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding, which predictively codes the input image with reference to one or more previously encoded images from the video sequence designated as "reference images". In this way, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image that can be selected as a predictive reference for the input image.

[0071] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture based on symbols created by the source encoder (330). The operation of the encoding engine (332) can advantageously be lossy. When the encoded video data can be decoded by the video decoder (333), Figure 3 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process performed on the reference image by the video decoder and can store the reconstructed reference image in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0072] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata such as reference image motion vectors, block shapes, etc. that can be used as appropriate prediction references for the new image. The predictor (335) can operate on a pixel-by-pixel basis based on the sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).

[0073] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, the setting of parameters and subgroup parameters for encoding video data.

[0074] The outputs of all the functional units mentioned above can undergo entropy encoding in the entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into an encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0075] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0076] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded image type to each encoded image, which may affect the encoding techniques that can be applied to the corresponding image. For example, images can typically be assigned to one of the following image types:

[0077] Intra-frame pictures (I-pictures) can be encoded and decoded without using any other pictures in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0078] Predictive images (P-images) can be encoded and decoded using intra-frame or inter-frame prediction that utilizes motion vectors and reference indices to predict sample values ​​for each block.

[0079] Bidirectional predictive images (B-images) can be encoded and decoded using intra-frame or inter-frame predictions that utilize two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multi-predictive images can use more than two reference images and associated metadata for the reconstruction of a single block.

[0080] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined by the coding assignments of the corresponding images applied to the block. For example, blocks of image I can be non-predictively coded, or blocks of image I can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of image P can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of image B can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.

[0081] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In the operation of the video encoder (303), various compression operations can be performed, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0082] In this implementation, the transmitter (340) may transmit additional data along with the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, SEI messages, VUI parameter set fragments, etc.

[0083] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simply called intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, a specific image during encoding / decoding—referred to as the current image—is segmented into blocks. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference images.

[0084] In some implementations, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image that both precede the current image in the video in decoding order (but may be past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. A block can be predicted using a combination of the first and second reference blocks.

[0085] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0086] According to some embodiments of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC (High Efficiency Video Coding) standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression, with CTUs in the pictures having the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU comprises three Coding Tree Blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively split into one or more Coding Units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type used for that CU, such as inter-frame prediction or intra-frame prediction. Based on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In implementations, prediction operations in encoding / decoding are performed on a block-by-block basis. Using a luma prediction block as an example, this block comprises a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0087] Note that any suitable technology can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In one implementation, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another implementation, one or more processors that execute software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).

[0088] Various inter-frame prediction modes can be used in video coding. For example, in VVC, for a CU (Cumulative Unit) in inter-frame prediction, motion parameters may include MV (Motion Vector Model), one or more reference picture indices, reference picture list usage indices, and additional information to be used for generating certain coded features for inter-frame prediction samples. Motion parameters can be explicitly or implicitly signaled. When a CU is encoded in skip mode, the CU may be associated with a PU (Programmable Unit) and may not have valid residual coefficients, encoded motion vector increments or MV differences (e.g., MVD (MV Difference, MVD)), or reference picture indices. A merge mode can be specified, where the motion parameters of the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally include additional information such as that introduced in VVC. The merge mode can be applied to inter-frame prediction CUs, not just skip mode. In the example, an alternative to the merge mode is explicit transmission of motion parameters, where the MV, the corresponding reference picture index for each reference picture list, and reference picture list usage flags and other information are explicitly signaled for the CU.

[0089] In implementations, such as in VVC, the VVC Test Model (VTM) reference software includes one or more refined inter-frame predictive coding tools, including: Extended Combined Prediction, Merge Motion Vector Difference (MMVD) mode, Adaptive Motion Vector Prediction (AMVP) mode with symmetric MVD signaling, Affine Motion Compensation Prediction, Subblock-based Temporal Motion Vector Prediction (SbTMVP), Adaptive Motion Vector Resolution (AMVR), Motion Field Storage (1 / 16th Luminance Sample MV Storage and 8×8 Motion Field Compression), Bi-Prediction with CU-Level Weight (BCW), Bi-directional Optical Flow (BDOF), Prediction Refinement using Optical Flow (PROF), and DecoderSide Motion Vector Refinement. Refinement (DMVR), Combined Inter and Intra Prediction (CIIP), Geometric Partitioning Mode (GPM), etc. Inter-frame prediction and related methods are described in detail below.

[0090] In some examples, extended merge predictions can be used. In examples, such as VTM4 (Versatile VideoCoding Test Model 4), the merge candidate list is constructed by including the following five types of candidates in sequence: spatial motion vector predictor (MVP) from spatially neighboring CUs, temporal MVP from juxtaposed CUs, history-based MVP (HMVP) from a first-in-first-out (FIFO) table, pairwise average MVP, and zero MV.

[0091] The size of the merge candidate list can be signaled in the slice header. In the example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU encoded in merge mode, the index of the best merge candidate (e.g., the merge index) can be encoded using truncated unary binary binarization (TU). The first CU of the merge index can be encoded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for other CUs.

[0092] Below are some examples of the processing for generating merge candidates for each category. In the implementation, spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In the example, in the location... Figure 4 You can select up to four merge candidates from the candidates depicted in the diagram.

[0093] Figure 4 The locations of spatial merging candidates according to an embodiment of this disclosure are shown. (Refer to...) Figure 4 The resulting order is B1, A1, B0, A0, and B2. Position B2 is considered only if any CU at positions A0, B0, B1, and A1 is unavailable (e.g., because the CU belongs to another slice or another tile) or is intra-coded. After adding the candidate at position A1, the addition of the remaining candidates undergoes redundancy checking, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0094] To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only... Figure 5 The pairs are linked by arrows, and a candidate is added to the candidate list only if the corresponding candidate used for redundancy check does not have the same motion information.

[0095] Figure 5 Candidate pairs considered for redundancy checking in spatial merging candidates according to an embodiment of this disclosure are shown. (Refer to...) Figure 5 The pairs linked by the corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0096] In this implementation, time candidates are derived as follows. In the example, only one time merging candidate is added to the candidate list. Figure 6 An exemplary motion vector scaling for temporal merging candidates is shown. To derive a temporal merging candidate for the current CU (611) in the current image (601), the scaled MV (621) can be derived based on the co-located CU (612) belonging to the collocated reference image (604) (e.g., by...). Figure 6 (As shown by the dashed line in the image). The list of reference images used to derive the corresponding CU (612) can be explicitly notified by a signal in the slice header. The scaling MV (621) for the time merging candidates can be obtained, as shown in the image. Figure 6 The dashed lines in the diagram indicate that scaling the MV (621) can be performed from the MV of the co-located CU (612) using Picture Order Count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) and the current picture (601). The POC distance td can be defined as the POC difference between the juxtaposed reference picture (604) and the co-located picture (603). The reference picture index for the temporal merging candidate can be set to zero. The juxtaposed picture is used as a reference picture for the source picture derived from the temporal motion information. The juxtaposed picture can be identified in one of two lists, referred to as List 0 or List 1. In some examples, the encoder can determine the juxtaposed picture and signal it using appropriate syntax techniques.

[0097] Figure 7 Exemplary candidate positions (e.g., C0 and C1) for the current CU's time-merging candidate are shown. The position of the time-merging candidate can be selected from candidate positions C0 and C1. Candidate position C0 is located at the lower right corner of the current CU's sibling CU (710). Candidate position C1 is located at the center of the current CU's sibling CU (710). If the CU at candidate position C0 is unavailable, intra-coded, or outside the current line of the CTU, candidate position C1 is used to derive the time-merging candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and in the current line of the CTU, candidate position C0 is used to derive the time-merging candidate. The time-merging candidate can specify motion information from the Temporal Motion Vector Predictor (TMVP).

[0098] In some examples (e.g., VVC), a technique called Geometric Partitioning Mode (GPM) is used. Specifically, in VVC, GPM is used for inter-frame prediction. In the examples, GPM is only applied to CUs of 8×8 or larger. GPM can be signaled as a merging mode using CU-level flags, where other merging modes include, for example, regular merging mode, Merging with Motion Vector Difference (MMVD) mode, Combined Inter-Frame and Intra-Frame Prediction (CIIP) mode, and sub-block merging mode.

[0099] When using GPM mode on a CU, one of several partitioning methods is used to divide the CU into two geometrically shaped partitions via partition edges. In some examples, 64 different partitioning methods are used. The partitioning method can be distinguished by 24 angles (non-uniformly quantized between 0 and 360°) and for each angle by up to four edges relative to the center of the CU. A partition edge is a line that intersects the boundary of the CU and divides it into two partitions.

[0100] Figure 8 The diagram shows some examples of the 24 angles used under GPM. In some examples, angles can be identified using angle indices, such as angle indices 0 to 23.

[0101] Figure 9 The diagram shows a possible partition edge for angle index 3 in the example. Figure 9 In this context, four possible partition edges can be associated with angle index 3. Note that for some angle indices, three possible partition edges can be associated with each angle index.

[0102] In some examples, each geometric partition in the CU uses its own motion for inter-frame prediction. In these examples, only unidirectional prediction is allowed per partition; that is, each partition has one motion vector and one reference image index. Unidirectional prediction motion constraints are applied to ensure that, similar to bidirectional prediction, two motion-compensated predictions are needed for each CU.

[0103] In some examples, when the GPM is used for the current CU, signals are further used to indicate the geometric partition index (e.g., indicating angles and edges) and two merge indices (one merge index for each partition, such as a first merge index and a second merge index). In the example, the number of maximum GPM candidate sizes is explicitly indicated by signals at the slice level, and syntax binarization for the GPM merge index is specified.

[0104] In some examples, after prediction is made for each of the two geometric partitions, a blending process with adaptive weights is used to adjust the sample values ​​along the edges of the geometric partitions.

[0105] Figure 10 A graph illustrating some examples of blending processes is shown. The blending process uses a blending intensity, or blending region width, referred to as GPM. The parameters. For all the different content, the blending intensity. It can be fixed. In Figure 10 In the example, the blended area is shown by the shaded portion (1010) in block (1000).

[0106] In some examples, the weight values ​​in the blending mask used for blending can be given by a ramp function, for example, according to equation (1):

[0107] Equation (1)

[0108] In the example, in fixed In the case of pixels, the ramp function can be quantized according to equation (2):

[0109] Equation (2)

[0110] The result of the mixing process is used as the prediction signal for the entire CU, and transform and quantization processes can be further applied to the entire CU in a manner similar to other prediction modes. Finally, the motion field (e.g., motion information) of the CU predicted using GPM is stored. In some examples (e.g., VVC), the motion information of the CU is stored in 4×4 units (e.g., for each 4×4 luminance sample). The stored motion information is used for MV prediction and merging list construction for the next encoded CU. In GPM, three types of motion information are covered and these are stored in 4×4 units. The three types of motion information can include motion information from two partitions and motion information from the mixed region. For example, two geometric partitions P0 and P1 include their own unidirectional MV, and the mixed region between P0 and P1 is predicted using motion information from the two geometric partitions P0 and P1. Therefore, the motion information of GPM is stored according to partitions.

[0111] Motion information for the GPM is signaled in merge mode. To avoid additional memory bandwidth access, only unidirectional prediction is allowed for each partition in the GPM. In some examples, regular merge candidates can be either unidirectional or bidirectional predictions and cannot be directly used as the GPM merge list. To minimize implementation complexity, an index parity-based approach can be used to extract GPM merge candidates directly from the regular merge list without pruning. For example, for candidates with even-valued GPM merge indices, MV0 from reference list 0 and its corresponding regular merge index are used as GPM merge candidates. If MV0 is unavailable, MV1 from reference list 1 is used instead. Conversely, MV1 is selected as the default GPM merge candidate for odd-valued GPM merge indices.

[0112] Based on some aspects of this disclosure, additional technologies for GPM beyond VVC have been developed.

[0113] In some examples, a technique called Geometric Partitioning Mode (GPM) combined with Motion Vector Difference (MMVD) can be used; this technique is also known as GPM-MMVD. GPM in VVC is an extension of the existing GPM unidirectional MV by applying motion vector refinement. For example, the CUs in GPM are first signaled with a flag to indicate whether GPM mode is used. When using mode GPM, each geometric partition of the CU in GPM can decide whether to signal MVD. When signaling MVD for a geometric partition, after selecting a GPM merging candidate, the motion of the partition is further refined using the signaled MVD information. All other processes remain the same as in GPM.

[0114] In some examples, MVD in GPM is signaled as a pair of distances and directions, similar to MMVD. In one example, GPM combined with MMVD (GPM-MMVD) involves nine candidate distances, such as (¼ pixel, ½ pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (e.g., four horizontal / vertical directions and four diagonal directions). Additionally, in the example, when a flag (e.g., pic_fpel_mmvd_enabled_flag) is equal to 1, MVD is shifted left by 2 bits, as in MMVD.

[0115] In some examples, a technique called Geometric Partitioning Pattern (GPM) combined with adaptive mixing is used. In some examples (e.g., VVC), the final predicted samples are generated by mixing the predictions of the two prediction signals using a weighted average. Two integer mixing matrices (W0 and W1) are used. In some examples, the weights in the GPM mixing matrix are derived from a ramp function based on the displacement from the predicted sample location to the GPM partition boundary. In some examples, the mixing region size is fixed at two (e.g., two samples on each side of the GPM partition split boundary).

[0116] In some examples, the blending process is improved by adding additional blending region sizes, such as four blending region sizes that are one-quarter, half, twice, and four times the existing region size.

[0117] Figure 11 A plot of the ramp function for the weights of GPM mixing based on the displacement (d) from the predicted sample location to the GPM partition boundary and the mixing region size (τ) is shown. Figure 11 In the example, the first curve (1110) corresponds to the ramp function of the regular mixed region (also known as the existing region size) – for example, 2 samples on each side of the GPM partition split boundary; the second curve (1120) corresponds to the ramp function of a quarter mixed region of the regular mixed region; the third curve (1130) corresponds to the ramp function of a half mixed region of the regular mixed region; the fourth curve (1140) corresponds to the ramp function of a twice mixed region of the regular mixed region; and the fifth curve (1150) corresponds to the ramp function of a four times mixed region of the regular mixed region.

[0118] In some examples, CU-level flags are encoded to signal the size of the selected blending region. Additionally, extended weighted precision can be utilized, where the maximum value of the weights is changed from 8 (in VVC) to 32 to accommodate the extended blending region size.

[0119] In some examples, to accommodate the increased width of the GPM blending region, the maximum value of the weights is changed from 8 to 32. In the example, the weights are calculated as in equation (3):

[0120] Equation (3)

[0121] The width of the mixing region (e.g., A selection is allowed from a set of predefined values. In the example, the predefined values ​​are... It can be {½, 1, 2, 4, 8}.

[0122] In some examples, a technique called Geometric Partitioning Pattern (GPM) combined with Template Matching (TM) is used; in this example, the technique is referred to as GPM-TM.

[0123] In some examples, to apply template matching to GPM, when GPM mode is enabled for CU, a CU-level flag is signaled to indicate whether TM should be applied to both geometric partitions. TM can be used to refine the motion information for each geometric partition. When TM is selected, a template is constructed using left neighbor samples, upper neighbor samples, or a combination of left neighbor samples and upper neighbor samples, depending on the segmentation angle.

[0124] Table 1 shows the templates used for the first and second geometric partitions.

[0125] Table 1

[0126] Segmentation angle 0 2 3 4 5 8 11 12 13 14 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A Segmentation angle 16 18 19 20 21 24 27 28 29 30 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A

[0127] In Table 1, A indicates the use of the top sample, L indicates the use of the left sample, and L+A indicates the use of both the left and top samples.

[0128] In some examples, motion information is then refined by minimizing the difference between the current template and the template in the reference image using the same search pattern of the merging mode when the half-pixel interpolation filter is disabled.

[0129] In some examples, a GPM candidate list can be constructed. For instance, in the first step, interleaved lists of MV candidates 0 and MV candidates 1 are derived directly from the regular merged candidate list, where list 0 MV candidates have a higher priority than list 1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates. In the second step, interleaved lists of MV candidates 1 and MV candidates 0 are further derived directly from the regular merged candidate list, where list 1 MV candidates have a higher priority than list 0 MV candidates. The same pruning method with an adaptive threshold is applied again to remove redundant MV candidates. In the third step, zero MV candidates are filled until the GPM candidate list is full.

[0130] In some examples, GPM-MMVD and GPM-TM are exclusively enabled for a single GPM CU. In the example, signaling for the GPM-MMVD syntax is performed first. When both GPM-MMVD control flags are equal to false (e.g., GPM-MMVD is disabled for both of the two GPM partitions), the GPM-TM flag is signaled to indicate whether template matching should be applied to both GPM partitions. In other cases (e.g., at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.

[0131] In some examples, a technique known as GPM with inter-frame prediction and intra-frame prediction is used. Under GPM with inter-frame prediction and intra-frame prediction, the final prediction sample is generated by weighting the inter-frame prediction samples and intra-frame prediction samples from each GPM-separated region. The inter-frame prediction samples are derived from the inter-frame GPM, while the intra-frame prediction samples are derived from an index and an Intra Prediction Mode (IPM) candidate list signaled by the encoder. The IPM candidate list size is predefined as 3.

[0132] Figures 12A to 12D Plots of GPM with inter-frame prediction and intra-frame prediction are shown in some examples. Figures 12A to 12C A graph showing the available IPM candidates is presented. Figure 12A A diagram showing the parallel angle pattern (parallel pattern) relative to the GPM block boundary (e.g., the split boundary); Figure 12B A diagram showing the vertical angle pattern (vertical pattern) relative to the GPM block boundary (e.g., the split boundary) is provided; and Figure 12C A diagram of the planar pattern is shown. Figure 12D A diagram is shown illustrating GPM with intra-frame prediction and intra-frame prediction. In some examples, GPM with intra-frame prediction and intra-frame prediction is limited to reduce the signaling overhead of IPM and avoid increasing the size of the intra-frame prediction circuitry on the hardware decoder. Additionally, direct motion vectors and IPM storage are introduced on the GPM mixing region to further improve coding performance.

[0133] In some examples, parallel modes are registered first in both Decoder Side Intra Mode Derivation (DIMD) and neighbor-based IPM derivation. Therefore, if no identical IPM candidate exists in the list, the two largest IPM candidates from both the Decoder Side Intra Mode Derivation (DIMD) method and / or neighbor-based block derivation can be registered. Regarding neighbor-based block derivation, there are at most five available neighbor block locations, but these locations are limited by the angle of the GPM block boundaries, as shown in the error! Reference source not found. These locations have already been used for GPM combined with template matching (GPM-TM).

[0134] Table 2 shows the locations of available neighboring blocks derived from IPM candidates based on the angle of the GPM block boundary.

[0135] Table 2

[0136] GPM perspective 0 2 3 4 5 8 11 12 13 14 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A Segmentation angle 16 18 19 20 21 24 27 28 29 30 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A

[0137] In Table 2, A represents the top prediction block, L represents the left prediction block, and L+A represents the left and top prediction blocks.

[0138] In some examples, a reordering technique known as template-matching-based GPM splitting pattern can be used.

[0139] Typically, template matching (TM) is used to refine motion on the decoder side. In TM mode, motion is refined by constructing a template based on neighboring reconstructed samples to the left and / or above and finding the closest match between the template in the current image and the reference frame.

[0140] Template matching can be applied to GPM. When the CU is encoded in GPM, a decision on whether to use TM for refinement can be made based on the motion of each geometric partition. When TM is selected, a template is constructed using neighboring samples to the left and / or above, and the motion is then refined by finding the best match between the current template and a reference region in the reference frame that has the same template pattern. The refined motion is used to perform motion compensation for the geometric partitions and is stored in the motion field.

[0141] In some examples, to reduce the signaling cost of the GPM split pattern index, a reordering of the GPM split pattern index can be used by employing template matching costs. The GPM split pattern reordering method is a two-step process performed after generating the corresponding reference templates in the two GPM partitions within the coding unit. For example, the first step can use the corresponding weights of the split patterns to blend the reference templates of the two GPM partitions (e.g., generating 64 blended reference templates) and calculate the TM cost value for each blended reference template; the second step can reorder the GPM split patterns in ascending order based on the TM cost values ​​of the blended reference templates and mark the best 32 GPM split patterns as available split patterns.

[0142] In some examples, edges on the template can be obtained by extending from the (GPM split) edges of the current CU.

[0143] Figure 13 The diagram illustrates an expansion of the edges on the template by performing GPM splitting on some examples. The corresponding weights used in the template blending process are calculated using the same GPM weights derived from the process, except that the weights are mapped to 0 and 8 depending on the nearest neighbor before use.

[0144] In some examples, after reordering in ascending order using TM cost, Golomb-Rice encoding (division by 4) is used to signal the index to indicate the use of the GPM split mode.

[0145] In some examples, a technique known as GPM with bidirectional predictive motion vectors is used. In some examples (e.g., ECM (Enhanced Compression Model), GPM relies on unidirectional predictive motion vectors to generate motion-compensated prediction samples for each inter-frame partition. In some examples, the use of bidirectional predictive motion vectors in GPM is permitted. Furthermore, GPM-MMVD and GPM-TM have been modified to incorporate the use of bidirectional predictive motion vectors. GPM with bidirectional predictive motion vectors may include certain elements, such as the following four elements in some examples.

[0146] The first element conditionally invokes the extraction process to extract unidirectional predicted motion vectors from the initial list. In the example, the extraction process is invoked only for small blocks such as 8×8, 16×8, and 8×16. For other larger blocks, the extraction process is bypassed in the example. The initial list is generated in the same way as before (i.e., normal merge list generation without any candidate reordering), except that when generating the initial list for larger blocks (i.e., blocks that bypass the extraction process), the motion vector difference threshold used to control whether candidates can be added to the initial list is increased to a full sample distance.

[0147] The second element modifies GPM-MMVD to support bidirectional predictive motion vectors as base vectors. For low-latency images, MVD, signaled by the signal, is applied over the L0 and L1 motion vectors, as in existing merged MMVD designs. For non-low-latency images, the bidirectional predictive motion vectors are first converted to unidirectional predictive motion vectors, and then MVD is applied over them.

[0148] The third element modifies GPM-TM to also support bidirectional prediction of motion vectors. When the image is not a low-latency image and the best template cost obtained using bidirectional prediction exceeds 75% of the best template cost obtained using unidirectional prediction, the refined unidirectional prediction motion vector is then determined as the final refined motion vector. Otherwise, the refined bidirectional prediction motion vector is determined as the final refined motion vector.

[0149] The fourth element enables 8×8 BDOF on top of the associated bidirectional predictive motion vector for each inter-frame partition (i.e., as in the existing design of multi-pass DMVR).

[0150] In inter-frame prediction, merging modes can be used to improve coding efficiency. In merging mode, motion vectors can be derived from neighboring blocks and directly used for motion compensation. To increase the accuracy of motion vectors (MVs) in merging mode, decoder-side motion vector refinement (DMVR) based on bilateral matching (BM) can be applied, for example, in VVC. In bidirectional prediction operations, refined MVs can be searched around the initial MVs in reference image lists L0 and L1. BM calculates the distortion between two candidate blocks in reference image lists L0 and L1.

[0151] Figure 14 Examples of schematic views illustrating BM-based decoder-side motion vector refinement are shown. Figure 14 As shown, the current image (1402) may include the current block (1408). The current image may have a first reference image (1404) from (reference image) list L0 and a second reference image (1406) from (reference image) list L1. For the current block (1408), a pair of reference blocks are identified in the first and second reference images based on the initial motion vectors MV0 and MV1. For example, the initial reference block (1412) in the first reference image (1404) can be located based on the initial motion vector MV0, and the initial reference block (1414) in the second image (1406) can be located based on the initial motion vector MV1. Search processing can be performed around the initial MV0 in the first reference image (1404) and the initial MV1 in the second reference image (1406). For example, adjustments will be made. The initial MV0 and MV1 are applied in the opposite direction to obtain MV candidates, such as MV0' and MV1'. Based on the MV candidates, a pair of candidate reference blocks are identified in the first reference image and the second reference image. For example, a candidate reference block (1410) can be identified in the first reference image (1404) based on MV0', and a candidate reference block (1416) can be identified in the second reference image (1406) based on MV1'. In some examples, bilateral matching (BM) refers to the operation of calculating a distortion metric between a pair of reference blocks in the corresponding reference images of the current image, such as calculating the sum of absolute differences (SAD) between a pair of reference blocks as the distortion metric for that pair of reference blocks. For example, the BM method calculates an initial SAD between a pair of initial reference blocks (1412) and (1414), and calculates a second SAD between a pair of candidate reference blocks (1410) and (1416). The initial SAD is associated with the initial MVs (e.g., MV0 and MV1), and the second SAD is associated with the MV candidates (e.g., MV0' and MV1'). Similarly, the BM method can compute the SAD of multiple MV candidates around the initial MV. The MV candidate with the lowest SAD can become the refined MV and is used to generate a bidirectional prediction signal to predict the current block (1408).

[0152] In some examples (e.g., VVC), the application of DMVR is restricted and is only applied to CUs encoded based on patterns and features that meet certain conditions. The DMVR algorithm is invoked if a block meets certain conditions. For example, conditions (also referred to as DMVR requirements or a set of conditions for DMVR) may include: (1) a CU-level merging mode with bidirectional prediction MV; (2) one reference image is past and the other is future relative to the current image; (3) the distances (e.g., POC differences) from the two reference images to the current image are the same; (4) the two reference images are short-term reference images; (5) the CU has more than 64 luma samples; (6) both the CU height and CU width are greater than or equal to 8 luma samples; (7) the weight index of the bidirectional prediction with CU-Level Weight (BCW) with CU-level weights indicates equal weights; (8) weighted prediction (WP) is not enabled for the current block; and (9) the combined inter-frame and intra-frame prediction (CIIP) mode is not used for the current block.

[0153] Note that the refined MV obtained through DMVR processing is used to generate inter-frame prediction samples and can be used in temporal motion vector prediction for future image encoding. In some examples, the original MV is used in unblocking processing and also in spatial motion vector prediction for future CU encoding.

[0154] In DVMR, the search point revolves around the initial MV, and the MV offset follows the MV difference mirroring rule. Any point represented by a candidate MV pair (MV0', MV1') that passes DMVR verification follows this rule. and .in, This represents the thinning offset between the initial MV (e.g., (MV0, MV1)) of one of the reference images and the thinned MV. In some examples, the thinning search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample thinning phase.

[0155] In some examples (e.g., VVC), decoder-side motion vector refinement (DMVR) is applied to the CU encoded in a regular merging mode. The MV pairs obtained from the regular merging candidates are used as input to the DMVR process. DMVR applies bilateral matching (BM) to refine the input MV pairs {MV0, MV1}, and then processes the refined MV pairs {MV0, MV1}... refinedL0 MV refinedL1} Used for motion compensation prediction of both luminance and chrominance components, such as Figure 4 As shown. The output MV of DMVR can be referred to as the refined MV pair, and can be expressed by equation (4):

[0156] Equation (4)

[0157] Because the input MV pair points to two different reference images, the motion vector difference Δmv is applied to the input MV pair to obtain a refined MV pair by using the MVD mirror property. The two different reference images have an equal difference in picture order count (POC) with the current image and the two reference images are in different time directions.

[0158] In some examples, DMVR can be applied at the sub-block level, dividing the luminance-coded block into 16×16 sub-blocks for MV refinement. Δmv is derived independently for each sub-block.

[0159] In some examples, the motion vector refinement search range is two integer luminance samples from the initial MV. The search for the motion vector can be performed in two steps, namely, a first step of, for example, an integer sample offset search stage (also known as integer precision motion search) and a second step of a fractional sample refinement stage (also known as a fractional motion search step or fractional sample offset search).

[0160] In some examples, a 25-point full search can be applied to the integer sample offset search, as shown in equation (5):

[0161] Equation (5)

[0162] Where (i, j) represent the coordinates of the search points surrounding the initial MV pair, and i and j are integer values ​​between -2 and 2 (inclusive). First, calculate the SAD of the initial MV pair, for example, according to equation (6):

[0163] Equation (6)

[0164]

[0165] Where W and H are the weight and height of the sub-block.

[0166] If the SAD of the initial MV pair is less than a threshold, the integer sample phase of the DMVR terminates. Otherwise, the SAD of the remaining 24 points is calculated and verified in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search phase. To reduce the penalty for uncertainty in DMVR refinement, the original MVs (e.g., initial MV candidates MV0 and MV1) may be preferred during DMVR processing. The SAD between reference blocks referenced by the initial MV candidates is reduced, for example, by 1 / 4 of the SAD value, to make the initial MV candidates preferred.

[0167] In some examples, fractional sample refinement follows the integer sample search. In others, it is performed using fractional sample offsets, such as ½-pixel offsets in the vertical and horizontal directions. Still others, to reduce computational complexity, precipitate fractional sample refinement is derived using a parametric error surface equation (also known as a quadratic prediction-based approach) instead of an additional search with SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. For example, fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration at a specific integer position (also known as the center) with the minimum SAD.

[0168] In subpixel offset estimation based on the parametric error surface, the cost at the center location (the center location is the point with the minimum SAD in the integer sample offset search) and the costs at the four neighboring locations from the center (e.g., (-1,0), (0,-1), (1,0), (0,1)) from the center location) are used to fit the 2-D parabolic error surface equation, such as equation (7):

[0169] Equation (7)

[0170] in, The fractional position corresponds to the position with the minimum cost, and C corresponds to the minimum cost value. The above equations are solved by using the cost values ​​of the five search points, calculated according to equations (8) and (9). :

[0171] Equation (8)

[0172] Equation (9)

[0173] Since all cost values ​​are positive and the minimum value is ,therefore and The value can be automatically constrained between 8 and 8. In VVC, the constraint corresponds to a half-pixel offset with 1 / 16th pixel MV accuracy. The calculated score... The integer distance refinement MV is added to obtain the subpixel accuracy refinement increment MV. In equations (8) and (9), , , and This represents the cost value at five points (the center location and four neighboring locations).

[0174] A technique known as bidirectional optical flow (BDOF) can be used, for example, in VVC. BDOF was previously known as BIO in JEM (Joint Exploration Model). Compared to the JEM version, the BDOF in VVC can be a simpler version, requiring less computation, especially in terms of the number of multiplications and multiplier size.

[0175] BDOF can be used to refine the bidirectional prediction signal of the CU at the 4×4 sub-block level and is also referred to as conventional BDOF or sub-block-based BDOF to distinguish it from the sample-based BDOF described in this disclosure. BDOF can be applied to a CU if the following conditions are met (also referred to as the requirements of BDOF or a set of conditions for BDOF): (1) the CU is encoded using a “true” bidirectional prediction mode, i.e., one of the two reference images is displayed before the current image and the other reference image is displayed after the current image; (2) the distances (e.g., POC differences) from the two reference images to the current image are the same; (3) the two reference images are short-term reference images; (4) the CU is not encoded using an affine mode or an SbTMVP merging mode; (5) the CU has more than 64 luminance samples; (6) both the CU height and CU width are greater than or equal to 8 luminance samples; (7) the BCW weight index indicates equal weights; (8) weighted prediction (WP) is not enabled for the current CU; and (9) the CIIP mode is not used for the current CU.

[0176] In some examples, BDOF is applied only to the luma component. As the name BDOF suggests, the BDOF mode can be based on an optical flow concept that assumes the motion of the object is smooth. For each 4×4 sub-block, motion refinement can be computed by minimizing the difference between the L0 and L1 prediction samples. Then, motion refinement can be used to adjust the bidirectional prediction sample values ​​in the 4×4 sub-blocks. BDOF may include the following steps.

[0177] First, the horizontal and vertical gradients of the two predicted signals from reference lists L0 and L1 can be calculated by directly calculating the difference between two neighboring samples. and , The horizontal and vertical gradients can be provided in the following equations (10) and (11):

[0178] Equation (10)

[0179] Equation (11)

[0180] in, It can be a list , Coordinates of the predicted signal in The sample value at that location, and shift1 can be calculated based on the luminance bit depth bitDepth, such as shift1=max(6,bitDepth-6).

[0181] Then, the autocorrelation and cross-correlation of gradients S1, S2, S3, S5 and S6 can be calculated according to the following equations (12) to (16):

[0182] Equation (12)

[0183] Equation (13)

[0184] Equation (14)

[0185] Equation (15)

[0186] Equation (16)

[0187] in, , and It can be provided in equations (17) to (19) respectively.

[0188] Equation (17)

[0189] Equation (18)

[0190] Equation (19)

[0191] in, It can be a 6x6 window surrounding a 4x4 sub-block, and The value and The values ​​can be set to equal to and .

[0192] Then, using cross-correlation and autocorrelation terms, the motion refinement can be derived using the following equations (20) and (21). :

[0193] Equation (20)

[0194]

[0195] Equation (21)

[0196] in, , , . It is the floor function, and Based on motion refinement and gradients, adjustments can be calculated for each sample in the 4×4 sub-block based on equation (22):

[0197] Equation (22)

[0198] Finally, the BDOF samples of CU can be calculated by adjusting the bidirectional prediction samples in the following equation (23):

[0199] Equation (23)

[0200] Values ​​can be selected such that the multiplier in BDOF processing does not exceed 15 bits, and the maximum bit width of intermediate parameters in BDOF processing can be kept within 32 bits.

[0201] Figure 15 A graph illustrating some calculations in the bidirectional optical flow (BDOF) example is shown. To derive the gradient values, a list needs to be generated outside the current CU boundary. ( Some predicted samples in ) .like Figure 15 As shown, BDOF in VVC can use an extended row / column (1502) around the boundary (1506) of CU (1504). To control the computational complexity of generating out-of-bounds predicted samples, the extended region can be generated by directly taking reference samples at nearby integer positions (e.g., using floor() operations on coordinates) without interpolation. Figure 15 The predicted samples are located in the non-shaded areas of the CU, and a normal 8-tap motion-compensated interpolation filter can be used to generate the predicted samples within the CU (e.g., ...). Figure 15 (The shaded area in the diagram). Expanded sample values ​​can be used solely for gradient computation. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, the samples and gradient values ​​can be filled (e.g., repeated) from their nearest neighbors.

[0202] In some examples, sample-based BDOF (S-BDOF) can be used instead of block-based BDOF (also known as regular BDOF or sub-block-based BDOF). In sample-based BDOF, motion refinement is derived on a block-by-block basis instead of... Motion refinement is performed for each sample. The encoded block is divided into 8×8 sub-blocks. For each sub-block, the SAD between two reference sub-blocks is checked against a threshold to determine whether BDOF should be applied. If it is decided to apply BDOF to the sub-block, for each sample in the sub-block, a sliding 5×5 window is used and the existing BDOF processing is applied to each sliding window to obtain v. x and vy The resulting motion refinement is applied. This is used to adjust the bidirectional predicted sample values ​​of the center sample in the window.

[0203] In some examples, multiple passes of DMVR can be used. In one example, in the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16×16 sub-block within the coded block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MV is stored for both spatial motion vector prediction and temporal motion vector prediction.

[0204] Specifically, the first pass performs block-based bilateral matching MV refinement. In this first pass, the refined MV is derived by applying the BM to the coded block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, a refined MV is searched around two initial MVs (MV0 and MV1) in the reference image lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. The bilateral matching cost can be calculated using any suitable error measurement metric that measures the error between the two reference blocks in L0 and L1. In this example, the bilateral matching cost consists of a term for the sum of absolute differences (SAD) between corresponding samples in the two reference blocks in L0 and L1.

[0205] BM can perform a local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to iterate through the search range [–sHor, sHor] in the horizontal direction and the search range [–sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension and the maximum value of sHor and sVer is 8.

[0206] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW×cbH is greater than 64, the Mean Removed SAD (MRSAD) cost function is applied to remove the distortion DC effect between reference blocks. The local search intDeltaMV terminates when bilCost at the center point of the 3×3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.

[0207] The existing fractional samples are further refined to derive the final deltaMV. The refined MV after the first pass is then obtained as:

[0208] Equation (24)

[0209] Equation (25)

[0210] In the second pass, sub-block-based bilateral matching MV refinement is performed. Specifically, in the second pass, the refined MV is derived by applying the BM to 16×16 grid sub-blocks. For each sub-block, a refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference image lists L0 and L1. The refined MV (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) is derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.

[0211] For each sub-block, BM performs a full search to derive integer sample precision intDeltaMV. The full search has a search range in the horizontal direction [–sHor, sHor] and a search range in the vertical direction [–sVer, sVer], where the values ​​of sHor and sVer are determined by the block dimension and the maximum value of sHor and sVer is 8.

[0212] The bilateral matching cost is calculated as: bilCost = satdCost × costFactor, by applying a cost factor to the sum of absolute transformed differences (SATD) between the two reference sub-blocks. In some examples, the search region (2 × sHor + 1) × (2 × sVer + 1) is divided into up to 5 diamond-shaped search regions.

[0213] Figure 16 The search region (1600) is shown in some examples. The search region (1600) is divided into 5 search regions (1601) to (1605). The shape of the search regions is similar to a rhombus.

[0214] In some examples, each search region is assigned a costFactor, determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond-shaped region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order from the upper left to the lower right of the region. An integer-pixel full search terminates if the minimum bilCost within the current search region is less than or equal to a threshold of sbW × sbH; otherwise, the integer-pixel full search continues to the next search region until all search points have been checked. Additionally, the search process terminates if the difference between the previous minimum cost and the current minimum cost in an iteration is less than or equal to a threshold of the block area.

[0215] In some examples, fractional sample refinement (e.g., DMVR fractional sample refinement in VVC) is further applied to derive the final deltaMV(sbIdx2). Then, the refined MV at the second pass is derived as:

[0216] Equation (26)

[0217] Equation (27)

[0218] In the third pass, sub-block-based bidirectional optical flow (MV) refinement can be performed. Specifically, in the third pass, the refined MV is obtained by applying BDOF to 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to obtain scaled Vx and Vy, without starting clipping from the refined MV of the parent-child blocks in the second pass. Rounded to 1 / 16 of the sample precision and clipped between -32 and 32. The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) at the third pass are obtained as follows:

[0219] Equation (28)

[0220] Equation (29)

[0221] In some examples, a technique known as high-precision MV refinement for BDOF is used. In some examples, BDOF sample adjustments can yield motion refinement for 4×4 sub-blocks. And the samples are adjusted individually. In some examples (e.g., ECM), two BDOFs can be used: one as a BDOF for MV refinement (the third stage of DMVR processing) and the other as a BDOF for sample adjustment (similar to VVC, but deriving motion adjustments for each sample separately). ).

[0222] In some examples, high-precision equations such as equations (30) and (31) are used to derive the BDOF MV refinement parameters:

[0223] Equation (30)

[0224] Equation (31)

[0225] Where Gx / Gy is the sum of the two horizontal / vertical gradients derived for each reference block. It is a weighted sum, where the weights depend on the target region. The position within the range. In other cases, weights can also be applied to derive vx / vy.

[0226] Furthermore, the sub-block size of BDOF DMVR is adaptively selected based on width × height. For blocks smaller than 256, a 4×4 sub-block size is used, and in other cases, an 8×8 sub-block size is used.

[0227] Note that in some relevant examples, an 8×8 BDOF (e.g., a sub-block-based BDOF) is applied to a GPM partition with bidirectional predicted motion vectors.

[0228] Some aspects of this disclosure provide techniques for adaptively applying sample-based BDOF (S-BDOF) to GPM segmentation blocks. For example, an encoder / decoder can determine that the current block in the current image is in Geometric Segmentation Mode (GPM), and when the first GPM partition has bidirectional predicted motion vectors, determine whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least the first GPM partition of the current block. When it is determined that S-BDOF motion refinement should be applied to the first GPM partition, the encoder / decoder can apply S-BDOF motion refinement to one or more samples in the first GPM partition, and reconstruct the current block using one or more samples reconstructed based on the S-BDOF motion refinement.

[0229] Some aspects of this disclosure provide techniques for adaptively applying S-BDOF to one or two GPM partitions with bidirectional predictive motion vectors.

[0230] In some examples, S-BDOF is always applied to all GPM partitions with bidirectional predicted motion vectors. Therefore, when a GPM partition has bidirectional predicted motion vectors, S-BDOF motion refinement is applied to one or more samples within that GPM partition.

[0231] In some examples, S-BDOF is never applied to any GPM partition within a GPM partition. For instance, S-BDOF motion refinement is disabled for any GPM partition.

[0232] In some examples, S-BDOF is applied to all GPM partitions with bidirectional predicted motion vectors, but only to sub-blocks (e.g., 8×8 or 4×4) that are not adjacent to GPM partition boundaries (also known as GPM split boundaries). In the examples, when the sub-block is not in a blending region (e.g., ...), Figure 10 When a sub-block is located within the shaded area (in the image) and does not overlap with the blending region, it is considered not adjacent to the GPM partition boundary. Then, S-BDOF motion refinement can be applied to the samples within the sub-block.

[0233] In some examples, regular sub-block-based BDOF is applied only to sub-blocks that do not have GPM segmentation boundaries, while S-BDOF is applied only to sub-blocks that do have GPM segmentation boundaries. In some examples, when GPM segmentation boundaries do not intersect with sub-blocks, regular sub-block-based BDOF motion refinement can be applied to the sub-blocks. However, in some examples, when GPM segmentation boundaries intersect with sub-blocks, S-BDOF motion refinement can be applied to the samples of the sub-blocks.

[0234] In some examples, high-level flags are signaled to determine whether S-BDOF should be applied to all GPM partitions with bidirectional predictive motion vectors. For example, when the flag is 1, S-BDOF motion refinement can be applied to all GPM partitions with bidirectional predictive motion vectors; and when the flag is 0, S-BDOF motion refinement is disabled for all GPM partitions with bidirectional predictive motion vectors. High-level flags can be signaled in high-level syntaxes including, but not limited to, SPS (Sequence Parameter Set), PPS (Picture Parameter Set), PH (Picture Header), slice headers, etc.

[0235] In some examples, S-BDOF is applied to GPM partitions with bidirectional predicted motion vectors, where the number of samples in the GPM partition is greater than (or equal to) a threshold. The threshold can be a predefined value or signaled in advanced syntaxes including but not limited to SPS, PPS, PH, slice headers, etc.

[0236] In some examples, S-BDOF is applied to GPM partitions where the number of sub-blocks (e.g., 8×8 or 4×4) in the GPM partition is greater than (or equal to) a threshold. The threshold can be a predefined value or signaled in advanced syntaxes including but not limited to SPS, PPS, PH, slice headers, etc.

[0237] In some examples, S-BDOF is applied to the following GPM partitions, where the weight values ​​of the corresponding samples in the corresponding portion of the blending mask (for the GPM partition) are higher than (or equal to) a threshold. The threshold can be a predefined value or signaled in advanced syntaxes including but not limited to SPS, PPS, PH, slice headers, etc.

[0238] In some examples, S-BDOF is applied to the following GPM partitions, where the corresponding sample weight values ​​in the corresponding portion of the blending mask (for GPM partitions) are below (or equal to) a threshold. The threshold can be a predefined value or signaled in advanced syntaxes including but not limited to SPS, PPS, PH, slice headers, etc.

[0239] In some examples, whether S-BDOF is applied to GPM partitions depends on the additional block-level syntax element that determines the applicability of S-BDOF. In the examples, the additional block-level syntax element can have the following four semantics: 1) apply S-BDOF to both GPM partitions; 2) do not apply S-BDOF to both GPM partitions; 3) apply S-BDOF only to the first GPM partition; 4) apply S-BDOF only to the second GPM partition.

[0240] In another example, the appended block-level syntax element can have the following three semantics: 1) apply S-BDOF to two GPM partitions; 2) do not apply S-BDOF to two GPM partitions; 3) apply S-BDOF to the GPM partition that uses the first GPM merge index, while the other GPM partition (which uses the second GPM merge index) does not use S-BDOF motion refinement.

[0241] Figure 17 A flowchart outlining a process (1700) according to an embodiment of this disclosure is shown. The process (1700) can be used in a video decoder. In various embodiments, the process (1700) is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video decoder (110), a processing circuitry system that performs the functions of a video decoder (210), etc. In some embodiments, the process (1700) is implemented as software instructions, so that the processing circuitry system executes the process (1700) when the software instructions are executed. The process begins at (S1701) and proceeds to (S1710).

[0242] At (S1710), an encoded video bitstream including encoding information of one or more images is received.

[0243] At (S1720), based on the encoding information, it is determined that the current block in the current image is in geometric segmentation mode (GPM).

[0244] At (S1730), it is determined whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least the first GPM partition of the current block, the first GPM partition having bidirectional predicted motion vectors.

[0245] At (S1740), when it is determined that S-BDOF motion refinement will be applied to the first GPM partition, S-BDOF motion refinement will be applied to one or more samples in the first GPM partition.

[0246] At (S1750), the current block is reconstructed using one or more samples based on S-BDOF motion refinement reconstruction.

[0247] In some examples, S-BDOF motion refinement is applied to any GPM partition in the current block that has bidirectional predicted motion vectors.

[0248] In some examples, the first GPM partition is divided into sub-blocks. Then, for the first sub-block, it is determined whether the first sub-block is adjacent to the GPM's partition boundary. For example, if at least one corner of the first sub-block is in the blending region of the partition boundary (e.g., ...), then the sub-block is considered adjacent to the GPM's partition boundary. Figure 10 When a sample is within the shaded region (in the diagram), the first sub-block is considered an adjacent segmentation boundary; and when none of the samples in the first sub-block are in the mixed region, the first sub-block is considered a non-adjacent segmentation boundary. In some examples, when the first sub-block is not adjacent to a segmentation boundary, S-BDOF motion refinement is applied to the first sample in the first sub-block.

[0249] In some examples, the first GPM partition is divided into sub-blocks. Then, it is determined whether the GPM partition boundary intersects with the first sub-block within the sub-block. When the partition boundary does not intersect with the first sub-block, regular BDOF, such as sub-block-based BDOF motion refinement, is applied to the first sub-block; and when the partition boundary intersects with the first sub-block, S-BDOF motion refinement is applied to the first sub-block.

[0250] In some examples, a flag is decoded from the encoded video bitstream indicating whether S-BDOF motion refinement is applied to each partition with bidirectional predictive motion vectors. For example, a flag of 1 indicates that S-BDOF motion refinement is applied to each GPM partition with bidirectional predictive motion vectors; and a flag of 0 indicates that S-BDOF motion refinement is not applied to any GPM partition. In some examples, the flag is signaled in a high-level syntax such as the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0251] In some examples, the number of samples in the first GPM partition is compared with a threshold to obtain a comparison result. Based on the comparison result, it is determined whether S-BDOF motion refinement should be applied to the first GPM partition.

[0252] In some examples, the number of sub-blocks in the first GPM partition is compared with a threshold to obtain a comparison result. Based on the comparison result, it is determined whether S-BDOF motion refinement should be applied to the first GPM partition.

[0253] In some examples, the weight values ​​in a portion of the GPM blending mask are compared to a threshold to obtain a comparison result; this portion corresponds to a first GPM partition. Based on the comparison result, it is determined that S-BDOF motion refinement will be applied to one or more samples in the first GPM partition. In one example, it is determined that S-BDOF motion refinement will be applied to one or more samples in the first GPM partition when each weight value in the portion of the blending mask is greater than or equal to the threshold. In another example, it is determined that S-BDOF motion refinement will be applied to one or more samples in the first GPM partition when each weight value is less than or equal to the threshold.

[0254] Note that the threshold can be predefined, or it can be signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0255] In some examples, such as decoding block-level syntax elements associated with the current block from an encoded video bitstream. A block-level syntax element indicates one of the following: applying S-BDOF motion refinement to both GPM partitions of the current block; not applying S-BDOF motion refinement to any GPM partition of the current block; applying S-BDOF motion refinement only to the first GPM partition; and applying S-BDOF motion refinement only to the second GPM partition of the current block.

[0256] In some examples, such as decoding block-level syntax elements associated with the current block from an encoded video bitstream. A block-level syntax element indicates at least one of the following: applying S-BDOF motion refinement to two GPM partitions of the current block; not applying S-BDOF motion refinement to any GPM partition of the current block; and applying S-BDOF motion refinement to the GPM partition in the current block using the first GPM merge index, but not to another GPM partition of the current block.

[0257] Then, the process proceeds to (S1799) and ends.

[0258] Process (1700) can be adjusted as appropriate. Steps in process (1700) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0259] Figure 18 A flowchart outlining a process (1800) according to an embodiment of this disclosure is shown. The process (1800) can be used in a video encoder. In various embodiments, the process (1800) is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video encoder (103), a processing circuitry system that performs the functions of a video encoder (303), etc. In some embodiments, the process (1800) is implemented as software instructions, so that the processing circuitry system executes the process (1800) when the software instructions are executed. The process begins at (S1801) and proceeds to (S1810).

[0260] At (S1810), it is determined that the current block in the current image will be encoded using the geometric segmentation mode (GPM).

[0261] At (S1820), it is determined whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least the first GPM partition of the current block. The first GPM partition has bidirectional predicted motion vectors.

[0262] At (S1830), when it is determined that S-BDOF motion refinement will be applied to the first GPM partition, S-BDOF motion refinement will be applied to one or more samples in the first GPM partition.

[0263] At (S1840), the current block is reconstructed using one or more samples based on S-BDOF motion refinement reconstruction.

[0264] In some examples, to determine whether to apply S-BDOF motion refinement, S-BDOF motion refinement is applied to each GPM partition in the current block that has a bidirectional predicted motion vector.

[0265] In some examples, to apply S-BDOF motion refinement, the first GPM partition is divided into sub-blocks, and it is determined whether the first sub-block is adjacent to the GPM's segmentation boundary. In the example, when the first sub-block is not adjacent to the segmentation boundary, S-BDOF motion refinement is applied to the first sample in the first sub-block.

[0266] In some examples, to apply S-BDOF motion refinement, the first GPM partition is divided into sub-blocks. It is determined whether the GPM's partition boundary intersects with the first sub-block within the sub-block. When the partition boundary does not intersect with the first sub-block, sub-block-based BDOF motion refinement is applied to the first sub-block. When the partition boundary intersects with the first sub-block, S-BDOF motion refinement is applied to the first sub-block.

[0267] In some examples, the number of samples in the first GPM partition is compared with a threshold to obtain a comparison result. Based on the comparison result, the application of S-BDOF motion refinement is determined.

[0268] In some examples, the number of sub-blocks in the first GPM partition is compared with a threshold to obtain a comparison result. Based on the comparison result, the application of S-BDOF motion refinement is determined.

[0269] In some examples, a flag is encoded in the encoded video bitstream that includes encoding information of the current block in the current picture. This flag indicates whether S-BDOF motion refinement is to be applied to each partition with bidirectional predictive motion vectors. This flag is signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0270] In some examples, to determine whether to apply S-BDOF motion refinement, the weight values ​​in a portion of the GPM blending mask are compared with a threshold to obtain a comparison result, which corresponds to the first GPM partition. Then, based on the comparison result, it is determined whether to apply S-BDOF motion refinement to one or more samples in the first GPM partition.

[0271] In one example, S-BDOF motion refinement is applied to one or more samples in the first GPM partition when each weight value in a portion of the blending mask is greater than or equal to a threshold. In another example, S-BDOF motion refinement is applied to one or more samples in the first GPM partition when each weight value is less than or equal to a threshold.

[0272] Note that in the example, the threshold can be predefined; and in another example, the threshold is signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

[0273] In some examples, such as in an encoded video bitstream, a signal is used to notify the block-level syntax element associated with the current block. The block-level syntax element indicates one of the following: applying S-BDOF motion refinement to both GPM partitions of the current block; not applying S-BDOF motion refinement to any GPM partition of the current block; applying S-BDOF motion refinement only to the first GPM partition; and applying S-BDOF motion refinement only to the second GPM partition of the current block.

[0274] In some examples, the block-level syntax element associated with the current block is signaled in the encoded video bitstream. The block-level syntax element indicates at least one of the following: applying S-BDOF motion refinement to two GPM partitions of the current block; not applying S-BDOF motion refinement to any GPM partition of the current block; and applying S-BDOF motion refinement to the GPM partition in the current block using the first GPM merge index, but not to another GPM partition of the current block.

[0275] Then, the process proceeds to (S1899) and ends.

[0276] Process (1800) can be adjusted as appropriate. Steps in process (1800) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0277] Some aspects of this disclosure provide methods for processing video media data. These methods include processing a bitstream of visual media data according to format rules. The bitstream includes encoded information for one or more images. The format rules specify: determining whether a current block in the current image is in a geometric segmentation mode (GPM); and whether, when a first GPM partition has bidirectional predicted motion vectors, sample-based bidirectional optical flow (S-BDOF) motion refinement is applied to at least the first GPM partition in the current block. In some examples, the format rules also specify: when applying S-BDOF motion refinement to the first GPM partition, determining the spatial relationship between sub-blocks within the first GPM partition and the GPM segmentation boundary. Furthermore, the format rules specify: when a sub-block is not adjacent to a segmentation boundary, applying sub-block-based BDOF to the sub-block; and when a segmentation boundary intersects with a sub-block, applying S-BDOF motion refinement to samples within the sub-block.

[0278] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 19 A computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0279] Computer software can be coded using any suitable machine code or computer language. Machine code or computer language can be subjected to mechanisms such as assembly, compilation, and linking to create code that includes instructions. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.

[0280] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0281] Figure 19 The components shown for the computer system (1900) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in the exemplary embodiments of the computer system (1900).

[0282] Computer systems (1900) may include certain human-computer interface input devices. Such human-computer interface input devices can respond to input made by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of data gloves), audio input (e.g., speech, tapping), visual input (e.g., gestures), and olfactory input (not depicted). Human-computer interface devices can also be used to capture certain media that are not necessarily directly related to conscious input made by humans, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0283] Human-machine interface input devices may include one or more of the following (only one of each is depicted): keyboard (1901), mouse (1902), touchpad (1903), touch screen (1910), data glove (not shown), joystick (1905), microphone (1906), scanner (1907), and camera device (1908).

[0284] The computer system (1900) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: haptic output devices (e.g., haptic feedback via a touchscreen (1910), data gloves (not shown), or joystick (1905), but haptic feedback devices that are not used as input devices may also exist); audio output devices (e.g., speakers (1909), headphones (not depicted)); visual output devices (e.g., screens (1910), including CRT (Cathode Ray Tube) screens, LCD (Liquid Crystal Display) screens, plasma screens, OLED (Organic Light Emitting Diode) screens, each with or without touchscreen input capability, each with or without haptic feedback capability—some of the screens may be able to output two-dimensional visual output or more than three-dimensional output in a manner such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and ashtrays (not depicted)); and printers (not depicted).

[0285] Computer systems (1900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory, ROM) / RW (1920) with media such as CD / DVD (1921), thumb drives (1922), removable hard disk drives or solid-state drives (1923), conventional magnetic media such as magnetic tape and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit, ASIC) / PLD (Programable Logic Device, PLD) such as security dongles (not depicted), etc.

[0286] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0287] Computer systems (1900) may also include interfaces (1954) to one or more communication networks (1955). Networks can be, for example, wireless, wired, or optical. Networks can also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include: local area networks such as Ethernet and wireless LANs; cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; cable or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CANBus (Controller Area Network Bus), etc. Some networks typically require external network interface adapters that attach to certain general-purpose data ports or peripheral buses (1949) (such as, for example, the USB (Universal Serial Bus, USB) port of a computer system (1900); other networks are typically integrated into the core of the computer system (1900) by attaching to system buses as described below (e.g., to an Ethernet interface in a PC (Personal Computer, PC) system or a cellular network interface in a smartphone computer system). Using any of these networks, the computer system (1900) can communicate with other entities. Such communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., to a CANBus device), or bidirectional, such as to other computer systems using local area networks or wide area digital networks. Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0288] The human-computer interface devices, human-accessible storage devices and network interfaces mentioned above can be attached to the core (1940) of a computer system (1900).

[0289] The core (1940) may include one or more central processing units (CPU) (1941), graphics processing units (GPUs) (1942), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1943), hardware accelerators for certain tasks (1944), graphics adapters (1950), etc. These devices, along with read-only memory (ROM) (1945), random access memory (1946), and internal mass storage devices (1947) such as internal non-user-accessible hard disk drives, SSDs (Solid-State Drives), etc., can be connected via the system bus (1948). In some computer systems, the system bus (1948) may be accessed in the form of one or more physical plugs to allow for expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1948) or attached to the core's system bus (1948) via a peripheral bus (1949). In the example, a screen (1910) may be connected to a graphics adapter (1950). Peripheral bus architectures include PCI (Peripheral Component Interconnect), USB, etc.

[0290] CPUs (1941), GPUs (1942), FPGAs (1943), and accelerators (1944) can execute certain instructions that, when combined, constitute the computer code mentioned above. This computer code can be stored in ROM (1945) or RAM (1946). Transient data can also be stored in RAM (1946), while permanent data can be stored, for example, in an internal mass storage device (1947). Fast storage and retrieval to any memory device in the memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage devices (1947), ROMs (1945), RAMs (1946), etc.

[0291] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0292] By way of example and not limitation, a computer system with an architecture (1900), and in particular a core (1940), can be functionalized by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage devices as described above, and certain storage devices of the non-transitory core (1940), such as internal mass storage devices (1947) or ROM (1945). Software implementing various embodiments of this disclosure can be stored in such devices and executed by the core (1940). Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software can cause the core (1940), and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1946) and modifying such data structures according to the processes defined by the software. Alternatively or as an alternative, a computer system may be provided with functionality by means of logic hardwired or otherwise embodied in circuitry (e.g., an accelerator (1944)), which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and references to logic may also include software. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0293] The use of “at least one of…” or “one of…” in this disclosure is intended to include any one or a combination thereof of the elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include either A or B or (A and B). The use of “one of…” does not exclude any combination of the elements where applicable, such as when the elements are not mutually exclusive.

[0294] While this disclosure has described several exemplary embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. Therefore, it will be appreciated that those skilled in the art will be able to conceive of many systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.

Claims

1. A method for processing visual media data, the method comprising: The bitstream of visual media data is processed according to format rules, where: The bitstream includes encoded information for one or more images, the encoded information indicating that the current block in the current image is in Geometric Partitioning (GPM) mode; and The formatting rules specify: Encode the current block in the current image using Geometric Partitioning (GPM); Whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least a first GPM partition of the current block, the first GPM partition having bidirectional predicted motion vectors; When the S-BDOF motion refinement is applied to the first GPM partition, the spatial relationship between the sub-blocks in the first GPM partition and the GPM partitioning boundary is determined. When the sub-block is not adjacent to the segmentation boundary, the sub-block-based BDOF is applied to the sub-block; and When the segmentation boundary intersects with the sub-block, the S-BDOF motion refinement is applied to the samples in the sub-block.

2. An apparatus for video decoding, the apparatus comprising a processing circuit system configured to: Receive encoded video bitstreams that include encoded information for one or more images; Based on the encoded information, it is determined that the current block in the current image is in geometric segmentation mode (GPM). Determine whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least a first GPM partition of the current block, the first GPM partition having bidirectional predicted motion vectors; When it is determined that the S-BDOF motion refinement should be applied to the first GPM partition, the S-BDOF motion refinement should be applied to one or more samples in the first GPM partition; and The current block is reconstructed using one or more samples based on the S-BDOF motion refinement reconstruction.

3. The apparatus according to claim 2, wherein, The processing circuit system is configured to: Determine whether to apply the S-BDOF motion refinement to each GPM partition in the current block that has a bidirectional predicted motion vector.

4. The apparatus according to claim 2, wherein, The processing circuit system is configured to: Divide the first GPM partition into sub-blocks; Determine whether the first sub-block is adjacent to the partition boundary of the GPM; and When the first sub-block is not adjacent to the segmentation boundary, the S-BDOF motion refinement is applied to the first sample in the first sub-block.

5. The apparatus according to claim 2, wherein, The processing circuit system is configured to: Divide the first GPM partition into sub-blocks; Determine whether the GPM's segmentation boundary intersects with the first sub-block in the sub-block; When the segmentation boundary does not intersect with the first sub-block, the sub-block-based BDOF motion refinement is applied to the first sub-block; as well as When the segmentation boundary intersects with the first sub-block, the S-BDOF motion refinement is applied to the first sub-block.

6. The apparatus according to any one of claims 2 to 3, wherein, The processing circuit system is configured to: Decode a flag from the encoded video bitstream, the flag indicating whether the S-BDOF motion refinement is applied to each partition with bidirectional predictive motion vectors, the flag being signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

7. The apparatus according to claim 2, wherein, The processing circuit system is configured to: The number of samples in the first GPM partition is compared with a threshold to obtain the comparison result; and Based on the comparison results, the application of the S-BDOF motion refinement is determined.

8. The apparatus according to claim 2, wherein, The processing circuit system is configured to: The number of sub-blocks in the first GPM partition is compared with a threshold to obtain a comparison result; and Based on the comparison results, the application of the S-BDOF motion refinement is determined.

9. The apparatus according to claim 2, wherein, The processing circuit system is configured to: The weight values ​​in a portion of the GPM blending mask are compared with a threshold to obtain a comparison result, the portion corresponding to the first GPM partition; as well as Based on the comparison results, it is determined that the S-BDOF motion refinement will be applied to one or more samples in the first GPM partition.

10. The apparatus according to claim 9, wherein, The processing circuit system is configured to: When each of the weight values ​​is higher than or equal to the threshold, it is determined that the S-BDOF motion refinement will be applied to one or more samples in the first GPM partition.

11. The apparatus according to claim 9, wherein, The processing circuit system is configured to: When each of the weight values ​​is less than or equal to the threshold, it is determined that the S-BDOF motion refinement will be applied to one or more samples in the first GPM partition.

12. The apparatus according to any one of claims 7 to 11, wherein, The threshold is predefined, or signaled in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), and Slice Header.

13. The apparatus according to claim 2, wherein, The processing circuit system is configured to: Decode the block-level syntax element associated with the current block, the block-level syntax element indicating at least one of the following: The S-BDOF motion refinement is applied to the two GPM partitions of the current block; The S-BDOF motion refinement is not applied to any GPM partition of the current block; The S-BDOF motion refinement is applied only to the first GPM partition; and The S-BDOF motion refinement is applied only to the second GPM partition of the current block.

14. The apparatus according to claim 2, wherein, The processing circuit system is configured to: Decode the block-level syntax element associated with the current block, the block-level syntax element indicating at least one of the following: The S-BDOF motion refinement is applied to the two GPM partitions of the current block; The S-BDOF motion refinement is not applied to any GPM partition of the current block; and The S-BDOF motion refinement is applied to the GPM partitions in the current block that use the first GPM merge index, but not to another GPM partition in the current block.

15. A method for video encoding, comprising: Determine the encoding method for the current block in the current image using the Geometric Partitioning Mode (GPM); Determine whether to apply sample-based bidirectional optical flow (S-BDOF) motion refinement to at least a first GPM partition of the current block, the first GPM partition having bidirectional predicted motion vectors; When it is determined that the S-BDOF motion refinement should be applied to the first GPM partition, the S-BDOF motion refinement should be applied to one or more samples in the first GPM partition; and The current block is reconstructed using one or more samples based on the S-BDOF motion refinement reconstruction.