Bidirectional optical flow for GPM with bidirectional predictive motion vector

By employing a geometric partitioning pattern and a bidirectional motion vector prediction method in video coding, and utilizing bidirectional optical flow and decoder-side motion vector refinement techniques, the problem of inaccurate motion vector prediction in existing technologies is solved, thereby improving video coding efficiency and quality.

CN121533022APending Publication Date: 2026-02-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480027226.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-24
Filing Date
2024-04-22
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively utilize spatial and temporal redundancy for efficient compression when processing video data. In particular, the accuracy and refinement of motion vector prediction are insufficient in inter-frame prediction, resulting in low coding efficiency.

Method used

The block partitioning method is adopted using geometric partitioning mode (GPM), and bidirectional predictive motion vectors and sub-block-based motion refinement techniques are used, including bidirectional optical flow (BDOF) and decoder-side motion vector refinement (DMVR). The sub-block size is determined by hybrid masking and the corresponding refinement method is applied to improve the prediction accuracy of motion vectors.

Benefits of technology

It improves the efficiency and quality of video encoding, reduces the amount of data, enhances the accuracy of motion vector prediction, and improves encoding efficiency and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121533022A_ABST
    Figure CN121533022A_ABST
Patent Text Reader

Abstract

Some aspects of the present disclosure provide a video decoding apparatus. The apparatus comprises processing circuitry configured to: receive an encoded video bitstream comprising encoded information for one or more pictures; determining, from the encoded information, that a current block in the current picture is in a geometric partition mode (GPM), wherein at least a first GPM partition has a bidirectional prediction motion vector; and applying sub-block-based motion refinement with bidirectional motion to the first sub-block and the second sub-block of at least the first GPM partition. The first sub-block and the second sub-block have different sub-block sizes. Processing circuitry reconstructs the current block according to sub-block-based motion refinement.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 461,580, filed April 24, 2023, entitled “Bi-Directional Optical Flow on GPM with Bi-Predictive Motion Vector,” which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure generally describes implementation methods related to video encoding and decoding. Background Technology

[0004] The background description provided herein is for the purpose of presenting the overall context of this disclosure. To the extent that the work described in this background section is presented, neither the work of the currently identified inventors nor any aspect of the description that may not have been otherwise considered prior art at the time of filing is expressly or implicitly acknowledged as prior art to this disclosure.

[0005] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. For instance, a video codec can use a technique called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image in the reconstruction to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by motion vectors (MV). Summary of the Invention

[0006] This disclosure includes methods and apparatus for video encoding / decoding.

[0007] Some aspects of this disclosure provide a method for processing visual media data. The method includes processing a bitstream of visual media data according to format rules. The bitstream includes encoded information of one or more images, the encoded information indicating that a current block in the current image is encoded in a geometric partitioning (GPM) mode, wherein at least a first GPM partition has bidirectional predictive motion vectors. The format rules specify: splitting the current block into GPM partitions according to the GPM mode; the first GPM partition having bidirectional predictive motion vectors; and dividing the first GPM partition into larger sub-blocks of size N×N, where N is a positive number and the N×N size is greater than or equal to the maximum sub-block size supported by sub-block-based motion refinement. The format rules also specify: determining the sub-block size of each of the larger N×N sub-blocks based on a mixing mask of the current block; dividing the larger sub-blocks into sub-blocks based on their respective sub-block sizes; and applying sub-block-based motion refinement to a specific sub-block based on the value of the mixing mask in that specific sub-block. Sub-block-based motion refinement is one of bidirectional optical flow (BDOF) motion refinement and decoder-side motion vector refinement (DMVR) refinement.

[0008] Some aspects of this disclosure provide a video decoding apparatus. The apparatus includes a processing circuitry configured to: receive an encoded video bitstream comprising encoded information of one or more images; determine, based on the encoded information, that a current block in the current image is in a geometric partitioning mode (GPM), wherein at least a first GPM partition has bidirectional predictive motion vectors; and apply sub-block-based motion refinement with bidirectional motion to a first sub-block and a second sub-block of at least the first GPM partition. The first and second sub-blocks have different sub-block sizes. The processing circuitry reconstructs the current block based on the sub-block-based motion refinement.

[0009] In some examples, sub-block-based motion refinement is bidirectional optical flow (BDOF) motion refinement, where the first sub-block is a first BDOF sub-block and the second sub-block is a second BDOF sub-block.

[0010] In some examples, the processing circuitry is configured to divide the first GPM partition into larger sub-blocks of size N×N, where N is a positive number and the N×N size is greater than or equal to the maximum supported BDOF sub-block size. The processing circuitry is also configured to: determine the corresponding BDOF sub-block size for each of the larger N×N sub-blocks; divide the larger sub-blocks into BDOF sub-blocks based on the corresponding BDOF sub-block sizes; and apply BDOF motion refinement to the BDOF sub-blocks.

[0011] In some examples, the processing circuitry is configured to determine the supported BDOF subblock size based on the size of the current block.

[0012] In some examples, the processing circuitry is configured to: determine a first BDOF sub-block size for the first larger sub-block based on the value of the mixing mask in the first larger sub-block of size N×N; divide the first larger sub-block into first BDOF sub-blocks based on the first BDOF sub-block size; and apply BDOF motion refinement to the first BDOF sub-blocks.

[0013] In the example, the processing circuitry is configured to set the size of the first BDOF sub-block to the supported maximum BDOF sub-block size when all mask values ​​in the first larger sub-block correspond to either the maximum or minimum weight value. In the example, the processing circuitry is configured to set the size of the first BDOF sub-block to be smaller than the supported maximum BDOF sub-block size and greater than or equal to the supported minimum BDOF sub-block size when no mask value in the first larger sub-block corresponds to the maximum weight value. In the example, the processing circuitry is configured to set the size of the first BDOF sub-block to be smaller than the supported maximum BDOF sub-block size and greater than or equal to the supported minimum BDOF sub-block size when no mask value in the first larger sub-block corresponds to the minimum weight value. In the example, the processing circuitry is configured to set the size of the first BDOF sub-block to be smaller than the supported maximum BDOF sub-block size and greater than or equal to the supported minimum BDOF sub-block size when no mask value in the first larger sub-block corresponds to either the maximum or minimum weight value.

[0014] In some examples, the processing circuitry is configured to determine whether to apply BDOF refinement to that portion of the first larger sub-block based on the mask value in at least a portion of the first larger sub-block.

[0015] In some examples, the processing circuitry is configured to: determine to apply BDOF refinement to the portion of the first larger sub-block when all mask values ​​in that portion correspond to either the maximum or minimum weight value; determine to apply BDOF refinement to the portion of the first larger sub-block when no mask value in that portion is zero; determine to apply BDOF refinement to the portion of the first larger sub-block when all mask values ​​in that portion are above a threshold; and determine to apply BDOF refinement to the portion of the first larger sub-block when all mask values ​​in that portion are below a threshold.

[0016] In some examples, sub-block-based motion refinement is decoder-side motion vector refinement (DMVR) refinement, where the first sub-block is the first DMVR sub-block and the second sub-block is the second DMVR sub-block.

[0017] In some examples, the processing circuitry is configured to: divide the first GPM partition into multiple sub-blocks; and determine whether to apply DMVR refinement to a particular sub-block based on the mask value of the mixing mask in that particular sub-block.

[0018] In some examples, the processing circuitry is configured to determine the supported DMVR sub-block size based on the size of the current block.

[0019] In some examples, the processing circuitry is configured to determine when to apply DMVR refinement to a particular sub-block if all mask values ​​in that sub-block correspond to either the maximum or minimum weight value.

[0020] In some examples, the processing circuitry is configured to apply multiple passes of DMVR to a specific sub-block when it is determined that DMVR refinement should be applied to that sub-block. In one example, the processing circuitry is configured to determine the application of multiple passes of DMVR based on the GPM split pattern index.

[0021] In the example, the processing circuitry is configured to determine the application of DMVR multiple times when the GPM angle is either horizontal or vertical.

[0022] In some examples, the processing circuitry is configured to check whether the GPM partition boundary intersects with a specific sub-block, apply sub-block-based motion refinement with bidirectional motion to the specific sub-block when the GPM partition boundary does not intersect with the specific sub-block, and disable sub-block-based motion refinement for the specific sub-block when the GPM partition boundary intersects with the specific sub-block.

[0023] Some aspects of this disclosure provide a video coding method. The method includes: determining that a GPM mode is to be used for a current block in a current image; determining that a first GPM partition has bidirectional predicted motion vectors; and dividing the first GPM partition into larger sub-blocks of size N×N, where N is a positive number and the size N×N is greater than or equal to the maximum sub-block size supported by sub-block-based motion thinning. The method further includes: determining the sub-block size of each of the larger N×N sub-blocks based on a blending mask of the current block; dividing the larger sub-blocks into sub-blocks based on their respective sub-block sizes; and determining whether to apply sub-block-based motion thinning to a specific sub-block based on the value of the blending mask in that specific sub-block.

[0024] In some examples, sub-block-based motion refinement is bidirectional optical flow (BDOF) motion refinement.

[0025] In some examples, sub-block-based motion refinement is decoder-side motion vector refinement (DMVR) refinement.

[0026] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes a processing circuitry system. The processing circuitry system can be configured to perform any of the described methods for video decoding / encoding.

[0027] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. Attached Figure Description

[0028] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0029] Figure 1 It is a schematic diagram of an exemplary block diagram of a communication system.

[0030] Figure 2 This is a schematic illustration of an exemplary block diagram of a decoder.

[0031] Figure 3 This is a schematic illustration of an exemplary block diagram of an encoder.

[0032] Figure 4 The locations of spatial merging candidates according to an embodiment of this disclosure are shown.

[0033] Figure 5 Candidate pairs for redundancy checking of spatial merging candidates are shown according to an embodiment of this disclosure.

[0034] Figure 6 An exemplary motion vector scaling for time merging candidates is shown.

[0035] Figure 7 An exemplary candidate position for the time merging candidate of the current block is shown.

[0036] Figure 8 A diagram showing angles used in Geometric Partitioning (GPM) mode in some examples is shown.

[0037] Figure 9 A diagram showing possible partition edges in the example is provided.

[0038] Figure 10 A diagram illustrating the hybrid processing in some examples is shown.

[0039] Figure 11 A plot of the weights for GPM mixing is shown, based on a ramp function of displacement from the predicted sample location to the GPM partition boundary and the size of the mixing region.

[0040] Figures 12A to 12DPlots of GPM using inter-frame prediction and intra-frame prediction are shown in some examples.

[0041] Figure 13 The diagram illustrates, in some examples, the expansion of the GPM split edge to obtain the edge on the template.

[0042] Figure 14 Exemplary schematic diagrams of decoder-side motion vector refinement based on bilateral matching are shown in some examples.

[0043] Figure 15 A graph showing some calculations in the bidirectional optical flow (BDOF) example is presented.

[0044] Figure 16 The search areas are shown in some examples.

[0045] Figure 17 A flowchart outlining the decoding process according to some embodiments of this disclosure is shown.

[0046] Figure 18 A flowchart is shown that outlines the coding process according to some embodiments of this disclosure.

[0047] Figure 19 It is a schematic diagram of a computer system according to an implementation method. Detailed Implementation

[0048] Figure 1 Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an example of the application of the disclosed subject matter, video encoder, and video decoder in a streaming environment. The disclosed subject matter can also be applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs (Compact Discs), DVDs (Digital Versatile Discs), memory sticks, etc.

[0049] The video processing system (100) includes a capture subsystem (113) that may include a video source (101), such as a digital camera device, which creates, for example, an uncompressed video picture stream (102). In the example, the video picture stream (102) includes samples captured by the digital camera device. The video picture stream (102) is depicted as a thick line to emphasize the high data volume when compared with encoded video data (104) (or encoded video bitstream), which may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize its lower data volume when compared to the video picture stream (102). This encoded video data (104) (or encoded video bitstream) can be stored on a streaming server (105) for future use. One or more streaming client subsystems, for example... Figure 1 Client subsystems (106) and (108) can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of encoded video data and creates an outgoing video picture stream (111) that can be presented on a display (112) (e.g., a screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., video bitstream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T (International Telecommunication Union-Telecommunication Standardization Sector, ITU-T) Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The topics that are exposed can be used in the context of VVC.

[0050] Note that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may also include a video encoder (not shown).

[0051] Figure 2An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuitry system). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the example.

[0052] The receiver (231) can receive, for example, one or more encoded video sequences included in a bitstream, to be decoded by the video decoder (210). In an implementation, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective user entities (not depicted). The receiver (231) can separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) can be external to the video decoder (210) (not depicted). In other applications, a buffer memory (not depicted) may exist outside the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may exist inside the video decoder (210) to handle broadcast timing, for example. The buffer memory (215) may not be necessary, or may be small, when the receiver (231) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous synchronization network. For the purpose of utilizing packet networks such as the Internet, a buffer memory (215) may be required, which may be relatively large and advantageously have an adaptive size, and may be implemented at least partially in the operating system or in a similar element (not depicted) outside the video decoder (210).

[0053] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (210) and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to it, such as... Figure 2 As shown. The control information used for presenting the device can be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of the encoded video sequence based on at least one parameter corresponding to the group. The subgroup can include Group of Picture (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0054] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0055] Depending on the type of encoded video images or a subset of encoded video images (e.g., inter-frame images and intra-frame images, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (220). For clarity, the flow of such subgroup control information between the parser (220) and the following multiple units is not depicted.

[0056] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0057] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) and control information from the parser (220), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values, which can be input to the aggregator (255).

[0058] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to intra-coded blocks. Intra-coded blocks are blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information obtained from the current picture buffer (258) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current picture buffer (258) buffers the partially reconstructed current image and / or the fully reconstructed current image. In some cases, the aggregator (255) adds the prediction information already generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.

[0059] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to inter-frame encoded blocks and potentially to motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols (221) belonging to the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The address in the reference image memory (257) from which the motion compensation prediction unit (253) obtains the predicted samples can be controlled by motion vectors, which can be obtained by the motion compensation prediction unit (253) in the form of symbols (221), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values ​​obtained from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0060] The output samples of the aggregator (255) can undergo various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, as well as sample values ​​in response to previous reconstruction and loop filtering.

[0061] The output of the loop filter unit (256) can be a sample stream, which can be output to the presentation device (212) and stored in the reference image memory (257) for future inter-frame image prediction.

[0062] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and that encoded image (by, for example, the parser (220)) has been identified as the reference image, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before reconstructing subsequent encoded images begins.

[0063] The video decoder (210) can perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T Recommendation H.265. An encoded video sequence can conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and in the sense of a configuration file documented in the video compression technology or standard. Specifically, the configuration file can select certain tools from all available tools in the video compression technology or standard as tools that are only available under that configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata managed by the HRD buffer used for signaling in the encoded video sequence.

[0064] In this implementation, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The video decoder (210) may use the additional data to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be, for example, in the form of temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0065] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuitry system). The video encoder (303) can be used in place of... Figure 1 The video encoder (103) in the example.

[0066] The video encoder (303) can obtain data from the video source (301) (which is located in...). Figure 3 In one example, the video source (301), which is not part of the electronic device (320), receives a video sample. This video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0067] The video source (301) can provide a sequence of source video in the form of a digital video sample stream to be encoded by the video encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera device capturing local image information as a video sequence. The video data can be provided as multiple individual pictures that are given motion when viewed sequentially. The pictures themselves can be organized as spatial pixel arrays, where each pixel can include one or more samples, depending on the sampling structure, color space, etc., used. The following description focuses on samples.

[0068] According to the implementation, the video encoder (303) can encode and compress images of the source video sequence into an encoded video sequence (343) in real time or under any other time constraints as required. Implementing an appropriate encoding rate is a function of the controller (350). In some implementations, the controller (350) controls and is functionally coupled to other functional units as described below. For brevity, the coupling is not depicted. Parameters set by the controller (350) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions belonging to the video encoder (303) optimized for a particular system design.

[0069] In some implementations, the video encoder (303) is configured to operate within an encoding / decoding loop. As a simplified description, in this example, the encoding / decoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder would also create them. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-accurate between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values ​​the decoder will "see" when using prediction during decoding. The basic principle of reference image synchronicity (and drift that occurs, for example, due to channel errors) is also used in some related techniques.

[0070] The operation of the "local" decoder (333) can be combined with that of the "remote" decoder, as already mentioned above. Figure 2 The operation of the video decoder (210) described in detail is the same. However, a brief reference is also provided. Figure 2 Since the symbols are available and the encoding of the symbols into an encoded video sequence by the entropy encoder (345) and the decoding of the symbols by the parser (220) can be lossless, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), can be fully implemented in the local decoder (333).

[0071] In the implementation, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. Since the encoder techniques are inverses of the fully described decoder techniques, the description of the encoder techniques can be simplified. A more detailed description is provided below in certain sections.

[0072] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding, which predictively codes the input image by referencing one or more previously encoded images from the video sequence designated as "reference images." In this way, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as the predictive reference for the input image.

[0073] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture based on symbols created by the source encoder (330). The operation of the encoding engine (332) can be advantageously for lossy processing. When the encoded video data can be decoded by the video decoder (333), Figure 3 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process performed on the reference image by the video decoder and can store the reconstructed reference image in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission error) as the reconstructed reference image to be obtained by the remote video decoder.

[0074] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or specific metadata such as reference image motion vectors, block shapes, etc. that can be used as appropriate prediction references for the new image. The predictor (335) can operate pixel-by-pixel based on the sample blocks to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).

[0075] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0076] The outputs of all the functional units mentioned above can be entropy encoded in the entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into an encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0077] The transmitter (340) may cache encoded video sequences, such as those created by the entropy encoder (345), in preparation for transmission via a communication channel (360), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) may combine encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0078] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, which may affect the encoding techniques that can be applied to the corresponding image. For example, an image can typically be assigned to one of the following image types:

[0079] Intra-frame pictures (I-pictures) can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0080] Predictive images (P-images) can be encoded and decoded using intra-frame or inter-frame prediction that uses motion vectors and reference indices to predict sample values ​​for each block.

[0081] Bidirectional predictive images (B-images) can be encoded and decoded using intra-frame or inter-frame predictions that predict sample values ​​for each block using two motion vectors and reference indices. Similarly, multiple predictive images can be used for the reconstruction of a single block using more than two reference images and associated metadata.

[0082] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded on a block-by-block basis. These blocks can be predictively encoded with reference to other (already encoded) blocks, determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of image I can be non-predictively encoded, or blocks of image I can be predictively encoded (spatial or intra-frame prediction) with reference to already encoded blocks of the same image. Pixel blocks of image P can be predictively encoded with reference to a previously encoded reference image via spatial or temporal prediction. Blocks of image B can be predictively encoded with reference to one or two previously encoded reference images via spatial or temporal prediction.

[0083] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In the operation of the video encoder (303), various compression operations can be performed, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.

[0084] In this implementation, the transmitter (340) may transmit additional data along with the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, SEI messages, VUI parameter set fragments, etc.

[0085] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simply called intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, a specific image during encoding / decoding—referred to as the current image—is partitioned into blocks. Where a block in the current image resembles a reference block in a previously encoded and still cached reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference images.

[0086] In some implementations, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image that both precede the current image in the video in decoding order (but may be past and future in display order, respectively). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.

[0087] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0088] According to some embodiments of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC (High-Efficiency Video Coding) standard, pictures in a video picture sequence are partitioned into Coding Tree Units (CTUs) for compression. The CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU comprises three Coding Tree Blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively split into one or more Coding Units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type used for that CU, such as inter-frame prediction or intra-frame prediction. Based on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In implementations, prediction operations in encoding / decoding are performed on a block-by-block basis. Using a luma prediction block as an example, a prediction block comprises a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0089] Note that any suitable technology can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In one implementation, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another implementation, one or more processors that execute software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).

[0090] Various inter-frame prediction modes can be used for video coding. For example, in VVC, for an inter-frame prediction CU, motion parameters may include MV (Motion Vector), one or more reference picture indices, reference picture list usage indices, and additional information to be used for generating specific coded features for inter-frame prediction samples. Motion parameters can be signaled explicitly or implicitly. When a CU is encoded in skip mode, the CU may be associated with a PU and may not have valid residual coefficients, encoded motion vector increments or MV differences (e.g., MVD (MV Difference)), or reference picture indices. A merge mode can be specified, where the motion parameters of the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally include additional information such as that introduced in VVC. The merge mode can be applied to inter-frame prediction CUs, not just skip mode. In the example, an alternative to the merge mode is explicit transmission of motion parameters, where the MV, the corresponding reference picture index for each reference picture list, and reference picture list usage flags and other information are explicitly signaled for the CU.

[0091] In implementations, such as in VVC, the VVC Test Model (VTM) reference software includes one or more refined inter-frame predictive coding tools, including: Extended Combined Prediction, Merge Motion Vector Difference (MMVD) mode, Adaptive Motion Vector Prediction (AMVP) mode with symmetric MVD signaling, Affine Motion Compensation Prediction, Subblock-based Temporal Motion Vector Prediction (SbTMVP), Adaptive Motion Vector Resolution (AMVR), Motion Field Storage (1 / 16th Luminance Sample MV Storage and 8×8 Motion Field Compression), Bi-Prediction with CU-Level Weight (BCW), Bi-Directional Optical Flow (BDOF), Prediction Refinement using Optical Flow (PROF), and Decoder Side Motion Vector Refinement. Refinement (DMVR), Combined Inter and Intra Prediction (CIIP), Geometric Partitioning Mode (GPM), etc. Inter-frame prediction and related methods are described in detail below.

[0092] In some examples, extended merge predictions can be used. In examples, such as VTM4, the merge candidate list is constructed by sequentially including the following five types of candidates: spatial motion vector predictors (MVPs) from spatially neighboring CUs, temporal MVPs from co-located CUs, history-based MVPs (HMVPs) from first-in-first-out (FIFO) tables, pairwise average MVPs, and zero MVs.

[0093] The size of the merge candidate list can be signaled in the slice header. In the example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU encoded in merge mode, truncated unary binary binarization (TU) can be used to encode the index of the best merge candidate (e.g., the merge index). The first CU of the merge index can be encoded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for other CUs.

[0094] Below are some examples of the processing for generating merge candidates for each category. In the implementation, spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as the derivation of spatial merge candidates in HEVC. In the example, in the location... Figure 4 Select up to four merged candidates from the candidates depicted at the location.

[0095] Figure 4 The locations of spatial merging candidates according to an embodiment of this disclosure are shown. (Refer to...) Figure 4 The derived order is B1, A1, B0, A0, and B2. Position B2 is considered only if any CU at positions A0, B0, B1, and A1 is unavailable (e.g., because the CU belongs to another slice or another tile) or is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates undergoes redundancy checking, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0096] To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only... Figure 5 The pairs connected by arrows are added to the candidate list only if the corresponding candidates used for redundancy check do not have the same motion information.

[0097] Figure 5 Candidate pairs considered for redundancy checking in spatial merging candidates according to an embodiment of this disclosure are shown. (Refer to...) Figure 5 The pairs connected by the corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0098] In this implementation, time candidates are exported as follows. In the example, only one time merging candidate is added to the candidate list. Figure 6 An exemplary motion vector scaling for temporal merging candidates is shown. To derive temporal merging candidates for the current CU (611) in the current image (601), the scaled MV (621) can be derived based on the co-located CU (612) belonging to the co-located reference image (604) (e.g., by...). Figure 6 (As shown by the dashed line in the image). The list of reference images used to export the corresponding CU (612) can be explicitly indicated by signals in the slice header. For example... Figure 6 As shown by the dashed lines, the scaled MV (621) for time merging candidates can be obtained. The scaled MV (621) can be scaled from the MV of the co-occurring CU (612) using the Picture Order Count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) and the current picture (601). The POC distance td can be defined as the POC difference between the co-occurring reference picture (604) and the co-occurring picture (603). The reference picture index of the time merging candidate can be set to zero. The co-occurring picture is the reference picture used as the source picture for inference of temporal motion information. The co-occurring picture can be identified in one of two lists called List 0 or List 1. In some examples, the encoder can determine the co-occurring picture and use appropriate syntax techniques to signal the co-occurring picture.

[0099] Figure 7 Exemplary candidate positions (e.g., C0 and C1) for the current CU's time-merging candidate are shown. The position of the time-merging candidate can be selected between candidate positions C0 and C1. Candidate position C0 is located at the lower right corner of the current CU's sibling CU (710). Candidate position C1 is located at the center of the current CU's sibling CU (710). If the CU at candidate position C0 is unavailable, intra-coded, or outside the current line of the CTU, candidate position C1 is used to derive the time-merging candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and in the current line of the CTU, candidate position C0 is used to derive the time-merging candidate. The time-merging candidate can specify motion information from the Temporal Motion Vector Predictor (TMVP).

[0100] In some examples (e.g., VVC), a technique called Geometric Partitioning Mode (GPM) is used. Specifically, in VVC, GPM is used for inter-frame prediction. In the examples, GPM is only applied to CUs of 8×8 or larger. GPM as a merging mode can be signaled using CU-level flags, while other merging modes include regular merging mode, merged motion vector difference (MMVD) mode, combined inter- and intra-frame prediction (CIIP) mode, and sub-block merging mode.

[0101] When using GPM mode for a CU, one of several partitioning methods is used to split the CU into two geometrically shaped partitions via partition edges. In some examples, 64 different partitioning methods are used. Partitioning methods can be distinguished by 24 angles (non-uniformly quantized between 0 and 360°) and up to four edges for each angle relative to the center of the CU. Partition edges are lines that intersect the boundary of the CU and split the CU into two partitions.

[0102] Figure 8 A diagram showing 24 angles used under GPM in some examples is illustrated. In some examples, angles may be identified using angle indices, such as angle indices 0 to 23.

[0103] Figure 9 A diagram showing possible partition edges for angle index 3 in the example is provided. Figure 9 In this context, four possible partition edges can be associated with angle index 3. Note that for some angle indices, three possible partition edges may be associated with each angle index.

[0104] In some examples, the motion of each geometric partition within a CU is used for inter-frame prediction. In these examples, only unidirectional prediction is allowed per partition; that is, each partition has one motion vector and one reference image index. Unidirectional prediction motion constraints are applied to ensure—similar to bidirectional prediction—that each CU requires two motion-compensated predictions.

[0105] In some examples, when using GPM for the current CU, signals are further used to indicate the geometric partition index (e.g., indicating angles and edges) and two merge indexes (one for each partition). In the example, the maximum number of GPM candidate sizes is explicitly signaled at the slice level, and syntax binarization for the GPM merge index is specified.

[0106] In some examples, after predicting each of the two geometric partitions, a blending process with adaptive weights is used to adjust the sample values ​​along the edges of the geometric partitions.

[0107] Figure 10The diagram shows some examples of the blending process. The blending process uses a parameter called the blending intensity or blending region width θ, which is called GPM. For all the different content, the blending intensity θ can be fixed.

[0108] In some examples, the weight values ​​in the blending mask used for blending can be given by a ramp function, for example, according to formula (1):

[0109] Formula (1)

[0110] In the example, with a fixed θ = 2 image elements, the ramp function can be quantized according to formula (2):

[0111] Formula (2)

[0112] The result of the mixing process is used as the prediction signal for the entire CU, and transform and quantization processes can be further applied to the entire CU in a manner similar to that used in other prediction modes. Finally, the motion field (e.g., motion information) of the CU predicted using GPM is stored. In some examples (e.g., VVC), the motion information of the CU is stored in 4×4 cells (e.g., for each 4×4 lumen sample). The stored motion information is used for MV prediction and merging list construction for the next encoded CU. In GPM, three types of motion information are covered and stored in 4×4 cells. The three types of motion information can include motion information from two partitions and motion information from the mixed region. For example, two geometric partitions P0 and P1 include their own unidirectional MV, and the mixed region between P0 and P1 is predicted using motion information from the two geometric partitions P0 and P1. Therefore, the motion information of GPM is stored according to partitions.

[0113] In merge mode, motion information for the GPM is signaled. To avoid additional memory bandwidth access, only unidirectional prediction is allowed per partition under GPM. In some examples, regular merge candidates can be either unidirectional or bidirectional predictions and cannot be directly used as the GPM merge list. To minimize implementation complexity, an index parity-based approach can be used to directly extract GPM merge candidates from the regular merge list without pruning. For example, for candidates with even-valued GPM merge indices, MV0 from reference list 0 and its corresponding regular merge index are used as GPM merge candidates. If MV0 is unavailable, MV1 from reference list 1 is used instead. Conversely, MV1 is selected as the default GPM merge candidate for odd-valued GPM merge indices.

[0114] Based on some aspects of this disclosure, additional technologies for GPM beyond VVC have also been developed.

[0115] In some examples, a technique known as Geometric Partitioning Mode (GPM) with Merged Motion Vector Difference (MMVD) can be used, also referred to as GPM-MMVD. GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. For example, the CUs under GPM are first signaled with a flag to indicate whether GPM mode is used. When using mode GPM, each geometric partition of the CU under GPM can decide whether to signal MVD. When MVD is signaled for a geometric partition, the motion of the partition is further refined using the signaled MVD information after selecting a GPM merge candidate. All other procedures remain the same as in GPM.

[0116] In some examples, the MVD in GPM is signaled as a pair of distances and directions, similar to that in MMVD. In the examples, there exists a GPM with MMVD (GPM-MMVD) involving nine candidate distances, such as (¼ image element, ½ image element, 1 image element, 2 image element, 3 image element, 4 image element, 6 image element, 8 image element, 16 image element) and eight candidate directions (e.g., four horizontal / vertical directions and four diagonal directions). Additionally, in the examples, when a sign (e.g., When ) equals 1, MVD is shifted left by 2 bits as in MMVD.

[0117] In some examples, a technique known as Geometric Partitioning Pattern (GPM) utilizing adaptive mixing is used. In some examples (e.g., VVC), the final predicted samples are generated by mixing the predictions of the two prediction signals using a weighted average. Two integer mixing matrices (W0 and W1) are used. In some examples, the weights in the GPM mixing matrix are derived based on the displacement from the predicted sample location to the GPM partition boundary using a ramp function. In some examples, the mixing region size is fixed at two (e.g., two samples on each side of the GPM partition split boundary).

[0118] In some examples, the blending process is improved by adding additional blending region sizes, such as adding four blending region sizes that are one-quarter, half, twice, and four times the existing region size.

[0119] Figure 11 A plot of the weights for GPM mixing is shown, based on a ramp function of the displacement (d) from the predicted sample location to the GPM partition boundary and the mixing region size (τ). Figure 11In the example, the first curve (1110) corresponds to the ramp function of the regular mixed region (also known as the existing region size), such as two samples on each side of the GPM partition split boundary; the second curve (1120) corresponds to the ramp function of a quarter-mixed region of the regular mixed region; the third curve (1130) corresponds to the ramp function of a half-mixed region of the regular mixed region; the fourth curve (1140) corresponds to the ramp function of a twice-mixed region of the regular mixed region; and the fifth curve (1150) corresponds to the ramp function of a four-mixed region of the regular mixed region.

[0120] In some examples, CU-level flags are encoded to signal the selected blending region size. Additionally, extended weighted precision can be utilized, where the maximum value of the weights is changed from 8 (in VVC) to 32 to accommodate the extended blending region size.

[0121] In some examples, to accommodate the increased width of the GPM blending region, the maximum value of the weights is changed from 8 to 32. In the examples, the weights are calculated according to formula (3):

[0122] Formula (3)

[0123] Allows selection of the width of the blending region from a set of predefined values ​​(e.g., In the example, the predefined value θ could be... .

[0124] In some examples, a technique called Geometric Partitioning Pattern (GPM) with Template Matching (I) is used, which is referred to as GPM-TM in the examples.

[0125] In some examples, to apply template matching to GPM, when GPM mode is enabled for CU, a signal is used to notify a CU-level flag indicating whether TM should be applied to both geometric partitions. TM can be used to refine the motion information for each geometric partition. When TM is selected, a template is constructed using left, top, or left and top neighbor samples based on the partition angle.

[0126] Table 1 shows templates for the first and second geometric partitions.

[0127] Table 1

[0128]

[0129] In Table 1, A indicates the use of the top sample, L indicates the use of the left sample, and L+A indicates the use of both the left and top samples.

[0130] In some examples, motion information is then refined by minimizing the difference between the current template and the template in the reference image using the same search pattern of the merging mode when the half-image element interpolation filter is disabled.

[0131] In some examples, a GPM candidate list can be constructed. For instance, in the first step, interleaved lists of MV candidates 0 and MV candidates 1 are derived directly from the regular merge candidate list, where list 0 MV candidates have a higher priority than list 1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates. In the second step, interleaved lists of MV candidates 1 and MV candidates 0 are further derived directly from the regular merge candidate list, where list 1 MV candidates have a higher priority than list 0 MV candidates. The same pruning method with an adaptive threshold is also applied to remove redundant MV candidates. In the third step, zero MV candidates are populated until the GPM candidate list is full.

[0132] In some examples, GPM-MMVD and GPM-TM are exclusively enabled for a single GPM CU. In these examples, signaling for the GPM-MMVD syntax is executed first. When both GPM-MMVD control flags are false (e.g., GPM-MMVD is disabled for both of the two GPM partitions), the GPM-TM flag is signaled to indicate whether template matching should be applied to both GPM partitions. Otherwise (e.g., at least one GPM-MMVD flag is true), the value of the GPM-TM flag is inferred as false.

[0133] In some examples, a technique known as GPM utilizing inter-frame prediction and intra-frame prediction is used. Under GPM utilizing inter-frame prediction and intra-frame prediction, the final prediction samples are generated by weighting the inter-frame prediction samples and intra-frame prediction samples from the regions separated by each GPM. The inter-frame prediction samples are derived from the inter-frame GPM, while the intra-frame prediction samples are derived from the Intra Prediction Mode (IPM) candidate list and from an index signaled by the encoder. The IPM candidate list size is predefined as 3.

[0134] Figures 12A to 12D Plots of GPM using inter-frame prediction and intra-frame prediction are shown in some examples. Figures 12A to 12C A graph showing the available IPM candidates is presented. Figure 12A A diagram showing the parallel angle pattern (parallel pattern) relative to the GPM block boundary (e.g., partition boundary); Figure 12B A diagram showing the vertical angle pattern (vertical pattern) relative to the GPM block boundary (e.g., partition boundary) is provided; and Figure 12C A diagram of the planar pattern is shown. Figure 12DA diagram is shown illustrating the use of intra-frame prediction and intra-frame prediction-based GPM. In some examples, the use of intra-frame prediction and intra-frame prediction-based GPM is limited to reduce the signaling overhead of IPM and avoid increasing the size of the intra-frame prediction circuitry on the hardware decoder. Additionally, direct motion vectors and IPM storage on the GPM mixing region are introduced to further improve coding performance.

[0135] In some examples, parallel modes are registered first in both Decoder Side Intra Mode Derivation (DIMD) and IPM derivation based on neighboring modes. Therefore, if no identical IPM candidates exist in the list, up to two IPM candidates derived from the Decoder Side Intra Mode Derivation (DIMD) method and / or neighboring blocks can be registered. For neighboring mode derivation, there are up to five available neighboring block locations, but these are limited by the GPM block boundary angles (as shown in Error! Reference source not found), which have already been used to utilize template-matched GPM (GPM-TM).

[0136] Table 2 shows the locations of available neighboring blocks for IPM candidate derivation based on the angle of the GPM block boundary.

[0137] Table 2

[0138]

[0139] In Table 2, A represents the top side of the prediction block, L represents the left side of the prediction block, and L+A represents the left and top sides of the prediction block.

[0140] In some examples, a template-matching-based reordering technique known as GPM splitting pattern can be used.

[0141] Typically, template matching (TM) is used to refine motion on the decoder side. In TM mode, motion is refined by constructing a template from neighboring samples on the left and / or top and finding the closest match between the template in the current image and the reference frame.

[0142] Template matching can be applied to GPM. When the CU is encoded in GPM, the decision to use TM for refinement can be made based on the motion of each geometric partition. When TM is selected, a template is constructed using left and / or top neighbor samples, and the motion is then refined by finding the best match between the current template and a reference region in the reference frame that has the same template pattern. The refined motion is used to perform motion compensation for the geometric partition and is stored in the motion field.

[0143] In some examples, to reduce the signaling cost of the GPM split pattern index, the index can be reordered using template matching costs. The reordering method for GPM split patterns is a two-step process performed after generating reference templates for each of the two GPM partitions in the coding unit. For example, the first step can use the weights of each split pattern to blend the reference templates of the two GPM partitions (e.g., generating 64 blended reference templates) and calculate the TM cost value for each of the blended reference templates; the second step can reorder the GPM split patterns in ascending order based on their TM cost values ​​and mark the best 32 GPM split patterns as available split patterns.

[0144] In some examples, the edges on the template can be obtained by extending from the (GPM split) edge of the current CU.

[0145] Figure 13 The diagram illustrates how, in some examples, the GPM split edges are expanded to obtain the edges on the template. Except that the weights are mapped to 0 and 8 (depending on which is closer) before use, the same GPM weight derivation process is used to compute the corresponding weights used in the template blending process.

[0146] In some examples, after reordering in ascending order using TM cost, Golomb-Rice code (with a divisor of 4) is used to signal the index to indicate the use of the GPM splitting pattern.

[0147] In some examples, a technique known as GPM with bidirectional predictive motion vectors is used. In some examples (e.g., ECM (Enhanced Compression Model), GPM relies on unidirectional predictive motion vectors to generate motion-compensated prediction samples for each inter-frame partition. In some examples, the use of bidirectional predictive motion vectors in GPM is permitted. Furthermore, GPM-MMVD and GPM-TM have been modified to incorporate the use of bidirectional predictive motion vectors. GPM with bidirectional predictive motion vectors may include specific elements, such as the following four elements in some examples.

[0148] The first element conditionally invokes the extraction process to extract unidirectional predicted motion vectors from the initial list. In the example, the extraction process is invoked only for small blocks such as 8×8, 16×8, and 8×16. For other larger blocks, the extraction process is bypassed in the example. The generation of the initial list is the same as before (i.e., normal merge list generation without any candidate reordering), except that when generating the initial list for larger blocks (i.e., blocks where the extraction process is bypassed), a motion vector difference threshold is added to control whether candidates can be added to the initial list, equal to a full sample distance.

[0149] The second element modifies GPM-MMVD to support bidirectional predictive motion vectors as base vectors. For low-latency images, the MVD, signaled by the IMPORTANT algorithm, is applied over the L0 and L1 motion vectors, as in existing merged MMVD designs. For non-low-latency images, the bidirectional predictive motion vectors are first converted to unidirectional predictive motion vectors, and then the MVD is applied over them.

[0150] The third element modifies GPM-TM to also support bidirectional prediction motion vectors. When the image is not a low-latency image and the cost of the best template using bidirectional prediction exceeds 75% of the cost of the best template using unidirectional prediction, then the refined unidirectional prediction motion vector is determined as the final refined motion vector. Otherwise, the refined bidirectional prediction motion vector is determined as the final refined motion vector.

[0151] The fourth element enables 8×8 BDOF on top of the associated bidirectional predictive motion vector for each inter-frame partition (i.e., as in the existing design of multi-pass DMVR).

[0152] In inter-frame prediction, merging modes can be used to improve coding efficiency. In merging mode, motion vectors can be derived from neighboring blocks and directly used for motion compensation. To improve the accuracy of motion vectors (MVs) in merging mode, decoder-side motion vector refinement (DMVR) based on bilateral matching (BM) can be applied, for example, in VVC. In bidirectional prediction operations, refined MVs can be searched around the initial MVs in reference image lists L0 and L1. BM calculates the distortion between two candidate blocks in reference image lists L0 and L1.

[0153] Figure 14 Exemplary schematic diagrams of BM-based decoder-side motion vector refinement are shown in some examples. For instance... Figure 14As shown, the current image (1402) may include the current block (1408). The current image may have a first reference image (1404) from (reference image) list L0 and a second reference image (1406) from (reference image) list L1. For the current block (1408), a pair of reference blocks are identified in the first and second reference images based on the initial motion vectors MV0 and MV1. For example, the initial reference block (1412) in the first reference image (1404) can be located based on the initial motion vector MV0, and the initial reference block (1414) in the second image (1406) can be located based on the initial motion vector MV1. Search processing can be performed around the initial MV0 in the first reference image (1404) and the initial MV1 in the second reference image (1406). For example, the MV0 can be adjusted. diff The initial MV0 and MV1 are applied in the opposite direction to obtain MV candidates, such as MV0' and MV1'. Based on the MV candidates, a pair of candidate reference blocks are identified in the first reference image and the second reference image. For example, a candidate reference block (1410) can be identified in the first reference image (1404) based on MV0', and a candidate reference block (1416) can be identified in the second reference image (1406) based on MV1'. In some examples, bilateral matching (BM) refers to the operation of calculating a distortion metric between a pair of reference blocks in the corresponding reference images of the current image, such as calculating the sum of absolute differences (SAD) between a pair of reference blocks as a distortion metric for that pair of reference blocks. For example, the BM method calculates an initial SAD between a pair of initial reference blocks (1412) and (1414), and a second SAD between a pair of candidate reference blocks (1410) and (1416). The initial SAD is associated with the initial MVs (e.g., MV0 and MV1), and the second SAD is associated with the MV candidates (e.g., MV0' and MV1'). Similarly, the BM method can compute the SAD of multiple MV candidates around the initial MV. The MV candidate with the lowest SAD can become the refined MV and is used to generate a bidirectional prediction signal to predict the current block (1408).

[0154] In some examples (e.g., VVC), the application of DMVR is restricted and applied only to CUs encoded using patterns and features that meet specific conditions. The DMVR algorithm is invoked if a block meets specific conditions. For example, conditions (also referred to as DMVR requirements or a set of conditions for DMVR) may include: (1) a CU-level merging pattern with bidirectional prediction MV; (2) one reference image is past and the other is future relative to the current image; (3) the distances (e.g., POC differences) of the two reference images to the current image are the same; (4) the two reference images are short-term reference images; (5) the CU has more than 64 luminance samples; (6) both the CU height and CU width are greater than or equal to 8 luminance samples; (7) the weight index of the bidirectional prediction with CU-Level Weights (BCW) with CU-level weights indicates equal weights; (8) weighted prediction (WP) is not enabled for the current block; and (9) the combined inter-frame and intra-frame prediction (CIIP) pattern is not used for the current block.

[0155] Note that the refined MV derived from DMVR processing is used to generate inter-frame prediction samples and can be used for temporal motion vector prediction for future image encoding. In some examples, the original MV is used for deblocking and also for spatial motion vector prediction for future CU encoding.

[0156] In DVMR, the search point revolves around the initial MV, and the MV offset follows the MV difference mirror rule. This is checked by DMVR and by candidate MV pairs. Any point indicated complies with and .in, Indicates the initial MV in one of the reference images (e.g., The refinement offset between the initial MV and the refined MV. In some examples, the refinement search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase.

[0157] In some examples (e.g., VVC), decoder-side motion vector refinement (DMVR) is applied to the CU encoded in a regular merging mode. The MV pairs obtained from the regular merging candidates are used as input to the DMVR process. DMVR applies bilateral matching (BM) to refine the input MV pairs {MV0, MV1} and processes the refined MV pairs {MV0, MV1}. 经细化的L0 MV 经细化的L1} Used for motion compensation prediction of both luminance and chrominance components, such as Figure 4As shown. The output MV of DMVR can be referred to as the refined MV pair, and can be expressed by formula (4):

[0158] Formula (4)

[0159] The motion vector difference Δmv is applied to the input MV pair to obtain a refined MV pair by using the MVD mirror property, since the input MV pair points to two different reference images that have an equal difference from the current image in picture order count (POC) and that the two reference images are in different time directions.

[0160] In some examples, DMVR can be applied at the sub-block level, where the luminance-coded block is divided into 16×16 sub-blocks for MV refinement. Δmv is derived independently for each sub-block.

[0161] In some examples, the motion vector refinement search range is two integer luminance samples starting from the initial MV. The motion vector search can be performed in two steps, for example, the first step is an integer sample offset search stage (also known as integer precision motion search) and the second step is a fractional sample refinement stage (also known as fractional motion search step or fractional sample offset search).

[0162] In some examples, a 25-point full search can be applied to the integer sample offset search, as shown in Equation (5):

[0163] Formula (5)

[0164] in, Let i and j represent the coordinates of the search points surrounding the initial MV pair, where i and j are integer values ​​between -2 and 2 (inclusive). First, calculate the SAD of the initial MV pair, for example, according to formula (6):

[0165] Formula (6)

[0166] as well as

[0167] Where W and H are the weight and height of the sub-block.

[0168] If the SAD of the initial MV pair is less than a threshold, the integer sample phase of the DMVR terminates. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search phase. To reduce the adverse effects of the uncertainty in DMVR refinement, the original MVs (e.g., initial MV candidates MV0 and MV1) can be preferred during DMVR processing. The SAD between reference blocks referenced by the initial MV candidates is reduced, for example, by 1 / 4 of the SAD value, to make the initial MV candidates preferred.

[0169] In some examples, fractional sample refinement follows the integer sample search. In some examples, fractional sample refinement is performed using fractional sample offsets, such as offsetting half of the image elements in the vertical and horizontal directions. In some examples, to save computational complexity, fractional sample refinement is derived using a parametric error surface equation (also known as a quadratic prediction-based method) instead of an additional search using SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. For example, fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration at a specific integer position (also known as the center) with the minimum SAD.

[0170] In subpixel offset estimation based on the parametric error surface, the cost at the center location (the center location is the point with the minimum SAD in the integer sample offset search) and the costs at four neighboring locations from the center (e.g., from the center locations (-1, 0), (0, -1), (1, 0), (0, 1)) are used to fit the 2-D parabolic error surface formula, such as formula (7):

[0171] Formula (7)

[0172] in, The fractional position corresponds to the position with the minimum cost, and C corresponds to the minimum cost value. The above formula is solved by using the cost values ​​of the five search points, calculated according to formulas (8) and (9). :

[0173] Formula (8)

[0174] Formula (9)

[0175] x min and y min The value can be automatically constrained between -8 and 8 because all cost values ​​are positive and the minimum is E(0, 0). The constraint corresponds to the half-image element offset of the 1 / 16-th-image element MV accuracy in the VVC. The calculated score... The integer distance thinning MV is added to obtain the subpixel-accurate thinning increment MV. In formulas (3) and (4), , , , and This represents the cost value at five points (the center location and four neighboring locations).

[0176] A technique known as bidirectional optical flow (BDOF) can be used in examples such as VVC. BDOF was previously known as BIO in JEM (Joint Exploration Model). The BDOF in VVC can be a simpler version compared to the JEM version, requiring less computation, particularly in terms of the number of multiplications and the size of the multiplier.

[0177] BDOF can be used to refine the bidirectional prediction signal of the CU at the 4×4 sub-block level. BDOF can be applied to the CU if the following conditions are met (also known as the requirements of BDOF or a set of conditions for BDOF): (1) The CU is encoded using a “true” bidirectional prediction mode, i.e., one of the two reference images is displayed before the current image and the other is displayed after the current image; (2) The distances (e.g., POC differences) from the two reference images to the current image are the same; (3) The two reference images are short-term reference images; (4) The CU is not encoded using an affine mode or an SbTMVP merging mode; (5) The CU has more than 64 luminance samples; (6) Both the CU height and CU width are greater than or equal to 8 luminance samples; (7) The BCW weight index indicates equal weights; (8) Weighted prediction (WP) is not enabled for the current CU; and (9) CIIP mode is not used for the current CU.

[0178] In some examples, BDOF is applied only to the luminance component. As the name BDOF suggests, the BDOF mode can be based on an optical flow concept that assumes smooth motion of the object. For each 4×4 sub-block, motion refinement (v) can be computed by minimizing the difference between the L0 and L1 prediction samples. x , v y Then, motion refinement can be used to adjust the bidirectional prediction sample values ​​in the 4×4 sub-blocks. BDOF may include the following steps.

[0179] First, the horizontal and vertical gradients of the two predicted signals from reference lists L0 and L1 can be calculated by directly calculating the difference between two neighboring samples. and , The horizontal and vertical gradients can be provided in the following formulas (10) and (11):

[0180] Formula (10)

[0181] Formula (11)

[0182] in, It can be a list , Coordinates of the predicted signal in The sample value at the position, and the shift of 1 can be calculated based on the brightness bit depth bitDepth, such as shift 1 = max (6, bitDepth-6).

[0183] Then, the gradient can be calculated according to the following formulas (12) to (16). , , , and Autocorrelation and cross-correlation:

[0184] Formula (12)

[0185] Formula (13)

[0186] Formula (14)

[0187] Formula (15)

[0188] Formula (16)

[0189] in, , and It can be provided in formulas (17) to (19) respectively.

[0190] Formula (17)

[0191] Formula (18)

[0192] Formula (19)

[0193] Where Ω can be a 6×6 window surrounding a 4×4 sub-block, and n a The value of n bThe values ​​can be set to min(1, -itDepth-11) and min(4, -itDepth-8) respectively.

[0194] Then, using cross-correlation and autocorrelation terms, the motion refinement (v) can be derived using the following formulas (20) and (21). x , v y ):

[0195] Formula (20)

[0196] Formula (21)

[0197] in, , , . It is the floor function, and Based on motion refinement and gradients, adjustments can be calculated for each sample in the 4×4 sub-block using formula (22):

[0198] Formula (22)

[0199] Finally, the BDOF samples of CU can be calculated by adjusting the bidirectional predicted samples in the following formula (23):

[0200] Formula (23)

[0201] Values ​​can be selected such that the multiplier in BDOF processing does not exceed 15 bits, and the maximum bit width of intermediate parameters in BDOF processing can be kept within 32 bits.

[0202] Figure 15 A graph illustrating some calculations in the bidirectional optical flow (BDOF) example is shown. To derive the gradient values, a list needs to be generated outside the current CU boundary. ( Some predicted samples in ) .like Figure 15 As shown, BDOF in VVC can use an extended row / column (1502) around the boundary (1506) of CU (1504). To control the computational complexity of generating out-of-bounds predicted samples, the extended region can be generated by directly taking reference samples at nearby integer positions (e.g., using floor() operations on coordinates) without interpolation (e.g., Figure 15 The predicted samples are located in the non-shaded areas of the CU, and a normal 8-tap motion-compensated interpolation filter can be used to generate the predicted samples within the CU (e.g., Figure 15 (The shaded area in the diagram). Expanded sample values ​​can be used solely for gradient computation. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, the samples and gradient values ​​can be filled (e.g., repeated) from their nearest neighbors.

[0203] In some examples, sample-based BDOF can be used instead of block-based BDOF. In sample-based BDOF, motion refinement (v) is derived on a block-by-block basis instead of on a block-by-block basis. x , v y This is performed on a sample-by-sample basis. The encoded block is divided into 8×8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, for each sample in the sub-block, a sliding 5×5 window is used, and the existing BDOF processing is applied to each sliding window to derive v. x and v y The derived motion refinement (v) x , v y This is applied to adjust the bidirectional predicted sample values ​​of the center sample of the window.

[0204] In some examples, multiple passes of DMVR can be used. In one example, in the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16×16 sub-block within the coded block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MV is stored for both spatial motion vector prediction and temporal motion vector prediction.

[0205] Specifically, the first pass performs block-based bilateral matching MV refinement. In the first pass, the refined MV is derived by applying the BM to the coded block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, a refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MV is derived around the initial MV based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. and The cost of bilateral matching can be calculated using any suitable error measurement metric that measures the error between the two reference blocks in L0 and L1. In the example, bilateral matching includes the sum of absolute differences (SAD) between corresponding samples in the two reference blocks in L0 and L1.

[0206] BM can perform a local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern that iterates through the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0207] The cost of bilateral matching is calculated as follows: When the block size When the value is greater than 64, the Mean Removed SAD (MRSAD) cost function is applied to remove the DC (Direct Current) effect that distorts the reference blocks. The local search for intDeltaMV terminates when the bilCost at the center point of the 3×3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.

[0208] The existing fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as follows:

[0209] Formula (24)

[0210] Formula (25)

[0211] In the second pass, sub-block-based bilateral matching MV refinement is performed. Specifically, in the second pass, the refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the two MVs obtained in the first pass are shown in reference image lists L0 and L1. and The refined MV is searched around the reference sub-blocks in L0 and L1. The refined MV is derived based on the minimum bilateral matching cost between two reference sub-blocks in L0 and L1. and ).

[0212] For each sub-block, BM performs a full search to derive integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0213] The bilateral matching cost is calculated by applying a cost factor to the sum of the absolute transformation difference (SATD) costs between the two reference subblocks: In some examples, the search area It is divided into up to 5 diamond-shaped search areas.

[0214] Figure 16 The search region (1600) is shown in some examples. The search region (1600) is divided into 5 search regions (1601) to (1605). The shape of the search regions is similar to a rhombus.

[0215] In some examples, each search region is assigned a costFactor, determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond-shaped region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order from the top left corner to the bottom right corner. The minimum bilCost within the current search region is less than or equal to... If the threshold is reached, the full search of integer image elements terminates; otherwise, the full search of integer image elements continues to the next search region until all search points have been checked. Additionally, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than or equal to a threshold for the block area, the search process terminates.

[0216] In some examples, fractional sample refinement, such as DMVR fractional sample refinement in VVC, is further applied to derive the final fractional sample refinement. Then, the refined MV from the second pass was exported as:

[0217] Formula (26)

[0218] Formula (27)

[0219] In the third pass, sub-block-based bidirectional optical flow (MV) refinement can be performed. Specifically, in the third pass, the refined MV is derived by applying BDOF to 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy, without clipping starting from the refined MV of the parent-child blocks in the second pass. The derived... Rounded to 1 / 16 of the sample precision and clipped between -32 and 32. The refined MV at the third pass (…). and ) was exported as:

[0220] Formula (28)

[0221] Formula (29)

[0222] In some examples, a technique known as high-precision MV refinement for BDOF is used. In some examples, BDOF sample adjustments can derive motion refinement for 4×4 sub-blocks. And adjust the samples individually. In some examples (e.g., ECM), two types of BDOF can be used: one as a BDOF for MV refinement (the third stage of DMVR processing), and the other as a BDOF for sample adjustment (similar to VVC, but deriving motion adjustments for each sample separately). ).

[0223] In some examples, high-precision formulas are used to derive the BDOF MV refinement parameters, such as formulas (30) and (31):

[0224] Formula (30)

[0225] Formula (31)

[0226] Here, Gx / Gy is the sum of the two horizontal / vertical gradients derived for each reference block. The sum (Σ) is a weighted sum, where the weights depend on the location within the target region Ω. In other cases, weights can also be applied to derive vx / vy.

[0227] In addition, the sub-block size of BDOF DMVR is adaptively selected based on width × height. For blocks smaller than 256, a 4×4 sub-block size is used, and in other cases, an 8×8 sub-block size is used.

[0228] Note that in some relevant examples, an 8×8 BDOF is applied to a GPM with bidirectional predicted motion vectors. However, applying BDOF along the boundaries of GPM partitions can lead to visual artifacts.

[0229] Some aspects of this disclosure provide techniques for applying variable BDOF sub-block sizes and adaptive switching of BDOF on / off to GPM partitions with bidirectional predictive motion vectors. For example, an encoder / decoder can determine a geometric partitioning pattern (GPM) for a current block in a current image, wherein at least a first GPM partition has bidirectional predictive motion vectors. The encoder / decoder can apply sub-block-based motion refinement with bidirectional motion to at least a first and a second sub-block of the first GPM partition. The first and second sub-blocks have different sub-block sizes. The current block can be reconstructed based on the sub-block-based motion refinement.

[0230] In some implementations, the coded blocks of each GPM partition are split into Sub-blocks, where N is a positive number and greater than or equal to the maximum supported BDOF sub-block size (e.g., the width or height of a square-shaped sub-block) of the encoded block (e.g., also referred to as the encoded block, the current block). In the example, based on The value of the blending mask in each sub-block Determine the size of BDOF sub-blocks at the sub-block level.

[0231] In some examples, the supported BDOF sub-block size can vary depending on the size of the encoded block. In these examples, the maximum supported (also known as the maximum supported) BDOF sub-block size can vary based on the size of the encoded block.

[0232] In some examples, the BDOF sub-block size is determined by... The corresponding mask value (e.g., the value of the mixing mask) in the sub-block is determined. Therefore, in some examples, the size of the BDOF sub-block can be variable within the coded block, and the sub-block size can depend on which sub-partition the sub-block belongs to and / or whether it is located at a partition boundary.

[0233] In some examples, when When all mask values ​​(of the mixed mask) within a sub-block are the maximum or minimum weight values, the BDOF sub-block size is equal to the maximum supported BDOF sub-block size.

[0234] In some examples, when If all mask values ​​within a sub-block are not the maximum weight value and N / 2 is greater than or equal to the minimum supported BDOF sub-block size, the BDOF sub-block size can be further divided into smaller sub-block sizes, for example... .

[0235] In some examples, when If all mask values ​​within a sub-block are not the minimum weight value and N / 2 is greater than or equal to the minimum BDOF sub-block size, the BDOF sub-block size can be further divided into smaller sub-block sizes, for example... .

[0236] In some examples, when If all mask values ​​within a sub-block are not the maximum or minimum weight values ​​and N / 2 is greater than or equal to the minimum supported BDOF sub-block size, the BDOF sub-block size can be further divided into smaller sub-block sizes, for example... .

[0237] In some examples, when the size of the split sub-blocks is larger than the minimum BDOF sub-block size, the BDOF sub-block size can be recursively split into smaller sub-blocks.

[0238] In some examples, The corresponding mask value in the sub-block is used to determine whether it is in the sub-block. BDOF is applied at the sub-block. And n can be N, (N / 2), (N / 4), etc.

[0239] In the example, when When all mask values ​​within a sub-block are the maximum (or minimum) weight value, such as 32 (e.g., in ECM), BDOF is applied to that sub-block. sub-block.

[0240] In the example, when When all mask values ​​within a sub-block are non-zero, apply BDOF to that sub-block. sub-block.

[0241] In the example, when BDOF is applied to the sub-block when all mask values ​​within it are greater than and / or equal to a predefined threshold T. Sub-block. The predefined threshold T can be a fixed predefined value, or it can be signaled using high-level syntax, including but not limited to SPS (Sequence Parameter Set), PPS (Picture Parameter Set), PH (Picture Header), slice header, etc.

[0242] In the example, when When all mask values ​​within a sub-block are greater than or equal to a given threshold, BDOF is applied to that sub-block. Sub-block. Example mask values ​​include the maximum weight value, or the maximum weight value minus a predefined offset value.

[0243] In the example, when When all mask values ​​within a sub-block are less than or equal to a given threshold, BDOF is applied to that sub-block. Sub-block. Example mask values ​​include 0 or a predefined positive offset value.

[0244] Some aspects of this disclosure provide techniques for determining DMVR subblock sizes and whether to apply DMVR to GPM partitions with bidirectional predictive motion vectors.

[0245] In some examples, the encoded blocks of each GPM partition are split into Sub-blocks, and each The corresponding mask value in the sub-block is used to determine whether it is in the sub-block. Apply DMVR at the sub-block.

[0246] In the example, the supported DMVR sub-block size can vary depending on the size of the encoded block.

[0247] In some examples, The corresponding mask value of the mixing mask in the sub-block is used to determine whether it is in the sub-block. Apply DMVR at the sub-block.

[0248] In the example, when When all mask values ​​within a sub-block are of maximum weight, such as 32 (e.g., in ECM), DMVR is applied to that sub-block. Sub-block. Otherwise, do not apply DMVR to this. sub-block.

[0249] In some examples, when it is determined that... When applying DMVR to a sub-block, multiple passes of DMVR can be applied, including but not limited to the first pass of DMVR, the second pass of DMVR, and BDOF MV refinement.

[0250] In some examples, it can be determined whether to apply multiple-pass DMVR on a GPM partition with bidirectional predictive MV based on the GPM split pattern index. In the examples, multiple-pass DMVR is applied when the GPM angle is horizontal or vertical.

[0251] In some examples, the first step is to check whether the sub-block contains a GPM partition boundary (e.g., a GPM partition boundary intersects with the sub-block). If the sub-block does not contain a GPM partition boundary, then sub-block-based DMVR and / or BDOF are applied to the partition with bidirectional prediction MV; otherwise, when the sub-block contains a GPM partition boundary, both sub-block-based DMVR and BDOF are disabled to avoid visual artifacts.

[0252] Figure 17 A flowchart outlining a process (1700) according to an embodiment of this disclosure is shown. Process (1700) can be used in a video decoder. In various embodiments, process (1700) is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video decoder (110), a processing circuitry system that performs the functions of a video decoder (210), etc. In some embodiments, process (1700) is implemented as software instructions, so that when the processing circuitry system executes the software instructions, the processing circuitry system executes process (1700). Processing begins at (S1701) and proceeds to (S1710).

[0253] At (S1710), an encoded video bitstream including encoded information of one or more images is received.

[0254] At (S1720), it is determined from the encoded information that the current block in the current image is in the geometric partitioning mode (GPM), wherein at least the first GPM partition has bidirectional predicted motion vectors.

[0255] At (S1730), the sub-block-based motion refinement with bidirectional motion is applied to the first and second sub-blocks of at least the first GPM partition, the first and second sub-blocks having different sub-block sizes.

[0256] At (S1740), the current block is reconstructed based on sub-block-based motion refinement.

[0257] In some examples, sub-block-based motion refinement is bidirectional optical flow (BDOF) motion refinement, where the first sub-block is a first BDOF sub-block and the second sub-block is a second BDOF sub-block.

[0258] In some examples, the first GPM partition is divided into larger N×N sub-blocks, where N is a positive number and the N×N size is greater than or equal to the maximum supported BDOF sub-block size. The corresponding BDOF sub-block size for each larger N×N sub-block is determined. The larger sub-block is then divided into BDOF sub-blocks based on the corresponding BDOF sub-block size. BDOF motion refinement is applied to the BDOF sub-blocks.

[0259] In some examples, the supported BDOF sub-block size is determined based on the size of the current block.

[0260] In some examples, the size of the first BDOF sub-block is determined based on the value of the mixing mask in the first larger sub-block of size N×N. The first larger sub-block is then divided into first BDOF sub-blocks according to the first BDOF sub-block size. BDOF motion thinning is then applied to the first BDOF sub-blocks.

[0261] In one example, when all mask values ​​in the first larger sub-block correspond to either the maximum or minimum weight value, the size of the first BDOF sub-block is set to the maximum supported BDOF sub-block size. In another example, when no mask value in the first larger sub-block corresponds to the maximum weight value, the size of the first BDOF sub-block is set to be smaller than the maximum supported BDOF sub-block size and greater than or equal to the minimum supported BDOF sub-block size. (This is repeated three times in the original text.)

[0262] In some examples, the application of BDOF refinement to a portion of the first larger sub-block is determined based on the mask values ​​in at least a portion of that sub-block. In one example, BDOF refinement is applied to that portion of the first larger sub-block when all mask values ​​in that portion correspond to either the maximum or minimum weight value. In another example, BDOF refinement is applied to that portion of the first larger sub-block when no mask value in that portion is zero. In yet another example, BDOF refinement is applied to that portion of the first larger sub-block when all mask values ​​in that portion are above a threshold. In yet another example, BDOF refinement is applied to that portion of the first larger sub-block when all mask values ​​in that portion are below a threshold.

[0263] In some examples, sub-block-based motion refinement is decoder-side motion vector refinement (DMVR) refinement, where the first sub-block is the first DMVR sub-block and the second sub-block is the second DMVR sub-block.

[0264] In the example, the first GPM partition is divided into multiple sub-blocks, and whether to apply DMVR refinement to a specific sub-block is determined based on the mask value of the blending mask in that specific sub-block.

[0265] In some examples, the supported DMVR sub-block size is determined based on the size of the current block.

[0266] In the example, DMVR refinement is applied to a specific sub-block when all mask values ​​in that sub-block correspond to either the maximum or minimum weight value.

[0267] In the example, when it is determined that DMVR refinement is to be applied to a specific sub-block, DMVR is applied multiple times to that specific sub-block.

[0268] In the example, the application of DMVR is determined based on the GPM split pattern index.

[0269] In the example, DMVR is applied multiple times when the GPM angle is either horizontal or vertical.

[0270] In the example, it checks whether the GPM partition boundary intersects with a specific sub-block. When the GPM partition boundary does not intersect with a specific sub-block, sub-block-based motion refinement with bidirectional motion is applied to the specific sub-block. When the GPM partition boundary intersects with a specific sub-block, sub-block-based motion refinement is disabled for the specific sub-block.

[0271] Then, the process proceeds to (S1799) and terminates.

[0272] Process (1700) can be adjusted as appropriate. Steps in process (1700) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0273] Figure 18 A flowchart outlining the process (1800) according to an embodiment of this disclosure is shown. The process (1800) can be used in a video encoder. In various embodiments, the process (1800) is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video encoder (103), a processing circuitry system that performs the functions of a video encoder (303), etc. In some embodiments, the process (1800) is implemented as software instructions, so that the processing circuitry system executes the process (1800) when the software instructions are executed. The process begins at (S1801) and proceeds to (S1810).

[0274] At (S1810), it is determined that GPM mode will be used for the current block in the current image.

[0275] At (S1820), the first GPM partition with bidirectional predicted motion vectors is determined.

[0276] At (S1830), the first GPM partition is divided into larger sub-blocks of size N×N, where N is a positive number and the size of N×N is greater than or equal to the maximum sub-block size supported by motion refinement based on sub-blocks.

[0277] At (S1840), the size of each of the larger N×N sub-blocks is determined based on the mixing mask of the current block.

[0278] At (S1850), the larger sub-block is divided into sub-blocks based on their respective sub-block sizes.

[0279] At (S1860), it is determined whether to apply sub-block-based motion refinement to a specific sub-block based on the value of the blending mask in that specific sub-block.

[0280] In some examples, sub-block-based motion refinement is bidirectional optical flow (BDOF) motion refinement.

[0281] In some examples, sub-block-based motion refinement is decoder-side motion vector refinement (DMVR) refinement.

[0282] Then, the process proceeds to (S1899) and terminates.

[0283] Process (1800) can be adjusted as appropriate. Steps in process (1800) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.

[0284] Some aspects of this disclosure provide a method for processing visual media data. The method includes processing a bitstream of visual media data according to format rules. The bitstream includes encoded information of one or more images, the encoded information indicating that a current block in the current image is encoded in a geometric partitioning (GPM) mode, wherein at least a first GPM partition has bidirectional predictive motion vectors. The format rules specify: splitting the current block into GPM partitions according to the GPM mode; the first GPM partition of the GPM partition having bidirectional predictive motion vectors; and dividing the first GPM partition into larger sub-blocks of size N×N. N is a positive number and the N×N size is greater than or equal to the maximum sub-block size supported by sub-block-based motion refinement. The format rules also specify: determining the sub-block size of each of the larger N×N sub-blocks based on a mixing mask of the current block; dividing the larger sub-blocks into sub-blocks based on their respective sub-block sizes; and applying sub-block-based motion refinement to a specific sub-block based on the value of the mixing mask in that specific sub-block. Sub-block-based motion refinement is one of bidirectional optical flow (BDOF) motion refinement and decoder-side motion vector refinement (DMVR) refinement.

[0285] The techniques described above can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 19 A computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0286] Computer software can be coded using any suitable machine code or computer language. The machine code or computer language can be subjected to mechanisms such as assembly, compilation, and linking to create code that includes instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.

[0287] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0288] Figure 19 The components shown for the computer system (1900) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. Nor should the configuration of the components be construed as having any dependency or requirement relating to any one or a combination of the components shown in the exemplary embodiments of the computer system (1900).

[0289] Computer system 1900 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input made by one or more human users through, for example, tactile input (e.g., keystrokes, swiping, data glove movement), audio input (e.g., voice, tapping), visual input (e.g., gestures), and olfactory input (not depicted). Human-machine interface devices can also be used to capture certain media that are not necessarily directly related to conscious input made by humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images acquired from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0290] Input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1901), mouse (1902), touchpad (1903), touch screen (1910), data glove (not shown), joystick (1905), microphone (1906), scanner (1907), and camera device (1908).

[0291] The computer system (1900) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: haptic output devices (e.g., haptic feedback via a touchscreen (1910), data gloves (not shown), or joystick (1905), but haptic feedback devices that are not used as input devices may also exist); audio output devices (e.g., speakers (1909), headphones (not depicted)); visual output devices (e.g., screens (1910), including CRT (Cathode Ray Tube) screens, LCD (Liquid Crystal Display) screens, plasma screens, OLED (Organic Light-Emitting Diode) screens, each screen may or may not have touchscreen input capability, each screen may or may not have haptic feedback capability—some of the screens may be able to output two-dimensional visual output or more than three-dimensional output in a manner such as stereoscopic graphics output; virtual reality glasses (not depicted); holographic displays and smoke generators (not depicted)); and printers (not depicted).

[0292] Computer systems (1900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory, ROM) / RW (1920) with media such as CD / DVD (1921), thumb drives (1922), removable hard disk drives or solid-state drives (1923), conventional magnetic media such as magnetic tape and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit, ASIC) / PLD (Programable Logic Device, PLD) such as security dongles (not depicted), etc.

[0293] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transient signals.

[0294] Computer systems (1900) may also include interfaces (1954) to one or more communication networks (1955). Networks can be, for example, wireless, wired, or optical. Networks can also be local area, wide area, metropolitan area, vehicle-mounted and industrial, real-time, latency-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LANs (Local Area Networks); cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and industrial networks including CANBus (Controller Area Network-BUS), etc. Some networks typically require external network interface adapters that attach to certain general-purpose data ports or peripheral buses (1949) (such as, for example, the USB (Universal Serial Bus, USB) port of a computer system (1900); other networks are typically integrated into the core of the computer system (1900) by attaching to system buses as described below (e.g., to an Ethernet interface in a PC (Personal Computer, PC) system or a cellular network interface in a smartphone computer system). The computer system (1900) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., to a CAN bus of a certain CAN bus device), or bidirectional, for example, to other computer systems using local area digital networks or wide area digital networks. As described above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.

[0295] The aforementioned human-computer interface devices, human-accessible storage devices, and network interfaces could be attached to the core (1940) of a computer system (1900).

[0296] The core (1940) may include one or more central processing units (CPU) (1941), graphics processing units (GPUs) (1942), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1943), hardware accelerators for certain tasks (1944), graphics adapters (1950), etc. These devices, along with read-only memory (ROM) (1945), random access memory (1946), and internal mass storage devices (1947) such as internal non-user-accessible hard disk drives, SSDs (Solid-State Drives), etc., can be connected via the system bus (1948). In some computer systems, the system bus (1948) may be accessed as one or more physical plugs to allow for expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1948) or may be attached to the core's system bus (1948) via a peripheral bus (1949). In the example, the screen (1910) can be connected to the graphics adapter (1950). Peripheral bus architectures include PCI (Peripheral Component Interconnect), USB, etc.

[0297] The CPU (1941), GPU (1942), FPGA (1943), and accelerator (1944) can execute certain instructions, which, when combined, can form the aforementioned computer code. This computer code can be stored in ROM (1945) or RAM (1946). Transient data can also be stored in RAM (1946), while permanent data can be stored, for example, in an internal mass storage device (1947). Fast storage and retrieval of any memory device in the memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage devices (1947), ROMs (1945), RAMs (1946), etc.

[0298] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type known and available to those skilled in the art of computer software.

[0299] By way of example and not limitation, a computer system with an architecture (1900), and in particular a core (1940), can be functionalized by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage devices as described above, and with certain storage devices of the core (1940) having non-transitory properties, such as a mass storage device (1947) or ROM (1945) within the core. Software implementing various embodiments of this disclosure can be stored in such devices and executed by the core (1940). Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software can cause the core (1940), and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (RandomAccess Memory, RAM) (1946) and modifying such data structures according to the processes defined by the software. Alternatively or as an alternative, a computer system may be functional as a result of hard-wired logic or other means embodied in circuitry (e.g., an accelerator (1944)), which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, a reference to software may include logic, and a reference to logic may also include software. Where appropriate, a reference to a computer-readable medium may include circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry implementing logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0300] The use of “at least one of…” or “one of…” in this disclosure is intended to include any one or a combination thereof. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include either A or B or (A and B). The use of “one of…” does not exclude any combination of the elements described where applicable, such as when the elements are not mutually exclusive.

[0301] While this disclosure has described several exemplary embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. Therefore, it will be appreciated that those skilled in the art will be able to conceive of many systems and methods that, while not explicitly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.

Claims

1. A method for processing visual media data, the method comprising: The bitstream of visual media data is processed according to format rules, where: The bitstream includes encoded information for one or more images, the encoded information indicating that the current block in the current image is encoded in a geometric partitioning (GPM) mode, wherein at least the first GPM partition has bidirectional predicted motion vectors; and The formatting rules specify: Split the current block into GPM partitions according to the GPM mode; The first GPM partition in the GPM partition has bidirectional predicted motion vectors; The first GPM partition is divided into larger sub-blocks of size N×N, where N is a positive number and the size of N×N is greater than or equal to the maximum sub-block size supported by motion refinement based on sub-blocks; The size of each of the larger N×N sub-blocks is determined based on the mixing mask of the current block; The larger sub-block is divided into sub-blocks based on their respective sub-block sizes; and The sub-block-based motion refinement is applied to the specific sub-block based on the value of the mixing mask in the specific sub-block, wherein the sub-block-based motion refinement is one of bidirectional optical flow (BDOF) motion refinement and decoder-side motion vector refinement (DMVR) refinement.

2. A video decoding apparatus, comprising a processing circuit system, the processing circuit system being configured to: Receive encoded video bitstreams containing encoded information of one or more images; Based on the encoded information, it is determined that the current block in the current image is in the geometric partitioning mode (GPM), wherein at least the first GPM partition has bidirectional predicted motion vectors; Sub-block-based motion refinement with bidirectional motion is applied to at least the first and second sub-blocks of the first GPM partition, wherein the first and second sub-blocks have different sub-block sizes; and The current block is reconstructed based on the sub-block-based motion refinement.

3. The apparatus according to claim 2, wherein, The sub-block-based motion refinement is bidirectional optical flow (BDOF) motion refinement, where the first sub-block is a first BDOF sub-block and the second sub-block is a second BDOF sub-block.

4. The apparatus according to claim 3, wherein, The processing circuit system is configured to: The first GPM partition is divided into larger sub-blocks of size N×N, where N is a positive number and the size of N×N is greater than or equal to the maximum supported BDOF sub-block size; Determine the corresponding BDOF sub-block size for the larger of the N×N sub-blocks; The larger sub-block is divided into BDOF sub-blocks based on the corresponding BDOF sub-block size; and The BDOF motion refinement is applied to the BDOF sub-block.

5. The apparatus according to any one of claims 2 to 4, wherein, The processing circuit system is configured to: The supported BDOF sub-block size is determined based on the size of the current block.

6. The apparatus according to claim 4, wherein, The processing circuit system is configured to: The size of the first BDOF sub-block is determined based on the value of the mixing mask in the first larger sub-block of size N×N; The first larger sub-block is divided into first BDOF sub-blocks according to the size of the first BDOF sub-block; and The BDOF motion refinement is applied to the first BDOF sub-block.

7. The apparatus according to claim 6, wherein, The processing circuitry is configured to perform at least one of the following operations: When all mask values ​​in the first larger sub-block correspond to the maximum weight value or the minimum weight value, the size of the first BDOF sub-block is set to the maximum supported BDOF sub-block size. When no mask value in the first larger sub-block corresponds to the maximum weight value, the size of the first BDOF sub-block is set to be smaller than the maximum supported BDOF sub-block size and greater than or equal to the minimum supported BDOF sub-block size. When no mask value in the first larger sub-block corresponds to the minimum weight value, the size of the first BDOF sub-block is set to be smaller than the maximum supported BDOF sub-block size and greater than or equal to the minimum supported BDOF sub-block size. And / or When no mask value in the first larger sub-block corresponds to the maximum weight value or the minimum weight value, the size of the first BDOF sub-block is set to be smaller than the maximum supported BDOF sub-block size and greater than or equal to the minimum supported BDOF sub-block size.

8. The apparatus according to claim 6, wherein, The processing circuit system is configured to: Whether to apply BDOF refinement to said portion of the first larger sub-block is determined based on the mask value in at least a portion of the first larger sub-block.

9. The apparatus according to claim 8, wherein, The processing circuit system is configured to: When all mask values ​​in the portion of the first larger sub-block correspond to the maximum weight value or the minimum weight value, it is determined that the portion of the first larger sub-block will be refined using the BDOF. When there is no mask value of zero in the portion of the first larger sub-block, it is determined that the portion of the first larger sub-block is to be refined by the BDOF. When all mask values ​​in the portion of the first larger sub-block are higher than a threshold, it is determined that the portion of the first larger sub-block will be refined using the BDOF. as well as When all mask values ​​in the portion of the first larger sub-block are less than a threshold, it is determined that the portion of the first larger sub-block will be refined using the BDOF refinement.

10. The apparatus according to claim 2, wherein, The sub-block-based motion refinement is decoder-side motion vector refinement (DMVR) refinement, where the first sub-block is a first DMVR sub-block and the second sub-block is a second DMVR sub-block.

11. The apparatus according to claim 10, wherein, The processing circuit system is configured to: Divide the first GPM partition into multiple sub-blocks; and Whether to apply DMVR refinement to a specific sub-block is determined based on the mask value of the mixing mask in that specific sub-block.

12. The apparatus according to claim 10, wherein, The processing circuit system is configured to: The supported DMVR sub-block size is determined based on the size of the current block.

13. The apparatus according to claim 11, wherein, The processing circuit system is configured to: When all mask values ​​in the specific sub-block correspond to the maximum weight value or the minimum weight value, it is determined that the DMVR refinement is applied to the specific sub-block.

14. The apparatus according to claim 11, wherein, The processing circuit system is configured to: When it is determined that the DMVR refinement is to be applied to the specific sub-block, the DMVR is applied multiple times to the specific sub-block.

15. A video coding method, comprising: Determine whether to use GPM mode for the current block in the current image; The first GPM partition is determined to have bidirectional predicted motion vectors; The first GPM partition is divided into larger sub-blocks of size N×N, where N is a positive number and the size of N×N is greater than or equal to the maximum sub-block size supported by motion refinement based on sub-blocks; The size of each of the larger N×N sub-blocks is determined based on the mixing mask of the current block; The larger sub-block is divided into sub-blocks based on their respective sub-block sizes; and Whether to apply the sub-block-based motion refinement to the specific sub-block is determined based on the value of the blending mask in the specific sub-block.