Adaptive use of decoder-side motion vector refinement and bi-directional optical flow

By applying the motion refinement technology of a bidirectional motion predictor in the video decoder, the problem of low motion refinement efficiency in the prior art is solved, and more efficient video frame processing and lower code flow distortion are achieved.

CN119999201APending Publication Date: 2025-05-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004251.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-20
Filing Date
2024-04-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When existing video decoding technologies process video frames, it is difficult to effectively utilize bidirectional motion predictors, resulting in low motion refinement efficiency.

Method used

By introducing a motion refinement method based on a bidirectional motion predictor in the video decoder, including decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BDOF), and determining whether to continue to apply these technologies based on the specific information of the current block (such as time identification, quantization parameters, content difference measurement, etc.).

Benefits of technology

It improves the motion refinement efficiency of video frames, reduces distortion in the code stream, and improves the overall performance of video decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999201A_ABST
    Figure CN119999201A_ABST
Patent Text Reader

Abstract

Some aspects of the present disclosure provide a video decoding apparatus. The apparatus comprises processing circuitry configured to: receive an encoded video bitstream comprising encoded information of one or more pictures; according to the coded information, determining that a current block in the current picture meets a qualification condition of motion refinement based on a bidirectional motion predictor; beginning applying the bidirectional motion predictor-based motion refinement on at least a portion of the current block; obtaining specific information used during the application of the motion refinement based on the bidirectional motion predictor; and determining whether to continue to apply the motion refinement based on the bidirectional motion predictor according to the specific information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 460,876, filed on April 20, 2023, entitled “ADAPTIVE USAGE OF DECODER SIDE MOTION VECTOR REFINEMENT AND BI-DIRECTIONAL OPTICAL FLOW,” the entire contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure describes embodiments that relate generally to video decoding. Background Art

[0004] The background description provided herein is intended to present the background of the present disclosure as a whole. The extent to which the work of the presently named inventors described in the background section and various aspects of this specification is performed does not indicate that it is prior art at the time of filing this disclosure, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.

[0005] Image / video compression helps to transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, the video codec can use a technology called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current picture under reconstruction for sample prediction. In another example, the video codec can use a technology called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can predict samples in the current picture based on previously reconstructed pictures through motion compensation. Motion compensation can usually be represented by a motion vector (MV). Summary of the invention

[0006] Aspects of the present disclosure include methods and apparatus for video encoding / decoding.

[0007] Some aspects of the present disclosure provide a method for processing visual media data. The method includes processing a code stream of visual media data according to a format rule. The code stream includes encoded information of one or more pictures, and the one or more pictures include a current picture. The format rule specifies that a current block in the current picture is determined to meet a predefined condition for motion refinement based on a bidirectional motion predictor, and motion refinement based on a bidirectional motion predictor is applied on at least a portion of the current block. The format rule also specifies that a difference metric between a forward predictor and a backward predictor of the current block is obtained during the application of motion refinement based on a bidirectional motion predictor. The difference metric includes at least one of the following: a time identification difference between a forward predictor of the current block and a backward predictor of the current block; a quantization parameter difference between a forward predictor of the current block and a backward predictor of the current block; a content difference metric between a forward predictor of the current block and a backward predictor of the current block; and a scene change between a forward predictor of the current block and a backward predictor of the current block. Determine whether to continue to apply motion refinement based on a bidirectional motion predictor based on the difference metric.

[0008] Some aspects of the present disclosure provide a video decoding device. The device includes a processing circuit, which is configured to: receive an encoded video code stream including encoded information of one or more pictures; determine, based on the encoded information, whether a current block in the current picture meets the eligibility conditions for motion refinement based on a bidirectional motion predictor; start applying motion refinement based on a bidirectional motion predictor on at least a portion of the current block; obtain specific information used during the application of motion refinement based on a bidirectional motion predictor; and determine, based on the specific information, whether to continue applying motion refinement based on a bidirectional motion predictor.

[0009] In some examples, the bi-directional motion predictor based motion refinement includes at least one of decoder side motion vector refinement (DMVR) and / or bi-directional optical flow (BDOF).

[0010] In some examples, the specific information includes a time identifier of the current block, and the processing circuit is configured to: compare the time identifier of the current block with a threshold to obtain a comparison result; and disable motion refinement based on the bidirectional motion predictor based on the comparison result. In addition, the processing circuit is configured to perform at least one of the following: when the time identifier is greater than the threshold, disable motion refinement based on the bidirectional motion predictor; when the time identifier is less than the threshold, disable motion refinement based on the bidirectional motion predictor; or when the time identifier is less than a first threshold and greater than a second threshold, disable motion refinement based on the bidirectional motion predictor.

[0011] In some examples, the specific information includes a quantization parameter of the current block, and the processing circuit is configured to: compare the quantization parameter of the current block with a threshold to obtain a comparison result; and disable motion refinement based on the bidirectional motion predictor based on the comparison result. In addition, the processing circuit is configured to perform at least one of the following: when the quantization parameter is greater than the threshold, disabling motion refinement based on the bidirectional motion predictor; when the quantization parameter is less than the threshold, disabling motion refinement based on the bidirectional motion predictor; or when the quantization parameter is less than a first threshold and greater than a second threshold, disabling motion refinement based on the bidirectional motion predictor.

[0012] In some examples, the specific information includes a time stamp difference between a first time stamp of a forward predictor of the current block and a second time stamp of a backward predictor of the current block, and the processing circuit is configured to: compare the time stamp difference with a threshold to obtain a comparison result; and disable motion refinement based on the bi-directional motion predictor based on the comparison result. In some examples, the processing circuit is configured to perform at least one of the following: when the time stamp difference is greater than the threshold, disabling motion refinement based on the bi-directional motion predictor; when the time stamp difference is less than the threshold, disabling motion refinement based on the bi-directional motion predictor; or when the time stamp difference is less than the first threshold and greater than the second threshold, disabling motion refinement based on the bi-directional motion predictor.

[0013] In some examples, the specific information includes a quantization parameter difference between a first quantization parameter of a forward predictor of the current block and a second quantization parameter of a backward predictor of the current block, and the processing circuit is configured to: compare the quantization parameter difference of the current block with a threshold to obtain a comparison result; and disable motion refinement based on a bidirectional motion predictor based on the comparison result. In some examples, the processing circuit is configured to perform at least one of the following: when the quantization parameter difference is greater than the threshold, disable motion refinement based on a bidirectional motion predictor; when the quantization parameter difference is less than the threshold, disable motion refinement based on a bidirectional motion predictor; or when the quantization parameter difference is less than a first threshold and greater than a second threshold, disable motion refinement based on a bidirectional motion predictor.

[0014] In some examples, the specific information includes a content difference measure between a forward predictor of the current block and a backward predictor of the current block, and the processing circuit is configured to: compare the content difference measure with a threshold to obtain a comparison result; and disable motion refinement based on the bidirectional motion predictor based on the comparison result. In some examples, the processing circuit is configured to perform at least one of the following: when the content difference measure is greater than the threshold, disable motion refinement based on the bidirectional motion predictor; when the content difference measure is less than the threshold, disable motion refinement based on the bidirectional motion predictor; or when the content difference measure is less than a first threshold and greater than a second threshold, disable motion refinement based on the bidirectional motion predictor.

[0015] In some examples, the specific information indicates whether a scene change occurs between a forward predictor of the current block and a backward predictor of the current block, and the processing circuit is configured to disable motion refinement based on the bidirectional motion predictor when a scene change occurs.

[0016] In some examples, the motion refinement based on the bidirectional motion predictor is BDOF, the specific information includes one or more intermediate BDOF offset values, and the processing circuit is configured to: determine whether to continue to apply the motion refinement based on the bidirectional motion predictor according to the one or more intermediate BDOF offset values. In some examples, the one or more intermediate BDOF offset values ​​include at least one of the following: a gradient value, a cross-correlation value, and an autocorrelation value.

[0017] In some examples, the motion refinement based on the bidirectional motion predictor is BDOF, the specific information includes a BDOF offset, and the processing circuit is configured to: compare the BDOF offset with a threshold to obtain a comparison result; and determine whether to apply the BDOF offset according to the comparison result.

[0018] Some aspects of the present disclosure provide a video encoding method. The method includes: determining that a current block in a current picture satisfies a qualification condition for motion refinement based on a bidirectional motion predictor; starting to apply motion refinement based on a bidirectional motion predictor on at least a portion of the current block; obtaining specific information used during the application of motion refinement based on a bidirectional motion predictor; and determining whether to continue to apply motion refinement based on a bidirectional motion predictor based on the specific information. In some examples, the motion refinement based on the bidirectional motion predictor includes at least one of DMVR and / or BDOF.

[0019] In some examples, the specific information includes at least one of: a time identifier of the current block; a quantization parameter of the current block; a time identifier difference between a forward predictor of the current block and a backward predictor of the current block; a quantization parameter difference between the forward predictor of the current block and the backward predictor of the current block; a content difference measure between the forward predictor of the current block and the backward predictor of the current block; and a scene change between the forward predictor of the current block and the backward predictor of the current block.

[0020] In some examples, the motion refinement based on the bidirectional motion predictor is BDOF, and the specific information includes at least one of the following: a gradient value, a cross-correlation value, an autocorrelation value, and a BDOF offset.

[0021] According to another aspect of the present disclosure, a device is provided, wherein the device comprises a processing circuit, and the processing circuit is configured to perform any of the video decoding / encoding methods described.

[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer, cause the computer to perform any of the described video decoding / encoding methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0024] Figure 1 is a schematic diagram of an exemplary block diagram of a communication system.

[0025] Figure 2 is a schematic diagram of an exemplary block diagram of a decoder.

[0026] Figure 3 is a schematic diagram of an exemplary block diagram of an encoder.

[0027] Figure 4 The positions of spatial merging candidates according to an embodiment of the present disclosure are shown.

[0028] Figure 5 Candidate pairs considered for redundancy check of spatial merging candidates according to an embodiment of the present disclosure are shown.

[0029] Figure 6 Exemplary motion vector scaling for temporal merging candidates is shown.

[0030] Figure 7 Exemplary candidate positions of temporal merging candidates for the current block are shown.

[0031] Figure 8 The SbTMVP process used in the subblock-based temporal motion vector prediction (SbTMVP) mode is shown.

[0032] Fig. 9 An exemplary schematic diagram of decoder-side motion vector refinement based on bilateral matching in some examples is shown.

[0033] Fig.10 A schematic diagram showing some calculations in BDOF in an example.

[0034] Fig.11 The search areas in some examples are shown.

[0035] Fig.12 A flow chart outlining a decoding process according to some embodiments of the present disclosure is shown.

[0036] Fig.13A flow chart outlining an encoding process according to some embodiments of the present disclosure is shown.

[0037] Fig.14 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION

[0038] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application for the disclosed subject matter, video encoders, and video decoders in a streaming environment. The disclosed subject matter can be equally applicable to other video-enabled applications including, for example, video conferencing, digital television, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0039] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101) such as a digital camera, which creates an uncompressed video picture stream (102). In an embodiment, the video picture stream (102) includes samples captured by the digital camera. Compared to the encoded video data (104) (or the encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (102) can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream (102), the encoded video data (104) (or the encoded video bitstream (104)) is depicted as a thin line to emphasize the lower amount of data of the encoded video data (104) (or the encoded video bitstream (104)), which can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a streaming server (105) to retrieve a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T H.265. In an embodiment, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application may be used in the context of the VVC standard.

[0040] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0041] Figure 2 is an example block diagram of a video decoder (210). The video decoder (210) may be disposed in an electronic device (230). The electronic device (230) may include a receiver (231) (eg, a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) of an embodiment.

[0042] The receiver (231) may receive one or more encoded video sequences, such as included in a bitstream, to be decoded by the video decoder (210). In one embodiment, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (431) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided outside the video decoder (210) (not shown). In other cases, a buffer memory (not shown) is provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided inside the video decoder (210) to, for example, handle broadcast timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, in order to be used on a service packet network such as the Internet, a buffer memory (215) may also be required, and the buffer memory may be relatively large and may have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not shown) outside the video decoder (410).

[0043] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The types of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown in . The control information for the display device may be a parameter set fragment (not indicated) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (220) may extract a subgroup parameter set of at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0044] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215), thereby creating symbols (221).

[0045] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser (220) from the coded video sequence. For the sake of brevity, such subgroup control information flow between parser (220) and the multiple units below is not described.

[0046] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0047] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) can output a block including sample values, which can be input into an aggregator (255).

[0048] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of a current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using reconstructed information extracted from the current picture buffer (458). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0049] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221), these samples may be added to the output of the scaler / inverse transform unit (251) (in this case referred to as residual samples or residual signals) by the aggregator (255) to generate output sample information. The acquisition of the predicted samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (253) in the form of the symbols (221), for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0050] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). In other embodiments, the video compression may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values.

[0051] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for subsequent inter-picture prediction.

[0052] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0053] The video decoder (210) may perform decoding operations according to, for example, the ITU-T H.265 standard or a predetermined video compression technology. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0054] In an embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0055] Figure 3 is an example block diagram of a video encoder (303). The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (eg, a transmission circuit). The video encoder (303) may be used to replace Figure 1 A video encoder (103) in an embodiment.

[0056] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In another embodiment, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source can collect video images to be encoded by the video encoder (503). In another embodiment, the video source (301) is a part of the electronic device (320).

[0057] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0058] According to an embodiment, the video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints as required. Implementing an appropriate encoding speed is a function of the controller (350). In some embodiments, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, couplings are not shown in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (group of pictures, GOP) layout, maximum motion vector search range, etc. The controller (350) can be used to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.

[0059] In some embodiments, the video encoder (303) operates in a coding loop. As a simple description, in embodiments, the coding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0060] The operation of the "local" decoder (333) may be combined with, for example, Figure 2 The "remote" decoder described in detail for the video decoder (210) is identical. However, additional brief reference is made to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0061] In one embodiment, the decoder technology, except the parsing / entropy decoding present in the decoder, is also present in the corresponding encoder in the same or substantially the same functional form. Therefore, the application focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0062] During operation, in some embodiments, the source encoder (330) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0063] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 3 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (334). In this way, the video encoder (303) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0064] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new picture. The predictor (335) may operate pixel-by-pixel based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (334).

[0065] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0066] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0067] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0068] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0069] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0070] A predictive picture (P picture) may be a picture that may be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and a reference index to predict sample values ​​of each block.

[0071] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0072] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or the blocks may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0073] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0074] In an embodiment, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0075] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0076] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block may be predicted by a combination of the first reference block and the second reference block.

[0077] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.

[0078] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Furthermore, each CTU can be split into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0079] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.

[0080] Various inter prediction modes can be used in video coding. For example, in VVC, for an inter-predicted CU, motion parameters may include (multiple) MVs, one or more reference picture indices, a reference picture list usage index, and additional information of certain coding features to be used for sample generation for inter prediction. Motion parameters may be written explicitly or implicitly. When a CU is encoded using skip mode, the CU may be associated with a PU and may have no significant residual coefficients, no encoded motion vector increments or MV differences (e.g., MVDs), or reference picture indices. A merge mode may be specified, in which the motion parameters of the current CU are obtained based on adjacent (multiple) CUs (including spatial candidates and / or temporal candidates) and optional additional information (e.g., information introduced in VVC). The merge mode may be applied to inter-predicted CUs, not just for skip mode. In the example, an alternative to the merge mode is to explicitly transmit motion parameters, where each CU explicitly writes (multiple) MVs, reference picture indices corresponding to each reference picture list, reference picture list usage flags, and other information.

[0081] In one embodiment, for example, in VVC, the VVC Test model (VTM) reference software includes one or more refined inter-frame prediction coding tools, including: extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD writing, affine motion compensation prediction, SbTMVP, adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luminance sample MV storage and 8×8 motion field compression), bi-prediction with CU-level weights (BCW), BDOF, prediction refinement using optical flow (PROF), DMVR, combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter-frame prediction and related methods are described in detail below.

[0082] In some examples, extended merge prediction can be used. In an example, for example in VTM4, a merge candidate list is constructed by including the following five types of candidates in order: (multiple) spatial motion vector predictors (MVPs) from (multiple) spatially adjacent CUs, (multiple) temporal MVPs from (multiple) co-located CUs, history-based MVPs (HMVPs) from first-in-first-out (FIFO) tables, (multiple) pairwise average MVPs, and (multiple) zero MVs.

[0083] The size of the merge candidate list can be written in the slice header. In the example, the maximum allowed size of the merge candidate list in VTM4 is 6. For each CU encoded in merge mode, the index of the best merge candidate (e.g., merge index) can be encoded using truncated unary binarization (TU). The first binary bit of the merge index can be encoded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), while the other binary bits can use bypass encoding.

[0084] Some examples of the generation process of merge candidates for each category are provided below. In one embodiment, the derivation of (multiple) spatial candidates is as follows. The derivation of spatial merge candidates in VVC can be the same as the derivation of spatial merge candidates in HEVC. In the example, at Figure 4 Select up to four merge candidates from the candidates for the positions depicted in .

[0085] Figure 4 FIG. 4 shows the positions of spatial merging candidates according to an embodiment of the present disclosure. Figure 4 , the order of derivation is B1, A1, B0, A0 and B2. Position B2 is considered only when any CU in positions A0, B0, B1 and A1 is unavailable (for example, because the CU belongs to another slice or another tile) or is intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0086] In order to reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only Figure 5 The pairs linked by arrows in , and a candidate is added to the candidate list only if the corresponding candidates used for redundancy check do not have the same motion information.

[0087] Figure 5 FIG. 4 shows candidate pairs considered for redundancy check of spatial merging candidates according to an embodiment of the present disclosure. Figure 5 , pairs linked by corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0088] In an embodiment, the derivation of the temporal candidate(s) is as follows. In the example, only one temporal merging candidate is added to the candidate list. Figure 6 An exemplary motion vector scaling for a temporal merge candidate is shown. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a temporal merge candidate may be derived based on a collocated CU (612) belonging to a collocated reference picture (604) (e.g., by Figure 6 The reference picture list used to derive the co-located CU (612) may be explicitly written in the slice header. Figure 6 As shown by the dotted line in , a scaled MV (621) for the temporal merge candidate can be obtained. The scaled MV (621) can be scaled according to the MV of the co-located CU (612) using picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the co-located reference picture (604) of the co-located picture (603) and the co-located picture (603). The reference picture index of the temporal merge candidate can be set to zero. A co-located picture is a reference picture that is used as a source picture for temporal motion information derivation. The co-located picture can be identified in one of two lists (referred to as list 0 or list 1). In some instances, the encoder can determine the co-located picture and write the co-located picture using appropriate syntax techniques.

[0089] Figure 7 Exemplary candidate positions (e.g., C0 and C1) of the temporal merge candidate for the current CU are shown. The position of the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is located at the lower right corner of the co-located CU (710) of the current CU. Candidate position C1 is located at the center of the co-located CU (710) of the current CU. If the CU at candidate position C0 is not available, intra-coded, or outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate. The temporal merge candidate can specify motion information of a temporal motion vector predictor (TMVP).

[0090] In order to improve coding efficiency and reduce the transmission overhead of (multiple) MVs, sub-block level MV refinement can be applied to extend CU-level temporal motion vector prediction (TMVP). In an example, a sub-block-based TMVP (SbTMVP) mode allows inheritance of sub-block level motion information from a co-located reference picture. A co-located reference picture can be indicated by a reference index in a syntax, such as a high-level syntax (e.g., a picture header, a slice header). Each of a plurality of sub-blocks in a current CU in a current picture (e.g., a current CU of a large size) can have corresponding motion information without explicitly sending a block partition structure or corresponding motion information. In the SbTMVP mode, the motion information of each sub-block can be obtained, for example, by the following three steps. In the first step, a displacement vector (DV) of the current CU can be derived. DV can indicate a block in a co-located reference picture, for example, DV points from a current block in a current picture to a block in a co-located reference picture. Therefore, the block indicated by DV is considered to be co-located with the current block and is referred to as a co-located block of the current block. In the second step, the availability of SbTMVP candidates can be checked, and then the center motion (e.g., the center motion of the current CU) can be derived. In the third step, the sub-block motion information can be derived according to the corresponding sub-block in the co-located block using DV. These three steps can be combined into one or two steps, and / or the order of the three steps can be adjusted.

[0091] Unlike TMVP candidate derivation that derives multiple temporal MVs from co-located blocks in a reference frame or reference picture, in SbTMVP mode, for each sub-block in the current CU in the current picture, a DV (e.g., a DV derived from the MV of the left neighboring CU of the current CU) can be applied to locate the corresponding sub-block in the co-located reference picture. In some examples, when the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set to the center motion of the co-located block.

[0092] The SbTMVP mode can be supported by various video coding standards including, for example, VVC. Similar to the TMVP mode in, for example, HEVC, in the SbTMVP mode, the motion field (also called the motion information field or MV field) in the co-located reference picture can be used to improve MV prediction and the merge mode of multiple CUs in the current picture. In the example, the same co-located reference picture as the co-located reference picture used by the TMVP mode is used in the SbTMVP mode. In the example, the SbTMVP mode differs from the TMVP mode in the following aspects: (i) the TMVP mode predicts motion information at the CU level, while the SbTMVP mode predicts motion information at the sub-CU level; (ii) the TMVP mode obtains multiple temporal MVs from a co-located block in the co-located reference picture (for example, the co-located block is the lower right block or the center block relative to the current CU), while the SbTMVP mode can apply motion shifting before obtaining temporal motion information from the co-located reference picture. In the example, the motion shift used in the SbTMVP mode is obtained from the MV of one of the multiple spatial neighboring blocks of the current CU.

[0093] Figure 8 The SbTMVP process used in the SbTMVP mode is shown. The SbTMVP process can predict multiple MVs of multiple sub-CUs (e.g., multiple sub-blocks) within the current CU (e.g., current block) (801) in the current picture (811), for example, in two steps. In the first step, check Figure 8 A spatial neighboring block (e.g., A1) of the current block (801) in FIG. 8 is selected. If the spatial neighboring block (e.g., A1) has an MV (821) that uses the co-located reference picture (812) as a reference picture for the spatial neighboring block (e.g., A1), then the MV (821) can be selected as the motion shift (or DV) to be applied to the current block (801). If no such MV is identified (e.g., an MV that uses the co-located reference picture (812) as a reference picture), the motion shift or DV can be set to a zero MV (e.g., (0,0)). In some examples, if no such MV is identified for the spatial neighboring block A1, then the MV(s) of multiple additional spatial neighboring blocks (e.g., A0, B0, B1, etc.) are checked.

[0094] In a second step, the motion shift or DV (821) identified in the first step may be applied to the current block (801) (e.g., DV (821) is added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., including multiple MVs and multiple reference indices) from the co-located reference picture (812). Figure 8In the example shown, a motion shift or DV (821) is set to the MV of a spatial neighbor A1 (e.g., block A1) of the current block (801). Motion information of multiple sub-CUs or multiple sub-blocks (831) in the current picture (811) can be derived using motion information of multiple sub-blocks in a corresponding co-located block (802) in a co-located reference picture (812). For example, after identifying the motion information of a co-located sub-CU (832) in the co-located block (802), the motion information of the co-located sub-CU (832) can be converted to motion information (e.g., (multiple) MVs and one or more reference indices) of the current sub-CU (831) using a scaling method, for example, in a manner similar to the TMVP process used in HEVC, where temporal motion scaling is applied to align multiple reference pictures of the multiple temporal MVs with multiple reference pictures of the current CU.

[0095] The motion field of the current block (801) derived based on the DV (821) may include motion information (e.g., (multiple) MVs and one or more associated reference indices) of each sub-block (831) in the current block (801). The motion field of the current block (801) may also be referred to as an SbTMVP candidate and corresponds to the DV (821).

[0096] For example, the motion information of the bidirectionally predicted subblock (831(1)) includes: a first MV, a first index indicating a first reference picture in reference picture list 0 (L0), a second MV, and a second index indicating a second reference picture in reference picture list 1 (L1). In the example, the motion information of the unidirectionally predicted subblock (831(2)) includes an MV and an index indicating a reference picture in L0 or L1.

[0097] In the example, DV (821) is applied to the center position of the current block (801) to locate the shifted center position in the co-located reference picture (812). If the block including the shifted center position is not inter-coded, the SbTMVP candidate is considered unavailable. Otherwise, if the block including the shifted center position (e.g., the co-located block (802)) is inter-coded, the motion information of the center position of the current block (801) can be derived based on the motion information of the co-located block (802) including the shifted center position in the co-located reference picture (812) (referred to as the center motion of the current block (801)). In the example, the center motion of the current block (801) can be derived based on the motion information of the co-located block (802) including the shifted center position in the co-located reference picture (812) using a scaling process. When SbTMVP candidates are available, for each sub-block (831) of the current block (801), DV (821) can be applied to find the corresponding sub-block (832) in the co-located reference picture (812). The motion information of the corresponding sub-block (832) can be used to derive the motion information of the sub-block (831) in the current block (801), for example, in the same manner as used to derive the center motion of the current block (801). In an example, if the corresponding sub-block (832) is not inter-coded, the motion information of the current sub-block (831) is set to the center motion of the current block (801).

[0098] In some examples, such as in VVC, when writing a sub-block based merge mode, a sub-block based merge list including a combination of SbTMVP candidates and (multiple) affine merge candidates is used. The SbTMVP mode can be enabled or disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP candidate (or SbTMVP predictor) can be added as the first entry of a sub-block based merge list, which includes a sub-block based merge candidate and followed by (multiple) affine merge candidates. The size of the sub-block based merge list can be written in the SPS. In the example, the maximum allowed size of the sub-block based merge list in VVC is 5. In the example, multiple SbTMVP candidates are included in the sub-block based merge list.

[0099] In some examples, such as in VVC, the sub-CU size used in SbTMVP mode is fixed to 8×8, which is the same size used in affine merge mode. In an example, SbTMVP mode is only applicable to CUs with width and height both greater than or equal to 8. The sub-block size (e.g., 8×8) can be configured to other sizes, such as 4×4 in the ECM software model for exploration outside of VVC. In an example, multiple co-located reference pictures (e.g., two co-located frames) are used to provide temporal motion information for SbTMVP and / or TMVP.

[0100] In inter-picture prediction, merge mode can be used to improve coding efficiency. In merge mode, a motion vector can be derived from multiple adjacent blocks, and the motion vector is directly used for motion compensation. In order to improve the accuracy of multiple MVs in merge mode, decoder-side motion vector refinement (DMVR) based on bilateral matching (BM) can be applied, for example, in VVC. In a bidirectional prediction operation, a refined MV can be searched around multiple initial MVs in reference picture list L0 and reference picture list L1. BM calculates the distortion between two candidate blocks in reference picture list L0 and list L1.

[0101] Fig. 9 FIG. 4 shows an exemplary schematic diagram of decoder-side motion vector refinement based on BM in some examples. Fig. 9 As shown, the current picture (902) may include a current block (908). The current picture may have a first reference picture (904) from a (reference picture) list L0 and a second reference picture (906) from a (reference picture) list L1. For the current block (908), a pair of reference blocks are identified in the first reference picture and the second reference picture based on initial motion vectors MV0 and MV1. For example, the initial reference block (912) in the first reference picture (904) may be located based on the initial motion vector MV0, and the initial reference block (914) in the second picture (906) may be located based on the initial motion vector MV1. A search process may be performed around the initial MV0 in the first reference picture (904) and the initial MV1 in the second reference picture (906). For example, the MV0 may be adjusted in opposite directions. diffApplied to initial MV0 and MV1 to obtain MV candidates, such as MV0' and MV1'. Based on the MV candidates, a pair of candidate reference blocks are identified in the first reference picture and the second reference picture. For example, a candidate reference block (910) can be identified in the first reference picture (904) based on MV0', and a candidate reference block (916) can be identified in the second reference picture (906) based on MV1'. In some examples, BM refers to an operation of calculating a distortion metric between a pair of reference blocks of corresponding reference pictures of the current picture, such as taking a sum of absolute differences (SAD) between a pair of reference blocks as a distortion metric for the pair of reference blocks. For example, the BM method calculates an initial SAD between a pair of initial reference blocks (912) and (914), and calculates a second SAD between a pair of candidate reference blocks (910) and (916). The initial SAD is associated with an initial MV (e.g., MV0 and MV1), and the second SAD is associated with an MV candidate (e.g., MV0' and MV1'). Similarly, the BM method can calculate the SAD of multiple MV candidates around the initial MV. The MV candidate with the lowest SAD may become the refined MV and used to generate a bi-directional prediction write to predict the current block (908).

[0102] In some examples (eg, VVC), the application of DMVR is restricted and applied only to multiple CUs coded with modes and features that meet certain conditions (also called DMVR eligibility conditions). If a block meets certain conditions, the DMVR algorithm is invoked. For example, conditions (also referred to as DMVR eligibility conditions, DMVR requirements, or a set of conditions for DMVR) may include: (1) a CU-level merge mode with a bidirectionally predicted MV; (2) relative to the current picture, one reference picture is before and the other reference picture is after; (3) the distances from the two reference pictures to the current picture (e.g., POC difference) are the same; (4) both reference pictures are short-term reference pictures (in the example, long-term and short-term are used to describe the buffer management of multiple reference pictures, long-term reference pictures can stay in the buffer without being overwritten, and short-term reference pictures may be overwritten by newly generated pictures in the buffer); (5) the CU has more than 64 luma samples; (6) the CU height and CU width are both greater than or equal to 8 luma samples; (7) the bidirectional prediction (BCW) weight index with CU-level weights indicates equal weights; (8) weighted prediction (WP) is not enabled for the current block; and (9) the joint inter- and intra-frame prediction (CIIP) mode is not used for the current block.

[0103] It should be noted that the refined MV derived by the DMVR process is used to generate multiple inter-frame prediction samples and can be used in temporal motion vector prediction for future picture encoding. In some examples, the original MV is used in the deblocking process and is also used in spatial motion vector prediction for future CU encoding.

[0104] In DVMR, the search point is around the initial MV, and the MV offset follows the MV difference mirror rule. Any point checked by DMVR (represented by the candidate MV pair (MV0', MV1')) follows MV0'=MV0+MV_offset and MV1'=MV1-MV_offset. Where MV_offset represents the refinement offset between the initial MV (e.g., (MV0, MV1) and the refined MV in one of the reference pictures. In some examples, the refinement search range is two integer luma samples from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage.

[0105] In some examples (e.g., VVC), decoder-side motion vector refinement (DMVR) is applied to CUs encoded in normal merge mode. The MV pairs obtained from normal merge candidates are used as input to the DMVR process. DMVR applies bilateral matching (BM) to refine the input MV pair {MV0, MV1} and uses the refined MV pair {MV 细化的L0 , MV 细化的L1} Perform motion compensation prediction on the luminance component and chrominance component, such as Figure 4 As shown. The multiple output MVs of DMVR can be called refined MV pairs and can be expressed by equation (1).

[0106]

[0107] By using the MVD mirroring property, the motion vector difference Δmv is applied to the input MV pair to obtain a refined MV pair. Because the input MV pair points to two different reference pictures, the two reference pictures have equal picture order count (POC) differences with the current picture, and the two reference pictures are in different temporal directions.

[0108] In some examples, DMVR may be applied at the sub-block level, with the luma coding block being partitioned into 16×16 sub-blocks for the MV refinement process. Δmv is derived independently for each sub-block.

[0109] In some examples, the motion vector refinement search range is two integer luma samples from the initial MV. The search for the motion vector may be performed in two steps, for example, the first step being an integer sample offset search phase (also referred to as integer precision motion search) and the second step being a fractional sample refinement phase (also referred to as a fractional motion search step or a fractional sample offset search).

[0110] In some examples, a 25-point full search may be applied to an integer sample offset search, as shown in equation (2):

[0111]

[0112] Where (i, j) represents the coordinates of the search points around the initial MV pair, and i and j are integer values ​​between -2 (inclusive) and 2 (inclusive). First, the SAD of the initial MV pair is calculated, for example, according to equation (3):

[0113]

[0114] diff m,n =abs(P0 i,j [m+i,2n+j]-P1 i,j [mi,2n-j])

[0115] and Among them, W is the weight of the sub-block, and H is the height of the sub-block.

[0116] If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. In order to reduce the penalty of uncertainty in DMVR refinement, the original MVs (e.g., initial MV candidates MV0 and MV1) can be preferred during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by, for example, 1 / 4 of the SAD value to make the initial MV candidates the preferred candidates.

[0117] In some examples, the integer sample search is followed by fractional sample refinement. In some examples, fractional sample refinement is performed by using fractional sample offsets, such as 1 / 2 pixel offsets in the vertical and horizontal directions, etc. In some examples, in order to save computational complexity, fractional sample refinement is derived by using parameter error surface equations (also known as quadratic prediction-based methods) rather than by additional searches and SAD comparisons. Fractional sample refinement is conditionally called based on the output of the integer sample search stage. For example, if the integer sample search stage terminates with a specific integer position (also known as the center) with the minimum SAD in the first iteration or the second iteration search, fractional sample refinement is further applied.

[0118] In the parametric error surface based sub-pixel offset estimation, the center position cost (the center position is the point with the minimum SAD in the integer sample offset search) and the costs of the four neighboring positions from the center (e.g., (-1,0), (0,-1), (1,0), (0,1) from the center position) are used to fit the 2-D parabolic error surface equation, such as equation (4):

[0119] E(x,y)=A(x min ) 2 +B(yy min ) 2 +C equation (4)

[0120] Among them, (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equations using the cost values ​​of the five search points, (x min ,y min ):

[0121] x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) Equation (5)

[0122] y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) Equation (6)

[0123] x min and min The value of can be automatically constrained between -8 and 8, since all cost values ​​are positive and the minimum is E(0,0). The constraint corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated score (x min ,y min ) is added to the integer distance refinement MV to obtain a refinement increment MV with sub-pixel accuracy. In equations (5) and (6), E(-1,0), E(1,0), E(1,0), E(0,-1), and E(0,0) represent the cost values ​​of five points (the center position and four neighboring positions).

[0124] A technique called bidirectional optical flow (BDOF) may be used, for example, in VVC. BDOF was previously known as BIO in JEM. BDOF in VVC may be a simpler version than the JEM version, requiring fewer computations, particularly in terms of the number of multiplications and the size of the multipliers.

[0125] BDOF can be used to refine the bidirectional prediction write of a CU at the 4×4 sub-block level. BDOF can be applied to a CU if the CU meets the following conditions (also called BDOF eligibility conditions, BDOF requirements, or a set of conditions for BDOF): (1) the CU is encoded using a "true" bi-prediction mode, i.e., one of the two reference pictures is earlier than the current picture in the display order, and the other is later than the current picture in the display order; (2) the distances from the two reference pictures to the current picture (e.g., POC difference) are the same; both reference pictures are short-term reference pictures; the CU is not encoded using the affine mode or the SbTMVP merge mode; (5) the CU has more than 64 luma samples; (6) the CU height and CU width are both greater than or equal to 8 luma samples; (7) the BCW weight index indicates equal weights; (8) weighted prediction (WP) is not enabled for the current CU; and (9) the CIIP mode is not used for the current CU.

[0126] In some examples, BDOF is applied only to the luma component. As the name of BDOF indicates, the BDOF mode can be based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ). The motion refinement can then be used to adjust the bi-directionally predicted sample values ​​in the 4×4 sub-block. BDOF may include the following steps.

[0127] First, the horizontal gradient of the two prediction signals from reference list L0 and reference list L1 can be calculated by directly calculating the difference between two adjacent samples and vertical gradient The horizontal gradient and the vertical gradient can be provided in equations (7) and (8) as follows:

[0128]

[0129] Among them, I (k) (i, j) may be the sample value at coordinate (i, j) of the predicted write in list k (k=0, 1), and shift1 may be calculated based on the luma bit depth bitDepth, since shift1=max(6, bitDepth-6).

[0130] Then, the autocorrelations and cross-correlations S1, S2, S3, S5, and S6 of the gradients can be calculated according to the following equations (9)-(13):

[0131] S1=∑ (i,j)∈Ω Abs(ψ x (i,j)) Equation (9)

[0132] S2=∑ (i,j)∈Ω ψ x (i,j)·Sign(ψ y (i,j)) Equation (10)

[0133] S3=∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) Equation (11)

[0134] S5=∑ (i,j)∈Ω Abs(ψ y (i,j)) Equation (12)

[0135] S6=∑ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) Equation (13)

[0136] where ψ can be provided in equations (14)-(16) respectively. x (i,j),ψ y (i,j) and θ(i,j).

[0137]

[0138] θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b ) Equation (16)

[0139] where Ω can be a 6×6 window around a 4×4 sub-block, and n a and n b The values ​​of can be set equal to min(1, bitDepth-11) and min(4, bitDepth-8) respectively.

[0140] The motion refinement (v) can then be derived using the cross-correlation and autocorrelation terms using equations (17) and (18) as follows x ,v y ):

[0141]

[0142] in, th′ BIO =2 max(5,BD-7) . is the floor function, and Based on the motion refinement and gradient, the adjustment can be calculated for each sample in the 4×4 sub-block based on equation (19):

[0143]

[0144] Finally, the BDOF samples of the CU can be calculated by adjusting the bidirectional prediction samples as follows:

[0145] pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+O offset )>>shift Equation (20)

[0146] The values ​​may be selected such that the multiplier in the BDOF process does not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process may be kept within 32 bits.

[0147] Fig.10 Some calculations in BDOF in the example are shown. In order to derive the gradient value, it is necessary to generate some prediction samples I in the list k (k = 0, 1) outside the current CU boundary (k) (i,j). Fig.10 As shown, BDOF in VVC can use an extended row / column (1002) around the boundary (1006) of the CU (1004). In order to control the computational complexity of generating prediction samples outside the boundary, the extended area (e.g., Fig.10 The prediction samples in the non-shaded area in ), and a conventional 8-tap motion compensation interpolation filter can be used to generate the CU (e.g., Fig.10 The extended sample values ​​can only be used for gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they can be obtained by filling (e.g., repeating) from the nearest neighbors of the samples and gradient values.

[0148] In some examples, a sample-based BDOF may be used instead of a block-based BDOF. In a sample-based BDOF, motion refinements (v x ,v y ), but is performed per sample. The coding block is divided into 8×8 sub-blocks. For each sub-block, it is determined whether to apply BDOF by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to a sub-block, a 5×5 sliding window is used for each sample in the sub-block, and the existing BDOF process is applied for each sliding window to derive v x and v y . The derived motion refinement (vx ,v y ) is applied to the bidirectionally predicted sample value of the center sample of the adjustment window.

[0149] In some examples, multi-stage DMVR can be used. In an example, in the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16×16 sub-block within the coding block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MV is stored for spatial motion vector prediction and temporal motion vector prediction.

[0150] Specifically, the first stage performs block-based bilateral matching MV refinement. In the first stage, a refined MV is derived by applying BM to the coding block. Similar to decoder-side motion vector refinement (DMVR), in a bidirectional prediction operation, the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. Based on the minimum bilateral matching cost between the two reference blocks in L0 and L1, the refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs. The bilateral matching cost can be calculated by any suitable error measurement metric that measures the error between the two reference blocks in L0 and L1. In the example, the bilateral matching cost includes a term that is the sum of absolute differences (SAD) between corresponding samples in the two reference blocks in L0 and L1.

[0151] BM can perform a local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to loop through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension and the maximum values ​​of sHor and sVer are 8.

[0152] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW×cbH is greater than 64, the mean removed SAD (MRSAD) cost function is applied to remove the DC effect of the distortion between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.

[0153] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first stage is then derived as:

[0154] MV0_pass1=MV0+deltaMV Equation (21)

[0155] MV1_pass1 = MV1 – deltaMV equation (22)

[0156] In the second stage, subblock-based bilateral matching MV refinement is performed. Specifically, in the second stage, refined MVs are derived by applying BM to 16×16 grid subblocks. For each subblock, the refined MVs are searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first stage in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference subblocks in L0 and L1.

[0157] For each sub-block, BM performs a full search to derive integer sample precision intDeltaMV. The search range of the full search in the horizontal direction is [-sHor, sHor], and the search range in the vertical direction is [-sVer, sVer], where the values ​​of sHor and sVer are determined by the block dimension, and the maximum values ​​of sHor and sVer are 8.

[0158] The bilateral matching cost is calculated by applying the cost factor to the sum of absolute transformed differences (SATD) cost between the two reference sub-blocks, such as: bilCost = satdCost × costFactor. In some examples, the search area (2×sHor+1)×(2×sVer+1) is divided into up to 5 diamond search areas.

[0159] Fig.11 The search area (1100) in some examples is shown. The search area (1100) is divided into five search areas (1101) to (1105). The shape of the search area is similar to a diamond.

[0160] In some examples, each search region is assigned a cost factor (costFactor) determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in order starting from the center of the search region. In each region, the search points are processed in a raster scan order starting from the upper left corner to the lower right corner of the region. When the minimum bilCost in the current search region is less than a threshold (the threshold is equal to sbW×sbH), the integer-pel full search is terminated, otherwise, the integer-pel full search continues to the next search region until all search points are checked. In addition, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold (the threshold is equal to the area of ​​the block), the search process is terminated.

[0161] In some examples, fractional sample refinement (e.g., DMVR fractional sample refinement in VVC) is further applied to derive the final deltaMV (sbIdx2). The refined MV of the second stage is then derived as:

[0162] MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2) equation (23)

[0163] MV1_pass2(sbIdx2)=MV1_pass1–deltaMV(sbIdx2) Equation (24)

[0164] In the third stage, sub-block based bidirectional optical flow MV refinement can be performed. Specifically, in the third stage, the refined MV is derived by applying BDOF to the 8×8 grid sub-block. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy without clipping from the refined MV of the parent sub-block in the second stage. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample accuracy and clipped between -32 and 32. The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) of the third stage are derived as:

[0165] MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv equation (25)

[0166] MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)–bioMv equation (26)

[0167] Note that in some related codec examples, DMVR / BDOF will be applied on a CU when the CU meets the DMVR / BDOF eligibility criteria, and there is no mechanism to adaptively disable these tools at the block level or sub-block level. Also note that DMVR / BDOF may introduce additional distortion into the prediction due to compression errors, and the use of DMVR / BDOF is not always beneficial.

[0168] According to some aspects of the present disclosure, when a CU meets the DMVR / BDOF eligibility conditions, DMVR / BDOF is then applied. During the application of DMVR / BDOF for refinement, some operations are performed to obtain some specific block information or sub-block information, for example, to obtain intermediate values ​​or intermediate parameters for DMVR / BDOF. The specific block information or sub-block information can be used to determine whether DMVR / BDOF will introduce distortion. Some aspects of the present disclosure provide techniques for adaptively using DMVR / BDOF on a block based on specific block information, or techniques for adaptively using DMVR / BDOF on a sub-block based on specific sub-block information. Therefore, although a block or sub-block is eligible for refinement by DMVR / BDOF, the application of DMVR / BDOF can be adaptively disabled based on specific block information or sub-block information. For example, the encoder / decoder can determine that the current block in the current picture meets the eligibility conditions for motion refinement based on a bidirectional motion predictor, and start applying motion refinement based on a bidirectional motion predictor on at least a portion of the current block. The encoder / decoder may obtain specific information used during application of the bi-directional motion predictor based motion refinement, and determine whether to continue to apply the bi-directional motion predictor based motion refinement according to the specific information.

[0169] Some aspects of the present disclosure provide an adaptive mechanism that allows deciding at a block level or sub-block level whether to use motion refinement based on a bidirectional motion predictor, such as DMVR and / or BDOF prediction improvement. It should be noted that when DMVR is mentioned, it generally refers to all multi-stage DMVR, excluding the DMVR stage related to BDOF; and when BDOF is mentioned, it generally refers to all multi-stage DMVR and (multiple) sample-based BDOF methods related to the BDOF process. In addition, the proposed method can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-temporary computer-readable medium.

[0170] According to one aspect of the present disclosure, a block (sub-block) level decision is made to determine whether to use DMVR and / or BDOF prediction improvement for the current block (sub-block) based on specific block (sub-block) information.

[0171] In some examples, block (sub-block) level decisions are made based on current block timestamp (eg, represented by TId in some examples) information.

[0172] In an example, when the TId of the current block is greater than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that in some examples, the threshold is pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0173] In another example, when the TId of the current block is less than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that in some examples, the threshold is predefined or derived or written in the bitstream at the sequence level or the picture level or the slice level.

[0174] In another example, when the TId of the current block is less than (or equal to) the first threshold and greater than (or equal to) the second threshold, DMVR and / or BDOF prediction improvement are disabled for the current block, otherwise DMVR and / or BDOF prediction improvement are enabled for the current block. It should be noted that in some examples, the first threshold and the second threshold may be pre-defined or derived or written in the bitstream at the sequence level, the picture level, or the slice level.

[0175] In some examples, block (sub-block) level decisions are made based on current block quantization parameter (QP) information.

[0176] In an example, when the current block QP is greater than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that in some examples, the threshold is pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0177] In another example, when the current block QP is less than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that in some examples, the threshold may be pre-defined, derived, or written in the bitstream at the sequence level, picture level, or slice level.

[0178] In another example, when the QP of the current block is less than (or equal to) the first threshold and greater than (or equal to) the second threshold, DMVR and / or BDOF prediction improvement are disabled for the current block, otherwise DMVR and / or BDOF prediction improvement are enabled for the current block. It should be noted that in some examples, the first threshold and the second threshold may be pre-defined, derived, or written in the bitstream at the sequence level, the picture level, or the slice level.

[0179] According to another aspect of the present disclosure, a block (sub-block) level decision is made to determine whether to use DMVR and / or BDOF prediction improvement for the current block (sub-block) based on information related to the forward predictor and backward predictor of the current block (sub-block).

[0180] In some examples, the block (sub-block) level decision is made based on the time identification (represented by TID in some examples) difference of the forward predictor and the backward predictor.

[0181] In the example, when the absolute value of the TId difference between the forward predictor and the backward predictor is greater than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold can be pre-defined or derived or written in the code stream at the sequence level, the picture level, or the slice level.

[0182] In another example, when the absolute value of the TId difference between the forward predictor and the backward predictor is less than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold can be pre-defined, derived, or written in the bitstream at the sequence level, the picture level, or the slice level.

[0183] In another example, when the absolute value of the TId difference between the forward predictor and the backward predictor is less than (or equal to) a first threshold and greater than (or equal to) a second threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold value can be pre-defined, derived, or written in the bitstream at the sequence level, the picture level, or the slice level.

[0184] In some examples, block (sub-block) level decisions are made based on quantization parameter (QP) differences of the forward and backward predictors.

[0185] In the example, when the absolute value of the QP difference between the forward predictor and the backward predictor is greater than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold may be pre-defined, derived, or written in the bitstream at the sequence level, the picture level, or the slice level.

[0186] In another example, when the absolute value of the QP difference between the forward predictor and the backward predictor is less than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold may be pre-defined, derived, or written in the bitstream at the sequence level, the picture level, or the slice level.

[0187] In another example, when the absolute value of the QP difference between the forward predictor and the backward predictor is less than (or equal to) the first threshold and greater than (or equal to) the second threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the first threshold and the second threshold can be pre-defined, derived, or written in the code stream at the sequence level, the picture level, or the slice level.

[0188] In some examples, block (sub-block) level decisions are made based on content information of the forward predictor and the backward predictor. For example, block (sub-block) level decisions are made based on a content difference measure between the forward predictor and the backward predictor. The content difference metric may be any suitable metric, such as mean squared error (MSE), root mean squared error (RMSE), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), universal quality image index (UQI), multi-scale structural similarity index (MS-SSIM), error relative GlobaleAdimensionnellede Synthèse (ERGAS), spatial correlation coefficient (SCC), relative average spectral error (RASE), spectral angle mapper (SAM), spectral distortion index (D_lambda), spatial distortion index (D_S), quality with no reference (QNR), visual information fidelity (VIF), block sensitive-peak signal-to-noise ratio (PSNR), and so on. signal-to-noise ratio (PSNR), etc. Note also that the content difference metric can be formed as a combination of the following: mean square error (MSE), root mean square error (RMSE), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), universal quality image index (UQI), multi-scale structural similarity index (MS-SSIM), relative overall dimensionless integrated error (ERGAS), spatial correlation coefficient (SCC), relative average spectral error (RASE), spectral angle mapper (SAM), spectral distortion index (D_lambda), spatial distortion index (D_S), no-reference quality (QNR), visual information fidelity (VIF), block-sensitive peak signal-to-noise ratio (PSNR), etc.

[0189] Although MSE is used in some examples, it is noted that similar techniques can also be applied to other content difference metrics.

[0190] In an example, when the mean square error (MSE) between the forward predictor and the backward predictor is less than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold can be pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0191] In another example, when the MSE metric between the forward predictor and the backward predictor is greater than (or equal to) a threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold can be pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0192] In another example, when the MSE metric between the forward predictor and the backward predictor is less than (or equal to) a first threshold and greater than (or equal to) a second threshold, DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block. It should be noted that the threshold value can be pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0193] In some examples, a block (sub-block) level decision is made based on the fact that the forward predictor and the backward predictor belong to the same scene within the video sequence. For example, when a scene change occurs between the forward predictor and the backward predictor (e.g., the forward predictor and the backward predictor belong to different scenes in the video sequence), DMVR and / or BDOF prediction improvement is disabled for the current block, otherwise DMVR and / or BDOF prediction improvement is enabled for the current block.

[0194] According to another aspect of the present disclosure, a block (sub-block) level decision is made to determine whether to use BDOF prediction improvement for the current block (sub-block) based on the intermediate BDOF offset calculation. In one example, the block (sub-block) level decision is made based on the gradient value. In another example, the block (sub-block) level decision is made based on the cross-correlation value and / or the autocorrelation value.

[0195] According to another aspect of the present disclosure, a block (sub-block) level decision is made to determine whether to use BDOF prediction improvement for the block (sub-block) based on the BDOF offset calculated for the current block (sub-block). In an example, when the BDOF offset calculated for a particular sample is greater than (or equal to) an absolute threshold, the BDOF offset is not applied to the sample. The threshold can be pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0196] In another example, when the BDOF offset calculated for a particular sample is greater than (or equal to) a relative threshold (e.g., a percentage relative to the sample), the BDOF offset is not applied to the sample. It should be noted that the threshold can be pre-defined or derived or written in the code stream at the sequence level or the picture level or the slice level.

[0197] Fig.12 A flow chart outlining a process (1200) according to an embodiment of the present disclosure is shown. The process (1200) may be used in a video decoder. In various embodiments, the process (1200) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some embodiments, the process (1200) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1200). The process starts at (S1201) and proceeds to (S1210).

[0198] At (S1210), an encoded video stream including encoded information of one or more pictures is received.

[0199] At (S1220), it is determined that the current block in the current picture satisfies the eligibility condition for motion refinement based on the bidirectional motion predictor according to the encoded information.

[0200] At (S1230), start applying motion refinement based on a bi-directional motion predictor on at least a portion of the current block. The portion of the current block may be the current block or may be a sub-block of the current block.

[0201] At (S1240), specific information used during application of bi-directional motion predictor based motion refinement is obtained.

[0202] At (S1250), it is determined whether to continue to apply the motion refinement based on the bidirectional motion predictor according to the specific information.

[0203] In some examples, the bi-directional motion predictor based motion refinement is a decoder-side motion vector refinement (DMVR). In some examples, the bi-directional motion predictor based motion refinement is a bi-directional optical flow (BDOF).

[0204] In some examples, the specific information includes a time identifier of the current block. For example, the time identifier of the current block is compared with a threshold to obtain a comparison result, and the motion refinement based on the bidirectional motion predictor is disabled based on the comparison result. In an example, when the time identifier is greater than the threshold, the motion refinement based on the bidirectional motion predictor is disabled; in another example, when the time identifier is less than the threshold, the motion refinement based on the bidirectional motion predictor is disabled. In another example, when the time identifier is less than a first threshold and greater than a second threshold, the motion refinement based on the bidirectional motion predictor is disabled.

[0205] In some examples, the specific information includes a quantization parameter of the current block. For example, the quantization parameter of the current block is compared with a threshold to obtain a comparison result. Motion refinement based on the bidirectional motion predictor is disabled based on the comparison result. In an example, when the quantization parameter is greater than the threshold, motion refinement based on the bidirectional motion predictor is disabled; in another example, when the quantization parameter is less than the threshold, motion refinement based on the bidirectional motion predictor is disabled. In another example, when the quantization parameter is less than a first threshold and greater than a second threshold, motion refinement based on the bidirectional motion predictor is disabled.

[0206] In some examples, the specific information includes a time stamp difference between a first time stamp of a forward predictor of the current block and a second time stamp of a backward predictor of the current block. For example, the time stamp difference is compared with a threshold to obtain a comparison result. Motion refinement based on the bidirectional motion predictor is disabled based on the comparison result. In an example, when the time stamp difference is greater than the threshold, the motion refinement based on the bidirectional motion predictor is disabled; in another example, when the time stamp difference is less than the threshold, the motion refinement based on the bidirectional motion predictor is disabled. In another example, when the time stamp difference is less than a first threshold and greater than a second threshold, the motion refinement based on the bidirectional motion predictor is disabled.

[0207] In some examples, the specific information includes a quantization parameter difference between a first quantization parameter of a forward predictor of the current block and a second quantization parameter of a backward predictor of the current block. For example, the quantization parameter difference of the current block is compared with a threshold to obtain a comparison result. Motion refinement based on a bidirectional motion predictor is disabled based on the comparison result. In an example, when the quantization parameter difference is greater than the threshold, motion refinement based on a bidirectional motion predictor is disabled; in another example, when the quantization parameter difference is less than the threshold, motion refinement based on a bidirectional motion predictor is disabled. In another example, when the quantization parameter difference is less than a first threshold and greater than a second threshold, motion refinement based on a bidirectional motion predictor is disabled.

[0208] In some examples, the specific information includes a content difference measure (e.g., a mean square error (MSE) measure) between a forward predictor of the current block and a backward predictor of the current block. For example, the content difference measure is compared with a threshold to obtain a comparison result. Motion refinement based on a bidirectional motion predictor is disabled based on the comparison result. In an example, when the content difference measure is greater than the threshold, motion refinement based on a bidirectional motion predictor is disabled; in another example, when the content difference measure is less than the threshold, motion refinement based on a bidirectional motion predictor is disabled. In another example, when the content difference measure is less than a first threshold and greater than a second threshold, motion refinement based on a bidirectional motion predictor is disabled.

[0209] In some examples, the specific information indicates whether a scene change occurs between a forward predictor of the current block and a backward predictor of the current block. When a scene change occurs, motion refinement based on a bidirectional motion predictor is disabled.

[0210] In some examples, the motion refinement based on the bidirectional motion predictor is a bidirectional optical flow (BDOF), and the specific information includes one or more intermediate BDOF offset values. For example, whether to continue to apply the motion refinement based on the bidirectional motion predictor is determined according to the one or more intermediate BDOF offset values. In an example, the one or more intermediate BDOF offset values ​​include at least one of the following: a gradient value, a cross-correlation value, and an autocorrelation value.

[0211] In some examples, the motion refinement based on the bidirectional motion predictor is a bidirectional optical flow (BDOF), and the specific information includes a BDOF offset calculated for the current block or a sub-block of the current block. The BDOF offset is compared with a threshold to obtain a comparison result. Whether to apply the BDOF offset is determined based on the comparison result.

[0212] Then, the process proceeds to (S1299) and terminates.

[0213] The process (1200) may be adjusted appropriately. The step(s) in the process (1200) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0214] Fig.13 A flow chart outlining a process (1300) according to an embodiment of the present disclosure is shown. The process (1300) may be used in a video encoder. In various embodiments, the process (1300) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), etc. In some embodiments, the process (1300) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1300). The process starts at (S1301) and proceeds to (S1310).

[0215] At (S1310), it is determined that a current block in a current picture satisfies an eligibility condition for motion refinement based on a bi-directional motion predictor.

[0216] At (S1320), an operation of applying motion refinement based on a bi-directional motion predictor on at least a portion of a current block begins.

[0217] At (S1330), specific information used during application of bi-directional motion predictor based motion refinement is obtained.

[0218] At (S1340), it is determined whether to continue to apply the motion refinement based on the bi-directional motion predictor according to the specific information. The motion refinement based on the bi-directional motion predictor includes at least one of DMVR and / or BDOF.

[0219] In some examples, the specific information includes at least one of: a time identifier of the current block; a quantization parameter of the current block; a time identifier difference between a forward predictor of the current block and a backward predictor of the current block; a quantization parameter difference between the forward predictor of the current block and the backward predictor of the current block; a content difference measure between the forward predictor of the current block and the backward predictor of the current block; and a scene change between the forward predictor of the current block and the backward predictor of the current block.

[0220] In some examples, the motion refinement based on the bidirectional motion predictor is bidirectional optical flow (BDOF), and the specific information includes at least one of: a gradient value, a cross-correlation value, an autocorrelation value, and a BDOF offset.

[0221] Then, the processing proceeds to (S1399) and terminates.

[0222] The process (1300) may be adjusted appropriately. The step(s) in the process (1300) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of implementation may be used.

[0223] Some aspects of the present disclosure provide a method for processing visual media data. The method includes processing a code stream of visual media data according to a format rule. The code stream includes encoded information of one or more pictures, and the one or more pictures include a current picture. The format rule specifies that a current block in the current picture is determined to meet a predefined condition for motion refinement based on a bidirectional motion predictor, and the motion refinement based on a bidirectional motion predictor is applied on at least a portion of the current block. The format rule also specifies that a difference metric between a forward predictor and a backward predictor of the current block is obtained during the application of the motion refinement based on the bidirectional motion predictor. The difference metric includes at least one of the following: a time stamp difference between the forward predictor of the current block and the backward predictor of the current block; a quantization parameter difference between the forward predictor of the current block and the backward predictor of the current block; a content difference metric between the forward predictor of the current block and the backward predictor of the current block; and a scene change between the forward predictor of the current block and the backward predictor of the current block. Determine whether to continue to apply the motion refinement based on the bidirectional motion predictor according to the difference metric.

[0224] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.14 A computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0225] Computer software may be encoded using any suitable machine code or computer language, which may be assembled, compiled, linked or similarly constructed to create code comprising instructions that may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0226] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0227] Fig.14 The components shown for the computer system (1400) are exemplary in nature and are not intended to limit the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should not be interpreted as having any dependency or requirement on any component or combination of components shown in the exemplary embodiment of the computer system (1400).

[0228] The computer system (1400) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from one or more human users through tactile input (e.g., keyboard input, sliding, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0229] The human-machine interface input device may include one or more of the following (only one of which is drawn): keyboard (1401), mouse (1402), touchpad (1403), touch screen (1410), data gloves (not shown), joystick (1405), microphone (1406), scanner (1407), camera (1408).

[0230] The computer system (1400) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (1410), a data glove (not shown), or a joystick (1405), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., screens (1410) including cathode ray tube (CRT) screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each of which has or does not have a touch screen input function, each of which has or does not have a tactile feedback function - some of which can output two-dimensional visual output or output of more than three dimensions by means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0231] The computer system (1400) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical disks (CD / DVD ROM / RW) (1420) with CD / DVD or similar media (1421), thumb drives (1422), removable hard disk drives or solid state drives (1423), traditional magnetic media such as tapes and floppy disks (not shown), special-purpose devices based on ROM / Application-Specific Integrated Circuit (ASIC) / Programmable Logic Device (PLD) such as security software protectors (not shown), and the like.

[0232] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0233] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). For example, the network may be wireless, wired, or optical. The network may also be a LAN, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. Examples of networks also include LANs such as Ethernet, wireless LANs, cellular networks (Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc.), television wired or wireless wide-area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle networks, and industrial networks (including controller area network buses (CANBus)), etc. Some networks typically require an external network interface adapter for connecting to some universal data ports or peripheral buses (1449) (e.g., the Universal Serial Bus (USB) port of the computer system (1400)). Other systems are typically integrated into the core of the computer system (1400) by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. The communication can be one-way, for reception only (e.g., wireless television), one-way, for transmission only (e.g., a CAN bus to certain CAN bus devices), or two-way (e.g., to other computer systems via a local or wide area digital network). Each of the above networks and network interfaces can use certain protocols and protocol stacks.

[0234] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core (1440) of the computer system (1400).

[0235] The core (1440) may include one or more CPUs (1441), GPUs (1442), dedicated programmable processing units in the form of Field Programmable Gate Areas (FGPAs) (1443), hardware accelerators for specific tasks (1444), etc., and a graphics adapter (1450). These devices, as well as read-only memory (ROM) (1445), random access memory (1446), internal mass storage (e.g., internal non-user accessible hard disk drives, solid-state drives, etc.) (1447), etc., may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessed in the form of one or more physical plugs so that it can be expanded by additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1448), or connected via a peripheral bus (1449). In one example, a screen (1410) may be connected to a graphics adapter (1450). The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), USB, etc.

[0236] The CPU (1441), GPUs (1442), FPGA (1443) and accelerator (1444) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Transient data can also be stored in RAM (1446), while permanent data can be stored in, for example, internal mass storage (1447). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage (1447), ROM (1445), RAM (1446), etc.

[0237] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purpose of this application, or may be medium and code well known and available to those skilled in the art of computer software.

[0238] As an example and not a limitation, a computer system having an architecture (1400), in particular a core (1440), can be provided as a processor (including CPUs, GPUs, FPGAs, accelerators, etc.) to execute software contained in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the core (1440) having non-volatility, such as a core internal mass storage (1447) or ROM (1445). Software implementing various embodiments of the present application can be stored in such a device and executed by the core (1440). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core (1440), in particular the processor therein (including CPUs, GPUs, FPGAs, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in a random access memory (Random Access Memory, RAM) (1446) and modifying such a data structure according to a software-defined process. Additionally or alternatively, the computer system may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., accelerator (1444)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an IC) storing execution software, circuitry containing execution logic, or both. The present application includes any suitable combination of hardware and software.

[0239] The use of "at least one" or "one" in this disclosure is intended to include any one or combination of the listed elements. For example, references to "at least one of A, B, or C"; "at least one of A, B, and C"; "at least one of A, B, and / or C"; and "at least one of A to C" are intended to include only A, only B, only C, or any combination thereof. References to "one of A or B" and "one of A and B" are intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0240] Although the present application has described a number of exemplary embodiments, various changes, arrangements and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design a variety of systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and therefore belong to the spirit and scope of the present application.

Claims

1. A method for processing visual media data, characterized in that: The method comprises: Process the code stream of visual media data according to the format rules, where: The code stream includes encoded information of one or more pictures, and the one or more pictures include a current picture; and The format rules specify: Determine a current block in the current picture as satisfying a predefined condition for motion refinement based on a bidirectional motion predictor, and apply the motion refinement based on a bidirectional motion predictor to at least a portion of the current block; Obtaining a difference metric between a forward predictor and a backward predictor of the current block during applying the bi-directional motion predictor based motion refinement, wherein the difference metric comprises at least one of the following: a time stamp difference between a forward predictor of the current block and a backward predictor of the current block; a quantization parameter difference between a forward predictor of the current block and a backward predictor of the current block; a content difference measure between a forward predictor of the current block and a backward predictor of the current block; and a scene change between a forward predictor of the current block and a backward predictor of the current block; and Whether to continue applying the bi-directional motion predictor based motion refinement is determined according to the difference metric.

2. A video decoding device, characterized in that: comprising a processing circuit, the processing circuit being configured to: Receiving an encoded video stream including encoded information of one or more pictures; determining, based on the encoded information, that a current block in the current picture satisfies a qualification condition for motion refinement based on a bidirectional motion predictor; Starting to apply the bidirectional motion predictor based motion refinement on at least a portion of the current block; obtaining specific information used during application of said bidirectional motion predictor based motion refinement; as well as Whether to continue applying the motion refinement based on the bidirectional motion predictor is determined according to the specific information.

3. The device according to claim 2, characterized in that The bi-directional motion predictor based motion refinement includes at least one of decoder side motion vector refinement (DMVR) and / or bi-directional optical flow (BDOF).

4. The device according to any one of claims 2 to 3, characterized in that The specific information includes a time mark of the current block, and the processing circuit is configured to: Compare the time mark of the current block with a threshold to obtain a comparison result; and The bi-directional motion predictor based motion refinement is disabled based on the comparison result.

5. The device according to claim 4, characterized in that The processing circuit is configured to perform at least one of the following: When the time stamp is greater than the threshold, disabling the motion refinement based on the bidirectional motion predictor; When the time stamp is less than the threshold, disabling the motion refinement based on the bidirectional motion predictor; or When the time stamp is less than a first threshold and greater than a second threshold, the bi-directional motion predictor based motion refinement is disabled.

6. The device according to any one of claims 2 to 3, characterized in that The specific information includes a quantization parameter of the current block, and the processing circuit is configured to: Comparing the quantization parameter of the current block with a threshold to obtain a comparison result; as well as The bi-directional motion predictor based motion refinement is disabled based on the comparison result.

7. The device according to claim 6, characterized in that The processing circuit is configured to perform at least one of the following: disabling the bidirectional motion predictor based motion refinement when the quantization parameter is greater than the threshold; When the quantization parameter is less than the threshold, disabling the motion refinement based on the bidirectional motion predictor; or When the quantization parameter is less than a first threshold and greater than a second threshold, the bi-directional motion predictor based motion refinement is disabled.

8. The device according to any one of claims 2 to 3, characterized in that The specific information includes a time mark difference between a first time mark of a forward predictor of the current block and a second time mark of a backward predictor of the current block, and the processing circuit is configured to: Comparing the time stamp difference with a threshold value to obtain a comparison result; and The bi-directional motion predictor based motion refinement is disabled based on the comparison result.

9. The device according to claim 8, characterized in that The processing circuit is configured to perform at least one of the following: disabling the motion refinement based on the bidirectional motion predictor when the time stamp difference is greater than the threshold; When the time stamp difference is less than the threshold, disabling the motion refinement based on the bidirectional motion predictor; or When the time stamp difference is less than a first threshold and greater than a second threshold, the bi-directional motion predictor based motion refinement is disabled.

10. The device according to any one of claims 2 to 3, characterized in that The specific information includes a quantization parameter difference between a first quantization parameter of a forward predictor of the current block and a second quantization parameter of a backward predictor of the current block, and the processing circuit is configured to: Compare the quantization parameter difference of the current block with a threshold value to obtain a comparison result; as well as The bi-directional motion predictor based motion refinement is disabled based on the comparison result.

11. The device according to any one of claims 2 to 3, characterized in that The specific information includes a content difference measure between a forward predictor of the current block and a backward predictor of the current block, and the processing circuit is configured to: Comparing the content difference metric with a threshold value to obtain a comparison result; as well as The bi-directional motion predictor based motion refinement is disabled based on the comparison result.

12. The device according to claim 12, characterized in that The processing circuit is configured to perform at least one of the following: disabling the bidirectional motion predictor based motion refinement when the content difference measure is greater than the threshold; disabling the bidirectional motion predictor based motion refinement when the content difference measure is less than the threshold; or When the content difference measure is less than a first threshold and greater than a second threshold, the bi-directional motion predictor based motion refinement is disabled.

13. The device according to any one of claims 2 to 3, characterized in that The specific information indicates whether a scene change occurs between a forward predictor of the current block and a backward predictor of the current block, and the processing circuit is configured to: When the scene change occurs, the bidirectional motion predictor based motion refinement is disabled.

14. The device according to claim 2, characterized in that The motion refinement based on the bidirectional motion predictor is bidirectional optical flow (BDOF), the specific information includes one or more intermediate BDOF offset values, and the processing circuit is configured to: Whether to continue applying the bi-directional motion predictor based motion refinement is determined according to the one or more intermediate BDOF offset values.

15. A video encoding method, characterized in that: include: determining that a current block in a current picture satisfies eligibility conditions for motion refinement based on a bidirectional motion predictor; Starting to apply the bidirectional motion predictor based motion refinement on at least a portion of the current block; obtaining specific information used during application of said bidirectional motion predictor based motion refinement; as well as Determine whether to continue to apply the bidirectional motion predictor-based motion refinement according to the specific information, wherein the bidirectional motion predictor-based motion refinement includes at least one of decoder-side motion vector refinement (DMVR) and / or bidirectional optical flow (BDOF).