Method and device for video processing and medium

By giving weight to the target area of the bidirectional optical flow (BDOF) process, the measurement processing at the sample points is optimized, and the problems of high computational complexity and low encoding and decoding efficiency in the prior art are solved, and more efficient encoding and decoding quality is achieved.

CN120391058APending Publication Date: 2025-07-29DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380087750.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-19
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially in the bidirectional optical flow (BDOF), where the calculation complexity is high and the encoding and decoding quality needs to be improved.

Method used

By assigning multiple weights to the metrics at each sample point in the target area of the bidirectional optical flow (BDOF) process, the gradient or difference of the sample point value is weighted and the measurement values are determined based on multiple reference video blocks of the current video block, thereby optimizing the BDOF process.

Benefits of technology

Improve the quality of encoding and decoding, reduce the computational complexity, and improve the encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120391058A_ABST
    Figure CN120391058A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method includes obtaining a plurality of weights for a conversion between a current video block of the video and a bitstream of the video, where the plurality of weights are used to weight a plurality of values of metrics at respective sample points in a target region of a bidirectional optical flow (BDOF) process applied to the current video block, the metric includes at least one of a gradient or a difference of sample values, and a plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block; and performing a conversion based on the plurality of weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to a bidirectional optical flow (BDOF) process. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the encoding / decoding efficiency of video encoding / decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: obtaining a plurality of weights for the conversion between a current video block of a video and a bitstream of the video, where the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to the current video block, the metric includes at least one of a gradient or a difference of a sample value, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; and performing the conversion based on the plurality of weights.

[0005] According to the method of the first aspect of the present disclosure, the plurality of weights are used to weight the plurality of values of the metric at each sample point in the target region of the BDOF process. Compared with a conventional solution without using weights, with the help of the plurality of weights, the proposed method can advantageously determine one or more parameters for the BDOF process by considering the importance of the sample points in the target region. Thereby, the encoding / decoding quality can be improved.

[0006] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a video processing device for a video. The method includes: obtaining a plurality of weights, where the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of the video, the metric includes at least one of a gradient or a difference of sample values, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; and generating a bitstream based on the plurality of weights.

[0009] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes: obtaining a plurality of weights, where the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of the video, the metric includes at least one of a gradient or a difference of sample values, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; generating a bitstream based on the plurality of weights; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram showing an exemplary video codec system according to some embodiments of the present disclosure is shown;

[0013] Figure 2 A block diagram showing a first exemplary video encoder according to some embodiments of the present disclosure is shown;

[0014] Figure 3 A block diagram showing an exemplary video decoder according to some embodiments of the present disclosure is shown;

[0015] Figure 4 An extended codec unit (CU) region used in BDOF is shown;

[0016] Figure 5 Decoder-side motion vector refinement is shown;

[0017] Figure 6 Shows a diamond-shaped area in the search region;

[0018] Figure 7 Shows the weights generated using an example Gaussian distribution;

[0019] Figure 8 Shows the weights generated using another example Gaussian distribution;

[0020] Figure 9 Shows the weights generated using yet another example Gaussian distribution;

[0021] Figure 10 Shows the weights generated using yet another example Gaussian distribution;

[0022] Figure 11 Shows different filter shapes applied to the data;

[0023] Figure 12 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and

[0024] Figure 13 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.

[0025] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description

[0026] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.

[0027] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.

[0028] As used in the present disclosure, the phrases "one embodiment", "an embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.

[0029] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0030] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Exemplary Environment

[0031] Figure 1 is a block diagram showing an exemplary video codec system 100 that may utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0032] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.

[0033] Video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 over the network 130A. The encoded video data may also be stored on the storage medium / server 130B for access by the destination device 120.

[0034] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.

[0035] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.

[0036] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.

[0037] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.

[0038] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.

[0039] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0040] In addition, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for explanatory purposes, these components are shown separately in Figure 2 the examples.

[0041] The segmentation unit 201 may segment a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0042] The mode selection unit 203 may select, for example, one coding mode among multiple coding modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select an Intra-Inter Joint Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0043] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.

[0044] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block. For example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture.

[0045] In some examples, the motion estimation unit 204 can perform uni-directional prediction on the current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index and a motion vector. The reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0046] Alternatively, in other examples, the motion estimation unit 204 can perform bi-directional prediction on the current video block. The motion estimation unit 204 can search the reference pictures in list 0 to find one reference video block for the current video block, and can also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 can then generate a plurality of reference indices and a plurality of motion vectors. The plurality of reference indices indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 can output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0047] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0048] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0049] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0050] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0051] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0052] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0053] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0054] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0055] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0056] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0057] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce block effect artifacts in the video block.

[0058] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0059] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.

[0060] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.

[0061] In Figure 3 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.

[0062] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-encoded video data, and the motion compensation unit 302 may determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information generally includes a horizontal motion vector displacement value and a vertical motion vector displacement value, one or two reference picture indices, and in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.

[0063] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.

[0064] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of a video block to compute the interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 may use the interpolation filter to generate a prediction block.

[0065] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.

[0066] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0067] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the buffer 307, and the buffer 307 provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on the display device.

[0068] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding, rather than limiting the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to the multi-functional video codec or other specific video codecs, the disclosed technology is also applicable to other video codec technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video / image codec technologies. Specifically, the technology is related to bidirectional optical flow. The technology can be applied to existing video codec standards such as HEVC, VVC, or future next-generation video codec standards such as ECM in VVC exploration. The technology is also applicable to future video codec standards or video codecs. 2. Introduction Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. As of July 2020, the Versatile Video Coding (VVC) standard has also been completed, aiming to further reduce the bitrate by 50% and provide a series of additional functions. After the completion of VVC, activities for future VVC have started. The description of additional tools beyond VVC tools has been summarized as follows: M. Coban, F. Léannec, K. Naser, and J. "Algorithm description of Enhanced Compression Model 5 (ECM5)", Document JVET-Z2025, 26th JVET meeting via teleconference on August 20 - 29, 2022, and its reference SW is named ECM. 2.1 Bidirectional Optical Flow (BDOF) in VVC The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (previously known as BIO) was included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version, which requires much less computation, especially in terms of the number of multiplications and the multiplier size. BDOF is used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if all of the following conditions are met: – The CU is decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other of the two reference pictures is after the current picture in display order. – The distance from the two reference pictures to the current picture (i.e., the POC difference) is the same. – Both of the two reference pictures are short-term reference pictures. – The CU is not decoded using the affine mode or the SbTMVP Merge mode. – The CU has more than 64 luma samples. – The CU height and CU width are both greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current CU. – The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As its name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bi-predicted sample values in the 4x4 sub-block. The following steps are applied during the BDOF process. First, by directly calculating the differences between two neighboring samples, the horizontal and vertical gradients of the two prediction signals, and k = 0, 1, are calculated, i.e., where I (k) (i,j) is the sample value at the prediction signal coordinates (i,j) in the list k, k = 0, 1, and shift1 is calculated based on the luma bit depth bitDepth as shift1 = max(6, bitDepth - 6). Then, the auto-correlations and cross-correlations S1, S2, S3, S5 and S6 of the gradients are calculated as follows: S1 = Σ (i,j)∈Ω Abs(ψ x (i,j)), S3 = ∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) S5 = ∑ (i,j)∈Ω Abs(ψ y (i,j)), S6 = ∑ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) where θ(i,j) = (I (1) (i,j) >> n b ) - (I (0) (i,j) >> n b ) where Ω is a 6×6 window around the 4×4 sub-block, and n a and n bThe values are respectively set to min(1, bitDepth - 11) and min(4, bitDepth - 8). Then, using the cross - correlation terms and auto - correlation terms, the motion refinement (v x , v y ) is derived using the following method: where th′ BIO = 2 max(5,BD-7) . is the floor function, and Based on the motion refinement and gradient, the following adjustment is calculated for each sample point in the 4x4 sub - block: Finally, the BDOF samples of the CU are calculated by adjusting the bi - directional prediction samples in the following manner: pred BDOF (x, y)=(I (0) (x, y)+I (1) (x, y)+b(x, y)+o offset ) >> shift These values are selected so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit - width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient values, some prediction samples I^((k))(i, j) in list k (k = 0, 1) outside the current CU boundary need to be generated. As Figure 4 depicted, the BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, the prediction samples (white positions) in the extended region are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without using interpolation, and the regular 8 - tap motion - compensated interpolation filter is used to generate the prediction samples (gray positions) inside the CU. These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it is partitioned into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries are considered as CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid additional complexity in SAD calculation, the SAD calculated in the DVMR process between the initial L0 prediction samples and the L1 prediction samples is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., the luma_weight_lx_flag of any one of the two reference pictures is 1, the BDOF is also disabled. When the CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.1.1 BDOF in ECM: Sample-based BDOF In sample-based BDOF, the motion refinement (Vx, Vy) is not derived based on blocks, but is performed for each sample. The encoded / decoded block is partitioned into 8×8 sub-blocks. For each sub-block, it is determined whether to apply BDOF by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, for each sample in the sub-block, a sliding 5x5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the central sample of the window. 2.2 Decoder-side Motion Vector Refinement (DMVR) in VVC To improve the accuracy of the Merge mode MV, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bidirectional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. As Figure 5 shown, the SAD between the red blocks for each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, the application of DMVR is restricted and is only applied to CUs encoded or decoded using the following modes and functions: – CU-level Merge mode with bidirectional predicted MVs. – One reference picture is past and the other reference picture is future with respect to the current picture. – The distances from the two reference pictures to the current picture (i.e., POC differences) are the same. – Both reference pictures are short-term reference pictures. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current block. – The CIIP mode is not used for the current block. The refined MVs derived by the DMVR process are used to generate inter-predicted samples and are also used for temporal motion vector prediction for future picture coding. The original MVs are used for the deblocking process and are also used for spatial motion vector prediction for future CU coding. Additional features of DMVR are mentioned in the following sub-articles. In DVMR, the search points are around the initial MV, and the MV offset obeys the MV difference mirroring rule. In other words, any point examined by DMVR represented by the candidate MV pair (MV0, MV1) follows the following two equations. MV0′ = MV0 + MV_offset MV1′ = MV1 - MV_offset where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample phase of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search phase. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV in the DMVR process. The SAD between the reference blocks pointed to by the initial MV candidate reference reduces the SAD value by 1 / 4. Integer sample point search is followed by fractional sample point refinement. To save computational complexity, fractional sample point refinement is derived using the parametric error surface equation instead of using SAD comparison for additional search. Fractional sample point refinement is conditionally invoked based on the output of the integer sample point search stage. When the integer sample point search stage ends at the center with the minimum SAD in the first iteration or the second iteration search, fractional sample point refinement is further applied. In sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form E(x, y) = A(x - x min ) 2 + B(y - y min ) 2 + C where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) y min = (E(0, -1) - E(0,1)) / (2((E(0, -1) + E(0,1) - 2E(0,0))) The values of x min and y min are automatically constrained between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined incremental MV. In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate the samples at fractional positions. In DMVR, the search points are centered around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, within a 2-sample search range, DMVR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples than the normal MC process, the samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV will be filled with samples from those available samples. When the width and / or height of a CU is greater than 16 luma samples, it is further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16. 2.3.Multi-pass Decoder-side Motion Vector Refinement (ECM) Multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coded / decoded block. In the second pass, BM is applied to each 16x16 sub-block within the coded / decoded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction. 2.3.1 The First Pass - Block-based Bilateral Matching MV Refinement In the first pass, the refined MV is derived by applying BM to the coded / decoded block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive the integer sample accuracy intDeltaMV. The local search applies a 3×3 square search pattern to loop within the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the mean removed SAD (MRSAD) cost function is applied to remove the DC effect of distortion between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search terminates. Otherwise, the current current minimum-cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. The current fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as: · MV0_pass1 = MV0 + deltaMV, · MV1_pass1 = MV1 – deltaMV. 2.3.2 Second Pass – Sub-block Based Bilateral Matching MV Refinement In the second pass, the refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, in the reference picture lists L0 and L1, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained through the first pass. The refined MV (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) is derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM performs a full search to derive the integer sample accuracy intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub-blocks, as: bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into no more than 5 diamond search areas, as Figure 6As shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond-shaped area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order starting from the upper left corner of the area to the lower right corner. When the minimum bilCost within the current search area is less than a threshold equal to sbW * sbH, the full pixel search is terminated; otherwise, the full pixel search continues to the next search area until all search points have been examined. Additionally, if the difference between the previous minimum cost and the current minimum cost in an iteration is less than a threshold equal to the area of the block, the search process is terminated. The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). Then, the refined MV for the second pass is derived as: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2) · MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2). 2.3.3 Third Pass – Sub-Block Based Bidirectional Optical Flow MVV Refinement In the third pass, the refined MV is derived by applying BDOF to an 8×8 grid sub-block. For each 8×8 sub-block, starting from the refined MV of the parent sub-block in the second pass, BDOF refinement is applied to derive the scaled Vx and Vy without clipping. The derived bioMv (Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32. The refined MV for the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as: · MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv, · MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv. In all of the foregoing sub-entries, when loop motion compensation is enabled, the motion vectors should be clipped to the loop offset being considered. 2.3.4 Adaptive Decoder-Side Motion Vector Refinement The adaptive decoder-side motion vector refinement method is an extension of multi-pass DMVR consisting of two new Merge modes to refine the MV for Merge candidates that meet the DMVR conditions only in one direction (L0 or L1) of bi-prediction. The multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vectors. However, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is set to zero. Merge candidates for the new Merge modes are derived from spatial neighboring coded / decoded blocks, TMVP, non-adjacent blocks, HMVP, paired candidates, similar to the conventional Merge mode. The difference is that only those Merge candidates that meet the DMVR conditions are added to the candidate list. The two new Merge modes use the same Merge candidate list. If the list of BM candidates includes inherited BCW weights and the DMVR process remains unchanged, if the weights are not equal and bi-prediction is weighted using BCW weights, MRSAD or MRSATD is used for distortion calculation. The Merge index is coded / decoded in the conventional Merge mode. 3. Problems There are several parts in the BDOF MV refinement / sample adjustment that can be improved. - The current formula used to derive the BDOF parameters is not an exact formula. - There is no weight to indicate the importance of each sample in the final formula. - There is no filtering process to smooth the finally derived MV refinement / sample adjustment. - There is no obvious difference in the conditions for applying BDOF for MV refinement / sample adjustment. Similarly, there is no difference in their formulas. 4. Detailed Solutions The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow way. In addition, these solutions can be combined in any way. The methods disclosed below can be applied to bi-directional optical flow, decoder-side motion vector refinement, and any of their extensions. Regarding the Derivation of BDOF MV Refinement Parameters In the following subsections, the general equations for deriving the BDOF parameters (vx and vy) are defined as: ∑Gx.Gx*vx + ∑Gx.Gy*vy = ∑dI.Gx. → s1*vx + s2*vy = s3, ∑Gx.Gy*vx + ∑Gy.Gy*vy = ∑dI.Gy → s2*vx + s5*vy = s6, where Gx and Gy respectively represent the sum of the horizontal gradient and the vertical gradient for two reference pictures. dI represents the difference between the two reference pictures. The summation (∑) is within a predefined region, which can be an N×M block around the current sample (for sample - adjusted BDOF) or around the current predictor block (for MV - refined BDOF). 1. It is proposed that different methods of deriving gradients for BDOF in VVC can be used to calculate the horizontal gradient and / or the vertical gradient. a. In one example, the gradient is calculated by directly computing the difference between two neighboring samples, i.e., b. In another example, the gradient is calculated by computing the difference between two shifted neighboring samples, i.e., i. shift1 and shift2 can be any integer such as 0, 1, 2, 6, etc. or can even be negative integers. c. In another example, the gradient can be calculated as a weighted sum using Nb samples before the current sample and Na samples after the current sample: i. The weights (i.e., wp) can be any integer such as -6, 0, 2, 7, etc. or any real number such as -6.3, -0.77, 0.1, 3.0, etc. ii. The weights used to calculate the horizontal gradient and the vertical gradient can be different from each other. (i) Alternatively, the weights used to calculate the horizontal gradient and the vertical gradient can be the same. iii. The weights can be signaled from the encoder to the decoder. iv. The weights can be derived using the decoded information. v. Nb and Na can be any integer, e.g., 0, 3, 10, etc. vi. Nb and Na can be different for calculating the gradients in the horizontal direction and the vertical direction. (i) Alternatively, Nb and Na can be the same for calculating the gradients in the horizontal direction and the vertical direction. 2. It is proposed that a complete linear equation formula can be used to derive the final MV refinement. a. In one example, after calculating all the gradients, s1, s2, s3, s5, and s6 are calculated as described above: ∑Gx.Gx*vx + ∑Gx.Gy*vy = ∑dI.Gx. → s1*vx + s2*vy = s3, ∑Gx.Gy*vx + ∑Gy.Gy*vy = ∑dI.Gy → s2*vx + s5*vy = s6. i. In one example, to derive the final MV of an M*N block, samples in the (M + K1)*(N + K2) region around the original block can be involved. For example, K1 and K2 can be any integers, such as 0, 2, 4, 7, 10, etc. b. In one example, after calculating all s1, s2, s3, s5, and s6, the determinant values D, Dx, and Dy are calculated as: D = (s1 >> shTem)*(s5 >> shTem) - (s2 >> shTem)*(s2 >> shTem), Dx = (s3 >> shTem)*(s5 >> shTem) - (s6 >> shTem)*(s2 >> shTem), Dy = (s1 >> shTem)*(s6 >> shTem) - (s3 >> shTem)*(s2 >> shTem). i. In one example, shTem can be any integer, such as 0, 1, 3, etc. c. In one example, after calculating D, Dx, and Dy, vx and vy can be derived as: vx = Dx / D and vy = Dy / D. i. In another example, if abs(D) is less than a predefined threshold, C, vx, and vy are set to zero. C can be a non - negative number, such as 0, 10, 17, etc. d. In one example, any number of shifts and clippings can be involved to derive the final vx and vy. i. In one example, the numerator and / or denominator can have additional shifts, which are shifted left by K overall, such that the finally derived vx and vy have higher precision. K can be any integer, such as 0, 1, 3, 4, 6, etc. ii. In one example, these shifts can occur in any order, such as having a shift at the beginning, and / or having a shift for intermediate variables and / or having a shift on the final MV. iii. In one example, the final vx and vy can be clipped to be between - B and B, where B can be any integer, such as 2, 10, 17, 32, 100, 156, 725, etc. e. In one example, the final vx and vy can be multiplied (or similarly divided) by a number before being used in the motion compensation process. i. In one example, vx and vy can be multiplied by R, where R is any real number, such as 1.25, 2, 3.1, 4, etc. ii. In another example, vx and vy can be divided by R, where R is any real number, such as 1.25, 2, 3.1, 4, etc. iii. In one example, the numerical values that multiply (or divide) the final vx and vy can be different for vx and vy. iv. In one example, the numerical values that multiply (or divide) the final vx and vy can depend on the block size, sequence resolution, block characteristics, etc. 3. It is proposed that the solution of the partial linear equation can be used to derive the final MV refinement. a. In one example, after calculating all the gradients, s1, s2, s3, s5, and s6 are calculated as described above: ∑Gx.Gx*vx + ∑Gx.Gy*vy = ∑dI.Gx. → s1*vx + s2*vy = s3, ∑Gx.Gy*vx + ∑Gy.Gy*vy = ∑dI.Gy → s2*vx + s5*vy = s6. b. In one example, after calculating all s1, s2, s3, s5, and s6, the approximate versions of vx and vy can be calculated as: vx = s3 / s1, vy = (s6 - s2*vx) / s5. c. In another example, after calculating vx similar to the above, a partial amount of vx can be placed in the second formula to derive vy. i. In one example, vy can be derived as vy = (s6 - s2*vx / T) / s5, where T can be any real number, such as 1.1, 2, 4, etc. d. In one example, after calculating all s1, s2, s3, s5, and s6, the approximate versions of vx and vy can be calculated as: Assume vx is zero: vy = s6 / s5, Substitute vy into the first formula: vx = (s3 - s2*vy) / s1. e. In another example, after calculating vy similar to the above, a partial amount of vy can be substituted into the second formula to derive vx. i. In one example, vx can be derived as vx = (s3 - s2*vy / T) / s1, where T can be any real number, such as 1.1, 2, 4, etc. f. In one example, after calculating all of s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as: Assume vy is zero: vx = s3 / s1. Assume vx is zero: vy = s6 / s5. 4. It is proposed that a simplified solution can be used to derive the final MV refinement. a. In one example, the method explained in the background subsection for VVC BDOF can be used to derive approximate versions of s1, s2, s3, s5, and s6. b. In one example, after calculating approximate versions of s1, s2, s3, s5, and s6, the values of the determinant D, Dx, and Dy are calculated as: D = (s1 >> shTem) * (s5 >> shTem) – (s2 >> shTem) * (s2 >> shTem), Dx = (s3 >> shTem) * (s5 >> shTem) – (s6 >> shTem) * (s2 >> shTem), Dy = (s1 >> shTem) * (s6 >> shTem) – (s3 >> shTem) * (s2 >> shTem). i. In one example, after calculating D, Dx, and Dy, vx and vy can be derived as: vx = Dx / D, and vy = Dy / D. ii. In one example, shTem can be any integer, such as 0, 1, 3, etc. c. In one example, after calculating approximate versions of s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as: Assume vy is zero: vx = s3 / s1, Substitute vx into the second formula: vy = (s6 - s2 * vx) / s5. i. Alternatively, the modified vx can be substituted into the second formula: vy = (s6 – s2 * vx / T) / s5, where T can be any real number, such as 1.1, 2, 4, etc. ii. Alternatively, it can be assumed that the first vx is zero, and vy can be derived. After that, vy or a scaled version of vy can be substituted into the first equation and vx can be derived. 5. It is proposed that any combination of the above methods can be used to derive the final MV refinement. a. In one example, any combination of the above methods (2, 3, and 4) can be combined and used together. Regarding the Derivation of BDOF Sample Adjustment Parameters 6. It is proposed that any of the methods described above for BDOF MV refinement can also be used for BDOF sample adjustment parameter derivation. a. In one example, after calculating all gradients, s1, s2, s3, s5, and s6 are calculated as described above: ∑Gx.Gx*vx + ∑Gx.Gy*vy = ∑dI.Gx. → s1*vx + s2*vy = s3, ∑Gx.Gy*vx + ∑Gy.Gy*vy = ∑dI.Gy → s2*vx + s5*vy = s6. i. In one example, the samples in the K×K region around the sample point can be involved in the derivation. K can be any integer, such as 1, 3, 4, 5, 7, 10, etc. b. In one example, after calculating s1, s2, s3, s5, and s6, the determinant values D, Dx, and Dy are calculated as: D = (s1 >> shTem)*(s5 >> shTem) - (s2 >> shTem)*(s2 >> shTem), Dx = (s3 >> shTem)*(s5 >> shTem) - (s6 >> shTem)*(s2 >> shTem), Dy = (s1 >> shTem)*(s6 >> shTem) - (s3 >> shTem)*(s2 >> shTem). i. shTem can be any integer, such as 0, 1, 3, etc. ii. In one example, after calculating D, Dx, and Dy, vx and vy can be derived as: vx = Dx / D, and vy = Dy / D. iii. In another example, if abs(D) is less than a predefined threshold, C, vx, and vy are set to zero. C can be a non - negative number, such as 0, 10, 17, etc. c. In one example, after calculating s1, s2, s3, s5, and s6, an approximate version of vx and vy can be calculated as: vx = s3 / s1, vy = (s6 - s2*vx) / s5. i. Or alternatively, the modified vx can be placed in the second formula: vy = (s6 – s2 * vx / T) / s5, where T can be any real number, such as 1.1, 2, 4, etc. ii. Alternatively, it can be assumed that the first vx is zero, and vy can be derived. After that, vy or a scaled version of vy can be substituted into the first equation and vx can be derived. d. In one example, the method explained in the background subsection for VVC BDOF can be used to derive approximate versions of s1, s2, s3, s5, and s6. e. In one example, the final vx and vy can be multiplied (or divided or shifted) by a number before being used in the sample point adjustment process. i. In one example, vx and vy can be multiplied by R, where R is any real number, such as 1.25, 2, 3.1, 4, etc. ii. In another example, vx and vy can be divided by R, where R is any real number, such as 1.25, 2, 3.1, 4, etc. iii. In one example, the value of the number by which the final vx and vy are multiplied (or divided) can be different for vx and vy can be different. iv. In one example, the value of the number by which the final vx and vy are multiplied (or divided) can depend on the block size, sequence resolution, block characteristics, position in the block, etc. Regarding the Application of Weights in Parameter Derivation 7. It is proposed that any weight can be applied before adding the BDOF intermediate parameters for MV refinement. a. In one example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, within the target region of Ω (the M_ext * N_ext region around the current block), all values are added with a similar weight (equal to 1). b. In another example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, within the target region of Ω (the M_ext * N_ext region around the current block), depending on the position of the value in the extended block (the target region of Ω), the value is added after being multiplied by a predefined weight. c. In one example, these predefined weights are defined as: w = (x >= (width / 2)? width - x : x + 1) * (y >= (height / 2)? height - y : y + 1) for x from 0 to width - 1 and y from 0 to height - 1. Width and height represent the width of the target region and the height of the target region. d. In another example, these predefined weights can be generated using some known probability distributions such as Gaussian distribution, which has any value of standard deviation (σ = 1, 1.5, 4 or any other real number) and a central position. i. In one example, these weights are generated for a 12×12 region using a Gaussian distribution with σ = 2.5, as Figure 7 shown. ii. In one example, these weights are generated for a 12×12 region using a Gaussian distribution with σ = 4, as Figure 8 shown. e. In another example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, within the target region of Ω (the M_ext*N_ext region surrounding the current block), depending on the position of the value within the extended block (the target region of Ω), the value is added after being shifted by a predefined value. f. In one example, the weight matrix can be represented as a left (or right) shift matrix, and depending on the matrix entries, the data is shifted (left or right) before summation. 8. In one example, different weights can be applied depending on the block size, block shape, block characteristics, sequence resolution, etc. i. Alternatively, depending on the block size, block shape, block characteristics, sequence resolution, etc., weights may not be applied. ii. The weight matrix can be explicitly coded and decoded in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or Slice Header (SH). 9. It is proposed that any weights can be applied before adding the BDOF intermediate parameters for sample point adjustment. a. In one example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, within the target region of Ω (the K1*K2 region surrounding the current sample point), all values are added with similar weights (equal to 1). K1 and K2 can be any integers, such as 1, 2, 3, 5, 8, etc. b. In another example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, within the target region of Ω (the K1*K2 region surrounding the current sample point), depending on the position of the value within the extended block (the target region of Ω), the value is added after being multiplied by a predefined weight. c. In one example, these predefined weights are defined as: w = (x >= (K1 / 2)? K1 - x : x + 1) * (y >= (K2 / 2)? K2 - y : y + 1), for x from 0 to K1 - 1 and y from 0 to K2 - 1. K1 and K2 represent the width of the target region and the height of the target region. d. In another example, these predefined weights can be generated using some known probability distributions such as Gaussian distribution, which has any value of standard deviation (σ = 1, 1.5, 2, 4 or any other real number) and any center position. i. In one example, these weights are generated using a Gaussian distribution with σ = 1 for a 5×5 region, as Figure 9 shown. ii. In one example, these weights are generated using a Gaussian distribution with σ = 2 for a 5×5 region, as Figure 10 shown. e. In one example, the weight matrix can be represented as a left (or right) shift matrix, and depending on the matrix entries, the data is shifted (left or right) before summation. f. In one example, different weights can be applied depending on the block size, block shape, block characteristics, sequence resolution, etc. i. Alternatively, depending on the block size, block shape, block characteristics, sequence resolution, etc., no weights may be applied. Regarding the Application of Filters on Final MV Refinement or Sample Adjustment 10. It is proposed that any type of filter can be applied to the finally derived MV refinements (vx and vy). Some examples are shown in Figure 11 as follows. a. In one example, any smoothing filter of any shape can be applied to all MVs derived by BDOF for each sub-block. b. In one example, all MVs within the PU can be used during filter application. c. In another example, during filter application, only the MVs with similar DMVR MVs in the second round can be used for those MVs. d. In one example, a shape filter with any weights can be applied to the MVs. i. In one instance, the weight for the center can be 8, and the weights for the four sides can be 1. ii. In one instance, the weight for the center can be 4, and the weights for the four sides can be 1. iii. In one instance, the weight for the center can be 4, and the weights for the four sides can be 2. iv. In one instance, the weight for the center can be 4, and the weights for the four sides can be 3. v. In one instance, the weight for the center can be 1, and the weights for the four sides can be 1. 11. It is proposed that any type of filter can be applied to the final derived BDOF sample MV adjustment or the final sample adjustment. Some examples are shown in Figure 11 as follows. a. In one example, the filter is applied to all (vx, vy) inside the sub-block or the final adjustment. b. In one example, a shape filter with any weight can be applied to (vx, vy) or the final adjustment. i. In one instance, the weight for the center can be 8, and the weight for the 4 sides can be 1. ii. In one instance, the weight for the center can be 4, and the weight for the 4 sides can be 1. iii. In one instance, the weight for the center can be 4, and the weight for the 4 sides can be 2. iv. In one instance, the weight for the center can be 4, and the weight for the 4 sides can be 3. v. In one instance, the weight for the center can be 1, and the weight for the 4 sides can be 1. Regarding the Conditions for Applying BDOF 12. It is proposed that there may be conditions for applying BDOF MV refinement or BDOF sample adjustment. a. In one example, the conditions for applying BDOF MV refinement can be similar to the conditions for applying BDOF sample adjustment. b. In another example, the conditions for applying BDOF MV refinement can be different from the conditions for applying BDOF sample adjustment. For example, BDOF MV refinement can be applied to CUs encoded / decoded using bi-prediction with unequal weights, while BDOF sample adjustment can only be applied to CUs encoded / decoded using bi-prediction with equal weights. 13. It is proposed that the cost of evaluating BDOF conditions can depend on the cost between 2 reference picture blocks. a. In one example, different cost functions can be used to derive the cost. i. In one example, the cost can be the sum of absolute differences (SAD) between 2 reference picture blocks. ii. In one example, the cost can be the sum of absolute transform differences (SATD) or any other cost metric between 2 reference picture blocks. iii. In one example, the cost can be the sum of absolute differences based on mean removal (MR-SAD) between 2 reference picture blocks. iv. In one example, the cost can be the sum of SAD / MR-SAD between 2 reference picture blocks Weighted average of SATD. v. In one example, the cost function between two reference picture blocks can be: (i) Sum of absolute differences (SAD) / Mean removed SAD (MR-SAD); (ii) Sum of absolute transform differences (SATD) / Mean removed SATD (MR- SATD); (iii) Sum of squared differences (SSD) / Mean removed SSD (MR-SSD); (iv) SSE / MR-SSE; (v) Weighted SAD / Weighted MR-SAD; (vi) Weighted SATD / Weighted MR-SATD; (vii) Weighted SSD / Weighted MR-SSD; (viii) Weighted SSE / Weighted MR-SSE; (ix) Gradient information. Regarding the BDOF MV Refinement Sub - block Size 14. It is proposed that any sub-block size can be used as the BDOF MV refinement sub-block size depending on the conditions. a. In one example, the sub-block size can be a fixed size, such as NxM, where N and M can be any positive integers, such as 1, 2, 3, 4, 5, 8, 12, 32, etc. b. In another example, the sub-block size can depend on the current PU size or CU size. As an example, for a block size of WxH, a sub-block size of W1xH1 can be used, where W1 and H1 depend on W and H, and W1 and H1 can be any positive integers. c. In one example, the sub-block size can depend on the reference picture characteristics. i. In one example, the sub-block size can be determined by the similarity of two prediction values from two reference pictures. If the two prediction values are similar, such as the SAD between the two prediction values is small, a large sub-block size can be applied; otherwise, a small sub-block size can be applied. ii. In one example, the sub-block size can be determined by the distribution of the differences between two prediction values. Those sub-blocks with difference energy (such as SAD or SSE) can be merged into larger units for MV refinement to reduce the computational complexity. d. In one example, the sub-block size can depend on the temporal gradient of two reference blocks. i. In one example, any cost function (such as SAD) can be used to calculate the gradients (or differences) of two reference blocks. e. In one example, the spatial domain gradients of the reference blocks can be used to determine the sub-block size. f. In one example, the sub-block size can depend on the prediction type. g. In one example, the sub-block size can depend on the DMVR first and / or second stage adjustment values. h. In one example, the sub-block size can depend on the sequence resolution. i. In one example, the sub-block size can be a function of all or some of the above parameters. General Aspects 15. In one example, the division operations disclosed in the document can be replaced by non-division operations, which can share the same or similar logic with the division replacement logic in CCLM or CCCM. 16. Whether to apply the above method and / or how to apply the above method can depend on the transcoded information. a. In one example, the transcoded information can include block size and / or temporal layer, and / or slice type / picture type, color component, etc. 17. Whether to apply the above method and / or how to apply the above method can be indicated in the bitstream. a. The indication of enabling / disabling or the method to be applied can be signaled at the sequence level / group of pictures level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. b. The indication of enabling / disabling or the method to be applied can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / piece / sub-picture / other types of regions including more than one sample or pixel.

[0069] More details of embodiments of the present disclosure related to the bidirectional optical flow (BDOF) process will be described below. The embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be construed in a narrow sense. In addition, these embodiments can be applied alone or combined in any way.

[0070] As used herein, the term "block" may refer to a color component, sub-picture, picture, stripe, slice, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), sub-blocks of a video block, sub-regions within a video block, a video processing unit including a plurality of samples / pixels, etc. A block may be rectangular or non-rectangular.

[0071] Figure 12 FIG. 1200 shows a flow chart of a method 1200 for video processing according to some embodiments of the present disclosure. Method 1200 may be implemented during the conversion between a current video block of a video and the bitstream of the video. As Figure 12 shown, method 1200 starts at 1202, where a plurality of weights are obtained. The plurality of weights are used to weight a plurality of values of a metric at each sample in a target region to which a BDOF process applied to the current video block is applied. The metric includes at least one of a gradient or a difference of sample values. Further, the plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block.

[0072] At 1204, a transformation is performed based on the plurality of weights. By way of example and not limitation, a set of parameters for the BDOF process may be determined based on the plurality of weights. Further, at least one offset may be determined based on the set of parameters.

[0073] In some embodiments, at least one offset may be used to refine the motion vector (MV) of the current video block. In this case, at least one offset may also be referred to as MV refinement. If the size of the current video block is M×N, the target region may include a region of size (M+K1)×(N+K2) around the current video block. Each of M, N, K1, and K2 may be an integer, such as 0, 2, 4, 7, 10, etc.

[0074] In some alternative embodiments, at least one offset may be used to adjust a current sample in the current video block. In this case, the BDOF process is also referred to as sample-based BDOF. The target region may include a region of size K3×K4 around the current sample. Each of K3 and K4 may be an integer, such as 1, 3, 4, 5, 7, 10, etc.

[0075] Further, the transformation may be performed based on at least one offset. In some embodiments, the transformation may include encoding the current video block into the bitstream. Alternatively or additionally, the transformation may include decoding the current video block from the bitstream. It should be understood that the above description and / or examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0076] In view of the above, multiple weights are used to weight multiple values of the metrics at each sample point in the target region of the BDOF process. Compared with the conventional solution without using weights, with the help of multiple weights, the proposed method can advantageously determine one or more parameters for the BDOF process by considering the importance of the sample points in the target region. Thereby, the coding and decoding quality can be improved.

[0077] In some embodiments, a set of parameters for the BDOF process can be determined based on multiple weights and a complete linear equation formula. Alternatively, a set of parameters can be determined based on multiple weights and a partial linear equation solution or a simplified solution. This will be described in detail below.

[0078] In some embodiments, a set of parameters can include a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter. Each of these parameters can be determined in a predetermined manner.

[0079] In some embodiments, the gradient can include at least one of a horizontal gradient or a vertical gradient. A set of parameters can be determined based on the following: s1 = ∑(Gx · Gx), s2 = ∑(Gx · Gy), s3 = ∑(dI · Gx), s5 = ∑(Gy · Gx), s6 = ∑(dI · Gy), where s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, Gx represents the sum of the values of the horizontal gradient determined for each reference video block among multiple reference video blocks, Gy represents the sum of the values of the vertical gradient determined for each reference video block among multiple reference video blocks, dI represents the difference in sample values between multiple reference video blocks, and ∑() represents the weighted sum within the target region based on multiple weights. In this case, these five parameters s1, s2, s3, s5, and s6 are determined based on the accurate formula for BDOF. Thereby, the coding and decoding quality can be advantageously improved.

[0080] Alternatively, a set of parameters can be determined based on the following: s1 = ∑ (i,j)∈n Abs(ψ x (i, j)), s2 = ∑ (i,j)∈Ω ψ x (i, j) · Sign(ψ y (i, j)), s3 = ∑ (i,j)∈Ω θ(i, j) · Sign(ψ x(i, j)), s4 = ∑ (i,j)eΩ Abs(ψ y (i, j)), s5 = ∑ (i,j)∈Ω θ(i, j)·Sign(ψ y (i, j)), θ(i, j) = (I (1) (i, j) >> nb) - (I (0) (i, j) >> nb), where s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, I (k) (i, j) represents the sample value at the coordinate (i, j) of the predicted signal of the reference video block of the current video block in list k, and k can be equal to 0 or 1, represents the value for the horizontal gradient at the coordinate (i, j) of the predicted signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at the coordinate (i, j) of the predicted signal of the reference video block of the current video block in list k, ∑ (i,j)∈Ω () represents the weighted sum within the target region Ω based on multiple weights, abs(z) represents the absolute value of the number z, sign(z) represents the sign of the number z, and each item in na and nb can be an integer, such as 0, 2, 6, etc.

[0081] In some embodiments, a set of determinant values is determined based on the set of parameters; in addition, at least one offset is determined based on the set of determinant values. By way of example and not limitation, the set of determinant values can be determined based on the following: D = (sl >> shTem) * (s5 >> shTem) - (s2 >> shTem) * (s2 >> shTem), Dx = (s3 >> shTem) * (s5 >> shTem) - (s6 >> shTem) * (s2 >> shTem), Dy = (s1 >> shTem) * (s6 >> shTem) - (s3 >> shTem) * (s2 >> shTem), where D represents the first determinant value in the set of determinant values, Dx represents the second determinant value in the set of determinant values, Dy represents the third determinant value in the set of determinant values, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and shTem can be an integer, 0, 1, 3, etc.

[0082] In some embodiments, at least one offset may be determined based on the following: vx = Dx / D, vy = Dy / D, where vx represents a first offset among the at least one offset, and vy represents a second offset among the at least one offset. In this case, the determined at least one offset is more accurate than conventional solutions. Thereby, the coding and decoding quality can be advantageously improved.

[0083] In some embodiments, if the absolute value of the first determinant value D is less than a predetermined threshold, each offset among the at least one offset (e.g., vx and vy) may be equal to a predetermined value, such as 0. For example, the predetermined threshold may be a non - negative number, such as 0, 10, 17, etc.

[0084] In some embodiments, the first offset among the at least one offset may be determined based on a set of parameters. Additionally, the second offset among the at least one offset may be determined based on the first offset and a set of parameters.

[0085] In one example, at least one offset may be determined based on the following: vx = s3 / s1, vy = (s6 - s2 * vx) / s5, where vx represents the first offset, vy represents the second offset, s1 represents a first parameter, s2 represents a second parameter, s3 represents a third parameter, s5 represents a fourth parameter, and s6 represents a fifth parameter. In this case, the at least one offset can be determined with a lower computational complexity. Thereby, the coding and decoding efficiency can be advantageously improved.

[0086] In another example, the first offset and the second offset may be determined based on the following: vx = s3 / s1, vy = (s6 - s2 * vx / T1) / s5, where vx represents the first offset, vy represents the second offset, s1 represents a first parameter, s2 represents a second parameter, s3 represents a third parameter, s5 represents a fourth parameter, s6 represents a fifth parameter, and T1 represents a real number, such as 1.1, 2, 4, etc. In this case, the at least one offset can be determined with a lower computational complexity. Thereby, the coding and decoding efficiency can be advantageously improved.

[0087] In some alternative embodiments, the second offset may be determined based on a set of parameters. Additionally, the first offset may be determined based on the second offset and a set of parameters.

[0088] In one example, the first offset and the second offset may be determined based on the following: vy = s6 / s5, vx = (s3 - s2 * vy) / s1, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter. In this case, at least one offset can be determined with a lower computational complexity. Thus, the encoding / decoding efficiency can be advantageously improved.

[0089] In another example, the first offset and the second offset can be determined based on the following: vy = s6 / s5, vx = (s3 - s2 * vy / T2) / s1, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and T2 represents a real number, such as 1.1, 2, 4, etc. In this case, at least one offset can be determined with a lower computational complexity. Thus, the encoding / decoding efficiency can be advantageously improved.

[0090] In some alternative embodiments, at least one offset can be determined based on the following: vx = s3 / s1, vy = s6 / s5, where vx represents the first offset among at least one offset, vy represents the second offset among at least one offset, s1 represents the first parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter. In this case, the computational complexity of at least one offset can be further reduced. Thus, the encoding / decoding efficiency can be advantageously further improved.

[0091] In some embodiments, a shift operation or a clipping operation can be used to determine at least one offset. For example, a shift operation can be applied to at least one of the numerator or denominator of a term in an equation for determining at least one offset. By way of example and not limitation, a right shift operation can be applied to the numerator of the term Dx / D, and a left shift operation can be applied to the numerator of the term Dy / D.

[0092] In some embodiments, the shift operation can include shifting left a predetermined number of bits, such as 0, 1, 3, 4, 6, etc. In some embodiments, the shift operation can be applied in a predetermined order, such as applying the shift operation at the beginning and / or applying the shift operation to an intermediate variable and / or applying the shift operation to the final output.

[0093] In some embodiments, a clipping operation may be applied to at least one offset to update the at least one offset. For example, the value of the updated at least one offset may be between a first value and a second value, such as an upper limit and a lower limit.

[0094] In some embodiments, at least one offset may be adjusted using at least one scaling factor, and a transformation may be performed based on the adjusted at least one offset. In one example, a first offset among the at least one offset may be adjusted by multiplying the first offset by a first scaling factor among the at least one scaling factor, and the first scaling factor may be a real number, such as 1.25, 2, 3.1, 4, etc. In another example, a second offset among the at least one offset may be adjusted by dividing the second offset by a second scaling factor among the at least one scaling factor, and the second scaling factor may be a real number, such as 1.25, 2, 3.1, 4, etc.

[0095] In some embodiments, different offsets among the at least one offset may be adjusted using different scaling factors among the at least one scaling factor. For example, the at least one scaling factor may depend on a block size, a sequence resolution, a block characteristic, etc.

[0096] In some embodiments, each weight among a plurality of weights may be equal to the same predetermined value. Alternatively, each weight among the plurality of weights may depend on the position of a corresponding sample point in a target region. For example, a first weight among the plurality of weights corresponding to a first sample point in the target region may depend on the position of the first sample point in the target region.

[0097] In some embodiments, a first weight may be determined based on the following: w1 = (x >= (wt / 2)? wt - x : x + 1) * (y >= (ht / 2)? ht - y : y + 1), where w1 represents the first weight, x represents the horizontal position of the first sample point in the target region, y represents the vertical position of the first sample point in the target region, wt represents the width of the target region, and ht represents the height of the target region. The logical operator (a? b : c) is defined as: if a is true, the output evaluates to the value of b; otherwise, the output evaluates to the value of c.

[0098] In some embodiments, the plurality of weights may be determined based on a predetermined probability distribution. By way of example and not limitation, the predetermined probability distribution may include a Gaussian distribution with a predetermined standard deviation. Figure 7 Weights generated using a Gaussian distribution with a standard deviation of 2.5 for a 12×12 region are shown. Figure 8 Weights generated using a Gaussian distribution with a standard deviation of 4 for a 12×12 region are shown. Figure 9Shows the weights generated using a Gaussian distribution with a standard deviation of 1 for a 5×5 region. Figure 10 Shows the weights generated using a Gaussian distribution with a standard deviation of 2 for a 5×5 region. It should be understood that the above examples are described only for the purpose of illustration. The scope of the present disclosure is not limited in this regard.

[0099] In some embodiments, multiple weights can be implemented using shift operations. For example, multiple weights can be represented as a left shift matrix or a right shift matrix.

[0100] In some embodiments, multiple weights can depend on the block size, block shape, block characteristics, and / or sequence resolution. In other embodiments, multiple weights can be indicated in a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header (SH).

[0101] In some embodiments, whether to apply the BDOF process for MV refinement to the current video block can depend on a first condition. Additionally or alternatively, whether to apply the BDOF process for sample adjustment to the current video block can depend on a second condition. In one example, the first condition can be the same as the second condition. In another example, the first condition can be different from the second condition. By way of example and not limitation, the first condition can include that the current video block is decoded using bi-predictive coding with unequal weights, and the second condition can include that the current video block is decoded using bi-predictive coding with equal weights.

[0102] In some embodiments, the sub-block size used as the BDOF MV refinement sub-block size can depend on conditions. For example, the sub-block size can be fixed. Alternatively, the sub-block size can depend on the size of the current prediction unit (PU) that can include the current video block, the size of the current coding unit (CU) that can include the current video block, the characteristics of multiple reference video blocks, the similarity of multiple prediction values from multiple reference video blocks, the difference distribution between multiple prediction values from multiple reference video blocks, the temporal gradient of multiple reference video blocks, the spatial gradient of multiple reference video blocks, the prediction type, the adjustment value determined in the first pass of multi-pass decoder-side motion vector refinement (DMVR), the adjustment value determined in the second pass of multi-pass DMVR, the sequence resolution, etc.

[0103] In some embodiments, the temporal gradient can be determined based on the sum of absolute differences (SAD). In other embodiments, the height of the sub-block size can depend on the height or width of the current video block. Additionally or alternatively, the width of the sub-block size can depend on the height or width of the current video block.

[0104] In some embodiments, a first cost for evaluating conditions for a BDOF process may depend on a second cost among a plurality of reference video blocks. Additionally, the second cost may be determined based on different cost functions.

[0105] For example, different cost functions may include SAD, mean-removed SAD (MR-SAD), sum of absolute transform differences (SATD), mean-removed SATD (MR-SATD), sum of squared differences (SSD), mean-removed SSD (MR-SSD), sum of squared errors (SSE), mean-removed SSE (MR-SSE), weighted SAD, weighted MR-SAD, weighted SATD, weighted MR-SATD, weighted SSD, weighted MR-SSD, weighted SSE, weighted MR-SSE, or gradient information.

[0106] In some embodiments, the second cost may be determined based on SAD among a plurality of reference video blocks, SATD among a plurality of reference video blocks, MR-SAD among a plurality of reference video blocks, or a weighted average of at least two of SAD, MR-SAD, or SATD. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0107] In some embodiments, a filter may be applied to at least one offset. Compared with traditional solutions that do not apply a filtering process, the proposed method may advantageously smooth at least one offset, thereby improving coding and decoding quality.

[0108] In some embodiments, a filter having a predetermined shape may be applied to the MVs determined by the BDOF process for sub-blocks of a current video block. Figure 11 Five example filter shapes are shown. In one example, all MVs within a PU may be used during the application of the filter. In another example, MVs having the same second-pass DMVR MVs may be used during the application of the filter.

[0109] In some embodiments, the filter has weights. For example, the weight for the center position in the filter may be equal to a third value (such as 1, 4, 8, etc.), and the weights for the four sides around the center position may be equal to a fourth value (such as 1, 2, 3, etc.). The third value may be the same as the fourth value. Alternatively, the third value may be different from the fourth value.

[0110] In some embodiments, the gradient at a sample point may be determined based on the difference between two neighboring sample points of the sample point. For example, the gradient may include a horizontal gradient and / or a vertical gradient. The horizontal gradient and the vertical gradient may be determined based on the following: where I(k) (i, j) represents the sample value at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, where k can be equal to 0 or 1. represents the value for the horizontal gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, and represents the value for the vertical gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k.

[0111] In some embodiments, the gradient at a sample can be determined based on the difference between two shifted neighboring samples of the sample. For example, the horizontal gradient and the vertical gradient can be determined based on the following: where I (k) (i, j) represents the sample value at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, where k can be equal to 0 or 1. represents the value for the horizontal gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, and each of shift1 and shift2 can be an integer, such as 0, 1, 2, 6, etc.

[0112] In some embodiments, the gradient at a sample can be determined based on a first number of samples before the sample and a second number of samples after the sample. In some embodiments, each of the first number and the second number can be an integer. As an example, the first number and / or the second number can be 0, 3, 10, etc. For example, the horizontal gradient and the vertical gradient can be determined based on the following: where I (k) (i, j) represents the sample value at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, where k can be equal to 0 or 1. represents the value for the horizontal gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, Na represents the second number, Nb represents the first number, and whp represents the weight for the sample at coordinate (i + p, j), and wvp represents the weight for the sample at coordinate (i, j + p).

[0113] In some embodiments, the weights used to determine the horizontal gradient may be the same as the weights used to determine the vertical gradient. Alternatively, the weights used to determine the horizontal gradient may be different from the weights used to determine the vertical gradient.

[0114] In some embodiments, the weights used to determine the horizontal gradient and / or the weights used to determine the vertical gradient may be indicated in the bitstream or determined based on the decoded information.

[0115] In some embodiments, the first number and the second number may be different for determining the gradients in the horizontal and vertical directions. Alternatively, the first number and the second number may be the same for determining the gradients in the horizontal and vertical directions.

[0116] In some embodiments, the division operation may be replaced by a non-division operation, which may employ the same or similar logic as the division replacement logic in the Cross-Component Linear Model (CCLM) or the Convolutional Cross-Component Model (CCCM).

[0117] In some embodiments, whether to apply the method and / or how to apply the method may depend on the decoded information of the current video block. By way of example and not limitation, the decoded information may include block size, temporal layer, color component, slice type, and / or picture type.

[0118] In some embodiments, whether to apply the method and / or how to apply the method may be indicated in the bitstream. In one example, an indication for indicating whether to enable the method or the method to be applied may be indicated at the sequence level, group of pictures level, picture level, slice level, slice group level, etc.

[0119] In some embodiments, an indication for indicating whether to enable the method or the method to be applied may be indicated in the sequence header, picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice header, or slice group header.

[0120] In some embodiments, an indication for indicating whether to enable the method or the method to be applied may be indicated at a region including more than one sample or pixel. For example, the region may include a Prediction Block (PB), Transform Block (TB), Coding Block (CB), Prediction Unit (PU), Transform Unit (TU), Coding Unit (CU), Virtual Pipeline Data Unit (VPDU), Coding Tree Unit (CTU), CTU row, slice, picture, or sub-picture.

[0121] It should be understood that the above diagrams and / or examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard. In addition, the specific values described herein are for illustrative purposes only and do not limit the scope of the present disclosure.

[0122] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing. In this method, a plurality of weights are obtained. The plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of a video. The metric includes at least one of a gradient or a difference of sample values. The plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block. In addition, the bitstream is generated based on the plurality of weights.

[0123] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, a plurality of weights are obtained. The plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of a video. The metric includes at least one of a gradient or a difference of sample values. The plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block. In addition, the bitstream is generated based on the plurality of weights and stored in a non-transitory computer-readable recording medium.

[0124] Embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.

[0125] Item 1. A method for video processing, comprising: obtaining a plurality of weights for conversion between a current video block of a video and a bitstream of the video, wherein the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to the current video block, the metric includes at least one of a gradient or a difference of sample values, and the plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block; and performing the conversion based on the plurality of weights.

[0126] Item 2. The method according to Item 1, wherein performing the conversion includes: determining a set of parameters for the BDOF process based on the plurality of weights; determining, based on the set of parameters, a motion vector (MV) for refining the current video block or at least one offset for adjusting a current sample point in the current video block; and performing the conversion based on the at least one offset.

[0127] Item 3. The method according to item 2, wherein the at least one offset is used to refine the motion vector of the current video block, the size of the current video block being M×N, the target region including a region of size (M + K1)×(N + K2) around the current video block, and each of M, N, K1, and K2 being an integer, or wherein the at least one offset is used to adjust the current sample, the target region including a region of size K3×K4 around the current sample, and each of K3 and K4 being an integer.

[0128] Item 4. The method according to any one of items 2 to 3, wherein the set of parameters for the BDOF process is determined based on the plurality of weights and a complete linear equation formula.

[0129] Item 5. The method according to any one of items 2 to 4, wherein the set of parameters includes a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter determined in a predetermined manner.

[0130] Item 6. The method according to item 5, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient, and the set of parameters is determined based on the following: s1 = ∑(Gx·Gx), s2 = ∑(Gx·Gy), s3 = ∑(dI·Gx), s5 = ∑(Gy·Gx), s6 = ∑(dI·Gy), where s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, Gx represents the sum of the values of the horizontal gradient determined for each of the plurality of reference video blocks, Gy represents the sum of the values of the vertical gradient determined for each of the plurality of reference video blocks, dI represents the difference in sample values between the plurality of reference video blocks, and ∑() represents the weighted sum within the target region based on the plurality of weights.

[0131] Item 7. The method according to item 6, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient, and the set of parameters is determined based on the following: s1 = ∑ (i,j)∈Ω Abs(ψ x (i, j)), s2 = ∑ (i,j)∈Ω ψ x (i, j)·Sign(ψ y (i, j)), s3 = ∑(i,j)∈Ω θ(i, j)·Sign(ψ x (f, j)), s4 = ∑ (i,j)∈Ω Abs(ψ y (i, j)), s5 = ∑ (i,j)∈Ω θ(i, j)·Sign(ψ y (i, j)), θ(i, j) = (I (1) (i, j) >> nb) - (I (0) (i, j) >> nb), where s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, I( k )(i, j) represents the sample value at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, k equals 0 or 1, represents the value for the horizontal gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at coordinate (i, j) of the prediction signal of the reference video block of the current video block in list k, ∑ (i,j)∈Ω () represents the weighted sum within the target region Ω based on the multiple weights, abs(z) represents the absolute value of the number z, sign(z) represents the sign of the number z, and each of na and nb is an integer.

[0132] Item 8. The method according to any one of Items 5 to 7, wherein determining the at least one offset includes: determining a set of determinant values based on the set of parameters; and determining the at least one offset based on the set of determinant values.

[0133] Item 9. The method according to Item 8, wherein the set of determinant values is determined based on the following: D = (s1 >> shTem) * (s5 >> shTem) - (s2 >> shTem) * (s2 >> shTem), Dx = (s3 >> shTem) * (s5 >> shTem) - (s6 >> shTem) * (s2 >> shTem), Dy = (s1 >> shTem) * (s6 >> shTem) - (s3 >> shTem) * (s2 >> shTem), Where D represents the first determinant value in the set of determinant values, Dx represents the second determinant value in the set of determinant values, Dy represents the third determinant value in the set of determinant values, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and shTem is an integer.

[0134] Item 10. The method according to item 9, wherein the at least one offset is determined based on: vx = Dx / D, vy = Dy / D, where vx represents the first offset in the at least one offset, and vy represents the second offset in the at least one offset.

[0135] Item 11. The method according to any one of items 9 to 10, wherein if the absolute value of the first determinant value is less than a predetermined threshold, each of the at least one offset is equal to a predetermined value.

[0136] Item 12. The method according to item 11, wherein the predetermined threshold is a non - negative number.

[0137] Item 13. The method according to any one of items 5 to 7, wherein the first offset in the at least one offset is determined based on the set of parameters, and the second offset in the at least one offset is determined based on the first offset and the set of parameters.

[0138] Item 14. The method according to item 13, wherein the at least one offset is determined based on: vx = s3 / s1, vy = (s6 - s2 * vx) / s5, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter.

[0139] Item 15. The method according to item 13, wherein the first offset and the second offset are determined based on: vx = s3 / s1, vy = (s6 - s2 * vx / T1) / s5, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and T1 represents a real number.

[0140] Item 16. The method according to any one of items 5 to 7, wherein the second offset in the at least one offset is determined based on the set of parameters, and the first offset in the at least one offset is determined based on the second offset and the set of parameters.

[0141] Item 17. The method according to Item 16, wherein the first offset and the second offset are determined based on: vy = s6 / s5, vx = (s3 - s2 * vy) / s1, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter.

[0142] Item 18. The method according to Item 16, wherein the first offset and the second offset are determined based on: vy = s6 / s5, vx = (s3 - s2 * vy / T2) / s1, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and T2 represents a real number.

[0143] Item 19. The method according to any one of Items 5 to 7, wherein the at least one offset is determined based on: vx = s3 / s1, vy = s6 / s5, where vx represents the first offset among the at least one offset, vy represents the second offset among the at least one offset, s1 represents the first parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter.

[0144] Item 20. The method according to any one of Items 2 to 19, wherein a shift operation or a clipping operation is used to determine the at least one offset.

[0145] Item 21. The method according to Item 20, wherein the shift operation is applied to at least one of the numerator or the denominator of a term in the equation for determining the at least one offset.

[0146] Item 22. The method according to any one of Items 20 to 21, wherein the shift operation includes shifting left by a predetermined number of bits.

[0147] Item 23. The method according to any one of Items 20 to 22, wherein the shift operation is applied in a predetermined order.

[0148] Item 24. The method according to any one of Items 20 to 23, wherein the clipping operation is applied to the at least one offset to update the at least one offset.

[0149] Item 25. The method according to Item 25, wherein the value of the updated at least one offset is between a first value and a second value.

[0150] Item 26. The method according to any one of Items 2 to 25, wherein performing the transformation based on the at least one offset includes: adjusting the at least one offset by using at least one scaling factor; and performing the transformation based on the adjusted at least one offset.

[0151] Item 27. The method according to Item 26, wherein a first offset among the at least one offset is adjusted by multiplying the first offset by a first scaling factor among the at least one scaling factor, and the first scaling factor is a real number.

[0152] Item 28. The method according to Item 26, wherein a second offset among the at least one offset is adjusted by dividing the second offset by a second scaling factor among the at least one scaling factor, and the second scaling factor is a real number.

[0153] Item 29. The method according to any one of Items 26 to 28, wherein different offsets among the at least one offset are adjusted by using different scaling factors among the at least one scaling factor.

[0154] Item 30. The method according to any one of Items 26 to 29, wherein the at least one scaling factor depends on at least one of the following: block size, sequence resolution, or block characteristics.

[0155] Item 31. The method according to any one of Items 1 to 30, wherein each of the plurality of weights is equal to the same predetermined value.

[0156] Item 32. The method according to any one of Items 1 to 30, wherein a first weight among the plurality of weights corresponding to a first sample point in the target region depends on the position of the first sample point in the target region.

[0157] Item 33. The method according to Item 32, wherein the first weight is determined based on the following: w1 = (x >= (wt / 2)? wt - x : x + 1) * (y >= (ht / 2)? ht - y : y + 1), where w1 represents the first weight, x represents the horizontal position of the first sample point in the target region, y represents the vertical position of the first sample point in the target region, wt represents the width of the target region, and ht represents the height of the target region.

[0158] Item 34. The method according to any one of Items 1 to 30, wherein the plurality of weights are determined based on a predetermined probability distribution.

[0159] Item 35. The method according to Item 34, wherein the predetermined probability distribution includes a Gaussian distribution having a predetermined standard deviation.

[0160] Item 36. The method according to any one of Items 1 to 30, wherein the plurality of weights are implemented using a shift operation.

[0161] Item 37. The method according to Item 36, wherein the plurality of weights are represented as a left shift matrix or a right shift matrix.

[0162] Item 38. The method according to any one of Items 1 to 37, wherein the plurality of weights depend on at least one of the following: block size, block shape, block characteristics, or sequence resolution.

[0163] Item 39. The method according to any one of Items 1 to 38, wherein the plurality of weights are indicated in one of the following: Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or Slice Header (SH).

[0164] Item 40. The method according to any one of Items 1 to 39, wherein whether the BDOF process for MV refinement is applied to the current video block depends on a first condition, or whether the BDOF process for sample adjustment is applied to the current video block depends on a second condition.

[0165] Item 41. The method according to Item 40, wherein the first condition is the same as the second condition.

[0166] Item 42. The method according to Item 40, wherein the first condition is different from the second condition.

[0167] Item 43. The method according to Item 42, wherein the first condition includes that the current video block is encoded and decoded using bi-directional prediction with unequal weights, and the second condition includes that the current video block is encoded and decoded using bi-directional prediction with equal weights.

[0168] Item 44. The method according to any one of Items 1 to 43, wherein the sub-block size used as the BDOF MV refinement sub-block size depends on a condition.

[0169] Item 45. The method according to Item 44, wherein the sub-block size is fixed.

[0170] Item 46. The method according to Item 44, wherein the sub-block size depends on at least one of the following: the size of the current prediction unit (PU) including the current video block; the size of the current coding unit (CU) including the current video block; the characteristics of the plurality of reference video blocks, the similarity of the plurality of prediction values from the plurality of reference video blocks, the distribution of the differences between the plurality of prediction values from the plurality of reference video blocks, the temporal gradient of the plurality of reference video blocks, the spatial gradient of the plurality of reference video blocks, the prediction type, the adjustment value determined in the first pass of decoder-side motion vector refinement (DMVR) in multiple passes, the adjustment value determined in the second pass of the multiple-pass DMVR, or the sequence resolution.

[0171] Item 47. The method according to Item 46, wherein the temporal gradient is determined based on the sum of absolute differences (SAD).

[0172] Item 48. The method according to Item 44, wherein the height of the sub-block size depends on the height or width of the current video block, or the width of the sub-block size depends on the height or width of the current video block.

[0173] Item 49. The method according to any one of Items 1 to 48, wherein the first cost for evaluating the condition for the BDOF process depends on the second cost between the plurality of reference video blocks.

[0174] Item 50. The method according to Item 49, wherein the second cost is determined based on different cost functions.

[0175] Item 51. The method according to Item 50, wherein the different cost functions include at least one of the following: SAD, mean-removed SAD (MR-SAD), sum of absolute transform differences (SATD), mean-removed SATD (MR-SATD), sum of squared differences (SSD), mean-removed SSD (MR-SSD), sum of squared errors (SSE), mean-removed SSE (MR-SSE), weighted SAD, weighted MR-SAD, weighted SATD, weighted MR-SATD, weighted SSD, weighted MR-SSD, weighted SSE, weighted MR-SSE, or gradient information.

[0176] Item 52. The method according to Item 49, wherein the second cost is determined based on at least one of the following: the SAD between the plurality of reference video blocks, the SATD between the plurality of reference video blocks, the MR-SAD between the plurality of reference video blocks, or the weighted average of at least two of the SAD, the MR-SAD, or the SATD.

[0177] Item 53. The method according to any one of Items 2 to 52, wherein a filter is applied to the at least one offset.

[0178] Item 54. The method according to any one of Items 2 to 52, wherein a filter having a predetermined shape is applied to the MVs determined by the BDOF process for sub - blocks of the current video block.

[0179] Item 55. The method according to any one of Items 53 to 54, wherein all MVs within the PU are used during the application of the filter.

[0180] Item 56. The method according to any one of Items 53 to 54, wherein MVs having the same second - pass DMVR MVs are used during the application of the filter.

[0181] Item 57. The method according to any one of Items 53 to 56, wherein the filter has weights.

[0182] Item 58. The method according to Item 57, wherein the weight for the central position in the filter is equal to a third value, and the weights for the four sides around the central position are equal to a fourth value.

[0183] Item 59. The method according to Item 58, wherein the third value is the same as the fourth value, or the third value is different from the fourth value.

[0184] Item 60. The method according to any one of Items 1 to 59, wherein the gradient at a sample point is determined based on the difference between two neighboring sample points of the sample point.

[0185] Item 61. The method according to Item 60, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient determined based on: where I (k) (i,j) represents the sample value at the coordinate (i,j) of the predicted signal of the reference video block of the current video block in list k, k is equal to 0 or 1, represents the value for the horizontal gradient at the coordinate (i,j) of the predicted signal of the reference video block of the current video block in list k, and represents the value for the vertical gradient at the coordinate (i,j) of the predicted signal of the reference video block of the current video block in list k.

[0186] Item 62. The method according to any one of Items 1 to 59, wherein the gradient at a sample point is determined based on the difference between two shifted neighboring sample points of the sample point.

[0187] Item 63. The method according to Item 62, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient determined based on the following: where I (k) (i,j) represents the sample value at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, where k is equal to 0 or 1, represents the value for the horizontal gradient at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, and each of shift1 and shift2 is an integer.

[0188] Item 64. The method according to any one of Items 1 to 59, wherein the gradient at the sample is determined based on a first number of samples before the sample and a second number of samples after the sample.

[0189] Item 65. The method according to Item 64, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient determined based on the following: where I (k) (i,j) represents the sample value at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, where k is equal to 0 or 1, represents the value for the horizontal gradient at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, Na represents the second number, Nb represents the first number, whp represents the weight for the sample at the coordinate (i + p,j), and wvp represents the weight for the sample at the coordinate (i,j + p).

[0190] Item 66. The method according to Item 65, wherein the weight for determining the horizontal gradient is the same as the weight for determining the vertical gradient, or the weight for determining the horizontal gradient is different from the weight for determining the vertical gradient.

[0191] Item 67. The method according to Item 66, wherein the weight for determining the horizontal gradient and / or the weight for determining the vertical gradient are indicated in the bitstream or determined based on the decoded information.

[0192] Item 68. The method according to any one of Items 64 to 67, wherein each of the first number and the second number is an integer.

[0193] Item 69. The method according to any one of Items 64 to 68, wherein the first number and the second number are different for determining the gradients in the horizontal and vertical directions, or the first number and the second number are the same for determining the gradients in the horizontal and vertical directions.

[0194] Item 70. The method according to any one of Items 1 to 69, wherein the division operation is replaced by a non-division operation.

[0195] Item 71. The method according to any one of Items 1 to 70, wherein whether to apply the method and / or how to apply the method depends on the decoded information of the current video block.

[0196] Item 72. The method according to Item 71, wherein the decoded information includes at least one of the following: block size, temporal layer, color component, slice type, or picture type.

[0197] Item 73. The method according to any one of Items 1 to 72, wherein whether to apply the method and / or how to apply the method is indicated in the bitstream.

[0198] Item 74. The method according to any one of Items 1 to 72, wherein the indication for indicating whether to enable the method or the method to be applied is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0199] Item 75. The method according to any one of Items 1 to 72, wherein the indication for indicating whether to enable the method or the method to be applied is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.

[0200] Item 76. The method according to any one of Items 1 to 72, wherein the indication for indicating whether to enable the method or the method to be applied is indicated at an area including more than one sample or pixel.

[0201] Item 77. The method according to Item 76, wherein the region includes one of the following: a prediction block (PB), a transform block (TB), a coding / decoding block (CB), a prediction unit (PU), a transform unit (TU), a coding / decoding unit (CU), a virtual pipeline data unit (VPDU), a coding / decoding tree unit (CTU), a CTU row, a stripe, a slice, or a sub-picture.

[0202] Item 78. The method according to any one of Items 1 to 77, wherein the conversion includes encoding the current video block into the bitstream.

[0203] Item 79. The method according to any one of Items 1 to 77, wherein the conversion includes decoding the current video block from the bitstream.

[0204] Item 80. An apparatus for video processing, including a processor and a non-transitory memory having instructions, wherein when the instructions are executed by the processor, the processor executes the method according to any one of Items 1 to 79.

[0205] Item 81. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Items 1 to 79.

[0206] Item 82. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method includes: obtaining a plurality of weights, wherein the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of the video, the metric including at least one of a gradient or a difference of sample values, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; and generating the bitstream based on the plurality of weights.

[0207] Item 83. A method for storing a bitstream of a video, including: obtaining a plurality of weights, wherein the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of the video, the metric including at least one of a gradient or a difference of sample values, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; generating the bitstream based on the plurality of weights; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0208] Figure 13FIG. 1300 is a block diagram of a computing device in which various embodiments of the present disclosure may be implemented. The computing device 1300 may be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or may be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).

[0209] It should be understood that Figure 13 the computing device 1300 shown in FIG. 1300 is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.

[0210] As Figure 13 shown, the computing device 1300 includes a general-purpose computing device 1300. The computing device 1300 may include at least one or more processors or processing units 1310, a memory 1320, a storage unit 1330, one or more communication units 1340, one or more input devices 1350, and one or more output devices 1360.

[0211] In some embodiments, the computing device 1300 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 1300 may support any type of interface to the user (such as a "wearable" circuitry, etc.).

[0212] The processing unit 1310 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in the memory 1320. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 1300. The processing unit 1310 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0213] The computing device 1300 generally includes various computer storage media. Such media can be any media accessible by the computing device 1300, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 1320 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 1330 can be any removable or non-removable media and can include machine-readable media such as a memory, flash drive, magnetic disk, or other media that can be used to store information and / or data and can be accessed within the computing device 1300.

[0214] The computing device 1300 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 13 , a disk drive for reading from and / or writing to a removable non-volatile magnetic disk and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk can be provided. In such a case, each drive can be connected to a bus (not shown) via one or more data media interfaces.

[0215] The communication unit 1340 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 1300 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the computing device 1300 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.

[0216] The input device 1350 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, and so on. The output device 1360 can be one or more of various output devices such as a display, speaker, printer, and so on. With the aid of the communication unit 1340, the computing device 1300 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 1300 can also communicate with one or more devices that enable a user to interact with the computing device 1300, or if needed, the computing device 1300 can also communicate with any device (such as a network card, modem, etc.) that enables the computing device 1300 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0217] In some embodiments, some or all components of computing device 1300 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across the locations of remote data centers. The cloud computing infrastructure may provide services through shared data centers, although to a user, they appear as a single access point. Thus, a cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0218] In an embodiment of the present disclosure, computing device 1300 may be used to implement video encoding / decoding. Memory 1320 may include one or more video codec modules 1325 having one or more program instructions. These modules are accessible and executable by processing unit 1310 to perform the functions of the various embodiments described herein.

[0219] In an example embodiment of performing video encoding, input device 1350 may receive video data as input 1370 to be encoded. The video data may be processed, for example, by video codec module 1325 to generate an encoded bitstream. The encoded bitstream may be provided as output 1380 via output device 1360.

[0220] In an example embodiment of performing video decoding, input device 1350 may receive the encoded bitstream as input 1370. The encoded bitstream may be processed, for example, by video codec module 1325 to generate decoded video data. The decoded video data may be provided as output 1380 via output device 1360.

[0221] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: Obtaining a plurality of weights for a conversion between a current video block of a video and a bitstream of the video, wherein the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to the current video block, the metric includes at least one of a gradient or a difference of sample values, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; And Performing the conversion based on the plurality of weights.

2. The method according to claim 1, wherein performing the conversion includes: Determining a set of parameters for the BDOF process based on the plurality of weights; Based on the set of parameters, determining a motion vector (MV) for refining the current video block or adjusting at least one offset of a current sample in the current video block; And Performing the conversion based on the at least one offset.

3. The method according to claim 2, wherein the at least one offset is used to refine the motion vector of the current video block, the size of the current video block is M×N, the target region includes a region of size (M+K1)×(N+K2) around the current video block, and each of M, N, K1, and K2 is an integer, or wherein the at least one offset is used to adjust the current sample, the target region includes a region of size K3×K4 around the current sample, and each of K3 and K4 is an integer.

4. The method according to any one of claims 2 to 3, wherein the set of parameters for the BDOF process is determined based on the plurality of weights and a complete linear equation formula.

5. The method according to any one of claims 2 to 4, wherein the set of parameters includes a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter determined in a predetermined manner.

6. The method according to claim 5, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient, and the set of parameters is determined based on the following: s1 = ∑(Gx·Gx), s2 = ∑(Gx·Gy), s3 = ∑(dI·Gx), s5 = ∑(Gy·Gx), s6 = ∑(dI·Gy), where s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, Gx represents the sum of the values of the horizontal gradient determined for each reference video block among the plurality of reference video blocks, Gy represents the sum of the values of the vertical gradient determined for each reference video block among the plurality of reference video blocks, dI represents the difference of the sample values between the plurality of reference video blocks, and ∑() represents the weighted sum within the target region based on the plurality of weights.

7. The method according to claim 6, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient, and the set of parameters is determined based on the following: s1 = ∑ (i,j)∈Ω Abs(ψ x (i, j)), s2 = ∑ (i,j)∈Ω ψ x (i, j)·Sign(ψ y (i, j)), s3 = ∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)), s4 = ∑ (i,j)∈Ω Abs(ψ y (i, j)) s5 = ∑ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)), θ(i,j) = (I (1) (i,j) >> nb) - (I (0) (i,j) >> nb), where s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, I (k) (i,j) represents the sample value at coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, where k is equal to 0 or 1, represents the value for the horizontal gradient at coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, ∑ (i,j)∈Ω () represents the weighted sum within the target region Ω based on the plurality of weights, abs(z) represents the absolute value of the number z, sign(z) represents the sign of the number z, and each of na and nb is an integer.

8. The method according to any one of claims 5 to 7, wherein determining the at least one offset comprises: Determining a set of determinant values based on the set of parameters; And Determining the at least one offset based on the set of determinant values.

9. The method according to claim 8, wherein the set of determinant values is determined based on the following: D = (s1 >> shTem) * (s5 >> shTem) - (s2 >> shTem) * (s2 >> shTem), Dx = (s3 >> shTem) * (s5 >> shTem) - (s6 >> shTem) * (s2 >> shTem), Dy = (s1 >> shTem) * (s6 >> shTem) - (s3 >> shTem) * (s2 >> shTem), where D represents the first determinant value in the set of determinant values, Dx represents the second determinant value in the set of determinant values, Dy represents the third determinant value in the set of determinant values, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and shTem is an integer.

10. The method according to claim 9, wherein the at least one offset is determined based on the following: vx = Dx / D, vy = Dy / D, where vx represents the first offset in the at least one offset, and vy represents the second offset in the at least one offset.

11. The method according to any one of claims 9 to 10, wherein if the absolute value of the first determinant value is less than a predetermined threshold, each offset in the at least one offset is equal to a predetermined value.

12. The method according to claim 11, wherein the predetermined threshold is a non - negative number.

13. The method according to any one of claims 5 to 7, wherein the first offset in the at least one offset is determined based on the set of parameters, and the second offset in the at least one offset is determined based on the first offset and the set of parameters.

14. The method according to claim 13, wherein the at least one offset is determined based on the following: vx = s3 / s1, vy = (s6 - s2 * vx) / s5, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter.

15. The method according to claim 13, wherein the first offset and the second offset are determined based on the following: vx = s3 / s1, vy = (s6 - s2 * vx / T1) / s5, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and T1 represents a real number.

16. The method according to any one of claims 5 to 7, wherein a second offset among the at least one offset is determined based on the set of parameters, and a first offset among the at least one offset is determined based on the second offset and the set of parameters.

17. The method according to claim 16, wherein the first offset and the second offset are determined based on the following: vy = s6 / s5, vx = (s3 - s2*vy) / s1, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter.

18. The method according to claim 16, wherein the first offset and the second offset are determined based on the following: vy = s6 / s5, vx = (s3 - s2*vy / T2) / s1, where vx represents the first offset, vy represents the second offset, s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, and T2 represents a real number.

19. The method according to any one of claims 5 to 7, wherein the at least one offset is determined based on the following: vx = s3 / s1, vy = s6 / s5, where vx represents the first offset among the at least one offset, vy represents the second offset among the at least one offset, s1 represents the first parameter, s3 represents the third parameter, s5 represents the fourth parameter, and s6 represents the fifth parameter.

20. The method according to any one of claims 2 to 19, wherein a shift operation or a clamping operation is used to determine the at least one offset.

21. The method according to claim 20, wherein the shift operation is applied to at least one of a numerator or a denominator of a term in an equation for determining the at least one offset.

22. The method according to any one of claims 20 to 21, wherein the shift operation comprises shifting left by a predetermined number of bits.

23. The method according to any one of claims 20 to 22, wherein the shift operation is applied in a predetermined order.

24. The method according to any one of claims 20 to 23, wherein the clamping operation is applied to the at least one offset to update the at least one offset.

25. The method according to claim 25, wherein a value of the updated at least one offset is between a first value and a second value.

26. The method according to any one of claims 2 to 25, wherein performing the conversion based on the at least one offset comprises: adjusting the at least one offset using at least one scaling factor; and performing the conversion based on the adjusted at least one offset.

27. The method according to claim 26, wherein a first offset among the at least one offset is adjusted by multiplying the first offset by a first scaling factor among the at least one scaling factor, and the first scaling factor is a real number.

28. The method according to claim 26, wherein a second offset among the at least one offset is adjusted by dividing the second offset by a second scaling factor among the at least one scaling factor, and the second scaling factor is a real number.

29. The method according to any one of claims 26 to 28, wherein different offsets among the at least one offset are adjusted using different scaling factors among the at least one scaling factor.

30. The method according to any one of claims 26 to 29, wherein the at least one scaling factor depends on at least one of the following: block size, sequence resolution, or block characteristics.

31. The method according to any one of claims 1 to 30, wherein each of the plurality of weights is equal to the same predetermined value.

32. The method according to any one of claims 1 to 30, wherein a first weight among the plurality of weights corresponding to a first sample point in the target region depends on the position of the first sample point in the target region.

33. The method according to claim 32, wherein the first weight is determined based on the following: w1 = (x >= (wt / 2)? wt - x : x + 1) * (y >= (ht / 2)? ht - y : y + 1), where w1 represents the first weight, x represents the horizontal position of the first sample point in the target region, y represents the vertical position of the first sample point in the target region, wt represents the width of the target region, and ht represents the height of the target region.

34. The method according to any one of claims 1 to 30, wherein the plurality of weights are determined based on a predetermined probability distribution.

35. The method according to claim 34, wherein the predetermined probability distribution includes a Gaussian distribution with a predetermined standard deviation.

36. The method according to any one of claims 1 to 30, wherein the plurality of weights are implemented using a shift operation.

37. The method according to claim 36, wherein the plurality of weights are represented as a left shift matrix or a right shift matrix.

38. The method according to any one of claims 1 to 37, wherein the plurality of weights depend on at least one of the following: block size, block shape, block characteristics, or sequence resolution.

39. The method according to any one of claims 1 to 38, wherein the plurality of weights are indicated in one of the following: Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or Slice Header (SH).

40. The method according to any one of claims 1 to 39, wherein whether to apply the BDOF process for MV refinement to the current video block depends on a first condition, or whether to apply the BDOF process for sample adjustment to the current video block depends on a second condition.

41. The method according to claim 40, wherein the first condition is the same as the second condition.

42. The method according to claim 40, wherein the first condition is different from the second condition.

43. The method according to claim 42, wherein the first condition includes that the current video block is encoded and decoded using bi - directional prediction with unequal weights, and the second condition includes that the current video block is encoded and decoded using bi - directional prediction with equal weights.

44. The method according to any one of claims 1 to 43, wherein the sub - block size used as the BD - OF MV refinement sub - block size depends on a condition.

45. The method according to claim 44, wherein the sub - block size is fixed.

46. The method according to claim 44, wherein the sub - block size depends on at least one of the following: the size of the current prediction unit (PU) including the current video block; the size of the current coding unit (CU) including the current video block; the characteristics of the plurality of reference video blocks; the similarity of the plurality of prediction values from the plurality of reference video blocks; the distribution of the differences between the plurality of prediction values from the plurality of reference video blocks; the temporal gradient of the plurality of reference video blocks; the spatial gradient of the plurality of reference video blocks; the prediction type; the adjustment value determined in the first pass of the decoder - side motion vector refinement (DMVR) in multiple passes; the adjustment value determined in the second pass of the multiple - pass DMVR, or the sequence resolution.

47. The method according to claim 46, wherein the temporal gradient is determined based on the sum of absolute differences (SAD).

48. The method according to claim 44, wherein the height of the sub - block size depends on the height or width of the current video block, or the width of the sub - block size depends on the height or width of the current video block.

49. The method according to any one of claims 1 to 48, wherein the first cost for evaluating the condition for the BD - OF process depends on the second cost between the plurality of reference video blocks.

50. The method according to claim 49, wherein the second cost is determined based on different cost functions.

51. The method according to claim 50, wherein the different cost functions include at least one of the following: SAD; mean - removed SAD (MR - SAD); sum of absolute transform differences (SATD); mean - removed SATD (MR - SATD); sum of squared differences (SSD); mean - removed SSD (MR - SSD); sum of squared errors (SSE); mean - removed SSE (MR - SSE); weighted SAD; weighted MR - SAD; weighted SATD; weighted MR - SATD; weighted SSD; weighted MR - SSD; weighted SSE; weighted MR - SSE; or gradient information.

52. The method according to claim 49, wherein the second cost is determined based on at least one of the following: the SAD between the plurality of reference video blocks; the SATD between the plurality of reference video blocks; The MR-SAD between the multiple reference video blocks, or The weighted average of at least two of the SAD, the MR-SAD, or the SATD.

53. The method according to any one of claims 2 to 52, wherein a filter is applied to the at least one offset.

54. The method according to any one of claims 2 to 52, wherein a filter having a predetermined shape is applied to the MVs determined by the BDOF process for sub-blocks of the current video block.

55. The method according to any one of claims 53 to 54, wherein during the application of the filter, all MVs within the PU are used.

56. The method according to any one of claims 53 to 54, wherein during the application of the filter, MVs having the same second-pass DMVR MV are used.

57. The method according to any one of claims 53 to 56, wherein the filter has weights.

58. The method according to claim 57, wherein the weight for the central position in the filter is equal to a third value, and the weights for the four sides around the central position are equal to a fourth value.

59. The method according to claim 58, wherein the third value is the same as the fourth value, or the third value is different from the fourth value.

60. The method according to any one of claims 1 to 59, wherein the gradient at a sample point is determined based on the difference between two neighboring sample points of the sample point.

61. The method according to claim 60, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient determined based on: where I (k) (i, j) represents the sample value at coordinates (i, j) of the prediction signal of the reference video block of the current video block in list k, where k is equal to 0 or 1, represents the value for the horizontal gradient at coordinates (i, j) of the prediction signal of the reference video block of the current video block in list k, and represents the value for the vertical gradient at coordinates (i, j) of the prediction signal of the reference video block of the current video block in list k.

62. The method according to any one of claims 1 to 59, wherein the gradient at a sample point is determined based on the difference between two shifted neighboring sample points of the sample point.

63. The method according to claim 62, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient determined based on: where I (k) (i,j) represents the sample value at coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, where k is equal to 0 or 1, represents the value for the horizontal gradient at coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, and each of shift1 and shift2 is an integer.

64. The method according to any one of claims 1 to 59, wherein the gradient at a sample point is determined based on a first number of sample points before the sample point and a second number of sample points after the sample point.

65. The method according to claim 64, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient determined based on: where I (k) (i,j) represents the sample value at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, where k is equal to 0 or 1, represents the value for the horizontal gradient at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, represents the value for the vertical gradient at the coordinate (i,j) of the prediction signal of the reference video block of the current video block in list k, Na represents the second number, Nb represents the first number, and whp represents the weight for the sample at the coordinate (i+p,j), and wvp represents the weight for the sample at the coordinate (i,j+p).

66. The method according to claim 65, wherein the weight for determining the horizontal gradient is the same as the weight for determining the vertical gradient, or the weight for determining the horizontal gradient is different from the weight for determining the vertical gradient.

67. The method according to claim 66, wherein the weight for determining the horizontal gradient and / or the weight for determining the vertical gradient are indicated in the bitstream or determined based on decoded information.

68. The method according to any one of claims 64 to 67, wherein each of the first number and the second number is an integer.

69. The method according to any one of claims 64 to 68, wherein the first number and the second number are different for determining the gradients in the horizontal and vertical directions, or the first number and the second number are the same for determining the gradients in the horizontal and vertical directions.

70. The method according to any one of claims 1 to 69, wherein the division operation is replaced by a non-division operation.

71. The method according to any one of claims 1 to 70, wherein whether the method is applied and / or how the method is applied depends on the information of the current video block that has been coded and decoded.

72. The method according to claim 71, wherein the coded and decoded information includes at least one of the following: block size, temporal layer, color component, slice type, or picture type.

73. The method according to any one of claims 1 to 72, wherein whether the method is applied and / or how the method is applied is indicated in the bitstream.

74. The method according to any one of claims 1 to 72, wherein the indication for indicating whether the method is enabled or the method to be applied is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.

75. The method according to any one of claims 1 to 72, wherein the indication for indicating whether the method is enabled or the method to be applied is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.

76. The method according to any one of claims 1 to 72, wherein the indication for indicating whether the method is enabled or the method to be applied is indicated at a region including more than one sample or pixel.

77. The method according to claim 76, wherein the region includes one of the following: prediction block (PB), transformation block (TB), coded and decoded block (CB), prediction unit (PU), transformation unit (TU), coded and decoded unit (CU), virtual pipeline data unit (VPDU), coded and decoded tree unit (CTU), CTU row, slice, tile, or sub-picture.

78. The method according to any one of claims 1 to 77, wherein the transformation includes encoding the current video block into the bitstream.

79. The method according to any one of claims 1 to 77, wherein the transformation includes decoding the current video block from the bitstream.

80. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 79.

81. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 79.

82. A non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing of a video, wherein the method includes: Obtaining a plurality of weights, wherein the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of the video, the metric includes at least one of a gradient or a difference of a sample value, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; And Generating the bitstream based on the plurality of weights.

83. A method for storing a bitstream of a video, including: Obtaining a plurality of weights, wherein the plurality of weights are used to weight a plurality of values of a metric at each sample point in a target region of a bidirectional optical flow (BDOF) process applied to a current video block of the video, the metric includes at least one of a gradient or a difference of a sample value, and the plurality of values of the metric are determined based on a plurality of reference video blocks of the current video block; Generating the bitstream based on the plurality of weights; And Storing the bitstream in a non-transitory computer-readable recording medium.