Method and device for video processing and medium

By combining sub-blocks with the same or similar motion vectors and applying a bidirectional optical flow process, the problem of high computational complexity in existing technologies is solved, achieving more efficient video encoding and decoding.

CN120937343APending Publication Date: 2025-11-11DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480021324.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-24
Filing Date
2024-03-22
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially when applying bidirectional optical flow processes, which have high computational complexity, leading to resource waste and low efficiency.

Method used

By combining multiple sub-blocks with the same or similar motion vectors into a combined sub-block and applying a bidirectional optical flow process, the independent processing of each sub-block is reduced, thereby improving encoding and decoding efficiency.

Benefits of technology

It improves the efficiency of video encoding and decoding, reduces computational complexity, and enhances resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937343A_ABST
    Figure CN120937343A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method includes determining, for a current sub-block of a current video block of a video, at least one sub-block from a plurality of sub-blocks of the current video block, a motion vector (MV) of each of the at least one sub-block and an MV of the current sub-block satisfying one of the following conditions: the MV of each of the at least one sub-block is the same as the MV of the current sub-block, or a difference measure between the MV of each of the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block; and performing the conversion based on applying the bi-directional optical flow process to the combined sub-blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to bidirectional optical flow (BDOF) processes. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: a conversion between a current video block and a bitstream of the video; determining at least one sub-block from a plurality of sub-blocks of the current video block for a current sub-block, wherein the motion vector (MV) of each sub-block in the at least one sub-block satisfies one of the following conditions: the MV of each sub-block in the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each sub-block in the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block by combining the at least one sub-block and the current sub-block; and performing a conversion based on applying a bidirectional optical flow (BDOF) process to the combined sub-block.

[0005] According to the method of the first aspect of this disclosure, more than one sub-block having the same or similar MV is combined, and the BDOF process is applied to the combined sub-blocks, rather than being applied individually to each sub-block within the more than one sub-block. Compared to conventional solutions where the BDOF process is applied individually to sub-blocks, the proposed method can advantageously perform the BDOF process more efficiently. In this way, encoding and decoding efficiency can be improved.

[0006] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining at least one sub-block from a plurality of sub-blocks of a current video block, wherein the motion vector (MV) of each sub-block in the at least one sub-block satisfies one of the following conditions: the MV of each sub-block in the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each sub-block in the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block by combining the at least one sub-block and the current sub-block; and generating a bitstream based on applying a bidirectional optical flow (BDOF) process to the combined sub-block.

[0009] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining at least one sub-block from a plurality of sub-blocks of a current video block, wherein the motion vector (MV) of each sub-block in the at least one sub-block satisfies one of the following conditions: the MV of each sub-block in the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each sub-block in the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block by combining the at least one sub-block and the current sub-block; generating a bitstream based on applying a bidirectional optical flow (BDOF) process to the combined sub-block; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown;

[0013] Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown;

[0014] Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown;

[0015] Figure 4 The Extended Codec Unit (CU) region used in BDOF is shown;

[0016] Figure 5 This shows the refinement of motion vectors on the decoding side;

[0017] Figure 6 The diamond-shaped area in the search region is shown;

[0018] Figure 7 The weights generated using the example Gaussian distribution are shown;

[0019] Figure 8 The weights generated using another example Gaussian distribution are shown;

[0020] Figure 9 The weights generated using yet another example of a Gaussian distribution are shown;

[0021] Figure 10 The weights generated using yet another example of a Gaussian distribution are shown;

[0022] Figure 11 Different filter shapes are shown for application to the data;

[0023] Figure 12 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and

[0024] Figure 13 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0025] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation

[0026] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0027] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0028] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0029] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example Environment

[0031] Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0032] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0033] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0034] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0035] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0036] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0037] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0038] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0039] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0040] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.

[0041] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0042] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0043] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0044] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0045] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0046] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0047] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can transmit the motion information of the current video block via a signal, referencing the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0048] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0049] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0050] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0051] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0052] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0053] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0054] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0055] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0056] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0057] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0058] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0059] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0060] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0061] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0062] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge pattern. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0063] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0064] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.

[0065] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0066] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0067] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0068] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates. 1. Brief Overview This disclosure relates to video / image coding and decoding technologies. Specifically, it relates to bidirectional optical flow. It can be applied to existing video coding and decoding standards such as HEVC and VVC, or to next-generation video coding and decoding standards such as ECM that go beyond VVC. It can also be applied to future video coding and decoding standards or video codecs. 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. As of July 2020, it had also completed the Multi-Functional Video Coding (VVC) standard, aiming to further reduce the bitrate by 50% and provide a range of additional features. Following the completion of VVC, activities beyond VVC have begun. These activities were undertaken by M. Coban, F. Léannec, K. Naser, and J. The description of the additional tools on top of the VVC tools has been summarized in “Algorithm Description of Enhanced Compression Model 5 (ECM 5)” (26th JVET Conference, April 20-29, 2022: via teleconference, document JVET-Z2025), and its reference software is named ECM. 2.1 Bidirectional Optical Flow (BDOF) in VVC The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) was included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires far fewer computations, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of the CU at the 4×4 sub-block level. BDOF is applied to the CU if all of the following conditions are met: -CU is encoded and decoded using a "true" bidirectional prediction mode, meaning that one of the two reference images is displayed before the current image in the order of display, and the other is displayed after the current image in the order of display. - The distances from the two reference images to the current image (i.e., the difference in point of view) are the same. Both reference images are short-term reference images. -CU is not encoded or decoded using affine mode or SbTMVP Merge mode. -CU has more than 64 luminance samples. - Both the CU height and CU width are greater than or equal to 8 luminance samples. -BCW weight index indicates equal weights. -WP is for cases where the current CU is not enabled. -CIIP mode is not used in the current CU. BDOF is applied only to the luminance component. As its name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4×4 sub-block, motion refinement (v) is calculated by minimizing the difference between the L0 and L1 predicted samples. x ,v y Motion refinement is then used to adjust the bidirectional prediction sample values ​​in the 4x4 sub-blocks. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two predicted signals are calculated by directly calculating the difference between two neighboring sample points. k = 0, 1, that is, Among them I (k) (i,j) is the sample value at coordinate (i,j) of the predicted signal in list k (k=0,1), and shift1 is calculated based on the luminance bit depth bitDepth as shift1=max(6,bitDepth-6). Then, the autocorrelation and cross-correlation of the gradients S1, S2, S3, S5, and S6 are calculated as follows: S1=∑ (i,j)∈Ω Abs(ψ x (i,j)),S3=∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) S5=∑ (i,j)∈Ω Abs(ψ y (i,j)),S6=∑ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) in θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b ) Where Ω is the 6×6 window surrounding the 4×4 sub-block, and n a and n b The values ​​were set to min(1, bitDepth-11) and min(4, bitDepth-8) respectively. Motion refinement (v) x ,v y Then, the following formula is derived using cross-correlation and autocorrelation terms: in th′ BIO =2 max(5,BD-7) . It is a floor function, and Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4×4 sub-block: Finally, the BDOF samples of CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift These values ​​were chosen so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. To derive the gradient values, it is necessary to generate some predicted sample points I in the list k (k = 0, 1) outside the current CU boundary. (k) (i,j). For example... Figure 4 As shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region (white area) are generated by directly taking reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and a normal 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values ​​are used only for gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled from their nearest neighbors (i.e., repeated). When the width and / or height of a CU is greater than 16 luminance samples, it will be divided into sub-blocks with a width and / or height equal to 16 luminance samples, and the sub-block boundaries will be considered as CU boundaries in the BDOF process. The maximum cell size for the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. The BDOF process is not applied to the sub-block when the SAD between the initial L0 and L1 predicted samples is less than a threshold. The threshold is set to equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 and L1 predicted samples calculated in the DVMR process is reused here. Bidirectional optical flow (BDOF) is disabled if BCW is enabled for the current block, meaning the BCW weight index indicates unequal weights. Similarly, BDOF is disabled if WP is enabled for the current block, meaning luma_weight_lx_flag is 1 for either of the two reference images. BDOF is also disabled when the CU is encoded / decoded using symmetric MVD mode or CIIP mode. 2.1.1 BDOF in ECM: Sample-based BDOF In sample-based BDOF, motion refinement is not performed on a block-based basis (Vx, Vy), but rather on a per-sample basis. The encoding / decoding block is divided into 8×8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD (Solution-Adjustment) between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, a sliding 5×5 window is used for each sample in the sub-block, and the existing BDOF procedure is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional predicted sample values ​​for the center sample of the window. 2.2 Decoder-Side Motion Vector Refinement (DMVR) in VVC To improve the accuracy of the motion vector refinement (MV) in the Merge mode, a decoder-side motion vector refinement based on bilateral matching (BM) is applied in the VVC. In the bidirectional prediction operation, a refined MV is searched around the initial MV in reference image lists L0 and L1. The BM method computes the distortion between two candidate blocks in reference image lists L0 and L1. Figure 5 As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, the application of DMVR is restricted and can only be used with CUs that are encoded and decoded using the following modes and features: - CU-level Merge pattern with bidirectional prediction of MV. - Relative to the current image, one reference image is from the past and the other is from the future. - The distances from the two reference images to the current image (i.e., the difference in point of view) are the same. Both reference images are short-term reference images. -CU has more than 64 luminance samples. - Both the CU height and CU width are greater than or equal to 8 luminance samples. -BCW weight index indicates equal weights. -WP is for blocks that are not enabled. -CIIP mode is not used in the current block. The refined motion vector (MV) derived through the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction in future image encoding and decoding. The original MV is used in the deblocking process and is also used for spatial motion vector prediction in future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-entries. In DVMR, the search point revolves around the initial MV, and the MV offset follows the MV difference mirror rule. In other words, any point examined by DMVR (represented by the candidate MV pair (MV0, MV1)) obeys the following two equations. MV0′=MV0+MV_offset MV1′ = MV1 - MV_offset Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference images. The refinement search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage. A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of the uncertainty in DMVR refinement, a bias towards the original MV is proposed during the DMVR process. The SAD between reference blocks referenced by the initial MV candidate references is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation, rather than through an additional search utilizing SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration with the minimum SAD at the center. In subpixel offset estimation based on parametric error surfaces, the cost at the center location and the costs at the four nearest neighbor locations are used to fit a two-dimensional parabolic error surface equation of the following form. E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C Where (x) min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost. The above equation is solved by using the costs of the five search points, (x) min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) Since all values ​​are positive and the minimum value is E(0,0), therefore x min . and y min The value is automatically constrained to between -8 and 8. This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fraction (x min ,y min An integer distance refinement MV is added to obtain a subpixel-precise refinement increment MV. In VVC, the resolution of the MV is 1 / 16 of a lumen sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search point surrounds the initial fractional pixel MV with an integer sample offset; therefore, for the DMVR search process, samples at those fractional positions need to be interpolated. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that, utilizing a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not needed by the interpolation process based on the original MV but are needed by the interpolation process based on the refined MV are filled from those available samples. When the width and / or height of a CU is greater than 16 luminance samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luminance samples. The maximum cell size for the DMVR search process is limited to 16x16. 2.3. Multi-pass decoder-side motion vector refinement (ECM) Multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for use in spatial and temporal motion vector prediction. 2.3.1 First pass - Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying the BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive the integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to iterate through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the mean-removed SAD (MRSAD) cost function is applied to remove the DC effect of distortion between reference blocks. The local search intDeltaMV is terminated when bilCost at the center point of the 3×3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: ●MV0_pass1=MV0+deltaMV, ●MV1_pass1=MV1–deltaMV. 2.3.2 Second pass - Sub-block-based bilateral matching MV refinement In the second pass, the refined MV is derived by applying the BM to 16×16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference image lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range [-sHor, sHor] in the horizontal direction and a search range [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference subblocks, as follows: bilCost = satdCost * costFactor. The search region (2 * sHor + 1) * (2 * sVer + 1) is divided into... Figure 6The diagram shows a maximum of five diamond-shaped search regions. Each search region is assigned a costFactor, determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order from the top left corner to the bottom right corner. A full integer-pixel search terminates when the minimum bilCost within the current search region is less than a threshold (equal to sbW*sbH); otherwise, the full integer-pixel search continues to the next search region until all search points have been checked. Additionally, the search process terminates if the difference between the previous minimum cost and the current minimum cost in an iteration is less than a threshold (equal to the area of ​​the block). The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). The MV of the second refinement is then derived as: ●MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2) ●MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2). 2.3.3 Third pass - Sub-block based bidirectional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy, without clipping from the refined MV of the parent block in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32. The third refined MV (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as follows: ●MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv, ●MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)–bioMv. In all of the aforementioned sub-entries, when surround motion compensation is enabled, the motion vector must be limited to account for surround offset. 2.3.4 Adaptive Decoder-Side Motion Vector Refinement The adaptive decoder-side motion vector refinement method is an extension of multi-pass DMVR, consisting of two new Merge modes for refining the motion vector only in one direction (L0 or L1) of the bidirectional predictions of Merge candidates that satisfy the DMVR conditions. The multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vector; however, in the first pass (i.e., PU level) of DMVR, either MVD0 or MVD1 is set to zero. The new Merge mode derives its Merge candidates from spatially adjacent encoded / decoded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates, similar to the regular Merge mode. The difference is that only those satisfying the DMVR conditions are added to the candidate list. Both new Merge modes use the same Merge candidate list. If the BM candidate list contains inherited BCW weights, the DMVR process remains unchanged, except for the use of MRSAD or MRSATD for distortion calculation when the weights are unequal and bidirectional prediction is weighted using BCW weights. The encoding and decoding of the Merge index is the same as in the regular Merge mode. 3. Problem Several aspects of BDOF MV refinement / sample adjustment can be improved. - The current formula used to drive the BDOF parameters is not an accurate formula. - There are no weights to indicate the importance of each sample point in the final formula. - There is no filtering process to smooth the final derived MV refinement / sample adjustment. - There is no clear distinction regarding the conditions for applying BDOF to refine MV / adjust sample points. Similarly, there is no distinction regarding their formulas. 4. Detailed Solution The detailed solutions below should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way. The methods disclosed below can be applied to bidirectional optical flow, decoder-side motion vector refinement, and any of their extensions. Derivation of BDOF MV refinement parameters In the following sections, the general equations used to derive the BDOF parameters (vx and vy) are defined as follows: ∑Gx.Gx*vx+∑Gx.Gy*vy=∑dI.Gx.→s1*vx+s2*vy=s3, ∑Gx.Gy*vx+∑Gy.Gy*vy=∑dI.Gy→s2*vx+s5*vy=s6, Where Gx and Gy represent the sum of the horizontal and vertical gradients for the two reference images, respectively. dI represents the difference between the two reference images. The summation (∑) is performed within a predefined region, which can be an NxM block around the current sample point (used for sample-adjusted BDOF) or an NxM block around the current predicted sub-block (used for MV-refined BDOF). 1. A method for deriving gradients, different from BDOF in VVC, is proposed, which can be used to calculate horizontal and / or vertical gradients. a. In one example, the gradient is calculated by directly computing the difference between two neighboring sample points, i.e., b. In another example, the gradient is calculated by computing the difference between two shifted neighboring samples, i.e., i. shift1 and shift2 can be any integer, such as 0, 1, 2, 6, ..., or even negative integers. c. In another example, the gradient can be calculated as a weighted sum of Nb samples before the current sample and Na samples after the current sample: i. The weight (i.e., wp) can be any integer, such as -6, 0, 2, 7… or any real number, such as -6.3, -0.77, 0.1, 3.0,… ii. The weights used to calculate the horizontal and vertical gradients can be different from each other. (i) Alternatively, the weights used to calculate the horizontal and vertical gradients can be the same. iii. Weights can be transmitted from the encoder to the decoder via signals. iv. Weights can be derived using the decoded information. v.Nb and Na can be any integer, such as 0, 3, 10, ... vi. Nb and Na can be used differently for calculating the gradients in the horizontal and vertical directions. (i) Alternatively, they can be the same for calculating the gradients in both the horizontal and vertical directions. vii. In one example, (i) Alternatively, additionally, the variable offset can be set to 0 or (1<<(shift-1)). 2. A complete linear equation formula is proposed that can be used to derive the final MV refinement. a. In one example, after calculating all gradients, s1, s2, s3, s5, and s6 are calculated as explained above: ∑Gx.Gx*vx+∑Gx.Gy*vy=∑dI.Gx.→s1*vx+s2*vy=s3, ∑Gx.Gy*vx+∑Gy.Gy*vy=∑dI.Gy→s2*vx+s5*vy=s6. i. In one example, to derive the final MV of an M*N block, samples from the (M+K1)*(N+K2) region surrounding the original block can be involved. For example, K1 and K2 can be any integers, such as 0, 2, 4, 7, 10, ... b. In one example, after calculating all s1, s2, s3, s5, and s6, the determinant values ​​D, Dx, and Dy are calculated as follows: D=(s1>>shTem)*(s5>>shTem)-(s2>>shTem)*(s2>>shTem), Dx=(s3>>shTem)*(s5>>shTem)-(s6>>shTem)*(s2>>shTem), Dy=(s1>>shTem)*(s6>>shTem)-(s3>>shTem)*(s2>>shTem). i. In one example, shTem can be any integer, such as 0, 1, 3, ... c. In one example, after calculating D, Dx, and Dy; vx and vy can be derived as: vx = Dx / D and vy = Dy / D. i. In another example, if abs(D) is less than a predefined threshold C, then vx and vy are set to zero. C can be any non-negative number, such as 0, 10, 17, ... d. In one example, shifting and limiting of any amount can be involved to derive the final vx and vy. i. In one example, the numerator and / or denominator may have an additional shift, such that they are shifted left by K overall, resulting in higher precision for the final derived vx and vy. K can be any integer, such as 0, 1, 3, 4, 6, ... ii. In one example, these shifts can occur in any order, such as having a shift at the beginning, and / or having a shift for intermediate variables and / or having a shift on the final MV. iii. In one example, the final vx and vy can be bounded between -B and B, where B can be any integer, such as 2, 10, 17, 32, 100, 156, 725, ... e. In one example, the final vx and vy can be multiplied (or similarly divided) by a number before being used in the motion compensation process. i. In one example, vx and vy can be multiplied by R, where R is any real number, such as 1.25, 2, 3.1, 4, ... ii. In another example, vx and vy can be divided by R, where R is any real number, such as 1.25, 2, 3.1, 4, ... iii. In one example, the values ​​of the numbers to be multiplied (or divided) by the final vx and vy can be different for vx and vy. iv. In one example, the value of the number to be multiplied (or divided) by the final vx, vy can depend on the block size, sequence resolution, block characteristics, etc. 3. It is proposed that the solutions to some linear equations can be used to derive the final MV refinement. a. In one example, after calculating all gradients, s1, s2, s3, s5, and s6 are calculated as explained above: ∑Gx.Gx*vx+∑Gx.Gy*vy=∑dI.Gx.→s1*vx+s2*vy=s3, ∑Gx.Gy*vx+∑Gy.Gy*vy=∑dI.Gy→s2*vx+s5*vy=s6. b. In one example, after calculating all s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as: vx = s3 / s1, vy=(s6–s2*vx) / s5. c. In another example, after calculating vx similarly to the above, a portion of vx can be put into the second formula to derive vy. i. In one example, vy can be derived as vy = (s6 – s2 * vx / T) / s5, where T can be any real number, such as 1.1, 2, 4, ... d. In one example, after calculating all s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as: Assuming vx is zero: vy = s6 / s5, Insert vy into the first formula: vx=(s3–s2*vy) / s1. e. In another example, after calculating vy similar to the above, a portion of vy can be inserted into the second formula to derive vx. i. In one example, vx can be derived as vx = (s3 – s2 * vy / T) / s1, where T can be any real number, such as 1.1, 2, 4, ... f. In one example, after calculating all s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as: Assume vy is zero: vx = s3 / s1. Assuming vx is zero: vy = s6 / s5. 4. A simplified solution is proposed that can be used to derive the final MV refinement. a. In one example, the methods for VVC BDOF explained in the background section can be used to derive approximate versions of s1, s2, s3, s5, and s6. b. In one example, after calculating approximate versions of s1, s2, s3, s5, and s6, the determinant values ​​D, Dx, and Dy are calculated as follows: D=(s1>>shTem)*(s5>>shTem)–(s2>>shTem)*(s2>>shTem), Dx=(s3>>shTem)*(s5>>shTem)–(s6>>shTem)*(s2>>shTem), Dy=(s1>>shTem)*(s6>>shTem)–(s3>>shTem)*(s2>>shTem). i. In one example, after calculating D, Dx, and Dy; vx and vy can be derived as: vx = Dx / D and vy = Dy / D. ii. In one example, shTem can be any integer, such as 0, 1, 3, ... c. In one example, after calculating approximate versions of s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as follows: Assuming vy is zero: vx = s3 / s1, Substituting vx into the second formula: vy=(s6–s2*vx) / s5. i. Alternatively, the modified vx can be inserted into the second formula: vy = (s6 – s2 * vx / T) / s5, where T can be any real number, such as 1.1, 2, 4, ... ii. Alternatively, firstly, vx can be assumed to be zero, and vy can be derived, then vy or its scaled version can be inserted into the first equation, and vx can be derived. 5. Any combination of the methods explained above can be used to derive the final MV refinement. a. In one example, any combination of the methods (2, 3, and 4) explained above can be combined and used together. Derivation of BDOF sample point adjustment parameters 6. Any of the methods for BDOF MV refinement explained above can also be used to derive BDOF sample adjustment parameters. a. In one example, after calculating all gradients, s1, s2, s3, s5, and s6 are calculated as explained above: ∑Gx.Gx*vx+∑Gx.Gy*vy=∑dI.Gx.→s1*vx+s2*vy=s3, ∑Gx.Gy*vx+∑Gy.Gy*vy=∑dI.Gy→s2*vx+s5*vy=s6. i. In one example, samples within a KxK region surrounding a sample point can be included in the derivation. K can be any integer, such as 1, 3, 4, 5, 7, 10, ... b. In one example, after calculating s1, s2, s3, s5, and s6, the determinant values ​​D, Dx, and Dy are calculated as follows: D=(s1>>shTem)*(s5>>shTem)-(s2>>shTem)*(s2>>shTem), Dx=(s3>>shTem)*(s5>>shTem)-(s6>>shTem)*(s2>>shTem), Dy=(s1>>shTem)*(s6>>shTem)-(s3>>shTem)*(s2>>shTem). i.shTem can be any integer, such as 0, 1, 3, ... ii. In one example, after calculating D, Dx, and Dy; vx and vy can be derived as: vx = Dx / D and vy = Dy / D. iii. In another example, if abs(D) is less than a predefined threshold C, then vx and vy are set to zero. C can be any non-negative number, such as 0, 10, 17, ... c. In one example, after calculating s1, s2, s3, s5, and s6, approximate versions of vx and vy can be calculated as: vx = s3 / s1, vy=(s6–s2*vx) / s5. i. Alternatively, the modified vx can be placed into the second formula: vy = (s6 – s2 * vx / T) / s5, where T can be any real number, such as 1.1, 2, 4, ... ii. Alternatively, firstly, vx can be assumed to be zero, and vy can be derived. Then, vy or its scaled version can be substituted into the first equation, and vx can be derived. d. In one example, the methods for VVC BDOF explained in the background section can be used to derive approximate versions of s1, s2, s3, s5, and s6. e. In one example, the final vx and vy can be multiplied (or divided or shifted) by a number before being used in the sample adjustment process. i. In one example, vx and vy can be multiplied by R, where R is any real number, such as 1.25, 2, 3.1, 4, ... ii. In another example, vx and vy can be divided by R, where R is any real number, such as 1.25, 2, 3.1, 4, ... iii. In one example, the values ​​of the numbers to be multiplied (or divided) by the final vx and vy can be different for vx and vy. iv. In one example, the value of the number to be multiplied (or divided) by the final vx, vy can depend on the block size, sequence resolution, block characteristics, position in the block, etc. Regarding the application of weights in parameter derivation 7. It was proposed that any weights can be applied before adding BDOF intermediate parameters for MV refinement. a. In one example, during the addition of parameters to obtain s1, s2, s3, s5 and s6, all values ​​are added with similar weights (1) within the target region Ω (the M_ext*N_ext region surrounding the current block). b. In another example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, these values ​​are added after being multiplied by a predefined weight depending on their position in the extended block (target region Ω) within the target region Ω. c. In one example, these predefined weights are defined as follows: w=(x>=(width / 2)?width-x:x+1)*(y>=(height / 2)?height-y:y+1) For x ranging from 0 to width-1 and y ranging from 0 to height-1. Width and height represent the width and height of the target area. d. In another example, these predefined weights can be generated using some known probability distribution, such as a Gaussian distribution with any value of standard deviation (σ = 1, 1.5, 4 or any other real number) and central location. i. In one example, such as Figure 7 As shown, these weights are generated using a Gaussian distribution with σ = 2.5 over a 12x12 region. ii. In one example, such as Figure 8 As shown, these weights are generated using a Gaussian distribution with σ=4 over a 12x12 region. e. In another example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, these values ​​are added after being shifted using predefined values ​​that depend on their position in the extended block (target region Ω) within the target region Ω. f. In one example, the weight matrix can be represented as a left (or right) shift matrix, and depending on the matrix entries, the data is shifted (left or right) before summing. 8. In one example, different weights can be applied based on block size, block shape, block characteristics, sequence resolution, etc. i. Alternatively, no weights can be applied based on block size, block shape, block characteristics, sequence resolution, etc. ii. The weight matrix can be explicitly encoded or decoded in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or Strip Header (SH). 9. It was proposed that any weights can be applied before adding BDOF intermediate parameters for sample adjustment. a. In one example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, all values ​​are added with similar weights (1) within the target region Ω (the K1*K2 region surrounding the current sample point). K1 and K2 can be any integers, such as 1, 2, 3, 5, 8, ... b. In another example, during the addition of parameters to obtain s1, s2, s3, s5, and s6, these values ​​are added within the target region Ω (the K1*K2 region surrounding the current sample point) after being multiplied by a predefined weight that depends on their position in the extended block (target region Ω). c. In one example, these predefined weights are defined as follows: w=(x>=(K1 / 2)?K1-x:x+1)*(y>=(K2 / 2)?K2-y:y+1), For x from 0 to K1-1 and y from 0 to K2-1. K1 and K2 represent the width and height of the target area. d. In another example, these predefined weights can be generated using some known probability distribution, such as a Gaussian distribution with any value of standard deviation (σ = 1, 1.5, 2, 4 or any other real number) and any central location. i. In one example, such as Figure 9 As shown, these weights are generated using a Gaussian distribution with σ=1 for a 5x5 region. ii. In one example, such as Figure 10 As shown, these weights are generated using a Gaussian distribution with σ=2 over a 5x5 region. e. In one example, the weight matrix can be represented as a left (or right) shift matrix, and depending on the matrix entries, the data is shifted (left or right) before summing. f. In one example, different weights can be applied based on block size, block shape, block characteristics, sequence resolution, etc. i. Alternatively, no weights can be applied based on block size, block shape, block characteristics, sequence resolution, etc. Regarding the application of filters for final MV refinement or sample adjustment 10. It is proposed that any type of filter can be applied to the final derived MV refinement (vx and vy). Some examples are provided in... Figure 11 It is depicted in the middle. a. In one example, any smoothing filter of any shape can be applied to all MVs derived by BDOF for each sub-block. b. In one example, all MVs inside the PU can be used during filter application. c. In another example, during the filter application, only the second round MV with similar DMVR can be used for those MVs. d. In one example, a shape filter with any weights can be applied to the MV. i. In one example, the weight of the center can be 8, and the weights of the four edges can be 1. ii. In one example, the weight of the center can be 4, and the weights of the four edges can be 1. iii. In one example, the center weight can be 4, and the weights of the four edges can be 2. iv. In one example, the center weight can be 4, and the weights of the four edges can be 3. v. In one example, the weight of the center can be 1, and the weights of the four edges can be 1. 11. It is proposed that any type of filter can be applied to the final derived BDOF sample MV adjustment or final sample adjustment. Some examples are provided in... Figure 11 It is depicted in the middle. a. In one example, the filter is applied to all (vx, vy) or finally adjusted within the sub-block. b. In one example, a shape filter with any weights can be applied to (vx, vy) or ultimately adjusted. i. In one example, the weight of the center can be 8, and the weights of the four edges can be 1. ii. In one example, the weight of the center can be 4, and the weights of the four edges can be 1. iii. In one example, the center weight can be 4, and the weights of the four edges can be 2. iv. In one example, the center weight can be 4, and the weights of the four edges can be 3. v. In one example, the weight of the center can be 1, and the weights of the four edges can be 1. Conditions for applying BDOF 12. It proposes conditions that can exist when applying BDOF MV refinement or BDOF sample adjustment. a. In one example, the conditions for applying BDOF MV refinement can be similar to the conditions for applying BDOF sample adjustment. b. In another example, the conditions for applying BDOF MV refinement may differ from the conditions for applying BDOF sample adjustment. For instance, BDOF MV refinement may be applied to bidirectional predictive codecs with unequal weights, while BDOF sample adjustment may be applied only to bidirectional predictive codecs with equal weights. 13. It is proposed that the cost for evaluating BDOF conditions can depend on the cost between two reference image blocks. a. In one example, different cost functions can be used to derive the cost. i. In one example, the cost could be the sum of absolute differences (SAD) between two reference image patches. ii. In one example, the cost could be the sum of absolute transformation differences (SATD) between two reference image patches or any other cost measure. iii. In one example, the cost could be the sum of absolute differences (MR-SAD) between two reference image patches based on mean removal. iv. In one example, the cost could be a weighted average of SAD / MR-SAD and SATD between two reference image patches. v. In one example, the cost function between two reference image patches could be: (i) Sum of absolute differences (SAD) / SAD after removing the mean (MR-SAD); (ii) Sum of absolute transformation differences (SATD) / SATD after mean removal (MR-SATD); (iii) Sum of squared differences (SSD) / SSD after mean removal (MR-SSD); (iv)SSE / MR-SSE; (v) Weighted SAD / Weighted MR-SAD; (vi) Weighted SATD / Weighted MR-SATD; (vii) Weighted SSD / Weighted MR-SSD; (viii) Weighted SSE / Weighted MR-SSE; (ix) Gradient information. Regarding BDOF MV refined sub-block size 14. It is proposed that any sub-block size, depending on the conditions, can be used as the BDOF MV refined sub-block size. a. In one example, the sub-block size can be a fixed size, such as N x M, where N and M can be any positive integer, such as 1, 2, 3, 4, 5, 8, 12, 32, ... b. In another example, the sub-block size can depend on the current PU or CU size. For example, for a block size WxH, a sub-block size W1xH1 can be used, where W1 and H1 depend on W and H, and can be any positive integer. i. In one example, for a block with a number of samples between C_i and C_(i+1) (i.e., width multiplied by height (W*H)), the sub-block size W_ix H_i can be used. C_i can be any non-negative number, such as 0, 4, 20, 128, 256, 951, 2048, 4100, ..., and W_ix H_i can be any positive integer pair, such as 2x2, 4x4, 8x4, 4x8, 8x8, 16x16, 19x15, ... ii. In one example, for a block having a width W between Cw_i and Cw_(i+1) and a height H between Ch_j and Ch_(j+1), the sub-block size W_i x H_j can be used. Cw_i and Ch_j can be any non-negative numbers, such as 0, 4, 20, 128, 256, 951, 2048, 4100, ..., and W_i x H_j can be any positive integer pair, such as 2x2, 4x4, 8x4, 4x8, 8x8, 16x16, 19x15, ... c. In one example, the sub-block size may depend on the color components and / or color format. d. In one example, the size of a sub-block can depend on the encoded / decoded information of the current block. i. In one example, the encoded and decoded information is residual information. ii. In one example, the encoded / decoded information is the encoding / decoding tool applied to the current block. e. In one example, the sub-block size can depend on the information of the predicted block. f. In one example, the size of the sub-block may depend on the characteristics of the reference image. i. In one example, the sub-block size can be determined by the similarity of two predicted values ​​from two reference images. If the two predicted values ​​are similar, such as having a smaller SAD between them, a larger sub-block size can be applied; otherwise, a smaller sub-block size can be applied. ii. In one example, the sub-block size can be determined by the distribution of the difference between two predictions. Those sub-blocks with difference energy (such as SAD or SSE) can be merged into larger cells for MV refinement, thereby reducing computational complexity. g. In one example, the sub-block size may depend on the temporal gradients of two reference blocks. i. In one example, any cost function (such as SAD) can be used to compute the gradient (or difference) between two reference blocks. h. In one example, the spatial gradient of a reference block can be used to determine the size of a sub-block. i. In one example, the sub-block size can depend on the quantization parameter (qp) value. i. In one example, for qp less than X, the sub-block size W_X x H_X can be used. ii. In one example, for qp greater than X, the sub-block size W_X x H_X can be used. iii. In one example, for qp X, the sub-block size W_X x H_X can be used. iv. In one example, X can be any non-negative integer, such as 10, 22, 27, 32, 37, 42, ..., and W_X and H_X can be any positive integer, such as 1, 2, 3, 4, 8, 10, ... v. In one example, qp can be the qp of the current CU, the qp of the current stripe, or the qp of the entire sequence. vi. In one example, the decision regarding the sub-block size can be made by the encoder, and it may or may not be signaled to the decoder. Similarly, it can be made by the decoder. vii. In one example, increasing or decreasing the sub-block size based on qp can be determined by either the encoder or the decoder. j. In one example, the sub-block size can depend on the prediction type. k. In one example, the sub-block size may depend on the DMVR first and / or second stage adjustment values. l. In one example, the sub-block size can depend on the sequence resolution. m. In one example, the sub-block size can depend on the encoding / decoding tool applied to the current block. In one example, the sub-block size can depend on the temporal layer. i. In one example, for a time-domain layer between Ti and Tj, the sub-block size W_ij x H_ij can be used. Ti and Tj can be any non-negative integers, such as 0, 1, 3, 4, ..., and W_ij and H_ij can be any positive integers, such as 2, 4, 6, 16, ... o. In one example, the sub-block size can be a function of all or some of the parameters mentioned above. p. In the above example, the sub-block size of the luminance and / or chroma block can be determined based on the above example. i. Alternatively, the sub-block size of the chroma block can be derived based on the sub-block size of the luma block and the color format and / or whether individual planar codecs are enabled. Regarding asymmetric BDOF 15. It was proposed that the MV adjustment for the first list and the second list does not have to be symmetrical. a. In one example, the MV refinement for reference image 0 could be (vx0, vy0), and the MV refinement for reference image 1 could be (-vx1, -vy1), where vx0, vy0, vx1, and vy1 can be any real numbers or integers. They may or may not have a relationship together. b. In one example, the general equations used to derive vx0, vy0, vx1, and vy1 can be written as the following four equations: ∑Gx0.Gx0*vx0+∑Gx1.Gx0*vx1+∑Gy0.Gx0*vy0+∑Gy1.Gx0*vy1=∑dI.Gx0. ∑Gx0.Gx1*vx0+∑Gx1.Gx1*vx1+∑Gy0.Gx1*vy0+∑Gy1.Gx1*vy1=∑dI.Gx1. ΣGx0.Gy0*vx0+ΣGx1.Gy0*vx1+ΣGy0.Gy0*vy0+ΣGy1.Gy0*vy1=ΣdI.Gy0. ΣGx0.Gy1*vx0+ΣGx1.Gy1*vx1+ΣGy0.Gy1*vy0+ΣGy1.Gy1*vy1=ΣdI.Gy1. Where Gx0, Gx1, Gy0, and Gy1 represent the horizontal gradient of reference image 0, the horizontal gradient of reference image 1, the vertical gradient of reference image 0, and the vertical gradient of reference image 1, respectively. dI represents the difference between the two reference images. The summation (∑) is performed within a predefined region, which can be an NxM block around the current sample point (used for sample-adjusted BDOF) or an NxM block around the current predicted sub-block (used for MV-refined BDOF). c. Alternatively, in matrix format they can be written as: The parameters in the matrix format match the parameters in the equation. d. In one example, the general formula for determinants can be used to solve the above linear equations. e. In one example, Gaussian elimination can be used to solve the linear equation above. f. In one example, any other method (including matrix factorization) can be used to solve the above linear equation. g. In one example, vx1 can be equal to k * vx0, and vy1 can be equal to k * vy0, where k can be any real number or integer, such as -0.3, 0, 0.1, 2, 3, ... i. In one example, any nonlinear method can be used to derive and solve nonlinear equations. h. In one example, any weighted sums described in the previous section can be used for summation. i. In one example, asymmetric BDOF can be applied to BDOF MV refinement and BDOF sample adjustment. j. In one example, asymmetric BDOF can be applied only to BDOF MV refinement. k. In one example, asymmetric BDOF can be applied only to BDOF sample adjustment. l. In one example, whether and / or how to apply asymmetric BDOF may depend on the POC or at least one POC distance. i. In one example, whether and / or how to apply asymmetric BDOF may depend on |POC_ref0-POC_cur| and / or |POC_ref1-POC_cur|, where POC_ref0 and POC_ref1 represent the POCs of two reference images, and POC_cur is the POC of the current image. m. In one example, whether and / or how to apply asymmetric BDOF can depend on the BCW weights. n. In one example, whether and / or how asymmetric BDOF is applied may depend on at least one template of the current block. i. Furthermore, whether and / or how asymmetric BDOF is applied may depend on at least one reference template of the template of the current block. Conditions for applying BDOF and its combination with other tools 16. It was proposed that BDOF and / or asymmetric BDOF (MV refinement or sample adjustment or both) can be used in combination with other tools or excluded from the use of other tools. a. In one example, BDOF can be applied to blocks encoded and decoded with unequal BCW weights. i. In one example, BDOF can be applied using BCW weights from a predefined set (such as {3}, or {3,5} or {-1,3}). b. In one example, BDOF can be applied to blocks of two reference images on the same side of the current frame. c. In one example, BDOF can be applied to a block of a reference image on the opposite side of the current frame. i. In one example, they can be at the same distance from the current frame. ii. In another example, they may have a different distance from the current frame. d. In one example, BDOF can be applied in combination with LIC. i. Alternatively, if the block uses a LIC, it can be turned off. e. In one example, BDOF can be applied in combination with OBMC. i. Alternatively, if the block uses OBMC, it can be turned off. f. In one example, BDOF can be applied in combination with CCIP. i. Alternatively, if the block uses CIIP, it can be turned off. g. In one example, BDOF can be applied in combination with SMVD. i. Alternatively, if the block uses SMVD, it can be turned off. h. In one example, BDOF DMVR or BDOF samples can be controlled separately. For example, BDOF DMVR is applied, but BDOF samples are not applied. i. Control can be at the sub-PU level, or the CU level, or the CTU level. Regarding merging sub-blocks with similar MVs before applying BDOF DMVR or BDOF samples. 17. It was proposed that multiple sub-blocks can share the same MV after / before an MV refinement process such as BDOF or DMVR. a. In one example, it is proposed to check the similar MVs of neighboring sub-blocks and combine them before applying the BDOF DMVR or BDOF sampling process. b. In one example, multiple sub - blocks sharing the same MV can perform motion compensation as a whole. c. In one example, N1 adjacent sub - blocks in a row with similar MVs can be merged. i. N1 can be any integer, such as 2, 3, 4, 10, … d. In one example, N2 adjacent sub - blocks in a column with similar MVs can be merged. i. N2 can be any integer, such as 2, 3, 4, 10, … e. In one example, all adjacent sub - blocks in a row with similar MVs up to the r - th round (first, second, …) of the DMVR sub - PU boundary can be merged. These sub - PU boundaries can occur every K - th pixel, where K can be any integer, such as 8, 16, 19, 32, … f. In one example, all adjacent sub - blocks in a column with similar MVs up to the r - th round (first, second, …) of the DMVR sub - PU boundary can be merged. These sub - PU boundaries can occur every K - th pixel, where K can be any integer, such as 8, 16, 19, 32, … g. In one example, the decision of row - based or column - based merging can depend on the width (W) and / or height (H) of the block size. i. In one example, if W >= H, a row - based merging method can be used. Otherwise, a column - based merging method can be used. ii. Alternatively, if W < H, a row - based merging method can be used. Otherwise, a column - based merging method can be used. h. In one example, for each sub - block within a larger M×N sub - block, the MV checking and merging process can be applied. M and N can be any integer, such as 4, 5, 10, 16, 32, … i. In one example, all 4x4 (or 8x8) sub - blocks within 16x16 or 8x8 or 16x8 or 32x32 can be merged. ii. In one example, these M and N can be variable or can be fixed. i. In one example, all sub - blocks within a PU or CU can be merged. j. In one example, in all the above scenarios, sub - blocks with almost similar MVs (and not necessarily the same) can also be merged. The almost - similar criterion can be defined as whether the first - order or second - order Euclidean distance of the MVs is less than a threshold. The merged MV can be the average, mode, … of all MVs. Or it can be the MV at the center or upper - left or lower - right or other positions. k. Alternatively, additionally, when two sub-blocks have similar motions, the motion information of one or both sub-blocks can be modified before use so that the two sub-blocks will use the same motion to perform the operation performed earlier. l. In one example, two motions are considered similar motions when they use the same reference image and their motion signatures (MVs) are similar (e.g., the MV difference is less than a threshold). m. The above examples can be applied to each prediction direction. i. Alternative locations, which can be applied together for all predicted directions. Regarding parallelization, code optimization, and applying shifts to BDOF. 18. A method for computing BDOF parameters using parallelization was proposed. a. In one example, the SIMD implementations of all relevant functions can be used. b. In one example, during the calculation of the sum of parameters for BDOF samples (or DMVR), the K sums of the sample sums at one iteration can be derived. K can be any integer, such as 2, 3, 4, 5, 8, ... c. In one example, the weighted sum can be implemented as a proper left shift. i. In one example, the multiplication using weight w_i can be replaced by a left shift of log2(1+w_i). 19. Several code optimizations for BDOF are proposed. a. In one example, the BDOF DMVR parameter will not always be calculated. Its calculation can be delayed and conditional if actually needed. i. In one example, it is only calculated if no BDOF samples are applied. b. In one example, the BDOF sample parameters will not be calculated continuously. If necessary, their calculation can be delayed and conditional. i. In one example, it is only calculated if the BDOF DMVR phase does not result in an MV update. ii. In one example, it is only calculated if BDOF DMVR is not applied. c. In one example, for a block where BDOF DMVR is off, different sub-block sizes can be used to check the application conditions of BDOF samples. i. In one example, the new sub-block size can be larger or smaller than the BDOF DMVR sub-block size. It can be M×N, where M and N can be any integer. Here are some examples for M×N sizes: 2×2, 4×4, 4×8, 8×4, 8×8, ... 20. It is proposed to add shifts (right shift or left shift) at different stages of BDOF parameter derivation in order to remove noise, avoid overflow, reduce data bandwidth, or improve the accuracy of the derived parameters. a. Shift operations can be right shift or left shift. b. Offsets can be added before and / or after the shift operation. c. The result of the shift can be limited to a range. d. In one example, the data can be shifted by Shift1 before / after the gradient is calculated. Shift1 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). e. In one example, the data can be shifted by Shift2 before / after calculating the difference in brightness. Shift2 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). f. In one example, the data can be shifted by Shift3 before / after multiplying by the gradient or brightness difference. Shift3 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). g. In one example, data can be shifted by Shift4 before / after the sum of the calculated parameters. Shift4 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). h. In one example, data can be shifted by Shift5 before / after calculating the determinant (multiplication of the final parameters). Shift5 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). i. In one example, the data can be shifted by Shift6 before / after the division of the determinant to obtain the final scaled MV adjustment. Shift6 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). j. In one example, the data can be shifted by Shift7 before / after the calculation of sample point adjustments. Shift7 can be any right or left shift with any integer value (such as 0, 1, 3, 4, 6, ...). k. In one example, the shift parameter can depend on the bit depth. General aspects 21. In one example, the division operation disclosed in this document may be replaced by a non-division operation, which may share the same or similar logic as the division replacement logic in CCLM or CCCM. 22. Whether and / or how the methods described above are applied may depend on the encoded / decoded information. a. In one example, the encoded / decoded information may include block size and / or temporal layer and / or strip / image type, color components, etc. 23. Whether and / or how the methods described above are applied can be indicated in the bitstream. a. Indicators for enabling / disabling or specifying which method to apply can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header. b. Indicators for enabling / disabling or indicating which method to apply can be transmitted via signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-pictures / other types of areas containing more than one sample or pixel.

[0069] Further details of embodiments of this disclosure relating to bidirectional optical flow (BDOF) processes will now be described. The embodiments of this disclosure should be considered as examples illustrating general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments may be applied individually or in any combination.

[0070] As used herein, the term "block" can refer to a color component, sub-picture, picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), sub-block of a video block, sub-region within a video block, video processing unit comprising multiple samples / pixels, etc. A block can be rectangular or non-rectangular.

[0071] Figure 12 A flowchart of a method 1200 for video processing according to some embodiments of the present disclosure is shown. Method 1200 can be implemented during the conversion between a current video block and a video bitstream. Figure 12 As shown, method 1200 begins at 1202, wherein for the current sub-block of the current video block, at least one sub-block is determined from among the multiple sub-blocks of the current video block.

[0072] In the first case, the motion vector (MV) of each sub-block in at least one sub-block satisfies the following condition with the MV of the current sub-block: the MV of each sub-block in at least one sub-block is the same as the MV of the current sub-block. In this case, only the sub-block(s) having the same MV as the current sub-block are selected.

[0073] In the second case, the MV of each sub-block in at least one sub-block satisfies the following condition with the MV of the current sub-block: the difference metric between the MV of each sub-block in at least one sub-block and the MV of the current sub-block is less than a threshold. In this case, in addition to sub-block(s) having the same MV as the current sub-block, sub-block(s) having similar MVs to the current sub-block can also be selected. By way of example rather than limitation, the difference metric can include first-order Euclidean distance, second-order Euclidean distance, etc.

[0074] At position 1204, the combined sub-block is obtained by combining at least one sub-block with the current sub-block. For example, at least one sub-block and the current sub-block can be merged to obtain a new sub-block with a larger size. Furthermore, motion information for the combined sub-block can be determined based on the motion information of at least one sub-block and the motion information of the current sub-block.

[0075] For example, in the first case described above, the MV of the combined sub-block can be the same as the MV of the current sub-block. In the second case described above, the MV of the combined sub-block can be determined based on the MV of the current sub-block and the MV of at least one sub-block. In one example embodiment, the MV of the combined sub-block can be determined as the average MV determined by averaging the MV of the current sub-block and the MV of at least one sub-block. In another example embodiment, the MV of the combined sub-block can be determined as the most frequently occurring MV among the MV of the current sub-block and the MV of at least one sub-block. In yet another example embodiment, the MV of the combined sub-block can be determined as the MV of the sub-block at a predetermined position. By way of example and not limitation, the predetermined position can be the center position, the upper left position, the lower right position, etc. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0076] At 1206, the conversion is performed based on applying a bidirectional optical flow (BDOF) process to the combined sub-block. For example, the BDOF process can be performed on the combined sub-block to obtain at least one offset. That is, the BDOF process is applied as a whole to at least one sub-block and the current sub-block, rather than being performed on each of the at least one sub-block and the current sub-block. Furthermore, the conversion can be performed based on at least one offset.

[0077] In some embodiments, the BDOF process can be a BDOF process for MV refinement. In this process, for example, at least one offset can be determined for refining the MV of a sub-block. Alternatively, the BDOF process can be a BDOF process for sample adjustment, also known as sample-based BDOF. In this process, for example, at least one offset can be determined for adjusting one or more samples in a sub-block. It should be noted that the BDOF process can be performed in several rounds. For example, the BDOF process can be performed as one or more rounds in the MV refinement process. By way of example, and not limitation, one or more rounds of decoder-side motion vector refinement (DMVR) can be performed during the MV refinement process. Then, the BDOF process for MV refinement can be performed in several rounds. Furthermore, one or more rounds of BDOF process for sample adjustment can be performed. The proposed method can be applied prior to at least one round of the BDOF process.

[0078] In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0079] Given the above, more than one sub-block with the same or similar MV is combined, and the BDOF process is applied to the combined sub-blocks, rather than being applied individually to each sub-block within the larger group. Compared to conventional solutions where the BDOF process is applied individually to sub-blocks, the proposed method can advantageously perform the BDOF process more efficiently. In this way, encoding and decoding efficiency can be improved.

[0080] In some embodiments, the multiple sub-blocks may include sub-blocks adjacent to the current sub-block. For example, adjacent and non-adjacent sub-blocks of the current sub-block may be examined for sub-block(s)(s) having the same MV as the current sub-block or sub-block(s) having similar MVs to the current sub-block. In some embodiments, motion compensation may be applied to combined sub-blocks. For example, for the combined sub-blocks as a whole, prediction may be determined based on motion information determined for the combined sub-blocks.

[0081] In some embodiments, at least one sub-block may include N neighboring sub-blocks in the same row as the current sub-block, and N may be an integer. For example, N1 neighboring sub-blocks in a row with similar MV may be merged, and N1 may be any integer, such as 2, 3, 4, 10, etc.

[0082] In some embodiments, at least one sub-block may include M neighboring sub-blocks in the same column as the current sub-block, and M may be an integer. For example, M1 neighboring sub-blocks in a column with similar values ​​MV may be merged. M1 may be any integer, such as 2, 3, 4, 10, etc.

[0083] In some embodiments, at least one sub-block may be within the sub-prediction unit (PU) boundary of the i-th round of decoder-side motion vector refinement (DMVR) process, and i may be an integer, such as 1, 2, etc. For example, only sub-blocks within the same sub-PU boundary are allowed to be combined. These sub-PU boundaries may appear every K-th pixel, where K may be any integer, such as 8, 16, 19, 32, etc.

[0084] In some embodiments, whether at least one sub-block is located in the same row or column as the current sub-block can depend on the size of the current video block. For example, if the width of the current video block is greater than or equal to its height, then at least one sub-block can be located in the same row as the current sub-block. That is, row-based combination is implemented, meaning only sub-blocks in the same row are allowed to be combined. If the width is less than the height, then at least one sub-block can be located in the same column as the current sub-block. That is, column-based combination is implemented, meaning only sub-blocks in the same column are allowed to be combined.

[0085] Alternatively, if the height of the current video block is greater than its width, at least one child block can be located in the same row as the current child block. That is, row-based combination is implemented. If the height is less than or equal to the width, at least one child block can be located in the same column as the current child block. That is, column-based combination is implemented.

[0086] In some embodiments, multiple sub-blocks are included in a sub-segment of the current video block, and the size of the sub-segment can be larger than the size of each of the multiple sub-blocks. For example, an MV checking and merging process can be applied to each sub-block within a larger M×N sub-block. M and N can be any integer, such as 4, 5, 10, 16, 32, etc. The M×N sub-block can be considered a sub-segment of the current video block. The size of the sub-segment can be variable or fixed. For example, the size of each of the multiple sub-blocks can be 4×4, and the size of the sub-segment can be one of 16×16, 8×8, 16×8, or 32×32. Alternatively, the size of each of the multiple sub-blocks can be 8×8, and the size of the sub-segment can be one of 16×16, 16×8, or 32×32. It should be understood that the specific values ​​listed herein are intended to be exemplary and not to limit the scope of this disclosure.

[0087] In some embodiments, at least one sub-block may be allowed to include all sub-blocks of the current video block. For example, all sub-blocks within a PU or CU may be allowed to be merged.

[0088] In some embodiments, if the difference metric between the MV of the first sub-block and the MV of the current sub-block is less than a threshold, and the reference image for the first sub-block can be the same as the reference image for the current sub-block, then the motion information of at least one sub-block in the first sub-block or the current sub-block can be modified so that the same motion information can be used for the first sub-block and the current sub-block in subsequent processes.

[0089] In some embodiments, the proposed method may be applied to at least one of two prediction directions. For example, the two prediction directions include a prediction direction corresponding to a first reference list (such as list L0) and a prediction direction corresponding to a second reference list (such as list L1).

[0090] In some embodiments, the BDOF parameters used in the BDOF process can be determined in parallel. For example, multiple functions related to the BDOF process can be implemented based on a Single Instruction Multiple Data (SIMD) scheme. SIMD is a parallel processing technique that enables the execution of a single instruction on multiple data elements simultaneously. In this way, hardware-level parallelism can be utilized to efficiently perform computations on arrays or vectors of data. By way of example, and not limitation, multiple summation operations can be performed at one iteration to determine the sum of the BDOF parameters. For example, instead of performing summation operations one by one, the summation operations can be performed in parallel, such as in groups of four. It should be understood that any other suitable parallel processing technique can also be implemented.

[0091] In some embodiments, multiplication can be implemented using a left shift operation. For example, multiplication using weight w can be replaced by a left shift of log2(1+w). In this way, multiplication can be implemented more efficiently.

[0092] In some embodiments, one or more parameters used in the BDOF process for MV refinement can be determined based on conditions. For example, one or more parameters can be determined if they will be used. Additionally or alternatively, one or more parameters can be determined if the BDOF process for sample adjustment is not applied.

[0093] Similarly, one or more parameters used in the BDOF procedure for sample adjustment can be determined based on conditions. For example, if one or more parameters will be used, then one or more parameters can be determined. Additionally or alternatively, if the BDOF procedure for MV refinement is not applied, then one or more parameters can be determined. Additionally or alternatively, if the result of the BDOF procedure for MV refinement indicates that MV updates are not needed, then one or more parameters can be determined.

[0094] In this way, the parameters used in the BDOF process can be determined only when they are actually needed, thereby saving computational resources and improving encoding and decoding efficiency.

[0095] In some embodiments, the first sub-block size used to determine whether the BDOF process for MV refinement should be applied may be different from the second sub-block size used to determine whether the BDOF process for sample adjustment should be applied. In one example, the second sub-block size may be larger than the first sub-block size. Alternatively, the second sub-block size may be smaller than the first sub-block size.

[0096] In some embodiments, the shift operation may be performed at one or more stages used in determining at least one BDOF parameter in the BDOF process. In one example, the shift operation may include a right shift operation. A right shift operation may be performed to remove noise, prevent overflow, or reduce data bandwidth. In another example, the shift operation may include a left shift operation. A left shift operation may be performed to improve the accuracy of the derived parameters.

[0097] Furthermore, an offset can be added to the operands and / or the result of the shift operation. Additionally or alternatively, the result of the shift operation can be bounded to a range of values. By way of example, this range may include upper and / or lower limits, which are predetermined or indicated in the bitstream.

[0098] In some embodiments, the shift operation may be performed before, during, or after any stage of parameter derivation. By way of example, the shift operation may be performed on at least one of the following: an input for determining the gradient, an intermediate result for determining the gradient, an output for determining the gradient, an input for determining the difference in brightness, an intermediate result for determining the difference in brightness, an output for determining the difference in brightness, an input for determining the product of gradients, an intermediate result for determining the product of gradients, an output for determining the product of gradients, an input for determining the product of brightness differences, an intermediate result for determining the product of brightness differences, an output for determining the product of brightness differences, an input for determining the sum of parameters, an intermediate result for determining the sum of parameters, and an intermediate result for determining the product of parameters. The following are examples of calculations: intermediate results, output of summation of parameters, input for determining the determinant, intermediate results for determining the determinant, output of determining the determinant, input for determining the product of parameters, intermediate results for determining the product of parameters, output of determining the product of parameters, input for determining the product of parameters, intermediate results for determining the product of parameters, output of determining the product of parameters, input for performing division of the determinant, intermediate results for performing division of the determinant, output of performing division of the determinant, input for determining sample point adjustment, intermediate results for determining sample point adjustment, or output of determining sample point adjustment. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0099] In some embodiments, the parameters of the shift operation may depend on the bit depth of the current video block. Alternatively, the parameters of the shift operation may be indicated in the bitstream.

[0100] For example, a right shift operation can be implemented, and an offset can be added to the operand of the right shift operation. By way of example, the horizontal and vertical gradients used in the BDOF process can be determined based on the following formula: Where I (k) (i,j). Represents the sample value at coordinates (i,j) of the predicted signal of the reference video block in list k, where k can be 0 or 1. This represents the value of the horizontal gradient of the predicted signal of the combined sub-block at coordinates (i,j) of the reference video block in list k. `i,j` represents the coordinates (i,j) of the predicted signal of the reference video block in list `k` for the combined sub-block. `whp` represents the value of the vertical gradient at coordinates (i+p,j). `wvp` represents the weight of the sample at coordinates (i,j+p), and each of `Na`, `Nb`, `A1`, `B1`, `A2`, and `B2` represents an integer. By way of example rather than constraint, `A1` can be equal to 0 or 1 << (B1-1), and `A2` can be equal to 0 or 1 << (B2-1).

[0101] In some embodiments, the size of the current sub-block may depend on the number of samples included in the current video block. For example, if the number of samples included in the current video block is within a certain range, the size of the current sub-block may be equal to the size corresponding to the specific range.

[0102] Alternatively, the size of the current sub-block can depend on the width and height of the current video block. For example, if the width and height of the current video block are within a certain range of values, the size of the current sub-block can be equal to the size corresponding to that specific range of values.

[0103] In some other embodiments, the size of the current sub-block may depend on the encoding / decoding tool applied to the current video block. In still other embodiments, the size of the current sub-block may depend on the temporal layer of the current video block or a reference video block of the current video block. For example, if the temporal layer is within a certain value range, the size of the current sub-block may be equal to the size corresponding to that specific value range.

[0104] In some embodiments, the sub-block size of the luma or chroma block for the current video block may depend on at least one of the following: including the size of the current prediction unit (PU) of the current video block, the size of the current codec unit (CU) of the current video block, characteristics of a plurality of reference video blocks of the current video block, similarity of a plurality of prediction values ​​from a plurality of reference video blocks, distribution of differences between a plurality of prediction values ​​from a plurality of reference video blocks, temporal gradients of a plurality of reference video blocks, spatial gradients of a plurality of reference video blocks, prediction type, adjustment value determined in the first pass of multi-pass decoder-side motion vector refinement (DMVR), adjustment value determined in the second pass of multi-pass DMVR, sequence resolution, color components of the current video block, color format of the current video block, codec information of the current video block, information of at least one prediction block of the current video block, or the value of a quantization parameter (QP) associated with the current video block. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0105] In one example embodiment, the encoded / decoded information may include residual information, the encoding / decoding tools applied to the current video block, etc. Additionally or alternatively, at least one prediction block may include multiple prediction blocks from multiple lists of reference images (such as list 0, list 1, etc.) of the current video block. Furthermore, information about at least one prediction block may include characteristics of at least one prediction block, the size of at least one prediction block, etc. Moreover, quantization parameters associated with the current video block may include quantization parameters of the current video block, quantization parameters of the current codec unit (CU) including the current video block, quantization parameters of the current stripe including the current video block, quantization parameters of the sequence including the current video block, etc.

[0106] In some embodiments, the sub-block size of the chroma block for the current video block may be determined based on the sub-block size of the luma block for the current video block and at least one of the following: the color format of the current video block, or information about whether individual planar codecs are enabled for the current video block.

[0107] In some embodiments, the BDOF process for MV refinement and the BDOF process for sample adjustment can be controlled separately. For example, the BDOF process for MV refinement can be applied, while the BDOF process for sample adjustment is not applied. Alternatively, the BDOF process for sample adjustment can be applied, while the BDOF process for MV refinement is not applied. For example, control can be performed at the sub-PU level, the codec unit (CU) level, the codec tree unit (CTU) level, etc.

[0108] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve encoding / decoding efficiency and encoding / decoding quality.

[0109] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. In this method, for a current sub-block of a current video block, at least one sub-block is determined from a plurality of sub-blocks of the current video block. The motion vector (MV) of each of the at least one sub-block satisfies one of the following conditions with the MV of the current sub-block: the MV of each of the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each of the at least one sub-block and the MV of the current sub-block is less than a threshold. Furthermore, combined sub-blocks are obtained by combining at least one sub-block and the current sub-block. The bitstream is generated based on applying a bidirectional optical flow (BDOF) process to the combined sub-blocks.

[0110] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, for a current sub-block of a current video block, at least one sub-block is determined from a plurality of sub-blocks of the current video block. The motion vector (MV) of each of the at least one sub-block satisfies one of the following conditions with the MV of the current sub-block: the MV of each of the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each of the at least one sub-block and the MV of the current sub-block is less than a threshold. Furthermore, combined sub-blocks are obtained by combining at least one sub-block and the current sub-block. The bitstream is generated based on applying a bidirectional optical flow (BDOF) process to the combined sub-blocks and is stored in a non-transitory computer-readable recording medium.

[0111] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0112] Item 1. A method for video processing, comprising: a conversion between a current video block and a bitstream of the video; determining at least one sub-block from a plurality of sub-blocks of the current video block for a current sub-block of the current video block, wherein the motion vector (MV) of each of the at least one sub-block satisfies one of the following conditions: the MV of each of the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each of the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block by combining the at least one sub-block and the current sub-block; and performing the conversion based on applying a bidirectional optical flow (BDOF) procedure to the combined sub-block.

[0113] Item 2. The method according to Item 1, wherein the difference metric includes first-order Euclidean distance or second-order Euclidean distance.

[0114] Item 3. The method according to any one of items 1-2, wherein if the MV of each of the at least one sub-blocks is the same as the MV of the current sub-block, then the MV of the combined sub-block is the same as the MV of the current sub-block, or if the difference metric is less than the threshold, then the MV of the combined sub-block is determined based on the MV of the current sub-block and the MV of the at least one sub-block.

[0115] Item 4. The method according to any one of items 1-3, wherein if the difference metric is less than the threshold, the MV of the combined sub-block is determined to be one of the following: an average MV determined by averaging the MV of the current sub-block and the MV of the at least one sub-block, the most frequently occurring MV among the MV of the current sub-block and the MV of the at least one sub-block, or the MV of a sub-block at a predetermined position.

[0116] Item 5. The method according to Item 4, wherein the predetermined position includes one of the following: a center position, an upper left position, or a lower right position.

[0117] Item 6. The method according to any one of items 1-5, wherein the BDOF process is a BDOF process for MV refinement or a BDOF process for sample adjustment.

[0118] Item 7. The method described in Item 6, wherein the BDOF process is one round of the MV refinement process.

[0119] Item 8. The method according to any one of items 1-7, wherein the plurality of sub-blocks includes sub-blocks adjacent to the current sub-block.

[0120] Item 9. The method according to any one of items 1-8, wherein motion compensation is applied to the combined sub-block.

[0121] Item 10. The method according to any one of items 1-9, wherein the at least one sub-block includes N neighboring sub-blocks of the current sub-block located in the same row as the current sub-block, and N is an integer.

[0122] Item 11. The method according to any one of items 1-10, wherein the at least one sub-block includes M neighboring sub-blocks of the current sub-block located in the same column as the current sub-block, and M is an integer.

[0123] Item 12. The method according to any one of Items 10-11, wherein the at least one sub-block is within the sub-prediction unit (PU) boundary of the i-th round decoder-side motion vector refinement (DMVR) process, and i is an integer.

[0124] Item 13. The method according to any one of items 1-12, wherein whether the at least one sub-block is located in the same row as the current sub-block or in the same column as the current sub-block depends on the size of the current video block.

[0125] Item 14. The method according to Item 13, wherein if the width of the current video block is greater than or equal to the height of the current video block, then the at least one sub-block is in the same row as the current sub-block, or if the width is less than the height, then the at least one sub-block is in the same column as the current sub-block.

[0126] Item 15. The method according to Item 13, wherein if the height of the current video block is greater than the width of the current video block, then the at least one sub-block is in the same row as the current sub-block, or if the height is less than or equal to the width, then the at least one sub-block is in the same column as the current sub-block.

[0127] Item 16. The method according to any one of items 1-15, wherein the plurality of sub-blocks are included in a sub-segment of the current video block, the size of the sub-segment being larger than the size of each of the plurality of sub-blocks.

[0128] Item 17. The method according to Item 16, wherein the size of each of the plurality of sub-blocks is 4×4, and the size of the sub-segment is one of 16×16, 8×8, 16×8 or 32×32, or the size of each of the plurality of sub-blocks is 8×8, and the size of the sub-segment is one of 16×16, 16×8 or 32×32.

[0129] Item 18. The method according to any one of Items 16-17, wherein the size of the sub-segment is variable or fixed.

[0130] Item 19. The method according to any one of items 1-18, wherein the at least one sub-block is allowed to include all sub-blocks of the current video block.

[0131] Item 20. The method according to any one of items 1-19, wherein if the difference metric between the MV of the first sub-block and the MV of the current sub-block is less than the threshold, and the reference image for the first sub-block is the same as the reference image for the current sub-block, then the motion information of at least one of the first sub-block or the current sub-block is modified such that the same motion information is used for the first sub-block and the current sub-block in subsequent processes.

[0132] Item 21. The method according to any one of items 1-20, wherein the method is applied to at least one of two prediction directions.

[0133] Item 22. The method according to Item 21, wherein the two prediction directions include a prediction direction corresponding to a first reference list and a prediction direction corresponding to a second reference list.

[0134] Item 23. The method according to any one of items 1-22, wherein the BDOF parameters used in the BDOF process are determined in parallel.

[0135] Item 24. The method according to any one of items 1-23, wherein multiple functions related to the BDOF procedure are implemented based on a single instruction multiple data (SIMD) scheme.

[0136] Item 25. The method according to any one of items 1-24, wherein a plurality of summation operations are performed at one iteration to determine the sum of the BDOF parameters.

[0137] Item 26. The method according to any one of items 1-25, wherein the multiplication operation is implemented using a left shift operation.

[0138] Item 27. The method described in Item 26, wherein the multiplication of weight w is replaced by a left shift of log2(1+w).

[0139] Item 28. The method according to any one of items 1-27, wherein one or more parameters used in the BDOF process for MV refinement are determined based on conditions.

[0140] Item 29. The method according to Item 28, wherein the one or more parameters are determined if they will be used, or if the BDOF procedure for sample adjustment is not applied.

[0141] Item 30. The method according to any one of items 1-29, wherein one or more parameters used in the BDOF process for sample adjustment are determined based on conditions.

[0142] Item 31. The method according to Item 30, wherein the one or more parameters are determined if they will be used, or if the BDOF process for MV refinement is not applied, or if the result of the BDOF process for MV refinement indicates that MV updates are not needed.

[0143] Item 32. The method according to any one of items 1-31, wherein the first sub-block size used to determine whether the BDOF process for MV refinement should be applied is different from the second sub-block size used to determine whether the BDOF process for sample adjustment should be applied.

[0144] Item 33. The method according to Item 32, wherein the size of the second sub-block is greater than the size of the first sub-block, or the size of the second sub-block is smaller than the size of the first sub-block.

[0145] Item 34. The method according to any one of items 1-33, wherein the shift operation is performed at one or more stages for determining at least one BDOF parameter used in the process of determining the BDOF.

[0146] Item 35. The method according to Item 34, wherein the shift operation includes a right shift operation or a left shift operation.

[0147] Item 36. The method according to any one of items 34-35, wherein an offset is added to at least one of: the operand of the shift operation, or the result of the shift operation.

[0148] Item 37. The method according to any one of items 34-36, wherein the result of the shift operation is limited to a range of values.

[0149] Item 38. The method according to any one of items 34-37, wherein the shift operation is performed on at least one of: an input for determining a gradient, an intermediate result for determining the gradient, an output for determining the gradient, an input for determining a difference in brightness, an intermediate result for determining the difference in brightness, an output for determining the difference in brightness, an input for determining a product of gradients, an intermediate result for determining the product of gradients, an output for determining the product of gradients, an input for determining a product of brightness differences, an intermediate result for determining the product of brightness differences, an output for determining the product of brightness differences, an input for determining the summation of parameters, and an input for determining the... Intermediate result of parameter summation, output of parameter summation, input for determining determinant, intermediate result of determining determinant, output of determining determinant, input for determining parameter product, intermediate result of determining parameter product, output of determining parameter product, input for determining parameter product, intermediate result of determining parameter product, output of determining parameter product, input for performing determinant division, intermediate result of performing determinant division, output of performing determinant division, input for determining sample point adjustment, intermediate result of determining sample point adjustment, or output of determining sample point adjustment.

[0150] Item 39. The method according to any one of items 34-38, wherein the parameters of the shift operation depend on the bit depth of the current video block.

[0151] Item 40. The method according to any one of items 1-39, wherein the horizontal and vertical gradients used in the BDOF process are determined based on the following formula: Among them I (k) (i,j). Represents the sample value at coordinates (i,j) of the predicted signal of the reference video block in list k, where k is equal to 0 or 1. The value of the horizontal gradient at coordinates (i,j) of the predicted signal of the reference video block in list k represents the combined sub-block. The coordinates (i,j) of the predicted signal of the reference video block in list k represent the vertical gradient value at coordinates (i+p,j), where whp represents the weight of the sample at coordinates (i+p,j), and wvp represents the weight of the sample at coordinates (i,j+p), where each of Na, Nb, A1, B1, A2, and B2 represents an integer.

[0152] Item 41. The method according to Item 40, wherein A1 equals 0 or 1 << (B1-1), and A2 equals 0 or 1 << (B2-1).

[0153] Item 42. The method according to any one of items 1-41, wherein the size of the current sub-block depends on the number of samples included in the current video block.

[0154] Item 43. The method according to Item 42, wherein if the number of samples included in the current video block is within a certain value range, then the size of the current sub-block is equal to the size corresponding to the value range.

[0155] Item 44. The method according to any one of items 1-41, wherein the size of the current sub-block depends on the width and height of the current video block.

[0156] Item 45. The method according to Item 44, wherein if the width and height of the current video block are within a certain range, then the size of the current sub-block is equal to the size corresponding to the range of values.

[0157] Item 46. The method according to any one of items 1-45, wherein the size of the current sub-block depends on the encoding / decoding tool applied to the current video block.

[0158] Item 47. The method according to any one of items 1-46, wherein the size of the current sub-block depends on the temporal layer of the current video block or a reference video block of the current video block.

[0159] Item 48. The method according to Item 47, wherein if the temporal layer is within a value range, then the size of the current sub-block is equal to the size corresponding to the value range.

[0160] Item 49. The method according to any one of items 1-48, wherein the sub-block size of the luma block or chroma block for the current video block depends on at least one of the following: the size of the current prediction unit (PU) including the current video block, the size of the current codec unit (CU) including the current video block, the characteristics of a plurality of reference video blocks of the current video block, the similarity of a plurality of prediction values ​​from the plurality of reference video blocks, the distribution of differences between the plurality of prediction values ​​from the plurality of reference video blocks, the temporal gradient of the plurality of reference video blocks, the spatial gradient of the plurality of reference video blocks, the prediction type, the adjustment value determined in the first pass of the multi-pass decoder-side motion vector refinement (DMVR), the adjustment value determined in the second pass of the multi-pass DMVR, the sequence resolution, the color components of the current video block, the color format of the current video block, the codec information of the current video block, the information of at least one prediction block of the current video block, or the value of the quantization parameter (QP) associated with the current video block.

[0161] Item 50. The method according to any one of items 1-48, wherein the sub-block size of the chroma block for the current video block is determined based on the sub-block size of the luma block for the current video block and at least one of the following: the color format of the current video block, or information regarding whether individual planar codecs are enabled for the current video block.

[0162] Item 51. The method according to any one of items 1-50, wherein the BDOF process for MV refinement and the BDOF process for sample adjustment are controlled separately.

[0163] Item 52. The method according to Item 51, wherein the BDOF process for MV refinement is applied, and the BDOF process for sample adjustment is not applied.

[0164] Item 53. The method according to any one of items 51-52, wherein the control is performed at one of the following levels: subPU level, codec unit (CU) level, or codec tree unit (CTU) level.

[0165] Item 54. The method according to any one of items 1-53, wherein the conversion includes encoding the current video block into the bitstream.

[0166] Item 55. The method according to any one of items 1-53, wherein the conversion includes decoding the current video block from the bitstream.

[0167] Item 56. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-55.

[0168] Item 57. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-55.

[0169] Item 58. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining at least one sub-block from a plurality of sub-blocks of a current video block for a current sub-block of the video, the motion vector (MV) of each of the at least one sub-block and the MV of the current sub-block satisfying one of the following conditions: the MV of each of the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each of the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block by combining the at least one sub-block and the current sub-block; and generating the bitstream based on applying a bidirectional optical flow (BDOF) process to the combined sub-block.

[0170] Item 59. A method for storing a bitstream of video, comprising: determining at least one sub-block from a plurality of sub-blocks of a current video block for a current sub-block of the video, wherein the motion vector (MV) of each of the at least one sub-block satisfies one of the following conditions: the MV of each of the at least one sub-block is the same as the MV of the current sub-block, or the difference metric between the MV of each of the at least one sub-block and the MV of the current sub-block is less than a threshold; obtaining a combined sub-block by combining the at least one sub-block and the current sub-block; generating the bitstream based on applying a bidirectional optical flow (BDOF) process to the combined sub-block; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0171] Figure 13A block diagram of a computing device 1300 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1300 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0172] It should be understood that, Figure 13 The computing device 1300 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0173] like Figure 13 As shown, computing device 1300 includes general-purpose computing device 1300. Computing device 1300 may include at least one or more processors or processing units 1310, memory 1320, storage unit 1330, one or more communication units 1340, one or more input devices 1350, and one or more output devices 1360.

[0174] In some embodiments, the computing device 1300 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 1300 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0175] Processing unit 1310 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 1320. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 1300. Processing unit 1310 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0176] Computing device 1300 typically includes various computer storage media. Such media can be any media accessible by computing device 1300, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 1320 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 1330 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 1300.

[0177] The computing device 1300 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 13 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0178] Communication unit 1340 communicates with another computing device via a communication medium. Additionally, the functionality of components in computing device 1300 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 1300 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0179] Input device 1350 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 1360 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 1340, computing device 1300 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 1300 can also communicate with one or more devices that enable a user to interact with computing device 1300, or, if necessary, with any device (e.g., network card, modem, etc.) that enables computing device 1300 to communicate with one or more other computing devices. Such communication can be performed via an input / output (I / O) interface (not shown).

[0180] In some embodiments, some or all of the components of computing device 1300 may be arranged in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although to users they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0181] In embodiments of this disclosure, computing device 1300 may be used to implement video encoding / decoding. Memory 1320 may include one or more video codec modules 1325 having one or more program instructions. These modules are accessible and executable by processing unit 1310 to perform the functions of the various embodiments described herein.

[0182] In an example embodiment of performing video encoding, input device 1350 may receive video data as input 1370 to be encoded. The video data may be processed, for example, by video codec module 1325 to generate an encoded bitstream. The encoded bitstream may be provided as output 1380 via output device 1360.

[0183] In an example embodiment of performing video decoding, input device 1350 may receive an encoded bitstream as input 1370. The encoded bitstream may be processed, for example, by video codec module 1325 to generate decoded video data. The decoded video data may be provided as output 1380 via output device 1360.

[0184] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, for the current sub-block of the current video block, at least one sub-block is determined from a plurality of sub-blocks of the current video block, wherein the motion vector (MV) of each of the at least one sub-blocks satisfies one of the following conditions with the MV of the current sub-block: The MV of each of the at least one sub-blocks is the same as the MV of the current sub-block, or The difference metric between the MV of each of the at least one sub-blocks and the MV of the current sub-block is less than a threshold; A combined sub-block is obtained by combining the at least one sub-block and the current sub-block; as well as The conversion is performed by applying a bidirectional optical flow (BDOF) process to the combined sub-block.

2. The method according to claim 1, wherein the difference metric includes first-order Euclidean distance or second-order Euclidean distance.

3. The method according to any one of claims 1-2, wherein if the MV of each of the at least one sub-block is the same as the MV of the current sub-block, then the MV of the combined sub-block is the same as the MV of the current sub-block, or If the difference metric is less than the threshold, the MV of the combined sub-block is determined based on the MV of the current sub-block and the MV of the at least one sub-block.

4. The method according to any one of claims 1-3, wherein if the difference metric is less than the threshold, the MV of the combined sub-block is determined to be one of the following: The average MV is determined by averaging the MV of the current sub-block and the MV of at least one sub-block. The most frequent MV among the MVs of the current sub-block and the MVs of at least one sub-block, or The MV of the sub-block at the predetermined position.

5. The method of claim 4, wherein the predetermined position includes one of the following: central location Top left position, or Bottom right position.

6. The method according to any one of claims 1-5, wherein the BDOF process is a BDOF process for MV refinement or a BDOF process for sample adjustment.

7. The method according to claim 6, wherein the BDOF process is one round of the MV refinement process.

8. The method according to any one of claims 1-7, wherein the plurality of sub-blocks includes sub-blocks adjacent to the current sub-block.

9. The method according to any one of claims 1-8, wherein motion compensation is applied to the combined sub-block.

10. The method according to any one of claims 1-9, wherein the at least one sub-block includes N neighboring sub-blocks of the current sub-block located in the same row as the current sub-block, and N is an integer.

11. The method according to any one of claims 1-10, wherein the at least one sub-block includes M neighboring sub-blocks of the current sub-block located in the same column as the current sub-block, and M is an integer.

12. The method according to any one of claims 10-11, wherein the at least one sub-block is within the sub-prediction unit (PU) boundary of the i-th round decoder-side motion vector refinement (DMVR) process, and i is an integer.

13. The method according to any one of claims 1-12, wherein whether the at least one sub-block is located in the same row as the current sub-block or in the same column as the current sub-block depends on the size of the current video block.

14. The method of claim 13, wherein if the width of the current video block is greater than or equal to the height of the current video block, then the at least one sub-block is located in the same row as the current sub-block, or If the width is less than the height, then the at least one sub-block is in the same column as the current sub-block.

15. The method of claim 13, wherein if the height of the current video block is greater than the width of the current video block, then the at least one sub-block is located in the same row as the current sub-block, or If the height is less than or equal to the width, then the at least one sub-block is in the same column as the current sub-block.

16. The method according to any one of claims 1-15, wherein the plurality of sub-blocks are included in a sub-segment of the current video block, and the size of the sub-segment is greater than the size of each of the plurality of sub-blocks.

17. The method of claim 16, wherein each of the plurality of sub-blocks has a size of 4×4, and the size of the sub-segment is one of 16×16, 8×8, 16×8, or 32×32, or Each of the plurality of sub-blocks has a size of 8×8, and the size of the sub-segment is one of 16×16, 16×8, or 32×32.

18. The method according to any one of claims 16-17, wherein the size of the sub-segment is variable or fixed.

19. The method according to any one of claims 1-18, wherein the at least one sub-block is permitted to include all sub-blocks of the current video block.

20. The method of any one of claims 1-19, wherein if the difference metric between the MV of the first sub-block and the MV of the current sub-block is less than the threshold, and the reference image for the first sub-block is the same as the reference image for the current sub-block, then the motion information of at least one of the first sub-block or the current sub-block is modified such that the same motion information is used for the first sub-block and the current sub-block in subsequent processes.

21. The method according to any one of claims 1-20, wherein the method is applied to at least one of two prediction directions.

22. The method of claim 21, wherein the two prediction directions include prediction directions corresponding to a first reference list and prediction directions corresponding to a second reference list.

23. The method according to any one of claims 1-22, wherein the BDOF parameters used in the BDOF process are determined in parallel.

24. The method according to any one of claims 1-23, wherein the plurality of functions associated with the BDOF process are implemented based on a single instruction multiple data (SIMD) scheme.

25. The method according to any one of claims 1-24, wherein a plurality of summation operations are performed at one iteration to determine the sum of the BDOF parameters.

26. The method according to any one of claims 1-25, wherein the multiplication operation is implemented using a left shift operation.

27. The method of claim 26, wherein the multiplication of the weight w is replaced by a left shift of log2(1+w).

28. The method according to any one of claims 1-27, wherein one or more parameters used in the BDOF process for MV refinement are determined based on conditions.

29. The method of claim 28, wherein if the one or more parameters will be used, then the one or more parameters are determined, or If the BDOF procedure for sample adjustment is not applied, then one or more parameters are determined.

30. The method according to any one of claims 1-29, wherein one or more parameters used in the BDOF process for sample adjustment are determined based on conditions.

31. The method of claim 30, wherein if the one or more parameters will be used, then the one or more parameters are determined, or If the BDOF procedure used for MV refinement is not applied, then one or more parameters are determined, or If the result of the BDOF process used for MV refinement indicates that MV updates are not needed, then one or more parameters are determined.

32. The method according to any one of claims 1-31, wherein the first sub-block size used to determine whether the BDOF process for MV refinement should be applied is different from the second sub-block size used to determine whether the BDOF process for sample adjustment should be applied.

33. The method of claim 32, wherein the size of the second sub-block is larger than the size of the first sub-block, or The size of the second sub-block is smaller than the size of the first sub-block.

34. The method according to any one of claims 1-33, wherein the shift operation is performed at one or more stages for determining at least one BDOF parameter used in the process of determining the BDOF.

35. The method of claim 34, wherein the shift operation includes a right shift operation or a left shift operation.

36. The method according to any one of claims 34-35, wherein the offset is added to at least one of: The operands of the shift operation, or The result of the shift operation.

37. The method according to any one of claims 34-36, wherein the result of the shift operation is limited to a range of values.

38. The method according to any one of claims 34-37, wherein the shift operation is performed on at least one of the following: The input used to determine the gradient. Intermediate results used to determine the gradient, Determine the output of the gradient. Input used to determine the difference in brightness. Intermediate results used to determine the difference in brightness, The output that determines the difference in brightness The input used to determine the product of gradients. Intermediate results used to determine the product of the gradients, Determine the output of the product of the gradients. The input used to determine the product of brightness differences. Intermediate results used to determine the product of the brightness differences, The output is determined by the product of the brightness differences. The input used to determine the summation of parameters, Intermediate results used to determine the summation of the parameters, Determine the output of the summation of the parameters. The input used to determine the determinant Intermediate results used to determine the determinant. Determine the output of the determinant. The input used to determine the product of the parameters. Intermediate results used to determine the product of the parameters. Determine the output of the product of the parameters. The input used to determine the product of the parameters. Intermediate results used to determine the product of the parameters. Determine the output of the product of the parameters. The input used to perform division of the determinant. Intermediate results used to perform the division of the determinant. The output of performing the division of the determinant, Inputs used to determine sample point adjustments Used to determine the intermediate results of the sample point adjustment, or Determine the output of the sample point adjustment.

39. The method according to any one of claims 34-38, wherein the parameters of the shift operation depend on the bit depth of the current video block.

40. The method according to any one of claims 1-39, wherein the horizontal and vertical gradients used in the BDOF process are determined based on the following formula: Where I (k) (i,j). Represents the sample value at coordinates (i,j) of the predicted signal of the reference video block in list k, where k is equal to 0 or 1. The value of the horizontal gradient at coordinates (i,j) of the predicted signal of the reference video block in list k represents the combined sub-block. The coordinates (i,j) of the predicted signal of the reference video block in list k represent the vertical gradient value at coordinates (i+p,j), where whp represents the weight of the sample at coordinates (i+p,j), and wvp represents the weight of the sample at coordinates (i,j+p), where each of Na, Nb, A1, B1, A2, and B2 represents an integer.

41. The method of claim 40, wherein A1 is equal to 0 or 1 << (B1-1), and A2 is equal to 0 or 1 << (B2-1).

42. The method according to any one of claims 1-41, wherein the size of the current sub-block depends on the number of samples included in the current video block.

43. The method of claim 42, wherein if the number of samples included in the current video block is within a certain range, then the size of the current sub-block is equal to the size corresponding to the range of values.

44. The method according to any one of claims 1-41, wherein the size of the current sub-block depends on the width and height of the current video block.

45. The method of claim 44, wherein if the width and height of the current video block are within a certain range, then the size of the current sub-block is equal to the size corresponding to the range of values.

46. ​​The method according to any one of claims 1-45, wherein the size of the current sub-block depends on the encoding / decoding tool applied to the current video block.

47. The method according to any one of claims 1-46, wherein the size of the current sub-block depends on the temporal layer of the current video block or a reference video block of the current video block.

48. The method of claim 47, wherein if the temporal layer is within a value range, the size of the current sub-block is equal to the size corresponding to the value range.

49. The method according to any one of claims 1-48, wherein the sub-block size of the luminance block or chroma block for the current video block depends on at least one of the following: Including the size of the current prediction unit (PU) of the current video block, Including the size of the current codec unit (CU) of the current video block, The characteristics of the multiple reference video blocks of the current video block, The similarity of multiple predicted values ​​from the plurality of reference video blocks, The distribution of differences between multiple predicted values ​​from the plurality of reference video blocks. The temporal gradients of the plurality of reference video blocks, The spatial gradient of the plurality of reference video blocks, Prediction type The adjustment values ​​determined in the first pass of the multi-pass decoder-side motion vector refinement (DMVR) are... The adjustment value determined in the second pass of the multiple passes of DMVR Sequence resolution, The color components of the current video block, The color format of the current video block, The encoded and decoded information of the current video block, Information of at least one predicted block of the current video block, or The value of the quantization parameter (QP) associated with the current video block.

50. The method according to any one of claims 1-48, wherein the sub-block size of the chroma block for the current video block is determined based on the sub-block size of the luma block for the current video block and at least one of the following: The color format of the current video block, or Information regarding whether individual planar encoding / decoding is enabled for the current video block.

51. The method according to any one of claims 1-50, wherein the BDOF process for MV refinement and the BDOF process for sample adjustment are controlled separately.

52. The method of claim 51, wherein the BDOF process for MV refinement is applied, and the BDOF process for sample adjustment is not applied.

53. The method according to any one of claims 51-52, wherein the control is performed at one of the following locations: Sub-PU level, Codec Unit (CU) level, or Code-decode tree unit (CTU) level.

54. The method according to any one of claims 1-53, wherein the conversion comprises encoding the current video block into the bitstream.

55. The method according to any one of claims 1-53, wherein the conversion comprises decoding the current video block from the bitstream.

56. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-55.

57. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-55.

58. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: For the current sub-block of the current video block, at least one sub-block is determined from a plurality of sub-blocks of the current video block, wherein the motion vector (MV) of each of the at least one sub-blocks satisfies one of the following conditions with the MV of the current sub-block: The MV of each of the at least one sub-blocks is the same as the MV of the current sub-block, or The difference metric between the MV of each of the at least one sub-blocks and the MV of the current sub-block is less than a threshold; A combined sub-block is obtained by combining the at least one sub-block and the current sub-block; as well as The bit stream is generated by applying the bidirectional optical flow (BDOF) process to the combined sub-blocks.

59. A method for storing a bitstream of video, comprising: For the current sub-block of the current video block, at least one sub-block is determined from a plurality of sub-blocks of the current video block, wherein the motion vector (MV) of each of the at least one sub-blocks satisfies one of the following conditions with the MV of the current sub-block: The MV of each of the at least one sub-blocks is the same as the MV of the current sub-block, or The difference metric between the MV of each of the at least one sub-blocks and the MV of the current sub-block is less than a threshold; A combined sub-block is obtained by combining the at least one sub-block and the current sub-block; The bit stream is generated by applying a bidirectional optical flow (BDOF) process to the combined sub-blocks; as well as The bitstream is stored in a non-transitory computer-readable recording medium.