Methods, apparatus, medium and computer program product for video coding
By applying bi-directional optical flow refinement and decoder-side motion vector refinement based on specific conditions, the method addresses the inefficiencies in existing video coding technologies, enhancing compression efficiency and video quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-03-26
Smart Images

Figure CN2025123362_26032026_PF_FP_ABST
Abstract
Description
METHODS, APPARATUS, MEDIUM AND COMPUTER PROGRAM PRODUCT FOR VIDEO CODINGCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application is based upon and claims priority to PCT Application No. PCT / CN2024 / 120381 filed on September 23, 2024, and PCT Application No. PCT / CN2024 / 125958 filed on October 19, 2024, the entire content thereof is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] This application is related to video coding and compression. More specifically, this application relates to methods and apparatus to improve the coding efficiency of bi-directional optical flow (BDOF) .BACKGROUND
[0003] Digital video is supported by a variety of electronic devices. The electronic devices transmit and receive or otherwise communicate digital video data across a communication network, and / or store the digital video data on a storage device. Due to a limited bandwidth capacity of the communication network and limited memory resources of the storage device, video coding may be used to compress the video data according to one or more video coding standards before it is communicated or stored, to generate encoded video data that uses a lower bit rate, while avoiding or minimizing degradations to video quality.SUMMARY
[0004] Embodiments of the present disclosure provide methods, apparatus, medium and computer program product for video coding.
[0005] According to a first aspect of the present disclosure, a method for video decoding is provided. The method includes determining a first one or more conditions for a current block in a current picture; determining a second one or more conditions, different from the first one or more conditions, for the current block; and in response to determining that the first one or more conditions and the second one or more conditions are met, determining whether to apply a Bi-directional optical flow refinement or a Decoder side motion vector refinement for the current block.
[0006] According to a second aspect of the present disclosure, a method for video encoding is provided. The method includes determining a first one or more conditions for a current block in a current picture; determining a second one or more conditions, different from the first one or more conditions, for the current block; and in response to determining that the first one or more conditions and the second one or more conditions are met, determining whether to apply a Bi-directional optical flow refinement or a Decoder side motion vector refinement for the current block.
[0007] According to a third aspect of the present disclosure, an apparatus for video coding is provided. The apparatus includes one or more processors; a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors, upon execution of the instructions, are configured to perform method according to the embodiments of the present application.
[0008] According to a fourth aspect of the present disclosure, a non-transitory computer readable storage medium is provided. The non-transitory computer readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform method according to the embodiments of the present application.
[0009] According to a fifth aspect of the present disclosure, a non-transitory computer readable storage medium is provided. The non-transitory computer readable storage medium stores a bitstream to be decoded by method according to the embodiments of the present application or a bitstream generated by method according to the embodiments of the present application.
[0010] According to a sixth aspect of the present disclosure, a method for storing a bitstream is provided. The method includes generating a bitstream by performing method according to the embodiments of the present application, and storing the bitstream on a storage medium.
[0011] According to a seventh aspect of the present disclosure, a computer program product is provided. The computer program product includes a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the method according to the embodiments of the present application.
[0012] It is to be understood that both the foregoing general description and the following detailed description are examples only and are not restrictive of the present disclosure.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0014] FIG. 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks in accordance with some implementations of the present disclosure.
[0015] FIG. 2 is a block diagram illustrating an exemplary video encoder in accordance with some implementations of the present disclosure.
[0016] FIG. 3 is a block diagram illustrating an exemplary video decoder in accordance with some implementations of the present disclosure.
[0017] FIGS. 4A through 4E are block diagrams illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes in accordance with some implementations of the present disclosure.
[0018] FIG. 5 is a diagram illustrating a computing environment coupled with a user interface, according to some implementations of the present disclosure.
[0019] FIG. 6 illustrates the Extended CU region used in BDOF.
[0020] FIG. 7 illustrates the process Decoding side motion vector refinement.
[0021] FIG. 8 illustrates the diamond regions in the search area in multi-pass DMVR.
[0022] FIG. 9 illustrates the proposed template matching based BDOF window size decision technique.
[0023] FIG. 10 is a flow chart illustrating a method for video decoding in accordance with some implementations of the present disclosure.
[0024] FIG. 11 is a flow chart illustrating a method for video encoding in accordance with some implementations of the present disclosure.DETAILED DESCRIPTION
[0025] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. But various alternatives may be used without departing from the scope of claims and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0026] It should be illustrated that the terms “first, ” “second, ” and the like used in the description, claims of the present disclosure, and the accompanying drawings are used to distinguish objects, and not used to describe any specific order or sequence. It should be understood that the data used in this way may be interchanged under an appropriate condition, such that the embodiments of the present disclosure described herein may be implemented in orders besides those shown in the accompanying drawings or described in the present disclosure.
[0027] FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel in accordance with some implementations of the present disclosure. As shown in FIG. 1, the system 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14. The source device 12 and the destination device 14 may comprise any of a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming device, or the like. In some implementations, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.
[0028] As shown in FIG. 1, the source device 12 includes a video source 18, a video encoder 20 and an output interface 22. The video source 18 may include a source such as a video capturing device, e.g., a video camera, a video archive containing previously captured video, a video feeding interface to receive video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources.
[0029] The captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may comprise a sequence of pictures, each of which may comprise one or more sample arrays, for example, luma (Y) only for monochrome; luma and two chroma in YCbCr or YCgCo domain; or green, blue, and red in GBR (also known as RGB) domain. For convenience of notation and terminology in this application, in some embodiments, variables and terms associated with each set of three sample arrays may be referred to as luma and chroma, where the two chroma arrays may be referred to as Cb and Cr, regardless of the actual color representation method in use. The video data may be in a chroma format of 4: 0: 0, 4: 2: 0, 4: 2: 2, or 4: 4: 4, but the present application is not limited thereto. A bit depth BitDepth of samples of sample arrays may be an integer in a range of 8 to 16. For example, a vlaue of BitDepth may be 8, 9, 10, 11, 12, 13, 14, 15 or 16. It should be illustrated that the value of BitDepth is not limited thereto, and may be any other value proposed in the future.
[0030] The encoded video data may be transmitted directly to the destination device 14 through the output interface 22 of the source device 12 via a link 16. The output interface 22 may include a modem and / or a transmitter. The link 16 may comprise any type of wireless communication medium or device and / or any type of wired communication medium or device capable of transmitting the encoded video data from the source device 12 to the destination device 14. The encoded video data may also (or alternatively) be stored onto a storage device 32 for later access by the destination device 14 via for example an input interface 28 or by other devices, for decoding and / or playback. The storage device 32 may include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Versatile Disks (DVDs) , Compact Disc Read-Only Memories (CD-ROMs) , flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.
[0031] The destination device 14 includes the input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and receive the encoded video data over the link 16. Alternatively, the destination device 14 may access the stored video data from the storage device 32 via streaming, downloading or a combination of both. The encoded video data may include a variety of syntax elements generated by the video encoder 20 for use by the video decoder 30 in decoding the video data. The display device 34 may be an integrated display device or an external display device that is configured to communicate with the destination device 14, and may display the decoded video data to a user.
[0032] The video encoder 20 and the video decoder 30 may operate (for example, encode and decode video data) according to proprietary or industry standards, such as Versatile Video Coding (VVC) , Joint Exploration test Model (JEM) , High-Efficiency Video Coding (HEVC / H. 265) , Advanced Video Coding (AVC / H. 264) , Moving Picture Expert Group (MPEG) coding, or extensions of such standards. It should be understood that the present application is not limited to a specific video encoding / decoding standard, and may be applicable to other current and future video encoding / decoding standards.
[0033] The video encoder 20 and the video decoder 30 each may be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, Digital Signal Processors (DSPs) , Application Specific Integrated Circuits (ASICs) , Field Programmable Gate Arrays (FPGAs) , discrete logic, software, hardware, firmware or any combinations thereof. When implemented partially in software, an electronic device may store instructions for the software in a suitable, non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in a respective device.
[0034] In some implementations, at least a part of components of the source device 12 and / or the destination device 14 (for example, components shown in Fig. 1, Fig. 2 and / or Fig. 3) may operate in a cloud computing service network which may provide software, platforms, and / or infrastructure, such as Software as a Service (SaaS) , Platform as a Service (PaaS) , or Infrastructure as a Service (IaaS) . In some implementations, one or more components in the source device 12 and / or the destination device 14 which are not included in the cloud computing service network may be provided in one or more client devices, and the one or more client devices may communicate with server computers in the cloud computing service network through a wireless or wired communication network. In an embodiment, at least a part of operations described herein may be implemented as cloud-based services provided by one or more server computers which are implemented by the at least a part of the components of the source device 12 and / or the destination device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by the one or more client devices. In some implementations, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. The terms such as “cloud, ” “cloud computing, ” “cloud-based” etc. herein may be used interchangeably as appropriate without departing from the scope of the present disclosure. It should be understood that the present disclosure is not limited to be implemented in the cloud computing service network described above. Instead, the present disclosure may also be implemented in any other type of computing environments currently known or developed in the future.
[0035] FIG. 2 is a block diagram illustrating an exemplary video encoder 20 in accordance with some implementations described in the present application.
[0036] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a Decoded Picture Buffer (DPB) 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partition unit 45, an intra prediction processing unit 46, and an Intra Block Copy (IBC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62 for video block reconstruction. An in-loop filter 63, such as a deblocking filter, may be positioned between the summer 62 and the DPB 64 to filter block boundaries to remove blockiness artifacts from reconstructed video. Another in-loop filter, such as Sample Adaptive Offset (SAO) filter, Cross Component Sample Adaptive Offset (CCSAO) filter and / or Adaptive in-Loop Filter (ALF) , may also be used in addition to the deblocking filter to filter an output of the summer 62. It should be illustrated that for the CCSAO technique, the present application is not limited to the embodiments described herein, and instead, the application may be applied to a situation where an offset is selected for a sample of any of a luma component and two chroma components (which may represent Y, Cb and Cr in YCbCr domain; Y, Cg and Co in YCgCo domain; or G, B and R in RGB domain for convenience of notation and terminology in this application as described above) according to one or more samples of any other of the luma component and the two chroma components to modify the sample of said any component based on the selected offset. Alternatively, there is also provided an SAO technique which is substantially the same as the CCSAO technique, except that for the SAO technique, an offset is selected for a sample of any of a luma component and two chroma components according to one or more samples of said any component to modify the sample of said any component based on the selected offset. Further, it should also be illustrated that a first component mentioned herein may be any of the luma component and the two chroma components, a second component mentioned herein may be any other of the luma component and the two chroma components, and a third component mentioned herein may be a remaining one of the luma component and the two chroma components. In some examples, the in-loop filters may be omitted, and the decoded video block may be directly provided by the summer 62 to the DPB 64. The video encoder 20 may take the form of a fixed or programmable hardware unit or may be divided among one or more of the illustrated fixed or programmable hardware units.
[0037] The video data memory 40 may store video data to be encoded by the components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18 as shown in FIG. 1. The DPB 64 is a buffer that stores reference video data (for example, reference frames or pictures) for use in encoding video data by the video encoder 20. The video data memory 40 and the DPB 64 may be formed by any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or off-chip relative to those components. It should be noted that the term “frame” may be used as synonyms for the term “image” or “picture” in the field of video coding.
[0038] As shown in FIG. 2, after receiving the video data, the partition unit 45 partitions the video data into video blocks. This partitioning may also include partitioning a video frame into slices, tiles (for example, sets of video blocks) , or other larger Coding Units (CUs) according to predefined splitting structures such as a Quad-Tree (QT) structure associated with the video data. The video frame is or may be regarded as a two-dimensional array or matrix of samples with sample values. A sample in the array may also be referred to as a pixel or a pel. A number of samples in horizontal and vertical directions (or axes) of the array or picture define a size and / or a resolution of the video frame. The video frame may be divided into multiple video blocks by, for example, using QT partitioning. The video block again is or may be regarded as a two-dimensional array or matrix of samples with sample values, although of smaller dimension than the video frame. A number of samples in horizontal and vertical directions (or axes) of the video block define a size of the video block. The video block may further be partitioned into one or more block partitions or sub-blocks (which may form again blocks) by, for example, iteratively using QT partitioning, Binary-Tree (BT) partitioning or Triple-Tree (TT) partitioning or any combination thereof. It should be noted that the term “block” or “video block” as used herein may be a portion, in particular a rectangular (square or non-square) portion, of a frame or a picture. With reference, for example, to HEVC and VVC, the block or video block may be or correspond to a Coding Tree Unit (CTU) , a CU, a Prediction Unit (PU) or a Transform Unit (TU) and / or may be or correspond to a corresponding block, e.g. a Coding Tree Block (CTB) , a Coding Block (CB) , a Prediction Block (PB) or a Transform Block (TB) and / or to a sub-block.
[0039] The prediction processing unit 41 may select one of a plurality of possible predictive coding modes, such as one of a plurality of intra or inter predictive coding modes, for the current video block based on error results (e.g., coding rate and the level of distortion) . The prediction processing unit 41 may provide the resulting intra or inter prediction coded block to the summer 50 to generate a residual block and to the summer 62 to reconstruct the encoded block for use as part of a reference frame subsequently. The prediction processing unit 41 also provides at least one of syntax elements, such as motion vectors, intra or inter mode indicators, partition information, and other such syntax information, to the entropy encoding unit 56.
[0040] In order to select an appropriate intra predictive coding mode for the current video block, the intra prediction processing unit 46 may perform intra predictive coding of the current video block relative to one or more neighbor blocks in the same frame as the current block to be coded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 perform inter predictive coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may perform multiple coding passes, e.g., to select an appropriate coding mode for each block of video data.
[0041] In some implementations, the motion estimation unit 42 generates a motion vector of the current block in a motion estimation process according to a predetermined pattern within a sequence of video frames. The motion vector may indicate displacement of a video block within a current frame relative to a predictive block within a reference frame relative to the current block being coded within the current frame The predetermined pattern may designate video frames in the sequence as P frames or B frames. In some implementations, a Motion Vector Predictor (MVP) of the current block which may be determined from motion information of spatially neighboring blocks and / or temporally co-located blocks of the current block is subtracted from an actual motion vector of the current block to produce a Motion Vector Difference (MVD) for the current block. Then, instead of encoding, into the video bitstream, the actual motion vector of the current block, information of the MVP and MVD may be encoded into the video bitstream. The IBC unit 48 may determine vectors, e.g., block vectors, for IBC coding in a manner similar to the determination of motion vectors by the motion estimation unit 42 for inter prediction, or may utilize the motion estimation unit 42 to determine the block vector. It is noted that an IBC mode may be regarded as either an intra prediction mode or a prediction mode other than the intra prediction mode and an inter prediction mode.
[0042] A predictive block for the video block may be or may correspond to a block or a reference block of a reference frame that is deemed as closely matching the video block to be coded in terms of pixel difference, which may be determined by Sum of Absolute Difference (SAD) , Sum of Square Difference (SSD) , or other difference metrics. In some implementations, the video encoder 20 may calculate values for sub-integer pixel positions of reference frames stored in the DPB 64. For example, the video encoder 20 may interpolate values of one-quarter pixel positions, one-eighth pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 may perform a motion search relative to the full pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision.
[0043] The motion estimation unit 42 determines motion vector information for a video block in an inter prediction coded frame by comparing the position of the video block to the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1) , each of which identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the determined motion vector information to the motion compensation unit 44 and then to the entropy encoding unit 56.
[0044] Motion compensation, performed by the motion compensation unit 44, may involve fetching or generating the predictive block based on the motion vector information determined by the motion estimation unit 42. Upon receiving the motion vector information for the current video block, the motion compensation unit 44 may locate a predictive block to which a motion vector points in one of the reference frame lists, retrieve the predictive block from the DPB 64, and forward the predictive block to the summer 50. The motion compensation unit 44 may also generate syntax elements associated with the video blocks of a video frame for use by the video decoder 30 in decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predictive block, any flags indicating the prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are illustrated separately for conceptual purposes.
[0045] In some implementations, the IBC unit 48 may generate vectors and fetch predictive blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, but with the predictive blocks being in the same frame as the current block being coded and with the vectors being referred to as block vectors as opposed to motion vectors.
[0046] In other examples, the IBC unit 48 may use the motion estimation unit 42 and the motion compensation unit 44, in whole or in part, to perform such functions for IBC prediction according to the implementations described herein. In either case, for Intra block copy, a predictive block may be a block that is deemed as closely matching the block to be coded, in terms of pixel difference, which may be determined by SAD, SSD, or other difference metrics, and identification of the predictive block may include calculation of values for sub-integer pixel positions.
[0047] The intra prediction processing unit 46 may intra-predict a current video block, as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra block copy prediction performed by the IBC unit 48, as described above. In particular, the intra prediction processing unit 46 may determine an intra prediction mode to encode a current block. The intra prediction processing unit 46 may provide information indicative of the selected intra-prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode in the bitstream.
[0048] After the prediction processing unit 41 determines the predictive block for the current video block, the summer 50 forms a residual block by subtracting pixel values of the predictive block from the pixel values of the current video block, forming pixel difference values. The pixel difference values may include luma or chroma component differences or both. The residual video data in the residual block may be included in one or more TUs and is provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using one or more transforms, such as a Discrete Cosine Transform (DCT) or a conceptually similar transform.
[0049] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 may then perform a scan of a matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 may perform the scan.
[0050] Following quantization, the entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, e.g., Context Adaptive Variable Length Coding (CAVLC) , Context Adaptive Binary Arithmetic Coding (CABAC) , Syntax-based context-adaptive Binary Arithmetic Coding (SBAC) , Probability Interval Partitioning Entropy (PIPE) coding or another entropy encoding methodology or technique. The encoded bitstream may then be transmitted to the video decoder 30 as shown in FIG. 1, or archived in the storage device 32 as shown in FIG. 1 for later transmission to or retrieval by the video decoder 30. The entropy encoding unit 56 may also entropy encode the motion vectors and the other syntax elements for the current video frame.
[0051] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transformation, respectively, to reconstruct the residual block in the pixel domain for generating a reference block for prediction of other video blocks.
[0052] The summer 62 adds the reconstructed residual block to the predictive block produced to produce a reference block for storage in the DPB 64.
[0053] FIG. 3 is a block diagram illustrating an exemplary video decoder 30 in accordance with some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, and a DPB 92. The prediction processing unit 81 includes a motion compensation unit 82, an intra prediction unit 84, and an IBC unit 85. The video decoder 30 may perform a decoding process generally reciprocal to the encoding process described above with respect to the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on motion vectors received from the entropy decoding unit 80, while the intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from the entropy decoding unit 80.
[0054] In some examples, a unit of the video decoder 30 may be tasked to perform the implementations of the present disclosure. Also, in some examples, the implementations of the present disclosure may be divided among one or more of the units of the video decoder 30.
[0055] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by the other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source, such as a camera, via wired or wireless network communication of video data, or by accessing physical data storage media (e.g., a flash drive or hard disk) . The video data memory 79 may include a Coded Picture Buffer (CPB) that stores encoded video data from an encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use in decoding video data by the video decoder 30. The video data memory 79 and the DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) , including Synchronous DRAM (SDRAM) , Magneto-resistive RAM (MRAM) , Resistive RAM (RRAM) , or other types of memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30, or off-chip relative to those components.
[0056] During the decoding process, the video decoder 30 receives an encoded video bitstream that represents video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 entropy decodes the bitstream to generate quantized coefficients, motion vector information or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to the prediction processing unit 81.
[0057] When the video frame is coded as an intra predictive coded (I) frame or for intra coded predictive blocks in other types of frames, the intra prediction unit 84 may generate prediction data for a video block of the current video frame based on a signaled intra prediction mode and reference data from previously decoded blocks of the current frame.
[0058] When the video frame is coded as an inter-predictive coded (i.e., B or P) frame, the motion compensation unit 82 produces one or more predictive blocks for a video block of the current video frame based on the motion vector information and other syntax elements received from the entropy decoding unit 80. Each of the predictive blocks may be produced from a reference frame within one of the reference frame lists. The video decoder 30 may construct the reference frame lists, List 0 and List 1, using default construction techniques based on reference frames stored in the DPB 92.
[0059] In some examples, when the video block is coded according to the IBC mode described herein, the IBC unit 85 produces predictive blocks for the current video block based on block vector information and other syntax elements received from the entropy decoding unit 80. The predictive blocks may be within a reconstructed region of the same picture as the current video block defined by the video encoder 20.
[0060] The motion compensation unit 82 and / or the IBC unit 85 determines prediction information for a video block of the current video frame by parsing the vector information and other syntax elements, and then uses the prediction information to produce the predictive blocks for the current video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode used to code video blocks of the video frame, an inter prediction frame type (e.g., B or P) , construction information for one or more of the reference frame lists for the frame, motion vectors for each inter predictive encoded video block of the frame, inter prediction status for each inter predictive coded video block of the frame, and other information to decode the video blocks in the current video frame.
[0061] Similarly, the IBC unit 85 may use some of the received syntax elements, e.g., a flag, to determine that the current video block was predicted using the IBC mode, construction information of which video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, block vectors for each IBC predicted video block of the frame, IBC prediction status for each IBC predicted video block of the frame, and other information to decode the video blocks in the current video frame.
[0062] The motion compensation unit 82 may also perform interpolation using the interpolation filters as used by the video encoder 20 during encoding of the video blocks to calculate interpolated values for sub-integer pixels of reference blocks. In this case, the motion compensation unit 82 may determine the interpolation filters used by the video encoder 20 from the received syntax elements and use the interpolation filters to produce predictive blocks.
[0063] Like the process of choosing a predictive block in a reference frame during inter-frame prediction of a video block, a set of rules needs to be adopted by both the video encoder 20 and the video decoder 30 for constructing a motion vector candidate list (also known as a “merge list” ) for a current block using those potential candidate motion vectors associated with spatially neighboring blocks and / or temporally co-located blocks of the current block and then selecting one member from the motion vector candidate list as a motion vector predictor for the current block. By doing so, there is no need to transmit the motion vector candidate list itself from the video encoder 20 to the video decoder 30 and an index of the selected motion vector predictor within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list for encoding and decoding the current block.
[0064] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameter calculated by the video encoder 20 for each video block in the video frame. The inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct the residual blocks in the pixel domain.
[0065] The summer 90 reconstructs decoded video block for the current video block by summing the residual block from the inverse transform processing unit 88 and a corresponding predictive block. An in-loop filter 91 such as deblocking filter, SAO filter, CCSAO filter and / or ALF may be positioned between the summer 90 and the DPB 92 to further process the decoded video block. In some examples, the in-loop filter 91 may be omitted, and the decoded video block may be directly provided by the summer 90 to the DPB 92. The decoded video blocks in a given frame are then stored in the DPB 92, which stores reference frames used for subsequent motion compensation of next video blocks. The DPB 92, or a memory device separate from the DPB 92, may also store decoded video for later presentation on a display device, such as the display device 34 of FIG. 1.
[0066] In a typical video coding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other instances, a frame may be monochrome and therefore includes only one two-dimensional array of luma samples.
[0067] As shown in FIG. 4A, the video encoder 20 (or more specifically the partition unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively in a raster scan order from left to right and from top to bottom. Each CTU is a largest logical coding unit and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set, such that all the CTUs in a video sequence have the same size being one of 128×128, 64×64, 32×32, and 16×16. But it should be noted that the present application is not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may comprise one CTB of luma samples, two corresponding CTBs of chroma samples, and syntax elements used to code the samples of the CTBs. The syntax elements describe properties of different types of units of a coded block of pixels and how the video sequence can be reconstructed at the video decoder 30, including inter or intra prediction, intra prediction mode, motion vectors, and / or other parameters. In monochrome pictures or pictures having three separate color planes, a CTU may comprise a single CTB and syntax elements used to code the samples of the CTB. A CTB may be an NxN block of samples.
[0068] To achieve a better performance, the video encoder 20 may recursively perform tree partitioning such as binary-tree partitioning, ternary-tree partitioning, quad-tree partitioning or a combination thereof on the CTBs of the CTU and divide the CTU into smaller CUs. As depicted in FIG. 4C, the 64x64 CTU 400 is first divided into four smaller CUs, each having a block size of 32x32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs of 16x16 by block size. The two 16x16 CUs 430 and 440 are each further divided into four CUs of 8x8 by block size. FIG. 4D depicts a quad-tree data structure illustrating the end result of the partition process of the CTU 400 as depicted in FIG. 4C, each leaf node of the quad-tree corresponding to one CU of a respective size ranging from 32x32 to 8x8. Like the CTU depicted in FIG. 4B, each CU may comprise a CB of luma samples and two corresponding CBs of chroma samples, and syntax elements used to code the samples of the CBs. In monochrome pictures or pictures having three separate color planes, a CU may comprise a single CB and syntax structures used to code the samples of the CB. It should be noted that the quad-tree partitioning depicted in FIGS. 4C and 4D is only for illustrative purposes and one CTU can be split into CUs to adapt to varying local characteristics based on quad / ternary / binary-tree partitions. In the multi-type tree structure, one CTU is partitioned by a quad-tree structure and each quad-tree leaf CU can be further partitioned by a binary and / or ternary tree structure. As shown in FIG. 4E, there are five possible partitioning types of a CB having a width W and a height H, i.e., quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.
[0069] In some implementations, the video encoder 20 may further partition a CB of a CU into one or more MxN PBs. A PB is a rectangular (square or non-square) block of samples on which the same prediction, inter, intra etc., is applied. A PU of a CU may comprise a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In monochrome pictures or pictures having three separate color planes, a PU may comprise a single PB and syntax structures used to predict the PB. The video encoder 20 may generate predictive luma, Cb, and Cr blocks for luma, Cb, and Cr PBs of each PU of the CU.
[0070] Furthermore, as illustrated in FIG. 4C, the video encoder 20 may use quad-tree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks respectively. A transform block is a rectangular (square or non-square) block of samples on which the same transform is applied. A TU of a CU may comprise a TB of luma samples, two corresponding TBs of chroma samples, and syntax elements used to transform the TBs. Thus, each TU of a CU may be associated with a TB, a Cb TB, and a Cr TB. In some examples, the luma TB associated with the TU may be a sub-block of the CU's luma residual block. The Cb TB may be a sub-block of the CU's Cb residual block. The Cr TB may be a sub-block of the CU's Cr residual block. In monochrome pictures or pictures having three separate color planes, a TU may comprise a single TB and syntax structures used to transform the samples of the TB.
[0071] In this section, bi-directional optical flow (BDOF) and its improvement methods are introduced. Principle of BDOF
[0072] Bi-directional optical flow (BDOF) technique is based on pixel level optical flow. Before introducing the design of BDOF in the VVC and ECM, the principle of BDOF is first introduced. Considering a pixel value at time t, the first order Taylor expansion is described as:
[0073] Under the optical flow assumption, the following equations are satisfied.
[0074] Let and equation (2) is rewritten as:
[0075] Equation (1) is converted as:
[0076] Noted that and are the motion speed along x and y directions and termed as:
[0077] Therefore, equation (4) becomes:
[0078] From the observation of equation (7) , to estimate the pixel value at time t from time t0, the and need to be estimated.
[0079] In the current VVC and ECM, BDOF is only applied to PU with true bi-prediction mode. Suppose that we have a forward reference picture at time t0 and a backward reference picture at time t1, and t-t0=t1-t=1.
[0080] So we have
[0081] For bi-prediction, the two reference samples are averaged as follow
[0082] Considering motion is along the trajectory, it is further assumed that and equation (9) becomes
[0083] To obtain Vx and Vy, bilateral matching is utilized by minimizing the following cost.
[0084] Where Ω is the region in the current PU. Design of BDOF in the VVC and ECM
[0085] The bi-directional optical flow (BDOF) tool is included in VVC. BDOF, previously referred to as BIO, was included in the JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of number of multiplications and the size of the multiplier.
[0086] BDOF is used to refine the bi-prediction signal of a CU at the 4×4 subblock level. BDOF is applied to a CU if it satisfies all the following conditions: · The CU is coded using “true” bi-prediction mode, i.e., one of the two reference pictures is prior to the current picture in display order and the other is after the current picture in display order · The distances (i.e., POC difference) from two reference pictures to the current picture are same · Both reference pictures are short-term reference pictures. · The CU is not coded using affine mode or the SbTMVP merge mode · CU has more than 64 luma samples · Both CU height and CU width are larger than or equal to 8 luma samples · BCW weight index indicates equal weight · WP is not enabled for the current CU · CIIP mode is not used for the current CU
[0087] BDOF is only applied to the luma component. As its name indicates, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 subblock, a motion refinement (vx, vy) is calculated by minimizing the difference between the L0 and L1 prediction samples. The motion refinement is then used to adjust the bi-predicted sample values in the 4x4 subblock. The following steps are applied in the BDOF process.
[0088] First, the horizontal and vertical gradients, and k=0, 1 , of the two prediction signals are computed by directly calculating the difference between two neighboring samples, i.e.,
[0089] where I (k) (i, j) are the sample value at coordinate (i, j) of the prediction signal in list k, k=0, 1, and shift1 is calculated based on the luma bit depth, bitDepth, as shift1=max (6, bitDepth-6) .
[0090] Then, the auto-and cross-correlation of the gradients, S1, S2, S3, S5 and S6, are calculated as where θ(i, j) = (I (1) (i, j) >>nb) - (I (0) (i, j) >>nb) (16)
[0091] where Ω is a 6×6 window around the 4×4 subblock, and the values of na and nb are set equal to min (1, bitDepth -11 ) and min (4, bitDepth -8 ) , respectively.
[0092] The motion refinement (vx, vy) is then derived using the cross-and auto-correlation terms using the following:
[0093] where th′BIO=2max (5, BD-7) . is the floor function, and
[0094] Based on the motion refinement and the gradients, the following adjustment is calculated for each sample in the 4×4 subblock:
[0095] Finally, the BDOF samples of the CU are calculated by adjusting the bi-prediction samples as follows: predBDOF (x, y) = (I (0) (x, y) +I (1) (x, y) +b (x, y) +οoffset) >>shift (20)
[0096] These values are selected such that the multipliers in the BDOF process do not exceed 15-bit, and the maximum bit-width of the intermediate parameters in the BDOF process is kept within 32-bit.
[0097] In order to derive the gradient values, some prediction samples I (k) (i, j) in list k (k=0, 1) outside of the current CU boundaries need to be generated. As depicted in FIG. 6, the BDOF in VVC uses one extended row / column around the CU’s boundaries. In order to control the computational complexity of generating the out-of-boundary prediction samples, prediction samples in the extended area (white positions) are generated by taking the reference samples at the nearby integer positions (using floor (. ) operation on the coordinates) directly without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate prediction samples within the CU (dotted positions) . These extended sample values are used in gradient calculation only. For the remaining steps in the BDOF process, if any sample and gradient values outside of the CU boundaries are needed, they are padded (i.e., repeated) from their nearest neighbors.
[0098] When the width and / or height of a CU are larger than 16 luma samples, it will be split into subblocks with width and / or height equal to 16 luma samples, and the subblock boundaries are treated as the CU boundaries in the BDOF process. The maximum unit size for BDOF process is limited to 16x16. For each subblock, the BDOF process could skipped. When the SAD of between the initial L0 and L1 prediction samples is smaller than a threshold, the BDOF process is not applied to the subblock. The threshold is set equal to (8 *W*H >> 1) , where W indicates the subblock width, and H indicates subblock height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 and L1 prediction samples calculated in DVMR process is re-used here.
[0099] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weight, then bi-directional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., the luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. When a CU is coded with symmetric MVD mode or CIIP mode, BDOF is also disabled. (1) Sample-based BDOF
[0100] In the ECM, sample-based BDOF is utilized. In the sample-based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, it is performed per sample.
[0101] The coding block is divided into 8×8 subblocks. For each subblock, whether to apply BDOF or not is determined by checking the SAD between the two reference subblocks against a threshold. If decided to apply BDOF to a subblock, for every sample in the subblock, a sliding 5×5 window is used and the existing BDOF process is applied for every sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bi-predicted sample value for the center sample of the window. Decoder side motion vector refinement (DMVR)
[0102] In order to increase the accuracy of the MVs of the merge mode, a bilateral-matching (BM) based decoder side motion vector refinement is applied in VVC. In bi-prediction operation, a refined MV is searched around the initial MVs in the reference picture list L0 and reference picture list L1. The BM method calculates the distortion between the two candidate blocks in the reference picture list L0 and list L1. As illustrated in FIG. 7, the SAD between the dashed blocks based on each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and used to generate the bi-predicted signal.
[0103] In VVC, the application of DMVR is restricted and is only applied for the CUs which are coded with following modes and features: · CU level merge mode with bi-prediction MV · One reference picture is in the past and another reference picture is in the future with respect to the current picture · The distances (i.e., POC difference) from two reference pictures to the current picture are same · Both reference pictures are short-term reference pictures · CU has more than 64 luma samples · Both CU height and CU width are larger than or equal to 8 luma samples · BCW weight index indicates equal weight · WP is not enabled for the current block · CIIP mode is not used for the current block
[0104] The refined MV derived by DMVR process is used to generate the inter prediction samples and also used in temporal motion vector prediction for future pictures coding. While the original MV is used in deblocking process and also used in spatial motion vector prediction for future CU coding. (1) Searching scheme
[0105] In DVMR, the search points are surrounding the initial MV and the MV offset obey the MV difference mirroring rule. In other words, any points that are checked by DMVR, denoted by candidate MV pair (MV0, MV1) obey the following two equations: MV0′=MV0+MVoffset (21) MV1′=MV1-MVoffset (22)
[0106] Where MVoffset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples from the initial MV. The searching includes the integer sample offset search stage and fractional sample refinement stage.
[0107] 25 points full search is applied for integer sample offset searching. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is smaller than a threshold, the integer sample stage of DMVR is terminated. Otherwise SADs of the remaining 24 points are calculated and checked in raster scanning order. The point with the smallest SAD is selected as the output of integer sample offset searching stage. To reduce the penalty of the uncertainty of DMVR refinement, it is proposed to favor the original MV during the DMVR process. The SAD between the reference blocks referred by the initial MV candidates is decreased by 1 / 4 of the SAD value.
[0108] The integer sample search is followed by fractional sample refinement. To save the calculational complexity, the fractional sample refinement is derived by using parametric error surface equation, instead of additional search with SAD comparison. The fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. When the integer sample search stage is terminated with center having the smallest SAD in either the first iteration or the second iteration search, the fractional sample refinement is further applied.
[0109] In parametric error surface based sub-pixel offsets estimation, the center position cost and the costs at four neighboring positions from the center are used to fit a 2-D parabolic error surface equation of the following form E(x, y) =A (x-xmin) 2+B (y-ymin) 2+C (23)
[0110] where (xmin, ymin) corresponds to the fractional position with the least cost and C corresponds to the minimum cost value. By solving the above equations by using the cost value of the five search points, the (xmin, ymin) is computed as: ymin= (E (0, -1) -E (0, 1) ) / (2 ( (E (0, -1) +E (0, 1) -2E (0, 0) ) ) (25)
[0111] The value of xmin and ymin are automatically constrained to be between –8 and 8 since all cost values are positive and the smallest value is E (0, 0) . This corresponds to half peal offset with 1 / 16th-pel MV accuracy in VVC. The computed fractional (xmin, ymin) are added to the integer distance refinement MV to get the sub-pixel accurate refinement delta MV. (2) Bilinear-interpolation and sample padding
[0112] In VVC, the resolution of the MVs is 1 / 16 luma samples. The samples at the fractional position are interpolated using a 8-tap interpolation filter. In DMVR, the search points are surrounding the initial fractional-pel MV with integer sample offset, therefore the samples of those fractional position need to be interpolated for DMVR search process. To reduce the calculation complexity, the bi-linear interpolation filter is used to generate the fractional samples for the searching process in DMVR. Another important effect is that by using bi-linear filter is that with 2-sample search range, the DVMR does not access more reference samples compared to the normal motion compensation process. After the refined MV is attained with DMVR search process, the normal 8-tap interpolation filter is applied to generate the final prediction. In order to not access more reference samples to normal MC process, the samples, which is not needed for the interpolation process based on the original MV but is needed for the interpolation process based on the refined MV, will be padded from those available samples. (3) Maximum DMVR processing unit
[0113] When the width and / or height of a CU are larger than 16 luma samples, it will be further split into subblocks with width and / or height equal to 16 luma samples. The maximum unit size for DMVR searching process is limit to 16x16. Multi-pass decoder-side motion vector refinement
[0114] A multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16x16 subblock within the coding block. In the third pass, MV in each 8x8 subblock is refined by applying bi-directional optical flow (BDOF) . The refined MVs are stored for both spatial and temporal motion vector prediction. (1) First pass - Block based bilateral matching MV refinement
[0115] In the first pass, a refined MV is derived by applying BM to a coding block. Similar to decoder-side motion vector refinement (DMVR) , in bi-prediction operation, a refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initiate MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.
[0116] BM performs local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to loop through the search range [–sHor, sHor] in horizontal direction and [–sVer, sVer] in vertical direction, wherein, the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0117] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW *cbH is greater than 64, mean-removal SAD (MRSAD) cost function is applied to remove the DC effect of distortion between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern and continue to search for the minimum cost, until it reaches the end of the search range.
[0118] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MVs after the first pass is then derived as: · MV0_pass1 = MV0 + deltaMV · MV1_pass1 = MV1 –deltaMV (2) Second pass - Subblock based bilateral matching MV refinement
[0119] In the second pass, a refined MV is derived by applying BM to a 16×16 grid subblock. For each subblock, a refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) , obtained on the first pass, in the reference picture list L0 and L1. The refined MVs (MV0_pass2 (sbIdx2) and MV1_pass2 (sbIdx2) ) are derived based on the minimum bilateral matching cost between the two reference subblocks in L0 and L1.
[0120] For each subblock, BM performs full search to derive integer sample precision intDeltaMV. The full search has a search range [–sHor, sHor] in horizontal direction and [–sVer, sVer] in vertical direction, wherein, the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0121] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference subblocks, as: bilCost = satdCost *costFactor. The search area (2*sHor + 1) * (2*sVer +1) is divided up to 5 diamond shape search regions shown on FIG. 8. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in the order starting from the center of the search area. In each region, the search points are processed in the raster scan order starting from the top left going to the bottom right corner of the region. When the minimum bilCost within the current search region is less than a threshold equal to sbW *sbH, the int-pel full search is terminated, otherwise, the int-pel full search continues to the next search region until all search points are examined. Additionally, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold that is equal to the area of the block, the search process terminates.
[0122] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2) . The refined MVs at second pass is then derived as: · MV0_pass2 (sbIdx2) = MV0_pass1 + deltaMV (sbIdx2) · MV1_pass2 (sbIdx2) = MV1_pass1 –deltaMV (sbIdx2) (3) Third pass - Subblock based bi-directional optical flow MV refinement
[0123] In the third pass, a refined MV is derived by applying BDOF to an 8×8 grid subblock. For each 8×8 subblock, BDOF refinement is applied to derive scaled Vx and Vy without clipping starting from the refined MV of the parent subblock of the second pass. The derived bioMv (Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0124] The refined MVs (MV0_pass3 (sbIdx3) and MV1_pass3 (sbIdx3) ) at third pass are derived as: · MV0_pass3 (sbIdx3) = MV0_pass2 (sbIdx2) + bioMv · MV1_pass3 (sbIdx3) = MV0_pass2 (sbIdx2) –bioMv
[0125] In all aforementioned sub-clauses, when wrap around motion compensation is enabled, the motion vectors shall be clipped with wrap around offset taken into consideration. (4) Fourth pass - Adaptive subblock based bi-directional optical flow MV refinement
[0126] In the fourth pass, a refined MV is derived by applying BDOF to an 4×4 or 8×8 or 16x16 grid subblock. When a block is smaller than 1024 pixels, the 4×4 grid subblock is used. Otherwise, 8×8 grid subblock is used. The MV of each subblock is refined in the same way as that used in third pass.
[0127] In all aforementioned sub-clauses, when wrap around motion compensation is enabled, the motion vectors shall be clipped with wrap around offset taken into consideration. It is noted that in ECM, the DMVR is extended to non-equal POC distance cases, and the mean removed equations are utilized to derive the BDOF MV refinement parameters as: (∑Gx. Gx+R1) *vx + ∑Gx. Gy *vy = ∑dI . Gx. → (ΣGx. Gx+R1) *vx + ΣGx. Gy *vy = ΣdI . Gx -dM . ΣGx ΣGx. Gy *vx + (ΣGy. Gy+R1) *vy= ΣdI . Gy → ∑Gx. Gy *vx + (∑Gy. Gy+R1) *vy = ΣdI . Gy -dM . ΣGy Affine subblock BDOF refinement
[0128] BDOF subblock MV refinement and sample adjustment is applied to an affine or SbTMVP coded block with subblock MC when BDOF condition is satisfied.
[0129] An affine coded block, e.g., affine regular merge mode, affine BM merge mode, affine AMVP mode, derives MVs for each 4×4 subblock from the affine model. The BDOF process starts with the 4×4 subblocks grouping with identical MVs. The first iteration of BDOF MV refinement is processed in 8x8 subblock grid as in ECM-10.0. When the grouped subblock size is less than 256, the second iteration of BDOF MV refinement is processed in 4×4 subblock grid, and otherwise in 8×8 subblock grid. When the grouped subblock size is 4xN or Nx4, the first iteration of BDOF MV refinement is bypassed.
[0130] In the current ECM, sample-based BDOF is utilized. For every sample in the subblock, a sliding 5×5 window is used and the BDOF process is applied for every sliding window to derive motion refinement. The following deficiencies that exist in the current BDOF technique are identified in this disclosure.
[0131] Firstly, in the ECM, fixed 5x5 sliding window is applied in BDOF which may not adapt to the diverse video characteristics.
[0132] Secondly, in the current VVC and ECM, BDOF is only applied to the prediction units satisfying the two conditions: The CU is coded using “true” bi-prediction mode, i.e., one of the two reference pictures is prior to the current picture in display order and the other is after the current picture in display order. Also, the distances (i.e., POC difference) from two reference pictures to the current picture are same. However, prediction sample refinement with optical flow for uni-prediction is not considered. Also, the BDOF is not applied to the cases of · True bi-prediction with non-equal distance between reference pictures to the current picture and · Bi-prediction with two reference pictures having smaller POC values than the current picture, i.e., low-delay case.
[0133] It should be noted that the following methods may be applied independently or combinedly. Adaptive BDOF window size
[0134] In this disclosure, adaptive sliding window size are proposed for BDOF. For each block, the sliding window size used to derive the motion refinement is adaptively decided and applied.
[0135] In one embodiment, the sliding window size are explicitly derived and signaled in the bitstream. At the encoder side, several sliding window size candidates are tested and selected for each prediction unit (PU) using rate-distortion optimization. The index for the optimal sliding window size is signaled in the bitstream.
[0136] In yet another embodiment, the sliding window size are implicitly derived at the decoder side with template matching, therefore no further signaling overhead is needed. The proposed template matching based sliding window size is illustrated in FIG. 9. Firstly, the predicted template of reference list 0 and list 1 are obtained ( and in FIG. 9) . Then BDOF process is applied to and with a certain sliding window size candidate. Finally, the template matching cost is calculated which measures the distance between the predicted block after BDOF process and the template of the current block XT. The sliding window size candidate which leads to the minimum template matching cost is selected and applied to the BDOF process of the current block. Uni-directional optical flow (UDOF)
[0137] In this disclosure, it is proposed to extend the optical flow-based sample refinement to uni-prediction, termed as uni-directional optical flow (UDOF) . According to equation (7) , to obtain the refined sample value at time t from time t0, and needs to be estimated. In the BDOF, bilateral matching is exploited to estimate and by minimizing the difference between the refined samples of the two directions. To enable optical flow-based sample refinement for uni-prediction, it is proposed to utilize template to estimate and
[0138] Denote the neighboring reconstructed template of the current PU at time t as the corresponding reference template at time t0 as IT (t0) . Applying equation (7) to the template, the refined reference template is obtained as follow.
[0139] Suppose that t-t0=1, equation (26) becomes
[0140] Then and are solved by minimizing the difference between IT (t) and
[0141] where ΩT represent the region in the template. The solved and are applied to the current PU, i.e., and
[0142] The refined samples for uni-prediction are obtained as follow. Template-based BDOF
[0143] In this disclosure, BDOF is extended to more general bi-prediction cases by exploiting template matching technique.
[0144] According to one or more embodiment of this disclosure, BDOF is applied to bi-prediction of low-delay cases, i.e., both the two reference pictures have smaller POC values than the current picture.
[0145] According to one or more embodiment of this disclosure, BDOF is applied to true bi-prediction with the distances from two reference pictures to the current picture are different.
[0146] According to one or more embodiment of this disclosure, it is proposed to apply template matching to solve Vx and Vy in equation (10) in BDOF. Denote the neighboring reconstructed template of the current PU at time t as the corresponding reference template at time t0 and t1 as IT (t0) and IT(t1) . By applying equation (10) to the template, the predicted template is obtained as follow.
[0147] Then and are solved by minimizing the difference between IT (t) and where ΩT represent the region in the template.
[0148] The solved and are applied to the current PU, i.e., and Then Vx and Vy are applied to equation (10) to obtain the refined samples.
[0149] It should be noted that the proposed TM-based BDOF is applicable to the following three cases. · Bi-prediction of low-delay cases, i.e., both the two reference pictures have smaller POC values than the current picture. · True bi-prediction with the distances from two reference pictures to the current picture are different. · True bi-prediction with the distances from two reference pictures to the current picture are the same. Adaptive optical model selection for BDOF
[0150] In the current VVC and ECM, the optical flow sample refinement processes of both directions are utilized for bi-predicted block by minimizing the bilateral matching cost. Due to the diversity of video content, such optical flow refinement method may be not always effective for certain prediction block.
[0151] According to the present disclosure, additional optical flow sample refinement models may be introduced to bi-prediction. More specifically, the optical flow sample refinement may be applied only to either the forward reference block or the backward reference block instead of both reference blocks, i.e., three optical flow sample refinement methods are defined. The bi-directional optical flow sample refinement model could be described as equation (9) . The uni-directional optical flow sample refinement method is conducted as the following equation.
[0152] where i indicates the direction the optical flow refinement is applied to.
[0153] In one embodiment, and are solved using bilateral matching as follow.
[0154] In yet another embodiment, and are solved using template matching as follow.
[0155] According to the present disclosure, the usage of the above-mentioned three optical flow sample refinement models could be explicitly signaled or implicitly derived. · In the explicit signaling method, the three models are checked at the encoder side and the model leading to the minimum rate-distortion cost is selected and signaled. · In the implicit derivation method, the three models are applied to the template, the model leading to minimum template cost is selected and applied to the current block. Template-based BDOF usage condition
[0156] According to the present disclosure, for prediction blocks satisfying the current BDOF condition, addition condition is added by using template. · Firstly, motion compensation is conducted for the template without BDOF using the motion information of the current block. The distance between the predicted template and the template of the current block is calculated as the template cost, denoted as cost1. · Secondly, motion compensation is conducted for the template with BDOF using the motion information of the current block. The distance between the predicted template and the template of the current block is calculated as the template cost, denoted as cost2. · If cost1 > cost2, BDOF is not applied to the current block and vice versa.
[0157] According to the present disclosure, some BDOF usage conditions may be removed or relaxed when template-based usage condition is exploited. · According to one additional example of the disclosure, the BDOF may be applicable for prediction block coded with reference pictures from the same directions. · According to the second additional example of the disclosure, the BDOF may be applicable for prediction block of which the distances (i.e., POC difference) from two reference pictures to the current picture are different. BDOF with unequal reference distance
[0158] According to one or more embodiments of this disclosure, BDOF with unequal reference picture distances is enabled, where the distances (i.e., POC differences) from two reference pictures to the current picture are different.
[0159] In the first method, POC distance is considered when solving and applying vx and vy in
[0160] equation (12) ~ equation (20) . Denote the POC of the current picture as poc, the POC values of the two reference pictures as poc0 and poc1. Two scaling factors are calculated using the two POC distances as follow. · if abs (poc1 -poc) > abs (poc0 -poc) , s0 = (abs (poc0 -poc) << S) / abs (poc1 -poc) and s1= (1<<S) · if abs (poc1 -poc) < abs (poc0 -poc) , s1 = (abs (poc1 -poc) << S) / abs (poc0 -poc) and s0= (1<<S)
[0161] where abs (. ) is used to calculate absolute value, S is the shift number to represent the scaling factors with integer.
[0162] and in equation (12) are left shifted by S.
[0163] In addition, θ (i, j) in equation (16) is also left shifted by S. θ(i, j) = ( (I (1) (i, j) >>nb) - (I (0) (i, j) >>nb) ) <<S (33)
[0164] The final prediction is obtained in the following manner. predBDOF (x, y) = ( (I (0) (x, y) +I (1) (x, y) +οoffset) <<S+b (x, y) ) >> (shift+S) (34)
[0165] In the second method, vx and vy are firstly solved using equation (12) ~ equation (18) as in the ECM and then scaled using POC distances. Two scaling factors are calculated using the two POC distances as follow. · if abs (poc1 -poc) > abs (poc0 -poc) , s0 = (abs (poc0 -poc) << S) / abs (poc1 -poc) and s1= (1<<S) · if abs (poc1 -poc) < abs (poc0 -poc) , s1 = (abs (poc1 -poc) << S) / abs (poc0 -poc) and s0= (1<<S)
[0166] where S is the shift number to represent the scaling factors with integer.
[0167] vx and vy are scaled with s0 and s1 to obtain the motion refinement for the two reference pictures.
[0168] The refinement value is calculated as follow. BDOF for Low-delay B
[0169] In the current VVC and ECM, BDOF is only applied to the cases that two reference pictures are from different directions, i.e., one of the two reference pictures is prior to the current picture in display order and the other is after the current picture in display order.
[0170] In this disclosure, BDOF is extended to the cases that two reference pictures are from the same directions. More specifically, both the two reference pictures are prior to the current picture in display order, i.e., low-delay case.
[0171] In one or more embodiments, BDOF is applied to blocks coded using non-true bi-prediction mode, i.e., both the two reference pictures are prior to the current picture in display order.
[0172] In one or more embodiments, BDOF is applied to blocks coded using non-true bi-prediction mode with different reference pictures, i.e., both the two reference pictures are prior to the current picture in display order and in the meantime the two reference pictures are different pictures.
[0173] To further optimize BDOF in low-delay B case, the following constraints are imposed. It should be noted that the following constraints could be utilized separately or combinedly.
[0174] In one or more embodiments, the affine subblock BDOF refinement is not enabled for low-delay B.
[0175] In one or more embodiments, affine subblock BDOF refinement is enabled for low-delay B, but BDOF is not enabled for subblock ATMVP merge mode.
[0176] In one or more embodiments, affine subblock BDOF refinement is enabled for low-delay B, but BDOF is only applied to subblock ATMVP merge mode, while not enabled for affine mode.
[0177] In one or more embodiments, affine subblock BDOF refinement is enabled for low-delay B, but BDOF is not enabled for affine BM merge mode.
[0178] In one or more embodiments, affine subblock BDOF refinement is enabled for low-delay B, but BDOF is not enabled for affine merge mode.
[0179] In one or more embodiments, affine subblock BDOF refinement is enabled for low-delay B with only sample refinement, the refined motion field is not saved.
[0180] In the current ECM, affine subblock BDOF refinement is adopted in which MV of each 4x4 subblock is refined in two iterations. The BDOF sample refinement is then applied to the subblock after MV refinement. In this disclosure, the MV refinement steps are reduced.
[0181] In one or more embodiments, the first iteration of subblock MV refinement is always bypassed.
[0182] In one or more embodiments, the second iteration of subblock MV refinement is always bypassed.
[0183] In one or more embodiments, both the first and second iterations of subblock MV refinement is always bypassed.
[0184] In addition, the following embodiments are proposed in this disclosure to improve the performance of BDOF. The following embodiments could be used separately or combinedly. Also, the following embodiments are not limited to the CU with reference pictures from different directions (one reference picture has smaller POC than the current picture and the other reference picture has larger POC than the current picture) .
[0185] In one or more embodiments, BDOF is not enabled if the CU is coded using GPM mode. In yet other embodiments, BDOF is enabled for CU coded with GPM mode.
[0186] In one or more embodiments, BDOF is not enabled if the CU is coded using BCW mode. In yet other embodiments, BDOF is enabled for CU coded with BCW mode.
[0187] In one or more embodiments, BDOF is not enabled if the CU is coded using MMVD merge mode. In yet other embodiments, BDOF is enabled for CU coded with MMVD merge mode.
[0188] In one or more embodiments, BDOF is not enabled if the CU is coded using AMVP merge mode. In yet other embodiments, BDOF is enabled for CU coded with AMVP merge mode.
[0189] In one or more embodiments, BDOF is not enabled if the CU is coded using AMVP with sbTMVP mode. In yet other embodiments, BDOF is enabled for CU coded with AMVP with sbTMVP mode.
[0190] In one or more embodiments, BDOF is not enabled if the CU is coded using template-based merge mode. In yet other embodiments, BDOF is enabled for CU coded with template-based merge mode. Block size constraint for BDOF LDB
[0191] To improve the performance of BDOF LDB, block size constraints are imposed to the blocks satisfying the BDOF conditions.
[0192] In one or more embodiments, BDOF LDB is only applied to the blocks of which the area is larger than a threshold.
[0193] In one or more embodiments, BDOF LDB is only applied to the blocks of which the area is smaller than a threshold.
[0194] In one or more embodiments, BDOF LDB is disabled to the blocks of which the width or height is larger than a threshold.
[0195] In one or more embodiments, BDOF LDB is disabled to the blocks of which the width or height is smaller than a threshold. BDOF LDB with subblock MV refinement
[0196] In the current ECM, BDOF is also applied to subblock mode, including affine mode and sbTMVP merge mode. More specifically, BDOF for affine and sbTMVP consists of two iterations of BDOF MV refinement process followed by one sample refinement process, while BDOF for blocks coded with non-subblock modes only consists of BDOF sample refinement process. To better exploit BDOF, it is proposed to apply subblock-based BDOF refinement for blocks with non-subblock modes. For a NxN block, it is firstly divided into several MxM subblocks. For each subblock, several iterations of BDOF MV refinement processes are firstly applied, followed by the BDOF sample refinement process.
[0197] In one or more embodiments of this disclosure, the subblock size is 4x4.
[0198] In one or more embodiments of this disclosure, the subblock size is the same with the size of the prediction block.
[0199] In one or more embodiment of this disclosure, one iteration of BDOF MV refinement process is applied followed by BDOF sample refinement.
[0200] In one or more embodiment of this disclosure, two iterations of BDOF MV refinement processes are applied followed by BDOF sample refinement. Syntax control for BDOF LDB
[0201] In this disclosure, embodiments are proposed to adaptively decide whether BDOF LDB is enabled.
[0202] In one or more embodiment of this disclosure, it is proposed to introduce additional syntax for BDOF LDB at Sequence Parameter Set (SPS) level. An additional flag is signaled in SPS. If the flag is true, BDOF LDB is disabled for the sequence, and vice versa.
[0203] In one or more embodiment of this disclosure, it is proposed to introduce additional syntax for BDOF LDB at picture level. An additional flag denoted as is signaled in the picture header. If the flag is true, BDOF LDB is disabled for the picture, and vice versa.
[0204] In one or more embodiment of this disclosure, it is proposed to introduce additional syntax for BDOF LDB at slice level. An additional flag denoted as is signaled in the slice header. If the flag is true, BDOF LDB is disabled for the slice, and vice versa.
[0205] In one or more embodiment of this disclosure, it is proposed to introduce additional syntax for BDOF LDB at CTU level. An additional flag denoted as is signaled in the CTU level. If the flag is true, BDOF LDB is disabled for the CTU, and vice versa.
[0206] In one or more embodiment of this disclosure, it is proposed to introduce additional syntax for BDOF LDB at CU level. An additional flag denoted as is signaled in the CU. If the flag is true, BDOF LDB is disabled for the CU, and vice versa. Decoder side BDOF LDB control
[0207] Due to the diversity of video characteristics, the current conditions of BDOF LDB may be not sufficient. To make the prediction of BDOF LDB more precise, it is proposed to implicitly decide whether BDOF LDB is enabled for a CU.
[0208] In one or more embodiment of this disclosure, if the distance between the prediction after and before BDOF operation is larger than a threshold, then BDOF is not applied to the CU. The distance could be measured with Sum of Square Error (SSE) , Sum of Abstract Difference (SAD) or other metrics. DMVR for Low-delay B
[0209] Similar to BDOF, in the current VVC and ECM, DMVR is only applied to the cases that the two reference pictures are from different directions, i.e., one of the two reference pictures is prior to the current picture in display order and the other is after the current picture in display order.
[0210] In this disclosure, to exploit the benefit of DMVR for bi-prediction, it is proposed to extend DMVR to low-delay cases, i.e., both reference pictures are from the same directions.
[0211] In one or more embodiments, DMVR is applied to blocks coded using non-true bi-prediction mode, i.e., both the two reference pictures are prior to the current picture in display order.
[0212] In yet another embodiment, DMVR is applied to blocks coded using non-true bi-prediction mode with different reference pictures, i.e., both the two reference pictures are prior to the current picture in display order and in the meantime the two reference pictures are different pictures. Additional conditions for DMVR and BDOF
[0213] In the current VVC and ECM, the usage of DMVR / BDOF is used under some conditions, like block size, prediction modes. However, due to the diversity of video content, these conditions maybe not enough. To further improve the compression performance of DMVR and BDOF, some additional conditions are proposed. It should be noted that these conditions can be applied separately or combinedly.
[0214] In the first embodiment of the additional conditions for DMVR / BDOF, motion vectors in both directions are utilized to decide whether DMVR / BDOF is applied for a certain PU. · In the one example, if abs (mvHor0 + mvHor1) > threshold or abs (mvVer0 + mvVer1) > threshold, then DMVR / BDOF is not applied to the current PU. Where mvHor0 and mvVer0 are the horizontal and vertical part of the motion vector in reference picture list 0; mvHor1 and mvVer1 are the horizontal and vertical part of the motion vector in reference picture list 1; · In yet another example, if abs (mvHor0 + mvHor1) + abs (mvVer0 + mvVer1) > threshold, then DMVR / BDOF is not applied to the current PU. Where mvHor0 and mvVer0 are the horizontal and vertical part of the motion vector in reference picture list 0; mvHor1 and mvVer1 are the horizontal and vertical part of the motion vector in reference picture list 1.
[0215] In the second embodiment of the additional conditions for DMVR / BDOF, the distance of the two reference blocks is used to decide whether DMVR / BDOF is applied to the current block. More specifically, if the distance between the two reference blocks is larger than a threshold, DMVR / BDOF is not applied to the current block.
[0216] In the third embodiment of the additional conditions for DMVR / BDOF, the usage of DMVR / BDOF is decided at the granularity of subblock instead of PU. Each PU is divided into several subblocks, for each subblock, the distance of the two reference blocks is used to decide whether DMVR / BDOF is applied to the current subblock. More specifically, if the distance two reference subblocks is larger than a threshold, DMVR / BDOF is not applied to the current subblock.
[0217] In the fourth embodiment of the additional conditions for DMVR / BDOF, the gradient of the two reference blocks is used to decide whether DMVR / BDOF is applied to the current block. More specifically, the horizontal and vertical gradients of each reference block are calculated, denoted as gX0, gY0, gX1 and gY1. The distances of the gradients at position k in the two reference blocks are calculated as abs (gX0 (k) -gX1 (k) ) + abs (gY0 (k) -gY1 (k) ) . The total distance of the gradients is calculated as the summation of the gradient distance for all the positions in the current block. If the total gradient distance is larger than a threshold, then DMVR / BDOF is not applied to the current block.
[0218] In the fifth embodiment of the additional conditions for DMVR / BDOF, the usage of DMVR / BDOF is decided at the granularity of subblock instead of PU. More specifically, the gradient of the two reference subblocks is used to decide whether DMVR / BDOF is applied to the current subblock. The horizontal and vertical gradients of each reference subblock are calculated, denoted as gX0, gY0, gX1 and gY1. The distances of the gradients at position k in the two reference blocks are calculated as abs (gX0 (k) -gX1 (k) ) + abs (gY0 (k) -gY1 (k) ) . The total distance of the gradients for each subblock is calculated as the summation of the gradient distance for all the positions in the current subblock. If the total gradient distance within the subblock is larger than a threshold, then DMVR / BDOF is not applied to the current subblock.
[0219] In one or more embodiment of this disclosure, if the current CU is coded with affine merge mode and the merge candidate is the TMVP merge candidate, BDOF / DMVR is not applied to the CU.
[0220] In one or more embodiment of this disclosure, BDOF LDB is decided based on the angle of two motion vectors of the two reference blocks. The angle of the two motion vectors is calculated as follow, where and are the motion vectors. · In the first example, if θ is larger than a threshold, BDOF is not applied to the current CU.
[0221] In the second example, if θ is larger than a threshold and the magnitude of the difference between and is larger than another threshold, BDOF is not applied to the current CU. Merge index based conditions for BDOF / DMVR
[0222] According to one or more embodiment of this disclosure, the usage BDOF / DMVR is decided based on the merge index. Below are some examples. · In the first example of merge index based BDOF / DMVR, if the merge index of the current PU has the parity of 0, BDOF / DMVR is not enabled for the PU. · In the second example of merge index based BDOF / DMVR, if the merge index of the current PU has the parity of 1, BDOF / DMVR is not enabled for the PU. · In the third example of merge index based BDOF / DMVR, if the merge index of the current PU has the parity of 0 and the current PU is coded with affine merge mode, BDOF / DMVR is not enabled for the PU. · In the fourth example of merge index based BDOF / DMVR, if the merge index of the current PU has the parity of 1 and the current PU is coded with affine merge mode, BDOF / DMVR is not enabled for the PU. · In the fifth example of merge index based BDOF / DMVR, if the merge index of the current PU has the parity of 0 and the current PU is coded with non-affine merge mode, BDOF / DMVR is not enabled for the PU. · In the sixth example of merge index based BDOF / DMVR, if the merge index of the current PU has the parity of 1 and the current PU is coded with non-affine merge mode, BDOF / DMVR is not enabled for the PU.
[0223] In the first method, merge index is always considered as additional conditions to decide the usage of BDOF / DMVR.
[0224] In the second method, if sps_alt_cost_enabled_flag is false, the usage of BDOF / MDVR does not depend on merge index. Otherwise, if sps_alt_cost_enabled_flag is true, the usage of BDOF / DMVR is decided based on merge index as described above.
[0225] FIG. 10 is a flow chart illustrating a method 1000 for video decoding in accordance with some implementations of the present disclosure. The method 1000 may be performed by a video decoder, for example, the video decoder 30.
[0226] As shown in FIG. 10, the method 1000 may include steps 1010 and 1020. However, the present disclosure does not limit to this. For example, any of steps 1010 and 1020 may be replaced with other steps as described as follows.
[0227] In step 1010, the video decoder may determine whether one or more conditions are met for a current block in a current picture. The current block may be or correspond to a Coding Tree Unit (CTU) , a Coding Units (CU) , a Prediction Unit (PU) or a Transform Unit (TU) and / or may be or correspond to a corresponding block, e.g. a Coding Tree Block (CTB) , a Coding Block (CB) , a Prediction Block (PB) or a Transform Block (TB) and / or to a sub-block.
[0228] The one or more conditions may also be referred to “additional conditions” in the sections before.
[0229] In some implementations, the step 1010 may include: determining an angle of two motion vectors of two reference blocks associated with the current block; and determining whether the one or more conditions are met based on the angle of the two motion vectors of the two reference blocks.
[0230] In some implementations, the angle may be determined as: wherein and are the two motion vectors, θ is the angle of the two motion vectors.
[0231] In some implementations, the step 1010 may further include: determining a magnitude of a difference between the two motion vectors; and determining whether the one or more conditions are met based on the magnitude of the difference between the two motion vectors.
[0232] In some implementations, the step 1010 may include: determining whether the one or more conditions are met based on a coded mode of the current block.
[0233] In some implementations, the step 1010 may include: determining whether the one or more conditions are met based on a merge index.
[0234] In some implementations, the determining whether the one or more conditions are met based on a merge index may include: in response to determining that a syntax element has a first syntax value, determining whether the one or more conditions are met based on the merge index. For example, if a syntax element has a first syntax value, whether the one or more conditions are met is determined based on the merge index. If the syntax element has a second syntax value, whether the one or more conditions are met is not determined based on the merge index.
[0235] In some implementations, the syntax element may include: a sps_alt_cost_enabled_flag.
[0236] In step 1020, in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block.
[0237] In some implementations, the step 1020 may include: in response to determining that the angle is larger than a first threshold, not applying the Bi-directional optical flow refinement for the current block.
[0238] In some implementations, the step 1020 may include: in response to determining that the angle is larger than a second threshold and the magnitude of the difference between the two motion vectors is larger than a third threshold, not applying the Bi-directional optical flow refinement for the current block.
[0239] In some implementations, the step 1020 may include: in response to determining that the current block is coded with an affine merge mode and a merge candidate of the current block is a TMVP merge candidate, not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block.
[0240] In some implementations, the step 1020 may include: not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following: the merge index of the current block has a parity of 0; the merge index of the current block has a parity of 1. In some implementations, the current block comprises a Prediction Unit.
[0241] In some implementations, the step 1020 may include: not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following: the merge index of the current block has a parity of 0 and the current block is coded with affine merge mode; the merge index of the current block has a parity of 1 and the current block is coded with affine merge mode; the merge index of the current block has a parity of 0 and the current block is coded with non-affine merge mode; and the merge index of the current block has a parity of 1 and the current block is coded with non-affine merge mode. In some implementations, the current block comprises a Prediction Unit.
[0242] In accordance with implementations of the present disclosure, the compression performance of DMVR and / or BDOF can be further improved.
[0243] Other implementation details of the method 1000 can be referred to the sections of “Additional conditions for DMVR and BDOF” and “Merge index based conditions for BDOF / DMVR” .
[0244] FIG. 11 is a flow chart illustrating a method 1100 for video encoding in accordance with some implementations of the present disclosure. The method 1100 may be performed by a video encoder, for example, the video encoder 20. As shown in FIG. 11, the method 1100 comprises steps 1110 and 1120. However, the present disclosure does not limit to this. For example, any of steps 1110 and 1120 may be replaced with other steps as described as follows.
[0245] In step 1110, the video encoder determines whether one or more conditions are met for a current block in a current picture. The current block may be or correspond to a CTU, a CU, a PU or a TU and / or may be or correspond to a corresponding block, e.g. a CTB, a CB, a PB or a TB and / or to a sub-block.
[0246] The one or more conditions may also be referred to “additional conditions” in the sections before.
[0247] In some implementations, the step 1110 may include: determining an angle of two motion vectors of two reference blocks associated with the current block; and determining whether the one or more conditions are met based on the angle of the two motion vectors of the two reference blocks.
[0248] In some implementations, the angle may be determined as: wherein and are the two motion vectors, θ is the angle of the two motion vectors.
[0249] In some implementations, the step 1110 may further include: determining a magnitude of a difference between the two motion vectors; and determining whether the one or more conditions are met based on the magnitude of the difference between the two motion vectors.
[0250] In some implementations, the step 1110 may include: determining whether the one or more conditions are met based on a coded mode of the current block k.
[0251] In some implementations, the step 1110 may include: determining whether the one or more conditions are met based on a merge index.
[0252] In some implementations, the determining whether the one or more conditions are met based on a merge index may include: in response to determining that a syntax element has a first syntax value, determining whether the one or more conditions are met based on the merge index. For example, if a syntax element has a first syntax value, whether the one or more conditions are met is determined based on the merge index. If the syntax element has a second syntax value, whether the one or more conditions are met is not determined based on the merge index.
[0253] In some implementations, the syntax element comprises a sps_alt_cost_enabled_flag.
[0254] In step 1120, in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block.
[0255] In some implementations, the step 1120 may include: in response to determining that the angle is larger than a first threshold, not applying the Bi-directional optical flow refinement for the current block.
[0256] In some implementations, the step 1120 may include: in response to determining that the angle is larger than a second threshold and the magnitude of the difference between the two motion vectors is larger than a third threshold, not applying the Bi-directional optical flow refinement for the current block.
[0257] In some implementations, the step 1120 may include: in response to determining that the current block is coded with an affine merge mode and a merge candidate of the current block is a TMVP merge candidate, not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block.
[0258] In some implementations, the step 1120 may include: not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following: the merge index of the current block has a parity of 0; the merge index of the current block has a parity of 1. In some implementations, the current block comprises a Prediction Unit.
[0259] In some implementations, the step 1120 may include: not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following: the merge index of the current block has a parity of 0 and the current block is coded with affine merge mode; the merge index of the current block has a parity of 1 and the current block is coded with affine merge mode; the merge index of the current block has a parity of 0 and the current block is coded with non-affine merge mode; and the merge index of the current block has a parity of 1 and the current block is coded with non-affine merge mode. In some implementations, the current block comprises a Prediction Unit.
[0260] In accordance with implementations of the present disclosure, the compression performance of DMVR and / or BDOF can be further improved.
[0261] Other implementation details of the method 1100 can be referred to the sections of “Additional conditions for DMVR and BDOF” and “Merge index based conditions for BDOF / DMVR” .
[0262] FIG. 5 shows a computing environment 1610 coupled with a user interface 1650. The computing environment 1610 can be part of a data processing server. The computing environment 1610 includes a processor 1620, a memory 1630, and an Input / Output (I / O) interface 1640.
[0263] The processor 1620 typically controls overall operations of the computing environment 1610, such as the operations associated with display, data acquisition, data communications, and image processing. The processor 1620 may include one or more processors to execute instructions to perform all or some of the steps in the above-described methods. Moreover, the processor 1620 may include one or more modules that facilitate the interaction between the processor 1620 and other components. The processor may be a Central Processing Unit (CPU) , a microprocessor, a single chip machine, a Graphical Processing Unit (GPU) , or the like.
[0264] The memory 1630 is configured to store various types of data to support the operation of the computing environment 1610. The memory 1630 may include predetermined software 1632. Examples of such data includes instructions for any applications or methods operated on the computing environment 1610, video datasets, image data, etc. The memory 1630 may be implemented by using any type of volatile or non-volatile memory devices, or a combination thereof, such as a Static Random Access Memory (SRAM) , an Electrically Erasable Programmable Read-Only Memory (EEPROM) , an Erasable Programmable Read-Only Memory (EPROM) , a Programmable Read-Only Memory (PROM) , a Read-Only Memory (ROM) , a magnetic memory, a flash memory, a magnetic or optical disk.
[0265] The I / O interface 1640 provides an interface between the processor 1620 and peripheral interface modules, such as a keyboard, a click wheel, buttons, and the like. The I / O interface 1640 can be coupled with an encoder and decoder.
[0266] In an embodiment, there is also provided a non-transitory computer-readable storage medium or a computer program product comprising a plurality of programs, for example, in the memory 1630, executable by the processor 1620, for performing the encoding or decoding method described above and / or storing a bitstream which is generated by the encoding method described above and / or is to be decoded by the decoding method described above. For example, the computer program product may include the non-transitory computer-readable storage medium. In one example, the plurality of programs may be executed by the processor 1620 to receive (for example, from the video encoder 20 in FIG. 2) a bitstream or data stream including encoded video information (for example, video blocks representing encoded video frames, and / or associated one or more syntax elements, etc. ) , and may also be executed by the processor 1620 to perform the decoding method described above according to the received bitstream or data stream. In another example, the plurality of programs may be executed by the processor 1620 to perform the encoding method described above to encode video information (for example, video blocks representing video frames, and / or associated one or more syntax elements, etc. ) into a bitstream or data stream, and may also be executed by the processor 1620 to transmit the bitstream or data stream (for example, to the video decoder 30 in FIG. 3) or store the bitstream or data stream. Alternatively, the non-transitory computer-readable storage medium or the computer program product may have stored therein the bitstream or data stream.
[0267] In an embodiment, there is provided a bitstream (for example, comprising encoded video information) which is generated by the encoding method described above and / or is to be decoded by the decoding method described above.
[0268] In an embodiment, there is also provided a computing device comprising one or more processors (for example, the processor 1620) ; and the non-transitory computer-readable storage medium or the memory 1630 having stored therein a plurality of programs executable by the one or more processors, wherein the one or more processors, upon execution of the plurality of programs, are configured to perform the above-described methods. In an example, the one or more processors, upon execution of the plurality of programs, are configured to perform the above-described encoding method to generate a bitstream, and the computing device may further comprise a transmitter, configured to transmit the bitstream. In an alternative example, the one or more processors, upon execution of the plurality of programs, are configured to perform the above-described encoding method to generate a bitstream, and transmit or store the bitstream. In an example, the bitstream is to be decoded by the above-described decoding method.
[0269] In an embodiment, the computing environment 1610 may be implemented with one or more ASICs, DSPs, Digital Signal Processing Devices (DSPDs) , Programmable Logic Devices (PLDs) , FPGAs, GPUs, controllers, micro-controllers, microprocessors, or other electronic components, for performing the above methods.
[0270] In an embodiment, there is also provided a method for storing a bitstream, comprising storing the bitstream on a non-transitory computer-readable storage medium, wherein the bitstream is generated by the encoding method described above and / or is to be decoded by the decoding method described above. In an embodiment, there is also provided a method for storing a bitstream or a method for encoding video data, comprising: performing the encoding method described above to generate a bitstream, and storing the bitstream on a non-transitory computer-readable storage medium. In an example, the bitstream is to be decoded by the decoding method described above.
[0271] In an embodiment, there is also provided a method for transmitting a bitstream which is generated by the encoding method described above and / or is to be decoded by the decoding method described above. In an embodiment, there is also provided a method for transmitting a bitstream or a method for encoding video data, comprising: performing the encoding method described above to generate a bitstream, and transmitting the bitstream to a decoder. In an example, the bitstream is to be decoded by the decoding method described above. In an embodiment, there is also provided a method for receiving a bitstream which is generated by the encoding method described above and / or is to be decoded by the decoding method described above.
[0272] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those of ordinary skill in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.
[0273] Unless specifically stated otherwise, an order of steps of the method according to the present disclosure is only intended to be illustrative, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but may be changed according to practical conditions. In addition, at least one of the steps of the method according to the present disclosure may be adjusted, combined or deleted according to practical requirements.
[0274] The examples were chosen and described in order to explain the principles of the disclosure and to enable others skilled in the art to understand various implementations of the disclosure and to best utilize the underlying principles and various implementations with various modifications as are suited to the particular use contemplated. Therefore, it is to be understood that the scope of the disclosure is not to be limited to the specific examples of the implementations disclosed and that modifications and other implementations are intended to be included within the scope of the present disclosure.
Claims
1.A method for video decoding, comprising:determining whether one or more conditions are met for a current block in a current picture; andin response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block.2.The method of claim 1, wherein the determining whether one or more conditions are met for a current block in a current picture comprises:determining an angle of two motion vectors of two reference blocks associated with the current block; anddetermining whether the one or more conditions are met based on the angle of the two motion vectors of the two reference blocks.3.The method of claim 2, wherein the angle is determined as: whereinandare the two motion vectors, θ is the angle of the two motion vectors.4.The method of claim 2, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:in response to determining that the angle is larger than a first threshold, not applying the Bi-directional optical flow refinement for the current block.5.The method of claim 2, wherein the determining whether one or more conditions are met for a current block in a current picture further comprises:determining a magnitude of a difference between the two motion vectors; anddetermining whether the one or more conditions are met based on the magnitude of the difference between the two motion vectors.6.The method of claim 5, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:in response to determining that the angle is larger than a second threshold and the magnitude of the difference between the two motion vectors is larger than a third threshold, not applying the Bi-directional optical flow refinement for the current block.7.The method of claim 1, wherein the determining whether one or more conditions are met for a current block in a current picture comprises:determining whether the one or more conditions are met based on a coded mode of the current block.8.The method of claim 7, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:in response to determining that the current block is coded with an affine merge mode and a merge candidate of the current block is a TMVP merge candidate, not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block.9.The method of claim 1, wherein the determining whether one or more conditions are met for a current block in a current picture comprises:determining whether the one or more conditions are met based on a merge index.10.The method of claim 9, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following:the merge index of the current block has a parity of 0;the merge index of the current block has a parity of 1.11.The method of claim 7, wherein determining whether one or more conditions are met for a current block in a current picture further comprises:determining whether the one or more conditions are met based on a merge index.12.The method of claim 11, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following:the merge index of the current block has a parity of 0 and the current block is coded with affine merge mode;the merge index of the current block has a parity of 1 and the current block is coded with affine merge mode;the merge index of the current block has a parity of 0 and the current block is coded with non-affine merge mode; andthe merge index of the current block has a parity of 1 and the current block is coded with non-affine merge mode.13.The method of claim 9 or 11, wherein the determining whether the one or more conditions are met based on a merge index comprises:in response to determining that a syntax element has a first syntax value, determining whether the one or more conditions are met based on the merge index.14.The method of claim 13, wherein the syntax element comprises a sps_alt_cost_enabled_flag.15.A method for video encoding, comprising:determining whether one or more conditions are met for a current block in a current picture; andin response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block.16.The method of claim 15, wherein the determining whether one or more conditions are met for a current block in a current picture comprises:determining an angle of two motion vectors of two reference blocks associated with the current block; anddetermining whether the one or more conditions are met based on the angle of the two motion vectors of the two reference blocks.17.The method of claim 16, wherein the angle is determined as: whereinandare the two motion vectors, θ is the angle of the two motion vectors.18.The method of claim 16, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:in response to determining that the angle is larger than a first threshold, not applying the Bi-directional optical flow refinement for the current block.19.The method of claim 16, wherein the determining whether one or more conditions are met for a current block in a current picture further comprises:determining a magnitude of a difference between the two motion vectors; anddetermining whether the one or more conditions are met based on the magnitude of the difference between the two motion vectors.20.The method of claim 19, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:in response to determining that the angle is larger than a second threshold and the magnitude of the difference between the two motion vectors is larger than a third threshold, not applying the Bi-directional optical flow refinement for the current block.21.The method of claim 15, wherein the determining whether one or more conditions are met for a current block in a current picture comprises:determining whether the one or more conditions are met based on a coded mode of the current block.22.The method of claim 21, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:in response to determining that the current block is coded with an affine merge mode and a merge candidate of the current block is a TMVP merge candidate, not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block.23.The method of claim 15, wherein the determining whether one or more conditions are met for a current block in a current picture comprises:determining whether the one or more conditions are met based on a merge index.24.The method of claim 23, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following:the merge index of the current block has a parity of 0;the merge index of the current block has a parity of 1.25.The method of claim 21, wherein determining whether one or more conditions are met for a current block in a current picture further comprises:determining whether the one or more conditions are met based on a merge index.26.The method of claim 25, wherein the in response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block comprises:not applying at least one of the Bi-directional optical flow refinement and the Decoder side motion vector refinement for the current block based on any one of the following:the merge index of the current block has a parity of 0 and the current block is coded with affine merge mode;the merge index of the current block has a parity of 1 and the current block is coded with affine merge mode;the merge index of the current block has a parity of 0 and the current block is coded with non-affine merge mode; andthe merge index of the current block has a parity of 1 and the current block is coded with non-affine merge mode.27.The method of claim 23 or 25, wherein the determining whether the one or more conditions are met based on a merge index comprises:in response to determining that a syntax element has a first syntax value, determining whether the one or more conditions are met based on the merge index.28.The method of claim 27, wherein the syntax element comprises a sps_alt_cost_enabled_flag.29.An apparatus for video coding, comprising:one or more processors; anda memory coupled to the one or more processors and configured to store instructions executable by the one or more processors,wherein the one or more processors, upon execution of the instructions, are configured to perform the method in any one of claims 1-14 or the method in any one of claims 15-28.30.Anon-transitory computer-readable storage medium storing a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the method of any one of claims 1-14 or the method of any one of claims 15-28.31.Acomputer readable storage medium storing a bitstream to be decoded by the method according to any one of claims 1-14.32.Acomputer readable storage medium storing a bitstream generated by the method according to any one of claims 15-28.33.A computer readable storage medium storing a bitstream formed by instructions which when executed by a computing device having one or more processors, cause the one or more processors to perform:determining whether one or more conditions are met for a current block in a current picture; andin response to determining that the one or more conditions are met, determining whether to apply at least one of a Bi-directional optical flow refinement and a Decoder side motion vector refinement for the current block of a method for video encoding, wherein the bitstream is to be decoded by the method according to any one of claims 1-14.34.Acomputer readable storage medium storing a bitstream formed by instructions which when executed by a computing device having one or more processors, cause the one or more processors to perform the method according to any one of claims 15-28.35.Amethod for storing a bitstream, comprising:generating a bitstream by performing the method of any one of claims 15-28; andstoring the bitstream on a storage medium.36.A computer program product comprising a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the method of any one of claims 1-14 or the method of any one of claims 15-28.
Citation Information
Patent Citations
Interaction between different decoder-side motion vector derivation modes
CN115842912A
Video decoding method, computing device and medium
CN116828185A
Inter-prediction method and device based on DMVR and bdof
US20220070466A1
Systems and methods for inter prediction compensation
US20220272323A1