Complexity Reduction and Bitwidth Control for Bidirectional Optical Flow

By applying bit shift operations on gradient arrays and correlation parameters, the method optimizes bi-directional optical flow in video coding, reducing complexity and enhancing efficiency, addressing the limitations of existing standards.

JP7744485B2Active Publication Date: 2025-09-25INTERDIGITAL VC HOLDINGS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024153291
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-01-07
Filing Date
2024-09-05
Publication Date
2025-09-25
Estimated Expiration
2039-09-17

AI Technical Summary

Technical Problem

Existing video coding standards like HEVC and VVC face challenges in achieving superior coding efficiency and complexity reduction, particularly with bi-directional optical flow (BIO) tools, which are computationally intensive and require high bit widths, limiting their performance and practical application.

Method used

The proposed method reduces bit width by performing right and downward bit shifts on gradient arrays and correlation parameters, generating reduced-bit-width intermediate parameters, and calculating motion refinements to predict video blocks using bi-directional optical flow, thereby optimizing computational complexity and efficiency.

Benefits of technology

This approach significantly reduces computational complexity and enhances coding efficiency for bi-directional optical flow, enabling improved video compression performance without sacrificing quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007744485000047
    Figure 0007744485000047
  • Figure 0007744485000048
    Figure 0007744485000048
  • Figure 0007744485000049
    Figure 0007744485000049
Patent Text Reader

Abstract

To provide systems and methods for reducing the complexity of using a bidirectional optical flow (BIO) in video coding.SOLUTION: In some embodiments, bit-width reduction steps are introduced in the BIO motion refinement process to reduce the maximum bit width used for BIO calculations. In some embodiments, simplified interpolation filters are used to generate predicted samples in an extended region around a current coding unit. In some embodiments, different interpolation filters are used for vertical and horizontal interpolations. In some embodiments, the BIO is disabled for coding units with small heights and / or for coding units that are predicted using a sub-block level inter prediction technique, such as advanced temporal motion vector prediction (ATMVP) or affine prediction.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a nonprovisional application claiming the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application No. 62 / 734,763 (filed September 21, 2018), U.S. Provisional Patent Application No. 62 / 738,655 (filed September 28, 2018), and U.S. Provisional Patent Application No. 62 / 789,331 (filed January 7, 2019), all of which are entitled "Complexity Reduction and Bit-Width Control for Bi-Directional Optical Flow," and are hereby incorporated by reference in their entireties. [Background technology]

[0002] Video coding systems are widely used to compress digital video signals to reduce the storage requirements and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block-based, wavelet-based, and object-based systems, block-based hybrid video coding systems are currently the most widely used and deployed. Examples of block-based video coding systems include MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and international video coding standards such as High Efficiency Video Coding (HEVC) developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG.

[0003] The first version of the HEVC standard was finalized in October 2013, providing approximately 50% bitrate savings or equivalent perceptual quality compared to the previous generation video coding standard, H.264 / MPEG AVC. While the HEVC standard offers significant coding improvements over its predecessor, there is evidence that superior coding efficiency to HEVC can be achieved using additional coding tools. Based on this, both VCEG and MPEG have begun work on exploring new coding techniques for future video coding standardization. In October 2015, the ITU-T VECG and ISO / IEC MPEG formed the Joint Video Exploration Team (JVET) to begin significant research into advanced technologies that could enable substantial improvements in coding efficiency. By integrating several additional coding tools onto the HEVC Test Model (HM), JVET developed reference software called the joint exploration model (JEM).

[0004] In October 2017, Non-Patent Document 1 was published by ITU-T and ISO / IEC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating a compression efficiency improvement of approximately 40% over HEVC. Based on these evaluation results, JVET launched a new project to develop a new generation video coding standard named Versatile Video Coding (VVC). In the same month, a reference software code base called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. Meanwhile, another reference software base called the Benchmark Set (BMS) was also created to facilitate the evaluation of new coding tools. The BMS code base includes a list of additional coding tools on the VTM that offer higher coding efficiency and moderate implementation complexity, and will be used as a benchmark for evaluating similar coding technologies during the VVC standardization process. Specifically, there are five JEM coding tools integrated into BMS-2.0, including 4x4 non-separable secondary transform (NSST), generalized bi-prediction (GBi), bi-directional optical flow (BIO), decoder-side motion vector refinement (DMVR), and current picture referencing (CPR). [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] joint call for proposals (CfP) on video compression with capability beyond HEVC, Oct. 2017 [Non-patent document 2] S. Jeong et al., “CE4 Ultimate motion vector expression (Test 4.5.4)”, JVET-L0054, Oct. 2018 [Non-patent document 3] S. Esenlik et al., “Simplified DMVR for inclusion in VVC”, JVET-L0670, Oct. 2018 [Non-patent document 4] M.-S. Chiang et al., “CE10.1.1: Multi-hypothesis prediction for improving AMVP mode, skip or merge mode, and intra mode”, JVET-L0100, Oct. 2018 [Non-patent document 5] M. Winken et al., “CE10-related: Multi-Hypothesis Inter Prediction with simplified AMVP process”, JVET-L0679, Oct. 2018 Summary of the Invention

[0006]

[0003] Embodiments described herein include methods for use in video encoding and decoding (collectively "coding"). In some embodiments, a video coding method is provided that includes, for at least one current block in a video coded using bidirectional optical flow, generating a first prediction signal array I from a first reference picture. (0) First horizontal gradient array based on (i,j)

[0007]

number

[0008] and calculating a second predicted signal array I from a second reference picture. (1)Second horizontal gradient array based on (i,j)

[0009]

number

[0010] and performing a right bit shift on the sum of (i) the first horizontal gradient array and (ii) the second horizontal gradient array, to generate a reduced-bit-width horizontal intermediate parameter array Ψ x (i,j) and calculating at least a horizontal motion refinement v based at least in part on the reduced bitwidth horizontal intermediate parameter array. x and calculating at least horizontal motion refinement v x and generating a prediction of the current block using bidirectional optical flow using

[0011] In some embodiments, the method includes generating a first predicted signal array I (0) (i,j) and the second predicted signal array I (1) calculating a signal difference parameter array θ(i,j) by a method including calculating the difference between (i) the signal difference parameter array θ(i,j) and (ii) the horizontal gradient intermediate parameter array Ψ x and calculating a signal horizontal gradient correlation parameter S3 by summing the components of the element-wise multiplication with (i,j), x The step of calculating the horizontal motion refinement v is performed by bit-shifting the signal horizontal gradient correlation parameter S3. x The method includes the step of obtaining:

[0012] In such an embodiment, the step of calculating the signal difference parameter array θ(i,j) includes calculating the first predicted signal array I (0) (i,j) and the second predicted signal array I (1)Before calculating the difference with (i,j), the first predicted signal array I (0) (i,j) and the second predicted signal array I (1) (i,j) and performing a right bit shift on each of them.

[0013] In some embodiments, the method includes generating a first prediction signal array I from a first reference picture. (0) First vertical gradient array based on (i,j)

[0014]

number

[0015] and calculating a second predicted signal array I from a second reference picture. (1) A second vertical gradient array based on (i,j)

[0016]

number

[0017] and performing a right bit shift on the sum of (i) the first vertical gradient array and (ii) the second vertical gradient array to generate a reduced bit-width vertical intermediate parameter array Ψ y (i,j) and the reduced bit-width horizontal intermediate parameter array Ψ x (i,j) and the reduced bit-width vertical intermediate parameter array Ψ y Based at least in part on (i,j), a vertical motion refinement v y and calculating the horizontal motion refinement v x and vertical motion refinement v y is generated using

[0018] Some such embodiments include: (i) a horizontal intermediate parameter array Ψ x (i,j) and (ii) the vertical intermediate parameter array Ψ y(i,j) and further comprising the step of calculating a cross-gradient correlation parameter S2 by a method comprising the step of summing components of an element-wise multiplication with (i,j), y The steps to calculate the horizontal motion refinement v x and (ii) a cross-gradient correlation parameter S2.

[0019] In some such embodiments, (i) horizontal motion refinement v x and (ii) the step of determining the product of the cross-gradient correlation parameter S2 is performed by dividing the cross-gradient correlation parameter S2 by the most significant bit MSB parameter portion S 2,m and the least significant bit (LSB) parameter part S 2,s and (i) horizontal motion refinement v x and (ii) the MSB parameter part S 2,m (i) determining the MSB product of v and v; and (ii) a horizontal motion refinement v x and (ii) the LSB parameter part S 2,S and performing a left bit shift of the MSB product to generate a bit-shifted MSB product; and adding the LSB product and the bit-shifted MSB product.

[0020] In some embodiments, generating a prediction for the current block using bidirectional optical flow comprises, for each sample in the current block:

[0021]

number

[0022] (v) Horizontal movement refinement v x , and (vi) vertical motion refinement v y and for each sample in the current block, calculating a bidirectional optical flow sample offset b based on at least a first predicted signal array I (0) (i,j), the second predicted signal array I (1)(i,j), and calculating the sum of the bidirectional optical flow sample offsets b.

[0023] In some embodiments, a gradient array

[0024]

number

[0025] The step of calculating each of the predicted signal array I (0) (i,j), I (1) Pad samples outside of (i,j) with their nearest boundary samples inside the predicted signal array.

[0026] In some embodiments, the step of calculating at least some values ​​of the signal difference parameter array θ(i,j) comprises calculating the predicted signal array I (0) (i,j), I (1) In some embodiments, the horizontal intermediate parameter array Ψ includes padding the samples outside (i,j) with their respective nearest boundary samples inside the predicted signal array. x The step of calculating at least some of the values ​​of (i,j) includes calculating a horizontal gradient array

[0027]

number

[0028] padding gradient values ​​outside of the horizontal gradient array with their nearest boundary samples inside the horizontal gradient array.

[0029] In some embodiments, the vertical intermediate parameter array Ψ y The step of calculating at least some of the values ​​of (i,j) includes calculating a vertical gradient array

[0030]

number

[0031] padding gradient values ​​outside of the vertical gradient array with their nearest boundary samples inside the vertical gradient array.

[0032] In some embodiments, the signal horizontal gradient correlation parameter S3 and the cross-gradient correlation parameter S2 are calculated for each sub-block in the current block.

[0033] The embodiments described herein may be performed by an encoder or by a decoder to generate a prediction of a video block.

[0034] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first gradient component (e.g., ∂I (0) / ∂x or ∂I (0) A second gradient component (e.g., ∂I / ∂y) is calculated based on a second prediction signal from a second reference picture. (1) / ∂x or ∂I (1) The first gradient component and the second gradient component are summed, and a downward bit shift is performed on the resulting sum to obtain a reduced bit-width correlation parameter (e.g., Ψ x or Ψ y ) is generated. A BIO motion refinement is calculated based at least in part on the reduced bitwidth correlation parameters. The calculated motion refinement is used to predict blocks using bidirectional optical flow.

[0035] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a second prediction signal (e.g., I (1) ) to generate a first predicted signal (e.g., I (0)) and performing a downward bit shift of the resulting difference to generate a reduced bit-width correlation parameter (e.g., θ). A BIO motion refinement is calculated based at least in part on the reduced bit-width correlation parameter. The calculated motion refinement is used to predict blocks using bidirectional optical flow.

[0036] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first prediction (e.g., I (0) ) signal, a first predicted signal of reduced bit width is generated by performing a downward bit shift on the signal. A second predicted signal from a second reference picture (e.g., I (1) ) to generate a reduced bit-width second prediction signal. A reduced bit-width correlation parameter (e.g., θ) is generated by subtracting the reduced bit-width first prediction signal from the reduced bit-width second prediction signal. A BIO motion refinement is calculated based at least in part on the reduced bit-width correlation parameter, and the calculated motion refinement is used to predict the block using bidirectional optical flow.

[0037] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first gradient component at a reduced bit width is calculated based on a first prediction signal at a reduced bit width from a first reference picture. A second gradient component at a reduced bit width is calculated based on a second prediction signal at a reduced bit width from a second reference picture. The first reduced bit width gradient component and the second reduced bit width gradient component are summed to generate a reduced bit width correlation parameter. A motion refinement is calculated based at least in part on the reduced bit width correlation parameter, and the block is predicted using bidirectional optical flow using the calculated motion refinement.

[0038] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first motion-compensated prediction signal and a second motion-compensated prediction signal are generated for samples in the current block, the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block being generated using a first interpolation filter having a first number of taps. The first motion-compensated prediction signal and the second motion-compensated prediction signal are also generated for samples in an extended region around the current block, and the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples outside the current block are generated using a second interpolation filter having a second number of taps that is less than the first number of taps. Motion refinement is calculated based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal, and the block is predicted using bidirectional optical flow using the calculated motion refinement.

[0039] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first motion-compensated prediction signal and a second motion-compensated prediction signal are generated, the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a horizontal interpolation filter having a first number of taps and a vertical interpolation filter having a second number of taps that is less than the first number of taps, a motion refinement is calculated based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal, and the block is predicted using the calculated motion refinement using bidirectional optical flow.

[0040] In some embodiments, a first motion-compensated prediction signal and a second motion-compensated prediction signal are generated for at least one current block in a video coded using bidirectional optical flow. The first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a horizontal interpolation filter having a first number of taps and a vertical interpolation filter having a second number of taps. The horizontal and vertical filters are applied in a predetermined order, with filters applied earlier in the order having a greater number of taps than filters applied later in the order. Motion refinement is calculated based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal, and the calculated motion refinement is used to predict the block using bidirectional optical flow.

[0041] In some embodiments, a method for coding a video including a plurality of coding units is provided. For a plurality of coding units in a video coded using bi-prediction, bi-directional optical flow is disabled for at least coding units having heights equal to or less than a threshold height (e.g., BIO may be disabled for coding units of height 4). For bi-predicted coding units for which bi-directional optical flow is disabled, bi-prediction without bi-directional optical flow is performed. For bi-predicted coding units for which bi-directional optical flow is not disabled (e.g., for at least one of the bi-predicted coding units for which bi-directional optical flow is not disabled), bi-prediction using bi-directional optical flow is performed.

[0042] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first motion-compensated prediction signal and a second motion-compensated prediction signal are generated for samples in the current block. The first and second values ​​relate to samples in an extended region around the current block, where the extended region does not include samples more than one row or column away from the current block. Motion refinement is calculated based at least in part on the first and second motion-compensated prediction signals and the first and second values ​​for samples in the extended region. The block is predicted using bidirectional optical flow using the calculated motion refinement.

[0043] In some embodiments, a method for coding a video including a plurality of coding units is provided. For a plurality of coding units in a video coded using bi-prediction, bi-directional optical flow is disabled for at least coding units predicted using sub-block-level inter prediction techniques (e.g., advanced temporal motion vector prediction and affine prediction, etc.). For bi-predicted coding units for which bi-directional optical flow is disabled, bi-prediction without bi-directional optical flow is performed. For bi-predicted coding units for which bi-directional optical flow is not disabled (e.g., for at least one of the bi-predicted coding units for which bi-directional optical flow is not disabled), bi-prediction using bi-directional optical flow is performed.

[0044] In some embodiments, for at least one current block in a video coded using bidirectional optical flow, a first motion-compensated prediction signal and a second motion-compensated prediction signal are generated for samples in the current block. The first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a horizontal interpolation filter having a first number of taps and a vertical interpolation filter having a first number of taps. The first motion-compensated prediction signal and the second motion-compensated prediction signal are also generated for samples in an extended region around the current block, and the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples outside the current block are generated using a horizontal interpolation filter having the first number of taps and a vertical interpolation filter having a second number of taps that is less than the first number of taps. Motion refinement is calculated based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal. Blocks are predicted using bidirectional optical flow using the calculated motion refinement.

[0045] In additional embodiments, encoder and decoder systems are provided for performing the methods described herein. The encoder or decoder systems may include a processor and a non-transitory computer-readable medium that stores instructions for performing the methods described herein. Additional embodiments may include a non-transitory computer-readable storage medium that stores video encoded using the methods described herein. [Brief explanation of the drawings]

[0046] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communications system shown in FIG. 1A, according to an embodiment. [Figure 2A]FIG. 1 is a functional block diagram of a block-based video encoder, such as the encoder used in VVC. [Figure 2B] FIG. 1 is a functional block diagram of a block-based video decoder, such as the decoder used in VVC. [Figure 3A] FIG. 10 is a diagram showing four-partitioning, which is block partitioning in a multi-type tree structure. [Figure 3B] FIG. 10 is a diagram illustrating vertical bipartitioning, which is block partitioning in a multi-type tree structure. [Figure 3C] FIG. 10 is a diagram illustrating horizontal bipartitioning, which is block partitioning in a multi-type tree structure. [Figure 3D] FIG. 10 is a diagram showing vertical tri-partitioning, which is block partitioning in a multi-type tree structure. [Figure 3E] FIG. 10 is a diagram showing horizontal three-partitioning, which is block partitioning in a multi-type tree structure. [Figure 4] Schematic diagram of prediction using bidirectional optical flow (BIO). [Figure 5] FIG. 1 illustrates a method for generating augmented samples for BIO using a simplified filter, according to some embodiments. [Figure 6] FIG. 1 illustrates a method for generating augmented samples for BIO using a simplified filter, according to some embodiments. [Figure 7] FIG. 10 illustrates sample and gradient padding to reduce the number of interpolated samples in an extended region of one BIO coding unit (CU) according to some embodiments. [Figure 8] FIG. 1 illustrates an example of a coded bitstream structure. [Figure 9] FIG. 1 illustrates an exemplary communication system. [Figure 10] FIG. 10 illustrates the use of integer samples as extended samples for BIO derivation. DETAILED DESCRIPTION OF THE INVENTION

[0047] Exemplary Network for Implementation of the Embodiments 1A illustrates an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communication system 100 enables the multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0048] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, the WTRUs 102a, 102b, 102c, 102d may all be referred to as “stations” and / or “STAs,” may be configured to transmit and / or receive wireless signals, and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, IoT devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain contexts), consumer electronics, devices operating on commercial and / or industrial wireless networks, etc. The WTRUs 102a, 102b, 102c, and 102d may all be referred to interchangeably as UEs.

[0049] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or the network 112. For example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNode B, a home Node B, a home eNode B, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0050] The base station 114a may be part of the RAN 104 / 113, which may include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, sometimes referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for wireless services for a particular geographic area, which may be relatively fixed or may change over time. A cell may be divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one for each sector of the cell. In an embodiment, the base station 114a may utilize MIMO technology and utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.

[0051] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0052] More specifically, as noted above, the communication system 100 may be a multiple-access system and may utilize one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a and the WTRUs 102a, 102b, and 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​UL Packet Access (HSUPA).

[0053] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0054] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.

[0055] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0056] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0057] 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a workplace, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by drones), a roadway, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may establish a picocell or femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not be required to access the Internet 110 via the CN 106 / 115.

[0058] The RAN 104 / 113 can communicate with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or VoIP services to one or more of the WTRUs 102a, 102b, 102c, and 102d. The data may have various quality of service (QoS) requirements, such as different throughput, latency, error tolerance, reliability, data throughput, and mobility requirements. The CN 106 / 115 can provide call control, billing services, mobile location services, prepaid calling, Internet connectivity, video distribution, and / or perform high-level security functions, such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 can communicate directly or indirectly with other RANs that utilize the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0059] The CNs 106 / 115 may also act as gateways for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as TCP, UDP, and / or IP in the TCP / IP Internet protocol suite. The networks 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs that may utilize the same RAT as the RAN 104 / 113 or a different RAT.

[0060] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that can utilize cellular-based wireless technology and a base station 114b that can utilize IEEE 802 wireless technology.

[0061] 1B is a system diagram of an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a GPS chipset 136, and / or other peripherals 138, etc. It will be understood that the WTRU 102 may include any sub-combination of the above elements while remaining consistent with an embodiment.

[0062] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, an ASIC, an FPGA circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0063] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0064] 1B is shown as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0065] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.

[0066] The processor 118 of the WTRU 102 is coupled to and may receive user input data from the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. The processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include RAM, ROM, a hard disk, or any other type of memory storage device. The removable memory 132 may include a SIM card, a memory stick, an SD memory card, or the like. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).

[0067] The processor 118 may be configured to receive power from the power source 134 and to distribute and / or control the power to other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0068] The processor 118 may be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or in place of information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may obtain location information by way of any suitable location-determination method while remaining consistent with an embodiment.

[0069] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a USB port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0070] The WTRU 102 may include a full-duplex radio, in which transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit for reducing or substantially eliminating self-interference through signal processing by hardware (e.g., a choke) or a processor (e.g., a separate processor (not shown) or processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)).

[0071] Although the WTRU is described in Figures 1A-1B as a wireless terminal, it is contemplated that in certain representative embodiments, such a terminal may use a wired communication interface (e.g., temporarily or permanently) with a communication network.

[0072] In an exemplary embodiment, the other network 112 may be a WLAN.

[0073] 1A-1B and corresponding description, one or more or all of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0074] The emulation device may be designed to implement one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation device may be directly coupled to another device for testing and / or may perform testing using wireless communication.

[0075] The one or more emulation devices may perform one or more functions, inclusive, without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in a test scenario in a test laboratory and / or in an undeployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may, for example, include one or more antennas) may be used to transmit and / or receive data by the emulation devices.

[0076] Detailed Description Block-Based Video Coding Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 2A shows a block diagram of a block-based hybrid video coding system. The input video signal 103 is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128 x 128 pixels. However, unlike HEVC, which partitions blocks based solely on a quadtree, VTM-1.0 partitions coding tree units (CTUs) into CUs based on a quadtree, binary tree, or ternary tree to adapt to various local characteristics. Furthermore, the concept of multiple partition unit types in HEVC has been removed, and the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transformation without further partitioning. In a multi-type tree structure, one CTU is first partitioned using a quadtree structure. Each quadtree leaf node can be further partitioned using a binary tree and a ternary tree structure. As shown in FIGS. 3A to 3E, there are five division types: four-division, two-horizontal division, two-vertical division, three-horizontal division, and three-vertical division.

[0077] As shown in FIG. 2A, spatial prediction (161) and / or temporal prediction (163) may be performed. Spatial prediction (or "intra-prediction") predicts the current video block using pixels from samples (called reference samples) of already coded neighboring blocks within the same video picture / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") predicts the current video block using pixels reconstructed from already coded video pictures. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. If multiple reference pictures are supported, a reference picture index is additionally transmitted and is used to identify which reference picture in the reference picture store (165) the temporal prediction signal comes from. After spatial prediction and / or temporal prediction, a mode decision block (181) in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block (117), and the prediction residual is decorrelated using a transform (105) and quantization (107). The quantized residual coefficients are inverse quantized (111) and inverse transformed (113) to form a reconstructed residual, which is then fed back to the prediction block (127) to form a reconstructed signal for the CU. In-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF), may be applied to the reconstructed CU (167) before it is placed in a reference picture store (165) and used to code future video blocks. To form the output video bitstream 121, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit (109) for further compression and packing to form the bitstream.

[0078] FIG. 2B shows a functional block diagram of a block-based video decoder. A video bitstream 202 is first unpacked and entropy decoded in an entropy decoding unit 208. Coding mode and prediction information is sent to either a spatial prediction unit 260 (if intra-coded) or a temporal prediction unit 262 (if inter-coded) to form a prediction block. Residual transform coefficients are sent to an inverse quantization unit 210 and an inverse transform unit 212 to reconstruct a residual block. The prediction block and residual block are added together at 226. The reconstructed block may further pass through in-loop filtering before being stored in a reference picture store 264. The reconstructed video in the reference picture store is sent to drive a display device and is also used to predict future video blocks.

[0079] As mentioned above, BMS-2.0 adheres to the same VTM-2.0 encoding / decoding workflow as shown in Figures 2A and 2B. However, some coding modules, especially those associated with temporal prediction, have been further enhanced to improve coding efficiency. This disclosure is directed to reducing the computational complexity and solving the large bit-width problem associated with existing BIO tools in BMS-2.0. Below, the main design aspects of the BIO tools are introduced, followed by a more detailed analysis of the computational complexity and bit-width of existing BIO implementations.

[0080] Bi-predictive prediction based on optical flow model Traditional bi-prediction in video coding is a simple combination of two temporally predicted blocks obtained from already reconstructed reference pictures. However, due to the constraints of block-based motion compensation (MC), there is a small residual motion observable between the samples of the two predicted blocks, which can reduce the efficiency of the motion-compensated prediction. To solve this problem, bi-directional optical flow (BIO) is applied in BMS-2.0 to reduce the impact of such motion on all samples within a block. Specifically, BIO is a sample-by-sample motion refinement performed in addition to block-based motion-compensated prediction when bi-prediction is used. In the current BIO design, the derivation of the refined motion vector for each sample within a block is based on the classical optical flow model. (k) Let (x, y) be the sample value at coordinates (x, y) of the predicted block obtained from reference picture list k (k=0, 1). (k) (x,y) / ∂x and ∂I (k) (x,y) / ∂y are the horizontal and vertical gradients of the sample. Given an optical flow model, the motion refinement (v x ,v y )teeth,

[0081]

number

[0082] can be derived by

[0083] In Figure 4, (MV x0 ,MV y0 ) and (MV x1 ,MV y1 ) is Block I (0) and I (1) Furthermore, we denote the block-level motion vectors used to generate the motion refinement (v x ,v y )teeth,

[0084]

number

[0085] As shown in Figure 4, the difference Δ between the values ​​of the samples after motion refinement compensation is calculated by minimizing the difference Δ (A and B in Figure 4).

[0086] Furthermore, to ensure the regularity of the derived motion refinement, the motion refinement is assumed to be consistent with respect to samples within one small unit (i.e., a 4x4 block). In BMS-2.0, (v x ,v y ) value is

[0087]

number

[0088] It is derived by minimizing Δ in a 6×6 window Ω around each 4×4 block, as in

[0089] To solve the optimization problem specified in equation (3), BIO uses an incremental method that optimizes refinement by first moving horizontally and then vertically.

[0090]

number

[0091] where:

[0092]

number

[0093] is a floor function that outputs the maximum value below the input, and th BIO is the motion refinement threshold to prevent error propagation due to coding noise and irregular local motion, which is 2 18-BDis equal to. The operator (?:) is a ternary conditional operator, and an expression of the form (a? b : c) is evaluated to b if the value of a is true, and otherwise it is evaluated to c. The function clip3(a, b, c) returns a if c < a, returns c if a ≤ c ≤ b, and returns b if b < c. The values of S1, S2, S3, S5, and S6 are further

[0094] [Number]

[0095] calculated as, where

[0096] [Number]

[0097] is.

[0098] In BMS-2.0, the BIO gradients in both the horizontal and vertical directions at (6) are directly obtained by calculating the difference between two adjacent samples at one sample position of each L0 / L1 prediction block (horizontally or vertically according to the direction of the derived gradient), for example, obtained as follows.

[0099] [Number]

[0100] In Equation (5), L is the bit-depth increase of the internal BIO process to maintain data accuracy, and is set to 5 in BMS-2.0. Further, to avoid division by a smaller value, the adjustment parameters r and m in Equation (4) are r = 500·4 BD-8 m = 700·4 BD-8 (8) where BD is the bit depth of the input video. Based on the motion refinement derived by equation (4), the final bi-predictive signal of the current CU can be calculated by interpolating L0 / L1 prediction samples along the motion trajectory based on optical flow equation (1), as specified below:

[0101]

number

[0102] where b is the bidirectional optical flow sample offset, shift is the right shift applied to combine the L0 and L1 predicted signals for bi-prediction, and may be set equal to 15-BD; offset is the bit depth offset, which can be set to 1 <<(14-BD)+2 · (1 <<13), and rnd( ·) is the rounding function that rounds the input value to the nearest integer value.

[0103] BIO Bit Width Analysis Similar to the aforementioned standard HEVC, for bi-predicted CUs in VVC, when MV refers to a fractional sample position, the L0 / L1 predicted signal, i.e., I (0) (x,y) and I (1) (x,y) are generated with intermediate precision (i.e., 16 bits) to maintain the precision of the subsequent averaging operation. Furthermore, if either of the two MVs is an integer, the precision of the corresponding predicted sample (taken directly from the reference picture) is increased to intermediate precision before averaging is applied. Given a bi-predictive signal at intermediate bit depth and assuming the input video is 10 bits, Table 1 summarizes the bit-widths of the intermediate parameters required at each stage of the BIO process as shown in the section "Bi-predictive Prediction Based on Optical Flow Model".

[0104] [Table 1]

[0105] As can be seen from Table 1, the maximum bit width of the entire BIO process is determined by the vertical motion refinement v in Eq. (4). y where S6 (42 bits) is V x (9 bits) and S2 (33 bits). Therefore, the maximum bit width of the existing BIO design is equal to 42+1=43 bits. Furthermore, the multiplication (i.e., v x S2) takes S2 as input, so the 33-bit multiplier is v y Therefore, the current direct implementation of BIO in BMS-2.0 requires a 33-bit multiplier and has a maximum bit width of 43 bits for the intermediate parameters.

[0106] Computational complexity analysis of BIO In this section, a computational complexity analysis is performed for the existing BIO design. Specifically, the number of operations (e.g., multiplications and additions) used to generate the final motion-compensated prediction with BIO applied is calculated according to the current BIO implementation of BMS-2.0. Furthermore, to facilitate the following discussion, we assume that the size of the current CU predicted by BIO is equal to W × H, where W is the width and H is the height of the CU.

[0107] Generating L0 and L1 prediction samples As shown in Equation (3), the local motion refinement (v x ,v y To derive v , both the required sample values ​​and gradient values ​​are calculated for all samples within a 6x6 surrounding window around the sample. Thus, the local motion refinement (v x ,v yTo derive (W+2)×(H+2), gradients of (W+2)×(H+2) samples are used by BIO. Furthermore, as shown in equation (7), both horizontal and vertical gradients are obtained by directly calculating the difference between two adjacent samples. Therefore, to calculate the (W+2)×(H+2) gradient value, the total number of prediction samples in both L0 and L1 prediction directions is equal to (W+4)×(H+4). Because current motion compensation is based on a 2D separable finite impulse response (FIR) 8-tap filter, the number of both multiplications and additions used to generate the L0 and L1 prediction samples is equal to ((W+4)×(H+4+7)×8+(W+4)×(H+4)×8)×2.

[0108] Gradient calculation As shown in equation (7), only one addition is required per sample since the gradients are calculated directly from the two adjacent predicted samples. Considering that both horizontal and vertical gradients are derived over an extended region of (W+2) x (H+2) for both L0 and L1, the total number of additions required to derive the gradients is equal to ((W+2) x (H+2)) x 2 x 2.

[0109] Correlation parameter calculation As shown in equations (5) and (6), there are five correlation parameters (i.e., S1, S2, S3, S5, and S6) calculated by BIO for every sample in the extended region (W+2)×(H). Furthermore, there are five multiplications and three additions used to calculate the five parameters at each sample position. Therefore, the total number of multiplications and additions to calculate the correlation parameters is equal to ((W+2)×(H+2))×5 and ((W+2)×(H+2))×3, respectively.

[0110] Sum As mentioned above, BIO motion refinement (v x ,v y) is derived separately for each 4x4 block in the current CU. To derive the motion refinement for each 4x4 block, the sum of five correlation parameters within a 6x6 surrounding region is calculated. Thus, at this stage, the summation of the five correlation parameters uses a total of (W / 4) x (H / 4) x 6 x 6 x 5 additions.

[0111] Motion refinement derivation As shown in Equation (4), the local motion refinement (v x ,v y ), there are two additions to add the tuning parameter r to S1 and S3. Furthermore, v y Therefore, to derive the motion refinement for all 4x4 blocks in a CU, the number of multiplications and additions used is equal to (W / 4) x (H / 4) and (W / 4) x (H / 4) x 3, respectively.

[0112] Bi-predictive signal generation Given the derived motion refinement, two more multiplications and six additions are used to derive the final predicted sample value at each sample location, as shown in equation (9). At this stage, a total of W x H x 2 multiplications and W x H x 6 additions are performed.

[0113] Issues Addressed in Some Embodiments As mentioned above, BIO can improve the efficiency of bi-predictive prediction by improving both the granularity and accuracy of the motion vectors used in the motion compensation stage. Although BIO can effectively improve coding performance, the complexity of actual hardware implementation increases significantly. In this disclosure, the following complexity problems exist in the current BIO design of BMS-2.0:

[0114] Large intermediate bit width and large multiplier for BIO Similar to the HEVC standard, when MVs point to fractional sample positions in a reference picture, 2D separable FIR filters are applied in the motion compensation stage to interpolate predicted samples of a prediction block. Specifically, first, one interpolation filter is applied horizontally to derive intermediate samples according to the horizontal fractional component of the MV, and then another interpolation filter is applied vertically on the above horizontal fractional samples according to the vertical fractional component of the MV. Assuming the input is 10-bit video (i.e., BD=10), Table 2 shows the bitwidth measurement of the motion compensation prediction process in VTM / BMS-2.0, assuming that the horizontal and vertical MVs point to half-sample positions corresponding to the worst-case bitwidth of the interpolated samples from the motion compensation process. Specifically, in the first stage, the values ​​of the input reference samples associated with the positive and negative filter coefficients are respectively adjusted to the maximum input value (i.e., 2 BD The worst-case bit-width of the intermediate data after the first interpolation process (horizontal interpolation) is calculated by setting the input value of the second interpolation to the worst possible value output from the first interpolation (i.e., -1) and the minimum input value (i.e., 0). The worst-case bit-width of the second interpolation process (vertical interpolation) is then obtained by setting the input value of the second interpolation to the worst possible value output from the first interpolation.

[0115] [Table 2]

[0116] As can be seen from Table 2, the maximum bit width of the motion compensation interpolation exists in the vertical interpolation process, where the input data is 15 bits and the filter coefficients are 7-bit signed values, therefore the bit width of the output data from the vertical interpolation is 22 bits. Furthermore, when the input data to the vertical interpolation process is 15 bits, a 15-bit multiplier is sufficient for generating intermediate fractional sample values ​​in the motion compensation stage.

[0117] However, as analyzed above, the existing BIO design requires a 33-bit multiplier and has a 43-bit intermediate parameter to maintain the accuracy of the intermediate data. Compared with Table 2, both figures are much larger than those of conventional motion compensation interpolation. In practice, such a large bit-width increase (especially the required multiplier bit-width increase) would be very expensive for both hardware and software, and would increase the implementation cost of BIO.

[0118] High computational complexity of BIO Based on the above complexity analysis, Tables 3 and 4 show the number of multiplications and additions that need to be performed per sample for different CU sizes according to the current BIO and compare them with the complexity statistics of a regular 4x4 bi-predicted CU, which corresponds to the worst-case computational complexity under VTM / BMS-2.0. For a 4x4 bi-predicted CU, given the length of the interpolation filter (e.g., 8), the total number of multiplications and additions is equal to (4x(4+7)x8+4x4x8)x2=960 (i.e., 60 per sample) and (4x(4+7)x8+4x4x8)x2+4x4x2=992 (i.e., 62 per sample).

[0119] [Table 3]

[0120] [Table 4]

[0121] As shown in Tables 3 and 4, the computational complexity shows a significant increase compared to the worst-case complexity of regular bi-prediction by enabling the existing BIO in BMS-2.0. The peak complexity increase occurs for 4x4 bi-predictive CUs, where the number of multiplications and additions with BIO enabled is 329% and 350% of the worst-case bi-prediction case.

[0122] Overview of Exemplary Embodiments To solve at least some of the above-mentioned problems, this section proposes methods to reduce the complexity of BIO-based motion compensation prediction while maintaining coding gain. First, to reduce implementation costs, this disclosure proposes bit-width control methods to reduce the internal bit-width used in hardware BIO implementation. In some proposed methods, BIO-enabled motion compensation prediction may be implemented using 15-bit multipliers and 32-bit intermediate values.

[0123] Second, a method is proposed to reduce the computational complexity of BIO by using simplified filters and reducing the number of extended prediction samples used in BIO motion refinement.

[0124] Furthermore, in some embodiments, it is proposed to disable CU-sized BIO operations, which result in a significant increase in computational complexity compared to normal bi-prediction. Based on the combination of these complexity reductions, the worst-case computational complexity (e.g., the number of multiplications and additions) of motion compensation prediction when BIO is enabled can be reduced to approximately the same level as the worst-case complexity of normal bi-prediction.

[0125] Exemplary BIO Bitwidth Control Method As pointed out above, the current implementation of BIO in BMS-2.0 uses a 33-bit multiplier and a 43-bit bitwidth of the intermediate parameters, which is much larger than that of the HEVC motion compensation interpolation implementation. Therefore, it is very costly to implement BIO in hardware and software. In this section, a bitwidth control method is proposed to reduce the bitwidth required for BIO. In an exemplary method, as shown below, the horizontal intermediate parameter array Ψ in Equation (6) is reduced to reduce the overall bitwidth of the intermediate parameters. x (i,j), the vertical intermediate parameter array Ψ y (i,j), and one or more of the signal difference parameter arrays θ(i,j), respectively, are first a bit and n b Bit-shifted downward:

[0126]

number

[0127] Furthermore, to further reduce the bit width, the original L-bit internal bit depth increase can be removed. With such a modification, the horizontal gradient correlation parameter (S1), cross-gradient correlation parameter (S2), signal-horizontal gradient correlation parameter (S3), vertical gradient correlation parameter (S5), and signal-vertical gradient correlation parameter (S6) can be implemented as follows by calculating with the equations in (5):

[0128]

number

[0129] Different numbers of right shifts (i.e., n a and n b ) is Ψ x (i,j), Ψ y (i,j), and θ(i,j), the values ​​of S1, S2, S3, S5, and S6 are downscaled by different factors, thereby reducing the derived motion refinement (v x ,v y ) may vary in width. Therefore, an additional left shift may be introduced in equation (4) to provide an accurate range of widths for the derived motion refinement. Specifically, in the exemplary method, the horizontal motion refinement v x and vertical motion refinement v y can be derived as follows:

[0130]

number

[0131] Note that, unlike equation (4), in this embodiment, the adjustment parameters r and m are not applied. x ,v y) in the original BIO design. BIO =2 18-BD A small motion refinement threshold th' compared to BIO =2 13-BD In equation (12), the product v x S2 is an input with a bit width greater than 16 bits, so v y To avoid this, the value of the cross-gradient correlation parameter S2 may be divided into two parts, the first part S 2,S is the lowest n S2 a second part S 2,m It is proposed that S2 includes other bits. Based on this, the value S2 can be expressed as follows:

[0132]

number

[0133] Then, substituting equation (13) into equation (12), the vertical motion refinement v y The calculation of is as follows:

[0134]

number

[0135]

number

[0136] Assuming that the input video is 10 bits, Table 5 summarizes the bit widths of the intermediate parameters when the example bit width control method is applied to BIO. As shown in Table 5, using the proposed example bit width control method, the internal bit width of the entire BIO process does not exceed 32 bits. The multiplication with the worst possible input is performed by using v in Equation (14). x S 2,m where the input S 2,m is 15 bits, and the input v x is 4 bits. Therefore, one 15-bit multiplier is sufficient when the exemplary method is applied for BIO.

[0137] [Table 5-1]

[0138] [Table 5-2]

[0139] Finally, in equation (10), the L0 and L1 prediction samples I (0) (i,j) and I (1) The BIO parameter θ(i,j) is calculated by applying a right shift to the difference between (i,j). (0) (i,j) and I (1)Both (i,j) values ​​are 16-bit, and their difference can be a single 17-bit value. Such a design may not be well suited for SIMD-based software implementation. For example, a 128-bit SIMD register can only process four samples in parallel. Therefore, another example suggests first applying a right shift before calculating the difference when calculating the signal difference parameter array θ(i,j), as follows: θ(i,j)=(I (1) (i,j)≫n b )-(I (0) (i,j)≫n b ) (16)

[0140] In such an embodiment, since the input value of each operation is 16 bits or less, more samples can be processed in parallel. For example, by using Equation (16), eight samples can be processed simultaneously by one 128-bit SIMD register. In some embodiments, a similar method is applied to the gradient calculation of Equation (7), where a 4-bit right shift is applied before calculating the difference between the L0 predicted sample and the L1 predicted sample to maximize the payload of each SIMD calculation. Specifically, by doing so, the gradient value can be calculated as follows:

[0141]

number

[0142] Exemplary Methods for Reducing BIO Computational Complexity As shown above, the existing BIO design in BMS-2.0 results in a large complexity increase (e.g., the number of multiplications and additions) compared to the worst-case computational complexity of normal bi-prediction. In the following, a method is proposed to reduce the worst-case computational complexity of BIO.

[0143] BIO complexity reduction by generating enhanced samples using simplified filters As mentioned above, assuming that the current CU is W×H, the gradients of samples in the extended region (W+2)×(H+2) are calculated to derive motion refinement for all 4×4 blocks in the CU. In existing BIO designs, the same interpolation filter (8-tap filter) used for motion compensation is used to generate these extended samples. As shown in Tables 3 and 4, the complexity due to the interpolation of samples in the extended region is the complexity bottleneck of BIO. Therefore, to reduce BIO complexity, instead of using an 8-tap interpolation filter, it is proposed to use a simplified interpolation filter with a shorter tap length for generating samples in the extended surrounding region of the BIO CU. On the other hand, generating extended samples requires accessing more reference samples from the reference picture, which may increase the memory bandwidth of BIO. To avoid memory bandwidth increase, reference sample padding, which is used in the current BIO of BMS-2.0, can be applied, in which reference samples outside the normal reference region of the normal motion compensation of a CU (i.e., (W+7) × (H+7)) are padded by the nearest boundary samples of the normal reference region. To calculate the size of the padded reference samples, assuming that the length of the simplified filter used to generate the extended samples is N, the number of padded reference samples along each of the top, bottom, left, and right boundaries of the normal reference region, M, is equal to:

[0144]

number

[0145] As shown in Equation (18), the reference samples on the boundary of the normal reference region are padded by two rows or two columns in each direction by using an 8-tap filter to generate extended predicted samples. FIG. 5 illustrates the use of simplified filter and reference sample padding to generate samples in the extended region by BIO. As shown in FIG. 5, the simplified filter is used only to generate predicted samples within the extended region. For positions within the region of the current CU, these predicted samples are generated by applying default 8-tap interpolation to maintain BIO coding efficiency. In particular, one embodiment of the present disclosure proposes using a bilinear interpolation filter (i.e., a 2-tap filter) to generate extended samples, thereby further reducing the number of operations used for BIO. FIG. 6 illustrates the case where a bilinear filter is used to interpolate extended samples for BIO. The predicted samples within the CU are interpolated by the default 8-tap filter. As shown in FIG. 6, due to the shortened filter length, the bilinear filter does not need to access additional reference samples outside the normal reference region to interpolate the required samples in the extended region. Therefore, in this case, padding of reference samples can be avoided, which can further reduce the complexity of BIO operations.

[0146] Furthermore, based on a comparison of the complexity statistics for different CU sizes in Tables 3 and 4, it can be seen that the complexity increase is greater for CU sizes with lower heights. For example, although an 8x4 CU and a 4x8 CU contain the same number of samples, they exhibit different complexity increase percentages. Specifically, for an 8x4 CU, the number of multiplications and additions increases by 149% and 172%, respectively, after enabling BIO, while for a 4x8 CU, the corresponding complexity increases are 126% and 149%, respectively. Such complexity differences are caused by the fact that in the current motion compensation design, a horizontal interpolation filter is applied first, followed by a vertical interpolation filter. When the applied MV points to a fractional position in the vertical direction, more intermediate samples are generated from the horizontal interpolation and used as inputs for the vertical interpolation. Therefore, the complexity effect of generating more reference samples in the expanded region is relatively more significant for CU sizes with lower heights.

[0147] In some embodiments, it is proposed to disable certain CU sizes with small heights to reduce worst-case BIO complexity. In addition to the above method of simply disabling certain CU sizes, another way to solve the increased number of operations in the vertical interpolation process is to simplify the interpolation filter used for vertical interpolation. In the current design, the same 8-tap interpolation filter is applied to both the horizontal and vertical directions. To reduce complexity, in some embodiments, when BIO is enabled, it is proposed to use different interpolation filters for the horizontal and vertical interpolation filters, and the filter size applied to the second filter process (e.g., vertical interpolation) is smaller than the filter size applied to the first filter process (e.g., horizontal interpolation). For example, a 4-tap chroma interpolation filter may be used to replace the current 8-tap interpolation filter for vertical interpolation. Doing so can reduce the complexity of generating predicted samples in the extended region by approximately half. Within a CU, samples may be generated using an 8-tap filter for vertical interpolation. To further reduce complexity, an even smaller-sized interpolation filter, such as a bilinear filter, may be used.

[0148] In one specific example, referred to herein as Option 1, to reduce worst-case BIO complexity, it is proposed to use a bilinear filter to generate sample values ​​in an extended region of the BIO and completely disable the BIO for CUs with height 4 (i.e., 4x4, 8x4, 16x4, 32x4, 64x4, and 128x4) as well as 4x8 CUs. Tables 6 and 7 show the number of multiplications and additions used per sample for various CU sizes by Option 1 and compare them with the worst-case number for normal bi-prediction. In Tables 6 and 7, highlighted rows represent CU sizes for which the BIO is disabled. In these rows, the operations associated with the corresponding BIO are set to 0, and the respective complexities are the same as for normal bi-prediction for CUs of the same size. As can be seen, in Option 1, the computational complexity peak occurs for 8x8 BIO CUs, where the number of multiplications and additions is 110% and 136% of the worst-case complexity for normal bi-prediction.

[0149] [Table 6]

[0150] [Table 7]

[0151] BIO complexity reduction by reducing the size of the expanded region As shown in Figures 5 and 6, the above-described BIO complexity reduction method still operates to interpolate two additional rows / columns of predicted samples around each boundary of the current CU. Although the simplified filter is used to reduce the number of operations, some complexity increase still occurs due to the number of samples that need to be interpolated. To further reduce BIO complexity, some embodiments propose a method to reduce the number of extended samples from two rows / columns to a single row / column at each CU boundary. Specifically, instead of using (W + 4) x (H + 4) samples as per the current BIO, some embodiments use only (W + 2) x (H + 2) samples for further complexity reduction. However, as shown in Equation (7), the gradient calculation for each sample uses the sample values ​​of both the left and right neighboring elements (for horizontal gradients) or the above and below neighboring elements (for vertical gradients). Therefore, by reducing the expanded region size to (W+2)×(H+2), the method can only calculate the gradient values ​​of samples within the CU, and therefore the existing BIO motion refinement cannot be performed directly on the 4×4 blocks located at the four corners of the CU region. To address this issue, some embodiments calculate the gradients of sample positions outside the CU (i.e.,

[0152]

number

[0153] ) and sample values ​​(i.e., I (k) A method is applied in which both (x,y) and (x,y) are set equal to their nearest neighbors within the CU. Figure 7 illustrates such a padding process for both sample values ​​and gradients. In Figure 7, the dark blocks represent predicted samples within the CU, and the white blocks represent predicted samples within the dilated region.

[0154] In the example shown in the figure, additional prediction samples for only a single row / column are generated in the extended region so that the gradients of all samples in the CU region (i.e., the dark blocks in FIG. 7) can be accurately derived. However, for the four corner sub-blocks of the CU (e.g., the sub-blocks surrounded by the thick black squares in FIG. 7), their BIO motion refinement is derived from a local region surrounding the sub-block (e.g., the region surrounded by the dashed black square in FIG. 7), so they use the gradient information of some samples in the extended region (e.g., the white blocks in FIG. 7), but that information is missing. To solve this problem, those missing gradients are padded by duplicating the gradient values ​​of the nearest boundary samples in the CU region, as indicated by the arrows in FIG. 7. Furthermore, if only gradients are padded, a problem may occur: the gradients and sample values ​​used for sample positions in the extended region are shifted, i.e., the sample values ​​are their true sample values, while the gradients are the gradients of their neighboring samples in the CU. This may reduce the accuracy of the derived BIO motion refinement. Therefore, to avoid such inconsistencies, both the sample values ​​and the gradients of the samples in the extended region are padded during the BIO derivation process.

[0155] To achieve even greater complexity reduction, in some embodiments, the proposed padding method is combined with the method of using a simplified interpolation filter and the method of disabling BIO for certain CU sizes mentioned above. In one specific example, referred to as Option 2, it is proposed to reduce the extended sample area to (W+2)×(H+2) by using padded samples and gradients for BIO derivation and applying a bilinear filter to generate extended samples in one additional row / column around the CU boundary. Furthermore, BIO is disabled for CUs with height 4 (i.e., 4×4, 8×4, 16×4, 32×4, 64×4, and 128×4) and 4×8 CUs. Tables 8 and 9 show the corresponding number of multiplications and additions used per sample for various CU sizes after applying such a method and compare them with the worst-case number for normal bi-prediction. As with Tables 6 and 7, the highlighted rows represent CU sizes for which BIO is disabled. As seen with option 2, the number of multiplications and additions is 103% and 129% of the worst-case complexity of regular bi-prediction.

[0156] [Table 8]

[0157] [Table 9]

[0158] In another embodiment, it is proposed to interpolate predicted samples within the extended region of a BIO CU using the default 8-tap filter. However, to reduce BIO complexity, the size of the extended region is reduced from (W+4)×(H+4) to (W+2)×(H+2), i.e., there is one additional row / column at each of the top, left, bottom, and right boundaries of the CU. As described in FIG. 7, to calculate missing gradients and avoid misalignment between predicted samples and gradients, both the sample values ​​and gradients of samples within the extended region are padded during the BIO derivation process. Furthermore, similar to options 1 and 2, certain block sizes (e.g., all CUs with a height equal to 4, and CUs with sizes of 4×8, 4×16, 8×8, and 16×8) can be disabled.

[0159] In another embodiment, it is proposed to discard all prediction samples in the extended area of ​​the BIO CU, so that the BIO process only concerns the interpolation of prediction samples within the current CU area. By doing so, the corresponding BIO operation of generating prediction samples becomes the same as the normal bi-prediction operation. However, due to the reduced number of interpolated samples, the gradients of boundary samples on the current CU cannot be derived by the normal BIO process. In such a case, it is proposed to pad the gradient values ​​of the intra-CU prediction samples to be the gradients of the samples at the CU boundary.

[0160] In another embodiment of the present disclosure, referred to herein as Option 3, it is proposed to reduce the extended sample area to (W+2)×(H+2) by using padded samples and gradients for BIO derivation, and apply the same 8-tap interpolation used in normal motion compensation to generate extended samples in one additional row / column around the CU boundary. Furthermore, in Option 3, BIO is disabled for CUs with height 4 (i.e., 4×4, 8×4, 16×4, 32×4, 64×4, and 128×4) as well as 4×8 CUs.

[0161] For the above-described methods, although the size of the extended region is reduced from (W+4)×(H+4) to (W+2)×(H+2), the methods may still operate to interpolate one additional row / column around the boundary of the BIO CU. As shown in Tables 8 and 9, such methods may still result in a non-negligible increase in overall BIO complexity. To further reduce BIO computational complexity, some embodiments propose directly using reference samples located at integer sample positions (without interpolation) and taken directly from the reference picture as samples within the extended region and using them to derive gradient values ​​of boundary samples of the current CU. Figure 10 shows an embodiment in which integer reference samples are used as extended samples for BIO derivation. As shown in Figure 10, samples within the CU region (shaded blocks) are generated by applying a default 8-tap interpolation filter. However, for samples in the extended region (unshaded blocks), instead of using an interpolation filter (e.g., a bilinear filter or an 8-tap interpolation filter), their sample values ​​are directly set equal to the corresponding sample values ​​at integer sample positions in the reference picture. By doing so, all operations introduced by interpolation of the extended samples can be avoided, thereby providing a significant complexity reduction of the BIO. In another embodiment, instead of using integer reference samples, it is proposed to directly set the samples in the extended region equal to the nearest neighbor samples at the CU boundary.

[0162] Because only a single row / column of additional prediction samples is used for BIO derivation in the above method, in some embodiments, a padding method such as that shown in FIG. 7 may be applied to pad both sample values ​​and gradients of samples on CU boundaries to the extended region during the BIO derivation process. In such embodiments, BIO may be disabled for certain CU sizes to reduce worst-case BIO complexity. For example, in some embodiments, BIO may be disabled for CUs with height 4 (i.e., 4×4, 8×4, 16×4, 32×4, 64×4, and 128×4) as well as 4×8 CUs.

[0163] Disable BIO for CUs predicted by sub-block mode In HEVC, each prediction unit has at most one MV for the prediction direction. In contrast, the current VTM / BMS-2.0 includes two sub-block level inter prediction techniques, including advanced temporal motion vector prediction (ATMVP) and affine prediction. In these coding modes, a video block is further divided into multiple small sub-blocks, and motion information for each sub-block is derived separately. The motion information for each sub-block is used to generate a prediction signal for the block in the motion compensation stage. On the other hand, the current BIO in BMS-2.0 can provide motion refinement at the 4x4 sub-block level in addition to CU-level motion compensation prediction. Due to the fine granularity of the motion field of sub-block coded CUs, the additional coding benefit obtained from the refined motion by BIO may be very limited. In some embodiments, BIO is disabled for CUs coded by sub-block mode.

[0164] Disable BIO for CUs using a predetermined prediction mode In VVC, some inter bi-prediction modes are based on the assumption that motion is linear and that motion vectors in list 0 and list 1 are symmetric. These modes include the merge with MVD mode (MMVD) described in Non-Patent Document 2 and the decoder-side MV derivation using bilateral matching described in Non-Patent Document 3. Because these modes generate predictions using symmetric motion, applying BIO to these predictions may not be efficient. To reduce complexity, in some embodiments, BIO is disabled for coding units predicted using symmetric modes such as MMVD or decoder-side MV derivation using bilateral matching.

[0165] Multi-hypothesis prediction for intra modes is described in Non-Patent Document 4. Multi-hypothesis prediction for intra modes combines one intra prediction and one inter merged indexed prediction. Because one prediction is obtained from the intra prediction, in some embodiments, BIO is disabled for coding units predicted using this combined inter and intra multiple hypothesis prediction.

[0166] Multi-hypothesis inter prediction is described in Non-Patent Document 5. In multi-hypothesis inter prediction, up to two additional MVs are signaled for one inter merge coding CU. There are up to four MVs per CU: two MVs from explicit signaling and two MVs from merge candidates indicated by the merge index. These multiple inter predictions are combined with a weighted average. In this case, the prediction may be sufficiently good. To reduce complexity, in some embodiments, BIO may be disabled for coding units predicted using this multi-hypothesis inter prediction mode.

[0167] Coded Bitstream Structure Figure 8 illustrates an example of a coded bitstream structure. Coded bitstream 1300 includes several NAL (Network Abstraction Layer) units 1301. NAL units may include coded sample data, such as coded slice data 1306, or high-level syntax metadata, such as parameter set data, slice headers 1305, or supplemental enhancement information data 1307 (sometimes referred to as SEI messages). Parameter sets are high-level syntax structures that include basic syntax elements that may apply to multiple bitstream layers (e.g., video parameter sets 1302 (VPS)), to a coded video sequence within a single layer (e.g., sequence parameter sets 1303 (SPS)), or to several coded pictures within a coded video sequence (e.g., picture parameter sets 1304 (PPS)). Parameter sets may be transmitted together with the coded pictures of the video bitstream or via other means (including out-of-band transmission using a reliable channel, hard coding, etc.). The slice header 1305 is a high-level syntax structure that may be relatively small or contain some picture-related information relevant only to a particular slice or picture type. The SEI message 1307 carries information that may not be needed by the decoding process but can be used for various other purposes, such as picture output timing or display and loss detection and concealment.

[0168] Communication Devices and Systems Figure 9 illustrates an example of a communication system. The communication system 1400 may include an encoder 1402, a communication network 1404, and a decoder 1406. The encoder 1402 can communicate with the communication network 1404 via a connection 1408, which may be a wired or wireless connection. The encoder 1402 may be similar to the block-based video encoder of Figure 2A. The encoder 1402 may include a single-layer codec (e.g., Figure 2A) or a multi-layer codec. The decoder 1406 can communicate with the communication network 1404 via a connection 1410, which may be a wired or wireless connection. The decoder 1406 may be similar to the block-based video decoder of Figure 2B. The decoder 1406 may include a single-layer codec (e.g., Figure 2B) or a multi-layer codec.

[0169] The encoder 1402 and / or decoder 1406 may be incorporated into a wide variety of wired communication devices and / or wireless transmit / receive units (WTRUs), including, but not limited to, digital televisions, wireless broadcast systems, network elements / terminals, servers such as content servers or web servers (e.g., Hypertext Transfer Protocol (HTTP) servers), personal digital assistants (PDAs), laptop or desktop computers, tablet computers, digital cameras, digital recording devices, video game devices, video game consoles, cellular or satellite radiotelephones, digital media players, and the like.

[0170] The communication network 1404 may be any suitable type of communication network. For example, the communication network 1404 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communication network 1404 enables multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication network 1404 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), and / or single-carrier FDMA (SC-FDMA). The communication network 1404 may include multiple connected communication networks. The communication network 1404 may include the Internet and / or one or more private commercial networks, such as cellular networks, WiFi hotspots, and / or Internet Service Provider (ISP) networks.

[0171] Encoder and decoder systems and methods In some embodiments, a method for encoding or decoding video is provided, the method including, for at least one current block in a video coded using bidirectional optical flow, calculating a first gradient component based on a first prediction signal from a first reference picture, calculating a second gradient component based on a second prediction signal from a second reference picture, summing the first gradient component and the second gradient component and performing a downward bit-shift of the resulting sum to generate a reduced bitwidth correlation parameter, calculating a motion refinement based at least in part on the reduced bitwidth correlation parameter, and predicting the block using the bidirectional optical flow using the calculated motion refinement.

[0172] In some embodiments, the first gradient component is ∂I (0) / ∂x and the second gradient component is ∂I (1) / ∂x and the reduced bitwidth correlation parameter is:

[0173]

number

[0174] In some embodiments, the first gradient component is ∂I (0) / ∂y and the second gradient component is ∂I (1) / ∂y, and the reduced bitwidth correlation parameter is:

[0175]

number

[0176] In some embodiments, a method for encoding or decoding video is provided, the method including: for at least one current block in a video coded using bidirectional optical flow, generating a reduced bitwidth correlation parameter by subtracting a first prediction signal based on a first reference picture from a second prediction signal based on a second reference picture and performing a downward bit-shift of the resulting difference; calculating a motion refinement based at least in part on the reduced bitwidth correlation parameter; and predicting the block using bidirectional optical flow using the calculated motion refinement.

[0177] In some embodiments, the first predicted signal is I (0) and the second predicted signal is I (1) and the reduced bitwidth correlation parameter is θ(i,j)=(I (1) (i,j)-I (0) (i,j))≫n b is.

[0178] In some embodiments, a method for encoding or decoding video is provided, the method comprising: for at least one current block in a video coded using bidirectional optical flow:

[0179]

number

[0180] calculating a horizontal motion refinement as

[0181]

number

[0182] calculating a vertical motion refinement as and predicting the block using bidirectional optical flow using the calculated horizontal and vertical motion refinement. S1=Σ (i,j)∈Ω Ψ x (i,j)·Ψ x (i,j), S2=Σ (i,j)∈Ω Ψ x (i,j)·Ψ y (i,j), S3=Σ (i,j)∈Ω θ(i,j)·Ψ x (i,j), S5=Σ (i,j)∈Ω Ψ y (i,j)·Ψ y (i,j), and S6=Σ (i,j)∈Ω θ(i,j)·Ψ y (i,j) is.

[0183] In some embodiments, a method for encoding or decoding video is provided. The method includes, for at least one current block in a video coded using bidirectional optical flow, generating a first predicted signal of a reduced bit width by performing a downward bit shift on a first predicted signal from a first reference picture, generating a second predicted signal of a reduced bit width by performing a downward bit shift on a second predicted signal from a second reference picture, generating a reduced bit width correlation parameter by subtracting the first predicted signal of the reduced bit width from the second predicted signal of the reduced bit width, calculating a motion refinement based at least in part on the reduced bit width correlation parameter, and predicting a block using bidirectional optical flow using the calculated motion refinement. In some such embodiments, the reduced bit width correlation parameter is θ(i,j)=(I (1) (i,j)≫n b )-(I (0) (i,j)≫n b ) is.

[0184] In some embodiments, a method for coding video is provided, the method including, for at least one current block in video coded using bidirectional optical flow, calculating a first gradient component for a reduced bit width based on a first prediction signal for the reduced bit width from a first reference picture, calculating a second gradient component for the reduced bit width based on a second prediction signal for the reduced bit width from a second reference picture, summing the first reduced bit width gradient component and the second reduced bit width gradient component to generate a reduced bit width correlation parameter, calculating motion refinement based at least in part on the reduced bit width correlation parameter, and predicting the block using bidirectional optical flow using the calculated motion refinement.

[0185] In some such embodiments, the first gradient component of the reduced bit width is ∂I (0) / ∂x, and the second gradient component of the reduced bit width is ∂I (1) / ∂x, and the reduced bit width correlation parameter is:

[0186]

number

[0187] In some embodiments, the first gradient component of the reduced bit width is ∂I (0) / ∂y, and the second gradient component of the reduced bit width is ∂I (1) / ∂y, and the reduced bit width correlation parameter is:

[0188]

number

[0189] In some embodiments, the step of calculating the first gradient component at the reduced bitwidth based on the first prediction signal at the reduced bitwidth from the first reference picture comprises:

[0190]

number

[0191] and the step of calculating the second gradient component of the reduced bit width based on the second prediction signal of the reduced bit width from the second reference picture includes:

[0192]

number

[0193] The method includes the step of calculating:

[0194] In some embodiments, the step of calculating the first gradient component at the reduced bitwidth based on the first prediction signal at the reduced bitwidth from the first reference picture comprises:

[0195]

number

[0196] Calculating The step of calculating a second gradient component of a reduced bit width based on a second prediction signal of a reduced bit width from a second reference picture includes:

[0197]

number

[0198] The method includes the step of calculating:

[0199] In some embodiments, a method of coding video is provided, the method including: for at least one current block in video coded using bidirectional optical flow, generating a first motion-compensated prediction signal and a second motion-compensated prediction signal for samples in the current block, where the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a first interpolation filter having a first number of taps; generating a first motion-compensated prediction signal and a second motion-compensated prediction signal for samples in an extended region around the current block, where the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples outside the current block are generated using a second interpolation filter having a second number of taps that is less than the first number of taps; calculating a motion refinement based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal; and predicting a block using bidirectional optical flow using the calculated motion refinement.

[0200] In some embodiments, the first interpolation filter is an 8-tap filter and the second interpolation filter is a 2-tap filter. In some embodiments, the second interpolation filter is a bilinear interpolation filter.

[0201] In some embodiments, a method for coding video is provided, the method including: generating, for at least one current block in video coded using bidirectional optical flow, a first motion-compensated prediction signal and a second motion-compensated prediction signal, wherein the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a horizontal interpolation filter having a first number of taps and a vertical interpolation filter having a second number of taps that is less than the first number of taps; calculating a motion refinement based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal; and predicting the block using bidirectional optical flow using the calculated motion refinement.

[0202] In some embodiments, a method of coding video is provided, the method including: generating, for at least one current block in video coded using bidirectional optical flow, a first motion-compensated prediction signal and a second motion-compensated prediction signal, wherein the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a horizontal interpolation filter having a first number of taps and a vertical interpolation filter having a second number of taps, the horizontal and vertical filters being applied in a predetermined order, with filters applied earlier in the order having a greater number of taps than filters applied later in the order; calculating a motion refinement based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal; and predicting the block using bidirectional optical flow using the calculated motion refinement.

[0203] In some embodiments, a method of coding a video including a plurality of coding units is provided, the method including, for a plurality of coding units in a video coded using bi-prediction, disabling bi-directional optical flow for coding units having at least a height of 4, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is not disabled. In some such embodiments, bi-directional optical flow is further disabled for coding units having a height of 8 and a width of 4.

[0204] In some embodiments, a method of coding a video including a plurality of coding units is provided, the method including, for a plurality of coding units in a video coded using bi-prediction, disabling bi-directional optical flow for coding units having heights less than or equal to a threshold height, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for the bi-predicted coding units for which bi-predicted optical flow is not disabled.

[0205] In some embodiments, a method of coding video is provided, the method including: for at least one current block in video coded using bidirectional optical flow, generating first and second motion-compensated prediction signals for samples in the current block, generating first and second values ​​for samples in an extended region around the current block, the extended region not including samples more than one row or column away from the current block, calculating motion refinement based at least in part on the first and second motion-compensated prediction signals and the first and second values ​​for samples in the extended region, and predicting a block using bidirectional optical flow using the calculated motion refinement.

[0206] In some such embodiments, generating first values ​​for the samples in the extended region includes setting each first sample value in the extended region equal to a first predicted sample value of its respective nearest neighbor element in the current block. In some embodiments, generating second values ​​for the samples in the extended region includes setting each second sample value in the extended region equal to a second predicted sample value of its respective nearest neighbor element in the current block.

[0207] Some embodiments further include generating first and second gradient values ​​at samples in an extended region around the current block, wherein generating the first gradient values ​​at samples in the extended region includes setting each first gradient value in the extended region equal to a gradient value calculated at its respective nearest neighbor in the current block using the first prediction signal, and generating the second gradient values ​​at samples in the extended region includes setting each second gradient value in the extended region equal to a gradient value calculated at its respective nearest neighbor in the current block using the second prediction signal.

[0208] In some embodiments, a method of coding a video including a plurality of coding units is provided, the method including, for a plurality of coding units in a video coded using bi-prediction, disabling bi-directional optical flow for at least coding units predicted using sub-block level inter prediction techniques, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is not disabled.

[0209] In some such embodiments, bi-prediction is disabled for at least coding units predicted using advanced temporal motion vector prediction (ATMVP).

[0210] In some embodiments, bi-prediction is disabled for at least coding units predicted using affine prediction.

[0211] In some embodiments, a method for coding a video including a plurality of coding units is provided, the method including, for a plurality of coding units in a video coded using bi-prediction, disabling bi-directional optical flow for coding units having a height of at least 4, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is not disabled, wherein performing bi-prediction with bi-directional optical flow for each current coding unit includes generating a first motion-compensated prediction signal and a second motion-compensated prediction signal for samples in the current coding unit, the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block. the first and second motion compensated prediction signals for samples in an extended region around the current coding unit are generated using a first interpolation filter having a first number of taps; generating the first and second motion compensated prediction signals for samples in an extended region around the current coding unit, wherein the first and second motion compensated prediction signals for samples outside the current coding unit are generated using a second interpolation filter having a second number of taps that is less than the first number of taps; calculating a motion refinement based at least in part on the first and second motion compensated prediction signals; and predicting the current coding unit using bidirectional optical flow using the calculated motion refinement.

[0212] In some such embodiments, the first interpolation filter is an 8-tap filter and the second interpolation filter is a 2-tap filter. In some embodiments, the second interpolation filter is a bilinear interpolation filter.

[0213] In some embodiments, bidirectional optical flow is also disabled for coding units having a height of eight and a width of four.

[0214] In some embodiments, a method for coding a video comprising multiple coding units is provided. The method includes, for a plurality of coding units in a video coded using bi-prediction, disabling bi-directional optical flow for coding units having a height of at least 4; performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled; and performing bi-prediction using bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is not disabled, wherein performing bi-prediction using bi-directional optical flow for each current coding unit includes generating first and second motion-compensated prediction signals for samples in the current coding unit; generating first and second values ​​for samples in an extended region around the current coding unit, the extended region not including samples multiple rows or columns away from the current coding unit; calculating motion refinement based at least in part on the first and second motion-compensated prediction signals and the first and second values ​​for samples in the extended region; and predicting the current coding unit using bi-directional optical flow using the calculated motion refinement.

[0215] In some embodiments, generating first values ​​for samples in the extended region includes setting each first sample value in the extended region equal to the first predicted sample value of its respective nearest neighbor in the current coding unit.

[0216] In some embodiments, generating second values ​​for the samples in the extended region includes setting each second sample value in the extended region equal to the second predicted sample value of its respective nearest neighbor element in the current coding unit.

[0217] Some embodiments further include generating first and second gradient values ​​at samples within an extended region around the current coding unit, wherein generating the first gradient values ​​at samples within the extended region includes setting each first gradient value within the extended region equal to a gradient value calculated at its respective nearest neighboring element within the current coding unit using the first prediction signal, and generating the second gradient values ​​at samples within the extended region includes setting each second gradient value within the extended region equal to a gradient value calculated at its respective nearest neighboring element within the current coding unit using the second prediction signal.

[0218] In some such embodiments, bidirectional optical flow is further disabled for coding units having a height of eight and a width of four.

[0219] In some embodiments, a method for coding video is provided, the method including: for at least one current block in video coded using bidirectional optical flow, generating a first motion-compensated prediction signal and a second motion-compensated prediction signal for samples in the current block, where the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples in the current block are generated using a horizontal interpolation filter having a first number of taps and a vertical interpolation filter having a first number of taps; generating a first motion-compensated prediction signal and a second motion-compensated prediction signal for samples in an extended region around the current block, where the first motion-compensated prediction signal and the second motion-compensated prediction signal for samples outside the current block are generated using a horizontal interpolation filter having the first number of taps and a vertical interpolation filter having a second number of taps that is less than the first number of taps; calculating a motion refinement based at least in part on the first motion-compensated prediction signal and the second motion-compensated prediction signal; and predicting a block using bidirectional optical flow using the calculated motion refinement.

[0220] In some embodiments, a method for coding a video including a plurality of coding units is provided, the method including, for a plurality of coding units in the video coded using bi-prediction, disabling bi-directional optical flow for at least coding units predicted using a symmetric prediction mode, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for bi-predicted coding units for which bi-predicted optical flow is not disabled.

[0221] In some embodiments, bi-prediction is disabled for coding units predicted using at least merge mode with MVD (MMVD). In some embodiments, bi-prediction is disabled for coding units predicted using decoder-side MV derivation with bilateral matching.

[0222] In some embodiments, a method of coding a video including a plurality of coding units is provided, the method including, for a plurality of coding units in the video coded using bi-prediction, disabling bi-directional optical flow for coding units predicted using multiple hypothesis prediction for at least an intra mode, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is not disabled.

[0223] In some embodiments, a method is provided for coding a video that includes a plurality of coding units, the method including, for a plurality of coding units in the video coded using bi-prediction, disabling bi-directional optical flow for at least coding units predicted using multiple-hypothesis inter prediction, performing bi-prediction without bi-directional optical flow for the bi-predicted coding units for which bi-directional optical flow is disabled, and performing bi-prediction with bi-directional optical flow for bi-predicted coding units for which bi-directional optical flow is not disabled.

[0224] It should be noted that one or more of the various hardware elements of the described embodiments are referred to as “modules” that perform (i.e., perform, execute, etc.) the various functions described herein with respect to the respective modules. As used herein, a module includes hardware deemed appropriate for a given implementation by one of ordinary skill in the art (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more ASICs, one or more FPGAs, one or more memory devices). It should also be noted that each described module may include executable instructions to perform one or more functions described as being performed by the respective module, which may take or include the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and may be stored on one or more suitable non-transitory computer-readable media, such as commonly referred to as RAM, ROM, etc.

[0225] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with the other features and elements. The methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, ROM, RAM, registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and DVDs. A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. Obtaining a first prediction signal array from a first reference picture; obtaining a first array of first gradient components, the first array including performing a right bit shift on two samples of the first predicted signal array and determining a difference between the two right bit shifted samples of the first predicted signal array; Obtaining a second prediction signal array from a second reference picture; obtaining a second array of first gradient components, the second array including performing a right bit shift on two samples of the second predicted signal array and determining a difference between the two right bit shifted samples of the second predicted signal array; obtaining an intermediate parameter array of first components, the intermediate parameter array including determining a sum of a first array of first gradient components and a second array of first gradient components; obtaining a motion refinement of at least a first component based at least in part on the intermediate parameter array of the first component; generating a prediction of a current block in the video using bidirectional optical flow using motion refinement of at least the first component; A video decoding method comprising:

2. 2. The method of claim 1, wherein obtaining the intermediate parameter array of the first components comprises performing a right bit shift on a sum of the first array of first gradient components and the second array of first gradient components.

3. 3. The method of claim 1, further comprising: obtaining a signal difference parameter array, the method comprising: performing a right bit shift on each of the first prediction signal array and the second prediction signal array before obtaining a difference between the first prediction signal array and the second prediction signal array; and wherein motion refinement of the first component is based at least in part on the signal difference parameter array.

4. the first component intermediate parameter array is a horizontal intermediate parameter array, and the first component motion refinement is a horizontal motion refinement, and the method comprises: obtaining a signal horizontal gradient correlation parameter by summing components of the element-wise multiplication of the signal difference parameter array and the horizontal mean parameter array; The method of claim 3 , wherein obtaining the horizontal motion refinement comprises bit-shifting the signal horizontal gradient correlation parameter to obtain the horizontal motion refinement.

5. the first gradient component is a horizontal gradient; For at least a plurality of horizontal gradients of the first array having coordinates (i, j), two samples of the first predicted signal array have coordinates (i+1, j) and (i-1, j); For at least a plurality of horizontal gradients of the second array having coordinates (i, j), two samples of the second predicted signal array have coordinates (i+1, j) and (i-1, j). The method of claim 4.

6. obtaining a first array of second gradient components, the first array including performing a right bit shift on two samples of the first predicted signal array and determining a difference between the two right bit shifted samples of the first predicted signal array; obtaining a second array of second gradient components, the second array including performing a right bit shift on two samples of the second predicted signal array and determining a difference between the two right bit shifted samples of the second predicted signal array; obtaining an intermediate parameter array of second components, the intermediate parameter array including determining a sum of the first array of second gradient components and the second array of second gradient components; obtaining a motion refinement of the second component based at least in part on the intermediate parameter array of the second component; further comprising The method of claim 3 , wherein the prediction of the current block further uses motion refinement of the second component.

7. the second gradient component is a vertical gradient; For at least a plurality of vertical gradients of the first array having coordinates (i, j), two samples of the first predicted signal array have coordinates (i, j+1) and (i, j-1); For at least a plurality of vertical gradients of the second array having coordinates (i, j), two samples of the second predicted signal array have coordinates (i, j+1) and (i, j-1). The method of claim 6.

8. Obtaining a first prediction signal array from a first reference picture; obtaining a first array of first gradient components, the first array including performing a right bit shift on two samples of the first predicted signal array and determining a difference between the two right bit shifted samples of the first predicted signal array; Obtaining a second prediction signal array from a second reference picture; obtaining a second array of first gradient components, the second array including performing a right bit shift on two samples of the second predicted signal array and determining a difference between the two right bit shifted samples of the second predicted signal array; obtaining an intermediate parameter array of first components, the intermediate parameter array including determining a sum of a first array of first gradient components and a second array of first gradient components; obtaining a motion refinement of at least a first component based at least in part on the intermediate parameter array of the first component; generating a prediction of a current block in the video using bidirectional optical flow using motion refinement of at least the first component; 1. A video decoder apparatus comprising one or more processors configured to execute:

9. 9. The apparatus of claim 8, wherein obtaining the intermediate parameter array of the first components comprises performing a right bit shift on a sum of the first array of first gradient components and the second array of first gradient components.

10. 10. The apparatus of claim 8 or claim 9, further comprising: obtaining a signal difference parameter array, the signal difference parameter array including performing a right bit shift on each of the first prediction signal array and the second prediction signal array before obtaining a difference between the first prediction signal array and the second prediction signal array; and wherein motion refinement of the first component is based at least in part on the signal difference parameter array.

11. the first component intermediate parameter array is a horizontal intermediate parameter array, and the first component motion refinement is a horizontal motion refinement, and the one or more processors: obtaining a signal horizontal gradient correlation parameter by summing components of the element-wise multiplication of the signal difference parameter array and the horizontal mean parameter array; The apparatus of claim 10 , wherein obtaining the horizontal motion refinement is further configured to perform bit-shifting the signal horizontal gradient correlation parameter to obtain the horizontal motion refinement.

12. the first gradient component is a horizontal gradient; For at least a plurality of horizontal gradients of the first array having coordinates (i, j), two samples of the first predicted signal array have coordinates (i+1, j) and (i-1, j); For at least a plurality of horizontal gradients of the second array having coordinates (i, j), two samples of the second predicted signal array have coordinates (i+1, j) and (i-1, j).

12. The apparatus of claim 11.

13. Obtaining a first prediction signal array from a first reference picture; obtaining a first array of first gradient components, the first array including performing a right bit shift on two samples of the first predicted signal array and determining a difference between the two right bit shifted samples of the first predicted signal array; Obtaining a second prediction signal array from a second reference picture; obtaining a second array of first gradient components, the second array including performing a right bit shift on two samples of the second predicted signal array and determining a difference between the two right bit shifted samples of the second predicted signal array; obtaining an intermediate parameter array of first components, the intermediate parameter array including determining a sum of a first array of first gradient components and a second array of first gradient components; obtaining a motion refinement of at least a first component based at least in part on the intermediate parameter array of the first component; generating a prediction of a current block in the video using bidirectional optical flow using motion refinement of at least the first component; A video encoding method comprising:

14. 14. The method of claim 13, wherein obtaining the intermediate parameter array of first components comprises performing a right bit shift on a sum of the first array of first gradient components and the second array of first gradient components.

15. 15. The method of claim 13 or 14, further comprising: obtaining a signal difference parameter array, the method comprising performing a right bit shift on each of the first prediction signal array and the second prediction signal array before obtaining the difference between the first prediction signal array and the second prediction signal array; and wherein motion refinement of the first component is based at least in part on the signal difference parameter array.

16. the first component intermediate parameter array is a horizontal intermediate parameter array, and the first component motion refinement is a horizontal motion refinement, and the method comprises: obtaining a signal horizontal gradient correlation parameter by summing components of the element-wise multiplication of the signal difference parameter array and the horizontal mean parameter array; The method of claim 15 , wherein obtaining the horizontal motion refinement comprises bit-shifting the signal horizontal gradient correlation parameter to obtain the horizontal motion refinement.

17. Obtaining a first prediction signal array from a first reference picture; obtaining a first array of first gradient components, the first array including performing a right bit shift on two samples of the first predicted signal array and determining a difference between the two right bit shifted samples of the first predicted signal array; Obtaining a second prediction signal array from a second reference picture; obtaining a second array of first gradient components, the second array including performing a right bit shift on two samples of the second predicted signal array and determining a difference between the two right bit shifted samples of the second predicted signal array; obtaining an intermediate parameter array of first components, the intermediate parameter array including determining a sum of a first array of first gradient components and a second array of first gradient components; obtaining a motion refinement of at least a first component based at least in part on the intermediate parameter array of the first component; generating a prediction of a current block in the video using bidirectional optical flow using motion refinement of at least the first component; 1. A video encoder apparatus comprising: one or more processors configured to execute:

18. 20. The apparatus of claim 17, wherein obtaining the intermediate parameter array of first components comprises performing a right bit shift on a sum of the first array of first gradient components and the second array of first gradient components.

19. 19. The apparatus of claim 17 or claim 18, further comprising: obtaining a signal difference parameter array, the signal difference parameter array including performing a right bit shift on each of the first prediction signal array and the second prediction signal array before obtaining a difference between the first prediction signal array and the second prediction signal array; and wherein motion refinement of the first component is based at least in part on the signal difference parameter array.

20. the first gradient component is a vertical gradient; For at least a plurality of vertical gradients of the first array having coordinates (i, j), two samples of the first predicted signal array have coordinates (i, j+1) and (i, j-1); For at least a plurality of vertical gradients of the second array having coordinates (i, j), two samples of the second predicted signal array have coordinates (i, j+1) and (i, j-1).

20. The apparatus of claim 19.

Citation Information

Patent Citations

  • JPP7311589B

  • JPP7553659B

  • Spatial prediction method, image decoding method, and image coding method

    US20130028530A1

  • Spatial prediction method, image decoding method, and image encoding method

    WO2011129084A1

  • Method and apparatus of motion refinement based on BI-directional optical flow for video coding

    WO2018166357A1