Precision refinement for motion compensation by optical flow

By refining motion-compensated predictions in video coding systems through the calculation and application of sample difference values derived from spatial gradients and motion vector refinements, the method addresses the limitations of existing systems, achieving improved accuracy and efficiency in video encoding and decoding.

JP7700337B2Active Publication Date: 2025-06-30INTERDIGITAL VC HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024141925
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-21
Filing Date
2024-08-23
Publication Date
2025-06-30
Estimated Expiration
2040-06-18

AI Technical Summary

Technical Problem

Existing video coding systems face limitations in achieving precise motion compensation due to the accuracy constraints of motion vectors, leading to inefficiencies in prediction error minimization and increased signaling overhead.

Method used

The method involves refining motion-compensated predictions by calculating a sample difference value using the scalar product of spatial gradients and motion vector refinements, and adding this value to the initial predicted sample value to generate a refined sample value.

Benefits of technology

This approach enhances prediction accuracy to pixel-level granularity without significantly increasing complexity, while maintaining the same worst-case memory access bandwidth as traditional block-level motion compensation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700337000021
    Figure 0007700337000021
  • Figure 0007700337000022
    Figure 0007700337000022
  • Figure 0007700337000023
    Figure 0007700337000023
Patent Text Reader

Abstract

To provide systems and methods for refining motion-compensated predictions in block-based video coding.SOLUTION: A video coding method comprises: obtaining an initial predicted sample value, based on motion-compensated prediction, for each of sample positions in a current block or subblock of samples; determining a motion vector refinement; determining, at each of the sample positions, a spatial gradient of sample values; determining a sample difference value, based on a scalar product of the spatial gradient and the motion vector refinement, for each of the sample positions; and modifying the initial predicted sample value based on the sample difference value.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application is a non - provisional application claiming the benefit under 35 U.S.C. § 119(e) from U.S. Provisional Patent Application No. 62 / 864,825, filed on June 21, 2019, entitled "Precision Refinement for Motion Compensation with Optical Flow", which is hereby incorporated by reference in its entirety.

Background Art

[0002] Video coding systems are widely used to compress digital video signals to reduce the storage requirements and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block - based, wavelet - based, and object - based systems, today, block - based hybrid video coding systems are the most widely used and deployed. Examples of block - based video coding systems include international video coding standards such as MPEG - 1 / 2 / 4 Part 2, H.264 / MPEG - 4 Part 10 AVC, VC - 1, and High Efficiency Video Coding (HEVC), which were developed by ITU - T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG's JCT - VC (Joint Collaborative Team on Video Coding).

[0003] In October 2017, the ITU-T and ISO / IEC issued a Call for Proposals (CfP) on video compression with capabilities beyond HEVC. In April 2018, 22 CfP responses regarding the standard dynamic range category were accepted and evaluated at the 10th JVET meeting, demonstrating that the compression ratio improvement over HEVC was approximately 40%. Based on such evaluation results, the Joint Video Experts Team (JVET) launched a new project to develop a new generation of video coding standard called Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate the reference implementation of the VVC standard. In the initial VTM-1.0, most coding modules including intra prediction, inter prediction, transform / inverse transform, quantization / inverse quantization, and in-loop filters follow the existing HEVC design, except that a multi-type tree-based block partitioning structure is used in the VTM. On the other hand, another reference software base called the Benchmark Set (BMS) was also generated to facilitate the evaluation of new coding tools. In the BMS codebase, a list of coding tools inherited from the Joint Exploration Model (JEM), which provides higher coding efficiency and moderate implementation complexity, is included on top of the VTM and is used as a benchmark when evaluating similar coding techniques in the VVC standardization process. Specifically, there are 9 JEM coding tools integrated into BMS-1.0, which include 65 angular intra prediction directions, modified coefficient coding, advanced multiple transform (AMT) + 4×4 non-separable second-order transform (NSST), affine motion model, generalized adaptive loop filter (GALF), advanced temporal motion vector prediction (ATMVP), adaptive motion vector accuracy, decoder-side motion vector refinement (DMVR), and linear model (LM) chroma mode.

Summary of the Invention

[0004] The embodiments described herein include methods used in video encoding and decoding (collectively referred to as "coding"). Systems and methods for refining motion-compensated prediction in block-based video coding are described. Video coding methods according to some embodiments include generating an initial predicted sample value for at least a first sample position in a current block of samples using motion-compensated prediction, determining a refinement of a motion vector associated with at least the first sample position, determining a spatial gradient of the sample values at the first sample position, determining a sample difference value by calculating a scalar product of the spatial gradient and the refinement of the motion vector, and adding the sample difference value to the initial predicted sample value to generate a refined sample value.

[0005] In some embodiments, the method is executed by a decoder, and determining a refinement of the motion vector includes decoding a refinement of the motion vector from a bitstream.

[0006] Some embodiments further include decoding refinement accuracy information from a bitstream, and determining the sample difference value includes scaling the scalar product by an amount indicated by the accuracy information. Scaling the scalar product may include bit-shifting the scalar product by an amount indicated by the accuracy information.

[0007] In some embodiments, the motion-compensated prediction is performed using at least one motion vector having an initial accuracy, and the refinement accuracy information includes an accuracy difference value representing a difference between the initial accuracy and the refinement accuracy.

[0008] In some embodiments, scaling the scalar product includes right-shifting the scalar product by a number of bits equal to the sum of the accuracy difference value and the initial accuracy.

[0009] In some embodiments, the method is performed by an encoder, and determining motion vector refinement comprises selecting motion vector refinement to substantially minimize the prediction error for an input video block, and the method further comprises encoding the motion vector refinement in a bitstream.

[0010] In some embodiments, the motion vector refinement is signaled in a bitstream as an index. In some embodiments, the index can identify one of a plurality of motion vector refinements from the group consisting of (0, -1), (1, 0), (0, 1), and (-1, 0). In some other embodiments, the index can identify one of a plurality of motion vector refinements from the group consisting of (0, -1), (1, 0), (0, 1), (-1, 0), (-1, -1), (1, -1), (1, 1), and (-1, 1).

[0011] In some embodiments, the motion vector refinement is associated with all sample positions within a current block (or current sub-block). In some other embodiments, the motion vector refinement is determined on a per-sample basis and can be different for different samples.

[0012] Some embodiments include at least one processor configured to perform any of the methods described herein. In some such embodiments, a computer-readable medium (e.g., a non-transitory medium) storing instructions that operate to perform any of the methods described herein is provided.

[0013] Some embodiments include a computer-readable medium (e.g., a non-transitory medium) storing video encoded using one or more of the methods disclosed herein.

[0014] An encoder or decoder system can include a processor and a non-transitory computer-readable medium storing instructions for performing the methods described herein.

[0015] One or more of the embodiments also provide a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods described above. The embodiments also provide a computer-readable storage medium (e.g., a non-transitory medium) storing a bitstream generated according to the methods described above. The embodiments also provide a method and an apparatus for transmitting a bitstream generated according to the methods described above. The embodiments also provide a computer program product including instructions for performing any of the described methods.

Brief Description of the Drawings

[0016]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 3A - 3B

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

[0017] Exemplary devices and networks for implementing embodiments FIG. 1A is a diagram showing an exemplary communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 can be a multi-access system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 can employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero tail unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtering OFDM, filter bank multicarrier (FBMC).

[0018] As shown in Figure 1A, communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, although it is understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) can be configured to transmit and / or receive wireless signals and can include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular telephones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), home appliance devices, commercial and / or industrial wirelessly operating device networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d can also be referred to synonymously as a UE.

[0019] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as CN106 / 115, the Internet 110, and / or other network 112. By way of example, base stations 114a, 114b may be a base transceiver station (BTS), Node-B, eNode B, home Node B, home eNode B, gNB, NR NodeB, a site controller, an access point (AP), a wireless router, etc. Although base stations 114a, 114b are each shown as a single element, it will be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0020] Base station 114a may be part of RAN104 / 113, which may also include other base stations and / or network elements (not shown) such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies that may be referred to as a cell (not shown). These frequencies may be a licensed spectrum, an unlicensed spectrum, or a combination of a licensed spectrum and an unlicensed spectrum. A cell may provide wireless service coverage to a particular geographic area that is relatively fixed or that may change over time. A cell may further be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0021] Base stations 114a and 114b can communicate with one or more WTRUs 102a, 102b, 102c, 102d via an air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).

[0022] More specifically, as described above, the communication system 100 can be a multi-connection system and can employ one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a of RAN104 / 113 and the WTRUs 102a, 102b, 102c can implement wireless technologies such as Universal Mobile Telecommunications System (UMTS) terrestrial radio access (UTRA), which can use wideband CDMA (WCDMA) to establish the air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).

[0023] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement wireless technologies such as evolved UMTS terrestrial radio access (E-UTRA), which can use Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro to establish the air interface 116.

[0024] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement wireless technologies such as NR radio access, which can use New Radio (NR) to establish the air interface 116.

[0025] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interfaces utilized by WTRUs 102a, 102b, 102c may be characterized by transmissions that are sent between multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).

[0026] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., wireless fidelity (WiFi)), IEEE 802.16 (i.e., worldwide interoperability for microwave access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, interim standard 2000 (IS-2000), interim standard 95 (IS-95), interim standard 856 (IS-856), global system for mobile communications (GSM), enhanced data rates for GSM evolution (EDGE), GSM EDGE (GERAN), etc.

[0027] The base station 114b in FIG. 1A can be, for example, a wireless router, a home Node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connection in a local area such as an office, a home, a vehicle, a campus, an industrial facility, an aerial corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in FIG. 1A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.

[0028] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more WTRUs 102a, 102b, 102c, 102d. The data can have various Quality of Service (QoS) requirements such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, mobile location-based services, prepaid calls, Internet connections, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same or a different Radio Access Technology (RAT) as RAN 104 / 113. For example, in addition to being connected to a RAN 104 / 113 that can utilize New Radio (NR) radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0029] CN106 / 115 may also function as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) of the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may use the same RAT or a different RAT as the RAN 104 / 113.

[0030] Some or all of the WTRUs 102a, 102b, 102c, 102d within the communication system 100 may include multimode functionality (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that may employ a cellular-based wireless technology and a base station 114b that may employ IEEE802 wireless technology.

[0031] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with one embodiment.

[0032] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, and the transceiver 120 can be coupled to the transmit / receive element 122. Although Figure 1B shows the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.

[0033] The transmit / receive element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 can be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0034] Although the transmit / receive element 122 is shown in FIG. 1B as a single element, the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.

[0035] The transceiver 120 can be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As noted above, the WTRU 102 can have a multimode function. Thus, the transceiver 120 can include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs such as, for example, NR and IEEE 802.11.

[0036] The processor 118 of the WTRU 102 may be coupled to and receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any suitable type of memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 may include a random access memory (RAM), a read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in a memory that is not physically located on the WTRU 102, such as a server or a home computer (not shown).

[0037] The processor 118 may receive power from a power source 134 and may be configured to distribute and / or control power to other components within the WTRU 102. The power source 134 may be any suitable device for supplying power to the WTRU 102. For example, the power source 134 may include one or more dry cells (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li ion), etc.), a solar cell, a fuel cell, etc.

[0038] Processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of WTRU 102. In addition to, or instead of, information from the GPS chipset 136, WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) via air interface 116 and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that WTRU 102 may obtain location information by any suitable location determination method while remaining consistent with one embodiment.

[0039] Processor 118 may further be coupled to other peripheral devices 138 that may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connections. For example, peripheral devices 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or video), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a Frequency Modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, an activity tracker, etc. Peripheral devices 138 may include one or more sensors, where the sensors may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geographical location information sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0040] The WTRU 102 may include a full-duplex radio in which transmission and reception of some or all of both UL (e.g., for transmission) and downlink (e.g., for reception) signals (e.g., associated with a particular subframe) can occur in parallel and / or simultaneously. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference via either hardware (e.g., a choke) or signal processing via a processor (e.g., via a separate processor (not shown) or the processor 118). In one embodiment, the WRTU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe of either UL (e.g., for transmission) or downlink (e.g., for reception)).

[0041] The WTRU is described as a wireless terminal in FIGS. 1A - 1B, but in certain representative embodiments, it is contemplated that such a terminal may use a wired communication interface to the communication network (e.g., temporarily or permanently).

[0042] In a representative embodiment, the other network 112 may be a WLAN.

[0043] In view of the overview of FIGS. 1A - 1B and the corresponding description, one or more, or all, of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more, or all, of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0044] An emulation device can be designed to implement one or more tests of other devices in a laboratory environment and / or an operator network environment. For example, one or more emulation devices can perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.

[0045] One or more emulation devices can perform one or more functions, including all, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device can be utilized in a test scenario in a test lab and / or a non-deployed (e.g., test) wired and / or wireless communication network to implement tests of one or more components. One or more emulation devices can be test equipment. Direct RF coupling and / or wireless communication via an RF circuit (which can include one or more antennas) can be used by an emulation device to transmit and / or receive data.

[0046] Exemplary system. Some embodiments are implemented using a system such as the system of FIG. 1C. FIG. 1C shows a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device that includes various components described below and is configured to execute one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 1000 can be embodied as a single integrated circuit, multiple ICs, and / or individual components, either alone or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or individual components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.

[0047] System 1000 includes at least one processor 1010 configured to execute loaded instructions, for example, to implement various aspects described in this document. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits, as known in the art. System 1000 includes at least one memory 1020 (e.g., volatile memory devices and / or non-volatile memory devices). System 1000 includes a storage device 1040 that can include non-volatile memory and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 can include, by way of non-limiting example, internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0048] System 1000 includes an encoder / decoder module 1030 configured to process data, for example, to provide encoded video or decoded video. The encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device that performs an encoding function and / or a decoding function. As is known, a device may include one or both of an encoding and a decoding module. Further, the encoder / decoder module 1030 can be implemented as a separate element of System 1000 or can be incorporated within the processor 1010 as a combination of hardware and software, as known to those skilled in the art.

[0049] The program code read into the processor 1010 or the encoder / decoder 1030 to execute the various aspects described in this document is stored in the storage device 1040 and can subsequently be read into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 can store one or more of the various items during the execution of the processes described in this document. Such stored items can include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and operation logic.

[0050] In some embodiments, the internal memory of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide a working memory for the processing required during encoding or decoding. However, in other embodiments, an external memory of the processing device (e.g., the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, for example, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM is used as a working memory for video encoding and decoding operations such as MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding and is also known as H.265 and MPEG-H Part2), or VVC (Versatile Video Coding, a new standard developed by JVET, i.e., Joint Video Experts Team).

[0051] Inputs to the elements of system 1000 can be provided through various input devices as shown in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives, for example, RF signals wirelessly transmitted by a broadcast station, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other embodiments not shown in FIG. 1C include composite video.

[0052] In various embodiments, the input device of block 1130 has respective input processing elements as known in the art. For example, the RF section can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a certain frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel in certain embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting the received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, down-converting, and filtering again to a desired frequency band an RF signal transmitted via a wired (e.g., cable) medium. In various embodiments, the order of the above (and other) elements is rearranged, some of these elements are removed, and / or other elements performing similar or different functions are added. Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0053] Furthermore, the USB and / or HDMI terminals may each include an interface processor for connecting the system 1000 to other electronic devices via a USB and / or HDMI connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, as needed, within a separate input processing IC or within the processor 1010. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within the processor 1010. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements including, for example, the processor 1010 and an encoder / decoder 1030 that operates in combination with memory and storage elements to process the data stream as needed for display on an output device.

[0054] The various elements of the system 1000 may be provided within an integrated housing, within which the various elements are interconnected and can transmit data between them using an internal bus well known in the art, including a suitable connection configuration 1140, such as an Inter-IC (I2C) bus, wiring, and a printed circuit board.

[0055] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or a network card, and the communication channel 1060 can be implemented, for example, within a wired and / or wireless medium.

[0056] In various embodiments, data is streamed to or otherwise provided to system 1000 using a Wi-Fi network, such as a wireless network like IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via communication channel 1060 and communication interface 1050 that are adapted for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, which enables streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that distributes data via an HDMI connection of input block 1130. Still other embodiments provide streamed data to system 1000 using an RF connection of input block 1130. As noted above, various embodiments provide data in a non-streaming manner. Further, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0057] System 1000 can provide output signals to various output devices, including display 1100, speaker 1110, and other peripheral devices 1120. The display 1100 in various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a mobile phone, or other devices. The display 1100 can also be integrated with other components (such as in a smartphone) or be separate (such as an external monitor for a laptop). The other peripheral devices 1120 include, in various examples of embodiments, one or more of a stand-alone digital video disk (or digital versatile disk) (both terms DVR), a disk player, a stereo system, and / or an illumination system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of system 1000. For example, a disk player performs the function of playing back the output of system 1000.

[0058] In various embodiments, between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120, signaling such as AV.Link, consumer electronics control (CEC), or other communication protocols that enable device - to - device control is used to communicate control signals, with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through their respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000 within an electronic device such as a television. In various embodiments, the display interface 1070 includes, for example, a display driver such as a timing controller (T Con) chip.

[0059] The display 1100 and the speaker 1110 can alternatively be separate from one or more of the other components, for example, if the RF portion of the input 1130 is part of a separate set-top box. In various embodiments where the display 1100 and the speaker 1110 are external components, output signals can be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output portion.

[0060] Embodiments can be implemented by computer software executed by the processor 1010, or by hardware, or by a combination of hardware and software. By way of non-limiting example, embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 can be of any type suitable for the technical environment and can include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.

Best Mode for Carrying Out the Invention

[0061] Block-based video coding. Similar to HEVC, VVC is built on a block-based hybrid video coding framework. FIG. 2A gives a block diagram of a block-based hybrid video encoding system 200. Although variations of this encoder 200 are contemplated, the encoder 200 is described below without explaining all the expected variations for clarity.

[0062] Before being symbolized, the video sequence may go through a pre - encoding process (204), for example, applying a color conversion (e.g., conversion from RGB4:4:4 to YCbCr 4:2:0) to the input color image, or performing remapping of the input image components (e.g., using histogram equalization of one of the color components) to obtain a more resilient signal distribution for compression. Metadata can be attached to the bitstream in association with the pre - processing.

[0063] The input video signal 202 including the image to be encoded is split (206) and processed block by block. Some blocks may be referred to as coding units (CUs). If the CUs are different, their sizes can also be different. In VTM - 1.0, a CU can be up to 128×128 pixels. However, different from HEVC which divides blocks based only on a quad - tree, in VTM - 1.0, a coding tree unit (CTU) is divided into CUs and adapts to various local characteristics based on a quad - / binary / ternary - tree. Additionally, the concept of multiple partition unit types in HEVC is removed, and the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC - 1.0. Instead, each CU is always used as a basic unit for both prediction and conversion without additional partitions. In a multi - type tree structure, the CTU is first divided by a quad - tree structure. Next, each quad - tree leaf node can be further divided by binary and ternary - tree structures. There are five splitting types: four - way split, vertical two - way split, horizontal two - way split, vertical three - way split, and horizontal three - way split.

[0064] In the encoder of FIG. 2A, spatial prediction (208) and / or temporal prediction (210) may be performed. Spatial prediction (or "intra prediction") uses samples (referred to as reference samples) of already encoded adjacent blocks within the same video image / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as "inter prediction" or "motion-compensated prediction") uses pixels reconstructed from already encoded video images to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU can be signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, if multiple reference images are supported, the reference image index can be additionally transmitted and used to identify from which reference image in the reference image store (212) the temporal prediction signal comes.

[0065] The mode decision block (214) in the symbolizer selects the best prediction mode, for example, based on the rate distortion optimization method. This selection can be made after spatial and / or temporal prediction has been performed. The intra / inter decision can be indicated, for example, by a prediction mode flag. The prediction block is subtracted from the current video block (216) to generate a prediction residual. The prediction residual is decorrelated using transform (218) and quantization (220). (In some blocks, the symbolizer may bypass both the transform and quantization, in which case the residual can be directly coded without applying the transform or quantization process.) The quantized residual coefficients are inverse quantized (222) and inverse transformed (224) to form a reconstructed residual, which is then fed back to the prediction block (226) to form the reconstructed signal of the CU. Further in-loop filtering, such as deblocking / SAO (sample adaptive offset) filtering, can be applied to the reconstructed CU (228) to reduce coding artifacts before it is placed in the reference picture store (212) and used to code future video blocks. To form the output video bitstream 230, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit (108), where they are further compressed and packed to form the bitstream.

[0066] FIG. 2B gives a block diagram of a block-based video decoder 250. In the decoder 250, as described below, the bitstream is decoded by decoder elements. The video decoder 250 generally performs a decoding path that is the reverse of the encoding path as described in FIG. 2A. The encoder 200 also generally performs video decoding as part of the encoding of video data.

[0067] In particular, the input to the decoder includes a video bitstream 252 that can be generated by the video encoder 200. The video bitstream 252 is first unpacked and entropy decoded by the entropy decoding unit 254 to obtain transform coefficients, motion vectors, and other coding information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (256). The coding mode and prediction information are sent either to the spatial prediction unit 258 (when intra-coded) or to the temporal prediction unit 260 (when inter-coded) to form a prediction block. The residual transform coefficients are sent to the inverse quantization unit 262 and the inverse transform unit 264 to reconstruct the residual block. Next, the prediction block and the residual block are added together at 266 to generate a reconstructed block. The reconstructed block can further pass through in-loop filtering 268 before being stored in the reference picture store 270 for use in predicting future video blocks.

[0068] The decoded picture 272 can further undergo post-decoding processing (274), such as inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (204). In the post-decoding processing, metadata derived in the pre-encoding processing and signaled in the bitstream can be used. The decoded and processed video can be sent to the display device 276. The display device 276 can be a device separate from the decoder 250, or the decoder 250 and the display device 276 can be components of the same device.

[0069] Using the various methods and other aspects described in this disclosure, the modules of the video encoder 200 or decoder 250 can be modified. Further, the systems and methods disclosed herein are not limited to VVC or HEVC, and can be applied to other standards and recommendations, as well as extended versions of any such standards and recommendations (including VVC and HEVC), whether existing or to be developed in the future. Unless specifically indicated or technically excluded, the aspects described in this disclosure can be used individually or in combination.

[0070] Inter prediction. Figures 3A and 3B are diagrams showing examples of motion prediction of video blocks (e.g., using the inter prediction module 210 or 260). Figure 3B, which shows an example of block-level motion within an image, shows an exemplary decoded picture buffer including reference pictures "Ref pic 0", "Ref pic 1", and "Ref pic 2". Blocks B0, B1, and B2 of the current picture can be predicted from blocks of reference pictures "Ref pic 0", "Ref pic 1", and "Ref pic 2", respectively. Motion prediction can use video blocks from adjacent video frames to predict the current video block. Motion prediction can utilize temporal correlation and / or remove temporal redundancy inherent in the video signal. For example, in H.264 / AVC and HEVC, temporal prediction can be performed on video blocks of various sizes (e.g., for the luminance component, the temporal prediction block size can vary from 16×16 to 4×4 in H.264 / AVC and from 64×64 to 4×4 in HEVC). According to the motion vector (mvx, mvy), temporal prediction can be performed as provided by the following equation. P(x,y)=ref(x-mvx,y-mvy) Here, ref(x, y) can be the pixel value at position (x, y) in the reference image, and P(x, y) can be the predicted block. The video coding system may support inter prediction with fractional pixel accuracy. When the motion vectors (mvx, mvy) have fractional pixel values, one or more interpolation filters may be applied to obtain pixel values at fractional pixel positions. A block-based video coding system may use multi-hypothesis prediction to improve temporal prediction. For example, the predicted signal may be formed by combining a number of predicted signals from different reference images. For example, H.264 / AVC and / or HEVC may use dual prediction that can combine two predicted signals. Dual prediction may combine two predicted signals from reference images respectively to form a prediction as in the following equation.

Equation

[0071] Affine mode In HEVC, only the translational motion model is applied to motion-compensated prediction. In the real world, there are various types of motions such as zoom-in / zoom-out, rotation, perspective motion, and other irregular motions. In VTM-2.0, affine motion-compensated prediction is applied. The affine motion model is either a 4-parameter or 6-parameter one. The first flag of each mutually coded block is signaled to indicate which of the translational motion model or the affine motion model is applied to inter prediction. In the case of the affine motion model, a second flag is transmitted to indicate which of the 4-parameter model and the 6-parameter model is being used.

[0072] The affine motion model with four parameters may have parameters including two parameters for translational motion in the horizontal and vertical directions, one parameter for zoom motion in both directions, and one parameter for rotational motion in both directions. Since the horizontal zoom parameter is equal to the vertical zoom parameter, a single zoom parameter is used. Since the horizontal rotation parameter is equal to the vertical rotation parameter, a single rotation parameter is used. The 4-parameter affine motion model is coded in VTM using two motion vectors at two control point positions defined at the upper left and upper right corners of the current block. As shown in Figure 4A, the affine motion field of the block is described by two control point motion vectors (V0, V1). Based on the motion of the control points, the motion field (v x , v y ) of the affine-coded block is described as follows. [Equation]

[0073] In Equation (1), (v 0x , v 0y ) is the motion vector of the control point at the upper left corner as shown in Figure 4A, and (v 1x , v 1y) is the motion vector of the control point in the upper right corner, and w is the width of the block. In VTM-2.0, the motion field of the affine-coded block is derived at the 4×4 sub-block level, and for each 4×4 sub-block (Figure 4B) within the current block, (v x , v y ) is derived and applied to the corresponding 4×4 sub-block. (Note that a set of samples that are referred to as a block in some contexts may be referred to as a sub-block in other contexts. For clarity, various terms may be used in context.)

[0074] These four parameters of the four-parameter affine model can be estimated iteratively. Let the MV pair at step k be

Number

Number

[0075] In Equation (2), at step k, (a, b) is the delta translation parameter, and (c, d) is the delta zoom and rotation parameter. The delta MV at the control point can be derived using its coordinates as shown in Equations (3) and (4). For example, (0, 0), (w, 0) are the coordinates of the upper left and upper right control points respectively.

Number

Number

[0076] Based on the optical flow equation, the relationship between the change in luminance and the spatial gradient and temporal motion is formulated as follows.

Equation

[0077] In Equation (2)

Equation

[0078] Since all samples within the block satisfy Equation (6), the parameter set (a, b, c, d) can be solved using the least squares error method. The motion vectors (MVs) at two control points in step (k + 1)

Equation

[0079] An affine motion model with six parameters can have two parameters for translational motion in the horizontal and vertical directions, one parameter for zoom motion, one parameter for rotational motion in the horizontal direction, one parameter for zoom motion, and one parameter for rotational motion in the vertical direction. The six-parameter affine motion model is encoded using three MVs at three control points. As shown in the example of FIG. 5, the three control points of the six-parameter affine coding block are defined at the upper left corner, upper right corner, and lower left corner of the block. The motion at the upper left control point is related to translational motion, the motion at the upper right control point is related to horizontal rotation and zoom motion, and the motion at the lower left control point is related to rotation and vertical zoom motion. In the case of a six-parameter affine motion model, the horizontal rotation and zoom motions may not be the same as the vertical motion. The motion vectors of each sub-block (v x ,v y ) can be derived using the three MVs at the control points as follows.

Equation

[0080] In Equation (7), (v 2x ,v 2y ) is the motion vector of the lower left control point, (x, y) is the center position of the sub-block, and w and h are the width and height of the block.

[0081] The six parameters of the six-parameter affine model can be estimated in the same way as used in the four-parameter model. Equation (2) can be modified to Equation (8) as follows.

Equation

[0082] In Equation (8), at step k, (a, b) are the delta translation parameters, (c, d) are the horizontal delta zoom and rotation parameters, and (e, f) are the vertical delta zoom and rotation parameters. Equation (6) is changed accordingly to obtain Equation (9). I’ k (i, j) - I(i, j) = (g x (i, j) * i) * c + g x (i, j) * j) * d + g y (i, j) * i) * e + (g y (i, j) * j) * f + g x (i, j) * a + g y (i, j) * b (9)

[0083] The parameter set (a, b, c, d, e, f) can be solved using the least squares method by considering all samples within the block. The top-left control point

Number

Number

Number

Number

Number

[0084] Refinement of Prediction by Affine Mode Optical Flow (PROF) To achieve a finer granularity of motion compensation, a method for refining sub-block-based affine motion-compensated prediction by optical flow has been proposed as described in Jiancong (Daniel) Luo, Yuwen He, "CE2-related: Prediction refinement with optical flow for affine mode", JVET-N0236, March 2019, Geneva, Switzerland. After sub-block-based affine motion compensation is performed, each luminance prediction sample is refined by adding the difference derived from the optical flow equation. The proposed PROF is described as including the following steps.

[0085] In the first step, sub-block-based affine motion compensation is performed to generate a sub-block prediction I(i,j).

[0086] In the second step, the spatial gradients g x (i,j) and g y (i,j) of the sub-block prediction are calculated at each sample position using a 3-tap filter [-1,0,1]. g x (i,j)=I(i + 1,j)-I(i - 1,j) g y (i,j)=I(i,j + 1)-I(i,j - 1) The sub-block prediction is extended by one pixel on each side for gradient calculation. To reduce memory bandwidth and complexity, the pixels on the extended boundary are copied from the nearest integer pixel position of the reference image. Thus, additional interpolation in the padding area is avoided.

[0087] In the third step, the refinement of the luminance prediction is calculated by the optical flow equation. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (12) Here, Δv(i,j) is the difference between the pixel MV calculated for the sample position (i,j) represented by b v(i,j)y and the sub-block MV of the sub-block to which the pixel (i,j) belongs, as shown in FIG. 6. Since the parameters of the affine model and the pixel positions relative to the centers of the sub-blocks do not change for each sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks within the same block. Let x and y be the horizontal and vertical offsets from the pixel position to the center of the sub-block, then Δv(x,y) can be derived from the following equation.

Equation

Equation

Equation

[0088] In the fourth step, the refinement of the luminance prediction is added to the sub-block prediction I(i,j). The final prediction I’ can be generated using the following equation. I’(i,j)=I(i,j)+ΔI(i,j) (13)

[0089] Problems Addressed in Some Embodiments The current motion compensation process is typically limited by the accuracy of the motion vectors. For example, the motion vectors indicated on the decoder side are used for motion-compensated prediction at the sample level, sub-block level, or block level, whereby an integer reference sample position and an interpolation filter at the fractional position can be determined. The accuracy of the associated motion vectors is a factor in the accuracy of the motion-compensated prediction at each sample. Using four additional bits for the fractional part of the motion vector can achieve a 1 / 16 PEL accuracy. However, this accuracy limitation causes potential problems. One problem is that when the motion-compensated prediction is already accurate enough, the extra bits become a waste of signaling overhead. Another problem is that in some cases, the number of additional bits provided may still not be sufficient, and higher accuracy may be desired. To provide a more efficient representation of accuracy, a more flexible and accurate method may be beneficial for improving the motion-compensated accuracy.

[0090] In VTM-5.0, when a specific accuracy granularity is required, the use of arbitrary accuracy is hindered because the interpolation filters corresponding to accuracies higher than that are predefined. For example, when the decoder achieves an accuracy of 1 / 32 PEL, which is finer than the 1 / 16 PEL accuracy of VTM-5.0, the interpolation filter with an accuracy higher than 1 / 32 PEL (e.g., 1 / 64 PEL) can be a predefined interpolation filter (e.g., the filter defined in the video standard specification).

[0091] Overview of some embodiments. The present disclosure describes a system and method for improving the prediction accuracy of a motion compensation (MC) prediction process using optical flow in a flexible manner for use with precision refinement. In some embodiments, after motion compensation is performed, the prediction at each sample is refined by adding the difference value derived from the optical flow equation. Such refinement may be referred to as motion compensation precision refinement by optical flow (MCPROF). The optical flow may be signaled as a refinement of the motion vector at the block level (which may be at the prediction unit level such as the CU level or sub-CU level). Since the difference values derived from the optical flow equation may represent different accuracies, a more refined accuracy may be achieved. Some embodiments described herein can achieve pixel-level granularity without significantly increasing complexity and can keep the worst-case memory access bandwidth the same as normal block-level motion compensation. The various embodiments described herein may be applied to any sub-block-based inter prediction mode and / or block-based inter prediction mode. The embodiments described herein can be applied to both single prediction and double prediction, and they can be applied to both interlaced mode and non-interlaced mode. One potential advantage of some embodiments is to provide precision refinement without requiring an additional interpolation filter.

[0092] Motion compensation precision refinement by optical flow (MCPROF) To achieve a more refined accuracy of motion compensation, in some embodiments, a method for refining the motion-compensated prediction using optical flow is adopted. After the motion compensation process is performed, the prediction of luminance and / or chrominance at each sample is refined by adding the difference derived from the optical flow equation. Examples of encoding methods using MCPROF are as follows.

[0093] In an exemplary method, after motion estimation in inter non-merge mode, a prediction I(i,j) is generated at each sample position (i,j) using a motion compensation process. The motion compensation process may be performed using existing inter prediction processes (including single prediction, dual prediction, and affine prediction). There may be one or more motion compensation processes performed at this step (e.g., multiple motion vector candidates are available). In the case of multiple motion compensation processes, as done in VTM5.0, one of the multiple motion compensation processes may be selected according to a predefined criterion (e.g., the motion compensation process with the minimum rate-distortion cost).

[0094] One or the selected motion compensation process is evaluated to determine whether to use accuracy refinement. For example, if the existing accuracy provided after the MC process (e.g., 1 / 4PEL) is accurate enough, a decision may be made not to proceed with accuracy refinement. Such a decision may be made, for example, when the residual value is close to 0 at the current motion compensation accuracy (e.g., less than a threshold of a predetermined magnitude).

[0095] However, depending on the situation, a decision may be made to proceed with accuracy refinement. In this case, an accuracy difference N used to convey the degree of refinement may be determined. For example, a determination that the substantially optimal (or accurate) residual value is about 9 / 16PEL may be made through the motion compensation process. To correspond to the optimal accuracy, a motion compensation accuracy of 1 / 16PEL may be desirable. However, when using an accuracy of 1 / 4PEL in the current motion compensation process (e.g., provided by an existing predefined interpolation filter), the accuracy difference N between the current accuracy and the desired accuracy is 2.

[0096] In some embodiments, the refinement of the motion-compensated prediction is determined using optical flow as follows. Let the existing motion vector(s) be MCP(i,j), the uncompressed original input sample value be O(i,j), the horizontal spatial gradient of MCP(i,j) be g x (i), and the vertical spatial gradient of MCP(i,j) be g y(i,j) shows the motion-compensated prediction. Refinement of additional motion vectors (Δmv x , Δmv y ) is used for the optical flow. The refinement of the motion vectors (Δmv x , Δmv y ) can be selected to substantially satisfy Equation 14. O(i,j) = MCP(i,j) + g x (i,j) * Δmv x + g y (i,j) * Δmv y (14)

[0097] In some embodiments, the refinement of the motion vectors (Δmv x , Δmv y ) is estimated by the least squares method as in Equation (15).

Equation

[0098] If it is determined to use a more refined accuracy (e.g., with respect to bit depth), the refined motion vector value Δmv(i,j) (i.e., (Δmv x (i,j), Δmv y (i,j))) can be signaled in the bitstream. In some embodiments, the associated accuracy difference N can also be signaled in the bitstream. To save signaling overhead, N can be signaled at different levels such as slice / image level, CTU level, or CU (or other block) level. Similarly, the refined motion vector values can be signaled at different levels such as slice / image level, CTU level, CU (or other block) level, or sample level. If the refined motion vector values are not signaled at the sample level, the refined motion vector values within a block or sub-block can be the same for each sample within that block or sub-block.

[0099] In some embodiments, the additional motion vector refinement value may be signaled in the form of specific adjacent sample positions. As shown in FIG. 7, the adjacent positions can be used to indicate the refinement value of the motion vector, which can be represented by one of the four closest adjacent positions or eight closest adjacent positions (e.g., 1 pixel distance) or even more adjacent positions (e.g., beyond 1 pixel distance). In the example shown in FIG. 7, when the adjacent sample position (i, j - 1) is signaled, it has the effect of signaling the refinement value of the motion vector Δmv(i, j) = (0, -1), where Δmv x (i, j) = 0, Δmv y (i, j) = -1. When the adjacent sample position (i - 1, j - 1) is signaled, it has the effect of signaling the refinement value of the motion vector Δmv(i, j) = (-1, -1), where Δmv x (i, j) = -1, Δmv y (i, j) = -1.

[0100] In some embodiments, the refinement value of the motion vector may be signaled in the form of an index value. For example, when the four closest adjacent positions are used, the corresponding upper, lower, left, and right adjacent positions are indexed as 0, 1, 2, 3 as an example. The index can be binary coded with a variable-length codeword. For example, if eight closest adjacent positions are permitted, the index of the four closest adjacent positions can be coded using a shorter codeword than the index of the other four adjacent positions in the two diagonal directions.

[0101] In some embodiments, the associated accuracy difference N can be an integer value equal to the bit depth difference between the current accuracy and the desired accuracy. For example, if the current accuracy of the motion compensated prediction process is 1 / 4 PEL and the required accuracy is 1 / 16 PEL, the signaled accuracy can be N = 2. In some embodiments, the encoder signals a flag indicating whether MCPROF is being used.

[0102] Prediction using optical flow for accuracy refinement. The accuracy difference N and the refinement value of the motion vector can be used by an encoder (e.g., at module 210) or a decoder (e.g., at module 260) in generating a prediction of a block or sub-block. When MCPROF is used (e.g., when an enable flag is signaled in the bitstream), the signaled values of Δmv(i,j) and N are obtained. The spatial gradients g x (i,j) and g y (i,j) are calculated for each sample position (i,j). The determination of the spatial gradient can be performed using a 3-tap filter as described above, or using other techniques.

[0103] An initial motion-compensated prediction I(i,j) is generated for the current block, for example, using single prediction, dual prediction, and / or affine prediction. The accuracy refinement is calculated using the scalar product of the refinement of the motion vector and the spatial gradient according to Equation 16. ΔI(i,j)=(gx(i,j)*Δmv x (i,j)+g y (i,j)*Δmv y (i,j))≫(N+existing MC precision) (16) Here, Δmv(i,j) is the refinement value of the signaled motion vector (e.g., received from the encoder), N is the difference in the signaled bit depth between the current accuracy and the desired accuracy, and g(i,j) is the spatial gradient calculated as described above.

[0104] The motion-compensated prediction at each sample is refined by adding a change in intensity (e.g., luminance or chrominance). The final prediction I’ can be generated according to the following equation. I’(i,j)=I(i,j)+ΔI(i,j) (17)

[0105] In the above exemplary embodiments, four adjacent positions or eight adjacent positions are considered. Each selected adjacent position is used to indicate the direction of improvement in accuracy and does not depend on waiting for the motion compensation process of adjacent samples to be completed.

[0106] An example of the method is shown in FIG. 8. In a video coder, based on motion-compensated prediction, an initial predicted sample value is obtained (802) for at least the first sample position within the current block of samples. For at least the first sample position, refinement of the motion vector is determined (804). The refinement of the motion vector can be encoded in the bitstream (806), for example, for storage or transmission.

[0107] The spatial gradient of the sample values is determined at the first sample position (808). The sample difference value is determined based on the scalar product of the spatial gradient and the refinement of the motion vector (810). In some embodiments, the determination of the sample difference value may include scaling (e.g., bit shifting) of the sample difference value, and the accuracy information indicating the amount of scaling can be encoded in the bitstream. The initial predicted sample value is corrected based on the sample difference value (812), for example, by adding the sample difference value to the initial predicted sample value to generate a refined sample value.

[0108] In some embodiments, the determination of the refinement of the motion vector (804) may include selecting the refinement of the motion vector to substantially minimize the prediction error for the input video block. In embodiments where the refinement of the motion vector is selected for each sample, the prediction error can be based on the difference (e.g., absolute difference or squared difference) between the refined sample value and the corresponding sample value of the input video block. In embodiments where the refinement of the motion vector is selected based on each block (or sub-block), the prediction error can be based on the sum of the differences (e.g., sum of absolute differences or squared differences) between the refined sample values and the sample values of the corresponding input video block over a plurality of sample positions within the block (or sub-block).

[0109] In some embodiments, the encoder may use the refined sample values generated at 812 when determining the prediction residual, and the prediction residual may also be encoded into the bitstream.

[0110] In a method performed by a video decoder, the decoder obtains (814) an initial predicted sample value for at least a first sample position within a current block of samples based on motion-compensated prediction. Refinement of the motion vector is determined (816) for at least the first sample position, for example, by decoding the refinement of the motion vector from the bitstream. The spatial gradient of the sample values is determined (818) at the first sample position. The sample difference value is determined (820) by calculating the scalar product of the spatial gradient and the refinement of the motion vector. In some embodiments, the determination of the sample difference value may further include scaling the sample difference value based on accuracy information decoded from the bitstream. The initial predicted sample value is corrected (822) based on the sample difference value. For example, the sample difference value may be added to the initial predicted sample value to generate a refined sample value.

[0111] In some embodiments, the decoder may further decode a prediction residual from the bitstream, and the prediction residual may be used when determining the sample values reconstructed at the first sample position. The reconstructed sample values may be displayed or communicated to another display device for display.

[0112] Exemplary communication system. FIG. 9 is a diagram showing an example of a communication system. The communication system 900 may include an encoder 902, a communication network 904, and a decoder 906. The encoder 902 may communicate with the network 904 via a connection 908 that may be a wired or wireless connection. The encoder 902 may be similar to the block-based video encoder of FIG. 2A. The encoder 902 may include a single-layer codec or a multi-layer codec. The decoder 906 may communicate with the network 904 via a connection 910 that may be a wired or wireless connection. The decoder 906 may be similar to the block-based video decoder of FIG. 2B. The decoder 906 may include a single-layer codec or a multi-layer codec.

[0113] The encoder 902 and / or the decoder 906 may be incorporated into a variety of wired communication devices and / or wireless transmit / receive units (WTRUs) such as, but not limited to, digital televisions, wireless broadcast systems, network elements / terminals, servers such as content or web servers (such as hypertext transfer protocol (HTTP) servers), personal digital assistants (PDAs), laptop or desktop computers, tablet computers, digital cameras, digital recording devices, video game devices, video game consoles, cellular phones or satellite radiotelephones, digital media players.

[0114] Communication network 904 can be a suitable type of communication network. For example, communication network 904 can be a multi-access system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. Communication network 904 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, communication network 904 can adopt one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), etc. Communication network 904 can include a plurality of connected communication networks. Communication network 904 can include the Internet and / or one or more private commercial networks such as, for example, cellular networks, WiFi hotspots, Internet service provider (ISP) networks, etc.

[0115] Additional embodiments. According to some embodiments, a block-based video coding method includes, in a decoder, generating a motion-compensated prediction of sample values in a current block of samples, decoding, from a bitstream, a precision difference value and refinement of a motion vector for the current block, determining a spatial gradient at a sample for each predicted sample in the current block, calculating a scalar product of the spatial gradient and the refinement of the motion vector, scaling the scalar product by an amount indicated by the precision difference value to generate a sample difference value, and adding the sample difference value to the predicted sample value to generate a refined sample value.

[0116] In some embodiments, the motion-compensated prediction of sample values in the current block is generated with a single prediction.

[0117] In some embodiments, the motion-compensated prediction of sample values in the current block is generated using a dual prediction.

[0118] In some embodiments, the refinement of the motion vector is signaled in a bitstream as an index. The index can identify one of the refinements of a plurality of motion vectors in the form of (i,j), where i and j are integers. The index can identify one of the refinements of a plurality of motion vectors from the group consisting of (0, -1), (1, 0), (0, 1), and (-1, 0). The index can identify one of the refinements of a plurality of motion vectors from the group consisting of (0, -1), (1, 0), (0, 1), (-1, 0), (-1, -1), (1, -1), (1, 1), and (-1, 1).

[0119] In some embodiments, the scaling of the scalar product includes a bit shift of the scalar product.

[0120] In some embodiments, the accuracy difference value is N, and the scaling of the scalar product includes right-shifting the scalar product by a number of bits equal to the sum of the signaled accuracy difference value N and the existing MC accuracy.

[0121] A block-based video coding method according to some embodiments includes, in an encoder, generating a motion-compensated prediction of sample values at a current block of samples for an input video block, selecting an accuracy difference value, determining respective spatial gradients at the samples, and determining a refinement of a motion vector for the current block, the refinement of the motion vector being selected to substantially minimize an error between (i) a scalar product of the spatial gradient and the refinement of the motion vector and (ii) a difference between the input video block and the motion-compensated prediction, and signaling, in a bitstream, the accuracy difference value and the refinement of the motion vector for the current block.

[0122] In some embodiments, the motion-compensated prediction of sample values within the current block is generated with a single prediction.

[0123] In some embodiments, the motion-compensated prediction of the sample values within a current block is generated using dual prediction.

[0124] In some embodiments, the refinement of the motion vector is signaled in a bitstream as an index. The index can identify one of the refinements of a plurality of motion vectors in the form of (i,j), where i and j are integers. The index can identify one of the refinements of a plurality of motion vectors from the group consisting of (0, -1), (1, 0), (0, 1), and (-1, 0). The index can identify one of the refinements of a plurality of motion vectors from the group consisting of (0, -1), (1, 0), (0, 1), (-1, 0), (-1, -1), (1, -1), (1, 1), and (-1, 1).

[0125] In some embodiments, the refinement of the motion vector is selected to substantially minimize the sum of the squared differences between (i) the scalar product of the spatial gradient and the refinement of the motion vector and (ii) the difference between the input video block and the motion-compensated prediction.

[0126] Some embodiments include a processor and a non-transitory computer-readable medium that operate to perform any of the functions described herein.

[0127] This disclosure describes a variety of aspects including tools, features, embodiments, models, approaches, etc. Many of these aspects are specifically described and may be described in a way that may seem limiting, at least to indicate individual characteristics. However, this is for clarity of explanation and does not limit the disclosure or scope of these aspects. In fact, all different aspects can be combined and exchanged to provide further aspects. Additionally, these aspects can be combined and exchanged with aspects described in previous applications.

[0128] Aspects described and contemplated in this application can be implemented in many different formats. Although some embodiments are specifically shown, other embodiments are contemplated, and consideration of a particular embodiment is not intended to limit the scope of implementation. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.

[0129] In this disclosure, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably.

[0130] Various methods are described herein, and each of those methods includes one or more steps or acts for achieving the described method. The order and / or use of particular steps and / or acts can be modified or combined, provided that a particular order of steps or acts is not required for the correct operation of the method. Further, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, acts, etc., such as "first decoding" and "second decoding". The use of such terms does not, except where particularly required, imply an order of modified acts. Thus, in this example, the first decoding need not be performed before the second decoding and can occur, for example, before, during, or overlapping a period with the second decoding.

[0131] For example, in the present disclosure, various numerical values can be used. The specific values are for illustrative purposes, and the described embodiments are not limited to these specific values.

[0132] The embodiments described herein can be executed by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The processor can be of any type suitable for the technical environment and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.

[0133] Various implementations involve decoding. The "decoding" used in this application can include, for example, all or part of a process executed on a received encoded sequence to generate a final output suitable for display. In various embodiments, such a process can include one or more of the processes typically executed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such a process can also or alternatively include processes executed by the decoders of the various implementations described in this application, such as extracting an image from a tiled (packed) image, determining an upsample filter to use, then upsampling the image, and flipping the image back to the intended orientation.

[0134] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will become clear based on the context of the specific description and is considered to be fully understood by those skilled in the art.

[0135] The various implementations involve encoding. Similar to the above considerations regarding "decoding", "encoding" as used in this application can include all or part of the processes performed on the input video sequence, for example, to generate an encoded bitstream. In various embodiments, such processes typically include one or more processes performed by an encoder, such as splitting, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes can also, or alternatively, include the processes performed by the encoders of the various implementations described in this application.

[0136] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will become clear based on the context of the particular description and is considered to be fully understood by those skilled in the art.

[0137] It should be understood that when a figure is presented as a flowchart, a block diagram of the corresponding apparatus is also provided. Similarly, when a figure is presented as a block diagram, a flowchart of the corresponding method / process is also provided.

[0138] Various embodiments refer to rate distortion optimization. In particular, during the encoding process, often considering computational complexity constraints, a balance or trade-off between rate and distortion is typically considered. Rate distortion optimization is usually formulated to minimize a rate distortion function that is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, an approach can be based on an extensive test of all encoding options, including all modes or coded parameter values being considered, involving a complete evaluation of the encoding cost and the associated distortion of the reconstructed signal after encoding and decoding. In particular, a faster approach can also be used to reduce the encoding complexity by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. These two approaches can also be used in combination, such as by using the approximate distortion for only some of the possible encoding options and the complete distortion for other encoding options. In other approaches, only a subset of the possible encoding options is evaluated. More generally, many approaches employ any of various techniques for performing the optimization, but the optimization does not necessarily involve a complete evaluation of both the encoding cost and the associated distortion.

[0139] The implementation forms and modes described in this specification can be implemented, for example, in a method or process, a device, a software program, a data stream, or a signal. Even if considered only in the context of a single implementation form (e.g., considered only as a method), the implementation form of the considered features can also be implemented in other forms (e.g., a device or a program). The device can be implemented, for example, with appropriate hardware, software, and firmware. Those methods can be implemented, for example, within a processor, which refers to generally processing devices including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a mobile phone, a portable / personal digital assistant ( "PDA"), and other devices that facilitate the communication of information between end users.

[0140] References to "one embodiment" or "an embodiment", or "one implementation form" or "an implementation form", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in relation to the embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" or "in one implementation form" or "in an implementation form", and the appearance of any other variations, which can be seen at various places throughout this application, do not necessarily all refer to the same embodiment.

[0141] Furthermore, this disclosure may refer to "determining" various parts of information. The determination of information may include, for example, one or more of the evaluation of information, the calculation of information, the prediction of information, or the retrieval of information from memory.

[0142] Furthermore, this application may refer to "accessing" various parts of information. The access to information may include, for example, one or more of the reception of information, the retrieval of information (e.g., from memory), the storage of information, the transfer of information, the copying of information, the calculation of information, the determination of information, the prediction of information, or the evaluation of information.

[0143] Furthermore, this application may refer to "receiving" various parts of information. Receiving is intended to be a broad term, similar to "accessing". Receiving information may include, for example, one or more of accessing the information or retrieving the information (e.g., from memory). Further, "receiving" typically involves, in some way, operations such as storing the information, processing the information, transmitting the information, moving the information, copying the information, deleting the information, calculating the information, determining the information, predicting the information, or evaluating the information.

[0144] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that any use of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). This can be extended to many listed items.

[0145] Also, as used herein, the word "signal" (when used as a verb) refers, among other things, to instructing something to a corresponding decoder. For example, in certain embodiments, an encoder signals a particular one of a plurality of parameters for refinement. In this way, in an embodiment, the same parameter is used on both the encoder side and the decoder side. Thus, for example, an encoder can send a particular parameter to a decoder (explicit signaling), and as a result, the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling can be used without performing a transmission (implicit signaling) to enable the decoder to easily recognize and select the particular parameter. By avoiding any actual function transmission, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to a corresponding decoder. The above relates to the verb form of the word "signal", but the word "signaling" can also be used as a noun herein.

[0146] Embodiments can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described embodiments. For example, a signal can be formatted to carry a bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0147] Some embodiments will be described. The features of these embodiments can be provided alone or in any combination across various claim categories and types. Further, the embodiments can include one or more of the following features, devices, or aspects alone or in any combination across various claim categories and types. · A bitstream or signal including a syntax for transmitting information generated according to any of the described embodiments. · Creating and / or transmitting and / or receiving and / or decrypting according to any of the described embodiments. · A method, process, device, medium for storing instructions, medium for storing data, or signal according to any of the described embodiments. · A TV, set-top box, mobile phone, tablet, or other electronic device that executes an encoding or decoding method according to any of the described embodiments. · A TV, set-top box, mobile phone, tablet, or other electronic device that executes a decoding method according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). · A TV, set-top box, mobile phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal including an encoded image according to any of the described embodiments and performs decoding. · A TV, set-top box, mobile phone, tablet, or other electronic device that wirelessly receives (e.g., using an antenna) a signal including an encoded image according to any of the described embodiments and performs decoding.

[0148] One or more of the various hardware elements of the described embodiments may be referred to herein as "modules" that perform (i.e., carry out, implement, etc.) the various functions described herein in connection with each module. As used herein, a module may include hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) that is considered suitable by those of ordinary skill in the relevant art for a given implementation. Each module described may also include executable instructions for performing one or more functions described as being performed by each module, and these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or may be stored in any suitable non-transitory computer-readable medium (s) generally referred to as RAM, ROM, etc.

[0149] Although features and elements are described above in particular combinations, one of ordinary skill in the art will understand that each feature or element may be used alone or in any combination with other features and elements. Further, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video encoding method comprising the steps of: obtaining an initial predicted sample value for each of a plurality of sample positions in a current block or sub-block of samples based on a motion compensated prediction; determining a refinement of a motion vector, the refinement of the motion vector having the same value for all of the plurality of sample positions in the current block or sub-block of samples; determining a spatial gradient of sample values ​​at each of the plurality of sample locations; determining a sample difference value for each of the plurality of sample locations based on a scalar product of the spatial gradient and a refinement of the motion vector; modifying the initial predicted sample values ​​based on the sample difference values; A method comprising:

2. The method of claim 1, wherein determining a refinement of the motion vector comprises selecting a refinement of the motion vector to substantially minimize a prediction error for the input video block; The method of claim 1 , further comprising signaling the refinement of the motion vectors in a bitstream.

3. The method of claim 2, wherein determining a refinement of the motion vector comprises selecting a refinement of the motion vector to substantially minimize a prediction error for the input video block; The method further comprises signaling the refinement of the motion vector in a bitstream; The method of claim 1 , wherein the motion vector refinement is signaled in the bitstream as an index.

4. A video decoding method, comprising: obtaining an initial predicted sample value for each of a plurality of sample positions in a current block or sub-block of samples based on a motion compensated prediction; determining a refinement of a motion vector, the refinement of the motion vector having the same value for all of the plurality of sample positions in the current block or sub-block of samples; determining a spatial gradient of sample values ​​at each of the plurality of sample locations; determining a sample difference value for each of the plurality of sample locations based on a scalar product of the spatial gradient and a refinement of the motion vector; modifying the initial predicted sample values ​​based on the sample difference values; A method comprising:

5. The method of claim 4, wherein determining the refinement of the motion vector includes decoding the refinement of the motion vector from a bitstream.

6. The method of claim 4, wherein determining the refinement of the motion vector includes decoding the refinement of the motion vector from a bitstream, the refinement of the motion vector being signaled in the bitstream as an index.

7. The method of claim 4, further comprising decoding refinement precision information from the bitstream, and determining the sample difference value comprises scaling the scalar product by an amount indicated by the precision information.

8. The method of claim 4, further comprising decoding refinement precision information from the bitstream, and determining the sample difference value comprises bit-shifting the scalar product by an amount indicated by the precision information.

9. The method of claim 4, further comprising decoding refinement precision information from the bitstream, wherein the motion compensated prediction is performed using at least one motion vector having an initial precision, the refinement precision information comprises a precision difference value representing a difference between the initial precision and the refinement precision, and determining the sample difference value comprises scaling a scalar product by an amount indicated by the precision information.

10. A video encoding device comprising: obtaining an initial predicted sample value for each of a plurality of sample positions in a current block or sub-block of samples based on a motion compensated prediction; determining a refinement of a motion vector, the refinement of the motion vector having the same value for all of the plurality of sample positions in the current block or sub-block of samples; determining a spatial gradient of sample values ​​at each of the plurality of sample locations; determining a sample difference value for each of the plurality of sample locations based on a scalar product of the spatial gradient and a refinement of the motion vector; modifying the initial predicted sample values ​​based on the sample difference values; 1. A video encoding apparatus comprising: a processor configured to:

11. The method of claim 1, wherein determining a refinement of the motion vector comprises selecting a refinement of the motion vector to substantially minimize a prediction error for the input video block; The video encoding apparatus of claim 10, further comprising signaling the motion vector refinement in a bitstream.

12. The method of claim 11, wherein determining a refinement of the motion vector comprises selecting a refinement of the motion vector to substantially minimize a prediction error for the input video block; signaling the motion vector refinement in a bitstream; The video encoding apparatus of claim 10 , wherein the motion vector refinement is signaled in the bitstream as an index.

13. The video encoding apparatus of claim 10, further comprising decoding refinement precision information from the bitstream, wherein the motion compensated prediction is performed using at least one motion vector having an initial precision, the refinement precision information comprising a precision difference value representing a difference between the initial precision and the refinement precision, and determining the sample difference value comprising scaling the scalar product by an amount indicated by the precision information.

14. A video decoding apparatus comprising: obtaining an initial predicted sample value for each of a plurality of sample positions in a current block or sub-block of samples based on a motion compensated prediction; determining a refinement of a motion vector, the refinement of the motion vector having the same value for all of the plurality of sample positions in the current block or sub-block of samples; determining a spatial gradient of sample values ​​at each of the plurality of sample locations; determining a sample difference value for each of the plurality of sample locations based on a scalar product of the spatial gradient and a refinement of the motion vector; modifying the initial predicted sample values ​​based on the sample difference values; 16. A video decoding apparatus comprising: a processor configured to:

15. The video decoding apparatus of claim 14, wherein determining the refinement of the motion vector comprises decoding the refinement of the motion vector from a bitstream.

16. The video decoding apparatus of claim 14, wherein determining the refinement of the motion vector includes decoding the refinement of the motion vector from a bitstream, the refinement of the motion vector being signaled in the bitstream as an index.

17. The video decoding apparatus of claim 14, further comprising: decoding refinement precision information from the bitstream, and determining the sample difference value comprises scaling the scalar product by an amount indicated by the precision information.

18. The video decoding apparatus of claim 14, further comprising: decoding refinement precision information from the bitstream, and determining the sample difference value comprises bit-shifting the scalar product by an amount indicated by the precision information.

19. The video decoding apparatus of claim 14, further comprising: decoding refinement precision information from the bitstream, wherein the motion compensated prediction is performed using at least one motion vector having an initial precision, the refinement precision information comprising a precision difference value representing a difference between the initial precision and the refinement precision, and determining the sample difference value comprising scaling the scalar product by an amount indicated by the precision information.

20. A method for generating an initial predicted sample value for each of a plurality of sample positions in a current block or sub-block of samples, the method comprising: obtaining an initial predicted sample value based on a motion compensated prediction; determining a refinement of a motion vector, the refinement of the motion vector having the same value for all of the plurality of sample positions in the current block or sub-block of samples; determining a spatial gradient of sample values ​​at each of the plurality of sample locations; determining a sample difference value for each of the plurality of sample locations based on a scalar product of the spatial gradient and a refinement of the motion vector; modifying the initial predicted sample values ​​based on the sample difference values; A computer-readable storage medium containing instructions for causing one or more processors to execute the above.

Citation Information

Patent Citations

  • Method and apparatus for video coding

    WO2020191034A1