Symmetric Motion Vector Difference Coding
By determining if SMVD is used in motion vector encoding and bypassing BDOF accordingly, the complexity of video coding systems is reduced, enhancing encoding and decoding efficiency.
Patent Information
- Application Number
- JP2021535909
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-22
- Filing Date
- 2019-12-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2039-12-19
AI Technical Summary
Existing video coding systems face increased complexity in calculations during encoding and decoding due to techniques like bidirectional motion-compensated prediction, which exploit temporal correlation to remove redundancy.
The proposed solution involves determining whether symmetric motion vector difference (SMVD) is used in motion vector encoding for a current encoding block, and if so, bypassing bidirectional optical flow (BDOF) for that block. This is done by obtaining an SMVD indication and using it to decide whether to reconstruct the block without performing BDOF.
This approach reduces the computational complexity by avoiding the BDOF process when SMVD is used, thereby improving encoding and decoding efficiency without compromising video quality.
Smart Images

Figure 0007697882000054 
Figure 0007697882000055 
Figure 0007697882000056
Abstract
Description
Technical Field
[0001] The present invention relates to symmetric motion vector differential coding.
Background Art
[0002] Cross - reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 783,437, filed on December 21, 2018; U.S. Provisional Patent Application No. 62 / 787,321, filed on January 1, 2019; U.S. Provisional Patent Application No. 62 / 792,710, filed on January 15, 2019; U.S. Provisional Patent Application No. 62 / 798,674, filed on January 30, 2019; and U.S. Provisional Patent Application No. 62 / 809,308, filed on February 22, 2019, the contents of which are hereby incorporated by reference in their entirety.
[0003] Video coding systems are used to compress digital video signals, for example, reducing the storage and / or transmission bandwidth required for such signals. Video coding systems can include block - based, wavelet - based, and / or object - based systems. The system uses video coding techniques such as bidirectional motion - compensated prediction (MCP) that can remove temporal redundancy by exploiting the temporal correlation between pictures.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Such techniques can increase the complexity of the calculations performed during encoding and / or decoding.
Means for Solving the Problems
[0005] Based on whether symmetric motion vector difference (SMVD) is used in motion vector encoding for the current encoding block, bidirectional optical flow (BDOF) can be bypassed for the current encoding block.
[0006] An encoding device (e.g., an encoder or a decoder) can determine that BDOF is available (enabled). The encoding device can determine whether to bypass BDOF for the current encoding block based at least in part on an SMVD indication for the current encoding block. The encoding device can obtain an SMVD indication indicating whether SMVD is used in motion vector encoding for the current encoding block. If the SMVD indication indicates that SMVD is used in motion vector encoding for the current encoding block, the encoding device can bypass BDOF for the current encoding block. If the encoding device determines to bypass BDOF for the current encoding block, it can reconstruct the current encoding block without performing BDOF.
[0007] The motion vector difference (MVD) for the current encoding block can indicate the difference between the motion vector predictor (MVP) for the current encoding block and the motion vector (MV) for the current encoding block. The MVP for the current encoding block can be determined based on the MVs of spatially adjacent blocks of the current encoding block and / or temporally adjacent blocks of the current encoding block.
[0008] If the SMVD indication indicates that SMVD is used for motion vector coding for the current coding block, the coding device can receive first motion vector coding information associated with the first reference picture list. Based on the first motion vector coding information associated with the first reference picture list and based on the symmetry between the MVD associated with the first reference picture list and the MVD associated with the second reference picture list, the coding device can determine second motion vector coding information associated with the second reference picture list.
[0009] In an example, if the SMVD indication indicates that SMVD is used for motion vector coding for the current coding block, the coding device can parse the first MVD associated with the first reference picture list in the bitstream. Based on the first MVD and based on the symmetry between the first MVD and the second MVD, the coding device can determine the second MVD associated with the second reference picture list.
[0010] If the coding device determines not to bypass BDOF for the current coding block, the coding device can improve the motion vectors of (e.g., each) sub-block of the current coding block based at least in part on the gradient associated with the position in the current coding block.
[0011] The coding device can receive a sequence-level SMVD indication indicating whether SMVD is available (enabled) for a sequence of pictures. If SMVD is enabled for a sequence of pictures, the coding device can obtain the SMVD indication associated with the current coding block based on the sequence-level SMVD indication.
Brief Description of the Drawings
[0012]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
[0013] A more detailed understanding can be obtained from the following detailed description, given by way of example, in conjunction with the drawings attached hereto.
[0014] FIG. 1A illustrates an exemplary communication system 100 in which one or more of the disclosed embodiments can be implemented. The communication system 100 can be a multi - access system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 can utilize one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single - carrier FDMA (SC - FDMA), zero - tail unique word DFT - spread OFDM (ZT UW DTS - s OFDM), unique word OFDM (UW - OFDM), resource block filtered OFDM, and filter bank multi - carrier (FBMC).
[0015] As shown in FIG. 1A, the communication system 100 can include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it is understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, any of them, which may sometimes be referred to as a “station” and / or “STA”, the WTRUs 102a, 102b, 102c, 102d can be configured to transmit and / or receive wireless signals and can include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular telephones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), home appliances, as well as devices operating on commercial and / or industrial wireless networks. Any of the WTRUs 102a, 102b, 102c, 102d may alternatively be referred to as a UE.
[0016] The communication system 100 can also include base station 114a and / or base station 114b. Each of base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as CN106 / 115, Internet 110, and / or other network 112. By way of example, base stations 114a, 114b can be a base transceiver station (BTS), Node B, eNode B, home Node B, home eNode B, gNB, New Radio (NR) Node B, site controller, access point (AP), and wireless router, among others. Although base stations 114a, 114b are each depicted as a single element, it will be understood that base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0017] The base station 114a can be part of the RAN 104 / 113, which can also include other base stations and / or network elements (not shown) such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or the base station 114b can be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies can be in a licensed spectrum, an unlicensed spectrum, or a combination of a licensed spectrum and an unlicensed spectrum. The cell can provide coverage for wireless services in a specific geographical area that can be relatively fixed or can change over time. The cell can further be divided into cell sectors. For example, the cell associated with the base station 114a can be divided into three sectors. Thus, in one embodiment, the base station 114a can include three transceivers, e.g., one for each sector of the cell. In an embodiment, the base station 114a can utilize multiple-input multiple-output (MIMO) technology and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0018] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).
[0019] More specifically, as mentioned above, the communication system 100 can be a multi-connection system and can utilize one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a within the RAN 104 / 113 and the WTRUs 102a, 102b, 102c can establish the air interface 116 using wideband CDMA (WCDMA) and implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0020] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can establish the air interface 116 using Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) and implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA).
[0021] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can establish the air interface 116 using New Radio (NR) and implement radio technologies such as NR radio access.
[0022] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interfaces utilized by the WTRUs 102a, 102b, 102c can be characterized by transmissions from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).
[0023] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement wireless technologies such as IEEE 802.11 (e.g., Wireless Fidelity (WiFi)), IEEE 802.16 (e.g., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0024] The base station 114b in FIG. 1A can be, for example, a wireless router, a home node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a localized area such as an office, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by a drone), and a roadway. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in FIG. 1A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0025] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, 102d. The data can have various Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN 106 / 115 can provide call control, billing services, mobile location-based services, prepaid outgoing calls, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that utilize the same RAT or a different RAT as RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113 which may be utilizing NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0026] CN106 / 115 can also serve as a gateway for the WTRU102a, 102b, 102c, 102d to access the PSTN108, the Internet 110, and / or other networks 112. The PSTN108 can include a circuit-switched telephone network that provides basic telephone service (POTS). The Internet 110 can include a worldwide system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) within the TCP / IP Internet protocol suite. The network 112 can include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 can include another CN that is connected to one or more RANs that can utilize the same RAT or a different RAT as the RAN104 / 113.
[0027] Some or all of the WTRU102a, 102b, 102c, 102d within the communication system 100 can include a multi-mode function (e.g., the WTRU102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU102c shown in Figure 1A can be configured to communicate with a base station 114a that can utilize cellular-based wireless technology and to communicate with a base station 114b that can utilize IEEE802 wireless technology.
[0028] Figure 1B is a system diagram illustrating an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 can include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU 102 can include any sub-combination of the above elements while maintaining compliance with the embodiments.
[0029] The processor 118 can be, for example, a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors cooperating with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, and the transceiver 120 can be coupled to the transmit / receive element 122. Although Figure 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0030] The transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmitting / receiving element 122 can be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmitting / receiving element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0031] In FIG. 1B, the transmitting / receiving element 122 is depicted as a single element, but the WTRU 102 can include any number of transmitting / receiving elements 122. More specifically, the WTRU 102 can utilize MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0032] The transceiver 120 can be configured to modulate signals to be transmitted by the transmitting / receiving element 122 and demodulate signals received by the transmitting / receiving element 122. As mentioned above, the WTRU 102 can have a multimode function. Thus, the transceiver 120 can include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.
[0033] The processor 118 of the WTRU 102 can be coupled to the speaker / microphone 124, keypad 126, and / or display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. In addition, the processor 118 can obtain information from and store data in any type of suitable memory, such as the non-removable memory 130 and / or the removable memory 132. The non-removable memory 130 can include random access memory (RAM), read only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card, among others. In other embodiments, the processor 118 can obtain information from and store data in a memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).
[0034] The processor 118 can receive power from the power supply 134 and can be configured to distribute power to and / or control the power to other components within the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 can include one or more dry cells (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell, among others.
[0035] Processor 118 can also be coupled to a GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 can receive location information on the air interface 116 from a base station (e.g., base stations 114a, 114b), and / or can determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 can acquire location information using any suitable location determination method while maintaining consistency with the embodiments.
[0036] Processor 118 can further be coupled to other peripheral devices 138, which can include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 can include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a Frequency Modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, and an activity tracker, among others. The peripheral devices 138 can include one or more sensors, which can be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0037] The WTRU 102 can include a full-duplex radio in which some or all of the transmission and reception of signals (associated with certain subframes for both, for example, the UL (for transmission) and the downlink (for reception)) can be parallel and / or simultaneous. The full-duplex radio can include an interference management unit 139 to reduce and / or substantially eliminate self-interference, either via hardware (such as a choke) or via signal processing through a processor (such as a separate processor (not shown) or processor 118). In embodiments, the WTRU 102 can include a half-duplex radio for some or all of the transmission and reception of signals (associated with certain subframes for either, for example, the UL (for transmission) or the downlink (for reception)).
[0038] Figure 1C is a system diagram illustrating the RAN 104 and the CN 106, according to an embodiment. As mentioned above, the RAN 104 can communicate with the WTRU 102a, 102b, 102c over the air interface 116 using E-UTRA radio technology. The RAN 104 can also communicate with the CN 106.
[0039] The RAN 104 can include eNodeBs 160a, 160b, 160c, although it will be understood that the RAN 104 can include any number of eNodeBs while maintaining consistency with the embodiments. The eNodeBs 160a, 160b, 160c can each include one or more transceivers for communicating with the WTRU 102a, 102b, 102c over the air interface 116. In one embodiment, the eNodeBs 160a, 160b, 160c can implement MIMO technology. Thus, the eNodeB 160a, for example, can transmit wireless signals to and / or receive wireless signals from the WTRU 102a using multiple antennas.
[0040] Each of the eNodeBs 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL. As shown in Figure 1C, the eNodeBs 160a, 160b, and 160c can communicate with each other over the X2 interface.
[0041] CN106 shown in Figure 1C can include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Each of the above elements is depicted as part of CN106, but it will be understood that any of these elements can be owned and / or operated by an entity different from the CN operator.
[0042] The MME 162 can be connected to each of the eNodeBs 160a, 160b, and 160c within the RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, and selecting a specific serving gateway during the initial attach of the WTRUs 102a, 102b, 102c. The MME 162 can provide control plane functions for exchanges between the RAN 104 and other RANs (not shown) that utilize other radio technologies such as GSM and / or WCDMA.
[0043] SGW164 can be connected to each of the eNodeBs 160a, 160b, and 160c within RAN104 via the S1 interface. SGW164 can generally route and transfer user data packets to / from WTRUs 102a, 102b, and 102c. SGW164 can perform other functions such as anchoring the user plane during eNodeB handover, triggering paging when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.
[0044] SGW164 can be connected to PGW166, and PGW166 can provide access to a packet-switched network, such as the Internet 110, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and IP-enabled devices.
[0045] CN106 can facilitate communication with other networks. For example, CN106 can provide access to a circuit-switched network, such as PSTN 108, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and conventional fixed-line telephone communication devices. For example, CN106 can include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN106 and PSTN 108. In addition, CN106 can provide access to other networks 112 to WTRUs 102a, 102b, and 102c, and other networks 112 can include other wired and / or wireless networks owned and / or operated by other service providers.
[0046] In FIGS. 1A-1D, the WTRU is described as a wireless terminal, but in certain representative embodiments, it is contemplated that such a terminal can (e.g., temporarily or permanently) use a wired communication interface to a communication network.
[0047] In some representative embodiments, another network 112 can be a WLAN.
[0048] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have an access or interface to a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic to an STA originating from outside the BSS can arrive through the AP and be delivered to the STA. Traffic transmitted from an STA to a destination outside the BSS can be sent to the AP for delivery to each respective destination. Traffic between STAs within the BSS can be sent through the AP. For example, a source STA can send the traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered peer-to-peer traffic and / or sometimes be referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent (e.g., directly) between a source STA and a destination STA using direct link setup (DLS). In certain representative embodiments, the DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) can communicate directly with each other. Communication in IBSS mode is sometimes referred to herein as "ad hoc" mode communication.
[0049] When using the operation of the 802.11ac infrastructure mode or the operation of a similar mode, the AP can transmit beacons on a fixed channel such as the primary channel. The primary channel can be of a fixed width (e.g., 20 MHz bandwidth), or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In one representative embodiment, for example, in an 802.11 system, carrier sense multiple access / collision avoidance (CSMA / CA) can be implemented. In the case of CSMA / CA, STAs including the AP (e.g., any STA) can sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA can back off. Within a given BSS, at any given time, one STA (e.g., only one station) can transmit.
[0050] A high throughput (HT) STA can use a 40 MHz wide channel for communication, for example, by combining the primary 20 MHz channel with adjacent or non - adjacent 20 MHz channels to form a 40 MHz wide channel.
[0051] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non - consecutive 80 MHz channels, which may be referred to as an 80+80 configuration. In the case of the 80+80 configuration, after channel encoding, the data can pass through a segment parser that can split the data into two streams. For each stream separately, an Inverse Fast Fourier Transform (IFFT) process and time - domain processing can be performed. The streams can be mapped onto two 80 MHz channels and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration can be reversed and the combined data can be transmitted to the Media Access Control (MAC).
[0052] The operation of the sub-1 GHz mode is supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using the non-TVWS spectrum. According to an exemplary embodiment, 802.11ah can support meter type control / machine type communication, such as MTC devices in a macro coverage area. The MTC devices can have limited functionality including certain functions, for example, support for a certain bandwidth and / or limited bandwidth (e.g., only their support). The MTC devices can include a battery having a battery life above a threshold (e.g., to maintain a very long battery life).
[0053] A WLAN system that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes channels that can be designated as primary channels. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs within a BSS. The bandwidth of the primary channel can be set and / or restricted by the STA that supports the minimum bandwidth operation mode among all STAs operating within the BSS. In the example of 802.11ah, for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode, even if the AP and other STAs within the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operation modes, the primary channel can be 1MHz wide. Carrier sensing and / or network allocation vector (NAV) setting can depend on the status of the primary channel. For example, if the primary channel is busy because an STA (that only supports the 1MHz operation mode) is transmitting to the AP, the entire available frequency band can be considered busy, even if most of the frequency band remains idle and available.
[0054] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz, depending on national regulations.
[0055] Figure 1D is a system diagram showing RAN 113 and CN 115 according to an embodiment. As mentioned above, RAN 113 can communicate with WTRUs 102a, 102b, 102c over air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0056] RAN 113 can include gNBs 180a, 180b, 180c, although it will be understood that RAN 113 can include any number of gNBs while maintaining consistency with the embodiment. Each of gNBs 180a, 180b, 180c can include one or more transceivers for communicating with WTRUs 102a, 102b, 102c over air interface 116. In one embodiment, gNBs 180a, 180b, 180c can implement MIMO technology. For example, WTRUs 102a, 108b can use beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, 180c. Thus, gNB 180a, for example, can use multiple antennas to transmit wireless signals to and / or receive wireless signals from WTRU 102a. In an embodiment, gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers can be in unlicensed spectrum, while the remaining component carriers can be in licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c can implement multi-point coordinated (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0057] The WTRUs 102a, 102b, 102c can communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol interval, and / or the OFDM sub-carrier interval can be different for different transmissions, different cells, and / or different portions of the radio transmission spectrum. The WTRUs 102a, 102b, 102c can communicate with the gNBs 180a, 180b, 180c using sub-frames or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or lasting for various lengths of absolute time).
[0058] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (such as eNodeBs 160a, 160b, and 160c for example). In a stand-alone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchor points. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals within an unlicensed band. In a non-stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with / connect to gNBs 180a, 180b, and 180c while also communicating with / connecting to another RAN such as eNodeBs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement the DC principle to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNodeBs 160a, 160b, and 160c. In a non-stand-alone configuration, eNodeBs 160a, 160b, and 160c can serve as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0059] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, and routing of control plane information to access and mobility management functions (AMFs) 182a, 182b. As shown in FIG. 1D, gNBs 180a, 180b, and 180c can communicate with each other over the Xn interface.
[0060] CN 115 shown in FIG. 1D can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and possibly data networks (DNs) 185a, 185b. Although each of the above elements is depicted as part of CN 115, it will be understood that any of these elements can be owned and / or operated by entities different from the CN operator.
[0061] AMF 182a and 182b can be connected to one or more of gNBs 180a, 180b, and 180c within RAN 113 via the N2 interface and can serve as control nodes. For example, AMF 182a and 182b can authenticate users of WTRUs 102a, 102b, and 102c, support network slicing (e.g., handling different PDU sessions with different requirements), select specific SMFs 183a and 183b, manage the registration area, terminate NAS signaling, and perform mobility management, etc. Network slicing can be used by AMF 182a and 182b to customize the CN support for WTRUs 102a, 102b, and 102c based on the type of service utilized by WTRUs 102a, 102b, and 102c. For example, different network slices can be established for different use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on high-speed large-capacity mobile broadband (eMBB) access, and / or services for machine type communication (MTC) access. AMF 182 can provide control plane functions for the exchange between RAN 113 and other RANs (not shown) that utilize other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.
[0062] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b and configure the routing of traffic through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and allocating WTRU or UE IP addresses, managing PDU sessions, enforcing policies and controlling QoS, and providing downlink data notifications. The PDU session type can be IP-based, non-IP-based, and Ethernet-based, etc.
[0063] UPF184a and 184b can be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N3 interface, and they can provide access to a packet-switched network such as the Internet 110 to WTRU102a, 102b, and 102c to facilitate communication between WTRU102a, 102b, and 102c and IP-corresponding devices. UPF184a and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homing PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0064] CN115 can facilitate communication with other networks. For example, CN115 can include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN115 and the PSTN 108. Additionally, CN115 can provide access to other networks 112 to the WTRUs 102a, 102b, 102c, where the other networks 112 can include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c can be connected to the local data networks (DNs) 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.
[0065] With reference to FIGS. 1A - 1D and the corresponding descriptions thereof, one or more of the functions described herein with respect to one or more of the WTRUs 102a - d, base stations 114a - b, eNodeBs 160a - c, MME 162, SGW 164, PGW 166, gNBs 180a - c, AMFs 182a - b, UPFs 184a - b, SMFs 183a - b, DNs 185a - b, and / or any other devices described herein can be performed by one or more emulation devices (not shown). An emulation device can be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device can be used to test other devices and / or to simulate network and / or WTRU functionality.
[0066] An emulation device can be designed to perform one or more tests on other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices can execute one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices can execute one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device can be directly coupled to another device for the purpose of conducting a test and / or can perform the test using over-the-air wireless communication.
[0067] One or more emulation devices can execute one or more functions including all functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device can be utilized in a test scenario in a test laboratory and / or in a non-deployed (e.g., test) wired and / or wireless communication network to perform tests on one or more components. One or more emulation devices can be test equipment. Direct RF coupling and / or wireless communication via an RF circuit (which can include one or more antennas for example) can be used by an emulation device to transmit and / or receive data.
[0068] A video encoding system can be used to compress a digital video signal, which can reduce the need for storage of the video signal and / or the transmission bandwidth. The video encoding system can include a block-based, wavelet-based, and / or object-based system. The block-based video encoding system can include MPEG1 / 2 / 4 Part 2, H.264 / MPEG4 Part 10 AVC, VC-1, High Efficiency Video Coding (HEVC), and / or Versatile Video Coding (VVC).
[0069] A block-based video encoding system can include a block-based hybrid video encoding framework. FIG. 2 is a diagram of an exemplary block-based hybrid video encoding framework for an encoder. The encoder can include a WTRU. An input video signal 202 can be processed block by block. The block size (e.g., an enlarged block size such as a coding unit (CU)) can compress high-resolution (e.g., 1080p and above) video signals. For example, a CU can include 64×64 pixels or more. A CU can be partitioned into prediction units (PUs) and / or can use separate predictions. For an input video block (e.g., a macroblock (MB) and / or a CU), spatial prediction 260 and / or temporal prediction 262 can be performed. Spatial prediction 260 (e.g., intra prediction) can use pixels from samples of encoded adjacent blocks (e.g., reference samples) in a video picture / slice to predict the current video block. Spatial prediction 260 can reduce spatial redundancy, which can be inherent in the video signal, for example. Motion prediction 262 (e.g., inter prediction and / or temporal prediction) can use pixels reconstructed from an encoded video picture to predict the current video block, for example. Motion prediction 262 can reduce temporal redundancy, which can be inherent in the video signal, for example. A motion prediction signal for a video block can be signaled by one or more motion vectors and / or can indicate the amount and / or direction of motion between the current block and / or a reference block of the current block. If multiple reference pictures are supported for each video block, a reference picture index for the video block can be sent. The reference picture index can be used to identify from which reference picture in a reference picture storage 264 a motion prediction signal can be derived.
[0070] After spatial prediction 260 and / or motion prediction 262, the mode decision block 280 in the encoder can determine a prediction mode (e.g., the best prediction mode), for example, based on rate-distortion optimization. The prediction block is subtracted 216 from the current video block and / or the prediction residual is decorrelated using transformation 204 and / or quantization 206 to achieve a bit rate such as a target bit rate. The quantized residual coefficients can be inverse quantized, for example, in quantization 210 and / or inverse transformed in transformation 212 to form a reconstructed residual, which can be added 226 to the prediction block, for example, to form a reconstructed video block. Before the reconstructed video block is put into the reference picture memory 264 and / or before it can be used to encode a video block (e.g., a future video block), in-loop filtering (e.g., deblocking filter and / or adaptive loop filter) can be applied to the video block to be reconstructed in loop filter 266. To form the output video bit stream 220, the coding mode (e.g., inter or intra), prediction mode information, motion information, and / or the quantized residual coefficients can be sent, for example, compressed and / or packed, to the entropy coding module 208 to form a bit stream (e.g., all can be sent).
[0071] Figure 3 is a diagram of an exemplary block-based video decoding framework for a decoder. The decoder can include a WTRU. A video bitstream 302 (e.g., the video bitstream 220 in FIG. 2) can be unpacked (e.g., first unpacked) and / or entropy decoded in an entropy decoding module 308. Encoding mode and prediction information can be sent to a spatial prediction module 360 (e.g., in the case of intra encoding) and / or a motion compensation prediction module 362 (e.g., in the case of inter encoding and / or temporal encoding) to form prediction blocks. Residual transform coefficients can be sent to an inverse quantization module 310 and / or an inverse transform module 312, for example, to reconstruct a residual block. The prediction block and / or the residual block can be added together at 326. The reconstructed block can pass through in-loop filtering in a loop filter 366, for example, before the reconstructed block is stored in a reference picture memory device 364. The reconstructed video 320 in the reference picture memory device 364 can be sent to a display device and / or used to predict video blocks (e.g., future video blocks).
[0072] In a video codec, using bidirectional motion compensation prediction (MCP) can remove temporal redundancy by exploiting temporal correlation between pictures. A bidirectional prediction signal can be formed by combining two unidirectional prediction signals using a weight value (e.g., 0.5). In a video, the illumination characteristics can change rapidly for each reference picture. Therefore, prediction techniques can compensate for illumination variations (e.g., fading transitions) over time by applying global or local weights and offset values to one or more sample values in the reference picture.
[0073] MCP in bi-prediction mode can be implemented using CU weight. As an example, MCP can be implemented using bi-prediction with CU weight. An example of bi-prediction with CU weight (BCW) can include generalized bi-prediction (GBi). Bi-prediction signal can be calculated based on one or more of weight, motion compensation prediction signal corresponding to motion vector associated with reference picture list, and / or the like. As an example, prediction signal of (given) sample x in bi-prediction mode can be calculated by Equation 1.
[0074] P[x]=w0*P0[x+v0]+w1*P1[x+v1] Formula 1 P[x] may denote the obtained prediction signal of sample x located at picture position x. i [x+v i ] is the motion vector (MV) v for the i-th list (e.g., list 0, list 1, etc.). i w0 and w1 may denote the motion compensation prediction signal of x using w0 and w1. w0 and w1 may denote two weight values applied to the prediction signal for the block and / or CU. As an example, w0 and w1 may denote two weight values shared across samples in the block and / or CU. By adjusting the weight values, various prediction signals can be obtained. As shown in Equation 1, various prediction signals can be obtained by adjusting the weight values w0 and w1.
[0075] Some configurations of weight values w0 and w1 may indicate prediction, such as unidirectional prediction and / or bidirectional prediction. For example, (w0,w1)=(1,0) may be used in association with unidirectional prediction using reference list L0. (w0,w1)=(0,1) may be used in association with unidirectional prediction using reference list L1. (w0,w1)=(0.5,0.5) may be used in association with bidirectional prediction using two reference lists (e.g., L1 and L2).
[0076] Weights can be signaled at the CU level. In an example, the weight values w0 and w1 can be signaled for each CU. Bidirectional prediction can be performed using the CU weights. Constraints on the weights can be applied to pairs of weights. The constraints can be pre-configured. For example, the constraints on the weights can include w0 + w1 = 1. The weights can be signaled. The signaled weights can be used to determine another weight. For example, using the constraints on the CU weights, only one weight can be signaled. The signaling overhead can be reduced. Examples of pairs of weights can include {(4 / 8, 4 / 8), (3 / 8, 5 / 8), (5 / 8, 3 / 8), (-2 / 8, 10 / 8), (10 / 8, -2 / 8)}.
[0077] Weights can be derived based on constraints on the weights, for example, when unequal weights are used. The encoding device can receive a weight indication and determine a first weight based on the weight indication. The encoding device can derive a second weight based on the determined first weight and the constraints on the weight.
[0078] Equation 2 can be used. In an example, Equation 2 can be created based on Equation 1 and the constraint w0 + w1 = 1.
[0079] P[x]=(1 - w1)*P0[x + v0]+w1*P1[x + v1] Equation 2 Weight values (e.g., w1 and / or w0) can be discretized. The overhead of weight signaling can be reduced. In an example, the bidirectional prediction CU weight value w1 can be discretized. The discretized weight value w1 can include, for example, one or more of -2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, and / or the like. A weight indication can be used to indicate, for example, the weights used for a CU in bidirectional prediction. Examples of the weight indication can include a weight index. In an example, each weight value can be indicated by an index value.
[0080] Figure 4 is a diagram of an exemplary video encoder that uses the support of BCW (e.g., GBi). The encoding device described in the example shown in Figure 4 can be or can include a WTRU. The encoder can include a mode decision module 404, a spatial prediction module 406, a motion prediction module 408, a transform module 410, a quantization module 412, an inverse quantization module 416, an inverse transform module 418, a loop filter 420, a reference picture memory 422, and an entropy encoding module 414. In the example, some or all of the modules or components of the encoder (e.g., the spatial prediction module 406) can be the same as or similar to those described with respect to Figure 2. Additionally, the spatial prediction module 406 and the motion prediction module 408 can be pixel region prediction modules. Thus, the input video bitstream 402 can be processed in a manner similar to the input video bitstream 202 to output a video bitstream 424. The motion prediction module 408 can further include support for bidirectional prediction with CU weights. In this way, the motion prediction module 408 can combine two separate prediction signals with a weighted average. Further, the selected weight index can be signaled in the input video stream 402.
[0081] FIG. 5 is a diagram of an exemplary module that supports bidirectional prediction using CU weights for an encoder. FIG. 5 shows a block diagram of an estimation module 500. The estimation module 500 can be used in a motion prediction module of an encoder, such as the motion prediction module 408. The estimation module 500 can be used with respect to BCW (e.g., GBi). The estimation module 500 can include a weight value estimation module 502 and a motion estimation module 504. The estimation module 500 can utilize a two-stage process to generate an inter prediction signal, such as a final inter prediction signal. The motion estimation module 504 can perform motion estimation by using a reference picture received from a reference picture storage device 506 and searching for two optimal motion vectors (MVs) that indicate (e.g., two) reference blocks. The weight value estimation module 502 can search for an optimal weight index to minimize the weighted bidirectional prediction error between the current block and the bidirectional prediction. The prediction signal of the generated bidirectional prediction can be calculated as a weighted average of two prediction blocks.
[0082] FIG. 6 is a diagram of an exemplary block-based video decoder that supports bidirectional prediction using CU weights. FIG. 6 shows a block diagram of an exemplary video decoder that can decode a bitstream from an encoder. The encoder can support BCW and / or share some similarities with the encoder described with respect to FIG. 4. The decoder described in the example shown in FIG. 6 can include a WTRU. As shown in FIG. 6, the decoder can include an entropy decoder 604, a spatial prediction module 606, a motion prediction module 608, a reference picture memory 610, an inverse quantization module 612, an inverse transform module 614, and a loop filter module 618. Some or all of the decoder modules can be the same as or similar to those described in connection with FIG. 3. For example, the prediction block and / or the residual block can be added together at 616. The video bitstream 602 can be processed to generate a reconstructed video 620, which can be sent to a display device and / or used to predict video blocks (e.g., future video blocks). The motion prediction module 608 can further include support for BCW. Encoding mode and / or prediction information can be used to derive a prediction signal using spatial prediction or MCP that supports BCW. For BCW, block motion information and / or weight values (e.g., in the form of an index indicating the weight value) can be received and decoded to generate a prediction block.
[0083] FIG. 7 is a diagram of an exemplary module that supports bidirectional prediction using CU weights for a decoder. FIG. 7 shows a block diagram of a prediction module 700. The prediction module 700 can be used in the motion prediction module of a decoder, such as the motion prediction module 608. The prediction module 700 can be used in relation to BCW. The prediction module 700 can include a weighted average module 702 and a motion compensation module 704, which can receive one or more reference pictures from a reference picture memory device 706. The prediction module 700 can use block motion information and weight values to calculate a prediction signal for BCW as a weighted average of (e.g., two) motion-compensated prediction blocks.
[0084] Bidirectional prediction in video coding can be based on a combination of multiple (e.g., two) temporal prediction blocks. In an example, CUs and blocks can be used interchangeably with each other. Temporal prediction blocks can be combined. In an example, two temporal prediction blocks obtained from a reconstructed reference picture can be combined using averaging. Bidirectional prediction can be based on block-based motion compensation. In bidirectional prediction, a relatively small motion can be observed between (e.g., two) prediction blocks.
[0085] For example, bidirectional optical flow (BDOF) can be used to compensate for the relatively small motion observed between prediction blocks. BDOF can be applied to compensate for such motion for samples inside the block. In an example, BDOF can compensate for such motion for individual samples inside the block. Doing so can improve the efficiency of motion-compensated prediction.
[0086] BDOF can include improvements to the motion vectors associated with a block. In an example, BDOF can include per-sample motion improvements that are performed in addition to block-based motion compensation prediction when bidirectional prediction is used. BDOF can include deriving improved motion vectors for samples. As an example of BDOF, the derivation of improved motion vectors for individual samples in a block can be based on an optical flow model.
[0087] BDOF can include improving the motion vectors of sub-blocks associated with a block based on one or more of the following, namely, the position in the block, the gradients associated with the position in the block (e.g., horizontal, vertical, and / or the like), the sample values associated with the reference picture list corresponding to that position, and / or the like. Equation 3 can be used to derive an improved motion vector for a sample. As shown in Equation 3, I (k) (x,y) represents the sample value at the coordinates (x,y) of the prediction block and can be derived from the reference picture list k (k = 0, 1). ∂I (k) (x,y) / ∂x, and ∂I (k) (x,y) / ∂y can be the horizontal and vertical gradients of the sample. The motion improvement (v x ,v y ) at (x,y) can be derived using Equation 3. Equation 3 can be based on the assumption that the optical flow model is valid.
[0088]
Equation
[0089] FIG. 8 shows an exemplary bidirectional optical flow. In FIG. 8, (MV x0 ,MV y0 ) and (MV x1 ,MV y1) can indicate block-level motion vectors. The block-level motion vectors can be used to generate prediction blocks I (0) and I (1) . The motion improvement parameters (v x , v y ) at the sample position (x, y) can be calculated, for example, by minimizing the difference Δ between the motion vector values of the samples after motion improvement (e.g., in FIG. 8, the motion vector A between the current picture and the backward reference picture, and the motion vector B between the current picture and the forward reference picture). The difference Δ between the motion vector values of the samples after motion improvement can be calculated, for example, using Equation 4.
[0090]
Equation
[0091] For motion improvement, it can be assumed, for example, that the samples inside one unit (e.g., a 4×4 block) are consistent. Such an assumption can support the derived regularity of motion improvement.
[0092]
Equation
[0093] The value of can be derived, for example, by minimizing Δ inside the 6×6 window Ω around each 4×4 block as shown in Equation 5.
[0094]
Equation
[0095] In the example, BDOF can include sequential techniques, which can optimize motion improvement horizontally (e.g., the first) and vertically (e.g., the second) as used in association with Equation 5. By doing so, Equation 6 is obtained.
[0096]
Number
[0097] In the formula,
[0098]
Number
[0099] can be a floor function that outputs the maximum value equal to or less than the input, th BIO can be, for example, a motion improvement value (e.g., a threshold value) for preventing error propagation caused by, for example, coding noise and irregular local motion. As an example, the motion improvement value can be 2 18-BD can be. For example, as shown in Equations 7 and 8, the values of S1, S2, S3, S5, and S6 can be calculated.
[0100]
Number
[0101] In the formula,
[0102]
Number
[0103] is.
[0104] The BDOF gradients in Equation 8 in the horizontal and vertical directions can be obtained by calculating the differences between a plurality of adjacent samples at the sample positions of the L0 / L1 prediction blocks. In the example, the differences can be calculated horizontally or vertically between two adjacent samples according to the direction of the gradient derived at one sample position of each L0 / L1 prediction block, for example, by using Equation 9.
[0105]
Number
[0106] In Equation 7, L can be, for example, the increment of the bit depth for the internal BDOF to maintain the accuracy of the data. L can be set to 5. The adjustment parameters r and m in Equation 6 can be defined as shown in Equation 10 (for example, to avoid division by a smaller value).
[0107] r = 500·4 BD-8 m = 700·4 BD-8 Equation 10 BD can be the bit depth of the input video. The current CU's bidirectional prediction signal (for example, the final bidirectional prediction signal) can be calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on, for example, the optical flow Equation 3 and the motion improvement derived from Equation 6. The prediction signal of the current CU can be calculated using Equation 11. The current CU's bidirectional prediction signal can be calculated using Equation 11.
[0108]
Number
[0109] shift and o offset can be the offset and right shift applied to combine the L0 and L1 prediction signals for bidirectional prediction, which can be set equal to 15 - BD and 1 << (14 - BD)+2·(1 << 13) respectively, and rnd(·) can be the rounding function that rounds the input value to the nearest integer value.
[0110] In a specific video, there can be various types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. A translation motion model, and / or an affine motion model can be applied to the MCP. The affine motion model can have 4 parameters, and / or 6 parameters. For example, a first flag for each inter-coded CU can be signaled to indicate whether a translation motion model is applied for inter prediction or an affine motion model is applied. When an affine motion model is applied, a second flag can be sent to indicate whether the model has 4 parameters or 6 parameters.
[0111] The 4-parameter affine motion model can include two parameters for parallel motion in the horizontal and vertical directions, one parameter for zoom motion in the horizontal and vertical directions, and / or one parameter for rotation motion in the horizontal and vertical directions. The horizontal zoom parameter can be equal to the vertical zoom parameter. The horizontal rotation parameter can be equal to the vertical rotation parameter. The 4-parameter affine motion model can be encoded using two motion vectors at two control point positions defined at the upper left and upper right corners of the (e.g., current) CU.
[0112] FIG. 9 shows an exemplary 4-parameter affine mode. FIG. 9 shows an exemplary affine motion field of a block. As shown in FIG. 9, the block is depicted by two control point motion vectors (V0, V1). Based on the motion of the control points, the motion field (v x , v y ) of one affine-coded block can be described by Equation 12.
[0113]
Equation
[0114] In Equation 12, (v 0x , v 0y ) can be the motion vector of the control point at the upper left corner. (v 1x , v 1y ) can be the motion vector of the control point at the upper right corner. w can be the width of the CU. The motion field of the affine-coded CU can be derived at the 4×4 block level. For example, (v x , v y ) is derived for each 4×4 block within the current CU and can be applied to the corresponding 4×4 block.
[0115] The four parameters can be repeatedly estimated. The pair of motion vectors at step k is
[0116]
Number
[0117] and can be denoted as such, where the original luminance signal is I(i, j) and the predicted luminance signal is I’ k (i, j). The spatial gradient
[0118]
Number
[0119] can be derived using the Sobel filters applied to the predicted signal I’ k (i, j) in the horizontal and vertical directions respectively. The differential coefficient of Equation 1 can be expressed by Equation 13.
[0120]
Number
[0121] In Equation 13, (a, b) can be the delta translation parameter, and (c, d) can be the delta zoom and rotation parameters at step k. The delta MV at the control point can be derived as Equations 14 and 15 using its coordinates. For example, (0, 0) and (w, 0) can be the coordinates for the upper left and upper right control points, respectively.
[0122]
Number
[0123]
Number
[0124] Based on the optical flow equation, the relationship between the changes in luminance and spatial gradient and the temporal motion can be formulated as Equation 16.
[0125]
Number
[0126]
Number
[0127] Substituting Equation 13 into it, Equation 17 for the parameters (a, b, c, d) can be generated.
[0128]
Number
[0129] When the sample in the CU satisfies Equation 17, for example, the parameter set (a, b, c, d) can be derived using the least squares calculation. The two control points at step (k + 1)
[0130]
Number
[0131] The motion vectors in are derived using Equations (14) and (15), which can be rounded to a specific accuracy (e.g., 1 / 4 pel). Using iteration, the motion vectors at two control points can be refined until the parameters (a, b, c, d) become zero or the number of iterations meets a predefined limit and converges.
[0132] The six-parameter affine motion model can include two parameters for translation in the horizontal and vertical directions, one parameter for zoom motion in the horizontal direction, one parameter for rotational motion, one parameter for zoom motion in the vertical direction, and / or one parameter for rotational motion. The six-parameter affine motion model can be encoded using three motion vectors at three control points. FIG. 10 shows an exemplary six-parameter affine mode. As shown in FIG. 10, the three control points for the six-parameter affine-coded CU can be defined at the upper left, upper right, and / or lower left corners of the CU. The motion at the upper left control point can be related to translational motion. The motion at the upper right control point can be related to rotational and zoom motion in the horizontal direction. The motion at the lower left control point can be related to rotational and zoom motion in the vertical direction. In the six-parameter affine motion model, the rotational and zoom motion in the horizontal direction may not be the same as these motions in the vertical direction. In the example, the motion vectors (v x , v y ) of each sub-block can be derived from Equations (18) and (19) using three motion vectors as control points.
[0133]
Equation
[0134]
Equation
[0135] In Equations 18 and 19, (v 2x , v 2y ) can be the motion vector of the bottom - left control point. (x, y) can be the center position of the sub - block. w and h can be the width and height of the CU.
[0136] The six parameters of the six - parameter affine model can be estimated, for example, in a similar way. For example, Equation 20 can be created based on Equation 13.
[0137]
Equation
[0138] In Equation 20, for step k, (a, b) can be the delta translation parameters. (c, d) can be the delta zoom and rotation parameters in the horizontal direction. (e, f) can be the delta zoom and rotation parameters in the vertical direction. For example, Equation 21 can be created based on Equation 16.
[0139]
Equation
[0140] The parameter set (a, b, c, d, e, f) can be derived using least - squares calculation by considering the samples within the CU. The motion vector of the top - left control point
[0141]
Equation
[0142] can be calculated using Equation 14. The motion vector of the top - right control point
[0143]
Equation
[0144] can be calculated using Equation 22. The motion vector of the upper right control point
[0145]
Number
[0146] can be calculated using Equation 23.
[0147]
Number
[0148] There may be a symmetric MV difference for bidirectional prediction. In some examples, the motion vectors in the forward reference picture and the backward reference picture may be symmetric, for example, due to the continuity of the motion trajectory in bidirectional prediction.
[0149] SMVD can be in an inter-coding mode. When using SMVD, the MVD of the first reference picture list (e.g., reference picture list 1) may be symmetric with respect to the MVD of the second reference picture list (e.g., reference picture list 0). The motion vector coding information (e.g., MVD) of one reference picture list can be signaled, and the motion vector information of another reference picture list may not be signaled. The motion vector information of another reference picture list can be determined based on, for example, the signaled motion vector information and the symmetry of the motion vector information of the reference picture list. In an example, the MVD of reference picture list 0 is signaled, and the MVD of list 1 may not be signaled. The MV encoded in this mode can be calculated using Equation 24A.
[0150]
Number
[0151] In the formula, the subscript characters indicate reference picture list 0 or 1, x indicates the horizontal direction, and y indicates the vertical direction.
[0152] As shown in Equation 24A, the MVD for the current coding block can represent the difference between the MVP for the current coding block and the MV for the current coding block. Those skilled in the art will understand that the MVP can be determined based on the MVs of spatially adjacent blocks of the current coding block and / or the MVs of temporally adjacent blocks of the current coding block. Equation 24A can be shown in FIG. 11. FIG. 11 shows an exemplary non-affine motion symmetry MVD (e.g., MVD = -MVD). As shown in Equation 24A and depicted in FIG. 11, the MV of the current coding block can be equal to the sum of the MVP for the current coding block and the MVD (or negative MVD depending on the reference picture list) of the current coding block. As shown in Equation 24A and depicted in FIG. 11, in the case of SMVD, the MVD (MVD1) of reference picture list 1 can be equal to the negative of the MVD (MVD0) of reference picture list 0. The MV predictor (MVP), (MVP0) of reference picture list 0 may or may not be symmetric with the MVP, (MVP1) of reference picture list 1. MVP0 may or may not be equal to the negative of MVP1. As shown in Equation 24A, the MV of the current coding block can be equal to the sum of the MVP of the current coding block and the MVD of the current coding block. Based on Equation 24A, the MV (MV0) of reference picture list 0 may not be equal to the negative of the MV (MV1) of reference picture list 1. The MV, MV0 of reference picture list 0 may or may not be symmetric with the MV, MV1 of reference picture list 1.
[0153] SMVD can be used for bidirectional prediction when reference picture list 0 contains forward reference pictures and reference picture list 1 contains backward reference pictures, or when reference picture list 0 contains backward reference pictures and reference picture list 1 contains forward reference pictures.
[0154] When using SMVD, the reference picture indices of reference picture list 0 and list 1 may not be signaled. They can be derived as follows. When reference picture list 0 contains forward reference pictures and reference picture list 1 contains backward reference pictures, the reference picture index in list 0 can be set to the forward reference picture closest to the current picture, and the reference picture index in list 1 can be set to the backward reference picture closest to the current picture. When reference picture list 0 contains backward reference pictures and reference picture list 1 contains forward reference pictures, the reference picture index in list 0 can be set to the backward reference picture closest to the current picture, and the reference picture index in list 1 can be set to the forward reference picture closest to the current picture.
[0155] In the case of SMVD, it is not necessary to signal the reference picture index for both lists. For one reference picture list (e.g., list 0), a set of MVDs can be signaled. This can reduce the signaling overhead for bidirectional prediction coding.
[0156] In merge mode, motion information can be derived and / or used (e.g., directly used) to generate the predicted samples of the current CU. Merge mode using motion vector difference (MMVD) can be used. A merge flag can be signaled to specify whether MMVD is used for a CU. The MMVD flag can be signaled after sending the skip flag.
[0157] In MMVD, after a merge candidate is selected, the merge candidate can be refined by MVD information. The MVD information can be signaled. The MVD information can include one or more of a merge candidate flag, a distance index for specifying the magnitude of motion, and / or an index for indicating the direction of motion. In MMVD, one of a plurality of candidates in the merge list (e.g., the first two candidates) can be selected for use on an MV basis. The merge candidate flag can indicate which candidate is used.
[0158] The distance index can specify motion magnitude information and / or can indicate a predefined offset from the starting point (e.g., from the candidate selected to be MV-based). FIG. 12 shows exemplary motion vector difference (MVD) search points. As shown in FIG. 12, the center point can be the starting point MV. As shown in FIG. 12, the pattern of points can indicate different search orders (e.g., from the point closest to the center MV to those far from the center MV). As shown in FIG. 12, an offset can be added to the horizontal and / or vertical components of the starting point MV. An exemplary relationship between the distance index and the predefined offset is shown in Table 1.
[0159] [Table 1]
[0160] The direction index can represent the direction of the MVD with respect to the starting point. The direction index can represent any one of the four directions shown in Table 2. The meaning of the MVD code can vary according to the information of one or more starting point MVs. When the starting point has a unidirectional prediction MV or a pair of bidirectional prediction MVs where both lists point to the same side of the current picture, the code in Table 2 can specify the sign of the MV offset added to one or more starting MVs. For example, when the picture order counts (POCs) of two references are both greater than the POC of the current picture or both less than the POC of the current picture, the code can specify the sign of the MV offset added to one or more starting MVs. When the starting point has a pair of bidirectional prediction MVs where both lists point to different sides of the current picture (for example, when the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture), the code in Table 2 can specify the sign of the MV offset added to the MV component of list 0 of the starting point MVs, and the sign of the MV offset added to the MVs of list 1 can have the opposite value.
[0161]
Table 2
[0162] The symmetric mode can be used for bidirectional prediction coding. One or more of the features described herein can be used in connection with the symmetric mode for bidirectional prediction coding. For example, in some examples, it can increase coding efficiency and / or reduce complexity. The symmetric mode can include SMVD. One or more of the features described herein can be associated with making a synergistic effect of SMVD with one or more other tools, such as bidirectional prediction using CU weights (BCW or BPWA), BDOF, and / or affine mode, etc. One or more of the features described herein can be used in coding (e.g., encoder optimization, etc.), which can include fast motion estimation for translational and / or affine motion.
[0163] The SMVD coding features (functions) can include one or more of restrictions, signaling, SMVD search features (functions), and / or the like.
[0164] The application of the SMVD mode can be based on the CU size. For example, the restriction can be that SMVD is not allowed for relatively small CUs (e.g., CUs having an area not exceeding 64). The restriction can be that SMVD is not allowed for relatively large CUs (e.g., CUs larger than 32×32). When the restriction does not allow SMVD for a CU, the symmetric MVD signaling is skipped or disabled for that CU, and / or the coding device (e.g., encoder) cannot search for the symmetric MVD.
[0165] The application of the SMVD mode can be based on the POC distance between the current picture and the reference picture. The coding efficiency of SMVD may decrease for a relatively large POC distance (e.g., a POC distance of 8 or more). When the POC distance between the reference picture (e.g., any reference picture) and the current picture is relatively large, SMVD may become inapplicable. When SMVD is inapplicable, the symmetric MVD signaling is skipped or becomes inapplicable, and / or the coding device (e.g., the encoder) may not search for the symmetric MVD.
[0166] The application of the SMVD mode can be restricted to one or more temporal layers. In an example, the lower temporal layer can refer to a reference picture having a large POC distance from the current picture in a hierarchical GOP structure. The coding efficiency of SMVD may decrease for the lower temporal layer. SMVD coding may not be allowed for relatively low temporal layers (e.g., temporal layers 0 and 1). When SMVD is not allowed, the symmetric MVD signaling can be skipped or become inapplicable, and / or the coding device (e.g., the encoder) may not search for the symmetric MVD.
[0167] In SMVD coding, one MVD of the reference picture list can be signaled (e.g., signaled explicitly). In an example, the coding device (e.g., the decoder) can parse the first MVD associated with the first reference picture list in the bitstream. The coding device can determine the second MVD associated with the second reference picture list based on the first MVD and determine that the first MVD and the second MVD are symmetric to each other.
[0168] The symbolization device (e.g., a decoder) can identify which reference picture list's MVD is signaled, for example, whether the MVD of reference picture list 0 is signaled or the MVD of reference picture list 1 is sent. In the example, the MVD of reference picture list 0 can be signaled (e.g., always signaled). The MVD of reference picture list 1 can be obtained (e.g., can be derived).
[0169] The reference picture list for which the MVD is signaled (e.g., explicitly signaled) can be selected. One or more of the following can be applied. An indication (e.g., a flag) can be signaled to indicate which reference picture list is selected. A reference picture list having a smaller POC distance with respect to the current picture can be selected. When the POC distances for the reference picture lists are the same, the reference picture lists can be pre-determined to break the balance. For example, when the POC distances for the reference picture lists are the same, reference picture list 0 can be selected.
[0170] The MVP index for a reference picture list (e.g., one reference picture list) can be signaled. In some examples, the indexes of the MVP candidates for both reference picture lists can be signaled (e.g., explicitly signaled). The MVP index for a reference picture list can be signaled (e.g., only the MVP index for one reference picture list) to reduce the signaling overhead, for example. The MVP index for other reference picture lists can be derived as described herein, for example. LX can be the reference picture list for which the MVP index is signaled (e.g., explicitly signaled), and i can be the signaled MVP index. mvp’ can be derived from the MVP of LX as shown in Equation 24.
[0171]
Number
[0172] In the formula, POC LX , POC 1-LX , and POC curr can be the POC of the reference picture in list LX, the reference picture in list (1-LX), and the current picture, respectively. Also, from the MVP list of the reference picture list (1-LX), the MVP closest to mvp’ can be selected, for example, as shown in Equation 25.
[0173]
Number
[0174] In the formula, j can be the MVP index of the reference picture list (1-LX). LX can be the reference picture list in which the MVD is signaled (for example, explicitly signaled).
[0175] Table 3 shows an exemplary CU syntax that can support symmetric MVD signaling for non-affine coding modes.
[0176]
Table 3-1
[0177]
Table 3-2
[0178] For example, an indication such as the sym_mvd_flag flag can indicate whether the SMVD is used in motion vector coding for the current coding block (for example, a bi-predictive coded CU).
[0179] Indications such as refIdxSymL0 can indicate the reference picture index in reference picture list 0. The refIdxSymL0 indication set to -1 can indicate that SMVD is not applicable and there is no sym_mvd_flag.
[0180] Indications such as refIdxSymL1 can indicate the reference picture index in reference picture list 1. The refIdxSymL1 indication having a value of -1 can indicate that SMVD is not applicable and there is no sym_mvd_flag.
[0181] In the example, SMVD may be applicable when reference picture list 0 contains a forward reference picture and reference picture list 1 contains a backward reference picture, or when reference picture list 0 contains a backward reference picture and reference picture list 1 contains a forward reference picture. In other cases, SMVD is not applicable. For example, when SMVD is not applicable, the signaling of the SMVD indication (e.g., the sym_mvd_flag flag at the CU level) may be skipped. The encoding device (e.g., the decoder) can perform one or more condition checks. As shown in Table 3, two condition checks can be performed (e.g., refIdxSymL0 > -1, and refIdxSymL1 > -1). One or more condition checks can be performed to determine whether the SMVD indication is used. As shown in Table 3, two condition checks are performed (e.g., refIdxSymL0 > -1, and refIdxSymL1 > -1) to determine whether the SMVD indication is received. In the example, in order for the decoder to examine these conditions, the decoder can wait for the reference picture lists (e.g., list 0 and list 1) of the current picture before performing a certain CU syntax analysis. In some examples, even if both of the two examined conditions (e.g., refIdxSymL0 > -1 and refIdxSymL1 > -1) are true, the encoder may not use SMVD for the CU (e.g., to save encoding complexity).
[0182] The SMVD indication is at the CU level and can be associated with the current coding block. The CU-level SMVD indication can be obtained based on the upper-level indication. In the example, the presence of the CU-level SMVD indication (e.g., the sym_mvd_flag) can be controlled by the upper-level indication (e.g., alternatively, or in addition). For example, the SMVD enabled indication such as the sym_mvd_enabled_flag can be signaled at the slice level, tile level, tile group level, or picture parameter set (PPS) level, sequence parameter set (SPS) level, and / or any syntax level, and the reference picture list is shared by the CU associated with the syntax level. For example, the slice-level flag can be placed in the slice header. In the example, the coding device (e.g., decoder) can receive a sequence-level SMVD indication indicating whether SMVD is available (enabled) for the picture sequence. If SMVD is available (enabled) for the sequence, the coding device can obtain the SMVD indication associated with the current coding block based on the sequence-level SMVD indication.
[0183] Using the upper-level SMVD enabled indication such as the sym_mvd_enabled_flag, the CU-level syntax analysis can be performed without examining one or more conditions (e.g., as described herein). SMVD can be made available or unavailable at a level higher than the CU level (e.g., at the discretion of the encoder). Table 4 shows an exemplary syntax that can support the SMVD mode.
[0184]
Table 4-1
[0185]
Table 4-2
[0186] There can be various ways for a symbolization device (e.g., an encoder) to determine whether to enable (make effective, enable) the SMVD and accordingly set the value of the sym_mvd_enabled_flag. One or more of the examples herein can be combined to determine whether to enable the SMVD and accordingly set the value of the sym_mvd_enabled_flag. For example, the encoder can determine whether to enable the SMVD by examining the reference picture list. If both a forward reference picture and a backward reference picture exist, the sym_mvd_enabled_flag can be set to be equal to true to enable the SMVD. For example, the encoder can determine whether to enable the SMVD based on the temporal distance between the current picture and the forward and / or backward reference pictures. When the reference picture is far from the current picture, the SMVD mode may not be effective. If the forward or backward reference picture is far from the current picture, the encoder can disable the SMVD. The symbolization device can set the value of the upper-level SMVD availability indication based on a value (e.g., a threshold value, etc.) for the temporal distance between the current picture and the reference picture. For example, to reduce the encoding complexity, the encoder can use (e.g., only use) the upper-level control flag to enable (make effective, enable) the SMVD for pictures with forward and backward reference pictures, so that the maximum temporal distance between these two reference pictures for the current picture is less than a value (e.g., a threshold value, etc.). For example, the encoder can determine whether to enable (make effective, enable) the SMVD based on the temporal layer (e.g., of the current picture). A relatively low temporal layer can indicate that the reference picture is far from the current picture, in which case the SMVD may not be effective.The encoder can determine that the current picture belongs to a relatively low temporal layer (e.g., lower than thresholds such as 1, 2, etc.), and can disable the use of SMVD for such a current picture. For example, the sym_mvd_enabled_flag can be set to false to disable SMVD. For example, the encoder can determine whether SMVD should be made available (enabled) based on the statistics of the previously encoded pictures in the same temporal layer as the current picture's temporal layer. The statistics can include the average POC distance for bi - directionally predicted coded CUs (e.g., the average of the distances between the current picture and the temporal centers of the two reference pictures of the current picture). R0 and R1 can be the reference pictures for bi - directionally predicted coded CUs. poc(x) can be the POC of picture x. The POC distance distance(CU) between the two reference pictures and the current picture (current_picture) can be calculated using Equation 26. i ) can be calculated using Equation 26. The average POC distance AvgDist for bi - directionally predicted coded CUs can be calculated using Equation 27.
[0187] Distance(CUi)=|2 * poc(current_picture)-poc(R0)-poc(R1)| Equation 26
[0188]
Number
[0189] The variable N can indicate the total number of bi - directionally predicted coded CUs that can have both forward and backward reference pictures. As an example, if AvgDist is smaller than a value (e.g., a pre - defined threshold), the sym_mvd_enabled_flag can be set to true by the encoder to make SMVD available (enabled), and in other cases, the sym_mvd_enabled_flag can be set to false to disable SMVD.
[0190] In some examples, the MVD value may be signaled. In some examples, a combination of a direction index and a distance index may be signaled and the MVD value may not be signaled. The exemplary direction tables and exemplary distance tables shown in Tables 1 and 2 can be used to signal and derive MVD information. For example, distance index 0 and direction index 0 can indicate MVD(1 / 2,0).
[0191] The search for symmetric MVD can be performed, for example, after a uni-directional prediction search and a bi-directional prediction search. The uni-directional prediction search can be used to search for the optimal MV pointing to a reference block for uni-directional prediction. The bi-directional prediction search can be used to search for two optimal MVs pointing to two reference blocks for bi-directional prediction. The search can be performed to find candidate symmetric MVDs, such as the best symmetric MVD. In an example, a set of search points can be repeatedly evaluated for the symmetric MVD search. The repetition can include the evaluation of the set of search points. The set of search points can form a search pattern centered around the best MV of the previous repetition, for example. For the first repetition, the search pattern can be centered around the first MV. The selection of the first MV can affect the overall result. A set of first MV candidates can be evaluated. The first MV for the symmetric MVD search can be determined, for example, based on the rate-distortion cost. In an example, the MV candidate having the lowest rate-distortion cost can be selected as the first MV for the symmetric MVD search. The rate-distortion cost can be estimated, for example, by summing the bi-directional prediction error of the MVD coding for reference picture list 0 and the weighted rate. The set of first MV candidates can include one or more of the MVs obtained from the uni-directional prediction search, the MVs obtained from the bi-directional prediction search, and the MVs from the advanced motion vector predictor (AMVP) list. At least one MV can be obtained for each reference picture from the uni-directional prediction search.
[0192] For example, in order to reduce complexity, early termination can be applied. Early termination can be applied (e.g., by an encoder) when the cost of bidirectional prediction is greater than a value (e.g., a threshold). In an example, for instance, before the first MV selection, if the rate-distortion cost for the MV obtained from the bidirectional prediction search is greater than a value (e.g., a threshold), the search for the symmetric MVD can be terminated. For example, the value can be set to a multiple (e.g., 1.1 times) of the unidirectional prediction cost. In an example, for instance, after the first MV selection, if the rate-distortion cost associated with the first MV is higher than a value (e.g., a threshold), the symmetric MVD search can be terminated. For example, the value can be set to a multiple (e.g., 1.1 times) of the lowest of the unidirectional prediction cost and the bidirectional prediction cost.
[0193] There may be an interaction between the SMVD mode and other coding tools. One or more of the following can be implemented, namely, combining symmetric affine MVD coding, bidirectional prediction using CU weights (BCW or BPWA) with SMVD, or combining BDOF with SMVD.
[0194] Symmetric affine MVD coding can be used. The affine motion model parameters can be represented by the motion vectors of the control points. A 4-parameter affine model can be represented by two control point MVs, and a 6-parameter affine model can be represented by three control point MVs. In the example shown in Equation 12, for example, it can be a 4-parameter affine model represented by two control point MVs, such as the top-left control point MV (v0) and the top-right control point MV (v1). The top-left control point MV can represent the translational motion. The top-left control point MV can have corresponding symmetric MVs associated with the forward and backward reference pictures following the motion trajectory, for example. The SMVD mode can be applied to the top-left control point. The other control point MVs can represent a combination of zoom, rotation, and / or shear mapping. SMVD may not be applied to the other control point MVs.
[0195] SMVD can be applied to the top-left control point (e.g., only to the top-left control point), and the other control point MVs can be set to their respective MV predictors.
[0196] The MVD of the control points associated with the first reference picture list can be signaled. The MVD of the control points associated with the second reference picture can be obtained based on the MVD of the control points associated with the first reference picture list, and the MVD of the control points associated with the first reference picture list is symmetric with respect to the MVD of the control points associated with the second reference picture list. In the example, when the symmetric affine MVD mode is applied, the MVD of the control points associated with reference picture list 0 can be signaled (e.g., only the control point MVDs of reference picture list 0 can be signaled). The MVD of the control points associated with reference picture list 1 can be derived based on the symmetric property. The MVD of the control points associated with reference list 1 may not be signaled.
[0197] A control point MV associated with a reference picture list can be derived. The control point MV for control point 0 (top left) of reference picture list 0 and reference picture list 1 can be derived, for example, using Equation 28.
[0198]
Number
[0199] Equation 28 can be shown in FIG. 13. FIG. 13 shows an exemplary affine motion symmetry MVD. As shown in Equation 28 and illustrated in FIG. 13, the control point MV at the top left of the current coding block can be equal to the sum of the MVP for the control point at the top left of the current coding block and the MVD (or negative MVD depending on the reference picture list) for the control point at the top left of the current coding block. As shown in Equation 28 and illustrated in FIG. 13, the top left control point MVD associated with reference picture list 1 can be equal to the negative of the top left control point MVD associated with reference picture list 0 for symmetric affine MVD coding.
[0200] The MV for other control points can be derived using at least affine MVP prediction, for example, as shown in Equation 29.
[0201]
Number
[0202] In Equations 28 to 30, the first element of the subscript can indicate the reference picture list. The second element of the subscript can indicate the control point index.
[0203] The translational MVD of the first reference picture list (e.g., mvdx 0,0 , mvdy 0,0 ) is the second reference picture list (e.g., mvdx 1,0 , mvdy 1,0) can be applied to the derivation of the top - left control point MV. The translational MVD of the first reference picture list (e.g., mvdx 0,0 , mvdy 0,0 ) cannot be applied to the derivation of other control point MVs of the second reference picture list (e.g., mvdx 1,j , mvdy 1,j ). In some examples of symmetric affine MVD derivation, the translational MVD of reference picture list 0 (mvdx 0,0 , mvdy 0,0 ) can be applied to the derivation of the top - left control point MV of reference picture list 1 (e.g., only if applicable). The other reference point MVs of list 1 can be the same as the corresponding predictors, for example, as shown in Equation 30.
[0204]
Number
[0205] Table 5 shows an exemplary syntax that can be used to signal the use of SMVD in combination with the affine mode (e.g., symmetric affine MVD coding).
[0206]
Table 5 - 1
[0207]
Table 5 - 2
[0208] For example, based on the inter-affine indication and / or the SMVD indication, it can be determined whether affine SMVD is used. For example, when the inter-affine indication inter_affine_flag is 1 and the SMVD indication sym_mvd_flag[x0][y0] is 1, affine SMVD can be applied. When the inter-affine indication inter_affine_flag is 0 and the SMVD indication sym_mvd_flag[x0][y0] is 1, non-affine motion SMVD can be applied. When the SMVD indication sym_mvd_flag[x0][y0] is 0, SMVD cannot be applied.
[0209] The MVD of the control points in the reference picture list (e.g., reference picture list 0) can be signaled. In the example, the MVD values of a subset of the control points in the reference picture list can be signaled. For example, the MVD of the upper left control point in reference picture list 0 can be signaled. Signaling of the MVD of other control points in reference picture list 0 can be skipped, and for example, the MVD of these control points can be set to 0. The MV of other control points can be derived based on the MVD of the upper left control point in reference picture list 0. For example, the MV of a control point for which the MVD is not signaled can be derived as shown in Equation 31.
[0210]
Number
[0211] Table 6 shows an exemplary syntax that can be used to signal information regarding the SMVD mode combined with the affine mode.
[0212]
Table 6-1
[0213]
Table 6-2
[0214] Indications such as the flag for only the top-left MVD can indicate whether only the MVD of the top-left control point in the reference picture list (e.g., reference picture list 0) is signaled, or whether the MVD of the control point in the reference picture list is signaled. This indication can be signaled at the CU level. Table 7 shows an exemplary syntax that can be used to signal information regarding the SMVD mode combined with the affine mode.
[0215]
Table 7-1
[0216]
Table 7-2
[0217] In the example, as described herein for the MMVD, a combination of a direction index and a distance index can be signaled. In Tables 1 and 2, an exemplary direction table and an exemplary distance table are shown. For example, the combination of direction index 0 and direction index 0 can indicate MVD(1 / 2,0).
[0218] In the example, an indication (e.g., a flag) is signaled to indicate which reference picture list's MVD is sent. The MVDs of other reference picture lists are not signaled and they can be derived.
[0219] One or more of the limitations described herein for translational motion symmetry MVD coding can be applied to affine motion symmetry MVD coding, for example, to reduce complexity and / or reduce signaling overhead.
[0220] Using symmetric affine MVD can reduce signaling overhead. Coding efficiency can be improved.
[0221] Bidirectional prediction motion estimation can be used to search for symmetric MVDs for the affine model. In the example, bidirectional prediction motion estimation can be applied after unidirectional prediction search to find a symmetric MVD for the affine model (e.g., the best symmetric MVD for the affine model). The reference pictures for reference picture list 0 and / or reference picture list 1 can be derived as described herein. The first control point MV can be selected from one or more of the results of unidirectional prediction search, the results of bidirectional prediction search, and / or the MVs from the affine AMVP list. The control point MV (e.g., the control point MV with the lowest rate-distortion cost) can be selected to be the first MV. The encoding device (e.g., an encoder, etc.) can examine one or more cases, i.e., in the first case, the symmetric MVD can be signaled for reference picture list 0, and the control point MV for reference picture list 1 can be derived based on symmetric mapping (using Equation 28 and / or Equation 29); in the second case, the symmetric MVD for reference picture list 1 can be signaled, and the control point MV for reference picture list 0 can be derived based on symmetric mapping. The first case can be used as an example in this specification. The symmetric MVD search technique can be applied based on the unidirectional prediction search results. When a control point MV predictor is given in reference picture list 1, iterative search using a predefined search pattern (e.g., a diamond pattern, a cube pattern, and / or the like, etc.) can be applied. In each iteration (e.g., each repetition), the MVD can be refined by the search pattern, and the control point MVs in reference picture list 0 and reference picture list 1 can be derived using Equation 28 and Equation 29. The bidirectional prediction errors corresponding to the control point MVs in reference picture list 0 and reference picture list 1 can be evaluated. For example, the rate-distortion cost can be estimated by summing the bidirectional prediction error of the MVD encoding for reference picture list 0 and the weighted rate.In an example, an MVD having a low (e.g., the lowest) rate-distortion cost during candidate search can be treated as the best MVD for a symmetric MVD search process. The MVs of other control points, such as the upper-right and lower-left control point MVs, can be improved, for example, using the optical flow-based techniques described herein.
[0222] Symmetric MVD search for symmetric affine MVD coding can be performed. In an example, a set of parameters, such as translation parameters, can first be searched, and then the search for non-translation parameters can be performed. In an example, the optical flow search can be performed by considering (e.g., together) the MVDs of reference picture list 0 and reference picture list 1. For a 4-parameter affine model, an exemplary optical flow equation for the MVD of list 0 can be shown in Equation 32.
[0223]
Number
[0224] In the formula,
[0225]
Number
[0226] can indicate the prediction of list 0 in the k-th iteration, and also
[0227]
Number
[0228] can indicate the spatial gradient of the list 0 prediction.
[0229] Reference picture list 1 can have a translational change (e.g., it can have only a translational change). The translational change has the same magnitude as that of reference picture list 0, but in the opposite direction, which is the condition for a symmetric affine MVD. The optical flow equation for reference picture list 1 MVD can be expressed by Equation 33.
[0230] [Number]
[0231] BCW weights w0 and w1 can be applied to list 0 prediction and list 1 prediction respectively. An exemplary optical equation for a symmetric affine model can be shown by Equation 34. I k ’(i,j) - I(i,j) = (G x (i,j)·i + G y (i,j)·j)·c + (-G x (i,j)·j + G y (i,j)·i)·d + H x (i,j)·a + H y (i,j)·b Equation 34 In the equation,
[0232] [Number]
[0233] is.
[0234] The parameters a, b, c, d can be estimated (e.g., by mean least square error calculation).
[0235] I k ’(i,j) - I(i,j) = G x (i,j)·i·c + G x (i,j)·j·d + G y (i,j)·j·e + G y (i,j)·j·f + H x (i,j)·a + Hy (i,j)·b Equation 35 The parameters a, b, c, d, e, and f can be estimated by calculating the mean least square error. When the joint optical flow search is performed, the affine parameters can be optimized together. The performance can be improved.
[0236] For example, in order to reduce complexity, truncation can be applied. In the example, for example, before making the first MV selection, if the bidirectional prediction cost becomes larger than a value (e.g., a threshold value), the search can be terminated. For example, the value can be set to a multiple of the unidirectional prediction cost, e.g., 1.1 times the unidirectional prediction cost. In the example, the encoding device (e.g., an encoder, etc.) can compare the current best affine motion estimation (ME) cost with the non-affine ME cost (considering unidirectional prediction and bidirectional prediction affine search) before the ME of the symmetric affine MVD starts. If the current best affine ME cost is larger than the non-affine ME cost multiplied by a value (e.g., a threshold value such as 1.1), the encoding device can skip the ME of the symmetric affine MVD. In the example, for example, after the first MV selection, the affine symmetric MVD search can be skipped if the cost of the first MV is higher than a value (e.g., a threshold value). For example, the value can be set to a multiple of the lowest of the unidirectional prediction cost and the bidirectional prediction cost (e.g., set to 1.1 times). In the example, the value can be set to a multiple (e.g., 1.1) of the non-affine ME cost.
[0237] SMVD can be combined with BCW. If BCW is available (enabled) for the current CU, SMVD can be applied in one or more ways. In some examples, SMVD can be made available when the weights of BCW are equal weights (e.g., 0.5, etc.) (e.g., only at that time), and for other BCW weights, SMVD may not be available. In such cases, the SMVD flag is signaled before the BCW weight index, and the signaling of the BCW weight index can be conditionally controlled by the SMVD flag. The SMVD indication (e.g., the SMVD flag, etc.) can have a value of 1, and the signaling of the BCW weight index can be skipped. The decoder can assume that the SMVD indication has a value of 0, which can correspond to equal weights for bidirectional prediction averaging. When the SMVD flag is 0, the BCW weight index can be encoded for the bidirectional prediction mode. In examples where SMVD is available when the BCW weights are equal weights and not available for other BCW weights, the BCW index may be skipped. In some examples, SMVD can be fully combined with BCW. The SMVD flag and the BCW weight index can be signaled for the explicit bidirectional prediction mode. The MVD search (e.g., of the encoder) for SMVD can consider the BCW weight index during bidirectional prediction averaging. The SMVD search can be based on the evaluation of one or more (e.g., all) possible BCW weights.
[0238] A symbolization tool (e.g., bidirectional optical flow (BDOF)) can be used in association with one or more other symbolization tools / modes. BDOF can be used in association with SMVD. Whether BDOF is applied to an encoding block may depend on whether SMVD is used. SMVD can be based on the assumption of symmetric MVD at the encoding block level. When implemented, BDOF can be used to improve sub-block MVs based on optical flow. The optical flow can be based on the assumption of symmetric MVD at the sub-block level.
[0239] BDOF can be used in association with SMVD. In an example, an encoding device (e.g., a decoder or an encoder) can receive one or more indications that SMVD and / or BDOF are available. BDOF be is available for the current picture (enabled) can be . The encoding device can determine whether to bypass or implement BDOF for the current encoding block. The encoding device can determine whether to bypass BDOF based on an SMVD indication (e.g., sym_mvd_flag[x0][y0]). BDOF can be used interchangeably with BIO in some examples.
[0240] The encoding device can determine whether to bypass BDOF for the current encoding block . For example For example, if the SMVD mode is used for the motion vector encoding of the current encoding block to reduce the decoding complexity, BDOF can be bypass for the current encoding block performed . If SMVD is not used for the motion vector encoding of the current encoding block, the encoding device can determine whether to enable BDOF for the current encoding block based on, for example, at least another condition.
[0241] The symbolization device can obtain an SMVD indication (e.g., sym_mvd_flag[x0][y0]). The SMVD indication can indicate whether SMVD is used in the motion vector for the current coding block.
[0242] The current coding block can be reconstructed based on a determination of whether to bypass BDOF. The MVD can be signaled at the CU level for the SMVD mode (e.g., explicitly signaled).
[0243] The symbolization device can be configured to perform motion vector coding using SMVD without using BDOF based on a determination to bypass BDOF.
[0244] Features and elements are described above in a particular combination, but those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein can be implemented by a computer program, software, or firmware incorporated into a computer-readable medium to be executed by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections), and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A device for video decoding, comprising: obtaining a symmetric motion vector difference indication (SMVD indication) associated with a video block of a picture, the SMVD indication indicating whether an SMVD mode is used in determining a motion vector (MV) for the video block; determining whether to bypass BD OF for the video block based on the SMVD indication associated with the video block, and determining to bypass BD OF for the video block based on a condition that the SMVD mode is used in the determination of the MV for the video block; a processor configured to decode the picture based on the determination of whether to bypass BD OF for the video block A device comprising the same.
2. The device according to claim 1, wherein SMVD uses a motion vector difference (MVD) for the video block, and the MVD indicates a difference between a motion vector predictor (MVP) for the video block and the MV for the video block.
3. The device according to claim 2, wherein the MVP for the video block is determined based on an MV of a spatially adjacent block or a temporally adjacent block of the video block.
4. Based on a condition that the SMVD indication indicates that SMVD is used in the determination of the MV for the video block, the processor: receives first MV information associated with a first reference picture list; determines second MV information associated with a second reference picture list based on the first MV information associated with the first reference picture list, and the motion vector difference (MVD) associated with the first reference picture list and the MVD associated with the second reference picture list are symmetric; The device according to claim 1, further configured as described above.
5. Based on a condition that the SMVD indication indicates that SMVD is used in the determination of the MV for the video block, the processor: Syntax-analyze a first motion vector difference (first MVD) associated with a first reference picture list in video data, Based on the first MVD, determine a second MVD associated with a second reference picture list, and the first MVD and the second MVD are symmetric to each other The device according to claim 1, further configured as such.
6. Based on the determination not to bypass BDOF for the video block, the processor The device according to claim 1, further configured to improve a motion vector (MV) associated with a sub-block of the video block based at least in part on a gradient associated with a location in the video block.
7. A method for video decoding, comprising: Obtaining a symmetric motion vector difference indication (SMVD indication) associated with a video block of a picture, the SMVD indication indicating whether an SMVD mode is used in determining a motion vector (MV) for the video block; Determining whether to bypass BDOF for the video block based on the SMVD indication associated with the video block, including bypassing BDOF for the video block based on a condition under which the SMVD mode is used in the determination of the MV for the video block; Decoding the picture based on the determination of whether to bypass BDOF for the video block A method comprising.
8. Based on a condition under which the SMVD indication indicates that SMVD is used in the determination of the MV for the video block, Receiving first MV information associated with a first reference picture list; Determining second MV information associated with a second reference picture list based on the first MV information associated with the first reference picture list, wherein a motion vector difference (MVD) associated with the first reference picture list and an MVD associated with the second reference picture list are symmetric to each other The method according to claim 7, further comprising.
9. Based on the condition that the SMVD indication indicates that SMVD is used in the determination of the MV for the video block, syntax-analyzing a first motion vector difference (first MVD) associated with a first reference picture list in video data; determining a second MVD associated with a second reference picture list based on the first MVD, wherein the first MVD and the second MVD are symmetric to each other The method according to claim 7, further comprising.
10. A computer-readable medium including instructions for causing one or more processors to perform the method according to any one of claims 7 to 9.
11. A device for video encoding, determining whether a symmetric motion vector difference (SMVD) mode is used for a video block in determining a motion vector (MV) for the video block, in the determination of the motion vector (MV) for the video block, determining whether to bypass BDOF for the video block based on the determination of whether the SMVD mode is used for the video block, and in the determination of the motion vector (MV) for the video block, being configured to determine to bypass BDOF for the video block based on the condition that the SMVD mode is used for the video block, a processor configured to encode the picture based on the determination of whether to bypass BDOF for the video block A device comprising.
12. Based on the condition that SMVD is enabled for the video block, the processor is obtaining a first motion vector difference (MVD) associated with a first reference picture list, wherein the first MVD associated with the first reference picture list is symmetric to a second MVD associated with a second reference picture list, including an indication of the first SMVD associated with the first reference picture list in the video data, and the video data does not include an indication of the second MVD associated with the second reference picture list The device according to claim 11, further configured as such.
13. The motion vector difference (MVD) for the video block indicates the difference between the motion vector predictor (MVP) for the video block and the MV for the video block, and the MVP for the video block is determined based on the MV of a spatially adjacent block of the video block or a temporally adjacent block of the video block. The device according to claim 11.
14. A method for video encoding, comprising: determining that a symmetric motion vector difference (SMVD) mode is used in determining a motion vector (MV) for a video block with respect to the video block of a picture; determining whether to bypass BDOF for the video block based on the determination of whether the SMVD mode is used for the video block of the picture in the determination of the motion vector (MV) for the video block, including bypassing BDOF for the video block based on the condition that the SMVD mode is used for the video block of the picture in the determination of the motion vector (MV) for the video block; encoding the picture based on the determination of whether to bypass BDOF for the video block A method comprising.
15. Based on the condition that SMVD is enabled for the video block, obtaining a first motion vector difference (MVD) associated with a first reference picture list, wherein the first MVD associated with the first reference picture list is symmetric with a second MVD associated with a second reference picture list; including an indication of the first SMVD associated with the first reference picture list in the video data, wherein the video data does not include an indication of the second MVD associated with the second reference picture list; The method according to claim 14, further comprising.
16. Based on the determination not to bypass BDOF for the video block, Improving the MV associated with the sub-blocks of the video block based at least in part on the gradient associated with the location in the video block The method according to claim 14, further comprising
Citation Information
Patent Citations
Constraining motion vector information derived by decoder-side motion vector derivation
JP2020511859A
Constraining motion vector information derived by decoder-side motion vector derivation
WO2018175720A1