Symmetric motion vector difference coding
By employing symmetric motion vector difference (SMVD) to bypass bidirectional optical flow (BDOF) in video coding, the complexity of calculations is reduced, enhancing the efficiency of video encoding and decoding processes.
Patent Information
- Application Number
- JP2025027949
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-02-22
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video coding systems increase the complexity of calculations during encoding and decoding due to techniques like bidirectional motion compensation prediction (MCP).
The use of symmetric motion vector difference (SMVD) in motion vector encoding allows for the bypassing of bidirectional optical flow (BDOF) for the current coding block, simplifying calculations by determining motion vector differences based on symmetry between reference picture lists.
This approach reduces the computational complexity of encoding and decoding by eliminating the need for bidirectional optical flow calculations when SMVD is used, thereby improving efficiency without compromising video quality.
Smart Images

Figure 2025081627000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to symmetric motion vector differential coding.
Background Art
[0002] Cross-reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 783,437, filed on December 21, 2018; U.S. Provisional Patent Application No. 62 / 787,321, filed on January 1, 2019; U.S. Provisional Patent Application No. 62 / 792,710, filed on January 15, 2019; U.S. Provisional Patent Application No. 62 / 798,674, filed on January 30, 2019; and U.S. Provisional Patent Application No. 62 / 809,308, filed on February 22, 2019, the contents of which are hereby incorporated by reference in their entirety.
[0003] Video coding systems are used to compress digital video signals, for example, reducing the storage and / or transmission bandwidth required for such signals. Video coding systems can include block-based, wavelet-based, and / or object-based systems. The system uses video coding techniques such as bidirectional motion compensation prediction (MCP) that can remove temporal redundancy by utilizing the temporal correlation between pictures.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Such techniques can increase the complexity of the calculations performed during encoding and / or decoding.
Means for Solving the Problems
[0005] Whether symmetric motion vector difference (SMVD) is used in motion vector encoding for the current coding block, bidirectional optical flow (BDOF) can be bypassed for the current coding block.
[0006] A coding device (e.g., an encoder or a decoder) can determine that BDOF is available (enabled). The coding device can determine whether to bypass BDOF for the current coding block based at least in part on an SMVD indication for the current coding block. The coding device can obtain an SMVD indication indicating whether SMVD is used in motion vector encoding for the current coding block. If the SMVD indication indicates that SMVD is used in motion vector encoding for the current coding block, the coding device can bypass BDOF for the current coding block. If the coding device determines to bypass BDOF for the current coding block, the coding device can reconstruct the current coding block without performing BDOF.
[0007] The motion vector difference (MVD) for the current coding block can indicate the difference between the motion vector predictor (MVP) for the current coding block and the motion vector (MV) for the current coding block. The MVP for the current coding block can be determined based on the MVs of spatially adjacent blocks of the current coding block and / or temporally adjacent blocks of the current coding block.
[0008] If the SMVD indication indicates that SMVD is used for motion vector coding for the current coding block, the coding device can receive first motion vector coding information associated with the first reference picture list. Based on the first motion vector coding information associated with the first reference picture list and based on the symmetry between the MVD associated with the first reference picture list and the MVD associated with the second reference picture list, the coding device can determine second motion vector coding information associated with the second reference picture list.
[0009] In the example, if the SMVD indication indicates that SMVD is used for motion vector coding for the current coding block, the coding device can parse the first MVD associated with the first reference picture list in the bitstream. Based on the first MVD and based on the symmetry between the first MVD and the second MVD, the coding device can determine the second MVD associated with the second reference picture list.
[0010] If the coding device determines not to bypass BDOF for the current coding block, the coding device can improve the motion vectors of the (e.g., each) sub-block of the current coding block based at least in part on the gradient associated with the position in the current coding block.
[0011] The coding device can receive a sequence-level SMVD indication indicating whether SMVD is available (enabled) for a sequence of pictures. If SMVD is enabled for a sequence of pictures, the coding device can obtain the SMVD indication associated with the current coding block based on the sequence-level SMVD indication.
Brief Description of the Drawings
[0012]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
[0013] A more detailed understanding can be obtained from the following detailed description given by way of example, in conjunction with the drawings attached hereto.
[0014] FIG. 1A is a diagram illustrating an exemplary communication system 100 in which one or more of the disclosed embodiments can be implemented. The communication system 100 can be a multi-connection system that provides content such as voice, data, video, messaging, broadcasting, etc. to a plurality of wireless users. The communication system 100 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 can utilize one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multicarrier (FBMC).
[0015] As shown in FIG. 1A, communication system 100 can include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, although the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, each of them, which may sometimes be referred to as a “station” and / or “STA”, WTRUs 102a, 102b, 102c, 102d can be configured to transmit and / or receive wireless signals and can be a user equipment (UE), mobile station, fixed or mobile subscriber unit, subscription-based unit, pager, cellular phone, personal digital assistant (PDA), smartphone, laptop, netbook, personal computer, wireless sensor, hotspot or Mi-Fi device, Internet of Things (IoT) device, watch or other wearable, head-mounted display (HMD), vehicle, drone, medical device and application (e.g., remote surgery), industrial device and application (e.g., robot and / or other wireless devices operating in industrial and / or automated processing chain scenarios), home appliance device, and devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, 102d may alternatively be referred to as a UE.
[0016] The communication system 100 can also include base station 114a and / or base station 114b. Each of base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as CN106 / 115, the Internet 110, and / or other network 112. By way of example, base stations 114a, 114b can be a base transceiver station (BTS), Node B, eNode B, home Node B, home eNode B, gNB, New Radio (NR) Node B, site controller, access point (AP), and wireless router, among others. Although base stations 114a, 114b are each depicted as a single element, it will be understood that base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0017] Base station 114a can be part of RAN104 / 113, which can also include other base stations and / or network elements such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes (not shown). Base station 114a and / or base station 114b can be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies can be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. The cell can provide coverage for wireless services in a specific geographic area that can be relatively constant or can change over time. The cell can further be divided into cell sectors. For example, the cell associated with base station 114a can be divided into three sectors. Thus, in one embodiment, base station 114a can include three transceivers, for example, one for each sector of the cell. In an embodiment, base station 114a can utilize multiple-input multiple-output (MIMO) technology and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0018] Base stations 114a, 114b can communicate with one or more of WTRUs 102a, 102b, 102c, 102d over air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).
[0019] More specifically, as mentioned above, the communication system 100 can be a multi-connection system and can utilize one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a within RAN 104 / 113 and the WTRUs 102a, 102b, 102c can establish the air interface 116 using wideband CDMA (WCDMA) and implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0020] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can establish the air interface 116 using Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) and implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA).
[0021] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can establish the air interface 116 using New Radio (NR) and implement radio technologies such as NR radio access.
[0022] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interfaces utilized by the WTRUs 102a, 102b, 102c can be characterized by transmissions from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).
[0023] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement wireless technologies such as IEEE 802.11 (e.g., Wireless Fidelity (WiFi)), IEEE 802.16 (e.g., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0024] The base station 114b in FIG. 1A can be, for example, a wireless router, a home node B, a home e-node B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a localized area such as an office, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by a drone), and a roadway. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in FIG. 1A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0025] RAN 104 / 113 can communicate with CN 106 / 115, and CN 106 / 115 can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, 102d. The data can have various Quality of Service (QoS) requirements such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN 106 / 115 can provide call control, billing services, mobile location-based services, prepaid outgoing calls, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that utilize the same or a different Radio Access Technology (RAT) as RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113 which may be utilizing New Radio (NR) wireless technology, CN 106 / 115 can also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi wireless technology.
[0026] CN106 / 115 can also serve as a gateway for WTRU102a, 102b, 102c, 102d to access the PSTN108, the Internet 110, and / or other networks 112. The PSTN108 can include a circuit-switched telephone network that provides basic telephone service (POTS). The Internet 110 can include a worldwide system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) within the TCP / IP Internet protocol suite. The network 112 can include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 can include another CN connected to one or more RANs that can utilize the same RAT or a different RAT as the RAN104 / 113.
[0027] Some or all of the WTRU102a, 102b, 102c, 102d within the communication system 100 can include a multimode functionality (e.g., the WTRU102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU102c shown in Figure 1A can be configured to communicate with a base station 114a that can utilize a cellular-based wireless technology and to communicate with a base station 114b that can utilize IEEE802 wireless technology.
[0028] Figure 1B is a system diagram illustrating an exemplary WTRU102. As shown in Figure 1B, the WTRU102 can include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU102 can include any sub-combination of the above elements while maintaining consistency with the embodiments.
[0029] The processor 118 can be, for example, a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors cooperating with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, among others. The processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, and the transceiver 120 can be coupled to the transmit / receive element 122. Although Figure 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0030] The transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmitting / receiving element 122 can be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmitting / receiving element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0031] In FIG. 1B, the transmitting / receiving element 122 is depicted as a single element, but the WTRU 102 can include any number of transmitting / receiving elements 122. More specifically, the WTRU 102 can utilize MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0032] The transceiver 120 can be configured to modulate the signals that are to be transmitted by the transmitting / receiving element 122 and demodulate the signals received by the transmitting / receiving element 122. As mentioned above, the WTRU 102 can have a multimode function. Thus, the transceiver 120 can include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11, for example.
[0033] The processor 118 of the WTRU 102 can be coupled to and receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can obtain information from and store data in any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 can include a random access memory (RAM), a read only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card, among others. In other embodiments, the processor 118 can obtain information from and store data in a memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0034] The processor 118 can receive power from a power source 134 and can be configured to distribute power to and / or control power to other components within the WTRU 102. The power source 134 can be any suitable device for powering the WTRU 102. For example, the power source 134 can include one or more dry cells (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell, among others.
[0035] Processor 118 can also be coupled to a GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 can receive location information on the air interface 116 from a base station (e.g., base stations 114a, 114b), and / or can determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 can obtain location information using any suitable location determination method while maintaining consistency with the embodiments.
[0036] Processor 118 can further be coupled to other peripheral devices 138, which can include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 can include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a Frequency Modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, and an activity tracker, among others. The peripheral devices 138 can include one or more sensors, which can be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0037] WTRU102 can include a full-duplex radio in which some or all of the transmission and reception of signals (associated with a particular subframe for both, e.g., UL (for transmission) and downlink (for reception)) can be parallel and / or simultaneous. The full-duplex radio can include an interference management unit 139 to reduce and / or substantially eliminate self-interference, either via hardware (e.g., a choke) or via signal processing through a processor (e.g., a separate processor (not shown) or processor 118). In embodiments, WTRU102 can include a half-duplex radio for some or all of the transmission and reception of signals (associated with a particular subframe for either, e.g., UL (for transmission) or downlink (for reception)).
[0038] FIG. 1C is a system diagram illustrating RAN104 and CN106, according to an embodiment. As mentioned above, RAN104 can communicate with WTRU102a, 102b, 102c over air interface 116 using E-UTRA radio technology. RAN104 can also communicate with CN106.
[0039] RAN104 can include eNodeBs 160a, 160b, 160c, although it will be understood that RAN104 can include any number of eNodeBs while maintaining consistency with the embodiments. Each of eNodeBs 160a, 160b, 160c can include one or more transceivers for communicating with WTRU102a, 102b, 102c over air interface 116. In one embodiment, eNodeBs 160a, 160b, 160c can implement MIMO technology. Thus, eNodeB 160a, for example, can transmit wireless signals to and / or receive wireless signals from WTRU102a using multiple antennas.
[0040] Each of the eNodeBs 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL. As shown in Figure 1C, the eNodeBs 160a, 160b, and 160c can communicate with each other over the X2 interface.
[0041] CN106 shown in Figure 1C can include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Each of the above elements is depicted as part of CN106, but it will be understood that any of these elements can be owned and / or operated by an entity different from the CN operator.
[0042] The MME 162 can be connected to each of the eNodeBs 160a, 160b, and 160c within the RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, and selecting a specific serving gateway during the initial attach of the WTRUs 102a, 102b, 102c. The MME 162 can provide control plane functions for exchanges between the RAN 104 and other RANs (not shown) that utilize other radio technologies such as GSM and / or WCDMA.
[0043] The SGW 164 can be connected to each of the eNodeBs 160a, 160b, and 160c within the RAN 104 via the S1 interface. The SGW 164 can generally route and transfer user data packets to / from the WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions such as anchoring the user plane during an eNodeB handover, triggering paging when DL data is available to the WTRUs 102a, 102b, and 102c, and managing and storing the contexts of the WTRUs 102a, 102b, and 102c.
[0044] The SGW 164 can be connected to the PGW 166, and the PGW 166 can provide access to a packet switched network, such as the Internet 110, to the WTRUs 102a, 102b, and 102c to facilitate communication between the WTRUs 102a, 102b, and 102c and IP-enabled devices.
[0045] The CN 106 can facilitate communication with other networks. For example, the CN 106 can provide access to a circuit switched network, such as the PSTN 108, to the WTRUs 102a, 102b, and 102c to facilitate communication between the WTRUs 102a, 102b, and 102c and conventional fixed line communication devices. For example, the CN 106 can include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 can provide access to other networks 112 to the WTRUs 102a, 102b, and 102c, and the other networks 112 can include other wired and / or wireless networks owned and / or operated by other service providers.
[0046] In FIGS. 1A-1D, the WTRU is described as a wireless terminal, but in certain representative embodiments, it is contemplated that such a terminal may (e.g., temporarily or permanently) use a wired communication interface to a communication network.
[0047] In some representative embodiments, the other network 112 can be a WLAN.
[0048] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have an access or interface to a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic to an STA originating from outside the BSS can arrive through the AP and be delivered to the STA. Traffic transmitted from an STA to a destination outside the BSS can be sent to the AP for delivery to each respective destination. Traffic between STAs within the BSS can be sent through the AP; for example, a source STA can send the traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered peer-to-peer traffic and / or sometimes be referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent (e.g., directly) between a source STA and a destination STA using direct link setup (DLS). In certain representative embodiments, the DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) can communicate directly with each other. Communication in IBSS mode is sometimes referred to herein as "ad hoc" mode communication.
[0049] When using the operation of 802.11ac infrastructure mode or the operation of a similar mode, the AP can transmit beacons on a fixed channel such as the primary channel. The primary channel can be of a fixed width (e.g., 20 MHz bandwidth), or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In one representative embodiment, for example, in an 802.11 system, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) can be implemented. In the case of CSMA / CA, STAs including the AP (e.g., any STA) can sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA can back off. Within a given BSS, at any given time, one STA (e.g., only one station) can transmit.
[0050] A high throughput (HT) STA can use a 40 MHz wide channel for communication, for example, by combining the primary 20 MHz channel with adjacent or non - adjacent 20 MHz channels to form a 40 MHz wide channel.
[0051] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non - consecutive 80 MHz channels, which may be referred to as an 80 + 80 configuration. In the case of the 80 + 80 configuration, after channel encoding, the data can pass through a segment parser that can split the data into two streams. For each stream separately, an Inverse Fast Fourier Transform (IFFT) process and time - domain processing can be performed. The streams can be mapped onto two 80 MHz channels and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80 + 80 configuration can be reversed and the combined data can be transmitted to the Media Access Control (MAC).
[0052] Operation in the sub-1 GHz mode is supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using the non-TVWS spectrum. According to an exemplary embodiment, 802.11ah can support meter type control / machine type communication, such as MTC devices in a macro coverage area. The MTC devices can have limited functionality including certain functions, for example, support for a certain bandwidth and / or limited bandwidth (e.g., only their support). The MTC devices can include a battery having a battery life above a threshold (e.g., to maintain a very long battery life).
[0053] A WLAN system that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes channels that can be designated as primary channels. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs within a BSS. The bandwidth of the primary channel can be set and / or restricted by the STA that supports the minimum bandwidth operating mode among all STAs operating within the BSS. In the example of 802.11ah, even if the AP and other STAs within the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode, the primary channel can be 1MHz wide. Carrier sensing and / or network allocation vector (NAV) setting can depend on the status of the primary channel. For example, if the primary channel is busy because an STA (that only supports the 1MHz operating mode) is transmitting to the AP, the entire available frequency band can be considered busy even though most of the frequency band remains idle and available.
[0054] In the United States, the available frequency band that can be used by 802.11ah is from 902MHz to 928MHz. In Korea, the available frequency band is from 917.5MHz to 923.5MHz. In Japan, the available frequency band is from 916.5MHz to 927.5MHz. The total bandwidth available for 802.11ah is from 6MHz to 26MHz, depending on national regulations.
[0055] Figure 1D is a system diagram showing RAN 113 and CN 115 according to an embodiment. As mentioned above, RAN 113 can communicate with WTRUs 102a, 102b, 102c over air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0056] RAN 113 can include gNBs 180a, 180b, 180c, although it will be understood that RAN 113 can include any number of gNBs while maintaining consistency with the embodiment. Each of gNBs 180a, 180b, 180c can include one or more transceivers for communicating with WTRUs 102a, 102b, 102c over air interface 116. In one embodiment, gNBs 180a, 180b, 180c can implement MIMO technology. For example, WTRUs 102a, 108b can transmit signals to and / or receive signals from gNBs 180a, 180b, 180c using beamforming. Thus, gNB 180a, for example, can transmit a radio signal to WTRU 102a and / or receive a radio signal from WTRU 102a using multiple antennas. In an embodiment, gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers can be in unlicensed spectrum, while the remaining component carriers can be in licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c can implement multi-site coordinated (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0057] WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol interval, and / or the OFDM sub-carrier interval can vary for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using sub-frames or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing various numbers of OFDM symbols and / or lasting for various lengths of absolute time).
[0058] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (such as eNodeBs 160a, 160b, and 160c). In a stand-alone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchor points. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals within an unlicensed band. In a non-stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with / connect to gNBs 180a, 180b, and 180c while also communicating with / connecting to another RAN such as eNodeBs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement the DC principle to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNodeBs 160a, 160b, and 160c. In a non-stand-alone configuration, eNodeBs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0059] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, and routing of control plane information to access and mobility management functions (AMFs) 182a, 182b. As shown in FIG. 1D, gNBs 180a, 180b, and 180c can communicate with each other over the Xn interface.
[0060] CN 115 shown in FIG. 1D can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and possibly data networks (DNs) 185a, 185b. Although each of the above elements is depicted as part of CN 115, it will be understood that any of these elements can be owned and / or operated by entities different from the CN operator.
[0061] AMF182a and 182b can be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N2 interface and can serve as control nodes. For example, AMF182a and 182b can authenticate users of WTRU102a, 102b, and 102c, support network slicing (e.g., handling different PDU sessions with different requirements), select specific SMF183a and 183b, manage the registration area, terminate NAS signaling, and perform mobility management, etc. Network slicing can be used by AMF182a and 182b to customize the CN support for WTRU102a, 102b, and 102c based on the type of services utilized by WTRU102a, 102b, and 102c. For example, different network slices can be established for different use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on high-speed large-capacity mobile broadband (eMBB) access, and / or services for machine-type communication (MTC) access. AMF182 can provide control plane functions for exchanges between RAN113 and other RANs (not shown) that utilize other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.
[0062] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b, and configure the routing of traffic through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and allocating WTRU or UE IP addresses, managing PDU sessions, implementing policies and controlling QoS, and providing downlink data notifications. The PDU session type can be IP-based, non-IP-based, and Ethernet-based, etc.
[0063] UPF184a and 184b can be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N3 interface, and they can provide access to a packet-switched network such as the Internet 110 to WTRU102a, 102b, and 102c, facilitating communication between WTRU102a, 102b, and 102c and IP-compatible devices. UPF184a and 184b can perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-homing PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0064] CN115 can facilitate communication with other networks. For example, CN115 can include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN115 and the PSTN108. In addition, CN115 can provide access to other networks 112 to WTRU102a, 102b, 102c, and other networks 112 can include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a, 102b, 102c can be connected to local data networks (DN) 185a, 185b through UPF184a, 184b via an N3 interface to UPF184a, 184b and an N6 interface between UPF184a, 184b and DN185a, 185b.
[0065] In view of FIGS. 1A - 1D and the corresponding descriptions thereof, one or more of the functions described herein with respect to one or more of WTRU102a - d, base stations 114a - b, eNodeB 160a - c, MME162, SGW164, PGW166, gNB180a - c, AMF182a - b, UPF184a - b, SMF183a - b, DN185a - b, and / or any other devices described herein can be performed by one or more emulation devices (not shown). An emulation device can be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device can be used to test other devices and / or to simulate network and / or WTRU functionality.
[0066] An emulation device can be designed to perform one or more tests on other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices can execute one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices can execute one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device can be directly coupled to another device for the purpose of conducting tests and / or can perform tests using over-the-air wireless communication.
[0067] One or more emulation devices can execute one or more functions including all functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device can be utilized in a test scenario in a test laboratory and / or in a non-deployed (e.g., test) wired and / or wireless communication network to perform tests on one or more components. One or more emulation devices can be test equipment. Direct RF coupling and / or wireless communication via an RF circuit (which can include one or more antennas for example) can be used by an emulation device to transmit and / or receive data.
[0068] A video encoding system can be used to compress a digital video signal, which can reduce the need for storage of the video signal and / or the transmission bandwidth. The video encoding system can include a block-based, wavelet-based, and / or object-based system. The block-based video encoding system can include MPEG1 / 2 / 4 Part 2, H.264 / MPEG4 Part 10 AVC, VC-1, High Efficiency Video Coding (HEVC), and / or Versatile Video Coding (VVC).
[0069] A block-based video encoding system can include a block-based hybrid video encoding framework. FIG. 2 is a diagram of an exemplary block-based hybrid video encoding framework for an encoder. The encoder can include a WTRU. An input video signal 202 can be processed block by block. The block size (e.g., an enlarged block size such as a coding unit (CU)) can compress high-resolution (e.g., 1080p and above) video signals. For example, a CU can include 64×64 pixels or more. A CU can be partitioned into prediction units (PUs) and / or can use separate predictions. For an input video block (e.g., a macroblock (MB) and / or a CU), spatial prediction 260 and / or temporal prediction 262 can be performed. Spatial prediction 260 (e.g., intra prediction) can use pixels from samples of encoded adjacent blocks (e.g., reference samples) in a video picture / slice to predict a current video block. Spatial prediction 260 can reduce spatial redundancy, which can be inherent in a video signal, for example. Motion prediction 262 (e.g., inter prediction and / or temporal prediction) can use pixels reconstructed from an encoded video picture to predict a current video block, for example. Motion prediction 262 can reduce temporal redundancy, which can be inherent in a video signal, for example. A motion prediction signal for a video block can be signaled by one or more motion vectors and / or can indicate the amount and / or direction of motion between a current block and / or a reference block of a current block. If multiple reference pictures are supported for a (e.g., each) video block, a reference picture index for the video block can be sent. The reference picture index can be used to identify from which reference picture in a reference picture storage 264 a motion prediction signal can be derived.
[0070] After spatial prediction 260 and / or motion prediction 262, the mode decision block 280 in the encoder can determine a prediction mode (e.g., the best prediction mode), for example, based on rate-distortion optimization. The prediction block is subtracted 216 from the current video block and / or the prediction residue is decorrelated using transformation 204 and / or quantization 206 to achieve a bit rate such as a target bit rate. The quantized residue coefficients can be inverse quantized in quantization 210 and / or inverse transformed in transformation 212 to form a reconstructed residue, which can be added 226 to the prediction block, for example, to form a reconstructed video block. Before the reconstructed video block is put into the reference picture memory 264 and / or before it can be used to encode a video block (e.g., a future video block), in-loop filtering (e.g., deblocking filter and / or adaptive loop filter) can be applied to the video block to be reconstructed in loop filter 266. To form the output video bit stream 220, the encoding mode (e.g., inter or intra), prediction mode information, motion information, and / or the quantized residue coefficients can be sent, for example, compressed and / or packed, to the entropy encoding module 208 to form a bit stream (e.g., all can be sent).
[0071] Figure 3 is a diagram of an exemplary block-based video decoding framework for a decoder. The decoder can include a WTRU. A video bitstream 302 (e.g., the video bitstream 220 in FIG. 2) can be unpacked (e.g., unpacked first) and / or entropy decoded in an entropy decoding module 308. Encoding mode and prediction information can be sent to a spatial prediction module 360 (e.g., in the case of intra encoding) and / or a motion compensation prediction module 362 (e.g., in the case of inter encoding and / or temporal encoding) to form prediction blocks. Residual transform coefficients can be sent to an inverse quantization module 310 and / or an inverse transform module 312, for example, to reconstruct residual blocks. The prediction block and / or the residual block can be added together at 326. The reconstructed block can pass through in-loop filtering in a loop filter 366, for example, before the reconstructed block is stored in a reference picture memory device 364. The reconstructed video 320 in the reference picture memory device 364 can be sent to a display device and / or used to predict video blocks (e.g., future video blocks).
[0072] In a video codec, using bidirectional motion compensation prediction (MCP) can remove temporal redundancy by exploiting temporal correlation between pictures. A bidirectional prediction signal can be formed by combining two unidirectional prediction signals using a weight value (e.g., 0.5). In a video, the illumination characteristics can change rapidly for each reference picture. Thus, prediction techniques can compensate for fluctuations in illumination (e.g., fading transitions) over time by applying global or local weights and offset values to one or more sample values in the reference picture.
[0073] The MCP in the bidirectional prediction mode can be implemented using CU weights. As an example, the MCP can be implemented using bidirectional prediction with CU weights. Examples of bidirectional prediction with CU weights (BCW) can include generalized bidirectional prediction (GBi). The bidirectional prediction signal can be calculated based on one or more of weights, a motion compensation prediction signal corresponding to a motion vector associated with a reference picture list, and / or the like. As an example, the prediction signal for a (given) sample x in the bidirectional prediction mode can be calculated by Equation 1.
[0074] P[x]=w 0 *P 0 [x+v 0 +w 1 *P 1 [x+v 1 Equation 1 P[x] can indicate the obtained prediction signal of sample x located at picture position x. P i [x+v i can indicate the motion compensation prediction signal of x using the motion vector (MV) v for the i-th list (e.g., list 0, list 1, etc.). w i and w 0 can indicate two weight values applied to the prediction signal for the block and / or CU. As an example, w 1 and w 0 can indicate two weight values shared across samples in the block and / or CU. By adjusting the weight values, various prediction signals can be obtained. As shown in Equation 1, by adjusting the weight values w 1 and w 0 and w 1 , various prediction signals can be obtained.
[0075] Some configurations of the weight values w 0 and w 1 can indicate predictions such as unidirectional prediction and / or bidirectional prediction. For example, (w 0 , w 1)=(1,0) can be used in association with unidirectional prediction using reference list L0. (w 0 ,w 1 )=(0,1) can be used in association with unidirectional prediction using reference list L1. (w 0 ,w 1 )=(0.5,0.5) can be used in association with bidirectional prediction using two reference lists (e.g., L1 and L2).
[0076] The weights can be signaled at the CU level. In the example, the weight values w 0 and w 1 can be signaled for each CU. Bidirectional prediction can be performed using the CU weights. Constraints on the weights can be applied to pairs of weights. The constraints can be preconfigured. For example, the constraints on the weights can include w 0 +w 1 =1. The weights can be signaled. The signaled weights can be used to determine another weight. For example, using the constraints on the CU weights, only one weight can be signaled. The signaling overhead can be reduced. Examples of pairs of weights can include {(4 / 8, 4 / 8), (3 / 8, 5 / 8), (5 / 8, 3 / 8), (-2 / 8, 10 / 8), (10 / 8, -2 / 8)}.
[0077] The weights can be derived based on the constraints on the weights, for example, when unequal weights are used. The encoding device can receive a weight indication and determine a first weight based on the weight indication. The encoding device can derive a second weight based on the determined first weight and the constraints on the weight.
[0078] Equation 2 can be used. In the example, Equation 2 can be created based on Equation 1 and the constraint w 0 +w 1 =1.
[0079] P[x]=(1 - w 1 )*P0 [x + v 0 + w 1 * P 1 [x + v 1 Formula 2 Weight value (e.g., w 1 and / or w 0 ) can be discretized. The overhead of weight signaling can be reduced. In the example, the bi - directional prediction CU weight value w 1 can be discretized. The discretized weight value w 1 can include, for example, one or more of - 2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, and / or the like. A weight indication can be used to indicate the weight used, for example, for a bi - directional prediction CU. An example of a weight indication can include a weight index. In the example, each weight value can be indicated by an index value.
[0080] Figure 4 is a diagram of an exemplary video encoder that uses the support of a BCW (e.g., GBi). The encoding device described in the example shown in Figure 4 can be, or can include, a WTRU. The encoder can include a mode decision module 404, a spatial prediction module 406, a motion prediction module 408, a transform module 410, a quantization module 412, an inverse quantization module 416, an inverse transform module 418, a loop filter 420, a reference picture memory 422, and an entropy encoding module 414. In the example, some or all of the modules or components of the encoder (e.g., the spatial prediction module 406) can be the same as, or similar to, those described with respect to Figure 2. Additionally, the spatial prediction module 406 and the motion prediction module 408 can be a pixel region prediction module. Thus, the input video bitstream 402 can be processed in a manner similar to the input video bitstream 202 to output a video bitstream 424. The motion prediction module 408 can further include support for bidirectional prediction with CU weights. In this way, the motion prediction module 408 can combine two separate prediction signals with a weighted average. Further, the selected weight index can be signaled in the input video stream 402.
[0081] FIG. 5 is a diagram of an exemplary module that supports bidirectional prediction using CU weights for an encoder. FIG. 5 shows a block diagram of an estimation module 500. The estimation module 500 can be used in the motion prediction module of an encoder, such as the motion prediction module 408. The estimation module 500 can be used with respect to BCW (e.g., GBi). The estimation module 500 can include a weight value estimation module 502 and a motion estimation module 504. The estimation module 500 can utilize a two-stage process to generate an inter prediction signal, such as a final inter prediction signal. The motion estimation module 504 can perform motion estimation by using a reference picture received from a reference picture memory device 506 and searching for two optimal motion vectors (MVs) that indicate (e.g., two) reference blocks. The weight value estimation module 502 can search for an optimal weight index to minimize the weighted bidirectional prediction error between the current block and the bidirectional prediction. The prediction signal of the generated bidirectional prediction can be calculated as a weighted average of two prediction blocks.
[0082] FIG. 6 is a diagram of an exemplary block-based video decoder that supports bidirectional prediction using CU weights. FIG. 6 shows a block diagram of an exemplary video decoder that can decode a bitstream from an encoder. The encoder can support BCW and / or share some similarities with the encoder described with respect to FIG. 4. The decoder described in the example shown in FIG. 6 can include a WTRU. As shown in FIG. 6, the decoder can include an entropy decoder 604, a spatial prediction module 606, a motion prediction module 608, a reference picture memory 610, an inverse quantization module 612, an inverse transform module 614, and a loop filter module 618. Some or all of the decoder modules can be the same as or similar to those described in connection with FIG. 3. For example, the prediction block and / or the residual block can be added together at 616. The video bitstream 602 can be processed to generate a reconstructed video 620, which can be sent to a display device and / or used to predict video blocks (e.g., future video blocks). The motion prediction module 608 can further include support for BCW. Encoding mode and / or prediction information can be used to derive a prediction signal using spatial prediction or MCP that supports BCW. For BCW, block motion information and / or weight values (e.g., in the form of an index indicating a weight value) can be received and decoded to generate a prediction block.
[0083] FIG. 7 is a diagram of an exemplary module that supports bidirectional prediction using CU weights for a decoder. FIG. 7 shows a block diagram of a prediction module 700. The prediction module 700 can be used in a motion prediction module of a decoder, such as the motion prediction module 608. The prediction module 700 can be used in relation to BCW. The prediction module 700 can include a weighted average module 702 and a motion compensation module 704, which can receive one or more reference pictures from a reference picture memory device 706. The prediction module 700 can calculate a prediction signal for BCW as a weighted average of (e.g., two) motion compensation prediction blocks using block motion information and weight values.
[0084] Bidirectional prediction in video coding can be based on a combination of multiple (e.g., two) temporal prediction blocks. In the example, CUs and blocks can be used interchangeably with each other. Temporal prediction blocks can be combined. In the example, two temporal prediction blocks obtained from reconstructed reference pictures can be combined using averaging. Bidirectional prediction can be based on block-based motion compensation. In bidirectional prediction, a relatively small motion can be observed between (e.g., two) prediction blocks.
[0085] For example, bidirectional optical flow (BDOF) can be used to compensate for the relatively small motion observed between prediction blocks. BDOF can be applied to compensate for such motion for samples inside the block. In the example, BDOF can compensate for such motion for individual samples inside the block. Doing so can improve the efficiency of motion compensation prediction.
[0086] BDOF can include improvements to the motion vectors associated with a block. In an example, BDOF can include per-sample motion improvements that are performed in addition to block-based motion compensation prediction when bidirectional prediction is used. BDOF can include deriving improved motion vectors for samples. As an example of BDOF, the derivation of improved motion vectors for individual samples in a block can be based on an optical flow model.
[0087] BDOF can include improving the motion vectors of sub-blocks associated with a block based on one or more of the following, namely, the position in the block, the gradients associated with the position in the block (e.g., horizontal, vertical, and / or the like), the sample values associated with the reference picture list corresponding to that position, and / or the like. Equation 3 can be used to derive an improved motion vector for a sample. As shown in Equation 3, I (k) (x,y) represents the sample value at the coordinates (x,y) of the predicted block and can be derived from the reference picture list k (k = 0,1). ∂I (k) (x,y) / ∂x, and ∂I (k) (x,y) / ∂y can be the horizontal and vertical gradients of the sample. The motion improvement (v x ,v y ) at (x,y) can be derived using Equation 3. Equation 3 can be based on the assumption that the optical flow model is valid.
[0088]
Equation
[0089] FIG. 8 shows an exemplary bidirectional optical flow. In FIG. 8, (MV x0 ,MV y0 ) and (MV x1 ,MV y1can indicate block-level motion vectors. Block-level motion vectors can be used to generate prediction blocks I (0) and I (1) . The motion refinement parameters (v x , v y ) at the sample position (x, y) can be calculated, for example, by minimizing the difference Δ between the motion vector values of the samples after motion refinement (e.g., in FIG. 8, motion vector A between the current picture and the backward reference picture, and motion vector B between the current picture and the forward reference picture). The difference Δ between the motion vector values of the samples after motion refinement can be calculated, for example, using Equation 4.
[0090] [Equation]
[0091] It can be assumed that motion refinement is consistent, for example, for samples inside one unit (e.g., a 4×4 block). Such an assumption can support the derived regularity of motion refinement.
[0092] [Equation]
[0093] The value of can be derived, for example, by minimizing Δ inside the 6×6 window Ω around each 4×4 block as shown in Equation 5.
[0094] [Equation]
[0095] In the example, BDOF can include sequential techniques, which can optimize motion refinement horizontally (e.g., the first) and vertically (e.g., the second) as used in association with Equation 5. By doing so, Equation 6 is obtained.
[0096] [Number]
[0097] In the formula,
[0098] [Number]
[0099] can be a floor function that outputs the maximum value equal to or less than the input, th BIO can be, for example, a motion improvement value (e.g., a threshold value) for preventing error propagation caused by encoding noise and irregular local motion. As an example, the motion improvement value can be 2 18-BD can be. For example, as shown in Equations 7 and 8, S 1 S 2 S 3 S 5 , and S 6 values can be calculated.
[0100] [Number]
[0101] In the formula,
[0102] [Number]
[0103] is.
[0104] The BDOF gradients in Equation 8 in the horizontal and vertical directions can be obtained by calculating the differences between a plurality of adjacent samples at the sample positions of the L0 / L1 prediction blocks. In the example, the differences can be calculated horizontally or vertically between two adjacent samples according to the direction of the gradient derived at one sample position of each L0 / L1 prediction block, for example, by using Equation 9.
[0105] [Number]
[0106] In Equation 7, L can be, for example, the increment of the bit depth for the internal BDOF to maintain the accuracy of the data. L can be set to 5. The adjustment parameters r and m in Equation 6 can be defined as shown in Equation 10 (for example, to avoid division by a smaller value).
[0107] r = 500·4 BD-8 m = 700·4 BD-8 Equation 10 BD can be the bit depth of the input video. The bi-directional prediction signal of the current CU (for example, the final bi-directional prediction signal) can be calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on, for example, the optical flow Equation 3 and the motion improvement derived from Equation 6. The prediction signal of the current CU can be calculated using Equation 11. The bi-directional prediction signal of the current CU can be calculated using Equation 11.
[0108] [Number]
[0109] shift and o offsetcan be the offset and right shift applied to combine the L0 and L1 prediction signals for bidirectional prediction, which can be set equal to 15 - BD and 1≪(14 - BD)+2·(1≪13) respectively, and rnd(·) can be a rounding function that rounds the input value to the nearest integer value.
[0110] In a specific video, there can be various types of motion, such as zoom in / out, rotation, perspective motion, and other irregular motions. A translation motion model and / or an affine motion model can be applied to the MCP. The affine motion model can have 4 parameters and / or 6 parameters. For example, a first flag for each inter-coded CU can be signaled to indicate whether a translation motion model or an affine motion model is applied for inter-prediction. When an affine motion model is applied, a second flag can be sent to indicate whether the model has 4 parameters or 6 parameters.
[0111] The 4 - parameter affine motion model can include two parameters for parallel motion in the horizontal and vertical directions, one parameter for zoom motion in the horizontal and vertical directions, and / or one parameter for rotation motion in the horizontal and vertical directions. The horizontal zoom parameter can be equal to the vertical zoom parameter. The horizontal rotation parameter can be equal to the vertical rotation parameter. The 4 - parameter affine motion model can be encoded using two motion vectors at two control point positions defined at the upper left and upper right corners of the (e.g., current) CU.
[0112] Figure 9 shows an exemplary 4 - parameter affine mode. Figure 9 shows an exemplary affine motion field of a block. As shown in Figure 9, the block has motion vectors (V 0 ,V1 ) is depicted by. Based on the movement of the control points, the motion field (v x , v y ) of one affine-encoded block can be described by Equation 12.
[0113]
Number
[0114] In Equation 12, (v 0x , v 0y ) can be the motion vector of the control point at the upper-left corner. (v 1x , v 1y ) can be the motion vector of the control point at the upper-right corner. w can be the width of the CU. The motion field of the affine-encoded CU can be derived at the 4×4 block level. For example, (v x , v y ) is derived for each 4×4 block within the current CU and can be applied to the corresponding 4×4 block.
[0115] The four parameters can be estimated iteratively. The pair of motion vectors at step k can be
[0116]
Number
[0117] and can be denoted as, the original luminance signal is I(i, j), and the predicted luminance signal is I’ k (i, j). The spatial gradient
[0118]
Number
[0119] can be derived using the Sobel filters applied to the predicted signal I’ k (i, j) in the horizontal and vertical directions respectively. The differential coefficients of Equation 1 can be represented by Equation 13.
[0120]
Number
[0121] In Equation 13, (a, b) can be the delta translation parameter, and (c, d) can be the delta zoom and rotation parameters at step k. The delta MV at the control point can be derived as Equations 14 and 15 using its coordinates. For example, (0, 0) and (w, 0) can be the coordinates for the upper left and upper right control points, respectively.
[0122]
Number
[0123]
Number
[0124] Based on the optical flow equation, the relationship between the change in luminance and spatial gradient and the temporal motion can be formulated as Equation 16.
[0125]
Number
[0126]
Number
[0127] Substituting Equation 13 into it, Equation 17 for the parameters (a, b, c, d) can be generated.
[0128]
Number
[0129] If the samples in the CU satisfy Equation 17, for example, the parameter set (a, b, c, d) can be derived using least squares calculation. The motion vectors at step (k + 1) for two control points
[0130]
Number
[0131] are derived using Equations 14 and 15, and they can be rounded to a specific accuracy (e.g., 1 / 4 pel). Using iteration, the motion vectors at the two control points can be improved until the parameters (a, b, c, d) become zero or the number of iterations meets a predefined limit and converges.
[0132] The six - parameter affine motion model can include two parameters for translation in the horizontal and vertical directions, one parameter for zoom motion in the horizontal direction, one parameter for rotation motion, one parameter for zoom motion in the vertical direction, and / or one parameter for rotation motion. The six - parameter affine motion model can be encoded using three motion vectors at three control points. Figure 10 shows an exemplary six - parameter affine mode. As shown in Figure 10, the three control points for the six - parameter affine - coded CU can be defined at the upper - left, upper - right, and / or lower - left corners of the CU. The motion at the upper - left control point can be related to translational motion. The motion at the upper - right control point can be related to rotational and zoom motions in the horizontal direction. The motion at the lower - left control point can be related to rotational and zoom motions in the vertical direction. In the six - parameter affine motion model, the rotational and zoom motions in the horizontal direction may not be the same as those in the vertical direction. In the example, the motion vectors (v x , v y ) of each sub - block can be derived from Equations 18 and 19 using three motion vectors as control points.
[0133] [Number]
[0134] [Number]
[0135] In Equation 18 and Equation 19, (v 2x , v 2y ) can be the motion vector of the lower left control point. (x, y) can be the center position of the sub-block. w and h can be the width and height of the CU.
[0136] The six parameters of the six-parameter affine model can be estimated, for example, in a similar way. For example, Equation 20 can be created based on Equation 13.
[0137] [Number]
[0138] In Equation 20, for step k, (a, b) can be the delta translation parameter. (c, d) can be the delta zoom and rotation parameter for the horizontal direction. (e, f) can be the delta zoom and rotation parameter for the vertical direction. For example, Equation 21 can be created based on Equation 16.
[0139] [Number]
[0140] The parameter set (a, b, c, d, e, f) can be derived using least squares calculation by considering the samples within the CU. The motion vector of the upper left control point
[0141] [Number]
[0142] can be calculated using Equation 14. The motion vector of the upper right control point
[0143]
Number
[0144] can be calculated using Equation 22. The motion vector of the upper right control point
[0145]
Number
[0146] can be calculated using Equation 23.
[0147]
Number
[0148] There may be a symmetric MV difference with respect to bidirectional prediction. In some examples, the motion vectors in the forward reference picture and the backward reference picture may be symmetric, for example, due to the continuity of the motion trajectory in bidirectional prediction.
[0149] SMVD can be in an inter-coding mode. When using SMVD, the MVD of the first reference picture list (e.g., reference picture list 1) can be symmetric with respect to the MVD of the second reference picture list (e.g., reference picture list 0). Motion vector coding information (e.g., MVD) of one reference picture list can be signaled, and the motion vector information of another reference picture list may not be signaled. The motion vector information of another reference picture list can be determined, for example, based on the signaled motion vector information and based on the fact that the motion vector information of the reference picture lists is symmetric. In the example, the MVD of reference picture list 0 is signaled, and the MVD of list 1 may not be signaled. The MV encoded in this mode can be calculated using Equation 24A.
[0150] [Number]
[0151] In the formula, the subscript indicates reference picture list 0 or 1, x indicates the horizontal direction, and y indicates the vertical direction.
[0152] As shown in Equation 24A, the MVD for the current coding block can indicate the difference between the MVP for the current coding block and the MV for the current coding block. Those skilled in the art will understand that the MVP can be determined based on the MV of spatially adjacent blocks of the current coding block and / or the MV of temporally adjacent blocks of the current coding block. Equation 24A can be shown in FIG. 11. FIG. 11 shows an exemplary non-affine motion symmetry MVD (e.g., MVD = -MVD). As shown in Equation 24A and depicted in FIG. 11, the MV of the current coding block can be equal to the sum of the MVP for the current coding block and the MVD (or negative MVD depending on the reference picture list) of the current coding block. As shown in Equation 24A and depicted in FIG. 11, in the case of SMVD, the MVD (MVD1) of reference picture list 1 can be equal to the negative of the MVD (MVD0) of reference picture list 0. The MV predictor (MVP), (MVP0) of reference picture list 0 may or may not be symmetric with the MVP, (MVP1) of reference picture list 1. MVP0 may or may not be equal to the negative of MVP1. As shown in Equation 24A, the MV of the current coding block can be equal to the sum of the MVP of the current coding block and the MVD of the current coding block. Based on Equation 24A, the MV (MV0) of reference picture list 0 may not be equal to the negative of the MV (MV1) of reference picture list 1. The MV, MV0 of reference picture list 0 may or may not be symmetric with the MV, MV1 of reference picture list 1.
[0153] SMVD can be used for bidirectional prediction when reference picture list 0 includes a forward reference picture and reference picture list 1 includes a backward reference picture, or when reference picture list 0 includes a backward reference picture and reference picture list 1 includes a forward reference picture.
[0154] When using SMVD, the reference picture indices of reference picture lists 0 and 1 may not be signaled. They can be derived as follows. If reference picture list 0 contains forward reference pictures and reference picture list 1 contains backward reference pictures, the reference picture index in list 0 can be set to the forward reference picture closest to the current picture, and the reference picture index in list 1 can be set to the backward reference picture closest to the current picture. If reference picture list 0 contains backward reference pictures and reference picture list 1 contains forward reference pictures, the reference picture index in list 0 can be set to the backward reference picture closest to the current picture, and the reference picture index in list 1 can be set to the forward reference picture closest to the current picture.
[0155] In the case of SMVD, it is not necessary to signal the reference picture index for both lists. For one reference picture list (e.g., list 0), a set of MVDs can be signaled. For bidirectional prediction coding, the signaling overhead can be reduced.
[0156] In merge mode, motion information can be derived and / or used (e.g., directly used) to generate the predicted samples of the current CU. Merge mode using motion vector difference (MMVD) can be used. A merge flag can be signaled to specify whether MMVD is used for a CU. The MMVD flag can be signaled after sending the skip flag.
[0157] In MMVD, after a merge candidate is selected, the merge candidate can be refined by MVD information. The MVD information can be signaled. The MVD information can include one or more of a merge candidate flag, a distance index for specifying the magnitude of motion, and / or an index for indicating the direction of motion. In MMVD, one of a plurality of candidates in the merge list (e.g., the first two candidates) can be selected for use on an MV basis. The merge candidate flag can indicate which candidate is used.
[0158] The distance index can specify motion magnitude information and / or can indicate a predefined offset from the starting point (e.g., from the candidate selected to be MV-based). FIG. 12 shows exemplary motion vector difference (MVD) search points. As shown in FIG. 12, the center point can be the starting point MV. As shown in FIG. 12, the pattern of points can indicate different search orders (e.g., from the point closest to the center MV to those far from the center MV). As shown in FIG. 12, an offset can be added to the horizontal and / or vertical components of the starting point MV. An exemplary relationship between the distance index and the predefined offset is shown in Table 1.
[0159] [Table 1]
[0160] The direction index can represent the direction of the MVD with respect to the starting point. The direction index can represent any one of the four directions shown in Table 2. The meaning of the MVD code can vary according to the information of one or more starting point MVs. When the starting point has a unidirectional prediction MV or a pair of bidirectional prediction MVs where both lists point to the same side of the current picture, the code in Table 2 can specify the sign of the MV offset added to one or more starting MVs. For example, when the picture order counts (POCs) of two references are both greater than the POC of the current picture, or both less than the POC of the current picture, the code can specify the sign of the MV offset added to one or more starting MVs. When the starting point has a pair of bidirectional prediction MVs where both lists point to different sides of the current picture (for example, when the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture, etc.), the code in Table 2 can specify the sign of the MV offset added to the MV component of list 0 of the starting point MVs, and the sign of the MV offset added to the MVs of list 1 can have the opposite value.
[0161]
Table 2
[0162] The symmetric mode can be used for bidirectional prediction coding. One or more of the features described herein can be used in connection with the symmetric mode for bidirectional prediction coding. For example, in some examples, it can increase coding efficiency and / or reduce complexity. The symmetric mode can include SMVD. One or more of the features described herein can be associated with making a synergistic effect of SMVD with one or more other tools, such as bidirectional prediction using CU weights (BCW or BPWA), BDOF, and / or affine mode, etc. One or more of the features described herein can be used in coding (e.g., encoder optimization, etc.), which can include fast motion estimation for translational and / or affine motion.
[0163] The SMVD coding features (functions) can include one or more of restrictions, signaling, SMVD search features (functions), and / or the like.
[0164] The application of the SMVD mode can be based on the CU size. For example, the restriction can be that SMVD is not allowed for relatively small CUs (e.g., CUs having an area not exceeding 64). The restriction can be that SMVD is not allowed for relatively large CUs (e.g., CUs larger than 32×32). When the restriction does not allow SMVD for a CU, the symmetric MVD signaling is skipped or disabled for that CU, and / or the coding device (e.g., encoder) cannot search for the symmetric MVD.
[0165] The application of the SMVD mode can be based on the POC distance between the current picture and the reference picture. The coding efficiency of SMVD may decrease for relatively large POC distances (e.g., POC distances of 8 or more). When the POC distance between a reference picture (e.g., any reference picture) and the current picture is relatively large, SMVD may become inapplicable. When SMVD is inapplicable, symmetric MVD signaling may be skipped or become inapplicable, and / or the coding device (e.g., the encoder) may not search for symmetric MVDs.
[0166] The application of the SMVD mode can be restricted to one or more temporal layers. In an example, a lower temporal layer can refer to a reference picture having a large POC distance from the current picture in a hierarchical GOP structure. The coding efficiency of SMVD may decrease for lower temporal layers. SMVD coding may not be allowed for relatively low temporal layers (e.g., temporal layers 0 and 1). When SMVD is not allowed, symmetric MVD signaling may be skipped or become inapplicable, and / or the coding device (e.g., the encoder) may not search for symmetric MVDs.
[0167] In SMVD coding, one MVD of the reference picture list can be signaled (e.g., explicitly signaled). In an example, the coding device (e.g., the decoder) can parse the first MVD associated with the first reference picture list in the bitstream. The coding device can determine the second MVD associated with the second reference picture list based on the first MVD and determine that the first MVD and the second MVD are symmetric to each other.
[0168] A symbolization device (e.g., a decoder) can identify which reference picture list's MVD is signaled, for example, whether the MVD of reference picture list 0 is signaled or the MVD of reference picture list 1 is sent. In an example, the MVD of reference picture list 0 can be signaled (e.g., always signaled). The MVD of reference picture list 1 can be obtained (e.g., can be derived).
[0169] It is possible to select the reference picture list for which the MVD is signaled (e.g., explicitly signaled). One or more of the following can be applied. An indication (e.g., a flag) can be signaled to indicate which reference picture list is selected. A reference picture list having a smaller POC distance with respect to the current picture can be selected. When the POC distances for the reference picture lists are the same, the reference picture lists can be pre-determined to break the balance. For example, when the POC distances for the reference picture lists are the same, reference picture list 0 can be selected.
[0170] The MVP index for a reference picture list (e.g., one reference picture list) can be signaled. In some examples, the indexes of the MVP candidates for both reference picture lists can be signaled (e.g., explicitly signaled). The MVP index for a reference picture list can be signaled (e.g., only the MVP index for one reference picture list) to reduce signaling overhead, for example. The MVP index for other reference picture lists can be derived as described herein, for example. LX can be the reference picture list for which the MVP index is signaled (e.g., explicitly signaled), and i can be the signaled MVP index. mvp’ can be derived from the MVP of LX as shown in Equation 24.
[0171]
Number
[0172] In the formula, POC LX , POC 1-LX , and POC curr can be the POC of the reference picture list LX reference picture, the list (1-LX) reference picture, and the current picture, respectively. Also, from the MVP list of the reference picture list (1-LX), the MVP closest to mvp' can be selected, for example, as shown in Equation 25.
[0173]
Number
[0174] In the formula, j can be the MVP index of the reference picture list (1-LX). LX can be the reference picture list in which the MVD is signaled (e.g., explicitly signaled).
[0175] Table 3 shows an exemplary CU syntax that can support symmetric MVD signaling for non-affine coding modes.
[0176]
Table 3-1
[0177]
Table 3-2
[0178] For example, an indication such as the sym_mvd_flag flag can indicate whether the SMVD is used in motion vector coding for the current coding block (e.g., a bi-predictive coded CU).
[0179] Indications such as refIdxSymL0 can indicate the reference picture index in reference picture list 0. A refIdxSymL0 indication set to -1 can indicate that SMVD is not applicable and that sym_mvd_flag is not present.
[0180] Indications such as refIdxSymL1 can indicate the reference picture index in reference picture list 1. A refIdxSymL1 indication having a value of -1 can indicate that SMVD is not applicable and that sym_mvd_flag is not present.
[0181] In the example, SMVD may be applicable when reference picture list 0 contains forward reference pictures and reference picture list 1 contains backward reference pictures, or when reference picture list 0 contains backward reference pictures and reference picture list 1 contains forward reference pictures. In other cases, SMVD is not applicable. For example, when SMVD is not applicable, the signaling of the SMVD indication (e.g., the sym_mvd_flag flag at the CU level) may be skipped. The encoding device (e.g., the decoder) can perform one or more condition checks. As shown in Table 3, two condition checks can be performed (e.g., refIdxSymL0 > -1, and refIdxSymL1 > -1). One or more condition checks can be performed to determine whether the SMVD indication is used. As shown in Table 3, two condition checks are performed (e.g., refIdxSymL0 > -1, and refIdxSymL1 > -1) to determine whether the SMVD indication is received. In the example, for the decoder to examine these conditions, the decoder can wait for the reference picture lists (e.g., list 0 and list 1) of the current picture before performing a certain CU syntax analysis. In some examples, even if both of the two examined conditions (e.g., refIdxSymL0 > -1 and refIdxSymL1 > -1) are true, the encoder may not use SMVD for the CU (e.g., to save encoding complexity).
[0182] The SMVD indication is at the CU level and can be associated with the current coding block. The CU-level SMVD indication can be obtained based on the upper-level indication. In the example, the presence of the CU-level SMVD indication (e.g., the sym_mvd_flag) can be controlled by the upper-level indication (e.g., alternatively, or in addition). For example, the SMVD enabled indication such as the sym_mvd_enabled_flag can be signaled at the slice level, tile level, tile group level, or picture parameter set (PPS) level, sequence parameter set (SPS) level, and / or any syntax level, and the reference picture list is shared by the CU associated with the syntax level. For example, the slice-level flag can be placed in the slice header. In the example, the coding device (e.g., decoder) can receive a sequence-level SMVD indication indicating whether SMVD is available (enabled) for the picture sequence. If SMVD is available (enabled) for the sequence, the coding device can obtain the SMVD indication associated with the current coding block based on the sequence-level SMVD indication.
[0183] Using the upper-level SMVD enabled indication such as the sym_mvd_enabled_flag, the CU-level syntax parsing can be performed without examining one or more conditions (e.g., as described herein). SMVD can be made available or unavailable at a level higher than the CU level (e.g., at the discretion of the encoder). Table 4 shows an exemplary syntax that can support the SMVD mode.
[0184]
Table 4-1
[0185]
Table 4-2
[0186] There can be various ways for a symbolization device (e.g., an encoder) to determine whether to make SMVD available (enabled), and accordingly set the value of sym_mvd_enabled_flag. One or more of the examples herein can be combined to determine whether to make SMVD available (enabled) and accordingly set the value of sym_mvd_enabled_flag. For example, the encoder can determine whether to make SMVD available (enabled) by examining the reference picture list. If both a forward reference picture and a backward reference picture exist, the sym_mvd_enabled_flag can be set to be equal to true to make SMVD available. For example, the encoder can determine whether to make SMVD available based on the temporal distance between the current picture and the forward and / or backward reference pictures. When the reference picture is far from the current picture, the SMVD mode may not be available. If the forward or backward reference picture is far from the current picture, the encoder can disable SMVD. The symbolization device can set the value of the upper-level SMVD availability indication based on a value (e.g., a threshold value, etc.) for the temporal distance between the current picture and the reference picture. For example, to reduce the encoding complexity, the encoder can use (e.g., only use) the upper-level control flag to make SMVD available (enabled) for pictures with forward and backward reference pictures, so that the maximum temporal distance between these two reference pictures for the current picture is less than a value (e.g., a threshold value, etc.). For example, the encoder can determine whether to make SMVD available (enabled) based on the temporal layer (e.g., of the current picture). A relatively low temporal layer can indicate that the reference picture is far from the current picture, in which case SMVD may not be available.The encoder can determine that the current picture belongs to a relatively low temporal layer (e.g., lower than thresholds such as 1, 2, etc.), and for such a current picture, it can disable the use of SMVD. For example, the sym_mvd_enabled_flag can be set to false to disable SMVD. For example, the encoder can determine whether SMVD should be made available (enabled) based on the statistics of the previously encoded pictures in the same temporal layer as the current picture. The statistics can include the average POC distance for bidirectionally predicted coded CUs (e.g., the average of the distances between the current picture and the temporal centers of the two reference pictures of the current picture). R0 and R1 can be the reference pictures for bidirectionally predicted coded CUs. poc(x) can be the POC of picture x. The POC distances distance(CU) between the two reference pictures and the current picture can be calculated using Equation 26. i ) can be calculated using Equation 26. The average POC distance AvgDist for bidirectionally predicted coded CUs can be calculated using Equation 27.
[0187] Distance(CUi)=|2 * poc(current_picture)-poc(R0)-poc(R1)| Equation 26
[0188]
Number
[0189] The variable N can indicate the total number of bidirectionally predicted coded CUs that can have both forward and backward reference pictures. As an example, if AvgDist is smaller than a value (e.g., a predefined threshold), the sym_mvd_enabled_flag can be set to true by the encoder to make SMVD available (enabled), and in other cases, the sym_mvd_enabled_flag can be set to false to disable SMVD.
[0190] In some examples, the MVD value may be signaled. In some examples, a combination of a direction index and a distance index may be signaled and the MVD value may not be signaled. The exemplary direction tables and exemplary distance tables shown in Tables 1 and 2 can be used to signal and derive MVD information. For example, distance index 0 and direction index 0 can indicate MVD(1 / 2,0).
[0191] Search for symmetric MVD can be performed, for example, after uni-directional prediction search and bi-directional prediction search. Uni-directional prediction search can be used to search for the optimal MV pointing to the reference block for uni-directional prediction. Bi-directional prediction search can be used to search for two optimal MVs pointing to two reference blocks for bi-directional prediction. The search can be performed to find candidate symmetric MVDs such as the best symmetric MVD. In an example, a set of search points can be repeatedly evaluated for symmetric MVD search. The repetition can include evaluation of the set of search points. The set of search points can form a search pattern centered around the best MV of the previous repetition, for example. For the first repetition, the search pattern can be centered around the first MV. The selection of the first MV can affect the overall result. A set of first MV candidates can be evaluated. The first MV for symmetric MVD search can be determined, for example, based on rate-distortion cost. In an example, the MV candidate having the lowest rate-distortion cost can be selected as the first MV for symmetric MVD search. The rate-distortion cost can be estimated, for example, by summing the bi-directional prediction error of MVD coding for reference picture list 0 and the weighted rate. The set of first MV candidates can include one or more of the MVs obtained from uni-directional prediction search, the MVs obtained from bi-directional prediction search, and the MVs from the advanced motion vector predictor (AMVP) list. At least one MV can be obtained for each reference picture from uni-directional prediction search.
[0192] For example, to reduce complexity, early termination can be applied. Early termination can be applied (e.g., by an encoder) when the cost of bidirectional prediction is greater than a value (e.g., a threshold). In an example, for instance, before the first MV selection, if the rate-distortion cost for the MV obtained from bidirectional prediction search is greater than a value (e.g., a threshold), the search for the symmetric MVD can be terminated. For example, the value can be set to a multiple (e.g., 1.1 times) of the unidirectional prediction cost. In an example, for instance, after the first MV selection, if the rate-distortion cost associated with the first MV is higher than a value (e.g., a threshold), the symmetric MVD search can be terminated. For example, the value can be set to a multiple (e.g., 1.1 times) of the lowest of the unidirectional prediction cost and the bidirectional prediction cost.
[0193] There may be an interaction between the SMVD mode and other coding tools. One or more of the following can be implemented, namely, combining symmetric affine MVD coding, bidirectional prediction using CU weights (BCW or BPWA) with SMVD, or combining BDOF with SMVD.
[0194] Symmetric affine MVD coding can be used. The affine motion model parameters can be represented by the motion vectors of the control points. A 4-parameter affine model can be represented by two control point MVs, and a 6-parameter affine model can be represented by three control point MVs. In the example shown in Equation 12, for example, it can be a 4-parameter affine model represented by two control point MVs, such as the upper left control point MV (v0) and the upper right control point MV (v1). The upper left control point MV can represent the translational motion. The upper left control point MV can have corresponding symmetric MVs associated with the forward and backward reference pictures following the motion trajectory, for example. The SMVD mode can be applied to the upper left control point. The other control point MVs can represent a combination of zoom, rotation, and / or shear mapping. SMVD may not be applied to the other control point MVs.
[0195] SMVD can be applied to the upper left control point (e.g., only to the upper left control point), and the other control point MVs can be set to their respective MV predictors.
[0196] The MVD of the control points associated with the first reference picture list can be signaled. The MVD of the control points associated with the second reference picture can be obtained based on the MVD of the control points associated with the first reference picture list, and the MVD of the control points associated with the first reference picture list is symmetric with respect to the MVD of the control points associated with the second reference picture list. In the example, when the symmetric affine MVD mode is applied, the MVD of the control points associated with reference picture list 0 can be signaled (e.g., only the control point MVD of reference picture list 0 can be signaled). The MVD of the control points associated with reference picture list 1 can be derived based on the symmetric property. The MVD of the control points associated with reference list 1 may not be signaled.
[0197] A control point MV associated with a reference picture list can be derived. The control point MV for control point 0 (upper left) of reference picture list 0 and reference picture list 1 can be derived, for example, using Equation 28.
[0198]
Number
[0199] Equation 28 can be shown in FIG. 13. FIG. 13 shows an exemplary affine motion symmetry MVD. As shown in Equation 28 and illustrated in FIG. 13, the control point MV at the upper left of the current coding block can be equal to the sum of the MVP for the control point at the upper left of the current coding block and the MVD (or a negative MVD depending on the reference picture list) for the control point at the upper left of the current coding block. As shown in Equation 28 and illustrated in FIG. 13, the upper left control point MVD associated with reference picture list 1 can be equal to the negative of the upper left control point MVD associated with reference picture list 0 for symmetric affine MVD coding.
[0200] The MV for other control points can be derived, for example, using at least affine MVP prediction as shown in Equation 29.
[0201]
Number
[0202] In Equations 28 to 30, the first element of the subscript can indicate the reference picture list. The second element of the subscript can indicate the control point index.
[0203] The translational MVD (e.g., mvdx 0,0 , mvdy 0,0 ) of the first reference picture list is for the second reference picture list (e.g., mvdx 1,0 , mvdy 1,0) can be applied to the derivation of the control point MV in the upper left corner. The translational MVD (e.g., mvdx 0,0 , mvdy 0,0 ) of the first reference picture list cannot be applied to the derivation of other control point MVs in the second reference picture list (e.g., mvdx 1,j , mvdy 1,j ). In some examples of symmetric affine MVD derivation, the translational MVD (mvdx 0,0 , mvdy 0,0 ) of reference picture list 0 can be applied to the derivation of the control point MV in the upper left corner of reference picture list 1 (e.g., it can only be applied). The other reference point MVs in list 1 can be the same as the corresponding predictors, for example, as shown in Equation 30.
[0204]
Number
[0205] Table 5 shows an exemplary syntax that can be used to signal the use of SMVD in combination with the affine mode (e.g., symmetric affine MVD coding).
[0206]
Table 5-1
[0207]
Table 5-2
[0208] For example, based on the inter-affine indication and / or the SMVD indication, it can be determined whether affine SMVD is used. For example, when the inter-affine indication inter_affine_flag is 1 and the SMVD indication sym_mvd_flag[x0][y0] is 1, affine SMVD can be applied. When the inter-affine indication inter_affine_flag is 0 and the SMVD indication sym_mvd_flag[x0][y0] is 1, non-affine motion SMVD can be applied. When the SMVD indication sym_mvd_flag[x0][y0] is 0, SMVD cannot be applied.
[0209] The MVD of the control points in the reference picture list (e.g., reference picture list 0) can be signaled. In the example, the MVD values of a subset of the control points in the reference picture list can be signaled. For example, the MVD of the upper left control point in reference picture list 0 can be signaled. Signaling of the MVD of other control points in reference picture list 0 can be skipped, and for example, the MVD of these control points can be set to 0. The MV of other control points can be derived based on the MVD of the upper left control point in reference picture list 0. For example, the MV of a control point for which the MVD is not signaled can be derived as shown in Equation 31.
[0210] [Number]
[0211] Table 6 shows an exemplary syntax that can be used to signal information regarding the SMVD mode combined with the affine mode.
[0212] [Table 6-1]
[0213]
Table 6-2
[0214] Indications such as the flag for only the top-left MVD can indicate whether only the MVD of the top-left control point in the reference picture list (e.g., reference picture list 0) is signaled, or whether the MVD of the control points in the reference picture list is signaled. This indication can be signaled at the CU level. Table 7 shows an exemplary syntax that can be used to signal information regarding the SMVD mode combined with the affine mode.
[0215]
Table 7-1
[0216]
Table 7-2
[0217] In the example, as described herein for the MMVD, a combination of a direction index and a distance index can be signaled. In Tables 1 and 2, an exemplary table of directions and an exemplary table of distances are shown. For example, the combination of direction index 0 and direction index 0 can indicate MVD(1 / 2,0).
[0218] In the example, an indication (e.g., a flag) is signaled to indicate which reference picture list's MVD is being sent. The MVDs of other reference picture lists are not signaled and can be derived.
[0219] One or more of the limitations described herein for translational motion symmetric MVD coding can be applied to affine motion symmetric MVD coding, for example, to reduce complexity and / or reduce signaling overhead.
[0220] Using a symmetric affine MVD can reduce signaling overhead. The coding efficiency can be improved.
[0221] Bidirectional prediction motion estimation can be used to search for symmetric MVDs for the affine model. In the example, after unidirectional prediction search, bidirectional prediction motion estimation can be applied to find a symmetric MVD for the affine model (e.g., the best symmetric MVD for the affine model). The reference pictures for reference picture list 0 and / or reference picture list 1 can be derived as described herein. The first control point MV can be selected from one or more of the results of unidirectional prediction search, the results of bidirectional prediction search, and / or the MVs from the affine AMVP list. The control point MV (e.g., the control point MV with the lowest rate-distortion cost) can be selected to be the first MV. The encoding device (e.g., an encoder, etc.) can examine one or more cases, i.e., in the first case, the symmetric MVD can be signaled for reference picture list 0, and the control point MV for reference picture list 1 can be derived based on symmetric mapping (using equation 28 and / or equation 29); in the second case, the symmetric MVD for reference picture list 1 can be signaled, and the control point MV for reference picture list 0 can be derived based on symmetric mapping. The first case can be used as an example herein. The symmetric MVD search technique can be applied based on the unidirectional prediction search results. When a control point MV predictor is given in reference picture list 1, iterative search using a predefined search pattern (e.g., diamond pattern, cube pattern, and / or the like, etc.) can be applied. In each iteration (e.g., each iteration), the MVD can be refined by the search pattern, and the control point MVs in reference picture list 0 and reference picture list 1 can be derived using equation 28 and equation 29. The bidirectional prediction errors corresponding to the control point MVs in reference picture list 0 and reference picture list 1 can be evaluated. For example, the rate-distortion cost can be estimated by summing the bidirectional prediction error of MVD encoding for reference picture list 0 and the weighted rate.In an example, an MVD having a low (e.g., the lowest) rate-distortion cost during candidate search can be treated as the best MVD for a symmetric MVD search process. The MVs of other control points, such as the upper-right and lower-left control point MVs, can be improved, for example, using the optical flow-based techniques described herein.
[0222] Symmetric MVD search for symmetric affine MVD coding can be performed. In an example, a set of parameters, such as translation parameters, can first be searched, and then the search for non-translation parameters can be performed. In an example, the optical flow search can be performed by considering (e.g., together) the MVDs of reference picture list 0 and reference picture list 1. For the case of a 4-parameter affine model, an exemplary optical flow equation for the MVD of list 0 can be expressed as Equation 32.
[0223]
Number
[0224] In the equation,
[0225]
Number
[0226] can represent the prediction of list 0 at the k-th iteration, and
[0227]
Number
[0228] can represent the spatial gradient of the list 0 prediction.
[0229] Reference picture list 1 can have a translational change (for example, it can have only a translational change). The translational change has the same magnitude as that of reference picture list 0 but is in the opposite direction, which is the condition for a symmetric affine MVD. The optical flow equation for reference picture list 1 MVD can be expressed by Equation 33.
[0230] [Number]
[0231] BCW weight w 0 and w 1 can be applied to list 0 prediction and list 1 prediction respectively. An exemplary optical equation for a symmetric affine model can be shown by Equation 34. I k ’(i,j) - I(i,j) = (G x (i,j) · i + G y (i,j) · j) · c + (-G x (i,j) · j + G y (i,j) · i) · d + H x (i,j) · a + H y (i,j) · b Equation 34 In the equation,
[0232] [Number]
[0233] is.
[0234] The parameters a, b, c, and d can be estimated (for example, by mean least square error calculation).
[0235] I k ’(i,j) - I(i,j) = G x (i,j) · i · c + G x (i,j) · j · d + G y (i,j) · j · e + G y (i,j) · j · f + Hx (i,j)·a + H y (i,j)·b Equation 35 The parameters a, b, c, d, e, and f can be estimated by calculating the average least square error. When the joint optical flow search is performed, the affine parameters can be optimized together. The performance can be improved.
[0236] For example, in order to reduce complexity, truncation can be applied. In the example, for example, before making the first MV selection, if the bi - directional prediction cost becomes larger than a value (for example, a threshold value), the search can be terminated. For example, the value can be set to a multiple of the uni - directional prediction cost, for example, 1.1 times the uni - directional prediction cost. In the example, the encoding device (for example, an encoder, etc.) can compare the current best affine motion estimation (ME) cost with the non - affine ME cost (considering uni - directional prediction and bi - directional prediction affine search) before the ME of the symmetric affine MVD starts. If the current best affine ME cost is larger than the non - affine ME cost multiplied by a value (for example, a threshold value such as 1.1), the encoding device can skip the ME of the symmetric affine MVD. In the example, for example, after the first MV selection, the affine symmetric MVD search can be skipped if the cost of the first MV is higher than a value (for example, a threshold value). For example, the value can be set to a multiple of the lowest of the uni - directional prediction cost and the bi - directional prediction cost (for example, set to 1.1 times). In the example, the value can be set to a multiple (for example, 1.1) of the non - affine ME cost.
[0237] SMVD can be combined with BCW. If BCW is available (enabled) for the current CU, SMVD can be applied in one or more ways. In some examples, SMVD can be made available when the weights of BCW are equal weights (e.g., 0.5, etc.) (e.g., only at that time), and for other BCW weights, SMVD may not be available. In such cases, the SMVD flag is signaled before the BCW weight index, and the signaling of the BCW weight index can be conditionally controlled by the SMVD flag. The SMVD indication (e.g., the SMVD flag, etc.) can have a value of 1, and the signaling of the BCW weight index can be skipped. The decoder can assume that the SMVD indication has a value of 0, which can correspond to equal weights for bidirectional prediction averaging. When the SMVD flag is 0, the BCW weight index can be encoded for the bidirectional prediction mode. In examples where SMVD is available when the BCW weights are equal weights and not available for other BCW weights, the BCW index may be skipped. In some examples, SMVD can be fully combined with BCW. The SMVD flag and the BCW weight index can be signaled for the explicit bidirectional prediction mode. The MVD search for SMVD (e.g., of the encoder) can consider the BCW weight index during bidirectional prediction averaging. The SMVD search can be based on the evaluation of one or more (e.g., all) possible BCW weights.
[0238] A symbolization tool (e.g., bidirectional optical flow (BDOF)) can be used in association with one or more other symbolization tools / modes. BDOF can be used in association with SMVD. Whether BDOF is applied to an encoding block may depend on whether SMVD is used. SMVD can be based on the assumption of symmetric MVD at the encoding block level. When implemented, BDOF can be used to improve sub-block MVs based on optical flow. The optical flow can be based on the assumption of symmetric MVD at the sub-block level.
[0239] BDOF can be used in association with SMVD. In an example, an encoding device (e.g., a decoder or an encoder) can receive one or more indications that SMVD and / or BDOF are available. BDOF can become available for the current picture. The encoding device can determine whether to bypass or implement BDOF for the current encoding block. The encoding device can determine whether to bypass BDOF based on an SMVD indication (e.g., sym_mvd_flag[x0][y0]). In some examples, BDOF can be used interchangeably with BIO.
[0240] The encoding device can determine whether to bypass BDOF for the current encoding block. BDOF can be bypassed for the current encoding block, for example, when the SMVD mode is used for the motion vector encoding of the current encoding block to reduce the complexity of decoding. If SMVD is not used for the motion vector encoding of the current encoding block, the encoding device can determine whether to enable BDOF for the current encoding block based on, for example, at least another condition.
[0241] The symbolization device can obtain the SMVD indication (e.g., sym_mvd_flag[x0][y0]). The SMVD indication can indicate whether SMVD is used in the motion vector for the current coding block.
[0242] The current coding block can be reconstructed based on a determination of whether to bypass BDOF. The MVD can be signaled at the CU level for the SMVD mode (e.g., explicitly signaled).
[0243] The symbolization device can be configured to perform motion vector coding using SMVD without using BDOF, based on a determination to bypass BDOF.
[0244] Features and elements are described above in a particular combination, but those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein can be implemented by a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections), and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. 1. A device for video encoding, comprising: obtaining, for a video block, a first motion vector differential (first MVD) associated with a reference picture for the video block; determining a second MVD based on the first MVD associated with the reference picture, the second MVD being symmetric with respect to the first MVD; obtaining a first predictive block of the video block based on the first MVD and a second predictive block of the video block based on the second MVD; determining a first weight and a second weight for bi-prediction with a CU weight for the video block; applying the first weighting to the first prediction block and the second weighting to the second prediction block to obtain another prediction block; Encoding the video block based on the another predicted block. Processor configured for A device with.
2. The processor, determining a first motion vector predictor (first MVP) associated with the video block and a second MVP associated with the video block; determining a first motion vector (MV) based on the first MVD and the first MVP associated with the video block, the first predictive block being determined based on the first MV of the video block; determining a second MV based on the second MVD and the second MVP associated with the video block, and the second predictive block is determined based on the second MV of the video block. The device of claim 1 further configured as follows:
3. 3. The device of claim 1 or 2, wherein the reference picture is associated with a first reference picture list for the video block, and the processor is further configured to determine the reference picture of the first reference picture list for the video block based on a picture order count (POC) of the reference picture and a POC of a current picture that includes the video block.
4. The processor, determining the second weight based on the first weight; Encode a weight index indicating the first weight.
4. The device of claim 1 further configured to:
5. The processor, determining, based on the determination that SMVD is enabled for the video block, to bypass bidirectional optical flow (BDOF) for the video block; Encoding the video block based on the determination to bypass the BDOF for the video block.
5. The device of claim 1 further configured to:
6. The processor, determining an initial motion vector (Initial MV) using an Initial MV candidate list, the Initial MV candidate list including an MV obtained using a unidirectional predictive search, an MV obtained using a bidirectional predictive search, and a plurality of MVs obtained using an Advanced Motion Vector Predictor (AMVP) list; performing an MVD search using said initial MV; determining, based on the MVD search, that the first MVD and the second MVD will be used to encode the video block; 6. The device of claim 1 further configured to:
7. The processor, performing said bidirectional predictive search; determining that a rate-distortion (RD) cost associated with the bidirectional predictive search is greater than a RD threshold; Terminating the bi-directional predictive search based on the determination that a RD cost associated with the bi-directional predictive search is greater than the RD threshold. The device of claim 6 further configured as follows:
8. 1. A device for video decoding, comprising: obtaining, for a video block, a first motion vector differential (first MVD) associated with a reference picture for the video block; determining a second MVD based on the first MVD associated with the reference picture, the second MVD being symmetric with respect to the first MVD; obtaining a first predictive block of the video block based on the first MVD and a second predictive block of the video block based on the second MVD; determining a first weight and a second weight for bi-prediction with a CU weight for the video block; applying the first weighting to the first prediction block and the second weighting to the second prediction block to obtain another prediction block; Decoding the video block based on the another predicted block. Processor configured for A device with.
9. The processor, 10. The device of claim 8, further configured to determine, based on an SMVD indicator indicating that SMVD is enabled for the video block, to bypass bidirectional optical flow (BDOF) for the video block, wherein the video block is decoded based on the determination to bypass BDOF for the video block.
10. The processor, determining a first motion vector predictor (first MVP) associated with the video block and a second MVP associated with the video block; determining a first motion vector (MV) based on the first MVD and the first MVP associated with the video block, the first predictive block being determined based on the first MV of the video block; determining a second MV based on the second MVD and the second MVP associated with the video block, and the second predictive block is determined based on the second MV of the video block.
10. A device according to claim 8 or 9.
11. The processor, determining the first weight for the video block based on a weight index; 11. The device of claim 8, wherein the second weight is determined based on the first weight.
12. 12. The device of claim 8, wherein the reference picture is associated with a first reference picture list for the video block, and the processor is further configured to determine the reference picture of the first reference picture list for the video block based on a picture order count (POC) of the reference picture and a POC of a current picture that includes the video block.
13. The device of claim 8 , wherein a symmetric MVD flag and a weight index are both signaled for an explicit bi-predictive mode.
14. 1. A method of encoding, comprising: obtaining, for a video block, a first motion vector differential (first MVD) associated with a reference picture for the video block; determining a second MVD based on the first MVD associated with the reference picture, the second MVD being symmetric with respect to the first MVD; obtaining a first predictive block of the video block based on the first MVD and a second predictive block of the video block based on the second MVD; determining a first weight and a second weight for bi-prediction at a CU weight for the video block; applying the first weighting to the first prediction block and the second weighting to the second prediction block to obtain another prediction block; encoding the video block based on the another predicted block; A method for providing the above.
15. determining the second weight based on the first weight; encoding a weight index indicative of the first weight; The method of claim 14 further comprising:
16. 1. A method for video decoding, comprising: obtaining, for a video block, a first motion vector differential (first MVD) associated with a reference picture for the video block; determining a second MVD based on the first MVD associated with the reference picture, the second MVD being symmetric with respect to the first MVD; obtaining a first predictive block of the video block based on the first MVD and a second predictive block of the video block based on the second MVD; determining a first weight and a second weight for bi-prediction at a CU weight for the video block; applying the first weighting to the first prediction block and the second weighting to the second prediction block to obtain another prediction block; decoding the video block based on the another predicted block; A method for providing the above.
17. 17. The method of claim 16, wherein a symmetric MVD flag and a weight index are both signaled for an explicit bi-predictive mode.
18. determining the first weight for the video block based on a weight index; determining the second weight based on the first weight; 18. The method of claim 16 or 17, further comprising:
Citation Information
Patent Citations
Bi-directional optical flow for video coding
US20170094305A1
Method and apparatus of motion compensation for video coding based on bi prediction optical flow techniques
US20180249172A1
Method and device for image decoding according to inter-prediction in image coding system
US20200314444A1
Systems and methods for generalized multi-hypothesis prediction for video coding
WO2017197146A1
Syntax reuse for affine mode with adaptive motion vector resolution
WO2020058890A1