Sub-block motion derivation and decoder-side motion vector refinement for merge mode
By identifying collocated pictures and refining motion vectors for sub-blocks in merge mode, the system enhances video coding efficiency and decoding performance.
Patent Information
- Application Number
- JP2025135257
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-06-27
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video coding systems face challenges in efficiently deriving motion vectors for sub-blocks and refining them in merge mode, leading to suboptimal compression and decoding performance.
The system identifies a collocated picture based on a slice header, selects candidate neighboring CUs with matching reference pictures, and performs temporal scaling and block validation to refine motion vectors for sub-blocks, enhancing the coding process.
This approach improves the accuracy and efficiency of motion vector derivation and refinement for sub-blocks, resulting in better video compression and decoding outcomes.
Smart Images

Figure 2025170304000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to sub-block motion derivation and decoder-side motion vector refinement for merge mode. [Background technology]
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 580184, filed November 1, 2017, U.S. Provisional Patent Application No. 62 / 623001, filed January 29, 2018, U.S. Provisional Patent Application No. 62 / 678576, filed May 31, 2018, and U.S. Provisional Patent Application No. 62 / 690661, filed June 27, 2018, the contents of which are incorporated herein by reference.
[0003] Video coding systems may be used to compress digital video signals, e.g., to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems may include block-based, wavelet-based, and / or object-based systems. Block-based hybrid video coding systems may be deployed. Summary of the Invention [Problem to be solved by the invention]
[0004] SUMMARY Disclosed herein are systems, methods, and means for sub-block motion derivation and motion vector refinement for merge mode. [Means for solving the problem]
[0005] Systems, methods, and means for sub-block motion derivation and motion vector refinement for merge mode may be disclosed herein. Video data may be coded (e.g., encoded and / or decoded). A collocated picture for a current slice of the video data may be identified. The collocated picture may be identified, for example, based on a collocated picture indication in a slice header. The current slice may include one or more coding units (CUs). One or more neighboring CUs may be identified for the current CU. The neighboring CUs (e.g., each neighboring CU) may correspond to a reference picture. A neighboring CU (e.g., one) may be selected to be a candidate neighboring CU based on the neighboring CU's reference picture and the collocated picture. A motion vector (MV) (e.g., a collocated MV) may be identified from the collocated picture based on the candidate neighboring CU's MV (e.g., a reference MV). The collocated MV may be a temporal MV, and the reference MV may be a spatial MV. The current CU may be coded (e.g., encoded and / or decoded) using the collocated MV.
[0006] Neighboring CUs may be selected to be candidate neighboring CUs based on the respective temporal differences between the neighboring CU's reference picture and the co-located picture. For example, a reference picture (e.g., each reference picture) may be associated with a Picture Order Count (POC), and a neighboring CU with the smallest POC difference from the co-located picture may be selected. The selected neighboring CU may have the same reference picture as the co-located picture. A neighboring CU with the same reference picture as the co-located picture may be selected without further consideration of other neighboring CUs.
[0007] For example, if the reference picture of the candidate neighboring CU is not the same as the co-located picture, temporal scaling may be performed on the reference MV. For example, the reference MV may be multiplied by a scaling factor. The scaling factor may be based on the temporal difference between the reference picture of the candidate neighboring CU and the co-located picture.
[0008] A collocated picture may include one or more collocated blocks. One or more of the collocated blocks may be valid collocated blocks. The valid collocated blocks may be contiguous and may form a valid collocated block region. The region may be identified, for example, based on the current slice. A collocated MV may be associated with a first collocated block, which may or may not be valid. If the first collocated block is not valid, a valid second collocated block may be selected. The collocated MV from the first collocated block may be replaced with a second collocated MV associated with the second collocated block. The second collocated MV may be used to code (e.g., encode and / or decode) the current CU. The second co-located block may be selected, for example, based on the second co-located block having the smallest distance to the first co-located block. For example, the second co-located block may be the closest valid block to the first co-located block.
[0009] The current CU may be subdivided into one or more sub-blocks. The sub-blocks (e.g., each sub-block) may correspond to a reference MV. Based on the reference MV for the sub-block, a co-located MV may be identified from a co-located picture for the sub-block (e.g., each sub-block). The size of the sub-block may be determined based on the temporal layer of the current CU. [Effects of the Invention]
[0010] A system, method, and means are provided for sub-block motion derivation and motion vector refinement for a novel merge mode. [Brief explanation of the drawings]
[0011] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system. [Figure 1B] 1B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communication system of FIG. 1A. [Figure 1C] FIG. 1B is a system diagram illustrating an example radio access network (RAN) and core network (CN) that may be used within the communication system of FIG. 1A. [Figure 1D] FIG. 1B is a system diagram illustrating a further example RAN and CN that may be used within the communication system of FIG. 1A. [Figure 2] FIG. 1 is an exemplary diagram of a block-based video encoder. [Figure 3] FIG. 2 is an exemplary block diagram of a video decoder. [Figure 4] FIG. 1 illustrates exemplary spatial merge candidates. [Figure 5] FIG. 1 illustrates an example of enhanced temporal motion vector prediction. [Figure 6] FIG. 1 illustrates an example of spatial-temporal motion vector prediction. [Figure 7] FIG. 10 illustrates an exemplary decoder-side motion vector refinement (DMVR) for normal merge mode. [Figure 8] A diagram showing an exemplary refreshing of a picture in which enhanced temporal motion vector prediction / spatial-temporal motion vector prediction block size statistics are reset to 0 when adaptively determining ATMVP / STMVP derivation granularity. [Figure 9] 10 is a flowchart of an example motion compensation for merge mode when DMVR early truncation is applied. [Figure 10] FIG. 1 illustrates an exemplary bi-prediction where the average of two prediction signals is calculated at an intermediate bit depth. [Figure 11] FIG. 10 illustrates an example of collocated block derivation for ATMVP. [Figure 12] A diagram showing unrestricted access of co-located blocks in ATMVP. [Figure 13] FIG. 10 illustrates a restricted region for deriving collocated blocks for an ATMVP coding unit. [Figure 14] FIG. 10 illustrates an example of collocated block derivation for ATMVP. [Figure 15] FIG. 10 illustrates an example of deriving an MV for a current block using a co-located picture. DETAILED DESCRIPTION OF THE INVENTION
[0012] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which:
[0013] 1A is a diagram illustrating an example communication system 100 in which one or more disclosed examples may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multicarrier (FBMC).
[0014] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed example may contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which may be referred to as a “station” and / or “STA,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspot or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain situations), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks. Any of the WTRUs 102a, 102b, 102c, 102d may be referred to interchangeably as a UE.
[0015] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNodeB, a Home Node B, a Home eNodeB, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each depicted as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0016] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, sometimes referred to as a cell (not shown). These frequencies may be in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrum. A cell may provide coverage for a wireless service in a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in an example, the base station 114a may include three transceivers, one for each sector of the cell. In an example, the base station 114a may utilize multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.
[0017] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0018] More specifically, as mentioned above, the communication system 100 may be a multiple-access system and may utilize one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a and the WTRUs 102a, 102b, and 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed Uplink (UL) Packet Access (HSUPA).
[0019] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE), and / or LTE Advanced (LTE-A), and / or LTE Advanced Pro (LTE-A Pro).
[0020] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.
[0021] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0022] In examples, the base station 114a and the WTRUs 102a, 102b, 102c may implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0023] 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a localized area such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by drones), and a roadway. In an example, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an example, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In an example, the base station 114b and the WTRUs 102c, 102d may establish a picocell or a femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in FIG. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 through the CN 106 / 115.
[0024] The RAN 104 / 113 may communicate with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput, delay, error resilience, reliability, data throughput, and mobility requirements. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 may communicate directly or indirectly with other RANs that utilize the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0025] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs that may utilize the same RAT as the RAN 104 / 113 or a different RAT.
[0026] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that may utilize cellular-based wireless technology and with a base station 114b that may utilize IEEE 802 wireless technology.
[0027] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any subcombination of the above elements.
[0028] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0029] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in an example, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an example, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In an example, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0030] 1B, the transmit / receive element 122 is depicted as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Thus, in an example, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0031] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.
[0032] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may obtain information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In examples, the processor 118 may obtain information from and store data in memory that is not physically located on the WTRU 102, such as located on a server or home computer (not shown).
[0033] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control the power to other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0034] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information over the air interface 116 from base stations (e.g., base stations 114a, 114b) and / or may determine its location based on the timing of signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 may acquire location information using any suitable location-determination method.
[0035] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0036] The WTRU 102 may include a full-duplex radio where transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and the downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit 139 to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing via a processor (e.g., a separate processor (not shown) or the processor 118). In an example, the WTRU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or the downlink (e.g., for reception)).
[0037] 1C is a system diagram illustrating an example RAN 104 and CN 106. As mentioned above, the RAN 104 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 utilizing E-UTRA radio technology. The RAN 104 may also communicate with the CN 106.
[0038] The RAN 104 may include eNodeBs 160a, 160b, 160c, although it will be understood that the RAN 104 may include any number of eNodeBs. The eNodeBs 160a, 160b, 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In examples, the eNodeBs 160a, 160b, 160c may implement MIMO technology. Thus, the eNodeB 160a, for example, may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.
[0039] Each of the eNodeBs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 1C, the eNodeBs 160a, 160b, 160c may communicate with one another over an X2 interface.
[0040] The CN 106 shown in FIG. 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166; although each of the above elements is depicted as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.
[0041] The MME 162 may be connected to each of the eNodeBs 160a, 160b, 160c in the RAN 104 via an S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.
[0042] The SGW 164 may be connected to each of the eNodeBs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during inter-eNodeB handover, triggering paging when DL data is available to the WTRUs 102a, 102b, 102c, and managing and storing the context of the WTRUs 102a, 102b, 102c.
[0043] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0044] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional landline communication devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. Additionally, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0045] Although in Figures 1A-1D the WTRU is described as a wireless terminal, it is contemplated that in some examples such a terminal may use a wired communication interface (e.g., temporary or permanent) with the communication network.
[0046] In an example, the other network 112 may be a WLAN.
[0047] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic originating from outside the BSS to a STA may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP for delivery to its respective destination. Traffic between STAs within the BSS may be sent through the AP; for example, a source STA may send traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent (e.g., directly) between a source STA and a destination STA using direct link setup (DLS). In examples, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all of the STAs) may communicate directly with each other. IBSS mode communication is sometimes referred to herein as "ad hoc" mode communication.
[0048] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In an example, for example, in an 802.11 system, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented. With CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. Within a given BSS, one STA (e.g., only one station) may transmit at any given time.
[0049] High-throughput (HT) STAs may use 40 MHz wide channels for communication, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.
[0050] A very high throughput (VHT) STA may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. A 40 MHz and / or 80 MHz channel may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining eight contiguous 20 MHz channels or two non-contiguous 80 MHz channels, which may be referred to as an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data may pass through a segment parser that may split the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing may be performed separately for each stream. The streams may be mapped onto two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration may be reversed, and the combined data may be transmitted to the medium access control (MAC).
[0051] Sub-1 GHz mode operation is supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. Following an example, 802.11ah may support meter-type control / machine-type communication, such as MTC devices in macro coverage areas. MTC devices may have limited functionality, including, for example, support for a certain bandwidth and / or limited bandwidths (e.g., only support for those bandwidths). MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).
[0052] WLAN systems, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, that may support multiple channels and channel bandwidths include a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by a STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the example of 802.11ah, for a STA (e.g., an MTC-type device) that supports (e.g., only supports) the 1 MHz mode, the primary channel may be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. For example, if the primary channel is busy because a STA (that only supports 1 MHz operating mode) is transmitting to the AP, the entire available frequency band may be considered busy, even though most of the frequency band may remain idle and available for use.
[0053] In the United States, the available frequency band that can be used by 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz, depending on country regulations.
[0054] 1D is a system diagram illustrating an example RAN 113 and CN 115. As mentioned above, the RAN 113 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 utilizing NR radio technology. The RAN 113 may also communicate with the CN 115.
[0055] The RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that the RAN 113 may include any number of gNBs. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In examples, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNB 180a, 180b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, and 180c. Thus, the gNB 180a may, for example, transmit wireless signals to and / or receive wireless signals from the WTRU 102a using multiple antennas. In an example, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on an unlicensed spectrum, while the remaining component carriers may be on a licensed spectrum. In an example, the gNBs 180a, 180b, and 180c may implement coordinated multipoint (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).
[0056] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., including different numbers of OFDM symbols and / or lasting for different lengths of absolute time).
[0057] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNodeBs 160a, 160b, 160c). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect to a gNB 180a, 180b, 180c while also communicating / connecting to another RAN, such as the eNodeBs 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement the DC principle to communicate with one or more gNBs 180a, 180b, 180c and one or more eNodeBs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNodeBs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.
[0058] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a, 184b, and routing of control plane information to Access and Mobility Management Functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D, the gNBs 180a, 180b, 180c may communicate with one another over the Xn interface.
[0059] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the above elements is depicted as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.
[0060] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize CN support for the WTRUs 102a, 102b, 102c based on the type of service utilized by the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases, such as services relying on Ultra-Reliable Low-Latency (URLLC) access, services relying on eMBB access, and / or services for Machine-Type Communication (MTC) access, etc. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies like WiFi.
[0061] The SMFs 183a, 183b may be connected to the AMFs 182a, 182b in the CN 115 via an N11 interface. The SMFs 183a, 183b may also be connected to the UPFs 184a, 184b in the CN 115 via an N4 interface. The SMFs 183a, 183b may select and control the UPFs 184a, 184b and configure the routing of traffic through the UPFs 184a, 184b. The SMFs 183a, 183b may perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notification. PDU session types may be IP-based, non-IP-based, Ethernet-based, etc.
[0062] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184a, 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multihoming PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0063] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In an example, the WTRUs 102a, 102b, 102c may be connected to local data networks (DNs) 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.
[0064] 1A-1D and the corresponding description thereof, one or more or all of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functionality.
[0065] The emulation device may be designed to perform one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation device may be directly coupled to another device for testing purposes and / or may perform tests using over-the-air wireless communication.
[0066] The one or more emulation devices may perform one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in an undeployed (e.g., test) wired and / or wireless communication network to perform tests of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, for example, one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0067] A block-based hybrid video coding framework may be provided. FIG. 2 provides a block diagram of an exemplary block-based hybrid video encoding framework. An input video signal 2 may be processed block by block. A block size (e.g., an extended block size, such as a coding unit (CU)) may compress a high-resolution (e.g., 1080p and above) video signal. For example, a CU may include 64×64 pixels. The CU may be partitioned into prediction units (PUs), for which separate predictions may be used. Spatial prediction 60 and / or temporal prediction 62 may be performed on (e.g., each) input video block (e.g., MB and / or CU). Spatial prediction (e.g., intra prediction) may use pixels from samples (e.g., reference samples) of coded neighboring blocks within a video picture / slice to predict a current video block. Spatial prediction may reduce spatial redundancy, for example, that may be inherent in a video signal. Temporal prediction (inter-prediction and / or motion-compensated prediction), for example, may use reconstructed pixels from an encoded video picture to predict a current video block. Temporal prediction may reduce temporal redundancy, for example, that may be inherent in a video signal. A temporal prediction signal for a video block may be signaled by one or more motion vectors (MVs). The MVs may indicate the amount and / or direction of motion between the current block and its reference block. If multiple reference pictures are supported for (e.g., each) video block, a reference picture index for the video block may be transmitted. The reference index may be used to identify which reference picture in reference picture store 64 the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block 80 within the encoder may determine a prediction mode (e.g., a best prediction mode), for example, based on rate-distortion optimization. The prediction block may be subtracted from the current video block at 16. The prediction residual may be decorrelated using transform 4 and / or quantization 6.The quantized residual coefficients may be dequantized at 10 and / or inverse transformed at 12, e.g., to form a reconstructed residual. The reconstructed residual may be added to a predictive block at 26, e.g., to form a reconstructed video block. Before the reconstructed video block is placed in reference picture store 64 and / or used to encode a video block (e.g., a future video block), an in-loop filter (e.g., a deblocking filter and / or an adaptive loop filter) may be applied to the reconstructed video block at 66. To form output video bitstream 20, the coding mode (e.g., inter or intra), prediction mode information, motion information, and / or the quantized residual coefficients (e.g., all) may be transmitted to entropy coding unit 8, e.g., compressed and / or packed to form the bitstream.
[0068] 3 shows a block diagram of an example block-based video decoder. A video bitstream 202 may be unpacked (e.g., first unpacked) and / or entropy decoded in an entropy decoding unit 208. To form a predictive block, coding mode and prediction information may be sent to a spatial prediction unit 260 (e.g., for intra-coding) and / or a temporal prediction unit 262 (e.g., for inter-coding). For example, to reconstruct a residual block, residual transform coefficients may be sent to an inverse quantization unit 210 and / or an inverse transform unit 212. The predictive block and residual block may be added together at 226. For example, the reconstructed block may pass through in-loop filtering before being stored in a reference picture store 264. The reconstructed video in the reference picture store may be transmitted to drive a display device and / or used to predict video blocks (e.g., future video blocks).
[0069] In motion-compensated prediction, for (e.g., each) inter-coded block, motion information (e.g., a motion vector (MV) and a reference picture index) may be used to track a corresponding matching block in a corresponding reference picture, which may be synchronized by, for example, an encoder and / or a decoder. Two modes (e.g., a merge mode and a non-merge mode) may be used to encode the motion information of an inter block. When a block is encoded by a non-merge mode, the MV may be encoded (e.g., differentially encoded) using an MV predictor. The difference between the MV and the MV predictor may be transmitted to the decoder. For (e.g., each) block encoded by a merge mode, the motion information of the block may be derived from spatial and / or temporal neighboring blocks. For example, a contention-based scheme may be applied to select a neighboring block (e.g., a best neighboring block) from available candidates. At the decoder, an index (e.g., just an index) of the candidate (e.g., the best candidate) may be transmitted to re-establish the motion information (e.g., the same motion information).
[0070] A merge mode may be implemented. Candidates (e.g., a set of candidates) in the merge mode may consist of one or more spatial neighboring candidates, e.g., temporal neighboring candidates, and one or more generated candidates. Figure 4 shows an example location of spatial candidates. To construct a list of merge candidates, spatial candidates may be checked and / or added to the list in the order of, e.g., A1, B1, B0, A0, and B2. If a spatially located block is intra-coded and / or outside the boundary of the current slice, the block may be unavailable. To reduce redundancy of spatial candidates, redundant entries (e.g., entries in which a candidate has the same motion information as an existing candidate) may be removed from the list. After including valid spatial candidates, temporal candidates may be generated from the motion information of co-located blocks in co-located reference pictures by temporal motion vector prediction (TMVP). The size (e.g., N) of the merge candidate list may be configured. For example, N may be 5. If the number of merge candidates (e.g., including spatial and temporal candidates) is greater than N, the first N-1 spatial candidates and / or temporal candidates may be kept in a list. For example, the first N-1 spatial candidates (e.g., only) and the temporal candidates may be kept in a list. If the number of merge candidates is less than N, for example, one or more candidates (e.g., the combined candidate and zero candidate) may be added to the candidate list until the number of merge candidates reaches N.
[0071] One or more candidates may be included in the merge candidate list. For example, the merge candidate list may include five spatial candidates as shown in Figure 4 and a TMVP candidate. One or more aspects of the motion derivation for the merge mode may be modified, for example, to include sub-block-based motion derivation and decoder-side motion vector refinement.
[0072] Sub-block-based motion derivation for the merge mode may be performed. (E.g., each) merge block may include a set of motion parameters (e.g., a motion vector and a reference picture index) for (e.g., each) prediction direction. One or more (e.g., two) merge candidates may be included in the merge candidate list, which may enable derivation of motion information at the sub-block level. Including merge candidates with sub-block-level motion information in the merge candidate list may be achieved by increasing the maximum size (e.g., N) of the merge candidate list, for example, from 5 to 7. When one or more of the candidates are selected, the encoder / decoder may divide the CU (e.g., the current CU) into 4x4 sub-blocks and derive motion information for (e.g., each) sub-block. Advanced temporal motion vector prediction (ATMVP) may be used. For example, ATMVP may divide the CU into sub-blocks (e.g., 4x4 sub-blocks). ATMVP may be built on TMVP and may enable a CU to obtain motion information of its sub-blocks from multiple small blocks belonging to temporal neighboring pictures (e.g., co-located reference pictures) of the current picture. In spatial-temporal motion vector prediction (STMVP), the motion parameters of a sub-block may be derived (e.g., recursively) by averaging the motion vectors of its temporal neighbors with the motion vectors of its spatial neighbors.
[0073] Advanced temporal motion vector prediction may be performed. ATMVP may enable a block to derive multiple motion information (e.g., motion vectors and / or reference indices) for sub-blocks within the block from multiple smaller blocks in temporal neighboring pictures of the current picture. For example, ATMVP may derive motion information for sub-blocks of a block as follows: Corresponding blocks (e.g., collocated blocks) of the current block may be identified in a temporal reference picture. The selected temporal reference picture may be a collocated picture. The current block may be divided into sub-blocks, and motion information for (e.g., each) sub-block may be derived from a corresponding smaller block in a collocated picture, as shown in FIG. 5.
[0074] The co-located block and / or co-located picture may be identified by motion information of spatial neighboring blocks of the current block. Available candidates (e.g., the first available candidate) in a merge candidate list may be considered. FIG. 5 shows an example of available candidates in a merge candidate list being considered. Block A may be identified as an available (e.g., the first available) merge candidate for a block (e.g., the current block), for example, based on the scanning order of the merge candidate list. To identify the co-located picture and / or co-located block, a motion vector (e.g., a corresponding motion vector) (e.g., MVA) of block A and / or a reference index of block A may be used. The location of the co-located block in the co-located picture may be determined by adding the motion vector (e.g., MVA) of block A to the coordinates of the current block.
[0075] For (e.g., each) sub-block in a block (e.g., the current block), motion information of the sub-block's corresponding small block in the co-located block (e.g., as indicated by an arrow in FIG. 5) may be used to derive motion information for the sub-block. For example, after motion information for (e.g., each) small block in the co-located block is identified, the motion information of the small block may be converted into a motion vector and / or reference index of the corresponding sub-block in the current block, e.g., in the same manner as temporal motion vector prediction (TMVP), in which temporal motion vector scaling may be applied.
[0076] Spatial-Temporal Motion Vector Prediction (STMVP) may be performed. In STMVP, motion information of sub-blocks within (e.g., one) coding block may be derived in a recursive manner. FIG. 6 shows an example of STMVP. For example, a current block may include one or more (e.g., four) sub-blocks, e.g., A, B, C, and D. Neighboring small blocks that are spatial neighbors to the current block may be labeled a, b, c, and d, respectively. Motion derivation for sub-block A may start by identifying the spatial neighbors (e.g., two spatial neighbors) of block A. The neighbor (e.g., the first neighbor) of block A may be the upper neighbor c. If the small block c is not available or is not intra-coded, the upper neighboring small blocks of the current block may be checked, for example, sequentially (e.g., from left to right). The neighbor (e.g., the second neighbor) of sub-block A may be the left neighbor b. If small block b is not available or is not intra-coded, neighboring small blocks to the left of the current block may be checked, for example, in order (e.g., from top to bottom). After obtaining motion information of spatial neighbors, motion information of temporal neighbors of sub-block A may be obtained in a manner similar to (e.g., the same as) TMVP. Motion information (e.g., all motion information) of available spatial and / or temporal neighbors (adjacent ones) (e.g., up to three) may be averaged and / or used as the motion information of sub-block A. For example, STMVP may be repeated based on raster scan order to derive motion information of sub-blocks (e.g., all other sub-blocks) within the current video block.
[0077] Decoder-side motion vector refinement for normal merge candidates may be performed. In merge mode, when a selected merge candidate is bi-predicted, for example, a prediction signal for the current CU may be formed by averaging two prediction blocks using two MVs associated with the candidate's reference lists L0 and L1. The motion parameters of spatial / temporal neighbors may be inaccurate and may not represent the true motion of the current CU. To refine the MV of a bi-predicted normal merge candidate, decoder-side motion vector refinement (DMVR) may be applied. For example, when a conventional merge candidate (e.g., a spatial merge candidate and / or a TMVP merge candidate) is selected, a bi-predictive template may be generated (e.g., initially generated) based on, for example, the motion vectors from the reference lists L0 and L1, respectively, as the average. For example, when weighted prediction is enabled, the average may be a weighted average. As described herein, local motion refinement based on template matching may be performed by DMVR around the initial MV using the bi-predictive template.
[0078] FIG. 7 illustrates exemplary motion refinement that may be applied in DMVR. DMVR may refine the MV of a normal merge candidate, for example, as follows: As shown in FIG. 7, a bi-predictive template may be generated by averaging predictive blocks (e.g., two predictive blocks) using initial MVs (e.g., MV0 and MV1) in L0 and L1 of the merge candidate. For (e.g., each) reference list (e.g., L0 or L1), a template matching-based motion search may be performed in a local area around the initial MV. For (e.g., each) motion vector (e.g., MV0 or MV1) of a corresponding reference list around the initial MV in the list, a cost value (e.g., sum of absolute differences (SAD)) between the bi-predictive template and the corresponding predictive block using the motion vector may be measured. For the prediction direction, the MV that minimizes the template cost in the prediction direction may be considered as the last MV in the reference list of the normal merge candidate. For the prediction direction, one or more (e.g., eight) neighboring MVs surrounding the initial MV (e.g., with an integer sample offset of 1) may be considered during motion refinement. To generate the final bi-predictive signal of the current CU, refined MVs (e.g., two refined MVs, such as MV0′ and MV1′ as shown in FIG. 7) may be used.
[0079] As described herein, sub-block-based motion derivation (e.g., ATMVP and STMVP) and / or DMVR may increase the efficiency of merge mode, for example, by improving the granularity and / or accuracy of derived motion vectors.
[0080] For ATMVP and / or STMVP, for example, motion parameters of the current CU may be derived based on a granularity of 4x4 blocks. For example, motion derivation may be repeated to generate motion information (e.g., all motion information) of the current CU. Reference samples may be obtained from a temporal reference picture. The encoder / decoder may switch memory access to one or more (e.g., different) regions within the reference picture.
[0081] In ATMVP and / or STMVP, granularity (e.g., 4x4 block size) may be applied and used to derive motion parameters of ATMVP / STMVP-coded CUs in one or more pictures. The motion of video blocks in different pictures may exhibit different characteristics. For example, based on the correlation between the current picture and the current picture's reference pictures, video blocks in one or more pictures (e.g., pictures in high temporal layers of a random access configuration) may exhibit stable motion. The motion of video blocks in one or more pictures (e.g., pictures in low temporal layers of a random access configuration) may be unstable. The granularity level for deriving motion parameters of ATMVP / STMVP-coded CUs may be adjusted, for example, depending on different pictures.
[0082] For example, DMVR may be used to compensate for motion inaccuracies caused by using the motion of spatial / temporal neighbors of the current CU. DMVR may be enabled for CUs coded by the normal merge mode. When the motion parameters provided by the normal merge candidates are accurate, the improvement achieved by DMVR may be negligible. For example, DMVR may be skipped.
[0083] The signaling may support picture / slice-level variation of the derivation granularity (e.g., sub-block size) for calculating motion parameters of ATMVP-coded CUs and / or STMVP-coded CUs. An optimal granularity for ATMVP and / or STMVP motion derivation for the current picture may be determined.
[0084] For motion derivation in the DMVR-based merge mode, early truncation may be performed. In the normal merge mode, DMVR may be applied. Two or more prediction signals (e.g., two prediction signals) may be generated from the normal merge candidate. For example, the similarity between the prediction signals may be measured to determine whether to skip DMVR.
[0085] Local motion refinement at the intermediate bit depth may be performed. Motion refinement of the DMVR may be performed at the input bit depth. Some bit shifting and rounding operations (e.g., unnecessary bit shifting and rounding operations) may be removed from the DMVR.
[0086] Sub-block-based motion derivation based on ATMVP and STMVP may be performed. For ATMVP and / or STMVP, motion derivation may be performed at a fixed granularity. The granularity may be signaled in the sequence parameter set (SPS) as the syntax element log2_sub_pu_tmvp_size. The same derivation granularity may be applied to ATMVP and STMVP and may be used to calculate motion parameters of ATMVP / STMVP-coded CUs in pictures in a sequence.
[0087] The motion fields generated by ATMVP and / or STMVP may provide different characteristics. As described herein, the motion parameters of sub-blocks of an STMVP-encoded CU may be recursively derived by averaging the motion information of spatial and / or temporal neighbors of (e.g., each) sub-block within the CU based on the raster scan order. The motion parameters of an ATMVP-encoded CU may be derived from the temporal neighbors of the sub-blocks within the CU. STMVP may result in stable motion, and the motion parameters of sub-blocks within a CU may be consistent. Different granularities may be used to derive motion parameters for ATMVP and STMVP.
[0088] ATMVP and / or STMVP may, for example, use motion parameters of temporal neighbors in a reference picture to calculate motion parameters for the current block. When small motion exists between the current block and its co-located block in the reference picture, ATMVP and / or STMVP may provide motion estimation (e.g., reliable motion estimation). For blocks with small motion between the current block and its co-located block (e.g., blocks in the highest temporal layer of a random access (RA) configuration), sub-block motion parameters generated by ATMVP and / or STMVP may be similar. For video blocks that exhibit large motion from the co-located block (e.g., blocks in the lowest temporal layer of an RA configuration), the motion parameters calculated by ATMVP and / or STMVP for (e.g., each) sub-block may deviate from those of the sub-block's spatially neighboring sub-blocks. Motion derivation may be performed on small sub-blocks. For example, a current coding unit (CU) may be subdivided into one or more sub-blocks, and each sub-block (e.g., each sub-block) corresponds to a MV (e.g., a reference MV). An MV (e.g., a co-located MV) may be identified from a co-located picture for a sub-block (e.g., each sub-block) based on the reference MV for that sub-block. Motion parameters may be derived from one or more (e.g., different) pictures at one or more (e.g., different) granularities for ATMVP / STMVP-coded CUs. The derivation granularity (e.g., sub-block size) for ATMVP and / or STMVP may be selected (e.g., adaptively selected), for example, at the picture / slice level. For example, the size of the sub-block may be determined based on the temporal layer of the current CU.
[0089] Signaling of adaptively selected ATMVP / STMVP derivation granularity at the picture / slice level may be performed. One or more (e.g., two) granularity flags may be signaled in the SPS. For example, a slice_atmvp_granularity_enabled_flag and / or a slice_stmvp_granularity_enabled_flag may be signaled in the SPS to indicate whether the derivation granularity of ATMVP and / or STMVP may be adjusted at the slice level, respectively. A value (e.g., 1) may indicate that the corresponding ATMVP / STMVP-based derivation granularity is signaled at the slice level. A value (e.g., 0) may indicate that the corresponding ATMVP / STMVP-based derivation granularity is not signaled at the slice level, and that a syntax element (sps_log2_subblk_atmvp_size, or sps_log2_subblk_stmvp_size) is signaled in the SPS to specify the corresponding ATMVP / STMVP-based derivation granularity that may be used for slices that reference the current SPS. Table 1 illustrates exemplary syntax elements that may be signaled in the SPS. The syntax elements in Table 1 may be used in other high-level syntax structures, such as a video parameter set (VPS) and / or a picture parameter set (PPS).
[0090] [Table 1]
[0091] The parameter slice_atmvp_granularity_enabled_flag may specify the presence or absence of the syntax element slice_log2_subblk_atmvp_size in the slice segment header of the slice that references the SPS. For example, a value of 1 may indicate that the syntax element slice_log2_subblk_atmvp_size is present, and a value of 0 may indicate that the syntax element slice_log2_subblk_atmvp_size is not present in the slice segment header of the slice that references the SPS.
[0092] The parameter sps_log2_subblk_atmvp_size may specify a value of the sub-block size that may be used to derive motion parameters for enhanced temporal motion vector prediction for slices that reference the SPS.
[0093] The parameter slice_stmvp_granularity_enabled_flag may specify the presence or absence of the syntax element slice_log2_subblk_stmvp_size in the slice segment header of the slice that references the SPS. For example, a value of 1 may indicate that the syntax element slice_log2_subblk_stmvp_size is present, and a value of 0 may indicate that the syntax element slice_log2_subblk_stmvp_size is not present in the slice segment header of the slice that references the SPS.
[0094] The parameter sps_log2_subblk_stmvp_size may specify a value of the sub-block size that may be used to derive motion parameters for spatial-temporal motion vector prediction for a slice that references an SPS.
[0095] In Table 1, the syntax elements sps_log2_subblk_atmvp_size and sps_log2_subblk_stmvp_size may be specified (e.g., specified once) and applied to a video sequence. One or more (e.g., different) values of sps_log2_subblk_atmvp_size and sps_log2_subblk_stmvp_size may be specified for pictures at one or more (e.g., different) temporal levels. For a current picture that references an SPS, the values of sps_log2_subblk_atmvp_size and sps_log2_subblk_stmvp_size may be determined and / or applied depending on the temporal level to which the current picture belongs. The syntax elements sps_log2_subblk_atmvp_size and sps_log2_subblk_stmvp_size may be applied to square-shaped sub-block units. The sub-block units may be rectangular. For example, if the sub-block unit is rectangular, the sub-block width and height for ATMVP and / or STMVP may be specified.
[0096] Slice-level adaptation of ATMVP / STMVP-based derivation granularity may be enabled. For example, slice_atmvp_granularity_enabled_flag and / or slice_stmvp_granularity_enabled_flag in the SPS may be set to a value (e.g., 1) indicating the presence of the syntax element. The syntax element may be signaled in a slice segment header of (e.g., each) slice that references the SPS to specify the corresponding granularity level of ATMVP / STMVP-based motion derivation for the slice. For example, Table 2 illustrates exemplary syntax elements that may be signaled in a slice segment header.
[0097] [Table 2]
[0098] The parameter slice_log2_subblk_atmvp_size may specify a value of the sub-block size that may be used to derive motion parameters for enhanced temporal motion vector prediction for the current slice.
[0099] The parameter slice_log2_subblk_stmvp_size may specify a value of a sub-block size that may be used to derive motion parameters for spatio-temporal motion vector prediction for the current slice.
[0100] As shown in Tables 1 and 2, one or more (e.g., two) sets of syntax elements may be used to (e.g., separately) control the sub-block granularity of ATMVP-based and / or STMVP-based motion derivation. The sub-block granularity of ATMVP-based and / or STMVP-based motion derivation may be controlled separately, for example, when the characteristics (e.g., motion regularity) of the motion parameters derived by ATMVP and / or STMVP are different. For example, to (e.g., jointly) control the derivation granularity of ATMVP and / or STMVP at the sequence level and the slice level, the set of syntax elements slice_atmvp_stmvp_granularity_enabled_flag and sps_log2_subblk_atmvp_stmvp_size may be signaled in the SPS, and slice_log2_subblk_atmvp_stmvp_size may be signaled in the slice segment header. Tables 3 and 4 show example syntax changes in the SPS and slice segment headers, for example, when the sub-block granularity of the ATMVP and STMVP-based derived granularity is adjusted (eg, jointly).
[0101] [Table 3]
[0102] The parameter slice_atmvp_stmvp_granularity_enabled_flag may specify the presence or absence of the syntax element slice_log2_subblk_atmvp_stmvp_size in the slice segment header of the slice that references the SPS. For example, a value of 1 may indicate that the syntax element slice_log2_subblk_atmvp_stmvp_size is present, and a value of 0 may indicate that the syntax element slice_log2_subblk_atmvp_stmvp_size is not present in the slice segment header of the slice that references the SPS.
[0103] The parameter sps_log2_subblk_atmvp_stmvp_size may specify a sub-block size value that may be used to derive motion parameters for ATMVP and / or STMVP for a slice that references the SPS.
[0104] [Table 4]
[0105] The parameter slice_log2_subblk_atmvp_stmvp_size may specify a sub-block size value that may be used to derive motion parameters for enhanced temporal motion vector prediction and / or spatio-temporal motion vector prediction for the current slice.
[0106] The sub-block granularity of ATMVP / STMVP-based motion derivation at the picture / slice level may be determined.
[0107] The granularity level of ATMVP / STMVP-based motion derivation for a picture / slice may be determined, for example, based on the temporal layer of the picture / slice. As described herein, given correlation between pictures in the same video sequence, the ATMVP / STMVP-derived granularity may be similar to that of neighboring pictures of the picture / slice, for example, in the same temporal layer. For example, for a picture in the highest temporal layer of RA, ATMVP / STMVP-based motion estimation may result in large block partitions. The granularity value may be adjusted (e.g., to a larger value). For a picture in the lowest temporal layer of RA, the motion parameters derived by ATMVP and / or STMVP may be less accurate. For example, to calculate the sub-block granularity that may be used for ATMVP / STMVP-based motion derivation in the current picture, the average size of CUs coded by ATMVP and / or STMVP from previously coded pictures in the same temporal layer may be used. For example, the current picture may be the i-th picture in the k-th temporal layer. There may be M CUs in the current picture, which may be coded by ATMVP and / or STMVP. The M CUs may be denoted by s0, s1, ..., s M-1 If the size of the CU is σ, then the average size of the CUs coded by ATMVP / STMVP in the current picture, e.g., σ k teeth,
[0108]
number
[0109] It can be calculated as follows:
[0110] Based on equation (1), when encoding the (i+1)th picture in the kth temporal layer, the corresponding sub-block size of ATMVP / STMVP-based motion derivation is
[0111]
number
[0112] teeth,
[0113]
number
[0114] can be determined by
[0115] For RA configurations, parallel encoding may be supported. When parallel encoding is enabled, a full-length sequence may be divided into multiple independent random access segments (RAS), which may span a sequence playback of, for example, a shorter duration (e.g., about 1 second), and each video segment may be encoded separately. Neighboring RASs may be independent. The result of sequential encoding (e.g., encoding the entire sequence in frame order) may be the same as parallel encoding. When adaptive sub-block granularity derivation is applied (e.g., in addition to parallel encoding), a picture may avoid using ATMVP / STMVP block size information of a picture from a preceding RAS, for example. For example, when encoding the first inter-picture in a RAS, σ is used for one or more (e.g., all) temporal layers (e.g., using 4x4 sub-blocks). k The value of may be reset to 0. Figure 8 illustrates an example where ATMVP / STMVP block size statistics may be reset to 0 to indicate the location of a refreshing picture (e.g., intra period equals 8) when parallel encoding is enabled. In Figure 8, blocks surrounded by dashed and / or solid lines may represent intra pictures and inter pictures, respectively, and pattern blocks may represent refreshing pictures. Once calculated,
[0116]
number
[0117] The log2() of the value may be transmitted in the bitstream in the slice header, for example, according to the syntax in Tables 1 and 2.
[0118] In an example, the sub-block size used for ATMVP / STMVP-based motion derivation may be determined at the encoder and transmitted to the decoder. In an example, the sub-block size used for ATMVP / STMVP-based motion derivation may be determined at the decoder. Adaptive determination of sub-block granularity may be used as a decoder-side technique. For example, ATMVP / STMVP block size statistics (e.g., as shown in equation (1)) may be maintained during encoding and / or decoding, and may be used to synchronize the encoder and decoder when they determine their respective sub-block sizes for ATMVP / STMVP-based motion derivation of a picture / slice, e.g., using equations (1) and (2). For example, if ATMVP / STMVP block size statistics are maintained during encoding and decoding,
[0119]
number
[0120] The value of may not be transmitted.
[0121] Early termination of DMVR based on the similarity of predicted blocks may be performed. DMVR may be performed for CUs coded using normal merge candidates (e.g., spatial candidates and / or TMVP candidates). When the motion parameters provided by the normal merge candidate are accurate, DMVR may be skipped without coding loss. To determine whether the normal merge candidate can provide accurate motion estimation for the current CU, the average difference between the two predicted blocks may be calculated, for example, as follows:
[0122]
number
[0123] where I (0) (x,y) and I (1) (x,y) may be the sample values at coordinates (x,y) of the L0 and L1 motion compensated blocks generated using the motion information of the merge candidate, B and N may be the set of sample coordinates and the number of samples defined within the current CU, respectively, and D may be a distortion measure (e.g., sum of squared errors (SSE), sum of absolute differences (SAD), and / or sum of absolute transformed differences (SATD)). Given Equation (3), for example, if the difference measure between two prediction signals is at most a predefined threshold (e.g., Diff≦D thres If the difference measure between the two prediction signals is greater than a predefined threshold, the prediction signals generated by the merge candidate may not be highly correlated, and DMVR may be applied. Figure 9 illustrates an example motion compensation after early truncation is applied to DMVR.
[0124] Sub-CU level motion derivation may be disabled (e.g., adaptively disabled), for example, based on prediction similarity. Before DMVR is performed, one or more (e.g., two) prediction signals using the MVs of the merge candidates may be available. The prediction signals may be used to determine whether DMVR should be disabled.
[0125] High-precision prediction for DMVR may be performed. A prediction signal of a bi-predictive block may be generated, for example, by averaging prediction signals from L0 and L1 at the precision of the input bit depth. If the MV points to a fractional sample position, a prediction signal may be obtained at intermediate precision (which may be higher than the input bit depth due to the interpolation operation, for example) using interpolation. The intermediate precision signal may be rounded to the input bit depth before the averaging operation. The input signal to the averaging operation may be shifted to a lower precision, which may introduce rounding errors into the generated bi-predictive signal, for example. For example, if a fractional MV is used for a block, two prediction signals at the input bit depth may be averaged at intermediate precision. If the MV corresponds to a fractional sample position, the interpolation filtering may not round the intermediate values to the input bit depth and may keep the intermediate values at a high precision (e.g., intermediate bit depth). In the case where one of the two MVs is integer motion (e.g., the corresponding prediction is generated without applying interpolation), the precision of the corresponding prediction can be increased to an intermediate bit depth before averaging is applied. For example, the precision of the two prediction signals can be the same. Figure 10 illustrates an exemplary bi-prediction when averaging two intermediate prediction signals with high precision, where
[0126]
number
[0127] and
[0128]
number
[0129] may refer to two prediction signals obtained from lists L0 and L1 at an intermediate bit depth (e.g., 14 bits), and BitDepth may indicate the bit depth of the input video.
[0130] In DMVR, a bi-predictive signal (e.g., I (0) (x,y) and I (1) (x,y)) may be defined with the precision of the input signal bit depth. The input signal bit depth may be, for example, 8 bits (e.g., if the input signal is 8 bits) or 10 bits (e.g., if the input signal is 10 bits). The prediction signal may be converted to a lower precision, for example, before motion refinement. Rounding errors may be introduced when measuring distortion costs. The conversion of the prediction signal from an intermediate bit depth to the input bit depth may include one or more rounding and / or clipping operations. DMVR-based motion refinement is performed using a prediction signal that may be generated at a high bit depth (e.g., at the intermediate bit depth in FIG. 10).
[0131]
number
[0132] and
[0133]
number
[0134] The corresponding distortion between two prediction blocks in equation (3) can be calculated with high accuracy as follows:
[0135]
number
[0136] where:
[0137]
number
[0138] and
[0139]
number
[0140] Diff may be high precision samples at coordinates (x, y) of the prediction blocks generated from L0 and L1, respectively. h may represent the corresponding distortion measure calculated at the intermediate bit depth. Due to the increased bit depth, the threshold (e.g., distortion measure threshold) that may be used to early terminate the DMVR may be adjusted so that the threshold is defined at the same bit depth as the predicted signal. If L1-norm distortion (e.g., SAD) is used, the following equation may be used to adjust the distortion threshold from the input bit depth to the intermediate bit depth:
[0141]
number
[0142] An enhanced temporal motion vector prediction (ATMVP) may be derived. For ATMVP, a co-located picture and a co-located block may be selected. A motion vector from a spatially neighboring CU may be added to a candidate list (e.g., a list of potential candidate neighboring CUs). For example, if a neighboring CU is available and the MV of the neighboring CU differs from one or more MVs in the existing candidate list, a motion vector from the spatially neighboring CU may be added. For example, as shown in FIG. 4, MVs from neighboring blocks may be added in the order A1, B1, B0, A0. The number of available spatial candidates may be represented by N0. ATMVP may be derived using N0 MVs.
[0143] N0 may be greater than 0. If N0 is greater than 0, an MV (e.g., the first available MV) may be used to determine the co-located picture and / or the offset for obtaining the motion. As shown in FIG. 5, the first available MV may be from neighboring CU A. The co-located picture for ATMVP may be a reference picture associated with the MV from CU A. The offset for obtaining the motion field may be derived from the MV. N0 may be equal to 0. If N0 is equal to 0, the co-located picture may be set to the co-located picture signaled in the slice header, and the offset for obtaining the motion field may be set to 0.
[0144] For example, when multiple reference pictures are used, the co-located pictures for ATMVP derivation of different CUs may be different. For example, when multiple reference pictures are used, the co-located pictures for ATMVP derivation of different CUs may be different because the determination of the co-located picture may depend on their neighboring CUs. For decoding of the current picture, the co-located picture for ATMVP may not be fixed, and ATMVP may refer to the motion fields of multiple reference pictures. The co-located picture for decoding of the current picture may be set to (e.g., one) reference picture, which may be signaled in the slice header, for example. The co-located picture may be identified. The reference picture of neighboring CU A may be different from the co-located picture. The reference picture of CU A may be R A and the collocated picture can be expressed as R col and the current picture may be represented as P. POC(x) may be used to indicate the POC of picture x. As calculated in equation (6), the MV of CU A may be calculated by subtracting the MV of picture R from the MV of picture R to obtain a prediction for the offset position. A to a co-located picture.
[0145] Music Video col =MV(A)×(POC(Rcol )-POC(P)) / (POC(R A )-POC(P)) (6) Scaled MV col For example, the colocated picture R col , L0, L1)。 The scaling in equation (6) may be based on the temporal distance of the pictures. For example, the first available MV from the spatial CU may be selected for scaling. An MV that may minimize a scaling error may be selected. For example, to minimize the scaling error, an MV to be scaled (e.g., the best MV) may be selected from N0 MVs in one or more (e.g., two) directions (e.g., list L0, list L1). For example, (e.g., each) neighboring block (e.g., neighboring CU) among the neighboring blocks may have a corresponding reference picture. A neighboring block may be selected to be a candidate neighboring block (e.g., a candidate neighboring CU) based on the difference between the neighboring block's reference picture and the co-located picture. A neighboring block selected to be a candidate neighboring block may have the smallest temporal distance between its reference picture and the co-located picture. For example, the reference picture (e.g., each reference picture) and the co-located picture may have a picture order count (POC), and a candidate neighboring block may be selected based on a difference (e.g., having the smallest difference) between the POC of the reference picture and the POC of the co-located picture. MVs (e.g., co-located MVs) may be identified from the co-located picture based on MVs (e.g., reference MVs) from the reference picture. A neighboring CU may have the same reference picture as the co-located picture, and when the reference picture and the co-located picture are determined to be the same, the neighboring CU may be selected (e.g., without consideration of other neighboring CUs). The MV from the reference picture may be a spatial MV, and the MV from the co-located picture may be a temporal MV.
[0146] For example, if the reference picture is not the same as the co-located picture, the MV from the reference picture may be scaled. For example, the MV from the reference picture may be multiplied by a scaling factor. The scaling factor may be based on the temporal difference between the reference picture and the co-located picture. The scaling factor may be ((POC(R col )-POC(P)) / (POC(R A )-POC(P). The scaling factor may have a value representing no scaling (e.g., 1). A scaling factor having a value representing no scaling may be defined as R A and R col may indicate that the MVs are the same picture. The scaling error may be measured in one or more of the following ways. For example, the scaling error may be measured as provided in equation (7) and / or as provided in equation (8). The absolute difference between the scale factor for a given MV and the scale factor value representing no scaling may be measured, for example, as provided in equation (7).
[0147] ScaleError=abs((POC(R col )-POC(R A )) / (POC(R A )-POC(P)) (7) The absolute difference between the reference picture of a given MV and the co-located picture may be measured, for example, as provided in equation (8).
[0148] ScaleError=abs((POC(R col )-POC(R A ))) (8) The search for the best MV may be terminated. For example, during the search for scaled MV candidates, if ScaleError is equal to 0 for a given MV, the search may be terminated (e.g., terminated early).
[0149] Neighboring MVs may be used to match the motion fields of co-located pictures for ATMVP derivation. Neighboring MVs may be selected, for example, by minimizing the accuracy loss caused by MV scaling. This may improve the accuracy of the scaled MVs. The presence of valid motion information in the reference block may be indicated by the selected neighboring MVs in the co-located picture. For example, when the reference block is an intra block (e.g., an intra block), ATMVP may be considered unavailable, for example, because there is no motion information associated with the reference block.
[0150] A best neighboring MV may be selected from the neighboring MVs that point to each inter-coded block in the co-located picture. MV scaling error may be minimized. For example, when determining the best neighboring MV (e.g., as shown in Equations (7) and (8)), a constraint may be imposed to ensure that the selected neighboring MV points to (e.g., one) inter-coded block in the co-located picture. The selection of the best neighboring MV may be formulated as provided in Equation (9).
[0151]
number
[0152] (x,y) may be the center position of the current CU. ColPic(x,y) may indicate a block at position (x,y) within a collocated picture. inter() may represent an indicator function that may indicate whether the input block is an inter block. S may indicate the set of available spatial neighboring blocks, for example, S={A1, B1, B0, A0}. ScaleError may indicate the MV scaling error, as calculated by equations (7) and (8).
[0153]
number
[0154] can be the scaled motion vector of neighboring block N. N * For example, based on ATMVP, it can represent a selected spatial neighbor whose MV is used to derive the motion field of the current CU. In the example, the current block can have three spatial neighbors A, B, and C that increase the scaling error (e.g., ScaleError(A) < ScaleError(B) < ScaleError(C)).
[0155]
Number
[0156] If it is false (e.g., after scaling, if the motion of block A identifies an intra-coded reference block in the collocated picture), the motion of block B (e.g., whose scaling error is the second smallest) can be used to identify the reference block in the collocated picture. If the scaled motion of B identifies an intra-coded reference block, the motion of block C can be used to identify the reference block in the collocated picture.
[0157] For example, the MV scaling error, as shown in Equation (7), may be used as a criterion for identifying the best neighboring block, as shown in Equation (9). The motion of the best neighboring block may be used to select a co-located block from a co-located picture. For example, the calculation of the MV scaling error, as shown in Equation (7), may include, for example, two subtractions, one division, and one absolute value operation. The division may be performed by multiplication (e.g., based on a LUT) and / or a right shift. To determine the co-located block for ATMVP, the best spatial neighboring block of the current CU may be selected. The co-located picture may be signaled at the slice and / or picture level. The co-located picture may be used for ATMVP derivation. The MVs of existing merge candidates may be examined in order (e.g., A1, B1, B0, A0 as shown in FIG. 4). To obtain a co-located block from a co-located picture, a first merge candidate MV that identifies a block (e.g., one) associated with the co-located picture and coded by inter prediction may be selected. For example, if no such candidate exists, zero motion may be selected. For example, if the selected MV points to an inter-coded co-located block (e.g., one), ATMVP may be enabled. For example, if the selected MV points to an intra-coded co-located block (e.g., one), ATMVP may be disabled. Figure 11 shows an example flowchart illustrating co-located block derivation as described herein.
[0158] The MVs of existing merge candidates may be examined in order. For example, referring to FIG. 4, the order may be A1, B1, B0, A0. To obtain a co-located block from a co-located picture, the MV of the first merge candidate associated with the co-located picture may be selected. ATMVP may be enabled based on the coding mode of the co-located block. For example, if the co-located block is intra-coded, ATMVP may be disabled because the co-located block may not provide motion information. For example, if none of the merge candidates are associated with the co-located picture, ATMVP may be disabled. Early termination of the check may be performed. For example, the check may be terminated as soon as the first merge candidate associated with the co-located picture is found. FIG. 14 shows a flowchart illustrating co-located block derivation for ATMVP using merge candidate checking.
[0159] Zero motion may be used to obtain the co-located block in the co-located picture. To determine whether to enable ATMVP, the block co-located with the current CU in the co-located picture may be checked. For example, if the block is inter-coded, ATMVP may be enabled. For example, if the block is intra-coded, ATMVP may be disabled.
[0160] The area for obtaining co-located blocks for ATMVP may be restricted. Co-located pictures for ATMVP derivation for different ATMVP blocks may be restricted to (e.g., one) reference picture. Corresponding co-located blocks may be indicated by the MVs of selected merge candidates of neighboring blocks. Corresponding co-located blocks may be distant from each other. An encoder or decoder may (e.g., frequently) switch access to motion (e.g., MVs and / or reference picture indices) of different regions within a co-located picture. Figure 12 shows an example of unrestricted access to co-located blocks for ATMVP. As shown in Figure 12, there may be one or more (e.g., three) ATMVP CUs within the current CTU, and the CUs (e.g., each CU) use motion offsets (e.g., different motion offsets), e.g., as indicated by different colors in Figure 12. The offsets may be used to find corresponding co-located blocks in the co-located picture, e.g., as indicated by dashed blocks in Figure 12. The co-located blocks may be found in different regions of the co-located picture, for example, due to motion offsets having different values.
[0161] Co-located blocks of an ATMVP CU (e.g., each ATMVP CU) may be derived within a constrained range (e.g., one). Figure 13 shows an example of applying a constrained region to derive co-located blocks for ATMVP. Figure 13 may show the same co-located picture and CTU as Figure 12. As shown in Figure 13, given the position of the current CU, a constrained area (e.g., region) within the co-located picture may be determined. For example, the constrained area may be determined based on the current slice. The position of the co-located block used for ATMVP derivation of the current CU may be within the area. Co-located blocks within the area may be valid co-located blocks. MVs from neighboring blocks (e.g., candidate neighboring blocks) that are not within the area may be replaced with MVs from valid co-located blocks. For example, as shown in Figure 13, the first co-located blocks of B1 and B2 (e.g., ColB1 and ColB2) may be within the constrained area. The initial co-located blocks ColB1 and ColB2 may be used for ATMVP derivation of B1 and B2. Because the first co-located block of B0 (e.g., ColB0) is outside the constrained area, a (e.g., one) co-located block (e.g., ColB0') may be generated by clipping the position of ColB0 toward the nearest boundary of the constrained area. For example, ColB0' may be the closest valid block to ColB0. The position of ColB0' may be set as the block located in the same location as the current block in the co-located picture (e.g., the motion of B0 may be set to zero). To code (e.g., encode and / or decode) the CU, the MV from ColB0' may be used (e.g., instead of the MV from ColB0). For example, when ColB0 is outside the boundary, ATMVP may be disabled.
[0162] The size of the restricted area for ATMVP co-located block derivation may be determined. For example, a fixed area (e.g., one) may be applied to CUs (e.g., all CUs) coded by ATMVP in a video sequence. The derivation of co-located blocks for a (e.g., one) ATMVP CU of a current CTU (e.g., the CTU containing the current CU) may be restricted to be within co-located CTUs of the same area within a co-located picture. For example, co-located blocks (e.g., only) within a CTU in a co-located picture that is co-located with the current CTU may be derived. The derivation of co-located blocks for a (e.g., one) CU's TMVP process may be restricted to be within the current CTU and within a (e.g., one) column of 4x4 blocks. A CTU may contain WxH (e.g., width x height) samples. The region for deriving TMVP co-located blocks may be (W + 4)xH. The same constrained area size (e.g., the current CTU plus (e.g., one) row of 4x4 blocks) may be used to derive co-located blocks for both ATMVP and TMVP. The size of the constrained area may be selected and signaled from the encoder to the decoder. Syntax elements may be added at the sequence and / or picture or slice level. Different profiles and / or levels may be defined for various application requirements. For example, syntax elements may be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS), or may be signaled in the slice header.
[0163] The selection of a co-located block for ATMVP and the area restriction for co-located block derivation may be combined. For example, as shown in FIG. 14, a co-located block may be selected based on the characteristics of one or more merge candidates. The MVs of existing merge candidates may be examined in order (e.g., A1, B1, B0, A0 as shown in FIG. 4). To obtain a co-located block from a co-located picture, the MV of the first merge candidate associated with the co-located picture may be selected. For example, if the co-located block is intra-coded or if none of the merge candidates is associated with the co-located picture, ATMVP may be disabled. A valid merge candidate may be found. For example, the valid merge candidate may be A1. The co-located block in the co-located picture corresponding to A1 may be denoted as ColA1. ColA1 may be outside the bounds of the restricted range. If ColA1 is outside the boundaries of the restricted range, ColA1 may be clipped back to the nearest boundary of the restricted area and set to a block located in the same location as the current block in the co-located picture (e.g., the motion of A1 may be set to zero), and / or ATMVP may be marked as disabled.
[0164] Figure 15 shows an example of deriving an MV for a current block using a co-located picture. The co-located picture for the current picture may be signaled, for example, in a slice header. The co-located picture may be used when performing ATMVP on the current picture. For example, reference pictures for neighboring blocks of the current block in the current picture may be compared with the co-located picture. A neighboring block may be selected based on the reference picture of the neighboring block that has the smallest temporal distance to the co-located picture. The temporal distance may be the POC difference. The reference picture of the neighboring block may be the same as the co-located picture. The MV from the reference picture for the selected neighboring block may be used to determine the MV from the co-located picture. The MV from the co-located picture may be used to code the current block. For example, if the reference picture of the selected neighboring block is not the same as the co-located picture, the MV from the co-located picture may be scaled.
[0165] Although features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware contained in a computer-readable medium, executed by a computer or processor. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer. [Industrial Applicability]
[0166] The present invention can be applied to video coding systems in general.
Claims
[Claim 1] 1. A device for video decoding, comprising: Get the colocated picture associated with the video block, determining a position associated with the co-located block in the co-located picture based on the position of the video block and a temporal motion vector (temporal MV), and a clipping operation constraining the position associated with the co-located block within a constrained region in the co-located picture; determining MVs of collocated sub-blocks in the collocated block; predicting a sub-block of the video block based on the MV of the co-located sub-block; Decoding the video block based on the predicted sub-blocks. Processor configured to A device with.