Adaptive control point selection for video coding based on affine motion models.

The method of clipping motion vectors in block-based video coding systems using affine motion mode addresses the issue of out-of-range values, enhancing prediction accuracy and efficiency by utilizing neighboring block vectors and defined ranges.

JP7794901B2Active Publication Date: 2026-01-06INTERDIGITAL VC HOLDINGS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024112738
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-11
Filing Date
2024-07-12
Publication Date
2026-01-06
Estimated Expiration
2039-06-27

AI Technical Summary

Technical Problem

In block-based video coding systems, motion vectors associated with video blocks can have values outside a particular range, leading to unintended results.

Method used

A method for clipping motion vectors based on affine motion mode, using control point affine motion vectors from neighboring blocks, and determining sub-block motion vectors within a defined motion field range, stored for spatial and temporal prediction.

Benefits of technology

Ensures accurate and efficient video coding by constraining motion vectors within valid ranges, improving prediction accuracy and reducing unintended outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794901000023
    Figure 0007794901000023
  • Figure 0007794901000024
    Figure 0007794901000024
  • Figure 0007794901000025
    Figure 0007794901000025
Patent Text Reader

Abstract

To provide motion vector clipping when an affine motion mode is enabled for a video block.SOLUTION: A video coding device is configured to: determine that an affine mode for a video block is enabled; determine a plurality of control point affine motion vectors associated with the video block; store the plurality of clipped control point affine motion vectors for motion vector prediction of a neighboring control point affine motion vector; derive a sub-block motion vector associated with a sub-block of the video block; clip the derived sub-block motion vector; and store it for spatial motion vector prediction or temporal motion vector prediction. For example, the video coding device may clip the derived sub-block motion vector based on a motion field range that may be based on a bit depth value.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and device for clipping motion vectors, and more particularly to a method and device for clipping motion vectors for video blocks. [Background technology]

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 691,770, filed June 29, 2018, U.S. Provisional Patent Application No. 62 / 734,728, filed September 21, 2018, and U.S. Provisional Patent Application No. 62 / 778,055, filed December 11, 2018, the contents of which are incorporated herein by reference.

[0003] Video coding systems may be used to compress digital video signals, e.g., to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems may include block-based systems, wavelet-based systems, and / or object-based systems. Block-based hybrid video coding systems may be deployed. In block-based video coding systems, motion vectors associated with sub-blocks of a video block may have values ​​that may be outside a particular range. Using such values ​​may produce unintended results. Summary of the Invention

[0004] Systems, methods, and means are disclosed for clipping of motion vectors when affine motion mode is enabled for a video block (e.g., a coding unit (CU)). A video coding device may determine that affine mode is enabled for a video block (e.g., a current video block). The video block may include multiple sub-blocks. The video coding device may determine multiple control point affine motion vectors associated with the video block. At least one of the control point affine motion vectors associated with the current video block may be determined using one or more control point affine motion vectors associated with one or more neighboring video blocks. The video coding device may clip the control point affine motion vector associated with the current video block. For example, the control point affine motion vector may be clipped based on a bit depth used for storing the motion field. The video coding device may store the clipped control point affine motion vector for motion vector prediction of neighboring control point affine motion vectors.

[0005] The video coding device may derive a sub-block motion vector associated with a sub-block. The video coding device may derive the sub-block motion vector based on one or more control point affine motion vectors. The video coding device may clip the derived sub-block motion vector. For example, the video coding device may clip the derived sub-block motion vector based on a motion field range. The motion field range may be used for motion field storage. The motion field range may be based on a bit depth value. The video coding device may store the clipped sub-block motion vector for spatial motion vector prediction or temporal motion vector prediction. The video coding device may predict the sub-block using the clipped sub-block motion vector.

[0006] The video coding device may determine control point positions associated with the control point affine motion vectors of the video block based on the shape of the video block. For example, the control point positions may be determined based on the length and / or width of the video block.

[0007] For example, the control point locations may include a top-left control point and a top-right control point, e.g., if the width of the current video block is greater than the length of the current video block. The video coding device may classify such a video block as a horizontal rectangular video block. For example, the control point locations may include a top-left control point and a bottom-left control point, e.g., if the width of the current video block is less than the length of the current video block. The video coding device may classify the current video block as a vertical rectangular video block. The control point locations may include a bottom-left control point and a top-right control point, e.g., if the width of the current video block is equal to the length of the current video block. The video coding device may classify the current video block as a square video block. [Brief explanation of the drawings]

[0008] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system. [Figure 1B] 1B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communication system shown in FIG. 1A. [Figure 1C] 1B is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system shown in FIG. 1A. [Figure 1D] FIG. 1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system shown in FIG. 1A. [Figure 2] FIG. 1 is an exemplary diagram of a block-based video encoder. [Figure 3] FIG. 2 is an exemplary block diagram of a video decoder. [Figure 4] FIG. 1 illustrates an exemplary block partition in a multi-type tree structure. [Figure 5] FIG. 10 is a diagram illustrating an example of an affine mode of four parameters. [Figure 6] FIG. 10 is a diagram illustrating examples of affine merge candidates. [Figure 7] FIG. 10 illustrates an exemplary motion vector derivation at control points for an affine motion model. [Figure 8] FIG. 1 illustrates an example of constructing an affine motion predictor. [Figure 9] FIG. 1 illustrates an example of temporal scaling of affine motion vectors (MVs) for generation of MV predictors. [Figure 10] FIG. 10 is a diagram illustrating an example of adaptive control point selection based on block shape. [Figure 11] FIG. 10 is a diagram illustrating an example of affine merge selection using maximum control point distance. [Figure 12] FIG. 1 illustrates an exemplary workflow for motion field generation for affine mode. [Figure 13] FIG. 1 illustrates an example workflow for generating prediction samples for an affine coding unit (CU) by reusing the motion field used for MV prediction and deblocking. [Figure 14] FIG. 1 illustrates an example workflow for MV prediction and deblocking for affine CUs by reusing a motion field to generate prediction samples. [Figure 15] FIG. 10 illustrates an example of modifying one or more control points MV to scale a reference block. [Figure 16] FIG. 10 illustrates an example of modifying control points MV to include a reference block. DETAILED DESCRIPTION OF THE INVENTION

[0009] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which:

[0010] 1A illustrates an example communication system 100 in which one or more disclosed examples may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0011] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed example may contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspot or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0012] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, e.g., the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node-B, an eNode-B, a Home Node-B, a Home eNode-B, a gNB, an NR Node-B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0013] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), e.g., a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for wireless service to a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in an example, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In an example, the base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.

[0014] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which may be any suitable wireless communications link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0015] More specifically, as described above, the communication system 100 may be a multiple-access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114a and the WTRUs 102a, 102b, 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​UL Packet Access (HSUPA).

[0016] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0017] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.

[0018] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement both LTE radio access and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0019] In an example, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology, such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.

[0020] 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a localized area, such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a roadway, etc. In an example, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an example, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In an example, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 1A, the base station 114b may have a direct connection to the Internet 110. Therefore, the base station 114b may not be required to access the Internet 110 via the CN 106 / 115.

[0021] The RAN 104 / 113 may be in communication with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, e.g., separate throughput, latency, error tolerance, reliability, data throughput, mobility, etc. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, e.g., user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs employing the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to a RAN 104 / 113 that may utilize NR radio technology, the CN 106 / 115 may be in communication with another RAN (not shown) that employs GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0022] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing Plain Old Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or IP in the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs that may employ the same RAT as the RAN 104 / 113 or a different RAT.

[0023] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with separate wireless networks over separate wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a, which may employ a cellular-based wireless technology, and with a base station 114b, which may employ an IEEE 802.11 wireless technology.

[0024] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the foregoing elements.

[0025] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0026] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in an example, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an example, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR signals, UV signals, or visible light signals. In an example, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0027] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in an example, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0028] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0029] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any type of suitable memory, for example, non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In examples, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, for example, on a server or home computer (not shown).

[0030] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0031] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. The WTRU 102 may receive location information over the air interface 116 from a base station (e.g., base stations 114a, 114b) in addition to or instead of information from the GPS chipset 136, and / or determine its location based on the timing of signals being received from two or more nearby base stations. It will be understood that the WTRU 102 may obtain location information through any suitable location determination method.

[0032] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0033] The WTRU 102 may include a full-duplex radio in which transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit 139 to reduce and or substantially eliminate self-interference through signal processing in hardware (e.g., a choke) or via a processor (e.g., via a separate processor (not shown) or via the processor 118). In an example, the WTRU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)).

[0034] 1C is a system diagram illustrating an example RAN 104 and CN 106. As described above, the RAN 104 may employ E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 may be in communication with the CN 106.

[0035] The RAN 104 may include eNode-Bs 160a, 160b, and 160c, although it will be understood that the RAN 104 may include any number of eNode-Bs. The eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In examples, the eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Thus, the eNode-B 160a may use multiple antennas, for example, to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.

[0036] Each of the eNode-Bs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 1C, the eNode-Bs 160a, 160b, 160c may communicate with each other via an X2 interface.

[0037] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. Although each of the above elements is shown as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0038] The MME 162 may be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during initial connection of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0039] The SGW 164 may be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions, such as fixing the user plane during handover between eNode Bs, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing the context of the WTRUs 102a, 102b, 102c, etc.

[0040] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0041] The CN 106 may facilitate communication with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communication between the WTRUs 102a, 102b, 102c and traditional landline communication devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. Additionally, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0042] Although the WTRU is depicted in Figures 1A-1D as a wireless terminal, it is contemplated that in certain examples such a terminal may use a wired communication interface with the communication network (e.g., temporarily or permanently).

[0043] In an example, the other network 112 may be a WLAN.

[0044] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that carries traffic to and from the BSS. Traffic to a STA originating from outside the BSS may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP for delivery to the respective destination. Traffic between STAs within a BSS may be sent through the AP, for example, where a source STA may send traffic to the AP, which may deliver the traffic to the destination STA. Traffic between STAs within a BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between (e.g., directly between) a source STA and a destination STA using direct link setup (DLS). In examples, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and the STAs within or using the IBSS (e.g., all of the STAs) may communicate directly with each other. The IBSS mode of communication is sometimes referred to herein as an "ad hoc" mode of communication.

[0045] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, the AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be of a fixed width (e.g., a 20 MHz wide bandwidth) or dynamically configured via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In an example, for example, in an 802.11 system, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented. With CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA may back out. One STA (e.g., only one station) may transmit in a given BSS at any given time.

[0046] A high-throughput (HT) STA may use a 40 MHz wide channel for communication, for example, via combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.

[0047] A very high throughput (VHT) STA may support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. A 40 MHz channel and / or an 80 MHz channel may be formed by combining adjacent 20 MHz channels. A 160 MHz channel may be formed by combining eight adjacent 20 MHz channels or by combining two non-adjacent 80 MHz channels (this may be referred to as an 80+80 configuration). For the 80+80 configuration, after channel encoding, the data may be passed through a segment parser, which may split the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing may be performed separately on each stream. The streams may be mapped onto two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration may be reversed, and the combined data may be sent to the medium access control (MAC).

[0048] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. By way of example, 802.11ah may support meter-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have limited capabilities, including support for (e.g., only support for) specific and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).

[0049] A WLAN system that may support multiple channels and channel bandwidths, e.g., 802.11n, 802.11ac, 802.11af, and 802.11ah, includes a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, for a STA (e.g., an MTC-type device) that supports (e.g., only supports) the 1 MHz mode, the primary channel may be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. If the primary channel is busy, for example due to STAs (that only support a 1 MHz mode of operation) transmitting to the AP, then the entire available frequency band may be considered busy even though most of those frequency bands remain idle and potentially available.

[0050] In the United States, the available frequency bands that can be used by 802.11ah are from 902 MHz to 928 MHz. In South Korea, the available frequency bands are from 917.5 MHz to 923.5 MHz. In Japan, the available frequency bands are from 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz depending on the country code.

[0051] 1D is a system diagram illustrating an example RAN 113 and CN 115. As described above, the RAN 113 may employ NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 may be in communication with the CN 115.

[0052] The RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that the RAN 113 may include any number of gNBs. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In examples, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a, 180b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, and 180c. Thus, the gNB 180a may use multiple antennas, for example, to transmit wireless signals to and / or receive wireless signals from the WTRU 102a. In an example, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on an unlicensed spectrum, while the remaining component carriers may be on a licensed spectrum. In an example, the gNBs 180a, 180b, and 180c may implement coordinated multipoint (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).

[0053] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may be different for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of varying or scalable lengths (e.g., including different numbers of OFDM symbols and / or different lengths of absolute time duration).

[0054] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing other RANs (e.g., eNode-Bs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect to a gNB 180a, 180b, 180c while also communicating / connecting to another RAN, such as an eNode-B 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement the DC principle to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNode-Bs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.

[0055] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, routing of control plane information to access and mobility management functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.

[0056] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the above elements is shown as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0057] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling separate PDU sessions with separate requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize CN support for the WTRUs 102a, 102b, 102c based on the type of service being utilized by the WTRUs 102a, 102b, 102c. For example, separate network slices may be established for separate use cases, such as services relying on Ultra-Reliable Low-Latency (URLLC) access, services relying on enhanced High-Capacity Mobile Broadband (eMBB) access, services related to Machine-Type Communications (MTC) access, etc. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.

[0058] The SMFs 183a, 183b may be connected to the AMFs 182a, 182b in the CN 115 via an N11 interface. The SMFs 183a, 183b may be connected to the UPFs 184a, 184b in the CN 115 via an N4 interface. The SMFs 183a, 183b may select and control the UPFs 184a, 184b and configure the routing of traffic through the UPFs 184a, 184b. The SMFs 183a, 183b may perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0059] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.

[0060] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In an example, the WTRUs 102a, 102b, 102c may be connected to local data networks (DNs) 185a, 185b through UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.

[0061] 1A-1D and the corresponding descriptions thereof, one or more or all of the functions described herein in connection with one or more of the WTRUs 102a-d, base stations 114a-b, eNode-Bs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-ab, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functions.

[0062] The emulation device may be designed to perform one or more tests of other devices in a lab environment and / or in an operator network environment. For example, one or more emulation devices may perform one or more, or all, functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communications network to test other devices in the network. One or more emulation devices may perform one or more, or all, functions while temporarily implemented / deployed as part of a wired and / or wireless communications network. The emulation device may be directly coupled to another device for testing purposes and / or may perform testing using over-the-air wireless communications.

[0063] The one or more emulation devices may perform one or more functions, including all functions, while not implemented / deployed as part of a wired and / or wireless communications network. For example, the emulation devices may be utilized in a testing laboratory and / or testing scenario in an undeployed (e.g., testing) wired and / or wireless communications network to perform testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, e.g., one or more antennas) may be used to transmit and / or receive data by the emulation devices.

[0064] Video coding systems may compress digital video signals to, for example, reduce the storage space and / or transmission bandwidth associated with storing and / or distributing such signals. Video coding systems may include block-based systems, wavelet-based systems, object-based systems, etc.

[0065] The video coding device may be based on a block-based hybrid video coding framework. A multi-type tree-based block partition structure may be adopted. Coding modules, such as one or more of an intra-prediction module, an inter-prediction module, a transform / inverse transform module, and a quantization / inverse quantization module, may be included. The video coding device may also include an in-loop filter.

[0066] A video coding device may include one or more coding tools such as, for example, a 65 angular intra prediction direction, modified coefficient coding, advanced multiple transform (AMT) + 4x4 non-separable secondary transform (NSST), affine motion model, generalized adaptive loop filter (GALF), advanced temporal motion vector prediction (ATMVP), adaptive motion vector precision, decoder-side motion vector refinement (DMVR), and / or linear model (LM) chroma mode.

[0067] An exemplary block-based video coding system may include a block-based hybrid video coding framework. FIG. 2 shows an exemplary block-based hybrid video encoding framework 200 for an encoder. As shown in FIG. 2, an input video signal 202 may be processed block by block. A block may be referred to as a coding unit (CU). A CU may also be referred to as a video block. For example, a CU may be up to 128×128 pixels in size. In the coding framework, a CU may be partitioned into prediction units (PUs) and / or separate prediction may be used. In the coding framework, a CU may be used as a basic unit for both prediction and transformation without further partitioning. A CTU may be partitioned into CUs to adapt various local features, for example, based on a quad / binary / ternary tree structure. In a multi-type tree structure, a CTU may be partitioned by a quad tree structure. The leaf nodes of the quad tree may be further partitioned by binary and ternary tree structures. As shown in FIG. 4, one or more partition types may be provided, including, for example, quarternary partition (FIG. 4(a)), horizontal binary partition (FIG. 4(c)), vertical binary partition (FIG. 4(b)), vertical ternary partition (FIG. 4(d)), and horizontal ternary partition (FIG. 4(e)).

[0068] As shown in FIG. 2, spatial prediction 260 and / or temporal prediction 262 may be performed on an input video block (e.g., a macroblock (MB) and / or a CU). Spatial prediction 260 (e.g., intra prediction) may predict a current video block using pixels from samples (e.g., reference samples) of neighboring blocks being coded in a video image / slice. Spatial prediction 260 may, for example, reduce spatial redundancy that may be inherent in a video signal. Motion prediction 262 (e.g., inter prediction and / or temporal prediction) may, for example, use reconstructed pixels from a video image being coded to predict the current video block. Motion prediction 262 may, for example, reduce temporal redundancy that may be inherent in a video signal. A motion prediction signal (e.g., a temporal prediction signal) for a video block (e.g., a CU) may be signaled by one or more motion vectors (MVs). An MV may indicate the amount and / or direction of motion between the current block and / or the current block's reference block or its temporal reference. If multiple reference pictures are supported for a (e.g., each) video block, a reference picture index for that video block may be sent by the encoder. The reference picture index may be used to identify which reference picture in reference picture store 264 from which the motion prediction signal may be derived.

[0069] After spatial prediction 260 and / or motion prediction 262, a mode decision block 280 at the encoder may determine a prediction mode (e.g., a best prediction mode), for example, based on rate-distortion optimization. The prediction block may be subtracted from the current video block at 216, and / or the prediction residual may be decorrelated using transform 204 and / or quantization 206 to achieve a bitrate, such as a target bitrate. The quantized residual coefficients may be inverse quantized at inverse quantization 210 and / or inverse transformed at transform 212, for example, to form a reconstructed residual, which may be added to the prediction block at 226, for example, to form a reconstructed video block. In-loop filtering (e.g., a deblocking filter and / or an adaptive loop filter) may be applied on the reconstructed video block at loop filter 266, after which the reconstructed video block may be placed in reference picture store 264 and / or used to code a video block (e.g., a future video block). To form the output video bitstream 220, the coding mode (e.g., inter or intra), prediction mode information, motion information, and / or quantized residual coefficients may be sent (e.g., all sent) to the entropy coding module 208, e.g., compressed and / or packed to form the bitstream.

[0070] 3 shows a block diagram of an example block-based video decoding framework for a decoder. A video bitstream 302 (e.g., video bitstream 220 in FIG. 2) may be unpacked (e.g., first unpacked) and / or entropy decoded in an entropy decoding module 308. Coding mode and prediction information may be sent to a spatial prediction module 360 ​​(e.g., if intra-coded) and / or a motion-compensated prediction module 362 (e.g., if inter-coded and / or temporally coded) to form a prediction block. Residual transform coefficients may be sent to an inverse quantization module 310 and / or to an inverse transform module 312, for example, to reconstruct a residual block. The prediction block and / or the residual block may be added together at 326. The reconstructed block may undergo in-loop filtering in a loop filter 366, for example, before the reconstructed block is stored in a reference image store 364. The reconstructed video 320 in the reference image store 364 may be sent to drive a display device and / or used to predict video blocks (e.g., future video blocks).

[0071] As described herein, affine motion compensation may be used as an inter-coding tool.

[0072] As described herein, various affine modes and affine motion models for video coding may be used. A translation motion model may be applied to motion-compensated prediction. Various types of motion (e.g., zoom-in or zoom-out, rotation, perspective motion, and / or other irregular motion) may exist. Motion-compensated prediction of an affine transformation (e.g., a simplified affine transformation) may be applied to the prediction. A flag for an inter-coded CU (e.g., each inter-coded CU) may be signaled to indicate, for example, whether translational motion or an affine motion model is applied to the inter prediction.

[0073] The simplified affine motion model may be a four-parameter model. Of the four parameters of this model, two parameters may be used for translational motion (e.g., in the horizontal and vertical directions), one parameter may be used for zoom motion, and one parameter may be used for rotational motion. The horizontal zoom parameter value may be equal to the vertical zoom parameter value. The horizontal rotation parameter value may be equal to the vertical rotation parameter value. The four-parameter motion model may be coded using two motion vectors as a motion vector pair at two control point locations, e.g., the top-left and top-right corner locations of the current video block or current CU. As shown in FIG. 5, the affine motion field of a CU or block is coded using two control point motion vectors (e.g.,

[0074]

number

[0075] ) based on the motion of the control points, the motion field (v x ,v y ) can be determined as follows:

[0076]

number

[0077] where (v 0x ,v 0y ) may be the motion vector of the upper left corner control point, and (v 1x ,v 1y ) may be the motion vector of the control point of the upper right corner. If the block is coded in affine mode, its motion field may be derived, for example, based on the granularity of sub-blocks. The motion vector of a sub-block (e.g., each sub-block) may be derived, for example, by calculating the motion vector of the center sample of the sub-block using equation (1). The motion vector may be rounded to a precision value (e.g., 1 / 16 pel precision). The derived motion vector may be used in a motion compensation stage to generate a prediction signal for a sub-block (e.g., each sub-block) within the current block. The size of the sub-block applied to affine motion compensation may be calculated using the following equation:

[0078]

number

[0079] where (v 2x ,v 2y ) may be the motion vector of the bottom-left control point, w and h may be, for example, the width and height of the CU calculated by equation (1), and M and N may be the width and height of the derived sub-block size.

[0080] Affine merge mode coding may be used to code a CU. Two sets of motion vectors associated with two control points for each reference picture list may be signaled with predictive coding. An affine merge mode may be applied, and the difference between a motion vector and its predictor may be coded using a lossless coding scheme. Signaling overhead, which may be significant (e.g., at low bit rates), may be signaled. For example, an affine merge mode may be applied to reduce signaling overhead by considering local continuity of the motion field. Motion vectors at two control points of the current CU may be derived. The motion vector of the current CU may be derived using the affine motion of affine merge candidates for the CU, which may be selected from its neighboring blocks.

[0081] As shown in FIG. 6, for example, a current CU coded in affine merge mode may have five neighboring blocks (N0 to N4). The neighboring blocks may be checked in the order of N0 to N4, i.e., N0, N1, N2, N3, and N4. The first affine-coded neighboring block may be used as an affine merge candidate. As shown in FIG. 7, a current CU may be coded in affine merge mode. A bottom-left neighboring block (e.g., N0) of the current CU may be selected as an affine merge candidate. The bottom-left neighboring block N0 may belong to a neighboring CU, CU0. The width and height of the CU containing block N0 may be denoted as nw and nh. The width and height of the current CU may be denoted as cw and ch. Position P i The MV in (v ix ,v iy ) at the control point P0. 0x ,v 0y ) may be derived according to the following formula:

[0082]

number

[0083] MV(v 1x ,v 1y ) may be derived according to the following formula:

[0084]

number

[0085] MV(v 2x ,v 2y ) may be derived according to the following formula:

[0086]

number

[0087] Once the MVs at two control points (e.g., P0 and P1) are determined, the MVs of sub-blocks (e.g., each sub-block) within the current CU may be derived. The derived MVs of the sub-blocks may be used for sub-block-based motion compensation and temporal motion vector prediction for future image coding.

[0088] Affine MV prediction may be performed. For non-merged affine-coded CUs, signaling MVs at control points may be associated with high signaling costs. Predictive coding may be used to reduce signaling overhead. An affine MV predictor may be generated from the motion of its neighboring coded blocks. Various types of predictors may be supported for MV prediction of affine-coded CUs. For example, an affine motion predictor generated from neighboring blocks of the control points and / or a translational motion predictor used for MV prediction. The translational motion predictor may be used as a complement to the affine motion predictor.

[0089] A set of MVs can be obtained and used to generate multiple affine motion predictors. As shown in FIG. 8, the MV set includes MVs from neighboring blocks {A, B, C} at corner P0 (which may include set S1, and {MV A ,MV B ,MV C}), MVs from neighboring blocks {D,E} at corner P1 (which may include set S2, {MV D ,MV E}), and / or MVs from neighboring blocks {F,G} at corner P2 (which may include set S3, denoted as {MV F ,MV G9 , the temporal distance between the current image 902 and the reference image 904 of the current CU may be denoted as TB. The temporal distance between the current image 902 and the reference image of the neighboring block 906 may be denoted as TD. The MV1 of the neighboring block may be scaled using the following:

[0090]

number

[0091] Here, MV2 may be used in the motion vector set.

[0092] The collocated blocks in the collocated reference image may be checked, for example, when the neighboring blocks are not inter-coding blocks. The MV may be scaled based on the temporal distance according to Equation (9), for example, when the temporally collocated blocks are inter-coding blocks. The MV in the neighboring blocks may be set to zero, for example, when the temporally collocated blocks are not inter-coding blocks.

[0093] An affine MV predictor may be generated by selecting an MV from a set of MVs. For example, there may be three sets of MVs, e.g., S1, S2, and S3. The sizes of S1, S2, and S3 may be 3, 2, and 2, respectively. In such an example, there may be 12 (e.g., 3×2×2) possible combinations. A candidate MV may be discarded, for example, if the magnitude of a parameter related to zoom or rotation represented by one or more MVs is greater than a threshold. The threshold may be predefined. A combination may be denoted as (MV0, MV1, MV2) with respect to three corners of a CU, e.g., top left, top right, and bottom left. A condition MV may be checked as follows: (|(v 1x -v 0x )│>T*w) or (|(v 1y -v 0y )│>T*h) or (|(v 2x -v 0x )│>T*w) or (|(v 2y -v 0y )│>T*h) (10) where T may be 1 / 2. A candidate MV may be discarded, for example, if a condition is met (e.g., too much zoom or rotation).

[0094] The remaining candidates may be sorted. A triplet of three MVs may represent a six-parameter motion model (e.g., including translation in horizontal and vertical directions, zoom, and rotation). The ordering criterion may be the difference between the six-parameter motion model and the four-parameter motion model, represented by (MV0,MV1). A candidate with a smaller difference may have a smaller index in the ordered candidate list. The difference between the affine motion represented by (MV0,MV1,MV2) and the affine motion model represented by (MV0,MV1) may be evaluated according to the following formula: D=|(v1x -v 0x )*h-(v 2y -v 0y )*w|+|(v 1y -v 0y )*h+(v 2x -v 0x )*w| (11)

[0095] An affine motion model may be used to improve coding efficiency. For example, for a large CU, MVs at two control points may be signaled. Motion vectors for sub-blocks within that CU may be interpolated. The motion for a sub-block (e.g., each sub-block) may differ due to, for example, zoom or rotational motion. The control points may be fixed in an affine motion model, for example, if the coding block chooses an affine motion model instead of a translational motion model. The control points used may be fixed, for example, to the upper-left and upper-right corners of the coding block. The motion vector precision for the affine MVs may be fixed (e.g., 1 / 4 pel). If the vertical position y of a sub-block is larger than the block width (w) in equation (1), upscaling (y / w) may be used.

[0096] While using the affine merge mode, the first available neighboring block from {N0, N1, N2, N3, N4} may not be the best block. From the affine MV derivation from the merge candidate (e.g., provided in equations () to ()), the accuracy may be related to the width of the merge candidate (e.g., indicated by "nw" in equations () to ()). The first affine merge candidate may not have the best accuracy in terms of the affine MV derivation. In affine MV prediction, the condition check based on equation (10) may discard candidates with large zoom or rotation. The discarded candidate may be added back to the list.

[0097] Systems, methods, and means for affine motion model-based coding may be disclosed herein. As disclosed herein, affine motion coding based on adaptive control point selection may be used. In affine motion coding based on adaptive control point selection, control point positions may be adaptively selected based on the shape of a block. For example, one or more control points may be selected based on whether the block is a horizontal rectangular block, a vertical rectangular block, or a square block. Affine merge candidates may be selected from neighboring blocks based on the distance between two control points. For example, an affine merge candidate with the largest control point distance may be selected. Affine predictors may be generated such that candidates with large zoom or rotation motions may be placed at the back of the predictor list.

[0098] Affine motion-based coding with adaptive control point selection may be used. For a video block coded in affine mode, for example, the upper-left and upper-right corners of the video block may be used as control points. The motion of a sub-block (e.g., each sub-block associated with a video block) may be derived using MVs at two control points, e.g., based on equation (1). The derivation accuracy may be related to the block width (e.g., the distance between the two control points). Some sub-blocks may be far from the two control points (e.g., P0 and P1 shown by video block 1010 in FIG. 10). Thus, the derived motion using MVs at P0 and P1 may be affected.

[0099] The selection of control points may be shape-dependent. Video blocks may be classified into categories, such as horizontal rectangular blocks, vertical rectangular blocks, or square blocks. For example, if the width of a block is greater than its height, the block may be classified as a horizontal rectangle. The control points for a horizontal rectangular block may be defined by an upper left corner (e.g., P0) and an upper right corner (e.g., P1), as shown, for example, by block 1010 in FIG. 10. If the width of a block is less than its height, the block may be classified as a vertical rectangle. The control points for a vertical rectangular block may be defined by an upper left corner (e.g., P0) and a lower left corner (e.g., P2), as shown, for example, by block 1020 in FIG. 10. For example, if the width of a block is equal to its height, the block may be classified as a square block. The control points for a square block may be defined by an upper right corner (e.g., P1) and a lower left corner (e.g., P2), as shown, for example, by block 1030 in FIG. 10.

[0100] For a horizontal rectangular block, control points P0 and P1 may be used. The MVs of the sub-blocks for the horizontal rectangular block may be derived based on equation (1).

[0101] For a vertical rectangular block, control points P0 and P2 may be used. The MV of a sub-block for a vertical rectangular block may be derived as follows: If the position of the center of the sub-block relative to the top left corner of the block is denoted by (x,y), and the MV of the sub-block centered at (x,y) is (v x ,v y ) Furthermore, let the block width be denoted as w, the block height be denoted as h, and let the MVs at P0 and P2 be (v 0x ,v 0y ), (v 2x ,v 2y), the MV of a sub-block of a horizontal regular block centered at (x,y) is derived as follows:

[0102]

number

[0103] For a square block, control points P1 and P2 may be used. The MVs of the sub-blocks belonging to the square block may be derived as follows: v x =v 1x +a*(xw)-b*y (14) v y =v 1y +b*(xw)+a*y (15) where a and b may be calculated as follows: a=(-(v 2x -v 1x )*w+(v 2y -v 1y )*h) / (w*w+h*h) (16) b=(-(v 2x -v 1x )*h-(v 2y -v 1y )*w) / (w*w+h*h) (17) If w is equal to h in the square block case, a and b may be simplified to: a=(-(v 2x -v 1x )+(v 2y -v 1y )) / (2w) (18) b=(-(v 2x -v 1x )-(v 2y -v 1y )) / (2w) (19)

[0104] A mode indicating the selection of control points for an affine-coded CU may be signaled. For example, a mode indicating which control points are used for an affine-coded CU may be signaled. For example, the mode may indicate that control points P0 and P1 are used in the horizontal direction, or that control points P0 and P2 are used in the vertical direction, or that control points P1 and P2 are used in the diagonal direction. The control point mode may be determined based on motion estimation cost or rate-distortion cost. For example, for a video block (e.g., each block), the encoder may attempt affine motion estimation using various control point selection modes to obtain a prediction error for each possible control point selection. The encoder may choose the mode with the lowest motion estimation cost, for example, by summing the motion prediction distortion and the control point MV bit cost.

[0105] Affine merge candidate selection using the maximum control point distance may be used. MVs at the control points of the current video block may be derived from the MVs of the merge candidates using equations (3)-(8). The accuracy of the motion vector derivation may depend on the distance between two control points of its neighboring blocks. The distance between two control points may be the width of the block. In the shape-dependent control point selection described herein, the distance between two control points may be measured based on the block shape. The square of the distance between two control points may be used to select an affine merge candidate from a neighboring block, e.g., {N0, N1, N2, N3, N4} shown in FIG. 6. The accuracy of the motion derivation for the current block may be higher, e.g., if the distance is larger. For example, as shown in FIG. 11, the affine merge candidate with the maximum control point distance may be selected to derive the MV. The affine merge candidates in the candidate list (e.g., all affine merge candidates) may be checked in order, e.g., as shown in FIG. 6. As shown in FIG. 11 , at 1102, a candidate neighboring block Nk may be selected from a list of available neighboring blocks. At 1104, the selected neighboring block Nk may be checked for affine mode. After checking that affine mode is enabled for the selected neighboring block Nk, a distance D between two control points may be calculated at 1106. The distance D may be calculated based on the block shape. At 1106, the merge candidate with the maximum control point distance may be selected as the affine merge candidate for the current block to derive MVs at the control points of the current block. At 1110, a check is made that all candidates have been evaluated.

[0106] An affine merge index may be signaled. The distance between two control points may be used to order the available merge candidates in the merge candidate list. The final affine merge candidate may be derived in the following manner: Available affine merge candidates (e.g., all available affine merge candidates) may be obtained from neighboring blocks. For candidates (e.g., each candidate) in the list, the distance between two control points may be calculated. The affine merge candidate list may be ordered, for example, in descending order of control point distance. The final affine merge candidate may be chosen from the ordered list using the merge index signaled for the coding block.

[0107] Affine MV prediction may be performed as described herein. The ordering of candidates in the generation of an affine MV predictor may be performed, for example, by checking condition (10) or by using the criteria provided in the following equation: D=max(|(v 1x -v 0x )*h-(v 2y -v 0y )*w|,|(v 1y -v 0y )*h+(v 2x -v 0x )*w|)+A1+A2 (20) Here, A1 and A2 may be adjustments if the zoom or rotation movement is too large. A1 and A2 may be calculated using the following formula:

[0108]

number

[0109] where T1 and T2 may be predefined thresholds (e.g., T1=3, T2=1 / 4), and w and h may be the width and height of the coding block. Using the ordering criteria provided in equation (20), candidates with large zoom or rotation motion may be placed later in the predictor list.

[0110] A unified control point MV for affine motion compensation, motion vector prediction, and / or deblocking may be used. As described herein, when affine mode is enabled, a CU may be divided into multiple sub-blocks (e.g., 4x4 sub-blocks) having equal sizes. A sub-block (e.g., each sub-block) may be assigned an MV (e.g., one unique MV) that may be derived using the affine mode. For example, the affine mode may be a four-parameter affine mode or a six-parameter affine mode. The affine mode may be signaled at the CU level. The center position of a sub-block (e.g., each sub-block) may be used to derive the corresponding MV of that sub-block based on the selected affine mode. MV of (i,j) sub-block

[0111]

number

[0112] may be derived from three control points MV, v0, v1, and v2, at the top-left, top-right, and bottom-left corners of the affine CU as follows:

[0113]

number

[0114] where (i,j) may be the horizontal and vertical index of a sub-block within a CU, and w sb and hsb may be the width and height of (e.g., one) sub-block (which may be equal to 4, for example). A CU may have one or more sub-blocks that may not include control point positions. For example, the top-left and top-right positions for a four-parameter affine mode and the top-left, top-right, and bottom-left positions for a six-parameter mode may not include control point positions. MVs in such cases may be calculated as provided in Equation (23). These MVs may be used to generate prediction samples for the sub-blocks during motion compensation. The MVs may be used to predict MVs of spatial and temporal neighboring blocks of the CU. The MVs may be used to calculate boundary strength values ​​used for the deblocking filter. For sub-blocks located at control point positions, their MVs may be used as seeds to derive control point MVs for those neighboring blocks through the affine merge mode. To preserve MV accuracy in the affine merge mode, the MVs in Equation (23) may be used in motion compensation for the control point sub-blocks (e.g., each control point sub-block). For spatial / temporal MV prediction and deblocking, the MVs may be replaced by the corresponding control point MVs. For example, for a CU coded by a four-parameter affine model, the MVs of its top-left and top-right sub-blocks, which may be used for MV prediction and deblocking, may be calculated as follows:

[0115]

number

[0116] For a CU coded in six-parameter affine mode, the MVs of the top-left, top-right, and bottom-left sub-blocks that may be used for MV prediction and / or deblocking may be calculated as follows:

[0117]

number

[0118] Figure 12 shows an example of generating a motion field for a CU that may be coded in affine mode. Based on the workflow shown in Figure 12, the MV accuracy of affine motion compensation and MV prediction may be maintained. The workflow shown in Figure 12 may be used in many ways. For example, for a sub-block containing a control point position of a CU (e.g., each sub-block associated with a CU), one or more different MVs may be derived and / or stored. In an example, the MV may be derived based on Equation (23) and may be used to generate a predicted sample of the sub-block. In an example, the MV may be derived based on Equations (24) and (25) and may be used for MV prediction and deblocking.

[0119] For a sub-block at a control point position (e.g., each sub-block), its MV may be set (e.g., initially set) to the corresponding control point MV. The MV may be set to the corresponding control point MV to derive MVs of its neighboring blocks in the analysis stage. In the motion compensation stage, the MV of the sub-block may be recalculated by using the center position as input to a selected affine model. For a sub-block at a control point position (e.g., each control point position), one or more different MVs may be stored. The MV for a sub-block at a control point position (e.g., each control point position) may be derived twice.

[0120] For a CU coded by the affine mode, the motion fields used in separate coding processes may be integrated. For example, as shown in FIG. 13, the MVs used for spatial / temporal MV prediction and deblocking (e.g., as shown by Equations (24) and (25)) may be reused to generate predicted samples of control point sub-blocks in an affine CU. For a sub-block located at the control point position of an affine CU, the MVs derived based on the center position of the sub-block (e.g., according to Equation (23)) may be reused in the motion compensation stage. The MVs may be MV predictors for spatial / temporal MV prediction. The MVs may be used to calculate boundary strengths for the deblocking process. FIG. 14 shows a workflow for deriving a motion field for an affine CU.

[0121] Motion vector clipping may be used. For example, if an affine mode associated with a video block or a CU is enabled, motion vector clipping may be used. If an affine mode is enabled, the CU may be divided into one or more sub-blocks. The sub-blocks associated with a CU may be equal in size (e.g., 4x4). The sub-blocks associated with a CU may be allocated MVs. For example, the MV allocated to each of the CU's MVs may be unique MVs. The allocated MVs may be derived, for example, by using a four-parameter affine mode or a six-parameter affine mode. The type of affine mode (four-parameter affine mode or six-parameter affine mode) may be signaled at the CU level. The derived MVs associated with a CU may be stored in a motion field and may be represented using a limited bit depth (e.g., 16 bits in VVC). When deriving sub-block MVs, the calculated MV values ​​may be outside the range of values ​​that may be represented based on the motion field bit depth. A calculated MV that is outside the range of values ​​can result in arithmetic underflow and / or overflow problems. Such underflow and / or overflow problems can occur even when the control point MV is within the range specified by the bit depth of the motion field. The MV may, for example, be clipped after derivation of the MV. Clipping the MV can result in similar behavior among various systems that may use various bit depth values. For example, a video encoding device may use a bit depth value that may be higher than the bit depth value used by a video decoding device, or vice versa.

[0122] MV of subblock (i,j)

[0123]

number

[0124] may be clipped according to Equation 26 as follows:

[0125]

number

[0126] where N may be the bit depth used for storing the motion field (e.g., N=16). As shown in equation (26), the MV of sub-block (i,j)

[0127]

number

[0128] may be clipped based on the range of the motion field. The range of the motion field may be a motion field storage bit depth (MFSBD) value. The MFSBD may be represented by a number of bits (e.g., 16 bits, 18 bits).

[0129] One or more control point MVs may be clipped based on a bit depth value that may be the same as the bit depth value used for storing the motion field. The control point MVs may be clipped. For example, the control point MVs may be clipped after derivation of the sub-block MVs. The control point MVs may be clipped to preserve the precision of the derived MVs. The control point MVs may have higher precision than the motion field storage bit depth. For example, the control point MVs used for sub-block derivation may have higher precision (e.g., may have more bits) than the range of values ​​that may be represented given the motion field storage bit depth. The control point MVs may be clipped and stored for affine merge derivation of neighboring blocks. For example, the control point MVs may be clipped and stored after derivation.

[0130] Various mechanisms may be used to derive sub-block MVs. For example, planar motion vector prediction and / or regression-based motion vector fields may be used. MVs associated with each sub-block in a CU may be derived from MVs of neighboring blocks of the CU. For example, MVs associated with each sub-block may be derived based on control point MVs of neighboring blocks of the CU. The derived sub-block MVs may be stored in a motion field for future coding. The derived MVs may be clipped based on a motion field storage bit depth value to avoid overflow and / or underflow issues.

[0131] The control point MVs and / or sub-block MVs of an affine-coded CU may be used in MV prediction, for example, in predicting neighboring blocks. The reference region pointed to by the MV may be outside and / or far from the image boundary, for example, even when the MV is clipped based on the motion field storage bit depth.

[0132] The affine control point MVs and / or affine sub-block MVs may be clipped within a range value. The range value may be specified by the image boundary plus a margin to allow for portions of the sub-blocks to be outside the image when deriving the MV of the CU being affine coded. For example, the control point MVs may be clipped and / or scaled (e.g., taking into account an additional margin) so that the resulting reference block after affine motion compensation is bounded by the image boundary. The sub-block MVs may be clipped so that the resulting reference sub-block after motion compensation and the reference image may overlap by at least one sample.

[0133] 15 illustrates an example of modifying one or more control points MV to scale a reference block. As shown in FIG. 15, one or more initial control points MV, v, associated with an initial reference block 1506 i may be modified to scale the reference block to be completely contained within the desired range. For example, the range may be based on the image boundary 1502 plus the margin 1504. As further shown in FIG. 15, the modified control points MV, v m may be determined based on the coordinates of the scaled reference block 1508. In an example, the initial control point MV may be modified to be scaled so that the reference block is completely contained within a range value, which may be based on the image boundary 1502 plus the margin 1504.

[0134] 16 shows an example of modifying the control points MV to include the reference block. As shown in FIG. 16, the initial control points MV, v2 i However, the initial reference block 1606 may be modified to exceed the valid area 1604. The valid area 1604 may be based on the image boundary 1602 + margin 1604. The initial control point MV, v2 imay be modified such that the bottom left corner of the modified reference block 1608 is selected as the intersection point between the initial reference block 1606 and the valid region 1604. The control point MVs (e.g., all control point MVs) may be modified using various mechanisms. Various techniques may be evaluated, and the technique that may yield the best performance may be selected. For example, the sub-block MVs (e.g., each sub-block MV) may be derived from the affine control point MVs. Clipping may be applied to the derived sub-block MVs. For example, clipping may be applied based on the location of the sub-block relative to the image boundary. In an example, the sub-block MVs (e.g., each sub-block MV) may be clipped such that the associated reference sub-block and the reference image overlap by one or more samples. For example, the horizontal component of the sub-block MV may be calculated by using equations (27) and (28):

[0135]

number

[0136] and

[0137]

number

[0138] can be clipped between

[0139]

number

[0140] where W pic and W SB x may be the image width and the sub-block width, respectively. SBmay be the horizontal coordinate of the upper left corner of the sub-block in the image. The vertical component of the sub-block MV can be calculated by using equations (29) and (30):

[0141]

number

[0142] and

[0143]

number

[0144] can be clipped between

[0145]

number

[0146] where H pic and H SB y may be the image height and sub-block height, respectively. o may be the offset for the filtering operation. SB may be the vertical coordinate of the top-left corner of the sub-block in the image. The top-left location of the CU relative to the image boundary may be used (e.g., instead of the location of the sub-block).

[0147] Although features and elements are described above in particular combinations, one of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer. [Explanation of symbols]

[0148] 1502 Image Border 1504 Margin 1506 Initial Reference Block 1508 Reference Blocks

Claims

1. determining that affine mode is enabled for a current block that includes multiple sub-blocks; obtaining a sub-block motion vector associated with a sub-block of the plurality of sub-blocks; clipping the sub-block motion vectors based on bit depth; decoding the sub-block based on the clipped sub-block motion vector; 1. A method for video decoding, comprising:

2. 2. The method of claim 1, wherein the sub-block motion vectors are based on control point motion vectors associated with the current block.

3. the control point motion vector is one of a plurality of control point motion vectors of the current block; The plurality of control point motion vectors are associated with a plurality of control point positions.

3. The method of claim 2.

4. 4. The method of claim 3, wherein the plurality of control point locations include a top-left control point and a top-right control point based on the width of the current block being greater than the height of the current block.

5. 4. The method of claim 3, wherein the plurality of control point locations include a top-left control point and a bottom-left control point based on the width of the current block being less than the height of the current block.

6. 4. The method of claim 3, wherein the plurality of control point locations include a bottom-left control point and a top-right control point based on the width of the current block being equal to the height of the current block.

7. 2. The method of claim 1, wherein the bit depth corresponds to a range of a motion field when storing the motion field on a storage device.

8. The method of claim 7, further comprising: storing the clipped sub-block motion vector for prediction of each sub-block of the plurality of sub-blocks. The method of claim 1 further comprising:

9. 4. The method of claim 3, wherein the bit depth is used for clipping the control point motion vectors.

10. determining that affine mode is enabled for a current block that includes multiple sub-blocks; obtaining a sub-block motion vector associated with a sub-block of the plurality of sub-blocks; clipping the sub-block motion vectors based on bit depth; encoding the sub-block based on the clipped sub-block motion vector; 1. A method for video encoding, comprising:

11. The method of claim 10, wherein the sub-block motion vectors are based on control point motion vectors associated with the current block.

12. the control point motion vector is one of a plurality of control point motion vectors of the current block; The plurality of control point motion vectors are associated with a plurality of control point positions.

12. The method of claim 11 .

13. 13. The method of claim 12, wherein the plurality of control point locations include a top-left control point and a top-right control point based on the width of the current block being greater than the height of the current block.

14. 13. The method of claim 12, wherein the plurality of control point locations include a top-left control point and a bottom-left control point based on the width of the current block being less than the height of the current block.

15. 13. The method of claim 12, wherein the plurality of control point locations include a bottom-left control point and a top-right control point based on the width of the current block being equal to the height of the current block.

16. 11. The method of claim 10, wherein the bit depth corresponds to a range of a motion field when storing the motion field on a storage device.

17. The method of claim 16, further comprising: storing the clipped sub-block motion vector for prediction of each sub-block of the plurality of sub-blocks. The method of claim 10 further comprising:

18. determining that affine mode is enabled for a current block that includes multiple sub-blocks; obtaining a sub-block motion vector associated with a sub-block of the plurality of sub-blocks; clipping the sub-block motion vectors based on bit depth; Decoding the sub-block based on the clipped sub-block motion vector.

1. A non-transitory computer-readable storage medium comprising instructions for causing one or more processors to:

19. The non-transitory computer-readable medium of claim 18, wherein the sub-block motion vectors are based on control point motion vectors associated with the current block.

20. the control point motion vector is one of a plurality of control point motion vectors of the current block; The plurality of control point motion vectors are associated with a plurality of control point positions.

20. The non-transitory computer-readable storage medium of claim 19.

Citation Information

Patent Citations

  • Video encoding / decoding method and apparatus utilizing motion vectors

    JP2014509479A

  • Buffering predictive data in video coding

    JP2014525198A

  • Affine motion vector derivation device, prediction image generation device, moving image decoding device, and moving image coding device

    WO2018061563A1

  • Encoding device, decoding device, encoding method, and decoding method

    WO2018190207A1