Adaptive motion vector precision for affine motion model-based video coding

Adaptive motion vector precision and affine motion models in video coding improve compression efficiency by up to 40% over HEVC, addressing limitations in existing standards through advanced motion vector estimation and modeling.

JP2026015327AActive Publication Date: 2026-01-29INTERDIGITAL VC HOLDINGS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025170197
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-31
Filing Date
2025-10-08
Publication Date
2026-01-29
Estimated Expiration
2039-08-28

Smart Images

  • Figure 2026015327000001_ABST
    Figure 2026015327000001_ABST
Patent Text Reader

Abstract

Systems and methods for video coding using an affine motion model with adaptive precision are described.SOLUTION: In an example, a block of video is encoded in a bitstream using an affine motion model, and the affine motion model is characterized by at least two motion vectors. A precision is selected for each of the motion vectors and the selected precision is signaled in the bitstream. In some embodiments, the precision is signaled by including information identifying the one of the plurality of elements in the selected predetermined precision set in the bitstream. The identified elements indicate respective precisions of motion vectors that characterize the affine motion model. In some embodiments, the used precision set is explicitly signaled in the bitstream, while in other embodiments, the precision set can be inferred from, for example, block size, block shape, or temporal layer.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a standard patent application claiming the benefit under 35 U.S.C. § 119(e) from U.S. Provisional Patent Application No. 62 / 724,500 (filed August 29, 2018), U.S. Provisional Patent Application No. 62 / 773,069 (filed November 29, 2018), and U.S. Provisional Patent Application No. 62 / 786,768 (filed December 31, 2018), all of which are entitled "Adaptive Motion Vector Precision for Affine Motion Model Based Video Coding," and each of these provisional applications is incorporated herein by reference in its entirety. [Background technology]

[0002] background

[0002] Video coding systems are widely used to compress digital video signals to reduce the storage requirements and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block-based systems, wavelet-based systems, and object-based systems, block-based hybrid video coding systems are currently the most widely used and deployed. Examples of block-based video coding systems include international video coding standards such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10, AVC, VC-1, and the latest video coding standard called High Efficiency Video Coding (HEVC). HEVC was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG.

[0003] The first version of the HEVC standard was completed in October 2013 and provides approximately 50% bitrate reduction or equivalent perceptual quality compared to the previous generation video coding standard, H.264 / MPEG AVC. While the HEVC standard offers significant coding improvements over its predecessor, there is evidence that better coding efficiency than HEVC can be achieved with additional coding tools. Based on this, both ITU-T VECG and MPEG have begun research into new coding techniques for future video coding standardization. In October 2015, ITU-T VECG and ISO / IEC MPEG formed the Joint Video Exploration Team (JVET) to begin significant research into advanced technologies that could enable significant improvements in coding efficiency over HEVC. In the same month, a software codebase called the Joint Exploration Model (JEM) was established for future video coding research. The JEM reference software is based on the HEVC Test Model (HM) developed by JCT-VC for HEVC. Any additional proposed coding tools can be integrated into the JEM software and tested using the JVET Common Test Conditions (CTC).

[0004] In October 2017, ITU-T and ISO / IEC issued a joint call for proposals (CfP) for video compression with features beyond HEVC. In April 2018, 22 CfP responses for the standard dynamic range category were accepted and evaluated at the 10th JVET meeting, demonstrating a compression efficiency improvement of approximately 40% over HEVC. Based on these evaluation results, the Joint Video Expert Team (JVET) launched a new project to develop a next-generation video coding standard named VVC (Versatile Video Coding). In the same month, a reference software code base called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. In the initial VTM-1.0 version, most coding modules, including intra-prediction, inter-prediction, transform / inverse transform and quantization / inverse quantization, and in-loop filters, follow the existing HEVC design, except that a multi-type tree-based block partitioning structure is used in VTM. Meanwhile, a separate reference software base called the Benchmark Set (BMS) has been created to facilitate the evaluation of new coding tools. The BMS code base includes a set of coding tools inherited from JEM on top of VTM, which offers higher coding efficiency and moderate implementation complexity. It will be used as a benchmark for evaluating similar coding technologies in the VVC standardization process. Specifically, nine JEM coding tools have been integrated into BMS-1.0, including 65 angular intra prediction directions, modified coefficient coding, advanced multiple transform (AMT) + 4x4 non-separable transform (NSST), affine motion model, generalized adaptive loop filter (GALF), advanced temporal motion vector prediction (ATMVP), adaptive motion vector accuracy, decoder-side motion vector refinement (DMVR), and linear model (LM) chroma mode. Summary of the Invention [Means for solving the problem]

[0005] overview

[0005] Embodiments described herein include methods for use in video encoding and decoding (collectively "encoding"). In some embodiments, a method is provided for decoding video from a bitstream, the method including, for at least one current block in the video, reading from the bitstream information identifying at least a first motion vector predictor and a second motion vector predictor, reading from the bitstream information identifying one of a plurality of precisions in a predetermined precision set, reading from the bitstream at least a first differential motion vector and a second differential motion vector, wherein the first and second differential motion vectors have the identified precisions, generating at least (i) a first control point motion vector from the first motion vector predictor and the first differential motion vector, and (ii) a second control point motion vector from the second motion vector predictor and the second differential motion vector, and generating a prediction of the current block using an affine motion model, wherein the affine motion model is characterized by at least the first control point motion vector and the second control point motion vector.

[0006]

[0006] The multiple precisions in the predetermined precision set may include 1 / 4 pixel precision, 1 / 16 pixel precision, and 1 pixel precision, and the predetermined precision set is different from the predetermined precision set used for non-affine inter coding of the same video.

[0007]

[0007] The affine motion model can be a four-parameter motion model or a six-parameter motion model. When the affine motion model is a six-parameter motion model, the method may further include reading, from the bitstream, information specifying a third motion vector predictor, reading, from the bitstream, a third differential motion vector having the specified accuracy, and generating a third control point motion vector from the third motion vector predictor and the third differential motion vector, where the affine motion model is characterized by the first control point motion vector, the second control point motion vector, and the third control point motion vector.

[0008]

[0008] Information specifying one of multiple precisions may be read from the bitstream on a block-by-block basis, allowing different blocks within an image to use different precisions.

[0009] In some embodiments, the motion vector predictor is rounded to a specified precision. Each of the control point motion vectors may be generated by adding a corresponding differential motion vector to the respective motion vector predictor.

[0010]

[0010] In some embodiments, a prediction of a current block is generated by using an affine motion model to determine a respective sub-block motion vector for each of a plurality of sub-blocks of the current block, and using the respective sub-block motion vectors to generate an inter prediction for each of the sub-blocks.

[0011]

[0011] In some embodiments, the method further includes reading a residual for the current block from the bitstream, and reconstructing the current block by adding the residual to a prediction of the current block.

[0012] Systems and methods for adaptively selecting the precision of affine motion vectors and for performing motion estimation for affine motion models are also described.

[0013]

[0013] In additional embodiments, encoder and decoder systems are provided for performing the methods described herein. The encoder or decoder system may comprise a processor and a non-transitory computer-readable medium storing instructions for performing the methods described herein. Further embodiments include a non-transitory computer-readable storage medium storing video encoded using any of the methods disclosed herein. [Brief explanation of the drawings]

[0014] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B]

[0015] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 2A]

[0016] 1 is a functional block diagram of a block-based video encoder, such as the encoder used in VVC. [Figure 2B]

[0017] FIG. 1 is a functional block diagram of a block-based video decoder, such as the decoder used in VVC. [Figure 3A]

[0018] FIG. 3A is a diagram showing block division using a multi-type tree structure, and is a diagram showing quadtree division. [Figure 3B] FIG. 3B is a diagram showing block division in a multi-type tree structure, and is a diagram showing vertical binary tree division. [Figure 3C] FIG. 3C is a diagram showing block division in a multi-type tree structure, and is a diagram showing horizontal binary tree division. [Figure 3D] FIG. 3D is a diagram showing block division in a multi-type tree structure, and is a diagram showing vertical ternary tree division. [Figure 3E] FIG. 3E is a diagram showing block division in a multi-type tree structure, and is a diagram showing horizontal ternary tree division. [Figure 4A]

[0019] FIG. 4A is a diagram showing a four-parameter affine motion model, which is a diagram showing an affine model. [Figure 4B] FIG. 4B is a diagram illustrating a four-parameter affine motion model, showing sub-block level motion derivation for affine blocks. [Figure 5]

[0020] 1 illustrates affine merge candidates, with candidate availability check order being N0, N1, N2, N3, and N4. [Figure 6]

[0021] FIG. 10 illustrates motion vector derivation at control points of an affine motion model. [Figure 7]

[0022] FIG. 1 illustrates the construction of affine motion vector predictors from motion vectors in blocks {A,B,C}, {D,E}, and {F,G}. [Figure 8]

[0023] FIG. 1 illustrates an example of affine motion vector (MV) temporal scaling for MV predictor generation. [Figure 9]

[0024] FIG. 10 shows neighboring blocks used in the context derivation of block BC. [Figure 10]

[0025] FIG. 10 is a diagram illustrating a mode determination method for CU encoding without division. [Figure 11]

[0026] FIG. 10 is a diagram showing a method for selecting a motion model and precision other than the default precision (1 / 4 pixel for a translational motion model, and (1 / 4 pixel, 1 / 4 pixel) for an affine motion model). [Figure 12]

[0027] FIG. 10 is a diagram illustrating an affine motion estimation method with an accuracy of (p0 pixel, p1 pixel). [Figure 13]

[0028] Figure 1 shows the refinement of MV0 using the 8 nearest positions. Step 1: Select the best position in {P1, P2, P3, P4}. Step 2: If MV0 was updated in step 1, select the best position from the two adjacent positions. [Figure 14]

[0029] 6-parameter affine mode, where V0, V1, and V2 are control points and (MVx,MVy) is the motion vector of the sub-block centered at position (x,y). [Figure 15]

[0030] FIG. 1 illustrates an example of an encoded bitstream structure. [Figure 16]

[0031] FIG. 1 illustrates an exemplary communication system. [Figure 17]

[0032] FIG. 10 illustrates motion vector derivation in sub-blocks of an 8x4 coding unit. [Figure 18]

[0033] FIG. 2 illustrates a method performed by a decoder in some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0015] Exemplary Network for Implementation of the Embodiments

[0034] 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), etc.

[0016]

[0035] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104, CN 106, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, although it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, which may all be referred to as “stations” and / or “STAs,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain contexts), consumer electronic devices, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as UEs.

[0017]

[0036] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node-B, an eNodeB, a Home NodeB, a Home eNodeB, a gNB, an NR NodeB, a site controller, an access point (AP), a wireless router, etc. While the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0018]

[0037] The base station 114a may be part of the RAN 104, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, sometimes referred to as a cell (not shown). These frequencies may be in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one for each sector of the cell. In one embodiment, the base station 114a may utilize multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in desired spatial directions.

[0019]

[0038] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0020]

[0039] More specifically, as mentioned above, the communication system 100 may be a multiple-access system and may utilize one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a and the WTRUs 102a, 102b, and 102c in the RAN 104 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 116 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​UL Packet Access (HSUPA).

[0021]

[0040] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).

[0022]

[0041] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.

[0023]

[0042] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using the principle of dual connectivity (DC). Thus, the radio interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to and from multiple types of base stations (e.g., eNBs and gNBs).

[0024]

[0043] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (e.g., Wireless Fidelity (Wi-Fi)), IEEE 802.16 (e.g., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0025]

[0044] 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point and may utilize any suitable RAT for facilitating wireless connectivity in a local area, such as a workplace, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), and a road. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106.

[0026]

[0045] The RAN 104 may communicate with the CN 106, which may be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput, delay, error tolerance, reliability, data throughput, and mobility requirements. The CN 106 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 and / or CN 106 may communicate directly or indirectly with other RANs that utilize the same RAT as the RAN 104 or a different RAT. For example, in addition to connecting to the RAN 104, which may utilize NR radio technology, the CN 106 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or Wi-Fi radio technology.

[0027]

[0046] The CN 106 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may utilize the same RAT as the RAN 104 or a different RAT.

[0028]

[0047] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that may utilize a cellular-based wireless technology and with a base station 114b that may utilize an IEEE 802.11 wireless technology.

[0029]

[0048] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any subcombination of the above elements while remaining consistent with an embodiment.

[0030]

[0049] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0031]

[0050] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0032]

[0051] 1B, the transmit / receive element 122 is shown as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0033]

[0052] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.

[0034]

[0053] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, but is instead located on a server or home computer (not shown), for example.

[0035]

[0054] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control the power to other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0036]

[0055] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may acquire location information using any suitable location determination method while remaining consistent with an embodiment.

[0037]

[0056] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0038]

[0057] The WTRU 102 may include a full-duplex radio in which transmission and reception of some or all of the signals associated with a particular subframe (e.g., on both the UL (e.g., for transmission) and downlink (e.g., for reception)) may occur in parallel and / or simultaneously. The full-duplex radio may include an interference management unit to reduce and or substantially eliminate self-interference either through hardware (e.g., a choke) or signal processing by a processor (e.g., via another processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe on either the UL (e.g., for transmission) or downlink (e.g., for reception)).

[0039]

[0058] Although the WTRUs are depicted in FIGS. 1A-1B as wireless terminals, it is contemplated that in certain representative embodiments such terminals may use a wired communication interface (e.g., temporarily or permanently) with a communication network.

[0040]

[0059] In a representative embodiment, the other network 112 may be a WLAN.

[0041]

[0060] 1A-1B and the corresponding description, one or more, or all, of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more, or all, of the functions described herein. For example, the emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0042]

[0061] The emulation device may be designed to perform one or more tests of other devices in a lab environment and / or a carrier network environment. For example, one or more emulation devices may be fully or partially implemented and / or deployed as part of a wired and / or wireless communication network and perform one or more, or all, of the functions to test other devices in the communication network. One or more emulation devices may be temporarily implemented / deployed as part of a wired and / or wireless communication network and perform one or more, or all, of the functions. The emulation device may be directly coupled to another device for testing purposes and / or may perform testing using wireless communication.

[0043]

[0062] One or more emulation devices may perform one or more functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in undeployed (e.g., testing) wired and / or wireless communication networks to perform testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, e.g., one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0044] Detailed Description Block-Based Video Coding

[0063] Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 2A shows a block diagram of an example of a block-based hybrid video coding system. An input video signal 103 is processed in blocks (called coding units (CUs)). In VTM-1.0, a CU can be up to 128x128 pixels. However, unlike HEVC, which divides blocks based only on quadtrees, VTM-1.0 divides one coding tree unit (CTU) into CUs based on quadtrees, binary trees, and ternary trees to adapt to various local characteristics. In addition, the concept of multiple division unit types in HEVC has been removed. That is, the distinction between CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC-1.0. Instead, each CU is always used as the basic unit for both prediction and transformation without further division. In the multi-type tree structure, one CTU is first divided by a quadtree structure. Then, each quadtree leaf node can be further divided by binary tree and ternary tree structures. As shown in Figures 3A-3E, there are five partitioning types: quadtree partitioning, horizontal binary tree partitioning, vertical binary tree partitioning, horizontal ternary tree partitioning, and vertical ternary tree partitioning. In Figure 2A, spatial prediction (160) and / or temporal prediction (162) may be performed. Spatial prediction (or "intra-prediction") predicts the current video block using pixels from samples of already coded neighboring blocks (called reference samples) within the same video image / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") predicts the current video block using reconstructed pixels from already coded video images. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference.If multiple reference pictures are supported, a reference picture index is additionally transmitted, and the reference picture index is used to identify which reference picture in the reference picture store (164) the temporal prediction signal comes from. After spatial and / or temporal prediction, a mode decision block (180) in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted (117) from the current video block, and the prediction residual is decorrelated using a transform (105) and quantized (107). The quantized residual coefficients are inverse quantized (111) and inverse transformed (113) to form a reconstructed residual, which is then added back (127) to the prediction block to form a reconstructed signal for the CU. Further in-loop filtering, such as a deblocking filter, may be applied to the reconstructed CU (166) before it is placed in the reference picture store (164) and used to encode future video blocks. To form the output video bitstream 121, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit (109) for further compression and packing to form the bitstream.

[0045]

[0064] 2B shows a block diagram of an example block-based video decoder. A video bitstream 202 is first unpacked and entropy decoded in an entropy decoding unit 208. The coding mode and prediction information are sent to a spatial prediction unit 260 (if intra-coded) or a temporal prediction unit 262 (if inter-coded) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit 210 and an inverse transform unit 212 to reconstruct a residual block. The prediction block and the residual block are then summed together at 226. The reconstructed block may further undergo in-loop filtering before it is stored in a reference picture store 264. The reconstructed video in the reference picture store may then be sent to drive a display device or used to predict future video blocks.

[0046]

[0065] As mentioned above, BMS-1.0 follows the same encoding / decoding workflow of VTM-1.0 as shown in Figures 2A and 2B. However, some coding modules, especially those related to temporal prediction, are further extended and enhanced. In the following, we briefly describe affine motion compensation as one inter-coding tool included in BMS-1.0 or the previous JEM.

[0047] Affine Mode

[0066] In HEVC, only the translational motion model is applied to motion compensated prediction. Meanwhile, in practice, there are many types of motion, such as zoom in / out, rotation, projective motion, and other irregular motion. In BMS, a simplified affine transformation motion compensated prediction is applied. A flag of each inter-coded CU is signaled to indicate whether the translational motion or the affine motion model is applied to inter prediction.

[0048]

[0067] The simplified affine motion model is a four-parameter model, namely, two parameters for translation in the horizontal and vertical directions, one parameter for zoom motion, and one parameter for rotation motion. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. The four-parameter affine motion model is coded in BMS using two motion vectors as a pair at two control point positions defined at the top-left and top-right corners of the current CU. As shown in Figure 4A, the affine motion field of a block is described by two control point motion vectors (V0, V1). Based on the control point motion, the motion field (v x ,v y ) is expressed as follows:

number

number

[0049]

[0068] The four affine model parameters can be estimated iteratively. The MV pair at step k is

number

number

number

[0050]

[0069] Based on the optical flow equation, the relationship between the change in luminance and the spatial gradient and temporal movement can be formulated as follows:

number

[0051]

[0070]

number

number

[0052]

[0071] Since all samples in the CU satisfy equation (7), the parameter set (a, b, c, d) can be solved using the least squares method.

number

[0053]

[0072] As shown in Figure 14, a 6-parameter affine coded CU has three control points: top left, top right, and bottom left. The movement of the top left control point is a translational movement, the movement of the top right control point is related to the rotation and zoom movement in the horizontal direction, and the movement of the bottom left control point is related to the rotation and zoom movement in the vertical direction. For a 4-parameter affine motion model, the rotation and zoom movement in the horizontal and vertical directions are the same. The motion vector (MV) of each sub-block is x ,MV y ) is derived using the three MVs at the control points as follows:

number

[0054] Affine Merge Mode

[0073] When a CU is coded in affine mode, two sets of motion vectors for two control points in each reference list are signaled in predictive coding. The difference between the MV and its predictor is losslessly coded, and this signaling overhead can be significant, especially at low bitrates. To reduce the signaling overhead, affine merge mode is also applied to BMS by considering the local continuity of the motion field. The motion vectors at the two control points of the current CU are derived using the affine motion of affine merge candidates selected from neighboring blocks. When the current CU is coded in affine merge mode, five neighboring blocks are checked in the order N0 to N4, as shown in Figure 5. The first affine-coded neighboring block is then used as the affine merge candidate. For example, as shown in Figure 6, the current CU is coded in affine merge mode, and its lower-left neighbor (N0) is selected as the affine merge candidate. The width and height of the CU containing block N0 are denoted by nw and nh. The width and height of the current CU are denoted by cw and ch. The position P i The MV in (v ix ,v iy ) at the control point P0. 0x ,v 0y ) is derived as follows:

number

[0055]

[0074] MV(v 1x ,v 1y ) is derived as follows:

number

[0056]

[0075] MV(v 2x ,v 2y ) is derived as follows:

number

[0057]

[0076] After the MVs at the two control points (P0 and P1) are derived, the MVs of each sub-block in the current CU are derived as described above, and the derived sub-block MVs can be used for sub-block-based motion compensation and temporal motion vector prediction for future image coding.

[0058] Affine MV prediction

[0077] For non-merged affine coded CUs, signaling MVs at control points is costly, so predictive coding is used to reduce the signaling overhead. In BMS, affine MV predictors are generated from the motion of neighboring coded blocks. There are two types of predictors for MV prediction of affine coded CUs: (a) affine motion generated from neighboring blocks of control points, and (b) translational motion used for traditional MV prediction. (b) is used only when the number of affine predictors from (a) is insufficient (less than 2 in BMS).

[0059]

[0078] Three sets of MVs are used to generate multiple affine motion predictors. As shown in Figure 7, the three MV sets are: (1) {MV A ,MV B ,MV C}, (2) MVs from adjacent blocks {A, B, C} at corner P0 composed of set S1, denoted as {MV D ,MV E}, and (3) MVs from adjacent blocks {D,E} in corner P1 composed of set S2, denoted as {MV F ,MV GMV1 is derived from the neighboring blocks {F, G} at corner P2, which are composed of set S3, denoted as {F, G}. The MV from the neighboring blocks is derived as follows: First, spatially neighboring blocks are checked. If the neighboring block is an inter-coded block, the MV is used directly. The reference image of the neighboring block is the same as the reference image of the current CU, or if the reference image of the neighboring block is different from the reference image of the current CU, the MV is scaled according to the temporal distance. As shown in Figure 8, the temporal distance between the current image and the reference image of the current CU is denoted as TB, and the temporal distance between the current image and the reference image of the neighboring block is denoted as TD. The MV1 of the neighboring block is scaled as follows:

number

[0060]

[0079] MV2 is used in the motion vector set.

[0061]

[0080] If the neighboring block is not an inter-coded block, the co-located block in the co-located reference image is checked. If the temporally co-located block is an inter-coded block, the MV is scaled based on the temporal distance using equation (18). If the temporally co-located block is not an inter-coded block, the MV of that neighboring block is set to zero.

[0062]

[0081] After the three sets of MVs are obtained, an affine MV predictor is generated by selecting one MV from each of the three sets. The sizes of S1, S2, and S3 are 3, 2, and 2, respectively. In total, 12 (3 x 2 x 2) combinations are obtained. In BMS, if the parameters related to zoom or rotation represented by the three MVs are greater than a predefined threshold, the candidate is discarded. For the three corners of the CU, i.e., top left, top right, and bottom left, one combination is denoted as (MV0, MV1, MV2). The following conditions are checked: (|(v 1x-v 0x )|>T*w) or (|(v 1y -v 0y )|>T*h) or (|(v 2x -v 0x )|>T*w) or (|(v 2y -v 0y )|>T*h) (17) T is (1 / 2). If the condition is met, i.e. the zoom or rotation is too large, the candidate is discarded.

[0063]

[0082] In BMS, all remaining candidates are sorted. A triplet of three MVs represents a six-parameter motion model including horizontal and vertical translation, zoom, and rotation. The ordering criterion is the difference between this six-parameter motion model and a four-parameter motion model, denoted by (MV0,MV1). The candidate with the smaller difference has a smaller index in the ordered candidate list. The difference between the affine motion represented by (MV0,MV1,MV2) and the affine motion model represented by (MV0,MV1) is evaluated by Equation (18): D=|(v 1x -v 0x )*h-(v 2y -v 0y )*w|+|(v 1y -v 0y )*h+(v 2x -v 0x )*w| (18)

[0064] Affine MV coding

[0083] When a CU is coded in affine mode, it can be in affine merge mode or affine non-merge mode. In the above affine merge mode, the affine MV at a control point is derived from the affine MV of a neighboring affine-coded CU. Therefore, there is no need to signal MV information for the affine merge mode. In the affine non-merge mode, the MV at the control point is coded using differential coding. The MV predictor is generated using neighboring MVs as described above, and the difference between the current MV and its predictor is coded. The signaled differential MV is called the MVD. Because the affine four-parameter model has two control points, two MVDs are used to signal unidirectional prediction, and four MVDs are used to signal bidirectional prediction. Because the affine six-parameter model has three control points, three MVDs are used to signal unidirectional prediction, and six MVDs are used to signal bidirectional prediction. MVDs are two-dimensional vectors (horizontal and vertical components) and are coded losslessly, making them difficult to compress. In the current VVC design (VTM-1.0 / BMS-1.0), the accuracy of the MVD for signaling is 1 / 4 pixel accuracy.

[0065] Adaptive MVD Accuracy

[0084] For CUs coded as non-merged and non-affine inter modes, the MVD between the current CU's MV and its predictor can be coded at different resolutions. This can be either quarter-pixel, full-pixel, or four-pixel precision. Quarter-pixel precision is fractional precision. Full-pixel and four-pixel precision both belong to integer precision. Precision is signaled by two flags for each CU to indicate the MVD precision. The first flag indicates whether the precision is quarter-pixel precision or not. If the precision is not quarter-pixel precision, the second flag is signaled to indicate full-pixel or four-pixel precision. In motion estimation, differential MVs are usually searched around an initial MV, which is treated as the starting position. The starting position can be selected from its spatial and temporal predictors. For ease of implementation, the starting MV is rounded to the precision for MVD signaling, and only MVD candidates with the desired precision are searched. The MV predictor is also rounded to the MVD precision. In the VTM / BMS reference software, the encoder checks the rate-distortion (RD) cost of different MVD precisions and selects the optimal MVD precision with the smallest RD cost. The RD cost is calculated by the weighted sum of the distortion of the sample values ​​and the coding rate, and is a measure of coding performance. A coding mode with a lower RD cost has better overall coding performance. To reduce signaling overhead, MVD precision-related flags are signaled only if the signaled MVD is non-zero. If the signaled MVD is zero, quarter-pixel precision is inferred.

[0066] MVD encoding

[0085] In VVC, the MVD entropy coding method is the same for both affine and non-affine coding modes. This method encodes two components independently. The MVD sign of each component is coded with one bit. The absolute value coding is done in two parts: (1) The values ​​0 and 1 are coded with flags. The first flag indicates whether the absolute value is greater than 0, and if the value is greater than 0, the second flag indicates whether the absolute value is greater than 1. (2) If the absolute value v is greater than 1, the remaining part (v-2) is binarized with a first-order Exponential-Golomb (EG) code, and these binarized bins are coded with a fixed-length code. For example, the binarization of the remaining part (v-2) using a first-order EG code is shown in Table 1.

[0067] [Table 1]

[0068] EG codes with different degrees may have different codeword lengths for the same value to be coded. The smaller the degree, the shorter the codeword length for small values ​​and the longer the codeword length for large values. In affine coding mode, the statistics of the MVD of the control points may be different. EG codes with the same degree may not be optimal for MVD coding of all control points.

[0069] Issues Addressed in Some Embodiments

[0086] As mentioned above, MVD signaling incurs significant signaling overhead for explicit affine-coded CUs compared to inter-CUs coded with a translational motion model due to the large number of signaled MVDs, i.e., two sets of MVDs for a four-parameter affine model and three sets of MVDs for a six-parameter affine model. Adaptive MVD precision for signaling helps achieve a better tradeoff between motion compensation efficiency and signaling overhead. However, the use of control point motion vectors in an affine model differs from that of a conventional translational motion model: MVs at control points are not used directly for motion compensation but are used to derive MVs for subblocks, which are then used for motion compensation for the subblocks.

[0070]

[0087] The motion estimation (ME) process for the affine motion model described above differs from the motion search method for a conventional translational motion model in VTM / BMS. The ME process used to find the optimal MV at two control points is based on the estimation of an optical flow field. The differential MV derived from the optical flow estimation differs for each iteration, making it difficult to control the step size for each iteration. In contrast, ME for a translational motion model to find the optimal MV for a coding block is typically a position-by-position search method within a specific range. Within a search range centered on a starting MV, the ME cost for each possible position, such as in a full search scheme, is evaluated and compared, and the optimal position with the smallest ME cost can then be selected. The ME cost is typically evaluated as a weighted sum of the prediction error and bits of MV-related signaling, including the reference image index and MVD. The prediction error can be measured by the sum of absolute differences (SAD) between the original signal and the predicted signal of the coding block.

[0071]

[0088] In this determination ME process for a translational motion model, there are many fast search methods for adaptively adjusting the search step size during the iteration. For example, the search can start with a coarse step size within a search window. Once an optimal position is obtained with the coarse step size, the step size can be reduced, and the search window is also reduced to a smaller window centered on the last optimal position obtained from the previous search window. This iterative search can be terminated when the search step size is reduced to a value equal to or less than a predefined threshold, or when the total number of searches meets a predefined threshold.

[0072]

[0089] The ME process for affine models is different from the ME process for translational models. This disclosure describes ME methods for affine models with different MVD precision.

[0073] Overview of Some Embodiments

[0090] To provide motion estimation for affine models, this disclosure describes an adaptive MVD accuracy method that improves coding efficiency of affine motion models. Some embodiments provide an improved tradeoff between signaling and motion compensation prediction efficiency. A decision method for adaptive MVD accuracy is also proposed.

[0074]

[0091] In some embodiments, the MVD precision for the affine model is adaptively selected from multiple precision sets for the two control points. The precision of the MVD at different control points may be different.

[0075]

[0092] In some embodiments, a MV search method for affine models at different MVD precisions is proposed to improve accuracy and reduce coding complexity.

[0076]

[0093] In some embodiments, the affine control point motion vector predictor (MVP) and MV are kept highly accurate, but the MVD is rounded to lower precision, which allows for improved accuracy of motion compensation using high-precision MVs.

[0077]

[0094] For simplicity, the following description illustrates the use of a four-parameter affine motion model as an example, although the proposed method can also be directly extended to a six-parameter affine motion model.

[0078] Adaptive MVD accuracy for affine models.

[0095] In VTM / BMS, MVD at control points for an affine model is always signaled with quarter-pixel precision. Fixed precision does not provide a good tradeoff between MVD signaling overhead and the efficiency of affine motion compensation. Increasing the precision of the MVD at those control points makes the MV derived by Equation (1) for each sub-block more accurate. Thus, motion prediction can be improved. However, it uses more bits for MVD signaling. In this disclosure, a method for adaptive MVD precision at control points is proposed. The motion of the top-left control point is related to the translational motion of each sub-block within that CU, and the motion differential between two control points is related to the zoom and rotational motion of each sub-block. Blocks coded with an affine motion model may have different motion characteristics. Some affine blocks may have translational and rotational motion with high precision, while some affine blocks may have translational motion with low precision. In some embodiments, the translational and rotational / zooming motions of an affine block may have different precisions. Based on this, some example embodiments signal different precision for MVD encoding at different control points.

[0079]

[0096] Signaling the precision for each control point separately increases the signaling overhead for affine-coded CUs. One embodiment is to signal the precision of two control points together. Only frequently used combinations are signaled. For example, a precision pair (prec0, prec1) can be used to indicate the precision "prec0" of the top-left control point and the precision "prec1" of the top-right control point. An exemplary embodiment uses four precision sets: S1{(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel)}, S2{(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 8 pixel)}, S3 {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 8 pixel)}, and S4{(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 8 pixel), (1 / 8 pixel, 1 / 8 pixel)}.

[0080]

[0097] (1 / 4 pixel, 1 / 4 pixel) precision is used for affine blocks as normal precision. (1 pixel, 1 / 4 pixel) is used for affine blocks with low precision and translation, but rotation / zoom still has normal precision. (1 / 4 pixel, 1 / 8 pixel) is used for affine blocks with high precision and rotation / zoom. (1 / 8 pixel, 1 / 8 pixel) is used for affine blocks with both translation and rotation / zoom with high precision. The precision set can be signaled, for example, in the sequence parameter set, picture parameter set, or slice header.

[0081]

[0098] In some embodiments, the precision of one control point is applied to the MVDs of the two lists if the current affine CU is coded in bi-predictive mode. In some embodiments, to reduce signaling redundancy, precision is signaled only if the MVD at that control point is non-zero. If the MVD at a control point is zero, precision information does not need to be signaled for that control point because precision does not affect the zero MVD. For example, if the top-left MVD is zero, (1 pixel, 1 / 4 pixel) precision is not valid for the current CU. Therefore, in this case, there is additional precision signaling when the precision set is S1. (1 / 4 pixel, 1 / 4 pixel) and (1 / 8 pixel, 1 / 8 pixel) are valid when the precision set is S3. The precision of the zero MVD may be inferred as a default precision, such as (1 / 4 pixel, 1 / 4 pixel). Another embodiment may always signal precision even if the MVD is zero because that predictor may lead to a high-precision MVD. For example, the MV predictor is derived from neighboring affine-coded CUs, and higher precision leads to higher precision MV predictors, and thus higher final MV precision.

[0082]

[0099] For binarization of these precision sets, Tables 2, 3, 4, and 5 are proposed, and the binarized bins are coded.

[0083] [Table 2]

[0084] [Table 3]

[0085] [Table 4]

[0086] [Table 5]

[0087]

[0100] Regarding precision encoding, we use S3 as an example. According to Table 4, there are two bins to be encoded for the S3 set after binarization. The second bin is encoded only if the first bin is 0. The bins are encoded with context-adaptive binary arithmetic coding (CABAC). The context of a bin in CABAC is used to record the probability of 0 or 1. The context of the first bin can be derived from its left and top neighbors, as shown in Figure 9. We define two functions: (1) Model(CU), which indicates whether the motion model of the CU is an affine model, and (2) Prec(CU), which indicates whether precision (full pixel, quarter pixel) is used for the CU.

number

[0088]

[0101] The accuracy of the neighboring CU and the current CU is compared, and two flags, namely, equalPrec(B L ), equalPrec(B A ) is obtained.

number

[0089]

[0102] The index of the context of the first bin is constructed as in equation (23). Context_idx(B C )=equalPrec(B A )+equalPrec(B L ) (twenty three)

[0090]

[0103] The second bin can be coded using one fixed context or can be coded with a 1-bit fixed length code.

[0091]

[0104] Alternatively, the one-pixel precision of the top-left control point in the precision pair-based signaling scheme described above can be replaced with half-pixel precision.

[0092]

[0105] Another embodiment is to signal the precision for each control point individually. For example, one precision selected from the set {1 pixel, 1 / 4 pixel, 1 / 8 pixel} is signaled for the top-left control point, and one precision selected from the set {1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel} is signaled for the top-right control point. The reason for the different precision sets for the two control points is that 1-pixel precision is too coarse for the top-right MV associated with rotation and zoom motion, because rotation and zoom motion have more complex warping effects than translation motion. If the affine block has translation motion with low precision, the top-left control point can select 1-pixel precision; if the affine block has translation motion with high precision, the top-left control point can select 1 / 8-pixel precision. If the affine block has rotation or zoom motion with high precision, the top-right control point can select 1 / 8-pixel precision. Based on the statistical values, the following binarization tables (Tables 6 and 7) can be used to encode the precision selected for the two control points. The binarized code is a codeword, which can be encoded using various entropy coding methods, such as CABAC. At the decoder side, the affine MV predictor at each control point is rounded to the precision of the MVD and scaled to high precision for MV field storage (e.g., 1 / 16 pixels for VVC). The decoded MVD is first scaled to high precision for MV field storage based on the precision. Then, the scaled MVD is added to the MV predictor to obtain a reconstructed MV with the precision used for motion field storage. The reconstructed MV at the control points is used to derive the MV of each subblock using equation (1) for motion compensation of each subblock to obtain the sample value prediction of that subblock.

[0093] [Table 6]

[0094] [Table 7]

[0095]

[0106] In another embodiment, the precision set of both control points may be the same {½ pixel, ¼ pixel, ⅛ pixel}, etc., but the binarization of the precision encoding of the two control points may be different. An example of the binarization of the precision encoding of the two control points is proposed in Table 8.

[0096] [Table 8]

[0097]

[0107] In some embodiments, control point precision control is applied only to large CUs to save signaling overhead. This is because affine motion models are typically used more frequently for large CUs. For example, in some embodiments, the MVD precision of control points may be signaled only if the CU has an area larger than a threshold (e.g., 16x16). For small CUs, precision may be estimated to (1 / 4 pixel) for both control points.

[0098]

[0108] In some embodiments, the precision set is changed at the image level. In a random access configuration, there may be different temporal layers, and different quantization parameters (QPs) may be used at different layers. For example, for low temporal layer images with a small QP, there may be more precision choices, and a higher precision, such as 1 / 8 pixel, may be preferred. Also, the precision set {1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel} may be used. For high temporal layer images with a large QP, there may be fewer precision choices, and a lower precision, such as 1 pixel, may be preferred. Also, the precision set {1 pixel, 1 / 4 pixel} or {1 pixel, 1 / 2 pixel, 1 / 4 pixel} may be used.

[0099]

[0109] For a six-parameter affine model, the motion at the top left is associated with translation, the difference in motion between the top right and top left is associated with horizontal rotation and zoom, and the difference in motion between the bottom left and top left is associated with vertical rotation and zoom. A precision triplet (p0, p1, p2) is specified for the six-parameter affine model, where p0, p1, and p2 are the precisions of the top left, top right, and bottom left control points. One embodiment is to set the same precision for MVD signaling for both the top right and bottom left control points. For example, the precision of the three control points can be one of the set {(1 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 8 pixel, 1 / 8 pixel)}. Another embodiment is to set different precisions for the top right and bottom left control points. To reduce signaling overhead, it is beneficial to have as few precision set choices as possible. In some embodiments, the precision set is selected based on the shape of the CU. If the width is equal to the height (i.e., for a square CU), the precision of the top right and bottom left can be the same, e.g., the precision set is {(1 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 8 pixel, 1 / 8 pixel)}. If the width is greater than the height (i.e., for a long CU), the precision of the top right control point can be equal to or greater than the precision of the bottom left control point, e.g., the precision set is {(1 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 8 pixel, 1 / 4 pixel)}. If the width is smaller than the height (i.e., for a tall CU), the accuracy of the top right control point can be less than or equal to the accuracy of the bottom left control point, for example, the accuracy set is {(1 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 4 pixel, 1 / 8 pixel)}.

[0100]

[0110] An example of a method performed by a decoder in some embodiments is shown in Figure 18. The decoder receives a bitstream (block 1802) and reads from the bitstream information identifying at least a first motion vector predictor (block 1804) and a second motion vector predictor (block 1806), information identifying one of a plurality of precisions in a predetermined precision set (block 1808), and at least a first differential motion vector (block 1810) and a second differential motion vector (block 1812). The first and second differential motion vectors have the precision identified by the information read in block 1808. The syntax and semantics by which the information is encoded into the bitstream may vary in different embodiments. The decoder generates a first control point motion vector (block 1814) from at least the first motion vector predictor and the first differential motion vector, and a second control point motion vector (block 1816) from the second motion vector predictor and the second differential motion vector. The decoder then generates a prediction for the current block using an affine motion model (block 1818). The affine motion model is characterized by at least a first control point motion vector and a second control point motion vector.

[0101] Motion estimation for affine motion models with adaptive MVD precision

[0111] When adaptive MVD precision is applied to the two affine control points, the encoder operates to determine an optimal precision that affects the coding performance of the affine motion model. The encoder also operates to apply a good motion estimation method at a given precision to determine the affine model parameters.

[0102]

[0112] In VVC, the CU mode decision flow chart is shown in Figure 10, where the encoder checks different coding modes and selects the best coding mode with the smallest RD cost. For explicit inter modes with different translational motion model precisions, i.e., 1 / 4 pixel, 1 pixel, and 4 pixel, there are three RD cost check processes. To reduce coding complexity, the 4-pixel precision-based RD cost is calculated only if the 1 pixel precision RD cost is smaller than or comparable to the 1 / 4 pixel RD cost. In the 1 / 4 pixel precision RD cost calculation process, the encoder compares the motion estimation costs of the translational motion model and the affine motion model and selects the motion model with the smallest ME cost. The precision of the affine motion model is (1 / 4 pixel, 1 / 4 pixel) for two control points.

[0103]

[0113] In some embodiments, more precision is introduced for the adaptive MVD precision of the affine motion model. For example, in addition to precision (¼ pixel, ¼ pixel), (1 pixel, ¼ pixel), (⅛ pixel, ⅛ pixel) are added to the affine model. The following description uses these three precisions for the affine model as an example. However, other embodiments may use other precisions or combinations of more precisions. The (¼ pixel, ¼ pixel) precision of the affine model may be used as the default precision. To reduce complexity, the ¼ pixel RD cost check process is maintained, in which the affine model with (¼ pixel, ¼ pixel) precision is evaluated. The remaining affine precision checks are added to the 1 pixel RD cost check.

[0104]

[0114] FIG. 11 shows a flow diagram of an embodiment using RD cost check with 1-pixel accuracy. One motion estimation for a translational motion model at 1-pixel accuracy (block 1102) and two affine motion estimations at (1-pixel, 1 / 4-pixel) accuracy (block 1104) and (1 / 8-pixel, 1 / 8-pixel) accuracy (block 1106) are performed, respectively. A motion model and corresponding accuracy are selected by comparing their ME costs (block 1108). To reduce coding complexity, affine motion estimation at these two accuracy levels is performed only if the current best mode is an inter-coding mode with an affine motion model, after the encoder has already checked the (1 / 4-pixel, 1 / 4-pixel) accuracy for the affine model. This is because different affine model accuracy levels are only valid when the current CU has affine motion. To further reduce coding complexity, in some embodiments, the encoder may check the ME cost of the affine model only if the current best coding mode is an affine non-merge mode or an affine non-skip mode. This is because merge and skip modes indicate that the current CU is already efficiently coded and the improvement may be very limited.

[0105]

[0115] The (1 pixel, 1 / 4 pixel) and (1 / 2 pixel, 1 / 4 pixel) precisions are lower than the default precision (1 / 4 pixel, 1 / 4 pixel). The precision of the upper-left control point is coarse, making it easy for the encoder to obtain a local minimum. Therefore, we propose a hybrid search method for this type of low precision. Figure 12 shows a flow chart of an example of the search method.

[0106]

[0116] The optical flow-based iterative search described above in the "Affine Mode" section is first applied. Next, (MV0, MV1) is obtained as input for the next step, where MV0 is the MV of the top-left control point and MV1 is the MV of the top-right control point (block 1202). The next step is to refine MV0 by checking the eight nearest neighboring positions (block 1204). An example is shown in Figure 13. If P0 is the position pointed to by MV0, P0 has eight nearest neighboring positions. The distances between P0 and P4 and P1 are at the accuracy of MV0, such as one pixel or half pixel. If MV0 is changed to point to a neighboring position, the corresponding MV1 is estimated using the optical flow-based search method, and the ME cost is calculated using the updated (MV0, MV1). These eight nearest neighboring positions are grouped into two groups. The first group is the four nearest neighboring positions {P1, P2, P3, P4}, and the second group is {P5, P6, P7, P8}. First, the ME cost at position P0 is compared with the ME costs of adjacent positions from {P1, P2, P3, P4}. If the cost of P0 is the smallest, the refinement of MV0 stops. If the ME costs of adjacent positions in the first group are lower than the ME cost of P0, the other two adjacent positions from {P5, P6, P7, P8} are further compared. For example, if the cost of P2 is the smallest in the first round, P5 and P6 are further checked. Thus, the maximum number of cost checks is 6, not 8.

[0107]

[0117] Once MV0 is determined, MV1 is further refined (block 1204). The refinement is an iterative search in a square pattern. For each iteration, there is a central position that was the best position in the last iteration. The encoder calculates the ME cost at eight neighboring positions, compares it with the current best ME cost, and moves the central position to a new position with the smallest ME cost among the central position and the eight neighboring positions. If a neighboring position has already been checked in the previous iteration, the check of that position is skipped in the current iteration. If there is no update in the current iteration, i.e., the central position is in the best position, the search ends. Alternatively, the search ends when the number of searches meets a predefined threshold (e.g., 8 or 16).

[0108]

[0118] For a six-parameter affine model, the search method proposed for the four-parameter affine model can be extended. Consider the case of searching for a six-parameter affine motion (MV0, MV1, MV2). The search can be performed using at least three steps: initial motion search, refinement of translational motion parameters, and refinement of rotational and zoom motion parameters. The first and second steps are the same as those for a four-parameter affine search. The third step is to refine both MV1 and MV2. To reduce the complexity of the search, these two can be refined using iterative refinement. For example, MV0 and MV2 can be fixed, and MV1 can be refined using the same scheme as for MV1 refinement for a four-parameter affine model. After MV1 is refined, MV0 and MV1 can be fixed, and MV2 can be refined using the same scheme. Then, MV1 is refined again. In this way, these two MVs related to rotation and zoom motion can be refined iteratively until one MV remains unchanged or the number of iterations meets a predefined threshold. To achieve rapid convergence, in this iterative refinement scheme, the starting MV for refinement can be selected as follows: The selection of MV1 or MV2 to refine first can depend on their respective accuracies. Usually, the MV with lower accuracy is refined first. If both have the same accuracy, the MV with a larger distance from the control point of the MV to the top-left control point can be selected.

[0109]

[0119] To further reduce coding complexity, CU size and temporal layer may be taken into consideration when the encoder tests various precisions at control points for affine model-based coding. Precision determination may be performed only for large CUs. For example, an exemplary precision determination method may be applied only to CUs with areas larger than a predefined threshold (e.g., 16x16). For CUs with areas smaller than the threshold, (1 / 4 pixel, 1 / 4 pixel) precision is used for the two control points. For different temporal layer images with different QP settings, the encoder may test only those possible precisions at each temporal layer. For example, only (1 pixel, 1 / 4 pixel) (1 / 4 pixel, 1 / 4 pixel) may be tested for images in higher temporal layers (e.g., images in the top temporal layer). Also, only (1 / 4 pixel, 1 / 4 pixel) (1 / 8 pixel, 1 / 8 pixel) may be tested for images in lower temporal layers (e.g., images in the bottom temporal layer). For these intermediate layer images, the entire precision set may be tested.

[0110] Subblock-based affine motion compensation and estimation.

[0120] Affine motion estimation is an iterative estimation process. At each iteration, the temporal difference, spatial gradient, and local affine parameters (a, b, c, and d in Equation (3)) between the original signal and the motion-compensated predicted signal using the current motion vector are related by the optical flow equation, as shown in Equation (7). However, to reduce memory access bandwidth at the decoder, affine motion compensation prediction is based on subblocks (e.g., 4x4) rather than samples. This is because, when a motion vector points to a fractional position, an interpolation filter typically exists to derive sample values ​​for motion compensation. This interpolation process can significantly improve prediction compared to directly using sample values ​​at the nearest integer positions. However, the interpolation references multiple neighboring samples at integer positions. Given the MVs at the control points, the MVs for each subblock can be derived based on the center position of the subblock using Equation (1). When the subblock size is 1x1, motion compensation is sample-based, meaning that the motion of each sample can be different. Assume there is a separable interpolation filter with tap length N and the subblock size is SxS. For one sample, it operates to fetch (S+N-1)x(S+N-1) integer samples surrounding the reference position pointed to by MV for interpolation in both the horizontal and vertical directions. On average, it operates to fetch ((S+N-1)x(S+N-1) / (SxS)) reference samples at integer positions per sample. For sample-based affine motion compensation where S is equal to 1, this is NxN. For example, in HEVC and VTM, N is 8, and for a sub-block size of 4x4, the memory accesses per sample are (121 / 16). For sample-based interpolation, the amount of memory accesses per sample is 64, which is 8.5 times larger than that of 4x4 sub-block-based motion compensation. Therefore, sub-block-based motion compensation is used for affine motion prediction. In the affine motion estimation method described in the "Affine Mode" section, sample-based prediction is used, and this sub-block-based motion compensation is not taken into account.From equation (3), we can see that the differential motion of each position is related to the position within the CU given the affine parameters. Therefore, if the motion of all samples within a subblock is derived using equation (3) using the center position of the subblock, samples belonging to a subblock will have the same differential motion. For example, if a sample's location is (i, j) within a CU, the center position of the subblock to which it belongs is evaluated as equation (24).

number

[0111]

[0121] Then, equation (3) is (i,j) with (i b ,j b ) to obtain equation (25).

number

[0112]

[0122] Using equation (25), in equation (6)

number

number

[0113]

[0123] In some embodiments, Equation (26) is used to estimate the optimal affine parameters (a, b, c, d) using the least squares method. In such embodiments for motion estimation, the differential motion of samples belonging to one sub-block is the same. Therefore, the final MV at the control point is more accurate for sub-block-based motion compensation prediction compared to the sample-based estimation method using Equation (7).

[0114]

[0124] In affine motion compensation, the positions used for sub-block MV derivation within a CU may not be the actual center positions of the sub-blocks. As shown in FIG. 17, the affine CU is 8x4, and the sub-block size for motion compensation is 4x4. The positions used for sub-block MV derivation can be calculated by Equation (24) given the sample position (i, j). These positions are P0 and P1 of the left 4x4 sub-block and the right 4x4 sub-block, respectively. Based on the coordinates of P0 and P1, the MVs are derived by Equation (1) for a four-parameter affine model, or by Equations (8) and (9) for a six-parameter affine model. However, when Equation (24) is used, P0 and P1 are not the centers of these two sub-blocks. MV0 and MV1 may not be accurate for motion-compensated prediction of the sub-blocks. In one embodiment, we propose to use Equation (27) to calculate the positions for sub-block MV derivation.

number

[0115]

[0125] According to equation (27), P0 is replaced by P0', and P0' is the center of the left 4x4 sub-block. Therefore, the corresponding MV0' is more accurate compared to MV0. In the affine motion estimation method described herein to improve the accuracy of affine motion estimation, equation (27) can be replaced by equation (24). Given the MVs at the control points of an affine-coded CU, the MVs of the sub-blocks of the chroma components can reuse the MVs of the luma components or can be derived separately using equation (27).

[0116] Affine MVD rounding

[0126] In some implementations of affine motion compensation, sub-block MVs derived by control point MVs have 1 / 16 pixel precision, but the control point MVs are rounded to 1 / 4 pixel precision. The control point MVs are derived by adding MVDs to MV predictors. MVDs are signaled with 1 / 4 pixel precision. The MV predictors are rounded to 1 / 4 pixel precision before being used to derive the control point MVs. With adaptive affine MVD precision, the MV predictors used to derive the control point MVs of the current coding block may have higher precision than the MV precision of the current CU. In this case, the MV predictors are rounded to lower precision. The rounding results in information loss. In some embodiments proposed herein, the control point MVs and MV predictors are kept at the highest precision, for example, 1 / 16 pixel, while the MVDs are rounded to the desired precision.

[0117]

[0127] In affine motion estimation, affine parameters may be estimated iteratively. For each iteration, differential control point MVs may be derived using the optical flow method as described in equations (4) and (5). In a VTM implementation, the control point MVs at step k are updated by:

number

number

number

number

number

[0118]

[0128] In an exemplary embodiment of the method provided herein, the control points MV in step k are updated by the following steps: The top-left control point MV is updated according to equations (29) to (31).

number

number

[0119]

[0129] In equations (29) to (34),

number

[0120]

[0130] MVP i is 1 / 16 pixel precision, so

number

[0121] Adaptive Affine MVD Coding

[0131] Affine MVDs with different precisions may have different characteristics. Control point MVDs may have different physical meanings. For example, the absolute value of the MVD may be smaller on average for (1 / 8 pixel, 1 / 8 pixel, 1 / 8 pixel) or (1 / 16 pixel, 1 / 16 pixel, 1 / 16 pixel) precision compared to (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel) precision. As explained in the "MVD Encoding" section above, the lengths of EG codes with different orders are different. In general, when the EG order is small, the length of EG codes for small values ​​is short and the length of EG codes for large values ​​is long. Some embodiments use an adaptive EG order for MVD encoding that considers the MVD precision and its physical motion meaning (e.g., rotation in different directions, zoom). Some embodiments use an adaptive EG order for MVD encoding that considers the MVD precision and its physical motion meaning (e.g., rotation in different directions, zoom). In some embodiments, the top-left MVD (MVD 0x ,MVD 0y ) is the MVD component MVD 0x and MVD 0y Since is a translational motion, it has the same EG order as the non-affine MVD encoding. For a six-parameter affine model, the MVD components MVD 1y and MVD 2x is related to the rotational motion, and the MVD component MVD 1x and MVD 2y is related to the zoom behavior. For a four-parameter affine model, the MVD components MVD 1y is related to the rotational motion, and the MVD component MVD 1x is related to the zoom action.

[0122]

[0132] In some embodiments, the MVD values ​​have different properties, so that the EG code has different orders for different MVD encodings. In some embodiments, the MVD associated with translational motion (MVD 0x ,MVD 0y ), the EG order is not signaled; instead, such MVD may use the same EG order (e.g., 1) as non-affine MVD encoding.

[0123]

[0133] In some embodiments, EG orders are signaled for the exponential-Golomb codes used for different MVD components corresponding to non-translational motion, such as the MVD components listed in Table 9 for three MVD precisions. In the embodiment of Table 9, six EG orders (EGOrder[0] to EGOrder[5]) are signaled in the bitstream. The EG orders range from 0 to 3 and use two bits for encoding. The MVD precision index indicates different MVD precisions. For example, an MVD precision index of "0" indicates (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel) precision, an MVD precision index of "1" indicates (1 / 16 pixel, 1 / 16 pixel, 1 / 16 pixel) precision, and an MVD precision index of "2" indicates (1 pixel, 1 pixel, 1 pixel) precision. These signaled EG orders are intended to indicate the EG orders used for EG binarization of different MVD components with different MVD precisions. For example, the EG order [0] is a state where the MVD accuracy index is "0" (i.e., (1 / 4 pixel, 1 / 4 pixel, 1 / 4 pixel) accuracy set), and the MVD component MVD 1y and MVD 2x For a four-parameter affine model, MVD 2x and MVD 2y There is no need to code MVD in Table 9. 1x and MVD 1y Only the .

[0124] [Table 9]

[0125]

[0134] Signaling of the EG order may be performed, for example, in a picture parameter set or a slice header. In an embodiment where the EG order is signaled in the slice header, the encoder may select the EG order based on previously coded pictures in the same temporal layer. After each inter picture is coded, the encoder may compare the total number of bins using different EG codes of different orders for all MVDs in that category. For example, all MVDs with MVD precision of "0" may be coded as follows: 1y and MVD 2xFor EG Order 0, EG Order 1, EG Order 2, and EG Order 3, the encoder compares the total number of bins and selects the order with the smallest value of the total number of bins. The selected order is then used for encoding the next image in the same temporal layer, and the selected order is also coded in the slice header of the next image in the same temporal layer.

[0126] Further embodiments

[0135] In some embodiments, a method of decoding video from a bitstream is provided. The method includes reading, for at least one block in the video, information from the bitstream that identifies one of a plurality of elements in a selected predetermined precision set, where the identified element of the selected predetermined precision set indicates at least a selected first precision and a selected second precision, and decoding the block using an affine motion model, where the affine motion model is characterized by at least a first motion vector having the selected first precision and a second motion vector having the selected second precision. The method may include reading, from the bitstream, information that indicates the first motion vector and the second motion vector. The information that indicates the first motion vector and the second motion vector may include a first differential motion vector and a second differential motion vector.

[0127]

[0136] In some embodiments, the information identifying one of the elements is read from the bitstream in blocks.

[0128]

[0137] In some embodiments, a first motion vector is associated with a first control point of the block, and a second motion vector is associated with a second control point of the block.

[0129]

[0138] In some embodiments, each element of the selected predetermined precision set includes a first available precision and a second available precision, where the second available precision may be greater than or equal to the first available precision.

[0130]

[0139] In some embodiments, information specifying a selected predetermined precision set from among multiple available predetermined precision sets is read from the bitstream. In some such embodiments, the information specifying the selected predetermined precision set is signaled in a picture parameter set, a sequence parameter set, or a slice header. Examples of predetermined position sets include: {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel)}, {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 8 pixel)}, {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 8 pixel)}, and {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 8 pixel), (1 / 8 pixel, 1 / 8 pixel)}.

[0131]

[0140] In some embodiments, the affine motion model is further characterized by a third motion vector having a selected third precision, and the identified element of the selected predetermined precision set further indicates the selected third precision.

[0132]

[0141] In some embodiments, information identifying one of the plurality of elements is coded in the bitstream using context-adaptive binary arithmetic coding.

[0133]

[0142] In some embodiments, a determination is made whether the size of the block is greater than a threshold size, and information identifying one of the plurality of elements is read from the bitstream for that block only if the size of the block is greater than the threshold size.

[0134]

[0143] In some embodiments, the selected predetermined precision set is selected based on the temporal layer of the image that contains the block.

[0135]

[0144] In some embodiments, the selected predetermined precision set is selected based on the shape of the block.

[0136]

[0145] In some embodiments, a method for decoding video in a bitstream is provided. The method includes reading, for at least one block in the video, from the bitstream: (i) first information indicating a first precision from a first predetermined set of available precisions and (ii) second information indicating a second precision from a second predetermined set of available precisions; decoding the block using an affine motion model characterized by at least a first motion vector having a selected first precision and a second motion vector having a selected second precision; and signaling in the bitstream: (i) the first information indicating the first precision from the first predetermined set of available precisions and (ii) the second information indicating the second precision from the second predetermined set of available precisions. The first predetermined set and the second predetermined set may be different.

[0137]

[0146] In some embodiments, the first predetermined set is {1 pixel, 1 / 4 pixel, 1 / 8 pixel} and the second predetermined set is {1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel}.

[0138]

[0147] In some embodiments, a first motion vector is associated with a first control point of the block, and a second motion vector is associated with a second control point of the block.

[0139]

[0148] In some embodiments, a method for encoding video in a bitstream is provided. The method includes, for at least one block in the video, encoding the block using an affine motion model, the affine motion model being characterized by at least a first motion vector having a selected first precision and a second motion vector having a selected second precision; and signaling in the bitstream information identifying one of a plurality of elements in a selected predetermined precision set, the identified element of the selected predetermined precision set indicating at least the selected first precision and the selected second precision. The method may further include signaling in the bitstream information indicating the first motion vector and the second motion vector. The information indicating the first motion vector and the second motion vector may include a first differential motion vector and a second differential motion vector.

[0140]

[0149] In some embodiments, information identifying one of the elements is transmitted in blocks.

[0141]

[0150] In some embodiments, a first motion vector is associated with a first control point of the block and a second motion vector is associated with a second control point of the block.

[0142]

[0151] In some embodiments, each element of the selected predetermined precision set includes a first available precision and a second available precision, and in some embodiments, the second available precision is greater than or equal to the first available precision.

[0143]

[0152] In some embodiments, the method includes signaling, in the bitstream, information identifying a selected predetermined precision set from a plurality of available predetermined precision sets. The information identifying the selected predetermined precision set may be signaled, for example, in a picture parameter set, a sequence parameter set, or a slice header.

[0144]

[0153] Examples of predetermined sets of locations include: {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel)}, {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 8 pixel)}, {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 8 pixel, 1 / 8 pixel)}, and {(1 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 4 pixel), (1 / 4 pixel, 1 / 8 pixel), (1 / 8 pixel, 1 / 8 pixel)}.

[0145]

[0154] In some embodiments, the affine motion model is further characterized by a third motion vector having a selected third precision, and the identified element of the selected predetermined precision set further indicates the selected third precision.

[0146]

[0155] In some embodiments, information identifying one of the plurality of elements is coded in the bitstream using context-adaptive binary arithmetic coding.

[0147]

[0156] In some embodiments, the method includes determining whether a size of the block is greater than a threshold size, and wherein information identifying one of the plurality of elements is signaled in the bitstream for the block only if the size of the block is greater than the threshold size.

[0148]

[0157] In some embodiments, the selected predetermined precision set is selected based on the temporal layer of the image that contains the block.

[0149]

[0158] In some embodiments, the selected predetermined precision set is selected based on the shape of the block.

[0150]

[0159] In some embodiments, a method for encoding video in a bitstream is provided, the method including: for at least one block in the video, encoding the block using an affine motion model, the affine motion model characterized by at least a first motion vector having a selected first precision and a second motion vector having a selected second precision; and signaling in the bitstream (i) first information indicating a first precision from a first predetermined set of available precisions and (ii) second information indicating a second precision from a second predetermined set of available precisions, where the first predetermined set and the second predetermined set may be different.

[0151]

[0160] In some embodiments, the first predetermined set is {1 pixel, 1 / 4 pixel, 1 / 8 pixel} and the second predetermined set is {1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel}.

[0152]

[0161] In some embodiments, a first motion vector is associated with a first control point of the block and a second motion vector is associated with a second control point of the block.

[0153]

[0162] Some embodiments include a method for encoding video in a bitstream, the method including: for at least one block in the video, determining a first rate-distortion cost of encoding the block using a translational motion model; determining a second rate-distortion cost of encoding the block using an affine prediction model with a first set of affine model precisions; determining whether the second rate-distortion cost is less than the first rate-distortion cost; in response to a determination that the second rate-distortion cost is less than the first rate-distortion cost, determining at least a third rate-distortion cost of encoding the block using an affine prediction model with a second set of affine model precisions; and encoding the block in the bitstream using a coding model associated with the determined lowest rate-distortion cost.

[0154]

[0163] In some embodiments, in response to determining that the second rate-distortion cost is less than the first rate-distortion cost, a fourth rate-distortion cost of encoding the block using an affine prediction model with a fourth set of affine model precision is determined.

[0155]

[0164] In some embodiments, a method for encoding video in a bitstream is provided, the method comprising: for at least one block in the video, calculating I' k (i,j)-I(i,j)=(g x (i,j)*i b +g y (i,j)*j b )*c+(-g x (i,j)*j b +g y (i,j)*i b )*d+g x (i,j)*a+g y (i,j)*b determining affine parameters a, b, c, and d using I(i,j) is the original luminance signal, and I' k (i,j) is the predicted luminance signal, and g x (i,j) and g y (i,j) is I' k is the spatial gradient applied to (i,j),

number

[0156]

[0165] In some embodiments, a method for encoding video is provided. The method includes: for at least one block in the video, determining a motion vector predictor (MVP) for at least one control point, where the motion vector predictor has a first precision; determining a differential motion vector (MVD) value for the control point, where the differential motion vector value has a second precision lower than the first precision; calculating a motion vector for the control point by adding at least the differential motion vector value to the motion vector predictor, where the calculated motion vector has the first precision; and predicting the block by affine prediction using the calculated motion vector for the at least one control point. The differential motion vector value can be signaled in a bitstream by an encoder or parsed from the bitstream by a decoder.

[0157]

[0166] In some embodiments, the method is performed by an encoder, and identifying the differential motion vector includes iteratively determining a motion vector differential for the control point based on the initial motion vector, updating the differential motion vector based on the motion vector differential, rounding the differential motion vector to a second precision, and adding the rounded differential motion vector to a motion vector predictor to generate an updated motion vector, the motion vector predictor, and an updated motion vector having the first precision.

[0158]

[0167] In some embodiments, the first precision is 1 / 16 pixel precision and the second precision is 1 / 4 pixel precision.

[0159]

[0168] In some embodiments, predicting a block by affine prediction is performed using two control points, a respective differential motion vector is determined for each control point, and each respective differential motion vector has a second precision.

[0160]

[0169] In some embodiments, predicting the block using affine prediction is performed using three control points, and a respective differential motion vector is determined for each control point, each differential motion vector having a second precision.

[0161]

[0170] In some embodiments, a method of decoding video from a bitstream is provided, the method including: for at least one block in the video, determining a respective coding order for each of a plurality of motion vector differential (MVD) components based at least in part on information encoded in the bitstream, reading each of the MVD components from the bitstream using the respective determined coding orders, and decoding the block using an affine motion model, where the affine motion model is characterized at least in part by the MVD components.

[0162]

[0171] In some embodiments, the method includes reading, from the bitstream, information specifying a precision of each of the MVD components, wherein a coding order of the MVD components is determined based in part on the precision of each. The MVD components may be coded using Exponential-Golomb coding, and the coding order may be an Exponential-Golomb coding order.

[0163]

[0172] Some embodiments include a method of decoding video from a bitstream, the method including: determining, for at least one block in the video, a respective coding order for each of a plurality of motion vector difference (MVD) components, the coding order of each of the MVD components being determined based on (i) a precision of the MVD component and (ii) whether the component is associated with a rotational motion or a zoom motion; reading each of the MVD components from the bitstream using the respective determined coding order; and decoding the block using an affine motion model, the affine motion model being at least partially characterized by the MVD component.

[0164]

[0173] Some embodiments involve reading order information from a bitstream, the order information comprising: (i) quarter-pixel accuracy and (ii) a first encoding order associated with rotational motion; (i) quarter pixel accuracy and (ii) a second coding order associated with zoom movements; (i) 1 / 16 pixel accuracy and (ii) a third coding order associated with rotational motion; (i) 1 / 16 pixel precision and (ii) a fourth coding order associated with zoom movements; (i) one pixel accuracy and (ii) a fifth coding order associated with rotational motion; (i) one pixel accuracy and (ii) a sixth coding order associated with zoom motion; and identifying the The respective encoding orders are performed using order information, which may be coded in the picture parameter set or slice header, for example.

[0165]

[0174] In some embodiments, the MVD components are encoded using Exponential-Golomb coding, and the coding order is the Exponential-Golomb coding order.

[0166]

[0175] In some embodiments, a method for encoding video in a bitstream is provided, the method including: selecting, for at least one block in the video, order information, the order information identifying a coding order of a differential motion vector (MVD) component based on (i) a precision of the MVD component and (ii) whether the component is associated with a rotational motion or a zoom motion; encoding the order information into the bitstream; and encoding the block using an affine motion model, the affine motion model being characterized at least in part by a plurality of MVD components, each of the plurality of MVD components being encoded into the bitstream using a coding order determined by the order information.

[0167]

[0176] In some embodiments, the order information is (i) quarter-pixel accuracy and (ii) a first encoding order associated with rotational motion; (i) quarter pixel accuracy and (ii) a second coding order associated with zoom movements; (i) 1 / 16 pixel accuracy and (ii) a third coding order associated with rotational motion; (i) 1 / 16 pixel precision and (ii) a fourth coding order associated with zoom movements; (i) one pixel accuracy and (ii) a fifth coding order associated with rotational motion; (i) one pixel accuracy and (ii) a sixth coding order associated with zoom motion; Identify. Determining the respective coding orders may be performed using order information, which may be coded in, for example, a picture parameter set or a slice header.

[0168]

[0177] In some embodiments, the MVD components are encoded using Exponential-Golomb coding, and the coding order is the Exponential-Golomb coding order.

[0169]

[0178] Some embodiments include a non-transitory computer-readable storage medium storing video encoded using any of the methods disclosed herein. Some embodiments include a non-transitory computer-readable storage medium storing instructions that operate to perform any of the methods disclosed herein.

[0170] Encoded Bitstream Structure

[0179] Figure 15 shows an example of a coded bitstream structure. The coded bitstream 1300 is composed of several NAL (Network Abstraction Layer) units 1301. The NAL units may contain coded sample data, such as coded slices 1306, or high-level syntax metadata, such as parameter set data, slice header data 1305, or supplemental side information data 1307 (which may also be referred to as SEI messages). A parameter set is a high-level syntax structure that contains essential syntax elements that can apply to multiple bitstream layers (e.g., video parameter sets 1302 (VPS)), or to coded video sequences within a single layer (e.g., sequence parameter sets 1303 (SPS)), or to several coded pictures within a single coded video sequence (e.g., picture parameter sets 1304 (PPS)). Parameter sets may be transmitted together with the coded pictures of the video bitstream or through other means (including out-of-band transmission using a reliable channel, hard coding, etc.). The slice header 1305 is also a high-level syntax structure that may be relatively small or may contain some image-related information that is relevant only to a particular slice or image type. The SEI message 1307 carries information that may not be needed by the decoding process, but may be used for various other purposes, such as image output timing or display and / or loss detection and concealment.

[0171] Communication devices and systems

[0180] FIG. 16 illustrates an example of a communication system. The communication system 1400 may include an encoder 1402, a communication network 1404, and a decoder 1406. The encoder 1402 may communicate with the communication network 1404 via a connection 1408, which may be a wired or wireless connection. The encoder 1402 may be similar to the block-based video encoder of FIG. 2A. The encoder 1402 may include a single-layer codec (e.g., FIG. 2A) or a multi-layer codec. The decoder 1406 may communicate with the communication network 1404 via a connection 1410, which may be a wired or wireless connection. The decoder 1406 may be similar to the block-based video decoder of FIG. 2B. The decoder 1406 may include a single-layer codec (e.g., FIG. 2B) or a multi-layer codec.

[0172]

[0181] The encoder 1402 and / or decoder 1406 may be incorporated into a wide variety of wired communication devices and / or wireless transmit / receive units (WTRUs), such as, but not limited to, digital televisions, wireless broadcast systems, network elements / terminals, servers such as content servers or web servers (e.g., Hypertext Transfer Protocol (HTTP) servers, etc.), personal digital assistants (PDAs), laptop or desktop computers, tablet computers, digital cameras, digital recording devices, video game devices, video game consoles, cellular or satellite wireless telephones, digital media players, and / or the like.

[0173]

[0182] The communications network 1404 may be any suitable type of communications network. For example, the communications network 1404 may be a multiple access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communications network 1404 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communications network 1404 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), and / or the like. The communications network 1404 may include multiple connected communications networks. The communications network 1404 may include the Internet and / or one or more commercial private networks, such as cellular networks, Wi-Fi hotspots, Internet Service Provider (ISP) networks, and / or the like.

[0174]

[0183] It should be noted that various hardware elements of one or more of the described embodiments are referred to as "modules," which perform (i.e., perform, execute, etc.) various functions described herein with respect to the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) deemed appropriate by one of ordinary skill in the art for a given implementation. It should be noted that each described module may also include executable instructions for performing one or more functions described as being performed by the respective module, which may take or include the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, and may be stored on any suitable non-transitory computer-readable medium or media, commonly referred to as RAM, ROM, etc.

[0175]

[0184] While features and elements have been described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. Additionally, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in conjunction with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. 1. A method for decoding video from a bitstream, the method comprising: for at least one current block in the video: reading information from the bitstream that identifies at least a first motion vector predictor and a second motion vector predictor; reading at least a first differential motion vector and a second differential motion vector from the bitstream, the first and second differential motion vectors having precision; reading, from the bitstream, information specifying one of a plurality of precisions within a predetermined precision set, the specified precision indicating the precision of the first and second differential motion vectors; generating at least (i) a first control point motion vector from the first motion vector predictor and the first differential motion vector, and (ii) a second control point motion vector from the second motion vector predictor and the second differential motion vector; generating a prediction of the current block using an affine motion model, the affine motion model being characterized by at least the first control point motion vector and the second control point motion vector; A method comprising:

2. The method of claim 1 , wherein the plurality of accuracies in the predetermined set of accuracies includes ¼ pixel accuracy, 1 / 16 pixel accuracy, and 1 pixel accuracy.

3. The method of claim 1 or 2, wherein the affine motion model is a four-parameter motion model.

4. the affine motion model is a six-parameter motion model, and the method comprises: reading information from the bitstream that identifies a third motion vector predictor; reading a third differential motion vector having the specified precision from the bitstream; generating a third control point motion vector from the third motion vector predictor and the third difference motion vector, the affine motion model is characterized by the first control point motion vector, the second control point motion vector, and the third control point motion vector; The method of claim 1 or 2, further comprising:

5. The method of any one of claims 1 to 4, wherein the information specifying one of the plurality of precisions is read from the bitstream in blocks.

6. The method of any one of claims 1 to 5, further comprising rounding at least one of the motion vector predictors to the specified precision.

7. A method according to any one of claims 1 to 6, wherein each of the control point motion vectors is generated by adding a corresponding differential motion vector to a respective motion vector predictor.

8. The method of any one of claims 1 to 7, wherein the predetermined precision set is different from the predetermined precision set used for non-affine inter coding of the video.

9. generating a prediction of the current block, determining a respective sub-block motion vector for each of a plurality of sub-blocks of the current block using the affine motion model; generating an inter prediction for each of the sub-blocks using the respective sub-block motion vectors; The method according to any one of claims 1 to 8, comprising:

10. reading a residual for the current block from the bitstream; reconstructing the current block by adding the residual to the prediction of the current block; The method of any one of claims 1 to 9, further comprising:

11. 1. A system for decoding video from a bitstream, the system comprising: a processor; and for at least one current block in the video: reading information from the bitstream that identifies at least a first motion vector predictor and a second motion vector predictor; reading at least a first differential motion vector and a second differential motion vector from the bitstream, the first and second differential motion vectors having precision; reading, from the bitstream, information specifying one of a plurality of precisions within a predetermined precision set, the specified precision indicating the precision of the first and second differential motion vectors; generating at least (i) a first control point motion vector from the first motion vector predictor and the first differential motion vector, and (ii) a second control point motion vector from the second motion vector predictor and the second differential motion vector; generating a prediction of the current block using an affine motion model, the affine motion model being characterized by at least the first control point motion vector and the second control point motion vector; and a non-transitory computer-readable medium storing instructions that operate to perform functions including:

Citation Information

Patent Citations

  • Motion vector prediction for affine motion models in video coding

    US20180098063A1

  • Adaptive motion vector precision for video coding

    US20180098089A1

  • Video signal processing method and apparatus using adaptive motion vector resolution

    WO2019235896A1

  • Syntax reuse for affine mode with adaptive motion vector resolution

    WO2020058890A1