Method and apparatus for reducing coding wait time for decoder-side motion refinement

JP2026137785APending Publication Date: 2026-08-27INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026101576
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-06-27
Filing Date
2026-06-18
Publication Date
2026-08-27

Smart Images

  • Figure 2026137785000001_ABST
    Figure 2026137785000001_ABST
Patent Text Reader

Abstract

A video coding system and method are proposed to reduce coding wait times caused by DMVR. [Solution] Using dual prediction, two unrefined motion vectors are identified for coding a first block of a sample (e.g., a first coding unit). One or both of the unrefined motion vectors are used to predict motion information for a second block of a sample (e.g., a second coding unit). The two unrefined motion vectors are refined using a DMVR, and the refined motion vectors are used to generate a prediction signal for the first block of the sample. Such embodiments allow the second block of a sample to be coded substantially in parallel with the first block without waiting for the completion of the DMVR on the first block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application is a patent application derived from U.S. Provisional Patent Application No. 62 / 690,507, filed on June 27, 2018, entitled “Methods and Apparatus for Reducing the Coding Latency of Decoder-Side Motion Refinement,” which is incorporated herein by reference in its entirety, and claims the benefit thereof. [Background technology]

[0002] Video coding systems are widely used to compress digital video signals, thereby reducing the need for storage and / or transmission bandwidth of such signals. Various types of video coding systems are most widely used and developed, including block-based systems, wavelet-based systems, object-based systems, and more recently, block-based hybrid video coding systems. Examples of block-based video coding systems include international video coding standards such as MPEG1 / 2 / 4 part2, H.264 / MPEG-4 part10 AVC, VC-1, and the latest video coding standard called High Efficiency Video Coding (HEVC), developed by the ITU-T / SG16 / Q.6 / VCEG and the JCT-VC (Joint Collaborative Team on Video Coding) of ISO / IEC / MPEG.

[0003] The first version of the HEVC standard was completed in October 2013, resulting in approximately 50% bitrate savings or equivalent perceived quality compared to the previous generation video coding standard, H.264 / MPEG AVC. While the HEVC standard offers a significant coding improvement over its predecessors, there is evidence that superior coding efficiency can be achieved with additional coding tools. Based on this, both VCEG and MPEG have begun work on pursuing new coding techniques for future video coding standardization. The Joint Video Exploration Team (JVET) was formed in October 2015 by ITU-TVECG and ISO / IECMPEG to initiate important research into advanced techniques that will enable significant enhancements to coding efficiency. Reference software, referred to as the Joint Exploration Model (JVET), is maintained by JVET by integrating several additional coding tools on top of the HEVC Test Model (HM).

[0004] In October 2017, a Joint Request for Proposals (CfP) for a video compression method superior to HEVC was issued by the ITU-T and ISO / IEC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating approximately 40% greater compression efficiency than HEVC. Based on these evaluation results, JVET launched a new project to develop a new generation video coding standard called Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. For the initial VTM-1.0, most of the coding modules, including intra-prediction, inter-prediction, transform / inverse transform, and quantization / dequantization, as well as in-loop filters, conform to the existing HEVC design, except that a multi-type tree-based block partitioning structure is used in VTM. Meanwhile, another reference software base called the Benchmark Set (BMS) has also been generated to facilitate the evaluation of the new coding tools. In the BMS codebase, a list of coding tools inherited from JEM, resulting in higher coding efficiency and moderate implementation complexity, is placed above VTM and used as a benchmark when evaluating similar coding techniques during the VVC standardization process. The JEM coding tools integrated into BMS-1.0 include 65 angular intra-prediction directions, modified coefficient coding, evolutionary multiple transformation (AMT) + 4x4 unseparable quadratic transformation (NSST), affine motion models, generalized adaptive loop filters (GALF), advanced temporal motion vector prediction (ATMVP), adaptive motion vector accuracy, decoder-side motion vector refinement (DMVR), and LM chroma modes. [Overview of the project]

[0005] Some embodiments include methods used in video coding and decoding (collectively, “coding”). Some embodiments of a block-based video coding method include the steps of: refining a first non-refined motion vector and a second non-refined motion vector in a first block to generate a first refined motion vector and a second refined motion vector; predicting motion information for a second block using one or both of the first non-refined motion vector and the second non-refined motion vector, wherein the second block is a spatial adjacency of the first block; and predicting the first block by biprediction using the first refined motion vector and the second refined motion vector.

[0006] In an embodiment of the video coding method, a first unrefined motion vector and a second unrefined motion vector associated with a first block are identified. Motion information for the second block adjacent to the first block is predicted using one or both of the first and second unrefined motion vectors. The first and second unrefined motion vectors are refined, for example, using decoder-side motion vector refinement (DMVR). The refined motion vectors are used to generate a first refined motion vector and a second refined motion vector that can be used for the dual prediction of the first block. Using the unrefined motion vector(s) to predict the motion information of the second block is performed using one or more techniques such as spatially evolved motion vector prediction (AMVP), temporal motion vector prediction (TMVP), or evolutionary temporal motion vector prediction (TMVP), and using the unrefined motion vector(s) as spatial merge candidates. In the case of spatial prediction, the second block may be spatially adjacent to the first block, and in the case of temporal prediction, the second block may be a juxtaposed block of a subsequently coded picture. In some embodiments, the deblocking filter strength for the first block is determined based on at least a portion of the first unrefined motion vector and the second unrefined motion vector.

[0007] In another embodiment of the video coding method, a first unrefined motion vector and a second unrefined motion vector associated with a first block are identified. The first unrefined motion vector and the second unrefined motion vector are refined, for example, using DMVR, to generate a first refined motion vector and a second refined motion vector. The motion information of the second block is predicted using either spatial motion prediction or temporal motion prediction. (i) When spatial motion prediction is used, one or both of the first unrefined motion vector and the second unrefined motion vector are used to predict the motion information. (ii) When temporal motion prediction is used, one or both of the first refined motion vector and the second refined motion vector are used to predict the motion information.

[0008] In another embodiment of the video coding method, at least one predictor for predicting the motion information of a current block is selected. The selection is made from a set of available predictors, and the available predictors include at least one unrefined motion vector from the spatial neighboring blocks of the current block and (ii) at least one refined motion vector from the collocated blocks of the current block.

[0009] In another embodiment of the video coding method, at least two non-overlapping regions within a slice are determined. A first unrefined motion vector and a second unrefined motion vector associated with a first block within the first region are identified. The first unrefined motion vector and the second unrefined motion vector are refined to generate a first refined motion vector and a second refined motion vector. In response to a determination that the motion information of a second block adjacent to the first block is predicted using the motion information of the first block, the motion information of the second block is predicted using (i) one or both of the first unrefined motion vector and the second unrefined motion vector when the first block is not on the lower boundary or the right boundary of the first region, and (ii) one or both of the first refined motion vector and the second refined motion vector when the first block is on the lower boundary or the right boundary of the first region.

[0010] In another embodiment of the video coding method, at least two non-overlapping regions within a slice are determined. A first unrefined motion vector and a second unrefined motion vector associated with a first block within the first region are identified. The first unrefined motion vector and the second unrefined motion vector are refined to generate a first refined motion vector and a second refined motion vector. In response to a determination that the motion information of a second block adjacent to the first block is predicted using the motion information of the first block, the motion information of the second block is predicted using (i) one or both of the first unrefined motion vector and the second unrefined motion vector when the second block is within the first region, and (ii) one or both of the first refined motion vector and the second refined motion vector when the second block is not within the first region.

[0011] In another embodiment of the video coding method, at least two non-overlapping regions within a slice are determined. A first unrefined motion vector and a second unrefined motion vector associated with a first block in the first region are identified. The first and second unrefined motion vectors are refined to produce a first refined motion vector and a second refined motion vector. Motion information for the second block is predicted using either spatial motion prediction or temporal motion prediction, where (i) if the first block is not on the lower or right boundary of the first region and spatial motion prediction is used, one or both of the first and second unrefined motion vectors are used to predict motion information; and (ii) if the first block is on the lower or right boundary of the first region and temporal motion prediction is used, one or both of the first refined motion vector and the second refined motion vector are used to predict motion information.

[0012] In another embodiment of the video coding method, at least two non-overlapping regions are defined within the slice. The available sets of predictors for predicting the motion information of the current block in the first region are determined, and the available sets of predictors are constrained so as not to include motion information of any block in the second region, which is different from the first region.

[0013] Several embodiments describe methods for refining motion vectors. In one embodiment, a first unrefined motion vector and a second unrefined motion vector for the current block are determined. First prediction I (0) It is generated using the first unrefined motion vector and the second prediction I (1) This is generated using the second, unrefined motion vector. Motion refinement relative to the current block.

[0014]

number

[0015] An optical flow model is used to determine this. The first and second unrefined motion vectors are refined using motion refinement to produce the first and second refined motion vectors. The current block is predicted by biprediction using the first and second refined motion vectors.

[0016] In another embodiment of the video coding method, a first unrefined motion vector and a second unrefined motion vector for the current block are determined. First prediction I (0) It is generated using the first unrefined motion vector and the second prediction I (1) This is generated using the second, unrefined motion vector. Motion refinement relative to the current block.

[0017]

number

[0018] This was determined,

[0019]

number

[0020] And θ is the set of coordinates for all samples in the current block,

[0021]

number

[0022] It is. The first unrefined motion vector and the second unrefined motion vector are refined using motion refinement so as to generate the first refined motion vector and the second refined motion vector. The current block is predicted by bi-prediction using the first refined motion vector and the second refined motion vector.

[0023] In another embodiment of the video coding method, a first motion vector and a second motion vector for a current block are determined. The first motion vector and the second motion vector are (a) generating a first prediction P 0 using the first motion vector and generating a second prediction P 1 using the second motion vector; (b) generating a bi-prediction template signal P 0 by averaging the first prediction P 1 and the second prediction P tmp ; (c) using an optical flow model to determine a first motion refinement (Δx,Δy) tmp 0 for the first motion vector and a second motion refinement (Δx,Δy) * 1 for the second motion vector based on the template signal P * ; (d) refining the first motion vector using the first motion refinement (Δx,Δy) * 0 and refining the second motion vector using the second motion refinement (Δx,Δy) * 1; and is refined by repeatedly executing steps including

[0024] Further embodiments include encoder and decoder (collectively, “codec”) systems configured to perform the methods described herein. Such systems may include a processor and a non-temporary computer storage medium, the non-temporary computer storage medium storing instructions operable to perform the methods described herein when executed on the processor. Additional embodiments include a non-temporary computer-readable medium storing video encoded using the methods described herein. [Brief explanation of the drawing]

[0025] [Figure 1A] This is a system diagram illustrating an exemplary communication system that can implement one or more of the disclosed embodiments. [Figure 1B] This is a system diagram illustrating an embodiment of a wireless transmit / receive unit (WTRU) that can be used in the communication system illustrated in Figure 1A, according to the embodiment. [Figure 2] This is a functional block diagram of a block-based video encoder, such as an encoder used for VVC. [Figure 3A] This shows block partitions and quaternary partitions in a multi-type tree structure. [Figure 3B] This shows block partitions and vertical binary partitions in a multi-type tree structure. [Figure 3C] This shows block partitions and horizontal binary partitions in a multi-type tree structure. [Figure 3D] This shows block partitions and vertical ternary partitions in a multi-type tree structure. [Figure 3E] This shows block partitions and horizontal ternary partitions in a multi-type tree structure. [Figure 4] This is a functional block diagram of a block-based video decoder, such as a decoder used for VVC. [Figure 5] An example of spatial motion vector prediction is provided. [Figure 6] An example of a time motion vector prediction (TMVP) is provided. [Figure 7] An example of an evolutionary time-motion vector prediction (ATMVP) is provided. [Figure 8A] An example of a decoder-side motion vector refinement (DMVR) is provided. [Figure 8B] An example of a decoder-side motion vector refinement (DMVR) is provided. [Figure 9] This example demonstrates parallel decoding for VTM-1.0. [Figure 10] This example illustrates the decoding wait time caused by the DMVR. [Figure 11] An embodiment is illustrated in which the refined MV from the DMVR is used solely to generate dual predictive signals. [Figure 12] This example illustrates an embodiment in which the refined MV from the DMVR is used for time motion prediction and deblocking, and the unrefined MV is used for spatial motion prediction. [Figure 13] This example illustrates an embodiment in which the refined MV from the DMVR is used for time motion prediction, and the unrefined MV is used for spatial motion prediction and deblocking. [Figure 14] The following examples illustrate parallel decoding after applying a latency reduction method to a DMVR according to several embodiments. [Figure 15] An embodiment is illustrated in which an unrefined MV is used for DMVR blocks within a picture segment for spatial motion prediction and deblocking. [Figure 16]This example illustrates an embodiment in which the current picture is divided into multiple segments, and the coding wait time is reduced for each block within each segment. [Figure 17] This example illustrates an embodiment in which the current picture is divided into multiple segments, and coding latency is reduced for blocks from different segments. [Figure 18] This is a flowchart of motion refinement processing using optical flow according to several embodiments. [Modes for carrying out the invention]

[0026] Network of the embodiment for implementation of the embodiment Figure 1A shows an exemplary communication system 100 that can implement one or more disclosed embodiments. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 100 can enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word discrete Fourier transform spread OFDM (ZT UW DTS-S-OFDM), unique-word OFDM (UW-OFDM), resource-block filtered OFDM, and filter-bank multi-carrier (FBMC).

[0027] As shown in Figure 1A, the communication system 100 may include radio transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it will be recognized that the disclosed embodiments consider any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a radio environment. For example, WTRU102a, 102b, 102c, and 102d, any of which may be referred to as “station” and / or “STA”, may be configured to transmit and / or receive radio signals and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, radio sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other radio devices operating in industrial and / or automated processing chain situations), consumer electronics devices, and devices operating on commercial and / or industrial radio networks. Any of WTRU102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0028] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as CN 106, the Internet 110, and / or other networks 112. For example, base stations 114a, 114b may be base transceiver stations (BTS), NodeBs, eNodeBs, home NodeBs, home eNodeBs, gNBs, NR NodeBs, site controllers, access points (APs), and wireless routers. Although each of the base stations 114a, 114b is represented as a single element, it will be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0029] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown) such as base station controllers (BSCs), radio network controllers (RNCs), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for radio services to a particular geographic area, which may be relatively fixed or change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In the embodiment, the base station 114a may utilize multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0030] Base stations 114a and 114b may communicate with one or more WTRUs 102a, 102b, 102c, and 102d over the air interface 116, and the air interface 116 may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0031] More specifically, as described above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a in RAN 104 / 113 and WTRU 102a, 102b, 102c may establish an air interface 116 using broadband CDMA (WCDMA®), or implement radio technologies such as Universal Mobile Communications System (UMTS) Terrestrial Radio Access (UTRA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​Uplink (UL) Packet Access (HSUPA).

[0032] In embodiments, base stations 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as Advanced UMTS Terrestrial Radio Access (E-UTRA) to establish an air interface 116 using Long-Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0033] In the embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as NR radio access, which may use New Radio (NR) to establish an air interface 116.

[0034] In embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, 102c may implement both LTE radio access and NR radio access, for example, using the Dual Connectivity (DC) principle. Thus, the air interface utilized by WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies, and / or multiple types of base stations (e.g., eNBs and gNBs) that are transmitted to and from them.

[0035] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), High-Speed ​​Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0036] In Figure 1A, base station 114b may be, for example, a wireless router, home NodeB, home eNodeB, or access point, and may utilize any suitable RAT to facilitate wireless connectivity in localized areas such as offices, homes, vehicles, campuses, industrial facilities, air corridors (used by drones, for example), and roadways. In one embodiment, base station 114b and WTRU 102c, 102d may establish a wireless local area network (WLAN) by implementing wireless technology such as IEEE 802.11. In another embodiment, base station 114b and WTRU 102c, 102d may establish a wireless personal area network (WPAN) by implementing wireless technology such as IEEE 802.15. In yet another embodiment, base station 114b and WTRU 102c, 102d may establish a picocell or femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not need to access the internet 110 via CN 106 / 115.

[0037] RAN104 / 113 may communicate with CN106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not shown in Figure 1A, it will be understood that RAN104 / 113 and / or CN106 / 115 may communicate directly or indirectly with other RANs that utilize the same or different RAT as RAN104 / 113. For example, in addition to being connected to RAN104 / 113 which may utilize NR radio technology, CN106 / 115 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0038] CN106 / 115 may also serve as a gateway for WTRU102a, 102b, 102c, and 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing basic telephone services (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols, such as the Transmit Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) within the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may utilize the same RAT as RAN104 / 113 or a different RAT.

[0039] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode functionality (for example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different radio networks on different radio links). For example, WTRU 102c shown in Figure 1A may be configured to communicate with base station 114a which may employ cellular-based radio technology, and with base station 114b which may utilize IEEE 802 radio technology.

[0040] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any subcombinations of the above elements while maintaining consistency with the embodiment.

[0041] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors working with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, and the transceiver 120 may be coupled to the transmit / receive element 122. Although Figure 1B shows the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0042] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In another embodiment, the transmit / receive element 122 may be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of radio signals.

[0043] In Figure 1B, the transmit / receive element 122 is represented as a single element, but the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Therefore, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals over the air interface 116.

[0044] The transceiver 120 may be configured to modulate the signal to be transmitted by the transmit / receive element 122 and to demodulate the signal received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multimode capabilities. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0045] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input data from them. The processor 118 may output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may obtain information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and may store data therein. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identification module (SIM) card, a memory stick, and a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from memory located on a server or home computer (not shown) or the like, which is not physically located on the WTRU 102, and may store data in such memory.

[0046] The processor 118 may receive power from the power supply 134 and may be configured to distribute power to and / or control power to other components within the WTRU 102. The power supply 134 may be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell.

[0047] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) on the air interface 116 and / or determine its own position based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information using any suitable location determination method while maintaining consistency with the embodiment.

[0048] The processor 118 may also be coupled to other peripherals 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, and an activity tracker. The peripheral device 138 may include one or more sensors, one or more of which are gyroscopes, accelerometers, Hall effect sensors, magnetometers, compass sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0049] WTRU102 may include a full-duplex radio in which the transmission and reception of some or all of the signals associated with specific subframes for both UL (e.g., for transmission) and downlink (e.g., for reception) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit 139 to reduce and / or substantially eliminate self-interference via hardware (e.g., chokes) or via signal processing via a processor (e.g., a separate processor (not shown) or processor 118). In embodiments, WTRU102 may include a half-duplex radio for the transmission and reception of some or all of the signals (e.g., associated with specific subframes for either UL (e.g., for transmission) or downlink (e.g., for reception).

[0050] In Figures 1A and 1B, the WTRU is described as a wireless terminal, but in a typical embodiment, such a terminal is intended to be able to use a wired communication interface with a communication network (for example, temporarily or permanently).

[0051] In a typical embodiment, the other network 112 may be a WLAN.

[0052] In view of Figures 1A to 1B and the corresponding descriptions, one or more or all of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0053] Emulation devices may be designed to perform one or more tests of other devices in a laboratory environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions, fully or partially implemented and / or deployed as part of a wired and / or wireless communication network, to test other devices in a communication network. One or more emulation devices may perform one or more or all functions, temporarily implemented / deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for testing purposes and / or perform tests using over-air radio communication.

[0054] One or more emulation devices may perform one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test scenario, in a test laboratory and / or in an undeployed (e.g., test) wired and / or wireless communication network, to perform testing of one or more components. One or more emulation devices may also be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., including one or more antennas) may be used by the emulation device to transmit and / or receive data.

[0055] Block-based video coding Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 2 is a functional block diagram of an embodiment of a block-based hybrid video coding system. The input video signal 103 is processed block by block. Blocks may be referred to as coding units (CUs). In VTM-1.0, CUs may be up to 128 × 128 pixels. However, compared to HEVC, which partitions blocks based solely on quadtrees, in VTM-1.0, coding tree units (CTUs) may be partitioned into CUs based on quad / binary / ternary trees to accommodate varying characteristics. In addition, the concept of multiple partition unit types in HEVC may be eliminated, and as a result, VVC does not use the separation of CUs, prediction units (PUs), and transformation units (TUs); instead, each CU may be used as the basic unit for both prediction and transformation without further partitioning. In a multi-type tree structure, the CTU is first partitioned by a quadtree structure. Each quadtree leaf node may then be further partitioned by binary and ternary tree structures. As shown in Figures 3A to 3E, there may be five types of partitioning: quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0056] In Figure 2, spatial prediction (161) and / or temporal prediction (163) may be performed. Spatial prediction (or "intra-prediction") uses pixels from samples of already coded adjacent blocks (referred to as reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (referred to as "inter-prediction" or "motion-compensated prediction") uses unconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. A temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent, which is used to identify which reference picture in the reference picture store (165) the temporal prediction signal originates from.

[0057] After spatial and / or temporal prediction, a mode determination block (181) in the encoder selects the best prediction mode, for example, based on a rate distortion optimization method. The prediction block is then subtracted from the current video block (117), and the prediction residual is decorrelated and quantized (107) using a transform (105). The quantized residual coefficients are inversely quantized (111) and inversely transformed (113) to form the reconstructed residual, which is then added back to the prediction block (127) to form the reconstructed signal of the CU. Furthermore, in-loop filtering, such as a deblocking filter, may be applied to the reconstructed CU before placing it in the reference picture store (165) and may be used to code subsequent video blocks. To form the output video bitstream 121, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropycoding unit (109) to be further compressed and packed to form the bitstream.

[0058] Figure 4 is a functional block diagram of a block-based video decoder. The video bitstream 202 is unpacked and entropy-decoded in the entropy decoding unit 208. The coding mode and prediction information are sent to either the spatial prediction unit 260 (if intra-coding is used) or the temporal prediction unit 262 (if intercoding is used) to form prediction blocks. The residual transformation coefficients are sent to the inverse quantization unit 210 and the inverse transformation unit 212 to reconstruct the residual blocks. The prediction blocks and residual blocks are then added together in 226, and the reconstructed blocks further pass through in-loop filtering before being stored in the reference picture store 264. The reconstructed video in the reference picture store is then sent out to be used to predict later video blocks and to drive the display device.

[0059] As previously mentioned, BMS-1.0 adheres to the same encoding / decoding workflow as VTM-1.0, as shown in Figures 2 and 4. However, several coding modules, particularly those associated with time prediction, are further extended and enhanced. Below, some intertools included in BMS-1.0 or the previous JEM are briefly described.

[0060] Motion vector prediction To reduce the overhead of signaling motion information, as in HEVC, both VTM and BMS include two modes for coding the motion information of each CU: merged mode and non-merged mode. In merged mode, the motion information of the current CU is derived directly from the spatially adjacent block and the temporally adjacent block, and a competition-based scheme is applied to select the best adjacent block from all available candidates. In response, only the index of the best candidate is sent to the decoder to reconstruct the motion information of the CU. When the intercoded PU is coded in non-merged mode, the MV is coded differently using an MV predictor derived from the Evolutionary Motion Vector Prediction (AMVP) technique. As in merged mode, AMVP derives the MV predictor from the spatially adjacent candidate and the temporally adjacent candidate. The difference between the MV predictor and the actual MV, and the index of the predictor, are then sent to the decoder.

[0061] Figure 5 shows an example of spatial MV prediction. In the current picture (CurrPic) to be coded, the current CU is a square CurrCU, which has the best matching block (CurrRefCU) in the reference picture (CurrRefPic). The MV of CurrCU, i.e., MV2, will be predicted. The spatial neighbors of the current CU are the CU adjacent above, the CU adjacent to the left, the CU adjacent to the upper left, the CU adjacent to the lower left, and the CU adjacent to the upper right. In Figure 5, the neighboring CUs are shown as the CU adjacent above, NeighbCU. Both the reference picture (NeighbRefPic) and MV (MV1) of NeighbCU are known because NeighbCU was coded before CurrCU.

[0062] Figure 6 shows an example of time-motion video prediction (TMVP). Four pictures (ColRefPic, CurrRefPic, ColPic, CurrPic) are shown in Figure 6. In the current picture to be coded (CurrPic), the current CU is a square CurrCU, which has the best matching block (CurrRefCU) in the reference picture (CurrRefPic). The MV of CurrCU, i.e., MV2, will be predicted. The time adjacency of the current CU is specified as a collocated CU (ColCU) in the adjacent picture (ColPic). Both the reference picture (ColRefPic) and MV (MV1) of ColCU are known because ColPic was coded before CurrPic.

[0063] For spatial motion vector predictions and temporal motion vector predictions, assuming that time and space are constrained, the MV between different blocks is treated as translational with a uniform velocity. In the embodiments of Figures 5 and 6, the temporal distance between CurrPic and CurrRefPic is TB, and the temporal distance between CurrPic and NeighbRefPic in Figure 5, or between ColPic and ColRefPic in Figure 6, is TD. The scaled MV predictor may be calculated as given by equation (1).

[0064]

number

[0065] In VTM-1.0, each merge block has at most one set of motion parameters (one motion vector and one reference picture index) for each prediction direction L0 and L1. In contrast, additional merge candidates based on Evolutionary Time Motion Vector Prediction (ATMVP) are included in BMS-1.0 to enable the derivation of motion information at the subblock level. Using such a mode, time motion vector prediction is improved by enabling a CU to derive multiple MVs for subblocks within the CU. Generally, ATMVP derives motion information for the current CU in two steps, as shown in Figure 7. The first step is to identify the corresponding block of the current block (referred to as the juxtaposed block) in the time reference picture. The selected time reference picture is referred to as the juxtaposed picture. The second step is to divide the current block into subblocks and derive motion information for each subblock from the corresponding small blocks in the juxtaposed picture.

[0066] In the first step, juxtaposed blocks and juxtaposed pictures are identified by the motion information of the spatially adjacent blocks of the current block. In the current design, the first available candidate in the merge candidate list is considered. Figure 7 illustrates this process. In particular, in the embodiment of Figure 7, block A is identified as the first available merge candidate of the current block based on the traversal order of the merge candidate list. Block A's reference index, along with its corresponding motion vector (MVA), is then used to identify juxtaposed pictures and juxtaposed blocks. The position of a juxtaposed block within a juxtaposed picture is determined by adding the motion vector (MVA) of block A to the coordinates of the current block.

[0067] In the second step, for each subblock within the current block, the motion information of its corresponding small block within the juxtaposed block (indicated by the small arrow in Figure 7) is used to derive the motion information of the subblock. In particular, after the motion information of each small block within the juxtaposed block is identified, it is converted into the motion vector and reference index of the corresponding subblock within the current block in the same manner as TMVP.

[0068] Decoder-side motion vector refinement (DMVR) For merge modes in VTM, when a selected merge candidate is bipredicted, the prediction signal for the current CU is formed by averaging two prediction blocks using two MVs associated with the candidate's reference lists L0 and L1. However, the motion information of the merge candidate (derived from either the spatial adjacency or spatial adjacency of the current CU) may not be accurate enough to represent the true motion of the current CU, thus degrading the efficiency of interpretation. To further improve the coding performance of merge modes, a decoder-side motion vector refinement (DMVR) method is applied in BMS-1.0 to refine the MVs of merge modes. In particular, when a selected merge candidate is bipredicted, the biprediction template is initially generated as the average of two prediction signals based on the MVs from reference lists L0 and L1, respectively. Block-matching based motion refinement is then performed locally around the initial MV, using the biprediction template as the target, as described below.

[0069] Figure 8A illustrates the motion refinement process applied in DMVR. Generally, DMVR refines the MV of a merge candidate in two steps. As shown in Figure 8A, in the first step, a dual prediction template is generated by averaging two prediction blocks using the initial MVs (i.e., MV0 and MV1) in L0 and L1 of the merge candidate. Next, for each reference list (i.e., L0 or L1), a block-matching based motion search is performed within the local region around the initial MV. For each MV, i.e., for MV0 or MV1 in that list around the initial MV in the corresponding reference list, a cost value (e.g., sum of absolute difference (SAD)) is measured between the dual prediction template and the corresponding prediction block using its motion vector. For each of the two prediction directions, the MV that minimizes the template in that prediction direction is considered the final MV in the reference list of the merge candidate. In the current BMS-1.0, for each prediction direction, eight adjacent MVs (with one integer sample offset) surrounding the initial MV are considered during the motion refinement process. Finally, two refined MVs (MV0' and MV1' shown in Figure 8A) are used to generate the final dual prediction signals for the current CU. In addition, in conventional DMVRs, to further improve coding efficiency, the refined MVs of a DMVR block are used to predict the motion information of its spatially adjacent blocks and temporally adjacent blocks (e.g., based on spatial AMVP, spatial merge candidate, TMVP, and ATMVP), as well as to calculate the boundary strength value of the deblocking filter applied to the current CU.Figure 8B is a flowchart of an example of DMVR processing, where "Spatial AMVP" and "Spatial Merge Candidate" refer to spatial MV prediction processing for spatially adjacent CUs that are in the current picture and are coded after the current CU according to the CU coding order, "TMVP" and "ATMVP" refer to temporal MV prediction processing for subsequent CUs in subsequent pictures (pictures coded after the current picture based on the picture coding order), and "Deblocking" refers to deblocking filtering processing for both the current block and its spatially adjacent blocks.

[0070] In the method shown in Figure 8B, a biprediction template is generated at 802. Motion refinement is performed on the L0 motion vector at 804, and motion refinement is performed on the L1 motion vector at 806. In 808, the final biprediction is generated using the refined L0 and L1 motion vectors. In the method of Figure 8B, the refined motion vectors are used to predict the motion of blocks to be coded subsequently. For example, the refined motion vectors are used for spatial AMVP (810), TMVP (814), and ATMVP (816). The refined motion vectors are also used as spatial merge candidates (812) and to calculate boundary intensity values ​​for the deblocking filter applied to the current CU (818).

[0071] Bidirectional optical flow In VTM / BMS-1.0, dual predictions are a combination of two time-prediction blocks obtained from a reference picture already reconstructed using averaging. However, due to the limitations of block-based motion compensation, they may be the remaining small motion that can be obtained between the two prediction blocks, thus reducing the efficiency of motion-compensated predictions. To address this problem, Bidirectional Optical Flow (BIO) is used in JEM to compensate for such motion on a sample-by-sample basis within a block. In particular, BIO is a sample-by-sample motion refinement performed on top of block-based motion-compensated predictions when dual predictions are used. The derivation of the refined motion vector for each sample within a block is based on a classical optical flow model. (k) (x,y) is the sample value in the prediction block matrix (x,y) derived from the reference picture list k (k=0,1), and ∂I (k) (x,y) / ∂x and ∂I (k) Let (x,y) / ∂y be the horizontal and vertical slopes of the sample. The dual prediction signal corrected by BIO is obtained as equation (2).

[0072]

number

[0073] τ0 and τ1 are I to the current picture. (0) and I (1) This is the time distance of the associated reference pictures Ref0 and Ref1. Furthermore, it is the motion refinement (v) at the sample position (x,y). x ,v y ) is calculated by minimizing the difference Δ between the sample values ​​after motion refinement compensation, as shown in equation (3).

[0074]

number

[0075] In addition, in order to bring about the regularity of the derived motion refinement, it is assumed that the motion refinement is consistent within a local surrounding area centered at (x,y), and therefore, as shown in equation (4), by minimizing the optical flow error metric Δ inside the 5×5 window Ω around the current sample at (x,y), (v x ,v y The value of ) is derived.

[0076]

number

[0077] Unlike DMVR, motion refinement is derived by BIO (v x ,v y It should be noted that this is applied only to enhance the biprediction signal and not to correct the motion information of the current CU. In other words, the MV used to predict the MV of spatially adjacent blocks and temporally adjacent blocks and to determine the deblocking boundary strength of the current CU is still the original MV (i.e., the block-based motion compensation signal I before BIO is applied). (0) (x,y) and I (1) (MV used to generate (x,y)).

[0078] DMVR coding waiting time Like HEVC and its predecessors, VTM-1.0 employs motion-compensated prediction (MCP) to efficiently reduce temporal redundancy between pictures, thereby achieving high intercoding efficiency. Because the MV used to generate the prediction signal for a single CU is either signaled in the bitstream or inherited from its spatial / temporal neighbors, there is no dependency between MCPs of spatially adjacent CUs. As a result, the MCP processing of all interblocks within the same picture / slice is independent of each other. Therefore, with VTM-1.0 and HEVC, decoding of multiple interblocks can be performed in parallel; for example, they may be assigned to different threads to take advantage of parallelism.

[0079] As described above, the DMVR tool is applied in BMS-1.0. To avoid generating extra signaling overhead, motion refinement is derived using two prediction signals associated with the original L0 and L1 MVs of the CU. Thus, when motion information for the CU is predicted from one of its spatial neighbors coded by the DMVR (e.g., by AMVP and merge mode), its decoding process waits until the MV of the neighboring block is fully reconstructed by the DMVR. This significantly complicates the pipeline design, particularly on the decoder side, and therefore leads to a significant increase in complexity for the hardware implementation.

[0080] To illustrate the coding latency caused by the DMVR, Figures 9 and 10 show an example comparing the decoding processes of VTM-1.0 and BMS-1.0. For ease of explanation, a case is described in which there are four CUs of equal block size, all four CUs are coded by the DMVR, each of them is decoded by a separate decoding thread, and the decoding complexity of each individual decoding module (e.g., MCP, DMVR, inverse quantization, and inverse transform) is assumed to be the same for all four CUs. As shown in Figure 9, because the four CUs can be decoded in parallel, the total decoding time of VTM-1.0 is equal to the decoding time of one CU, i.e., T MCP +T de-quant +T inv-trans This is equal to: Due to dependencies introduced by the DMVR, the decoding process for BMS-1.0 (shown in Figure 10) cannot call the decoding of each individual coding block until the DMVR of its spatially adjacent block has completely finished. Therefore, the total decoding time for four CUs for BMS-1.0 is T total =4 × (T MCP +T DMVR )+T de-quant +T inv-trans It is equivalent to this. As can be understood, the use of predictive samples to refine motion information by DMVR introduces dependencies within adjacent interblocks, and therefore significantly increases latency for both encoding and decoding processes.

[0081] Overview of methods to reduce waiting times This disclosure proposes a method for eliminating or reducing the encoding / decoding latency of a DMVR while preserving its primary coding performance. In particular, various embodiments of the disclosure include one or more of the following aspects:

[0082] Unlike the current DMVR method in BMS-1.0, where a refined DMVR motion of one block is always used to predict the motion of the spatial / temporal adjacent block and derive the deblocking filter structure, in some embodiments, it is proposed to use the unrefined MV of the DMVR block (the MV used to generate the original biprediction signal) completely or partially for MV prediction and deblocking processes. Assuming that the original MV can be directly obtained from parsing and motion vector reconstruction (the parsed motion vector difference in addition to the motion vector predictor) without DMVR, there is no dependency between adjacent blocks and the decoding of multiple interCUs can be performed in parallel.

[0083] Since unrefined MVs may have lower accuracy than refined MVs, this can lead to some degradation in coding performance. To mitigate such losses, in some embodiments, it is proposed to divide the picture / slice into multiple regions. Furthermore, in some embodiments, additional constraints are proposed to allow independent decoding of multiple CUs within the same region or multiple CUs from different regions.

[0084] In some embodiments, optical flow-based motion derivation methods are proposed to replace block-matching-based motion search for calculating the motion refinement of each DMVR CU. Compared to block-matching-based methods that perform motion search within a small local window, some embodiments directly calculate the motion refinement based on spatial sample derivatives and time sample derivatives. This can reduce computational complexity and increase motion refinement accuracy because the derived refined motion values ​​are not limited to the search window.

[0085] Using unrefined motion vectors to reduce DMVR latency As noted above, using the refined MV of one DMVR block as the MV predictor for its adjacent blocks is inappropriate for parallel coding / decoding to a true CODEC design, because coding / decoding of adjacent blocks is not performed until the refined MV of the current block is completely reconstructed through the DMVR. Based on such analysis, this chapter proposes a method for eliminating coding latency introduced by the DMVR. In some embodiments, the core design of the DMVR (e.g., block matching-based motion refinement) remains identical to existing designs. However, the MV of the DMVR block used to perform MV prediction (e.g., AMVP, merge, TMVP, and ATMVP) and deblocking is modified to eliminate dependencies between adjacent blocks introduced by the DMVR.

[0086] Use of unrefined motion vectors for spatial and temporal motion predictions In some embodiments, it is proposed to always perform MV prediction and deblocking using the unrefined motion of the DMVR block instead of using the refined motion. Figure 11 shows the modified DMVR processing after such a method has been applied. As shown in Figure 11, instead of using the refined MV, the unrefined MV (the original MV before the DMVR) is used to derive the MV predictor and determine the boundary strength of the deblocking filter. Only the refined MV is used to generate the final biprediction signal of the block. Such embodiments may be used to eliminate the encoding / decoding latency of the DMVR because there is no dependency between the refined MV of the current block and the decoding of its adjacent blocks.

[0087] Using unrefined motion vectors for spatial motion prediction In the embodiment shown in Figure 11, the unrefined MV of the DMVR block is used to derive a time-motion predictor for juxtaposed blocks in a later picture via TMVP and ATMVP, and to calculate the boundary strength for the deblocking filter between the current block and its spatial adjacencies. This can lead to some coding performance loss because the unrefined MV may be less accurate than the refined MV. On the other hand, time-motion prediction (TMVP and ATMVP) predicts the MV in the current picture using the MV of a previously decoded picture (in particular, a juxtaposed picture). Therefore, the refined MV of the DMVR CU in the juxtaposed picture has already been reconstructed before performing time-motion prediction on the current picture. A similar situation is also applicable to the deblocking filter process, which can only be invoked after the samples of the current block have been completely reconstructed through MC (including DMVR), inverse quantization, and inverse transform, because the deblocking filter is applied to reconstructed samples. Therefore, the refined MV is already available before deblocking is applied to the DMVR block.

[0088] In the method shown in Figure 11, at 1100, the unrefined motion vector for the first block is identified. The unrefined motion vector may be signaled for the first block using one of the various available MV signaling techniques. At 1102, the unrefined motion vector is used to generate a biprediction template. At 1104, motion refinement is performed on the L0 motion vector, and at 1106, motion refinement is performed on the L1 motion vector. At 1108, the refined L0 and L1 motion vectors are used to generate the final biprediction for the first block. In the method of Figure 11, the unrefined motion vector is used to predict the motion of subsequently coded blocks (e.g., a second block). For example, the unrefined motion vector is used for spatial AMVP (1110), TMVP (1114), and ATMVP (1116). The unrefined motion vectors are also used as spatial merge candidates (1112) and to calculate the boundary intensity values ​​for the deblocking filter (1118).

[0089] In another embodiment, to address these issues and achieve better coding performance, it is proposed to use different MVs (unrefined MV and refined MV) of the DMVR block for spatial motion prediction, temporal motion prediction, and deblocking filters. In particular, in this embodiment, only the unrefined MV is used to derive MV predictors for spatial motion prediction (e.g., spatial AMVP and spatial merge candidates), while the refined MV is used not only to derive the final prediction of the block but also to generate MV predictors for temporal motion prediction (TMVP and ATMVP) and to calculate boundary strength parameters for the deblocking filter. Figure 12 shows the DMVR processing according to this second embodiment.

[0090] In the method shown in Figure 12, at 1200, the unrefined motion vector for the first block is identified. The unrefined motion vector may be signaled for the first block using one of the various available MV signaling techniques. At 1202, the unrefined motion vector is used to generate a biprediction template. At 1204, motion refinement is performed on the L0 motion vector, and at 1206, motion refinement is performed on the L1 motion vector. At 1208, the refined L0 and L1 motion vectors are used to generate the final biprediction for the first block. In the method of Figure 12, the unrefined motion vector is used to predict the motion of a subsequent block coded in the same picture as the first block (e.g., a second block). For example, the unrefined motion vector is used for a spatial AMVP (1110) and as a spatial merge candidate (1212). For example, the refined motion vector is used to predict the motion of subsequent coded blocks (e.g., a third block) in another picture, using TMVP(1214) or ATMVP(1216). The refined motion vector is also used to calculate the boundary intensity values ​​of the deblocking filter(1218).

[0091] Use of unrefined motion vectors for spatial motion prediction and deblocking In the embodiment shown in Figure 12, different MVs of the DMVR block are used for spatial motion prediction and deblocking filters. However, unlike the MV used for time motion prediction (stored in external memory), the MVs used for spatial motion prediction and deblocking are often stored using on-chip memory for practical CODEC designs to increase data access speed. Therefore, some implementations of the method in Figure 12 require two different on-chip memories to store both the unrefined and refined MVs for each DMVR block. This can double the line buffer size used to cache the MVs, which can be undesirable for hardware implementations. To maintain the same total on-chip memory size for MV storage as in VTM-1.0, further embodiments propose using the unrefined MVs of the DMVR block for the deblocking process. Figure 13 shows an example of DMVR processing according to this embodiment. In particular, as in the method in Figure 12, refined DMVR MVs are also used to generate time motion predictors through TMVP and ATMVP, in addition to generating the final dual-prediction signal. However, in the embodiment shown in Figure 13, the unrefined MV is used not only to derive the spatial motion predictor (spatial AMVP and spatial merge), but also to determine the boundary strength for the deblocking filter of the current block.

[0092] In the method shown in Figure 13, at 1300, the unrefined motion vector for the first block is identified. The unrefined motion vector may be signaled for the first block using one of the various available MV signaling techniques. At 1302, the unrefined motion vector is used to generate a biprediction template. At 1304, motion refinement is performed on the L0 motion vector, and at 1306, motion refinement is performed on the L1 motion vector. At 1308, the refined L0 and L1 motion vectors are used to generate the final biprediction for the first block. In the method of Figure 13, the unrefined motion vector is used to predict the motion of a subsequent block coded in the same picture as the first block (e.g., a second block). For example, the unrefined motion vector is used for a spatial AMVP (1310) and as a spatial merge candidate (1312). For example, refined motion vectors are used to predict the movement of subsequent coded blocks (e.g., a third block) in other pictures, using TMVP(1314) or ATMVP(1318).

[0093] The embodiments in Figures 11 to 13 can reduce or eliminate the encoding / decoding latency caused by the DMVR, assuming that the dependency of decoding one block on the reconstruction of the refined MV of the spatially adjacent DMVR blocks does not exist in those embodiments. Based on the same embodiment in Figure 10, Figure 14 shows an embodiment of parallel decoding when one of the methods in Figures 11 to 13 is applied. As shown in Figure 14, there is no decoding latency between adjacent blocks because decoding of multiple DMVR blocks can be performed in parallel. Correspondingly, the total decoding time may be equal to the decoding of one block, which is T MCP +T DMVR +T de-quant +T inv-trans It may also be expressed as follows.

[0094] Segment-based methods for reducing DMVR latency As noted above, one cause of coding / decoding latency for DMVR is the dependency between the reconstruction of the refined MV of a DMVR block and the decoding of its adjacent blocks, which is incurred by spatial motion prediction (e.g., spatial AMVP and spatial merge modes). Methods such as those shown in Figures 11-13 can eliminate or reduce coding latency for DMVR, but this reduced latency may come at the cost of degraded coding efficiency due to the use of less accurate, unrefined MVs for spatial motion prediction. On the other hand, as shown in Figure 10, the worst-case coding / decoding latency caused by DMVR is directly related to the maximum number of consecutive blocks coded by the DMVR mode. To address these issues, in some embodiments, region-based methods are used to reduce coding / decoding latency and to mitigate coding losses resulting from the use of unrefined MVs for spatial motion prediction.

[0095] In particular, in some embodiments, the picture is divided into multiple non-overlapping segments, and the unrefined MV of each DMVR block within a segment is used as a predictor to predict the MV of its adjacent blocks within the same segment. However, when a DMVR block is located on the right or lower boundary of a segment, its unrefined MV is not used; instead, the refined MV of the block is used as a predictor to predict the MV of the block from adjacent segments for better spatial motion prediction efficiency.

[0096] Figure 15 shows an example of DMVR processing according to one embodiment, and Figure 16 shows an embodiment in which blank blocks represent DMVR blocks using unrefined MVs for spatial motion prediction, spatial merging, and deblocking, and patterned blocks represent DMVR blocks using refined MVs for spatial motion prediction, spatial merging, and deblocking. In the embodiment of Figure 16, encoding / decoding of different interblocks within the same segment may be performed independently of each other, while decoding of blocks from different segments is still dependent. For example, blocks on the left boundary of segment #2 cannot begin their decoding process until the DMVR of those adjacent blocks in segment #1 is complete, because they can use the refined MV of an adjacent DMVR block in segment #1 as a spatial MV predictor. In addition, as shown in Figure 15, similar to the method in Figure 13, the same MV of a single DMVR block is used for spatial motion prediction and deblocking filters to avoid increasing the on-chip memory for storing MVs. In another embodiment, it is proposed to always use a refined MV for the deblocking process.

[0097] In the method shown in Figure 15, at 1502, the unrefined motion vector for the first block is identified. The unrefined motion vector may be signaled for the first block using any of the various available MV signaling techniques. At 1504, the unrefined motion vector is used to generate a biprediction template. At 1506, motion refinement is performed on the L0 motion vector, and at 1508, motion refinement is performed on the L1 motion vector. At 1510, the refined L0 and L1 motion vectors are used to generate the final biprediction for the first block.

[0098] In step 1512, it is determined whether the first block is located on the right-side segment boundary or the lower-side segment boundary. If the first block is not located on the right-side segment boundary or the lower-side segment boundary, then the unrefined motion vector is used to predict the motion of subsequent blocks coded in the same picture as the first block (e.g., the second block). For example, the unrefined motion vector is used for the spatial AMVP (1514) and as a spatial merge candidate (1516). The unrefined motion vector is also used to calculate the boundary intensity value of the deblocking filter (1518). On the other hand, if the first block is located on the right-side segment boundary or the lower-side segment boundary, then the refined motion vector is used to predict the motion of subsequent blocks coded in the same picture as the first block (e.g., the second block) (e.g., by AMVP 1514 and spatial merge candidate 1516), and the refined motion vector is also used to calculate the boundary intensity value of the deblocking filter (1518). Regardless of the result of the determination in 1512, the refined motion vector is used, for example, TMVP(1520) or ATMVP(1522) to predict the motion of subsequent blocks coded in other pictures (e.g., a third block).

[0099] In the embodiment shown in Figure 16, only the refined MV is enabled for spatial motion prediction of blocks located on the left / upper boundary of a segment within a single picture. However, depending on the segment size, the overall proportion of blocks to which the refined MV can be applied for spatial motion prediction may be relatively small. This can still result in a performance degradation that is not negligible for spatial motion prediction. To further improve performance, in some embodiments, it is proposed that the refined MV of a DMVR block within a segment be allowed to predict the MV of an adjacent block within the same segment. However, as a result, it is not possible to decode multiple blocks within a single segment in parallel. To improve encoding / decoding parallelism, this method also proposes prohibiting the current block from using the MV (either unrefined or refined MV) of an adjacent block from another segment as a predictor for spatial motion prediction (e.g., spatial AMVP and spatial merge). In particular, in such a method, if an adjacent block is from a different segment to the current block, it is treated as unavailable for spatial motion vector prediction.

[0100] One such embodiment is shown in Figure 17. In Figure 17, blank blocks represent CUs that are permitted to use adjacent MVs for spatial motion prediction (adjacent MVs are refined MVs if the adjacent block is a single DMVR block, or unrefined MVs if not), and patterned blocks represent CUs that are prevented from using the MVs of their adjacent blocks from different segments for spatial motion prediction. The embodiment according to Figure 17 enables parallel decoding of interblocks that are not within a single segment but span across segments.

[0101] Generally, DMVR is only enabled for bidirectional predicted CUs that have both forward and backward prediction signals. In particular, DMVR requires the use of two reference pictures, one with a smaller picture order count (POC) and the other with a larger POC than the current picture. In contrast, low-latency (LD) pictures are predicted from reference pictures that both precede the current picture in display order and have POCs of all reference pictures in L0 and L1 smaller than the current picture's POC. Therefore, DMVR cannot be applied to LD pictures, and the coding latency caused by DMVR does not exist for LD pictures. Based on such analysis, in some embodiments, it is proposed that when DMVR is applied, only the above DMVR parallelism constraint (disabling spatial motion prediction across segment boundaries) be applied to non-LD pictures. For LD pictures, no constraint is applied, and it is still permissible to predict the MV of the current block based on its spatially adjacent MVs from another segment. In a further embodiment, the encoder / decoder determines whether the constraint applies based on checking the POC of all reference pictures in L0 and L1, without additional signaling. In another embodiment, it is proposed to add a picture / slice-level flag to indicate whether the DMVR parallelism constraint applies to the current picture / slice.

[0102] In some embodiments, the number of segments within a picture / slice and the position of each segment are selected by the encoder and signaled to the decoder. Signaling may also be performed on other parallelism tools in HEVC and JEM (e.g., slice, tile, and wavefront parallelism (WPP)). Various choices may lead to different trade-offs between coding performance and encoding / decoding parallelism. In one embodiment, it is proposed to set the size of each segment to be equal to the size of one CTU. From a signaling standpoint, syntax elements may be added at the sequence level and / or picture level. For example, the number of CTUs in each segment may be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS), or in the slice header. Other variations of the syntax elements may be used, for example, among other alternatives, the number of CTU rows in each picture / slice may be used, or the number of segments in each picture / slice may be used.

[0103] Examples of motion refinement methods Additional embodiments described herein serve to replace block-matching motion searches for calculating DMVR motion refinements. Compared to block-matching-based methods that perform motion searches within a small local window, embodiments of the embodiments directly calculate motion refinements based on spatial and time sample derivatives. Such embodiments can reduce computational complexity and increase refinement accuracy because the derived refined motion values ​​are not limited to a search window.

[0104] Motion refinement using block-level BIO As discussed above, when blocks are bipredicted, BIO is used in JEM to provide sample-by-sample motion refinement to the top of the block-based motion-compensated predictions. Based on the current design, BIO only enhances the motion-compensated predicted samples as a result of refinement without updating the MVs, which are stored in the MV buffer and used for spatial motion predictions, temporal motion predictions, and deblocking filters. This is in contrast to the current DMVR, where BIO does not introduce any encoding / decoding latency between adjacent blocks. However, in the current BIO design, motion refinement is derived in small units (e.g., 4x4). This leads to a computational complexity that cannot be ignored, especially on the decoder side. This is undesirable for hardware CODEC implementations. Therefore, in some embodiments, to address the latency of the DMVR while maintaining acceptable coding complexity, it is proposed to use a block-based BIO to compute local motion refinement for video blocks coded by the DMVR. In particular, in the proposed embodiment, the core design of the BIO (e.g., calculation of gradients and refined motion vectors) is maintained identically to that in existing designs for calculating motion refinement. However, to reduce complexity, the amount of motion refinement is derived based on the CU level, and a single value is aggregated for all samples within the CU and used to calculate a single motion refinement, so that all samples within the current CU share the same motion refinement. Based on the same notation used above for BIO, an example of the proposed block-level BIO motion refinement is derived as equation (5).

[0105]

number

[0106] θ is a set of coordinates of the sample in the current CU, and Δ(x,y) is the optical flow error metric as shown in equation (3) above.

[0107] As shown above, the motivation for BIO is to improve the accuracy of predicted samples based on local gradient information at each sample location within the current block. For large video blocks containing many samples, local gradients at different sample locations may exhibit highly variable characteristics. In such cases, the block-based BIO derivation described above may not yield reliable motion refinement for the current block, leading to a loss of coding performance. Based on such considerations, in some embodiments, it is proposed to enable only CU-based BIO motion derivation for DMVR blocks when the block size is small (e.g., below a given threshold). Otherwise, CU-based BIO motion derivation is disabled, and instead, existing block-matching-based motion refinement (in some embodiments, the proposed DMVR latency elimination / reduction method described above is applied) is used to derive local motion refinement for the current block.

[0108] Motion refinement using optical flow As mentioned above, BIO assumes that the derived L0 and L1 motion refinements at each sample position are symmetrical around the current picture, i.e.,

[0109]

number

[0110] and

[0111]

number

[0112] Based on this assumption, we estimate local motion refinement.

[0113]

number

[0114] and

[0115]

number

[0116] These are horizontal and vertical motion refinements associated with prediction lists L0 and L1. However, such assumptions may not apply to blocks coded by the DMVR. For example, in existing DMVRs (as shown in Figure 8A), two separate block matching-based motion searches are performed on L0 and L1 so that the MVs that minimize the template costs of the L0 and L1 prediction signals can be different. Due to such symmetric motion constraints, the motion refinements derived by the BIO may not always be accurate in improving (and sometimes even degrading) the prediction quality for the DMVR.

[0117] In some embodiments, an improved motion derivation method is used to calculate motion refinement for the DMVR. The classical optical flow model states that the brightness of the picture remains constant over time, as expressed by equation (6).

[0118]

number

[0119] x and y represent spatial coordinates, and t represents time. The right side of equation (6) may be expanded by a Taylor expansion with respect to (x,y,t). Then, the equation for optical flow becomes linearly equation (7).

[0120]

number

[0121] Using the camera acquisition time as the basic unit of time (for example, by setting dt=1), equation (7) may be discretized by changing the optical flow function from a continuous domain to a discrete domain. Let I(x,y) be the sample values ​​acquired from the camera, then equation (7) becomes equation (8).

[0122]

number

[0123] In various embodiments, one or more error metrics may be defined based on the degree to which the leftmost expression in equation (9) is not equal to zero. Motion refinement may be employed to substantially minimize the error metric.

[0124] In some embodiments, it is proposed to use a discretized optical flow model to estimate local motion refinements in L0 and L1. In particular, the biprediction template is generated by averaging two prediction blocks using the initial L0 and L1 MVs of the merge candidate. However, instead of performing a block-matching motion search within the local region, the optical flow model in equation (8) is used in some proposed embodiments to directly derive the refined MVs for each reference list L0 / L1, as shown in equation (9).

[0125]

number

[0126] P 0 and P 1P is the predicted signal generated using the original MV for reference lists L0 and L1, respectively. tmp This is a biprediction template signal,

[0127]

number

[0128] and

[0129]

number

[0130] The predicted signal P can be calculated based on different gradient filters, such as the Sobel filter or the 2D separable gradient filter used by BIO (described in J. Chen, E. Alshina, G. Sullivan, J. R. Ohm, J. Boyce, “Algorithm description of joint exploration test model 6”, JVET-G1001, Jul. 2017, Torino, Italy). 0 and P 1 This is the horizontal / vertical slope. Equation (9) gives one individual for it

[0131]

number

[0132] ,

[0133]

number

[0134] , and P tmp -P k Predicted signal P can be calculated. 0 or P 1This represents a set of equations for each sample in the given context. Two unknown parameters Δx k and Δy k Therefore, by minimizing the sum of squared errors in equation (9), as shown in equation (10), the over-determined problem can be solved.

[0135]

number

[0136]

number

[0137] Δx is the time difference between the L0 / L1 prediction signal and the dual prediction template signal, and θ is the set of coordinates in the coding block. By solving the linear least mean squared error (LLMSE) problem in equation (10), we obtain equation (11), (Δx,Δy) * k The parsed representation can be obtained.

[0138]

number

[0139] Based on equation (11), in some embodiments, in order to improve the accuracy of the derived MV, such methods involve motion refinement in a recursive formula (i.e., (Δx, Δy) * k ) can be selected. Such an embodiment uses the original L0 and L1 MV of the current block to generate an initial biprediction template signal and, based on equation (11), the corresponding delta motion (Δx, Δy) * k This can be done by calculating the refined MV, which is then used as a movement to generate a dual-prediction template sample along with new L0 and L1 prediction samples, and the dual-prediction template sample is then locally refined (Δx,Δy)* k This is used to update the value of . This process may be repeated until MV is no longer updated or until the maximum number of iterations is reached. One embodiment of such a process is summarized by the following procedure, as shown in Figure 18.

[0140] In step 1802, the counter l is initialized to l=0. In step 1804, the initial L0 and L1 prediction signals are generated.

[0141]

number

[0142] and

[0143]

number

[0144] , as well as the initial biprediction template signal

[0145]

number

[0146] This is the original MV of the block.

[0147]

number

[0148] and

[0149]

number

[0150] Generated using: Local L0 and L1 motion refinement based on equation (11) in 1806 and 1808.

[0151]

number

[0152] and

[0153]

number

[0154] , as well as the MV for Block,

[0155]

number

[0156] and

[0157]

number

[0158] It will be updated as follows.

[0159]

number

[0160] and

[0161]

number

[0162] If it is zero (determined at 1810), or l=l max If this is the case (determined in 1812), then in 1814, the final biprediction may be generated using the refined motion vector. Otherwise, in 1816, counter l is incremented, and the process is repeated, MV

[0163]

number

[0164] and

[0165]

number

[0166] Using L0 and L1 prediction signals

[0167]

number

[0168] and

[0169]

number

[0170] and dual prediction template signals

[0171]

number

[0172] It was updated (in 1806 and 1808).

[0173] Figure 18 shows an example of DMVR processing using the optical-flow-based motion derivation method of an embodiment for calculating motion refinement of a DMVR block. As shown in Figure 18, the optical MV of a single DMVR block is identified by iteratively modifying the original MV based on an optical flow model. While such a method can yield good motion estimation accuracy, it can also lead to a significant increase in complexity. To reduce the complexity of the derivation, in one embodiment of the disclosure, it is proposed to apply only one iteration to derive motion refinement using the proposed motion derivation method, for example, to apply the processing shown in 1804 to 1808 to derive the modified MV of the DMVR block.

[0174] The optical flow-based motion derivation model can be more efficient for smaller CUs than for larger CUs due to the high consistency in the characteristics of the sample within the small block. In some embodiments, it is proposed to enable the proposed optical flow-based motion derivation for DMVR blocks when the block size is small (e.g., below a given threshold). Otherwise, existing block matching-based motion refinement is used to derive local motion refinement for the current block (e.g., according to the proposed DMVR latency elimination / reduction method described herein).

[0175] It should be noted that one or more different hardware elements of the embodiments described refer to “modules” that perform various functions described herein in relation to each module. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that is recognized by those skilled in the art for a given implementation. It should be noted that each described module may also include instructions executable to perform one or more functions described as being performed by each module, and such instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, and / or software instructions, or may include them, and may be stored in any suitable non-temporary computer-readable medium commonly referred to as RAM, ROM, etc.

[0176] Although features and elements have been described above in specific combinations, those skilled in the art will recognize that each feature or element may be used alone or in any combination with other features and elements. In addition, the methods described herein may be implemented in computer programs, software, or firmware embedded on computer-readable media for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multipurpose disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a TRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video decoding method, The first unrefined motion vector and the second unrefined motion vector of the first block are refined to generate the first refined motion vector and the second refined motion vector, Using the first refined motion vector and the second refined motion vector, a prediction of the first block is generated, The boundary filter strength of the first block is determined based at least partially on the first unrefined motion vector and the second unrefined motion vector, Includes, The first unrefined motion vector and the second unrefined motion vector are not used to generate the prediction in the method.

2. The method according to claim 1, further comprising applying a filter to at least one boundary of the first block using the boundary filter strength determined above.

3. The method according to claim 1, further comprising predicting motion information of a second block using at least one of the first refined motion vector and the second refined motion vector, wherein the second block and the first block are juxtaposed blocks of different pictures.

4. The method according to claim 1, wherein refining the first unrefined motion vector and the second unrefined motion vector is performed using decoder-side motion vector refinement (DMVR).

5. The method according to claim 1, wherein refining the first unrefined motion vector and the second unrefined motion vector includes selecting the first refined motion vector and the second refined motion vector such that the sum of absolute differences substantially minimizes.

6. A video decoding device comprising one or more processors, The one or more processors described above are The first unrefined motion vector and the second unrefined motion vector of the first block are refined to generate the first refined motion vector and the second refined motion vector, Using the first refined motion vector and the second refined motion vector, a prediction of the first block is generated, The boundary filter strength of the first block is determined based at least partially on the first unrefined motion vector and the second unrefined motion vector, Configured to perform at least the following: The apparatus is not used to generate the prediction, the first unrefined motion vector and the second unrefined motion vector.

7. The apparatus according to claim 6, further configured to apply a filter to at least one boundary of the first block using the determined boundary filter strength.

8. The apparatus according to claim 6, further configured to predict motion information of a second block using at least one of the first refined motion vector and the second refined motion vector, wherein the second block and the first block are juxtaposed blocks of different pictures.

9. The apparatus according to claim 6, wherein refining the first unrefined motion vector and the second unrefined motion vector is performed using decoder-side motion vector refinement (DMVR).

10. The apparatus according to claim 6, wherein refining the first unrefined motion vector and the second unrefined motion vector includes selecting the first refined motion vector and the second refined motion vector such that the sum of absolute differences substantially minimizes.

11. A video encoding method, The first unrefined motion vector and the second unrefined motion vector of the first block are refined to generate the first refined motion vector and the second refined motion vector, Using the first refined motion vector and the second refined motion vector, a prediction of the first block is generated, The boundary filter strength of the first block is determined based at least partially on the first unrefined motion vector and the second unrefined motion vector, Includes, The first unrefined motion vector and the second unrefined motion vector are not used to generate the prediction in the method.

12. The method according to claim 11, further comprising applying a filter to at least one boundary of the first block using the boundary filter strength determined above.

13. The method according to claim 11, further comprising predicting motion information of a second block using at least one of the first refined motion vector and the second refined motion vector, wherein the second block and the first block are juxtaposed blocks of different pictures.

14. The method according to claim 11, wherein refining the first unrefined motion vector and the second unrefined motion vector is performed using decoder-side motion vector refinement (DMVR).

15. The method according to claim 11, wherein refining the first unrefined motion vector and the second unrefined motion vector includes selecting the first refined motion vector and the second refined motion vector such that the sum of absolute differences substantially minimizes.

16. A video encoding device comprising one or more processors, The one or more processors described above are The first unrefined motion vector and the second unrefined motion vector of the first block are refined to generate the first refined motion vector and the second refined motion vector, Using the first refined motion vector and the second refined motion vector, a prediction of the first block is generated, The boundary filter strength of the first block is determined based at least partially on the first unrefined motion vector and the second unrefined motion vector, Configured to perform at least the following: The apparatus is not used to generate the prediction, the first unrefined motion vector and the second unrefined motion vector.

17. The apparatus according to claim 16, further configured to apply a filter to at least one boundary of the first block using the determined boundary filter strength.

18. The apparatus according to claim 16, further configured to predict motion information of a second block using at least one of the first refined motion vector and the second refined motion vector, wherein the second block and the first block are juxtaposed blocks of different pictures.

19. The apparatus according to claim 16, wherein refining the first unrefined motion vector and the second unrefined motion vector is performed using decoder-side motion vector refinement (DMVR).

20. The apparatus according to claim 16, wherein refining the first unrefined motion vector and the second unrefined motion vector includes selecting the first refined motion vector and the second refined motion vector such that the sum of absolute differences substantially minimizes.