Block Boundary Prediction Refinement Using Optical Flow
Sub-block/block refinement techniques using optical flow enhance block boundary predictions, addressing inefficiencies in video coding systems by improving compression and transmission accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2026-03-04
AI Technical Summary
Existing video coding systems face challenges in accurately predicting block boundaries, leading to inefficiencies in compression and transmission of digital video signals.
Implementing sub-block/block refinement techniques, including block boundary prediction refinement with optical flow (BBPROF), which utilizes motion vectors and spatial gradients to refine predictions for pixel values, enhancing decoding accuracy.
Improves the precision of block boundary predictions, resulting in more efficient compression and transmission of digital video signals.
Smart Images

Figure 0007824352000020 
Figure 0007824352000021 
Figure 0007824352000022
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 62 / 856,519, entitled "Block Boundary Prediction Refinement with Optical Flow," filed June 3, 2019, the entire contents of which are incorporated by reference as if fully set forth herein. [Background technology]
[0002] Video coding systems may be used to compress digital video signals, e.g., to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems may include block-based, wavelet-based, and / or object-based systems. Block-based hybrid video coding systems may be deployed. Summary of the Invention
[0003] Systems, methods, and means for sub-block / block refinement, including sub-block / block boundary refinement such as block boundary prediction refinement with optical flow (BBPROF), are disclosed. A block including a current sub-block may be decoded based on a sample value obtained for a first pixel, which may be obtained based on, for example, a motion vector (MV) for the current sub-block, MVs for sub-blocks adjacent to the current sub-block, and sample values for a second pixel adjacent to the first pixel. The sub-block / block refinement may be applied in a decoder-side motion vector refinement (DMVR) mode, a sub-block-based temporal motion vector prediction (SbTMVP) mode, and / or an affine mode. BBPROF may include, for example, sub-block-based motion compensation to generate the sub-block-based prediction. A spatial gradient of the sub-block-based prediction may be calculated at one or more pixel / sample locations. An MV difference may be calculated between the current sub-block and one or more adjacent sub-blocks. The MV difference may be used to calculate a motion vector offset at one or more pixel / sample locations. An intensity change per pixel in the current sub-block may be calculated based on optical flow. A sample value offset may be used to indicate the intensity change per pixel. For example, by adding the calculated intensity change to the sub-block prediction, the prediction (e.g., motion compensated prediction) for the pixel or sample location may be refined.
[0004] In an example, a method for performing sub-block / block refinement may be implemented. The method may be implemented, for example, by an apparatus, which may include one or more processors configured to execute computer-executable instructions that, when executed by one or more processors, perform the method, may be stored on a computer-readable medium or a computer program product. Thus, the apparatus may include one or more processors configured to perform the method. The computer-readable medium or computer program product may include instructions that cause one or more processors to perform the method by executing the instructions. The computer-readable medium may include data content generated according to the method. The signal may include a residual generated based on an original image block and a block predicted using obtained sample values for a first pixel according to the method. The apparatus may include an accessing unit and a transmitter configured to perform a second method, which includes accessing data including a residual generated based on obtained sample values for a first pixel by an apparatus including one or more processors configured to implement the method (e.g., by executing instructions) and transmitting the data including the residual. A device, such as a television, cellular phone, tablet, or set-top box, may comprise an apparatus having one or more processors configured to implement the method (e.g., by executing instructions) and at least one of (i) an antenna configured to receive a signal, the signal including data representing an image, (ii) a band limiter configured to limit the received signal to a band of frequencies including the data representing the image, or (iii) a display configured to display the image.
[0005] For example, a method for performing subblock / block refinement may include obtaining a sample value for a first pixel, e.g., based on an MV for a current subblock, an MV for a subblock adjacent to the current subblock, and a sample value for a second pixel adjacent to the first pixel, and decoding a block including the current subblock based on the obtained sample value for the first pixel.
[0006] For example, a method of encoding a block including a current sub-block based on a obtained sample value for a first pixel may include obtaining a sample value for the first pixel based on, for example, an MV for the current sub-block, an MV for a sub-block adjacent to the current sub-block, and a sample value for a second pixel adjacent to the first pixel, and encoding the block including the current sub-block based on the obtained sample value for the first pixel.
[0007] The block may include, for example, a first pixel, a second pixel, and a third pixel adjacent to the first pixel. Obtaining a sample value for the first pixel may include, for example, determining that the first pixel is adjacent to a boundary of the current sub-block, and determining (i) a difference between a motion vector (MV) for the current sub-block and a motion vector (MV) for a sub-block adjacent to the current sub-block, (ii) a gradient for the first pixel based on the sample value for the second pixel and the sample value for the third pixel, and (iii) a sample value offset based on the determined gradient and a difference between the motion vector (MV) for the current sub-block and a motion vector (MV) for a sub-block adjacent to the current sub-block, and determining a sample value for the first pixel based on the determined sample value offset.
[0008] For example, a gradient may be determined based on the sample values for at least the second pixel, and the sample values for the first pixel may be obtained using this gradient.
[0009] For example, a gradient for an optical flow model may be determined based on the sample values for at least the second pixel, and the gradient may be used in the optical flow model to obtain the sample values for the first pixel.
[0010] The difference between the MV for the current sub-block and the MV for the sub-block adjacent to the current sub-block may be used to obtain the sample value for the first pixel.
[0011] The sub-block adjacent to the current sub-block may be a first sub-block. This block may include the first sub-block and a second sub-block adjacent to the current sub-block. The sample value for the first pixel may be obtained based (e.g., further) on the MV for the second sub-block.
[0012] A sample value for the first pixel may be obtained based on a determination that the first pixel is adjacent to a boundary of the current sub-block.
[0013] The first pixel and the second pixel may be within the current sub-block.
[0014] The weighting factor may be used to obtain a sample value for the first pixel, and the weighting factor may vary according to the distance of the first pixel from the corresponding boundary of the current sub-block.
[0015] A sample value offset for the first pixel may be determined based on, for example, an MV for the current sub-block, an MV for a sub-block adjacent to the current sub-block, and a sample value for a second pixel adjacent to the first pixel. A sample value for the first pixel may be obtained using the determined sample value offset and a sample value for the predicted first pixel.
[0016] The sample value for the first pixel may be obtained, for example, based on a determination that the first pixel is adjacent to a boundary of the current sub-block. The boundary of the current sub-block may include a common boundary between the current sub-block and an adjacent sub-block.
[0017] The first pixel may be located, for example, within four rows of pixels from the top boundary of the current sub-block, within four rows of pixels from the bottom boundary of the current sub-block, within four columns of pixels from the left boundary of the current sub-block, or within four columns of pixels from the right boundary of the current sub-block.
[0018] Each feature disclosed anywhere in this specification is described and can be implemented separately / individually and in any combination with any other feature disclosed herein and / or with any feature disclosed anywhere that may be implicitly or explicitly referenced herein or that may otherwise fall within the scope of the subject matter disclosed herein. [Brief explanation of the drawings]
[0019] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1C] 1B is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1D] FIG. 1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 2] FIG. 1 illustrates an exemplary video encoder. [Figure 3] FIG. 1 illustrates an example of a video decoder. [Figure 4] FIG. 1 illustrates an example of a system in which various aspects and examples may be implemented. [Figure 5] FIG. 1 is a diagram of an exemplary four-parameter affine mode model and sub-block level motion derivation for affine blocks. [Figure 6] 1 is a diagram of an exemplary six-parameter affine mode, where V0, V1, and V2 are control points, and (MVx, MVy) is the motion vector of a sub-block centered at position (x, y). [Figure 7] FIG. 1 is a diagram of an exemplary decode-side motion vector (MV) refinement. [Figure 8A] 1 is a diagram of an example of spatially adjacent blocks that can be used by sub-block-based temporal motion vector prediction (SbTMVP). [Figure 8B] 1 is a diagram of an exemplary derivation of a sub-coding unit (CU) motion field. [Figure 9] 1 is a diagram of an example sub-block to which overlapped block motion compensation (OBMC) is applied. [Figure 10] FIG. 10 illustrates an exemplary sub-block MV (VSB) and pixel Δv(i,j). [Figure 11] FIG. 2 is a diagram of an example of a method for sub-block / block refinement according to one or more of equations (1)-(25). [Figure 12] FIG. 10 is a diagram of an exemplary MV difference calculation from selected adjacent sub-blocks. DETAILED DESCRIPTION OF THE INVENTION
[0020] A detailed description of exemplary embodiments will now be described with reference to various figures. While this description provides detailed examples of possible implementations, it should be noted that these details are intended to be illustrative and in no way limit the scope of the present application.
[0021] 1A illustrates an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, and broadcast, to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may use one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), etc.
[0022] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which may be referred to as a “station” and / or “STA,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain contexts), consumer electronic devices, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.
[0023] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node-B, an eNode-B, a Home Node-B, a Home eNode-B, a gNB, an NR Node-B, a site controller, an access point (AP), a wireless router, etc. While the base stations 114a, 114b are each shown as a single element, it should be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0024] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, sometimes referred to as a cell (not shown). These frequencies may be in a licensed spectrum, an unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for wireless service in a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, the base station 114a may use multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.
[0025] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communications link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0026] More specifically, as noted above, the communication system 100 may be a multiple-access system and may use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114a and the WTRUs 102a, 102b, 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA+ (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0027] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0028] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.
[0029] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement both LTE and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0030] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.
[0031] 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a localized area, such as a workplace, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 1A, the base station 114b may have a direct connection to the Internet 110. Therefore, the base station 114b may not be required to access the Internet 110 via the CN 106 / 115.
[0032] The RAN 104 / 113 may be in communication with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have varying Quality of Service (QoS) requirements, such as different throughput, latency, error resilience, reliability, data throughput, mobility, etc. The CN 106 / 115 may provide call control, charging services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform higher-level security functions such as user authentication. Although not shown in FIG. 1A , it should be understood that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs that use the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also be in communication with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0033] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may use the same RAT as the RAN 104 / 113 or a different RAT.
[0034] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with both the base station 114a, which may use a cellular-based radio technology, and the base station 114b, which may use an IEEE 802 radio technology.
[0035] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It should be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0036] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it should be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0037] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It should be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0038] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) to transmit and receive wireless signals over the air interface 116.
[0039] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.
[0040] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Further, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).
[0041] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control the power to other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0042] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of when signals are received from two or more nearby base stations. It should be understood that the WTRU 102 may acquire location information by way of any suitable location-determination method while remaining consistent with an embodiment.
[0043] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0044] The WTRU 102 may include a full-duplex radio (e.g., where transmission and reception of some or all of the signals associated with a particular subframe for both the UL (e.g., for transmission) and downlink (e.g., for reception) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference through signal processing in hardware (e.g., a choke) or via a processor (e.g., via a separate processor (not shown) or via processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe for the UL (e.g., for transmission) or downlink (e.g., for reception)).
[0045] 1C is a system diagram illustrating the RAN 104 and the CN 106 according to one embodiment. As mentioned above, the RAN 104 may use E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 may also be in communication with the CN 106.
[0046] The RAN 104 may include eNode-Bs 160a, 160b, and 160c, although it should be understood that the RAN 104 may include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In an embodiment, the eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Thus, for example, the eNode-B 160a may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.
[0047] Each of the eNode-Bs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 1C, the eNode-Bs 160a, 160b, 160c may communicate with one another via an X2 interface.
[0048] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the foregoing elements is shown as part of the CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0049] The MME 162 may be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that use other radio technologies such as GSM and / or WCDMA.
[0050] The SGW 164 may be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available to the WTRUs 102a, 102b, 102c, managing and storing the context of the WTRUs 102a, 102b, 102c, etc.
[0051] The SGW 164 may be connected to a PGW 166 that may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0052] The CN 106 may facilitate communication with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communication between the WTRUs 102a, 102b, 102c and traditional land-line communication devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 106 and the PSTN 108. Additionally, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0053] Although the WTRU is shown in FIGS. 1A-1D as a wireless terminal, it is contemplated that in some representative embodiments such a terminal may use a wired communication interface (e.g., temporarily or permanently) with the communication network.
[0054] In an exemplary embodiment, the other network 112 may be a WLAN.
[0055] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that carries traffic into and / or from the BSS. Traffic to a STA originating from outside the BSS may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP to be delivered to the respective destination. Traffic between STAs within the BSS may be sent through the AP; for example, a source STA may send traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between (e.g., directly between) a source STA and a destination STA using direct link setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode is sometimes referred to herein as an "ad hoc" communication mode.
[0056] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In some representative embodiments, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. With CSMA / CA, STAs (e.g., every STA), including the AP, may sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, that particular STA may back off. One STA (e.g., only one station) may transmit at any given time within a given BSS.
[0057] High-throughput (HT) STAs may use 40 MHz wide channels for communication, for example, via a combination of a primary 20 MHz channel and adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.
[0058] A very high throughput (VHT) STA may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining eight contiguous 20 MHz channels or by combining two non-contiguous 80 MHz channels, sometimes referred to as an 80+80 configuration. For the 80+80 configuration, the data may be channel encoded and then passed through a segment parser, which may segment the data into two streams. Inverse fast Fourier transform (IFFT) processing and time-domain processing may be performed separately for each stream. The streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration may be reversed, and the combined data may be sent to the medium access control (MAC).
[0059] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah may support meter type control / machine-type communication (MTC) devices within macro coverage areas. MTC devices may have limited functionality, including some functionality, for example, support for some and / or limited bandwidths (e.g., only support for some bandwidths). MTC devices may include batteries with a higher-than-threshold battery life (e.g., to maintain very long battery life).
[0060] WLAN systems that may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In an 802.11ah example, the primary channel may be 1 MHz wide for a STA (e.g., an MTC-type device) that supports (e.g., only supports) the 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. If the primary channel is busy, for example, because a STA (that only supports 1 MHz mode of operation) is transmitting to the AP, the entire available frequency band may be considered busy, even though most of the frequency band may remain idle and available for use.
[0061] In the United States, the available frequency band that can be used by 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is 6MHz to 26MHz depending on the country code.
[0062] 1D is a system diagram illustrating the RAN 113 and the CN 115 according to one embodiment. As mentioned above, the RAN 113 may use NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 may also be in communication with the CN 115.
[0063] The RAN 113 may include gNBs 180a, 180b, and 180c, although it should be understood that the RAN 113 may include any number of gNBs while remaining consistent with one embodiment. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a, 180b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, and 180c. Thus, for example, the gNB 180a may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a. In one embodiment, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on an unlicensed spectrum, while the remaining component carriers may be on a licensed spectrum. In one embodiment, the gNBs 180a, 180b, 180c may implement coordinated multipoint (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNB 180a and the gNB 180b (and / or gNB 180c).
[0064] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., including varying numbers of OFDM symbols and / or lasting for varying lengths of absolute time).
[0065] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing any other RANs (e.g., eNode-Bs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect with a gNB 180a, 180b, 180c while also communicating / connecting with another RAN, such as an eNode-B 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement the DC principle and communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNode-Bs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput to serve the WTRUs 102a, 102b, 102c.
[0066] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 184a, 184b, routing of control plane information towards access and mobility management functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.
[0067] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements is shown as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0068] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize CN support for the WTRUs 102a, 102b, 102c based on the type of service utilized by the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases, such as services relying on Ultra-Reliable Low-Latency (URLLC) access, services relying on Enhanced Multimedia Broadcasting (eMBB) access, services for Machine Type Communications (MTC) access, and / or the like. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that use other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0069] The SMFs 183a, 183b may be connected to the AMFs 182a, 182b in the CN 115 via an N11 interface. The SMFs 183a, 183b may also be connected to the UPFs 184a, 184b in the CN 115 via an N4 interface. The SMFs 183a, 183b may select and control the UPFs 184a, 184b and configure the routing of traffic through the UPFs 184a, 184b. The SMFs 183a, 183b may perform other functions such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.
[0070] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.
[0071] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. Additionally, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to local data networks (DNs) 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.
[0072] 1A-1D and the corresponding description thereof, one or more or all of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, eNode-Bs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or simulate network and / or WTRU functionality.
[0073] The emulation device may be designed to implement one or more tests of other devices in a lab environment and / or in an operator network environment. For example, one or more emulation devices may perform one or more, or all, functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communications network to test other devices within the communications network. One or more emulation devices may perform one or more, or all, functions while temporarily implemented / deployed as part of a wired and / or wireless communications network. The emulation device may be directly coupled to another device for testing and / or may perform testing using over-the-air wireless communications.
[0074] The one or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communications network. For example, the emulation devices may be utilized in test labs and / or test scenarios within non-deployed (e.g., test) wired and / or wireless communications networks to implement testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may, for example, include one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0075] This application describes various aspects, including tools, features, examples or embodiments, models, techniques, and the like. Many of these aspects have been described in detail, and often in a manner that may seem limiting, at least to illustrate their individual characteristics. However, this is for clarity of explanation and does not limit the scope of the application or these aspects. In fact, all of the different aspects may be combined and interchanged to provide further aspects. Furthermore, these aspects may also be combined and interchanged with aspects described in prior applications.
[0076] The aspects described and contemplated in this application may be implemented in many different forms. While Figures 5-12 described herein may provide some embodiments, other embodiments are contemplated. A discussion of Figures 5-12 does not limit the breadth of their implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects may be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.
[0077] In this application, the terms "reconstruct" and "decode" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and "image," "picture," and "frame" may be used interchangeably.
[0078] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc. may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, “first decode” and “second decode.” The use of such terms does not imply an order to modified operations unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but may occur before, during, or in a period overlapping with the second decode.
[0079] Various methods and other aspects described herein may be used to modify modules of video encoder 200 and decoder 300, e.g., intra-prediction and entropy coding and / or decoding modules (260, 360, 245, 330), as shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein presents aspects not limited to VVC or HEVC and may apply to any type, format, or version of video coding, e.g., whether described in a standard or recommendation, whether existing or developed in the future, and an extension of any such standard and recommendation (including, e.g., VVC and HEVC). Unless otherwise indicated or technically excluded, aspects described herein may be used individually or in combination.
[0080] In the examples described herein, various numerical values are used, such as weighting factors such as {1 / 4, 1 / 8, 1 / 16, 1 / 32} or {3 / 4, 7 / 8, 15 / 16, 31 / 32}, filters such as a 3-tap filter [-1, 0, 1], etc. These and other specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0081] 2 illustrates an exemplary video encoder. While variations of the exemplary encoder 200 are contemplated, the encoder 200 is described below for clarity without describing all possible variations.
[0082] Before being encoded, the video sequence may undergo pre-encoding processing (201), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to make the signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with this pre-processing and attached to the bitstream.
[0083] In the encoder 200, a picture is encoded by the encoder elements as follows. The picture to be encoded is partitioned (202) and processed, for example, in units of coding units (CUs). Each unit is encoded, for example, using intra mode or inter mode. When a unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which one of intra mode or inter mode to use to encode the unit, and indicates the intra / inter decision, for example, by a prediction mode flag. For example, a prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0084] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.
[0085] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. Combining the decoded prediction residual with the predicted block (255) reconstructs an image block. An in-loop filter (265) is applied to the reconstructed picture, for example, to implement a deblocking / sample adaptive offset (SAO) filter to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).
[0086] 3 illustrates an example of a video decoder. In the exemplary decoder 300, a bitstream is decoded by decoder elements as follows. The video decoder 300 generally performs a reverse decoding path relative to the encoding path described in FIG. 2. The encoder 200 also generally performs video decoding as part of encoding the video data. For example, the encoder 200 may perform one or more of the video decoding steps presented herein. The encoder reconstructs the decoded image to maintain synchronization with the decoder, e.g., with respect to one or more of reference pictures, entropy coding contexts, and other decoder-related state variables.
[0087] In particular, the decoder's input includes a video bitstream, such as may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is divided. Thus, the decoder partitions (335) the picture according to the decoded picture partition information. The transform coefficients are inverse quantized (340) and inverse transformed (350) to decode the prediction residual. Combining (355) the decoded prediction residual with the predicted block reconstructs an image block. The prediction block may result from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0088] The decoded picture may further undergo post-decoding processing (385), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding process can use metadata derived in the pre-encoding process and signaled in the bitstream.
[0089] FIG. 4 illustrates an example of a system in which various aspects and embodiments described herein may be implemented. System 400 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 400, alone or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various embodiments, system 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 400 is configured to implement one or more of the aspects described herein.
[0090] The system 400 includes at least one processor 410 configured to execute loaded instructions to implement various aspects described herein, for example. The processor 410 may include embedded memory, input / output interfaces, and various other circuits, as known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a nonvolatile memory device). The system 400 includes a storage device 440, which may include nonvolatile and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. The storage device 440 may include, by way of non-limiting example, an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.
[0091] System 400 includes an encoder / decoder module 430 configured to process data to provide, for example, encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 430 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software, as is known to those skilled in the art.
[0092] Program code to be loaded onto the processor 410 or the encoder / decoder 430 to implement various aspects described herein may be stored in the storage device 440 and then loaded onto the memory 420 for execution by the processor 410. According to various embodiments, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 may store one or more of a variety of items during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0093] In some embodiments, memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be the memory 420 and / or the storage device 440, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used, for example, to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations such as, for example, MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by JVET, the Joint Video Experts Team).
[0094] Input to the elements of system 400 may be provided through various input devices, shown in block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals transmitted wirelessly, for example by a broadcaster, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples not shown in FIG. 4 include composite video.
[0095] In various embodiments, the input devices of block 445 have associated respective input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select a signal frequency band, which in some embodiments may be referred to as a channel (for example), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs a variety of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and re-filtering RF signals transmitted over a wired (e.g., cable) medium to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0096] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 400 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 410, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within processor 410, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including processor 410 and encoder / decoder 430, which operate in combination with memory and storage elements to process the data stream as desired for presentation on an output device.
[0097] The various elements of system 400 may be provided within a unified housing, where the various elements may be interconnected and transmit data between them using a suitable connection arrangement 425, for example, an internal bus known in the art, including an Inter-IC (I2C) bus, wiring, and printed circuit boards.
[0098] System 400 includes a communication interface 450 that enables communication with other devices over a communication channel 460. Communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 460. Communication interface 450 may include, but is not limited to, a modem or a network card, and communication channel 460 may be implemented, for example, in a wired and / or wireless medium.
[0099] In various embodiments, data is streamed or otherwise provided to system 400 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronic Engineers). The Wi-Fi signal in these examples is received via communication channel 460 and communication interface 450 adapted for Wi-Fi communication. Communication channel 460 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 400 using a set-top box that delivers data via an HDMI connection in input block 445. Still other embodiments provide streamed data to system 400 using an RF connection in input block 445. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.
[0100] System 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 475 may be integrated with other components (e.g., as in the case of a smartphone) or may be separate (e.g., an external monitor for a laptop). The other peripheral devices 495, in various example embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0101] In various embodiments, control signals are communicated between system 400 and display 475, speaker 485, or other peripheral device 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 using communication channel 460 via communication interface 450. Display 475 and speaker 485 may be integrated with other components of system 400 in a single unit, for example, within an electronic device such as a television. In various embodiments, display interface 470 includes a display driver, for example, a timing controller (T Con) chip.
[0102] Display 475 and speakers 485 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various embodiments in which display 475 and speakers 485 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0103] The embodiments may be implemented by computer software implemented by the processor 410 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 420 may be of any type appropriate to the technology environment, and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 may be of any type appropriate to the technology environment, and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0104] Various implementations include decoding. As used herein, "decoding" can encompass all or part of the processes performed on a received encoded sequence to generate, for example, a final output suitable for display. In various embodiments, such processes include one or more of processes typically performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of various implementations described herein, e.g., decoding a block including a current sub-block based on sample values obtained for a first pixel, where the sample values may be obtained based, for example, on a motion vector (MV) for the current sub-block, MVs for sub-blocks adjacent to the current sub-block, and sample values for a second pixel adjacent to the first pixel, etc.
[0105] As a further embodiment, in one example, "decoding" refers to entropy decoding only, in another example, "decoding" refers to differential decoding only, and in another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding" is intended to refer specifically to a subset of operations or to the broad decoding process generally will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.
[0106] Various implementations include encoding. Similar to the above discussion of “decoding,” “encoding,” as used herein, can encompass, for example, all or part of the processes performed on an input video sequence to generate an encoded bitstream. In various embodiments, such processes include one or more of processes typically performed by an encoder, e.g., partitioning, differential encoding, transforming, quantization, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by encoders of various implementations described herein, e.g., encoding a block including a current sub-block based on sample values obtained for a first pixel, where the sample values may be obtained based, for example, on a motion vector (MV) for the current sub-block, MVs for sub-blocks adjacent to the current sub-block, and sample values for a second pixel adjacent to the first pixel, etc.
[0107] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding" is intended to refer specifically to a subset of operations or to the broad encoding process generally will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.
[0108] It should be noted that the syntax elements used herein are descriptive terms, and therefore they do not preclude the use of other syntax element names.
[0109] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.
[0110] Given the constraints of computational complexity, a balance or tradeoff between rate and distortion is usually considered during the encoding process. Rate-distortion optimization is usually formulated as minimizing a rate-distortion function, which is a weighted sum of the rate and the distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these approaches may be based on extensive testing of all encoding options, including all considered modes or coding parameter values, with a thorough evaluation of their coding costs and associated distortions of the reconstructed signal after encoding and decoding. To reduce encoding complexity, faster approaches may also be used, particularly with the calculation of approximated distortions based on predicted or predicted residual signals rather than reconstructed ones. A mixture of these two approaches may also be used, such as by using approximated distortions for only some of the possible encoding options and full distortions for others. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches use any of a variety of techniques to perform optimization, but the optimization does not necessarily involve a thorough evaluation of both the coding costs and associated distortions.
[0111] The implementations and aspects described herein may be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even when discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed feature may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. For example, the methods may be implemented in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants (PDAs), and other devices that facilitate communication of information between end users.
[0112] References to "one embodiment," "an embodiment," "one example," "one implementation," or "an implementation," as well as other variations, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment," "in an embodiment," "in one example," "in one implementation," or "in an implementation," as well as any other variations appearing in various places throughout this application, are not necessarily all referring to the same embodiment or example.
[0113] Additionally, the application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0114] Additionally, the application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0115] Additionally, the application may refer to "receiving" various pieces of information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" is typically included in various respects, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0116] For example, it should be understood that the use of any of " / ," "and / or," and "at least one of" in the cases of "A / B," "A and / or B," and "at least one of A and B" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A and B and C). This may be expanded to as many items as are listed, as would be apparent to one of ordinary skill in this and related arts.
[0117] Also, as used herein, the term "signaling" refers to, among other things, indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals (e.g., to a decoder) a weight index or the like. In this way, in one embodiment, the same parameters are used on both the encoder and decoder sides. Thus, for example, an encoder can send a particular parameter to a decoder (explicit signaling), so that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter among others, signaling can be used without sending it (implicit signaling) so that the decoder can simply know and select the particular parameter. By avoiding sending any actual function, bit savings are achieved in various embodiments. It should be understood that signaling can be done in various ways. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. While the foregoing relates to the verb form of the word "signaling," the word "signal" can also be used as a noun herein.
[0118] As will be apparent to one skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0119] Bidirectional motion compensation prediction (MCP) may be implemented. MCP may provide high efficiency in removing temporal redundancy, for example, by exploiting temporal correlation between pictures. For example, a bidirectional prediction signal may be generated by combining two unidirectional prediction signals (e.g., using a weight value equal to 0.5). Combining unidirectional prediction signals may be suboptimal, for example, when illumination changes rapidly from one reference picture to another. Prediction techniques may compensate for illumination variations over time, for example, by applying global or local weights and / or offset values to one or more (e.g., respectively) of the sample values in the reference pictures.
[0120] Coding modules (eg, associated with temporal prediction) may be extended and / or enhanced. Affine motion compensation may be used as an inter-coding tool.
[0121] An implementation using affine mode may be described herein. A translational motion model may be applied for motion-compensated prediction. There may be many types of motion (e.g., zoom-in or zoom-out, rotation, perspective motion, and / or other irregular motion). Simplified affine transform motion-compensated prediction may be applied. A flag for an inter-coded CU (e.g., each inter-coded CU) may be signaled to indicate, for example, whether a translational motion or affine motion model is applied for inter prediction. A flag may be signaled (e.g., if affine motion is used) to indicate the number of parameters (e.g., 4 or 6) used in the affine motion model.
[0122] The affine motion model may be a four-parameter model. Two parameters may be used for translation (e.g., one each in the horizontal and vertical directions). One parameter may be used for zoom motion. One parameter may be used for rotation motion. The horizontal zoom parameter may be equal to the vertical zoom parameter. The horizontal rotation parameter may be equal to the vertical rotation parameter. The four-parameter motion model may be coded using two motion vectors (MVs) as a pair (e.g., 1) at two control point locations defined at the top-left and top-right corners of the current CU. Figure 5 is a diagram of an exemplary four-parameter affine motion model and sub-block level motion derivation for an affine block. As shown in Figure 5, the affine motion field of a block may be described by two control point motion vectors (V0, V1). Based on the motion of the control points, the motion field (v x ,v y ) can be described, for example, according to Eq.
[0123]
number
[0124] Here, as shown in Figure 5, (v 0x ,v 0y ) can be the motion vector of the control point in the upper left corner, and (v 1x ,v 1y ) may be the motion vector of the control point in the upper right corner, and w may be the width of the CU.
[0125] The affine motion model may be a six-parameter model. Two parameters may be used for translational movement (e.g., one each for the horizontal and vertical directions). Two parameters may be used for zoom movement (e.g., one each for the horizontal and vertical directions). Two parameters may be used for rotational movement (e.g., one each for the horizontal and vertical directions). The six-parameter motion model may be coded with three MVs at three control points. Figure 6 shows an example six-parameter affine model, where V0, V1, and V2 are the control points and (MV x ,MV y ) is the motion vector of the sub-block centered at position (x,y). As shown in FIG. 6, control points for a 6-parameter affine coded CU may be defined at the top-left corner, top-right corner, and bottom-left corner of the CU. The motion at the top-left control point may relate to translational motion. The motion at the top-right control point may relate to rotation and zoom motion in the horizontal direction. The motion at the bottom-left control point may relate to rotation and zoom motion in the vertical direction. The rotation and zoom motion in the horizontal direction may be different from the motion in the vertical direction. The MV(v x ,v y ) may be derived using three MVs at the control points, for example, according to Equations 2 and 3.
[0126]
number
[0127] where (v 2x ,v 2y ) may be the motion vector of the bottom-left control point, (x, y) may be the center position of the sub-block, and w and h may be the width and height of the CU, respectively.
[0128] A motion field for a block coded with an affine motion model may be derived, for example, based on the granularity of a sub-block. The MV of (e.g., each) sub-block may be derived, for example, by calculating the MV of the central sample of the sub-block (e.g., according to Equation (1)) (e.g., as shown in FIG. 5). The calculation may be rounded, for example, to 1 / 16 pel accuracy. The derived MV may be used in the motion compensation stage to generate a prediction signal for a sub-block (e.g., each sub-block) inside the current block. The sub-block size applied to affine motion compensation may be, for example, 4×4. The four parameters of the four-parameter affine model may be estimated iteratively, for example. For example, one or more MV pairs in step k may be
[0129]
number
[0130] The original luminance signal can be denoted as I(i,j). The predicted luminance signal can be denoted as I' k (i,j). The spatial gradient g x (i,j) and g y (i,j) are, for example, the predicted signal I' in the horizontal and vertical directions, respectively. k The derivative of Equation (1) can be expressed, for example, according to Equation 4:
[0131]
number
[0132] Here, in step k, (a, b) may be delta translation parameters, and (c, d) may be delta zoom and rotation parameters. The delta MVs at the control points may be derived in coordinates according to, for example, Equations 5 and 6. For example, (0, 0) and (w, 0) may be the coordinates for the top-left and top-right control points, respectively.
[0133]
number
[0134] The relationship between the change in luminance and the spatial gradient and temporal shift may be formulated, for example, according to Equation 7.
[0135]
number
[0136] where:
[0137]
number
[0138] and
[0139]
number
[0140] may be substituted for values in equation (4) to obtain an equation for parameters (a, b, c, d), for example, as shown in equation 8. I' k (i,j)-I(i,j)=(g x (i,j)*i+g y (i,j)*j)*c+(-g x (i,j)*j+g y (i,j)*i)*d+g x (i,j)*a+g y (i,j)*b (8) The parameter set (a, b, c, d) can be derived, for example, using the least squares method (e.g., so that the samples in the CU satisfy Equation 8). MV at the control point at step (k+1)
[0141]
number
[0142] may be solved with Equation 5 and Equation 6, which may be rounded to a particular precision (e.g., ¼ pel). The MVs at the two control points may be refined (e.g., using iterations) until the parameters (a, b, c, d) become (e.g., all) zero or the number of iterations performed reaches a (e.g., predefined) limit.
[0143] The six parameters of the six-parameter affine model may be estimated. Equation 4 may be modified, for example, according to Equation 9:
[0144]
number
[0145] where, in step k, (a, b) may be delta translation parameters, (c, d) may be delta zoom and rotation parameters for the horizontal direction, and (e, f) may be delta zoom and rotation parameters for the vertical direction. Equation 8 may be modified, for example, according to Equation 10. I' k (i,j)-I(i,j)=(g x (i,j)*i)*c+(g x (i,j)*j)*d+(g y (i,j)*i)*e+(g y (i,j)*j)*f+g x (i,j)*a+g y (i,j)*b (10) The parameter set (a, b, c, d, e, f) may be derived, for example, using a least squares method by considering a sample (e.g., multiple samples) within a CU.
[0146]
number
[0147] can be calculated using Equation 5. MV of the top right control point
[0148]
number
[0149] and the MV of the bottom left control point
[0150]
number
[0151] may be calculated, for example, according to Equations 11 and 12.
[0152]
number
[0153] Decoder-side motion vector refinement (DMVR) may be provided. Bidirectional matching (BM)-based DMVR may be applied, for example, to improve the accuracy of MV in merge mode. In bidirectional prediction operations, refined MVs may be searched around an initial MV in reference picture list L0 and / or reference picture list L1. The BM-based DMVR may calculate distortion between two candidate blocks in reference picture list L0 and list L1. Figure 7 shows exemplary decode-side motion vector (MV) refinement. As shown in Figure 7, the sum of absolute differences (SAD) between collocated blocks may be calculated, for example, based on one or more (e.g., each) MV candidates around the initial MV. The MV candidate with the lowest SAD may be the refined MV and may be used to generate a bidirectional prediction signal.
[0154] The refined MVs derived by the DMVR may be used, for example, to generate inter-predicted samples. The refined MVs derived by the DMVR may be used, for example, in temporal motion vector prediction for future picture encoding. The initial MVs may be used, for example, in deblocking and / or spatial motion vector prediction for future CU encoding to avoid MV dependency between the current CU and neighboring CUs.
[0155] 7, the search points surrounding the initial MV and MV offset may observe the MV difference mirroring (e.g., symmetry) rule. The points checked by the DMVR, indicated by the candidate MV pair (MV0, MV1), may obey Equation 13 and / or Equation 14.
[0156]
number
[0157] Music Video offset may represent a refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range may be, for example, two integer luma samples from the initial MV. A fast search method with an early termination mechanism may be applied, for example, to reduce the search complexity.
[0158] Sub-block-based temporal motion vector prediction (SbTMVP) may be provided. SbTMVP may use motion fields in a co-located picture to improve motion vector prediction and merge modes for CUs in a current picture. The same co-located picture used by temporal motion vector prediction (TMVP) may be used for SbTMVP. SbTMVP may differ from TMVP in one or more of the following aspects: TMVP may predict motion at the CU level. SbTMVP may predict motion at the sub-CU level. TMVP may fetch temporal motion vectors from a co-located block in a co-located picture. The co-located block may be a bottom-right or center block with respect to the current CU. SbTMVP may apply a motion shift, for example, before fetching temporal motion information from the co-located picture. The motion shift may be obtained, for example, from a motion vector from one of the spatially neighboring blocks of the current CU.
[0159] 8A and 8B illustrate an exemplary SbTMVP process. FIG. 8A illustrates exemplary spatially adjacent blocks that may be used in SbTMVP. FIG. 8B illustrates an exemplary derivation of a sub-coding unit (CU) motion field. As shown in FIG. 8B, a sub-CU motion field may be derived by applying a motion shift from a spatial neighbor and scaling motion information from a corresponding co-located sub-CU. SbTMVP may predict a motion vector for a sub-CU within a current CU. A spatial neighbor (e.g., A1 in FIG. 8A) may be examined. A motion vector for A1 may be selected, for example, if A1 has a motion vector that uses the co-located picture as its reference picture (e.g., for a motion shift to be applied). The motion shift may be selected as (0,0), for example, if no motion is identified. The selected motion shift may be applied, for example, to obtain sub-CU-level motion information (e.g., motion vectors and / or reference indexes, etc.) from the co-located picture. For example, the selected motion shift may be added to the coordinates of the current block. The motion shift may be set to the motion of block A1 (e.g., in the example shown by FIG. 8B). Motion information of a corresponding block (e.g., of each) sub-CU in the co-located picture (e.g., of the smallest motion grid covering the center sample) may be used to derive motion information for the sub-CU. The motion information of the co-located sub-CU (e.g., after being identified) may be converted into a motion vector and reference index of the current sub-CU. For example, temporal motion scaling may be applied to align the reference picture of the temporal motion vector to that of the current CU.
[0160] A combined sub-block-based merge list including SbTMVP candidates and affine merge candidates may be used to signal the sub-block-based merge mode. The SbTMVP mode may be enabled and / or disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP predictor may be added as the (e.g., first) entry in the list of sub-block-based merge candidates, followed by the affine merge candidates. The size of the sub-block-based merge list may be signaled in the SPS. The maximum allowed size of the sub-block-based merge list may be, for example, 5.
[0161] The sub-CU size used in SbTMVP may be fixed at, for example, 8 x 8. The SbTMVP mode may be applicable to (e.g., only) CUs with widths and heights greater than or equal to 8.
[0162] The encoding logic for the additional SbTMVP merge candidate may be the same as the encoding logic for other merge candidates. For example, an additional RD check may be performed for each CU in a P or B slice. The additional RD check may be used to determine whether to use the SbTMVP candidate.
[0163] Overlapping block motion compensation (OBMC) may be provided. OBMC may be switched on and off, for example, using syntax at the CU level. OBMC may be performed on motion compensation (MC) block boundaries (e.g., excluding the right and bottom boundaries of a CU). OBMC may be applied to luma-chroma components. An MC block may correspond to a coding block. A CU may be coded in a sub-CU mode (e.g., including sub-CU merge, affine, and FRUC modes). One or more sub-blocks (e.g., each sub-block) of a CU coded in a sub-CU mode may be an MC block. OBMC may be performed at the sub-block level for (e.g., all) MC block boundaries, for example, to process CU boundaries (e.g., uniformly). Figure 9 shows an example of sub-blocks to which OBMC is applied. The sub-block size may be set equal to 4x4, for example, as shown in Figure 9.
[0164] OBMC may be applied to the current sub-block. Motion vectors (e.g., in addition to the current motion vector) of four connected adjacent sub-blocks (e.g., if available and not identical to the current motion vector) may be used to derive a predictive block for the current sub-block. In one or more examples, "adjacent" may be used interchangeably with "adjacent." For example, multiple predictive blocks based on multiple motion vectors may be combined to generate a final predicted signal for the current sub-block.
[0165] The predicted block based on the motion vector of the adjacent sub-block is P N where N indicates the index for the neighboring sub-blocks above, below, left, and right of the current sub-block. A prediction block based on the motion vector of the current sub-block can be expressed as P C OBMC can be expressed as, for example, P N If P is based on the motion information of an adjacent sub-block that contains the same motion information as the current sub-block, it may be skipped. NOne or more (e.g., each) samples of P C can be added to the same sample at, for example, P N The four rows / columns of P C In the example, the weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are added to P N and the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C For example, if the height or width of the coding block is equal to 4, or if the CU is coded in sub-CU mode, P N For MC blocks with two small rows and / or columns of P C For small MC blocks, the weighting factors {1 / 4, 1 / 8} can be added to P N and the weighting factors {3 / 4, 7 / 8} are used for P C may be used for P, which is generated based on the motion vectors of vertically (e.g., and / or horizontally) adjacent sub-blocks. N In the case of P N Samples in the same row (e.g., and / or column) of P are weighted with the same coefficient. C For example, pixels in the overlapping area may use different weighting factors than those used for pixels in the non-overlapping area.
[0166] A CU-level flag may be signaled to indicate whether OBMC is applied for the current CU, e.g., a CU size of 256 luma samples or less. OBMC may be applied by default, e.g., for CUs with a size larger than 256 luma samples or not coded in AMVP mode. The impact of OMBC may be taken into account at the encoder, e.g., during the motion estimation stage. A predicted signal formed by OBMC using motion information of the upper and left adjacent blocks may be used to compensate for the upper and left boundaries of the original signal of the current CU. A (e.g., normal) motion estimation process may be applied (e.g., subsequently) (e.g., separately).
[0167] Prediction refinement using optical flow (PROF) may be applied to the affine mode. PROF may, for example, refine sub-block-based affine motion compensation prediction using optical flow to achieve finer granularity of motion compensation. Luma prediction samples may be refined (e.g., after sub-block-based affine motion compensation), for example, by adding a difference derived by an optical flow equation. PROF may include one or more of the following: Sub-block-based affine motion compensation may be performed to generate the sub-block prediction I(i,j); The spatial gradient of the sub-block prediction g x (i,j) and g y (i, j) may be calculated at one or more sample locations (e.g., each sample location). For example, spatial gradients may be calculated using one or more pixels that may or may not be partially or fully contiguous. In one example, a gradient for a first pixel may be based on a sample value for a second pixel and a sample value for a third pixel, where the second pixel and the third pixel are adjacent to the first pixel. In some examples, the first pixel for which the gradient is calculated may abut one or both of the second pixel or the third pixel adjacent to the first pixel. In other examples, the first pixel for which the gradient is calculated may be near but not abut one or both of the second pixel or the third pixel adjacent to the first pixel. These calculations may be performed using a 3-tap filter, such as [-1, 0, 1], as shown in Equations 15 and 16, for example. g x (i,j)=I(i+1,j)-I(i-1,j) (15) g y (i,j)=I(i,j+1)-I(i,j-1) (16) The sub-block prediction may be extended for gradient calculation (e.g., by one pixel on each side). Pixels on the extended boundary may be copied, for example, from the nearest integer pixel position in the reference picture. For example, if pixels on the extended boundary are copied from the nearest integer pixel position in the reference picture, additional interpolation for the padding region may be avoided. Figure 10 shows an example of a sub-block MV V SB and pixel Δv(i,j). The luminance prediction refinement may be calculated by an optical flow equation, for example, as shown in Equation 17. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (17) Here, Δv(i,j) is the difference between the pixel MV denoted by v(i,j) calculated for sample location (i,j) as shown in FIG. 10 and the sub-block MV of the sub-block to which pixel (i,j) belongs.
[0168] The affine model parameters and pixel location relative to the sub-block center may not change between sub-blocks. Δv(i,j) may be calculated for a (e.g., first) sub-block and reused for other sub-blocks (e.g., within the same CU). Let x and y be the horizontal and vertical offsets from the pixel location to the center of the sub-block, then Δv(x,y) may be derived, for example, according to Equation 18:
[0169]
number
[0170] where, for a four-parameter affine model, c and e can be determined according to Equation 19:
[0171]
number
[0172] where, for a six-parameter affine model, c, d, e, and f may be determined according to Equation 20:
[0173]
number
[0174] where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the top-left, top-right, and bottom-left control point motion vectors, respectively, and w and h are the width and height of the CU. Luma prediction refinement may be added to the sub-block prediction I(i,j). The final prediction I′ may be generated, for example, according to Equation 21. I'(i,j)=I(i,j)+ΔI(i,j) (21) DMVR and SbTMVP may be used in different prediction modes to improve the accuracy of the predicted MV. The refined MV after DMVR or SbTMVP may then be used (e.g., solely) to perform sub-block-based motion compensation. OBMC may include pixel-level refinement. OBMC may be used to reduce boundary discontinuities in sub-blocks of a CU or sub-CU. OBMC may include multiple motion compensation operations for one or more (e.g., each) sub-blocks. For example, the MVs of four connected adjacent sub-blocks may be used to derive a predicted block for the current sub-block if they are available and are not identical to the MV of the current sub-block.
[0175] A method for sub-block / block refinement, e.g., pixel-level refinement, may be provided. For example, the method may be used to reduce boundary discontinuities. Figure 11 provides an example of the method. The method described in Figure 11 may be applied in a decoder and / or an encoder.
[0176] FIG. 11 shows an example of a method for subblock / block refinement according to one or more of Equations (1) through (25). Examples disclosed herein, as well as other examples, may operate according to the example method 1100 shown in FIG. 11. Method 1100 includes 1102 and 1104. In 1102, a sample value for a first pixel may be obtained, for example, based on (1) a motion vector (MV) for the current subblock, (2) MVs for subblocks adjacent to the current subblock, and (3) a sample value for a second pixel adjacent to the first pixel. In 1104, a block including the current subblock may be encoded or decoded based on the obtained sample value for the first pixel. When the method described in FIG. 11 is applied to a decoder, 1104 in FIG. 11 may be performed by the decoder and may involve decoding the block including the current subblock based on the obtained sample value for the first pixel. When the method described in FIG. 11 is applied to an encoder, 1104 in FIG. 11 may be performed by the encoder, and 1104 may involve encoding a block including the current sub-block based on the sample value for the obtained first pixel.
[0177] An example method for sub-block / block refinement, e.g., for encoding and decoding, is provided. The example may refer to a "boundary," which includes different types of boundaries, such as a boundary of a block, a sub-block, a CU, and / or a PU. The example may refer to an "adjacent," which may include different types of adjacencies, such as spatial adjacencies and temporal adjacencies of a block, a sub-block, a CU, and / or a PU. The example may refer to an "adjacent," which includes different types of adjacencies, such as adjacent blocks, adjacent sub-blocks, adjacent pixels, and / or pixels adjacent to a boundary. Spatial adjacencies may be adjacent within the same frame, while temporal adjacencies may be at the same location in adjacent frames. For example, adjacent sub-blocks are sub-blocks that may be spatially or temporally adjacent. Boundary pixels are pixels adjacent to a boundary, and the boundary may be any type of boundary. For example, boundary pixels may be adjacent to a boundary of a block, a sub-block, a CU, and / or a PU.
[0178] Sub-block / block refinement may include sub-block / block boundary refinement. For example, MV differences between a current block and / or sub-block and an adjacent block and / or sub-block may be calculated and converted into sample value differences derived by an optical flow equation. Pixel intensities (e.g., luma and / or chroma) of boundary pixels of the current block and / or sub-block may be refined, for example, by adding the derived difference values. The derived difference values may be referred to as block boundary prediction refinement using optical flow (BBPROF). Sample value offsets for the boundary pixels may indicate the derived difference values. The sample values for the boundary pixels may indicate the pixel intensities of the boundary pixels. BBPROF (e.g., as described herein) may provide pixel-level granularity for sub-block and block boundary refinement. BBPROF (e.g., as described herein) may be applied to any sub-block-based inter prediction mode and / or CU-based inter prediction mode.
[0179] Boundary pixels may include pixels at the boundaries of blocks and / or sub-blocks. For example, a rectangular sub-block may have four boundaries, including left, right, top, and bottom boundaries. A boundary may include a common boundary shared between two sub-blocks. These two sub-blocks may abut each other at the boundary. A pixel may be located at a boundary when it is located near the boundary. For example, a pixel may be located at a boundary when it is a number of rows (e.g., 4, 3, 2, or 1) of pixels from the top boundary of the sub-block, a number of rows (e.g., 4, 3, 2, or 1) of pixels from the bottom boundary of the sub-block, a number of columns (e.g., 4, 3, 2, or 1) of pixels from the left boundary of the sub-block, or a number of columns (e.g., 4, 3, 2, or 1) of pixels from the right boundary of the sub-block. In some examples, a first pixel may be at or located at the common boundary of a first sub-block when the first pixel is located inside the first sub-block and abuts a pixel inside a second sub-block that shares a common boundary with the first sub-block. In some examples, a pixel may be located on a boundary of a sub-block, but outside the sub-block.
[0180] BBPROF may be applied, for example, in DMVR mode. BBPROF may reduce block boundary discontinuities of DMVR-based sub-block level motion compensation prediction. Pixel intensity changes may be applied by BBPROF. The pixel intensity changes may be derived, for example, from an optical flow equation. A sample value offset for a pixel may indicate the pixel intensity change. BBPROF may be used to perform one (e.g., only one) motion compensation operation per sub-block. Motion compensation in DMVR mode may perform one motion compensation operation per sub-block.
[0181] The refined motion vector for a sub-block within a CU may be derived, for example, by performing DMVR (e.g., as described herein). Sub-block-based motion compensation (e.g., as described herein) may be performed to generate the sub-block-based prediction.
[0182] Spatial gradient of subblock prediction g x (i,j) and g y (i,j) may be calculated at one or more (eg, each) pixel / sample locations (eg, those described herein).
[0183] the motion vector difference MV between the current sub-block and one or more neighboring sub-blocks under consideration (neighboring sub-blocks under consideration) diff may be calculated. This MV difference may be at the sub-block level. Neighboring sub-blocks of each candidate that are not far from (e.g., close or nearby) the current sub-block may be considered. MV diff Various quantities and / or locations of sub-blocks may be selected as adjacent sub-blocks to calculate MV. In an example, BBPROF in DMVR mode may select four adjacent sub-blocks (e.g., left, upper, right, and lower adjacent sub-blocks), two adjacent sub-blocks (e.g., upper and left adjacent sub-blocks), corner adjacent sub-blocks (e.g., upper left, lower right), or other quantities and locations of adjacent sub-blocks. diff can be used to calculate
[0184] In one example, four adjacent sub-blocks (e.g., the sub-blocks above, below, left, and right of the current sub-block) may be considered, and MV diff can be an MV difference set containing four different MV difference values. For example, MV diff MV diff ={MV diff (A),MV diff (B),MV diff (L),MV diff (R)}, where A, B, L, and R may represent the MV differences between the current sub-block and the upper, lower, left, and right sub-blocks, respectively.
[0185] 12 shows an example MV difference calculation from selected adjacent sub-blocks, for example, in DMVR mode. As shown in FIG. 12, a sub-block may have its own MV difference after DMVR. A current sub-block in a CU (e.g., the current DMVR sub-block in FIG. 12) may have, for example, four connected adjacent sub-blocks, excluding the current sub-block located at the boundary.
[0186] Calculated sub-block level MV difference MV diff may be used to calculate a motion vector offset Δv(i,j) at one or more (e.g., each) pixel / sample locations within the current sub-block, for example, as shown in Equation 21.
[0187]
number
[0188] where n may be an index for a particular adjacent sub-block, N may be the total number of adjacent sub-blocks under consideration, and w(i,j,n) is the number of adjacent MVs. diff (n) may be a weighting factor when applied for a particular pixel at location (i,j). N may be equal to 4, for example, if the left, upper, right, and lower adjacent sub-blocks are considered.
[0189] The set of weighting factors may be, for example, {1 / 4, 1 / 8, 1 / 16, 1 / 32}. The weighting factors may be used by four rows / columns of pixels on one or more (e.g., each) sides of the current sub-block, respectively. MV Difference MV diffmay be calculated, for example, based on the motion vectors of vertically and / or horizontally adjacent sub-blocks. Pixels in the same row and / or column of the current sub-block may use the same weighting factor. For example, pixels in the first column to the left of the current sub-block may use the same weighting factor (e.g., 1 / 4), pixels in the second column may use the same weighting factor (e.g., 1 / 8), etc. The weighting factor may be determined, for example, based on the distance from the current position to the block boundary between the current block and its adjacent blocks. The weighting factor may be smaller, for example, when the column and / or row is farther from the block boundary.
[0190] The weighting coefficients may be adjusted (e.g., dynamically) based on, for example, pixel location. In one example, the pixel may be located to the top left of the current sub-block. The MV differences from left and upper adjacent sub-blocks may be weighted / combined together to generate a final MV offset at those pixels, for example, if both left and upper adjacent sub-blocks exist and if the MV differences from the left and upper adjacent sub-blocks are both non-zero. For example, if the current sub-block is at the left or top boundary of the current CU, the left or upper adjacent sub-blocks may not be available.
[0191] The intensity change per pixel within the current sub-block may be calculated, for example, according to optical flow equation 22. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (22) Here, Δv(i,j) and g(i,j) may be the MV offset and spatial gradient at one or more (e.g., every) sample locations (i,j), which may be calculated, for example, in the previous step.
[0192] For example, by adding the calculated intensity change (e.g., luminance or chrominance) to the sub-block prediction, the prediction for the pixel or sample location may be refined. The refined prediction for the pixel or sample location may be associated with a certain reference picture list, e.g., list L0 or list L1. The final prediction I′ may be generated, for example, according to Equation 23. I'(i,j)=I(i,j)+ΔI(i,j) (23) When BBPROF is applied in DMVR mode, four adjacent sub-blocks (e.g., at most) may be considered. Interior sub-blocks may wait for the DMVR process of adjacent sub-blocks to complete. In an example, when BBPROF is applied in DMVR mode, two adjacent sub-blocks may be considered. In an example, for example, two neighbors (e.g., only the top and left neighbors) may be considered, so that the BBPROF of the current sub-block may depend on the DMVR process for the two adjacent sub-blocks.
[0193] BBPROF may be applied in SbTMVP mode. For example, BBPROF may reduce discontinuities at sub-block boundaries in SbTMVP-based sub-block-level motion compensation prediction. One or more examples herein for applying BBPROF in DMVR mode may be applicable to implementing BBPROF in SbTMVP mode. Applying BBPROF in SbTMVP mode may include performing one (e.g., only one) motion compensation operation per sub-block. SbTMVP motion compensation may perform one motion compensation operation per sub-block. Applying BBPROF in SbTMVP mode may include one or more of the following:
[0194] For example, by implementing SbTMVP (e.g., as described herein), a refined motion vector for one or more (e.g., each) of the sub-blocks within a CU may be derived. Motion information may be fetched from a co-located sub-CU. Appropriate temporal scaling may be applied to the motion information. For example, sub-block-based motion compensation may be performed to generate a sub-block-based prediction.
[0195] Spatial gradient of subblock prediction g x (i,j) and g y (i,j) may be calculated at one or more (eg, all) pixel / sample locations (eg, those described herein).
[0196] The motion vector difference MV between the current sub-block and, for example, one or more neighboring sub-blocks under consideration diff can be calculated.
[0197] The prediction refinement described herein may be applied in ATMVP. In one example, the prediction refinement may use the motion vectors of the four adjacent sub-blocks of the current sub-block shown in Figure 9. The motion vectors are the MV differences MV between the current sub-block and the spatially adjacent sub-blocks. diff can be used to derive
[0198] Calculated sub-block level MV difference MV diff may be used to calculate a motion vector offset Δv(i,j) at one or more (eg, each) of the pixel / sample locations within the current sub-block.
[0199] The intensity change per pixel within the current sub-block may be calculated by an optical flow equation (eg, based on Equation 22).
[0200] The prediction for one or more (e.g., each) of the reference picture lists may be refined, for example, by adding intensity changes (e.g., luminance or chrominance). A final prediction I′ may be generated, for example, according to Equation 24. I'(i,j)=I(i,j)+ΔI(i,j) (24) BBPROF may be applied in affine mode. BBPROF may be applied to affine-coded CUs (e.g., similar to SbTMVP). An affine-coded CU may include multiple sub-blocks. Block-level MVs for one or more (e.g., each) of the sub-blocks may be derived by an affine motion model (e.g., as described herein). Four parameters of a four-parameter affine model and / or six parameters of a six-parameter affine model may be estimated, for example, with two or three control point motion vectors. The four or six estimated affine model parameters may be used, for example, to derive block-level motion vectors for sub-blocks within the affine-coded CU. Sub-block-level MV difference MV at one or more (e.g., each) of pixel / sample locations within the sub-block diff and / or the motion vector offset Δv(i,j) may be calculated, for example, at different sub-block motion vectors, e.g., sub-block level MV difference MV diff and / or a motion vector offset Δv(i,j) may be calculated (eg, BBPROF may be implemented as described herein).
[0201] BBPROF may be applied to pixels adjacent to sub-block boundaries, for example, as described herein, and to pixels adjacent to CU boundaries. For example, BBPROF may be applied at the CU level. The reference pictures may be the same or different for adjacent CUs, for example.
[0202] CU level MV difference MV diffFor example, BBPROF may be calculated (e.g., directly) for a particular CU if the selected neighboring CUs (e.g., upper, lower, left, and right CUs) have the same reference picture as the current CU. The prediction of the boundary pixels of a particular CU may be refined (e.g., directly) by applying BBPROF, for example.
[0203] For example, if (i) one or more (e.g., all) of the selected neighboring CUs (e.g., upper, lower, left, and right CUs) and (ii) the current CU have different reference pictures in their reference picture lists, temporal motion scaling may be applied to the particular CU. Appropriate temporal motion scaling may align the reference pictures of the temporal motion vectors of the selected neighboring CUs to the reference picture of the particular CU. CU-level MV difference MV diff may be calculated (eg, based on the scaled MVs of selected neighboring CUs), and prediction refinement at CU boundaries may be achieved, for example, by applying BBPROF.
[0204] As additional examples, several implementation variations of BBPROF are provided. The quantity and location of adjacent (e.g., adjacent) sub-blocks that may be selected for sub-block / block refinement, such as sub-block / block boundary refinement (e.g., BBPROF), may not be limited to the examples described herein, such as the example described with reference to FIG. 12. Other quantities and / or locations of sub-blocks may be selected. Adjacent sub-blocks of each candidate that are not far from (e.g., adjacent or near) the current sub-block may be considered. For example, BBPROF may be implemented using MV diff To calculate , one may use adjacent sub-blocks at the corners (e.g., top left, bottom right), the four sub-blocks shown in FIG. 12, or other quantities and positions of adjacent sub-blocks.
[0205] The location of the adjacent sub-block under consideration may not be restricted to being within the same CU as the current sub-block. The aspect ratio of the adjacent sub-block under consideration from an adjacent CU may or may not be restricted to being the same as the current sub-block. Different aspect ratios may be allowed for sub-block / block refinement, such as sub-block / block boundary refinement (e.g., BBPROF).
[0206] An MV offset at a pixel / sample location may be derived, for example, based on a sub-block MV difference. The number of rows and / or columns of pixels on one or more (e.g., each) of the sides of the current sub-block may be configurable and / or dynamically changed, for example, based on one or more (e.g., predefined) criteria. For example, two or more columns of pixels on the left side of the current sub-block may include an MV difference from the adjacent sub-block to the left, instead of a default number of columns of pixels, such as four columns of pixels.
[0207] The MV difference may be based on vertically and / or horizontally adjacent sub-blocks. In an example, pixels in the same row and / or column of the current sub-block may use the same weighting factor. In an example, pixels in the same row and / or column of the current sub-block may use different weighting factors.
[0208] The weighting factor may vary, for example, the weighting factor may increase as the spatial distance between the pixel and the vertical or horizontal boundary decreases.
[0209] The intensity difference (eg, derived from Equation 22) may be multiplied by a weighting factor w, eg, before the intensity difference is added to the prediction, eg, as shown in Equation 25. I'(i,j)=I(i,j)+w·ΔI(i,j) (25) where w may be set to a value between 0 and 1. w may be signaled, for example, at the CU level or the picture level. For example, w may be signaled by a weight index. Equation 25 may be a variation of Equation 23 and / or Equation 24.
[0210] BBPROF may be used, for example, after DMVR-based L0 prediction and L1 prediction are combined with weights. BBPROF may be applied to, for example, one prediction, such as L0 or L1, for example, to reduce complexity. In an example, BBPROF may be applied to, for example, one prediction when the reference picture is closer to the current picture in the temporal domain. In an example, BBPROF may be applied to, for example, one prediction when the reference picture is farther from the current picture in the temporal domain.
[0211] Numerous embodiments are described herein. Features of the embodiments may be provided alone or in any combination across various claim categories and types. Furthermore, an embodiment may include one or more of the features, devices, or aspects described herein alone or in any combination across various claim categories and types, such as, for example, any of the following:
[0212] The method described in FIG. 11 may be applied to a decoder and / or an encoder. When the method described in FIG. 11 is applied to an encoder, 1104 in FIG. 11 may be performed by the encoder, and 1104 may involve encoding a block including the current sub-block based on a sample value for the obtained first pixel. The method described in FIG. 11 may be based on one or more of Equations (1) to (25). For example, the decoder may decode the current sub-block based on a sample value for the pixel. The pixel may be located on one of the boundaries of the current sub-block. The sample value may be a refined sample value obtained based on one or more of Equations (1) to (25). As shown in one or more of Equations (1) to (25), the decoder may obtain a sample value for a pixel based on, for example, an MV for the current sub-block, an MV for a sub-block adjacent to the current sub-block, and a sample value for a pixel adjacent to the pixel for which the sample value is obtained. The decoder may obtain a prediction of a sample value for a pixel, for example, before the decoder refines the prediction of the sample value. The prediction of the sample value for the pixel may be referred to as I(i,j), for example, as shown in Equation (23). As shown in Equation (23), the decoder may obtain a sample value for the pixel based on the sample value offset and the sample value prediction, for example, the sum of the sample value offset and the sample value prediction. The sample value offset for the pixel may be referred to as ΔI(i,j), for example, as shown in Equation (23). As shown in one or more of Equations (1) through (25), the sample value offset may be obtained based on, for example, the MV for the current sub-block, the MV for the sub-block adjacent to the current sub-block, and the sample values for the pixels adjacent to the pixel for which the sample value is obtained. The decoder may obtain an MV difference (e.g., the MV difference in Equation (21)) using the MV for the current sub-block and the MV for the sub-block adjacent to the current sub-block, as described herein.The MV difference is, for example, MV, as shown in Equation 21. diff (n). The decoder may obtain an MV difference using one or more MVs associated with one or more respective sub-blocks adjacent to the current sub-block. The decoder may obtain a gradient using sample values for pixels adjacent to the pixel for which the sample value is obtained, for example, as shown in Equation (15) and Equation (16). As an example shown in Equation (15) and Equation (16), the decoder may obtain a gradient using one or more sample values for one or more respective pixels adjacent to the pixel for which the sample value is obtained. The decoder may obtain a sample value for the pixel based on the gradient and the MV difference. In one example, the decoder may obtain an MV offset based on the MV difference, as shown in Equation (21). The decoder may use the MV offset and gradient to obtain a sample value for the pixel, as shown in Equation (22) and Equation (23). The decoder may obtain a sample value offset based on the gradient and MV difference, and use the sample value offset to obtain a sample value for the pixel. The decoder may determine a weighting factor and use the weighting factor to obtain a sample value for the pixel, for example, as shown in Equation (21). The decoder may decode the block, including the current sub-block, based on the obtained sample values for the pixels.
[0213] Decoding tools and techniques including one or more of entropy decoding, inverse quantization, inverse transform, and differential decoding may be used to enable the method described in Figure 11 at a decoder. These decoding tools and techniques may be used to enable one or more of sub-block / block refinement according to the method described in Figure 11, sub-block / block boundary refinement according to the method described in Figure 11, BBPROF according to the method described in Figure 11, sub-block / block refinement in DMVR mode, sub-block / block refinement in SbTMVP mode, sub-block / block refinement in affine mode, obtaining sample values according to the method described in Figure 11, obtaining sample value offsets according to the method described in Figure 11, obtaining gradients as described herein, obtaining MV differences as described herein, obtaining predictions for sample values, and other decoder behaviors related to any of the above.
[0214] The encoder may encode the current subblock based on a sample value for a pixel. The pixel may be located on one of the boundaries of the current subblock. The sample value may be a refined sample value obtained based on one or more of Equations (1) through (25). As shown in one or more of Equations (1) through (25), the encoder may obtain a sample value for a pixel based on, for example, an MV for the current subblock, an MV for a subblock adjacent to the current subblock, and a sample value for a pixel adjacent to the pixel for which the sample value is obtained. The encoder may obtain a prediction of the sample value for the pixel, for example, before the encoder refines the prediction of the sample value. The prediction of the sample value for the pixel may be referred to as I(i,j), for example, as shown in Equation (23). As shown in Equation (23), the encoder may obtain a sample value for the pixel based on a sample value offset and a prediction of the sample value, for example, a sum of the sample value offset and the prediction of the sample value. The sample value offset for a pixel may be referred to as ΔI(i,j), e.g., as shown in Equation (23). As shown in one or more of Equations (1) through (25), the sample value offset may be obtained, e.g., based on the MV for the current sub-block, the MV for the sub-block adjacent to the current sub-block, and the sample values for the pixel adjacent to the pixel for which the sample value is obtained. The encoder may obtain an MV difference (e.g., the MV difference in Equation (21)) using the MV for the current sub-block and the MV for the sub-block adjacent to the current sub-block, as described herein. The MV difference may be obtained, e.g., based on the MV diff(n). The encoder may obtain an MV difference using one or more MVs associated with one or more respective sub-blocks adjacent to the current sub-block. As shown in Equation (15) and Equation (16), the encoder may obtain a gradient using sample values for pixels adjacent to the pixel for which the sample value is obtained. As an example shown in Equation (15) and Equation (16), the encoder may obtain a gradient using one or more sample values for one or more respective pixels adjacent to the pixel for which the sample value is obtained. The encoder may obtain a sample value for the pixel based on the gradient and the MV difference. In one example, the encoder may obtain an MV offset based on the MV difference as shown in Equation (21). The encoder may use the MV offset and gradient to obtain a sample value for the pixel as shown in Equation (22) and Equation (23). The encoder may obtain a sample value offset based on the gradient and MV difference and use the sample value offset to obtain a sample value for the pixel. The encoder may determine a weighting factor and use the weighting factor to obtain a sample value for the pixel, for example, as shown in Equation (21). The encoder may encode the block that includes the current sub-block based on the obtained sample values for the pixels.
[0215] Encoding tools and techniques including one or more of quantization, entropy coding, inverse quantization, inverse transform, and differential encoding may be used to enable in an encoder the method described in Figure 11. These encoding tools and techniques may be used to enable one or more of sub-block / block refinement according to the method described in Figure 11, sub-block / block boundary refinement according to the method described in Figure 11, BBPROF according to the method described in Figure 11, sub-block / block refinement in DMVR mode, sub-block / block refinement in SbTMVP mode, sub-block / block refinement in affine mode, obtaining sample values according to the method described in Figure 11, obtaining sample value offsets according to the method described in Figure 11, obtaining gradients as described herein, obtaining MV differences as described herein, obtaining predictions for sample values, and other encoder behaviors related to any of the above.
[0216] For example, a syntax element may be inserted into the signaling to enable a decoder to identify an indication associated with performing, or for use, the method described in Figure 11. For example, the syntax element may include an indication of one or more of BBPROF, DMVR, SbTMVP mode, affine mode, e.g., to indicate to the decoder whether one or more of them are enabled or disabled. As an example, the syntax element may include an indication of one or more weighting factors described herein and / or an indication of parameters that the decoder uses to perform one or more examples herein.
[0217] The method described in Figure 11 may be selected and / or applied for application at a decoder, for example, based on syntax elements. For example, the decoder may receive an indication to enable BBPROF. Based on the indication, the decoder may perform the method described in Figure 11 for pixels located at or near sub-block boundaries.
[0218] The encoder may adjust the prediction residual based on one or more examples herein. The residual may be obtained, for example, by subtracting a predicted video block from an original image block. For example, the encoder may predict a video block based on sample values for pixels obtained as described herein. The encoder may obtain the original image block and subtract the predicted video block from the original image block to generate a prediction residual.
[0219] The bitstream or signal may include one or more of the described syntax elements or variations thereof. For example, the bitstream or signal may include a syntax element indicating that any of BBPROF, DMVR, SbTMVP mode, or affine mode is enabled or disabled.
[0220] A bitstream or signal may include syntax that carries information generated in accordance with one or more examples of this specification. For example, the information or data may be generated when implementing the example shown in Figure 11. The generated information or data may be carried within syntax included within the bitstream or signal.
[0221] Syntax elements may be inserted into the signal that allow the decoder to adjust the residual to correspond to that used by the encoder. For example, the residual may be generated using one or more examples herein.
[0222] A method, process, apparatus, instruction storage medium, data storage medium, or signal for creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof.
[0223] A method, process, apparatus, instruction storage medium, data storage medium, or signal for creating and / or transmitting and / or receiving and / or decoding according to any of the described examples.
[0224] and a method, process, apparatus, instruction storage medium, data storage medium, or signal according to one or more of the following, but not limited to: determining a spatial gradient of sub-block based prediction at one or more pixel / sample locations; using MV differences to calculate motion vector offsets at one or more pixel / sample locations; determining an intensity change per pixel within a current sub-block based on, for example, optical flow; refining a prediction for a reference picture list by, for example, adding the calculated intensity change to a sub-block prediction; determining that a first pixel is adjacent to a boundary of the current sub-block; determining a difference between an MV for the current sub-block and an MV for a sub-block adjacent to the current sub-block; determining a gradient for a first pixel based on a sample value for a second pixel adjacent to the first pixel and a sample value for a third pixel adjacent to the first pixel; determining a sample value offset based on a difference between an MV for the current sub-block and an MV for the sub-block adjacent to the current sub-block; obtaining a sample value for the first pixel based on the determined sample value offset, e.g., determining a gradient based on the sample values for at least the second pixel; determining a sample value for the first pixel using the gradient, e.g., determining a gradient for an optical flow model based on the sample values for the at least the second pixel; using the gradient in the optical flow model to obtain a sample value for the first pixel; using a difference between an MV for the current sub-block and an MV for the sub-block adjacent to the current sub-block to obtain a sample value for the first pixel, e.g., further based on an MV for the second sub-block adjacent to the current sub-block; obtaining a sample value for the first pixel based on determining that the first pixel is adjacent to a boundary of the current sub-block;using a weighting factor to obtain a sample value for a first pixel, the weighting factor may vary or not vary according to the distance of the first pixel from a boundary of a corresponding current sub-block; determining a sample value offset for the first pixel based on, for example, an MV for the current sub-block, an MV for a sub-block adjacent to the current sub-block, and a sample value for a second pixel adjacent to the first pixel; obtaining a sample value for the first pixel using the determined sample value offset and a predicted sample value for the first pixel; and obtaining a sample value for the first pixel based on determining that the first pixel is adjacent to the boundary of the current sub-block, for example, when the boundary of the current sub-block may include a common boundary between the current sub-block and a sub-block adjacent to the current sub-block.
[0225] A TV, set-top box, cell phone, tablet, or other electronic device that implements block / sub-block / CU refinement according to any of the described examples.
[0226] A TV, set-top box, cell phone, tablet, or other electronic device that performs block / sub-block / CU refinement according to any of the described examples and displays the resulting image (e.g., using a monitor, screen, or other type of display).
[0227] A TV, set-top box, cell phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal containing an encoded image and performs block / sub-block / CU refinement according to any of the described examples.
[0228] A TV, set-top box, cell phone, tablet, or other electronic device that receives a signal containing an encoded image wirelessly (e.g., using an antenna) and performs block / sub-block / CU refinement according to any of the described examples.
[0229] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer. [Explanation of symbols]
[0230] 100 Communication Systems 200 Encoder 300 decoder 400 System 1100 methods
Claims
1. Based on a sample value for a second pixel and based on a motion vector (MV) difference between a first video block and a second video block, a sample value for a first pixel located at a boundary of the first video block is obtained, the second video block being adjacent to the first video block, the second pixel being adjacent to the first pixel, Decoding the first video block based on the obtained sample value for the first pixel. A processor configured to 1. A device for video decoding, comprising:
2. Based on a sample value for a second pixel and based on a motion vector (MV) difference between a first video block and a second video block, a sample value for a first pixel located at a boundary of the first video block is obtained, the second video block being adjacent to the first video block, the second pixel being adjacent to the first pixel, encoding the first video block based on the obtained sample values for the first pixels; A processor configured to 1. A device for video encoding, comprising:
3. Obtaining a sample value for a first pixel located at a boundary of a first video block based on a sample value for a second pixel and based on a motion vector (MV) difference between the MV for a first video block and the MV for a second video block, wherein the second video block is adjacent to the first video block and the second pixel is adjacent to the first pixel; decoding the first video block based on the obtained sample value for the first pixel; and 1. A method for video decoding, comprising:
4. Obtaining a sample value for a first pixel located at a boundary of a first video block based on a sample value for a second pixel and based on a motion vector (MV) difference between the MV for a first video block and the MV for a second video block, wherein the second video block is adjacent to the first video block and the second pixel is adjacent to the first pixel; encoding the first video block based on the obtained sample values for the first pixels; and 1. A method for video encoding, comprising:
5. the first video block includes the first pixel, the second pixel, and a third pixel adjacent to the first pixel, and the processor: determining that the first pixel is a boundary pixel of the first video block; based on a determination that the first pixel is the boundary pixel of the first video block, determining the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block; determining a gradient for the first pixel based on the sample value for the second pixel and the sample value for the third pixel; determining a sample value offset based on the determined gradient and based on the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block, and the sample value for the first pixel is obtained based on the determined sample value offset; 10. The device of claim 1, further configured to:
6. 6. The device of claim 5, wherein the reference picture for the first video block and the reference picture for the second video block are the same.
7. the first video block is associated with a first reference picture and the second video block is associated with a second reference picture, and the processor: determining that the first reference picture and the second reference picture are different; Obtaining the MV for the second video block using scaling 6. The device of claim 5, further configured to:
8. 2. The device of claim 1, wherein a gradient for an optical flow model is determined based at least on the sample value for the second pixel, and the gradient is used in the optical flow model to obtain the sample value for the first pixel.
9. the first video block includes the first pixel, the second pixel, and a third pixel adjacent to the first pixel, and the processor: determining that the first pixel is a boundary pixel of the first video block; based on a determination that the first pixel is the boundary pixel of the first video block, determining the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block; determining a gradient for the first pixel based on the sample value for the second pixel and the sample value for the third pixel; determining a sample value offset based on the determined gradient and based on the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block, and the sample value for the first pixel is obtained based on the determined sample value offset; 3. The device of claim 2, further configured to:
10. the first video block includes the first pixel, the second pixel, and a third pixel adjacent to the first pixel, and obtaining the sample value for the first pixel includes: determining that the first pixel is a boundary pixel of the first video block; based on a determination that the first pixel is the boundary pixel of the first video block, determining the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block; determining a gradient for the first pixel based on the sample value for the second pixel and the sample value for the third pixel; determining a sample value offset based on the determined gradient and based on the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block; obtaining the sample value for the first pixel based on the determined sample value offset; 4. The method of claim 3, comprising:
11. 11. The method of claim 10, wherein the reference picture for the first video block and the reference picture for the second video block are the same.
12. the first video block includes the first pixel, the second pixel, and a third pixel adjacent to the first pixel, and obtaining the sample value for the first pixel includes: determining that the first pixel is a boundary pixel of the first video block; based on a determination that the first pixel is the boundary pixel of the first video block, determining the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block; determining a gradient for the first pixel based on the sample value for the second pixel and the sample value for the third pixel; determining a sample value offset based on the determined gradient and based on the MV difference between the MV for the first video block and the MV for the second video block adjacent to the first video block; obtaining the sample value for the first pixel based on the determined sample value offset; 5. The method of claim 4, comprising:
13. 13. The method of claim 3, wherein a weighting factor that varies according to the distance of the first pixel from a corresponding boundary of the first video block is used to obtain the sample value for the first pixel.
Citation Information
Patent Citations
Video signal processing apparatus and method, recording medium, and program
JP2003209811A
Image processing apparatus, image processing method, program, and storage medium
JP2013090034A