Block boundary optical flow prediction correction
By using block boundary optical flow prediction correction technology, the prediction at the block boundary of the video coding system is corrected by using motion vector difference and pixel gradient, which solves the problem of insufficient prediction accuracy in the existing technology and improves the efficiency and quality of video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTERDIGITAL VC HOLDINGS INC
- Filing Date
- 2020-06-02
- Publication Date
- 2026-05-19
AI Technical Summary
Existing video coding systems suffer from insufficient prediction accuracy in optical flow prediction at block boundaries, which affects video coding efficiency and quality.
The Block Boundary Optical Flow Prediction Correction (BBPROF) technique is adopted. By calculating the difference in motion vectors and pixel gradients between the current sub-block and its neighboring sub-blocks, and combining them with the optical flow model, the sample value prediction of pixels is corrected, thereby improving the prediction accuracy.
It improves the prediction accuracy of video coding systems at block boundaries, thereby enhancing video coding efficiency and quality.
Smart Images

Figure CN122069360A_ABST
Abstract
Description
[0001] This is a divisional application. The parent application is entitled "Block Boundary Optical Flow Prediction Correction", filed on June 2, 2020, with application number 202080046609.6.
[0002] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 856,519, filed June 3, 2019, entitled “Block Boundary Prediction Refinement with Optical Flow,” the entire contents of which are incorporated herein by reference as if fully set forth herein. Background Technology
[0003] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include block-based, wavelet-based, and / or object-based systems. Hybrid block-based video coding systems can be deployed. Summary of the Invention
[0004] Systems, methods, and tools for sub-block / block correction are disclosed, including sub-block / block boundary correction, such as Block Boundary Optical Flow Prediction Correction (BBPROF). A block including the current sub-block can be decoded based on sample values obtained for a first pixel, which can be obtained from, for example, the motion vector (MV) of the current sub-block, the MVs of neighboring sub-blocks, and sample values of a second pixel adjacent to a first pixel. Sub-block / block correction can be applied to decoder-side motion vector correction (DMVR) mode, sub-block-based temporal motion vector prediction (SbTMVP) mode, and / or affine mode. BBPROF can include, for example, sub-block-based motion compensation to generate sub-block-based predictions. Spatial gradients of sub-block-based predictions at one or more pixel / sample locations can be computed. MV differences between the current sub-block and one or more neighboring sub-blocks can be computed. MV differences can be used to compute motion vector offsets at one or more pixel / sample locations. Intensity changes for each pixel in the current sub-block can be computed based on optical flow. Sample value offsets can be used to indicate the intensity changes for each pixel. The prediction of pixel or sample location can be corrected, for example, by adding the calculated intensity change to the sub-block prediction (e.g., motion-compensated prediction).
[0005] In the example, a method may be implemented to perform sub-block / block correction. The method may be implemented, for example, by an apparatus that may include one or more processors configured to execute computer-executable instructions, which may be stored on a computer-readable medium or computer program product, and which, when executed by the one or more processors, execute the method. Therefore, the apparatus may include one or more processors configured to execute the method. The computer-readable medium or computer program product may include instructions that cause the one or more processors to execute the method by executing the instructions. The computer-readable medium may contain data content generated according to the method. The signal may include a residual generated based on an original image block and a block predicted according to the method using obtained sample values of a first pixel. The apparatus may include an access unit and a transmitter configured to execute a second method comprising: accessing data including the residual generated by the apparatus based on obtained sample values of a first pixel, the apparatus including one or more processors configured to implement the method (e.g., by executing instructions); and transmitting the data including the residual. Devices, such as televisions, mobile phones, tablets, or set-top boxes, may include: means having one or more processors configured to implement a method (e.g., by executing instructions); and at least one of: (i) an antenna configured to receive a signal including data representing an image; (ii) a band limiter configured to limit the received signal to a band including data representing an image; or (iii) a display configured to display an image.
[0006] Methods for performing sub-block / block correction may include, for example, obtaining a sample value of a first pixel based on, for example, the MV of the current sub-block, the MV of the sub-block adjacent to the current sub-block, and a sample value of a second pixel adjacent to the first pixel; and decoding a block including the current sub-block based on the obtained sample value of the first pixel.
[0007] The method of encoding a block including the current sub-block based on the obtained sample value of the first pixel may include, for example, obtaining the sample value of the first pixel based on, for example, the MV of the current sub-block, the MV of the sub-block adjacent to the current sub-block, and the sample value of the second pixel adjacent to the first pixel; and encoding the block including the current sub-block based on the obtained sample value of the first pixel.
[0008] A block may include, for example, a first pixel, a second pixel, and a third pixel adjacent to the first pixel. Obtaining a sample value of the first pixel may include, for example, determining that the first pixel is adjacent to the boundary of the current sub-block; determining: (i) the difference between the MV of the current sub-block and the MV of the sub-block adjacent to the current sub-block, (ii) the gradient of the first pixel based on the sample values of the second pixel and the third pixel, and (iii) a sample value offset based on the determined gradient and the difference between the MV of the current sub-block and the MV of the sub-block adjacent to the current sub-block; and obtaining a sample value of the first pixel based on the determined sample value offset.
[0009] The gradient can be determined, for example, based at least on the sample value of the second pixel. The gradient can be used to obtain the sample value of the first pixel.
[0010] The gradient of the optical flow model can be determined, for example, based at least on the sample value of the second pixel. The gradient can be used in the optical flow model to obtain the sample value of the first pixel.
[0011] The sample value of the first pixel can be obtained by using the difference between the MV of the current sub-block and the MV of the sub-block adjacent to the current sub-block.
[0012] The sub-block adjacent to the current sub-block can be the first sub-block. A block may include the first sub-block and a second sub-block adjacent to the current sub-block. The sample value of the first pixel can be obtained (for example, further) based on the MV of the second sub-block.
[0013] The sample value of the first pixel can be obtained based on determining that the first pixel is adjacent to the boundary of the current sub-block.
[0014] The first and second pixels can be in the current sub-block.
[0015] A weighting factor can be used to obtain the sample value of the first pixel. The weighting factor can vary based on the distance of the first pixel from the corresponding boundary of the current sub-block.
[0016] The sample value offset of the first pixel can be determined, for example, based on the MV of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and the sample value of the second pixel adjacent to the first pixel. The determined sample value offset of the first pixel and the predicted sample value can be used to obtain the sample value of the first pixel.
[0017] The sample value of the first pixel can be obtained, for example, based on determining that the first pixel is adjacent to the boundary of the current sub-block. The boundary of the current sub-block may include the common boundary between the current sub-block and the sub-blocks adjacent to the current sub-block.
[0018] The first pixel may be located, for example, in four rows of pixels from the top boundary of the current sub-block, in four rows of pixels from the left boundary of the current sub-block, in four columns of pixels from the left boundary of the current sub-block, or in four columns of pixels from the right boundary of the current sub-block.
[0019] Each feature disclosed anywhere in this document is described, and such feature may be implemented separately / individually and in any combination with any other feature disclosed herein and / or with any feature disclosed elsewhere, whether implicitly or explicitly mentioned herein or otherwise falling within the scope of the subject matter disclosed herein. Attached Figure Description
[0020] Figure 1A This is a system diagram illustrating an exemplary communication system in which one or more of the disclosed embodiments may be implemented.
[0021] Figure 1B This shows that, according to the implementation plan, it is possible to Figure 1A A system diagram of an exemplary wireless transmit / receive unit (WTRU) used within the communication system shown.
[0022] Figure 1C This illustrates the implementation scheme. Figure 1A The diagram shows an exemplary radio access network (RAN) and an exemplary core network (CN) used within the communication system.
[0023] Figure 1D This illustrates the implementation scheme. Figure 1A The system diagram shows another exemplary RAN and another exemplary CN used within the communication system shown.
[0024] Figure 2 This is a schematic diagram illustrating an exemplary video encoder.
[0025] Figure 3 This is a schematic diagram illustrating an example of a video decoder.
[0026] Figure 4 This is a schematic diagram illustrating an example of a system in which various aspects and examples can be implemented.
[0027] Figure 5 An exemplary four-parameter affine pattern model of an affine block and the derivation of sub-block-level motion are shown.
[0028] Figure 6 An exemplary six-parameter affine pattern is shown, where V0, V1, and V2 are control points, and (MV x MV y ) is the motion vector of the sub-block centered at position (x, y).
[0029] Figure 7 An exemplary decoding-side motion vector (MV) correction is shown.
[0030] Figure 8AAn example of a spatially neighboring block that can be used by sub-block-based temporal motion vector prediction (SbTMVP) is shown.
[0031] Figure 8B An exemplary derivation of the motion field of the sub-coding unit (CU) is shown.
[0032] Figure 9 An exemplary sub-block with Overlapping Block Motion Compensation (OBMC) applied is shown.
[0033] Figure 10 An example sub-block MV (V) is shown. SB ) and pixels .
[0034] Figure 11 Examples of methods for sub-block / block correction based on one or more of equations (1) to (25) are shown.
[0035] Figure 12 An exemplary calculation of the MV difference from a selected neighboring sub-block is shown. Detailed Implementation
[0036] A detailed description of exemplary embodiments will now be described with reference to the various accompanying drawings. Although this specification provides detailed examples of possible specific implementations, it should be noted that the details are intended to be exemplary and in no way limit the scope of this application.
[0037] Figure 1A This is a schematic diagram illustrating an exemplary communication system 100 that can be implemented in one or more of the disclosed embodiments. Communication system 100 can be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. Communication system 100 enables multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, communication system 100 can employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0038] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile user units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0039] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106 / 115, Internet 110, and / or other networks 112. As an example, base stations 114a and 114b may be base transceiver stations (BTS), Node Bs, evolved Node Bs, home Node Bs, home evolved Node Bs, gNBs, NR Node Bs, site controllers, access points (APs), wireless routers, etc. Although base stations 114a and 114b are each depicted as a single element, it should be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0040] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies (which may be referred to as cells (not shown)). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of radio services to a specific geographic area, which may be relatively fixed or changeable over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.
[0041] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.
[0042] More specifically, as noted above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRUs 102a, 102b, and 102c in RAN 104 / 113 may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may use Wideband CDMA (WCDMA) to establish air interfaces 115 / 116 / 117. WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0043] In one implementation, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as evolved UMTS terrestrial radio access (E-UTRA), which can use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.
[0044] In one implementation, base station 114a and WTRUs 102a, 102b, 102c can enable radio technologies such as NR radio access, which can use New Radio (NR) to establish air interface 116.
[0045] In one implementation, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can, for instance, use a dual connectivity (DC) principle to implement both LTE and NR radio access together. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0046] In other implementations, base station 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), GSM Enhanced Data Rate Evolution (EDGE), and GSM EDGE (GERAN).
[0047] Figure 1ABase station 114b can be, for example, a wireless router, a home node B, a home evolution node B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in local areas such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for use by drones), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1A As shown, base station 114b may have a direct connection to Internet 110. Therefore, base station 114b may not need to access Internet 110 via CN 106 / 115.
[0048] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although not explicitly stated... Figure 1A As shown, but it should be understood that RAN 104 / 113 and / or CN106 / 115 can communicate directly or indirectly with other RANs that use the same RAT as RAN 104 / 113 or a different RAT. For example, in addition to being connected to RAN 104 / 113 which can utilize NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA2000, WiMAX, E-UTRA or WiFi radio technology.
[0049] CN 106 / 115 may also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0050] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can employ cellular-based radio technology and with a base station 114b that can employ IEEE 802 radio technology.
[0051] Figure 1B This is a system diagram illustrating an exemplary WTRU 102. (See diagram below.) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmitting / receiving element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It should be understood that WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with the implementation.
[0052] Processor 118 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 can be coupled to transceiver 120, which can be coupled to transmitting / receiving element 122. Although Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it should be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0053] Transmitting / receiving element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 may be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 may be configured to transmit and / or receive both RF and optical signals. It should be understood that transmitting / receiving element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0054] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, WTRU 102 may employ MIMO technology. Therefore, in one embodiment, WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.
[0055] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As noted above, WTRU 102 may have multi-mode capability. Therefore, transceiver 120 may include multiple transceivers to enable WTRU 102 to communicate via various RATs such as NR and IEEE 802.11.
[0056] The processor 118 of WTRU 102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and may receive user input data therefrom. The processor 118 may also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Furthermore, the processor 118 may access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132) and store data in any type of suitable memory. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a user identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 may access information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)) and store data in that memory.
[0057] The processor 118 may receive power from the power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 may be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell battery packs (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0058] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It should be understood that, while remaining consistent with the implementation, the WTRU 102 may acquire location information using any suitable location determination method.
[0059] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functions, and / or wired or wireless connectivity. For example, peripheral device 138 may include an accelerometer, electronic compass, satellite transceiver, digital camera (for photos and / or video), Universal Serial Bus (USB) port, vibration device, television transceiver, hands-free headset, Bluetooth.® Modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0060] WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference through signal processing via hardware (e.g., a choke) or via a processor (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception) may be concurrent and / or simultaneous.
[0061] Figure 1C This is a system diagram illustrating RAN 104 and CN 106 according to one embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with CN 106.
[0062] RAN 104 may include evolved Nodes B 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of evolved Nodes B while remaining consistent with the implementation scheme. Evolved Nodes B 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one implementation, evolved Nodes B 160a, 160b, and 160c may implement MIMO technology. Therefore, evolved Node B 160a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.
[0063] Each of the evolved nodes B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL, etc. Figure 1C As shown, evolution nodes B 160a, 160b, and 160c can communicate with each other via the X2 interface.
[0064] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is depicted as part of the CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0065] The MME 162 can connect to each of the evolved nodes B 162a, 162b, and 162c in RAN 104 via the S1 interface and can be used as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0066] The SGW 164 can connect to each of the evolved Node Bs 160a, 160b, and 160c in RAN104 via the S1 interface. The SGW 164 typically routes and forwards user data packets to and from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions such as anchoring the user plane during inter-evolved Node B handovers, triggering paging when DL data is available for WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.
[0067] SGW 164 can be connected to PGW 166, which provides WTRU 102a, 102b, 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.
[0068] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to circuit-switched networks (such as PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and conventional landline communication equipment. For example, CN 106 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 106 and PSTN 108, or be able to communicate with such an IP gateway. Additionally, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0069] Despite WTRU in Figures 1A to 1D While described as a wireless terminal, it is conceivable that in some representative implementations, such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.
[0070] In a representative implementation, the other network 112 may be a WLAN.
[0071] A WLAN in Infrastructure Basic Services Set (BSS) mode may have an access point (AP) for the BSS and one or more sites (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or carries traffic out of the BSS. Traffic originating outside the BSS and destined for a STA can reach and be delivered to the STA via the AP. Traffic originating from a STA and destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP, for example, where a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as point-to-point traffic. Point-to-point traffic can be sent between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative implementations, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). WLANs using the Standalone BSS (IBSS) mode may not have an access point (AP), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as the "ad-hoc" communication mode in this document.
[0072] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of fixed width (e.g., a 20 MHz wide bandwidth) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative implementations, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. For CSMA / CA, each STA (including the AP) can listen to the primary channel. If the primary channel is listened to / detected and / or determined to be busy by a particular STA, that STA can back off. A single STA (e.g., only one station) can transmit in a given BSS at any given time.
[0073] High-throughput (HT) STAs can communicate using a 40MHz wide channel, for example, by combining a primary 20MHz channel with adjacent or non-adjacent 20MHz channels to form a 40MHz wide channel.
[0074] Very High Throughput (VHT) STAs support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels, or by combining two non-consecutive 80MHz channels (this can be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data can be split into two streams by a segment parser. Each stream can be processed individually using Inverse Fast Fourier Transform (IFFT) and time-domain processing. These streams can be mapped to two 80MHz channels, and data can be transmitted via a transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration can be reversed, and the combined data can be sent to Media Access Control (MAC).
[0075] 802.11af and 802.11ah support operating modes below 1 GHz. Compared to those used in 802.11n and 802.11ac, 802.11af and 802.11ah reduce channel operating bandwidth and carrier. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative implementations, 802.11ah may support instrument-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain bandwidths and / or limited bandwidths. MTC devices may include batteries with battery life above a threshold (e.g., to maintain a very long battery life).
[0076] WLAN systems supporting multiple channels, and channel bandwidths such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that can be designated as primary channels. A primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by STAs operating in the BSS (each supporting a minimum bandwidth operating mode). In the 802.11ah example, for STAs supporting (e.g., only supporting) a 1MHz mode (e.g., MTC type devices), the primary channel can be 1MHz wide, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example, because an STA (supporting only the 1MHz operating mode) is transmitting to the AP, the entire available band can be considered busy even if most of the band remains idle and potentially available.
[0077] In the United States, the available frequency bands for 802.11ah are 902MHz to 928MHz. In South Korea, the available frequency bands are 917.5MHz to 923.5MHz. In Japan, the available frequency bands are 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah ranges from 6MHz to 26MHz, depending on the country code.
[0078] Figure 1D This is a system diagram illustrating RAN 113 and CN 115 according to one implementation scheme. As noted above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c via air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0079] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the implementation scheme. Each of gNBs 180a, 180b, and 180c may include one or more transceivers for communication with WTRUs 102a, 102b, and 102c via air interface 116. In one implementation, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Therefore, gNB 180a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one implementation, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a may transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In one implementation, gNBs 180a, 180b, and 180c may implement Cooperative Multipoint (CoMP) technology. For example, WTRU 102a may receive cooperative transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0080] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable parameter sets. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary depending on different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing different numbers of OFDM symbols and / or continuously varying absolute time lengths).
[0081] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., evolved Node Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobility anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate or connect to gNBs 180a, 180b, and 180c, and also communicate or connect to other RANs (such as eNode-B160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more evolved Node Bs 160a, 160b, and 160c. In a non-standalone configuration, evolved Node Bs 160a, 160b, and 160c can be used as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0082] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0083] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. Although each of the foregoing elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0084] AMF 182a and 182b can connect to one or more of gNBs 180a, 180b, and 180c via the N2 interface in RAN113 and can be used as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the type of service used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services that rely on Ultra-Reliable Low Latency (URLLC) access, services that rely on Enhanced Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, etc. The AMF162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0085] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0086] UPF 184a and 184b can connect via the N3 interface to one or more of the gNBs 180a, 180b, and 180c in RAN 113. These gNBs can provide WTRU 102a, 102b, and 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multihomed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0087] CN 115 may facilitate communication with other networks. For example, CN 115 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 115 and PSTN 108, or may communicate with such an IP gateway. Additionally, CN 115 may provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c may be connected to DNs 185a and 185b via UPFs 184a and 184b through an N3 interface to UPFs 184a and 184b and an N6 interface between UPFs 184a and 184b and local data networks (DNs) 185a and 185b.
[0088] Given Figures 1A to 1D as well as Figures 1A to 1D The corresponding descriptions herein refer to one or more of the functions described below, which may be performed by one or more emulation devices (not shown): WTRU102a-d, base station 114a-b, evolved Node B160a-c, MME 162, SGW 164, PGW 166, gNB180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN185a-b, and / or any other device described herein. An emulation device may be one or more devices configured to mimic one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0089] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, the one or more simulation devices may perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices may perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices may be directly coupled to another device for testing purposes and / or may use over-the-air wireless communication to perform tests.
[0090] The one or more emulation devices may perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be used in test scenarios within a test laboratory and / or non-deployed (e.g., testing) wired and / or wireless communication networks to perform testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) may be used by the emulation devices to transmit and / or receive data.
[0091] This application describes multiple aspects, including tools, features, examples or implementations, models, methods, etc. Many of these aspects are described in a particular manner and, at least to illustrate individual characteristics, are generally described in a way that may sound restrictive. However, this is for clarity and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in earlier filings.
[0092] The aspects described and envisioned in this application can be implemented in many different forms. Figures 5 to 12 Some implementation schemes are available, but other implementation schemes are also envisioned. Figures 5 to 12 The discussion does not limit the breadth of specific implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.
[0093] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0094] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined. Furthermore, terms such as "first," "second," etc., are used in various embodiments to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a sequence of modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and may occur, for example, before, during, or in overlapping time periods of the second decoding.
[0095] The various methods and other aspects described in this application can be used to modify modules of the video encoder 200 and decoder 300 (e.g., intra-frame prediction and entropy coding and / or decoding modules (260, 360, 245, 330)), such as Figure 2 and Figure 3 As shown. Furthermore, the subject matter disclosed herein presents aspects not limited to VVC or HEVC and can be applied to, for example, any type, format, or version of video coding (whether described in standards or recommendations, whether pre-existing or future-developed), and any extensions to such standards and recommendations (e.g., including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.
[0096] Various numerical values, such as weighting factors (e.g., {1 / 4, 1 / 8, 1 / 16, 1 / 32} or {3 / 4, 7 / 8, 15 / 16, 31 / 32}), filters (e.g., 3-tap filters [-1, 0, 1]), etc., are used in the examples describing this application. These and other specific values are used for the purposes of describing the examples, and the aspects described are not limited to these specific values.
[0097] Figure 2 This is a schematic diagram illustrating an exemplary video encoder. Variations of the exemplary encoder 200 are contemplated, but encoder 200 is described below for clarity, without describing all anticipated variations.
[0098] Before encoding, the video sequence may undergo pre-coding processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a signal distribution more resilient to compression (e.g., histogram equalization using one of the color components). Metadata may be associated with pre-processing and appended to the bitstream.
[0099] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units such as coding units (CUs). For example, each unit is encoded using either an intra-frame mode or an inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame mode or inter-frame mode is used to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0100] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.
[0101] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).
[0102] Figure 3 This is a schematic diagram illustrating an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 generally performs operations similar to... Figure 2 The encoding process is the reverse of the decoding process. Encoder 200 may also perform video decoding as part of the encoding of video data. For example, encoder 200 may perform one or more video decoding steps as presented herein. The encoder may, for example, reconstruct the decoded image to maintain synchronization with the decoder relative to one or more of the following: a reference picture, entropy coding context, and other decoder-related state variables.
[0103] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image is partitioned. Therefore, the decoder can partition (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks. Prediction blocks (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).
[0104] The decoded image may also undergo post-decoding processing (385), such as inverse color transformation (e.g., a transformation from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (201). Post-decoding processing may utilize metadata derived in the pre-encoding process and signaled in the bitstream.
[0105] Figure 4 This is a schematic diagram illustrating an example of a system in which the various aspects and embodiments described herein may be implemented. System 400 may be embodied as a device that includes the various components described below and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various embodiments, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 400 is configured to implement one or more aspects described in this document.
[0106] System 400 includes at least one processor 410 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0107] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both of an encoding module and a decoding module. Alternatively, the encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software known to those skilled in the art.
[0108] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. According to various embodiments, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more items from various projects during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded or partially decoded video, bitstreams, matrices, variables, and intermediate or final results of processing equations, formulas, operations, and operational logic.
[0109] In some embodiments, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory may be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one implementation, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2) or VVC (Universal Video Coding, a new standard developed by the Joint Video Experts Group JVET).
[0110] As shown in block 445, inputs to the components of system 400 can be provided through various input devices. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 4 Other examples not shown include composite video.
[0111] In various embodiments, the input device of block 445 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) re-band-limiting the signal to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box implementation, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various implementations rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various implementations, the RF section includes an antenna.
[0112] Furthermore, USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 410. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, within a separate interface IC or within processor 410. Demodulated streams, error-corrected streams, and demultiplexed streams are provided to various processing elements, including, for example, processor 410 and encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.
[0113] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and data can be transferred between these components using a suitable connection arrangement 425 (e.g., internal buses known in the art, including inter-chip (I2C) buses, wiring, and printed circuit boards).
[0114] System 400 includes a communication interface 450 capable of communicating with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0115] In various implementations, wireless networks such as Wi-Fi networks, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), are used to stream or otherwise provide data to system 400. In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 460 in these implementations is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other cloud-based communications. Other implementations use a set-top box to provide streaming data to system 400, delivering data via an HDMI connection to input block 445. Still other implementations use an RF connection to input block 445 to provide streaming data to system 400. As mentioned above, various implementations provide data in a non-streaming manner. Furthermore, various implementations use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0116] System 400 can provide output signals to various output devices, including display 475, speaker 485, and other peripheral devices 495. Display 475 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 475 can be used in televisions, tablets, laptops, mobile phones, or other devices. Display 475 can also be integrated with other components (e.g., in a smartphone) or standalone (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 495 include one or more of a standalone digital video disc (or digital versatile disc, both terms being DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.
[0117] In various embodiments, control signals are transmitted between system 400 and display 475, speaker 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices are communicatively coupled to system 400 via dedicated connections through corresponding interfaces 470, 480, and 490. Alternatively, output devices can be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 may be integrated into a single unit with other components of system 400 in electronic devices such as televisions. In various embodiments, display interface 470 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0118] Alternatively, if the RF portion of input 445 is part of a separate set-top box, the display 475 and speaker 485 may be separate from one or more other components. In various embodiments where the display 475 and speaker 485 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0119] These implementations can be executed by computer software implemented by processor 410, by hardware, or by a combination of hardware and software. As a non-limiting example, these implementations can be implemented by one or more integrated circuits. Memory 420 can be of any type suitable for the technical environment, and as a non-limiting example, it can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 410 can be of any type suitable for the technical environment, and as a non-limiting example, it can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0120] Various specific implementations involve decoding. As used in this application, "decoding" may encompass all or part of a process performed, for example, on a received encoded sequence, to produce a final output suitable for display. In various implementations, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various implementations, such a process also or alternatively includes a process performed by a decoder of the various specific implementations described in this application, for example, decoding a block including the current sub-block based on sample values obtained for the first pixel, such as the motion vector (MV) of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and sample values of the second pixel adjacent to the first pixel.
[0121] As a further implementation, in one example, "decoding" refers only to entropy decoding; in another implementation, "decoding" refers only to differential decoding; and in yet another implementation, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" specifically refers to a subset of operations or broadly refers to a wider decoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.
[0122] Various specific implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream. In various implementations, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various implementations, such processes also or alternatively include processes performed by an encoder of the various specific implementations described herein, such as encoding a block comprising the current sub-block based on sample values obtained for the first pixel, such as the motion vector (MV) of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and sample values of a second pixel adjacent to the first pixel.
[0123] As a further example, in one implementation, "encoding" refers only to entropy encoding; in another implementation, "encoding" refers only to differential encoding; and in yet another implementation, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" specifically refers to a subset of operations or broadly refers to a wider encoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.
[0124] It should be noted that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.
[0125] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.
[0126] During the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate-distortion optimization is generally formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to solve the rate-distortion optimization problem. For example, these methods may be based on extensive testing of all encoding options (including all considered modes or values of encoding parameters) and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to reduce encoding complexity, particularly for calculating approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed residual signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.
[0127] The specific embodiments and aspects described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of specific embodiment (e.g., discussed only as a method), specific embodiments of the discussed features may be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. These methods may be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0128] The references to “an implementation,” “implementation,” “example,” or “a specific implementation,” or “specific implementation,” and their variations, mean that the specific features, structures, characteristics, etc., described in connection with the implementation are included in at least one implementation. Therefore, the appearance of the phrases “in an implementation,” “in an implementation,” “in an example,” or “in a specific implementation,” and any other variations appearing throughout this application, do not necessarily refer to the same implementation or example.
[0129] Additionally, this application may involve "determining" various types of information. Determining information may include, for example, one or more of estimated information, calculated information, predicted information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0130] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or more of these.
[0131] Furthermore, this application may relate to "receiving" various types of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or more. Moreover, "receiving" typically involves one or more of the following during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0132] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such phrases are intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as possible listed.
[0133] Moreover, as used herein, the term "signaling" refers to (among other things) instructing the corresponding decoder to do something. For example, in some implementations, the encoder signals (e.g., to the decoder) a weight index, etc. In this way, the same parameters are used at both the encoder and decoder sides in the implementation. Thus, for example, the encoder can transmit (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various implementations by avoiding the transmission of any actual function. It should be understood that signaling can be implemented in many ways. For example, in various implementations, information is signaled to the corresponding decoder using one or more syntax elements, flags, etc. Although the verb form of the term "signal" was mentioned above, the term "signal" can also be used as a noun in this document.
[0134] It will be apparent to those skilled in the art that the embodiments may produce various signals formatted to carry, for example, storable or transmissible information. The information may include, for example, instructions for performing a method or data generated by one of the embodiments. For example, the signal may be formatted to carry a bit stream of the embodiment. Such signals may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and using a modulated carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0135] Bidirectional motion compensation prediction (MCP) can be performed. MCP can be highly efficient in removing temporal redundancy, for example, by utilizing the temporal correlation between images. A dual prediction signal can be formed, for example, by combining two single prediction signals (e.g., using a weight value equal to 0.5). For example, combining single prediction signals may be suboptimal when illumination changes rapidly from one reference image to another. The prediction technique can compensate for changes in illumination over time, for example, by applying global or local weights and / or offset values to one or more (e.g., each) sample values in the reference images.
[0136] Coding modules can be extended and / or enhanced (e.g., associated with time prediction). Affine motion compensation can be used as an inter-frame coding tool.
[0137] This document describes a specific implementation using affine modes. Translational motion models can be applied to motion-compensated predictions. Many types of motion may exist (e.g., zooming in or out, rotation, perspective motion, and / or other irregular motion). Simplified affine transformation motion-compensated predictions can be applied. Flags for inter-frame coding CUs (e.g., per inter-frame coding CU) can be signaled, for example, to indicate whether a translational or affine motion model is applied to inter-frame prediction. Flags (e.g., if affine motion is used) can be signaled to indicate the number of parameters used in the affine motion model (e.g., four or six).
[0138] The affine motion model can be a four-parameter model. Two parameters can be used for translational movement (e.g., one parameter each in the horizontal and vertical directions). One parameter can be used for scaling motion. One parameter can be used for rotational motion. The horizontal scaling parameter can be equal to the vertical scaling parameter. The horizontal rotation parameter can be equal to the vertical rotation parameter. The four-parameter motion model can be encoded using two motion vectors (MV) as a pair (e.g., one) at two control point positions defined at the top left and top right corners of the current CU. Figure 5 An exemplary four-parameter affine pattern model of an affine block and the derivation of sub-block-level motion are shown. For example... Figure 5 As shown, the affine motion field of the block can be described by two control point motion vectors (V0, V1). Based on the control point motion, the motion field ( v x , v y This can be described, for example, according to Equation 1: (1) in( v 0x , v 0y ) can be the motion vector of the top-left control point, ( v 1x , v 1y () can be the motion vector of the upper right control point, such as Figure 5 As shown, and w It can be the width of the CU.
[0139] Affine motion models can be six-parameter models. Two parameters can be used for translational motion (e.g., one parameter each in the horizontal and vertical directions). Two parameters can be used for scaling motion (e.g., one parameter each in the horizontal and vertical directions). Two parameters can be used for rotational motion (e.g., one parameter each in the horizontal and vertical directions). The six-parameter motion model can be encoded using three MVs at three control points. Figure 6 An exemplary six-parameter affine pattern is shown, where V0, V1, and V2 are control points, and (MVx MV y ) is the motion vector of the sub-block centered at position (x, y). For example... Figure 6 As shown, control points for a six-parameter affine-coded CU can be defined at the top left, top right, and bottom left corners of the CU. Motion at the top left control point can be related to translational motion. Motion at the top right control point can be related to rotational and scaling motion in the horizontal direction. Motion at the bottom left control point can be related to rotational and scaling motion in the vertical direction. The rotational and scaling motion in the horizontal direction can differ from the motion in the vertical direction. The MV (Movement Value) of a sub-block (e.g., each sub-block) can be derived, for example, using the three MVs at the control points according to Equations 2 and 3. v x , v y ): (2) (3) in( v 2x , v 2y ) can be the motion vector of the lower left control point, ( x , y ) can be the center position of the sub-block, and w and h These can be the width and height of the CU, respectively.
[0140] The motion field for blocks encoded using an affine motion model can be derived, for example, at the granularity of sub-blocks. This can be achieved, for example, by calculating the MV of the center samples of the sub-blocks (e.g., as shown in the image). Figure 5 The MV (value) of each sub-block is derived (e.g., according to equation (1)). The calculation can be rounded to, for example, 1 / 16-pel accuracy. The derived MV can be used in the motion compensation stage to generate predicted signals for sub-blocks (e.g., each sub-block) within the current block. The size of the sub-blocks applied to affine motion compensation can be, for example, 4×4. The four parameters of the 4-parameter affine model can be estimated iteratively, for example, in step [step]. k One or more MVs at a location can be represented as { , The original illuminance signal can be represented as: The predicted illuminance signal can be represented as: Spatial gradient and It can be used, for example, in the horizontal and vertical directions to predict signals. The derivation of the Sobel filter on the given surface. The derivative of equation (1) can be expressed, for example, according to equation 4: (4) In step k, (a, b) can be Δ translation parameters, and (c, d) can be Δ scaling and rotation parameters. For example, according to equations 5 and 6, ΔMV at the control points can be derived using coordinates. For example, (0, 0) and (w, 0) can be the coordinates of the upper left and upper right control points, respectively.
[0141] (5) (6) The relationship between illuminance variation and spatial gradient and time shift can be adjusted, for example, according to Equation 7: (7) in and The values in equation (4) can be substituted, for example, to obtain an equation for the parameters (a, b, c, d), as shown in equation (8): (8) The parameter set (a, b, c, d) can be derived, for example, using the least squares method (e.g., because the samples in CU satisfy Equation 8). The MV at the two control points in step (k+1) { , Equations 5 and 6 can be used to solve the problem, and they are rounded to a specific precision (e.g., 1 / 4 pel). The MV at the two control points can be modified (e.g., using iteration) until the parameters (a, b, c, d) are (e.g., all) zero or the number of iterations has been performed (e.g., a predefined limit).
[0142] The six parameters of the six-parameter affine model can be estimated. Equation 4 can be modified, for example, according to Equation 9: (9) In step k, (a, b) can be Δ translation parameters, (c, d) can be Δ scaling and rotation parameters in the horizontal direction, and (e, f) can be Δ scaling and rotation parameters in the vertical direction. For example, Equation 8 can be modified according to Equation 10: (10) The parameter set (a, b, c, d, e, f) can be derived, for example, by considering samples within the CU (e.g., multiple samples) using least squares. The MV of the upper left control point can be calculated using equation (5). For example, the MV of the upper right control point can be calculated using equations 11 and 12. MV of the lower left control point : (11) (12) Decoder-side motion vector correction (DMVR) can be provided. For example, two-sided matching (BM) based DMVR can be applied to improve the accuracy of MV in merged patterns. In dual prediction operations, a corrected MV can be searched around the initial MV in reference image list L0 and / or reference image list L1. BM-based DMVR can compute the distortion between two candidate blocks in reference image list L0 and list L1. Figure 7 An exemplary decoding-side motion vector (MV) correction is illustrated. Figure 7 As shown, the sum of absolute differences (SAD) between juxtaposed blocks can be calculated, for example, based on one or more (e.g., each) MV candidates surrounding the initial MV. The MV candidate with the lowest SAD can become a modified MV and can be used to generate a double prediction signal.
[0143] The modified MV derived from DMVR can be used, for example, to generate inter-frame prediction samples. The modified MV derived from DMVR can be used for temporal motion vector prediction, for example, for future image encoding. The initial MV can be used for deblocking and / or spatial motion vector prediction for future CU encoding, for example, to avoid any MV dependencies between the current CU and neighboring CUs.
[0144] like Figure 7 As shown, the search points around the initial MV and MV offset can follow the MV difference mirroring (e.g., symmetric) rule. The points checked by DMVR, represented by the candidate MV pair (MV0, MV1), can be determined according to Equations 13 and / or 14: This can represent the corrected offset between the initial MV and the corrected MV in one of the reference images. The corrected search range can be, for example, two integer illuminance samples from the initial MV. Fast search methods with early termination mechanisms can be applied, for example, to reduce search complexity.
[0145] Sub-block-based temporal motion vector prediction (SbTMVP) is available. SbTMVP can use motion fields in juxtaposed images to improve motion vector prediction and merging patterns of CUs in the current image. The same juxtaposed image used by temporal motion vector prediction (TMVP) can be used for SbTMVP. SbTMVP may differ from TMVP in one or more of the following ways: TMVP predicts motion at the CU level. SbTMVP predicts motion at the sub-CU level. TMVP obtains temporal motion vectors from juxtaposed blocks in the juxtaposed image. The juxtaposed block can be the lower right or center block relative to the current CU. SbTMVP may apply a motion offset, for example, before obtaining temporal motion information from the juxtaposed image. The motion offset may be obtained from, for example, the motion vector of a spatially neighboring block from the spatially neighboring blocks of the current CU.
[0146] Figure 8A and Figure 8B An exemplary SbTMVP process is shown. Figure 8A An example spatial neighbor block that can be used in SbTMVP is shown. Figure 8B An exemplary derivation of the motion field of a sub-coding unit (CU) is shown. For example... Figure 8B As shown, the motion field of a sub-CU can be derived by applying motion offsets from spatial neighbors and scaling motion information from the corresponding juxtaposed sub-CU. SbTMVP can predict the motion vectors of sub-CUs within the current CU. Spatial neighbors can be checked (e.g., Figure 8A (A1 in the image). For example, if A1 has motion vectors that use the juxtaposed image as its reference image, then the motion vector of A1 can be selected (e.g., for the motion offset to be applied). For example, if no motion is identified, the motion offset can be selected as (0, 0). The selected motion offset can be applied, for example, to obtain sub-CU level motion information from the juxtaposed image (e.g., such as motion vectors and / or reference indices). For example, the selected motion offset can be added to the coordinates of the current block. The motion offset can be set as the motion of block A1 (e.g., in the image). Figure 8B (In the illustrated example). Motion information of the corresponding blocks (e.g., the smallest motion grid covering the center sample) of each sub-CU in the juxtaposed image can be used to derive the motion information of the sub-CU. The motion information of the juxtaposed sub-CU (e.g., after being identified) can be converted into the motion vector and reference index of the current sub-CU. For example, temporal motion scaling can be applied to align the reference image of the temporal motion vector with the reference image of the current CU.
[0147] A merge list based on combined sub-blocks, including SbTMVP candidates and affine merge candidates, can be used for signaling in sub-block-based merge modes. The SbTMVP mode can be enabled and / or disabled by the Sequence Parameter Set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor can be added as (e.g., the first) entry to the list of sub-block-based merge candidates, followed by affine merge candidates. The size of the sub-block-based merge list can be signaled in the SPS. For example, the maximum allowed size of the sub-block-based merge list can be 5.
[0148] The size of the sub-CU used in SbTMVP can be fixed, for example, it can be 8×8. The SbTMVP pattern can (for example, only) be applied to CUs with a width and height greater than or equal to 8.
[0149] The encoding logic for additional SbTMVP merge candidates can be the same as that used for other merge candidates. For example, for each CU in a P or B slice, an additional RD check can be performed. The additional RD check can be used to determine whether to use an SbTMVP candidate.
[0150] Overlapping Block Motion Compensation (OBMC) can be provided. OBMC can be enabled and disabled, for example, at the CU level using syntax. OBMC can be performed against motion compensation (MC) block boundaries (e.g., except for the right and bottom boundaries of the CU). OBMC can be applied to illuminance and chromaticity components. MC blocks can correspond to encoded blocks. CUs can be encoded using subCU modes (e.g., including subCU merging, affine, and FRUC modes). One or more sub-blocks of a CU encoded using subCU modes (e.g., each sub-block) can be MC blocks. OBMC can be performed at the sub-block level at (e.g., all) MC block boundaries, for example, to (e.g., consistently) process CU boundaries. Figure 9 An example of applying OBMC is shown. The sub-block size can be set to equal 4×4, for example, as shown below. Figure 9 As shown.
[0151] OBMC can be applied to the current sub-block. Four motion vectors connecting neighboring sub-blocks (e.g., in addition to the current motion vector) (e.g., if available and different from the current motion vector) can be used to derive the prediction block for the current sub-block. In one or more examples, "neighboring" can be used interchangeably with "adjacent". Multiple prediction blocks based on multiple motion vectors can be combined, for example, to generate the final prediction signal for the current sub-block.
[0152] The predicted block based on the motion vectors of neighboring sub-blocks can be represented as: P N ,in NIndicates the indices of neighboring sub-blocks above, below, to the left, and to the right of the current sub-block. A predicted block based on the motion vector of the current sub-block can be represented as... P C For example, if P N If the motion information is based on neighboring sub-blocks that include the same motion information as the current sub-block, then the OBMC can be skipped. Otherwise, it can be... P N One or more (e.g., each) samples added P C The same sample in the data, for example, can be four rows / four columns. P N Add to P C In the example, the weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} can be used. P N Furthermore, the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} can be used... P C For example, if the height or width of the coded block is equal to four (4) or the CU is encoded using a subCU mode, then for a small MC block, two rows and / or two columns can be used. P N Add to P C For small MC blocks, the weighting factors {1 / 4, 1 / 8} can be used. P N Furthermore, the weighting factors {3 / 4, 7 / 8} can be used... P C For motion vectors generated based on vertically (e.g., and / or horizontally) neighboring sub-blocks P N This can be done in the same row (e.g., and / or column). P N Samples from the same dataset are added to samples with the same weighting factor. P C For example, overlapping region pixels can use a different weighting factor than that used for non-overlapping region pixels.
[0153] A signal can be sent to a CU-level flag to indicate whether OBMC is applied to the current CU, such as a CU with a size less than or equal to 256 illuminance samples. For example, OBMC can be applied by default for CUs with a size greater than 256 illuminance samples or those not encoded using the AMVP pattern. For example, the effects of OMMC can be taken into account at the encoder during the motion estimation phase. The predicted signal formed by OBMC using motion information from the top neighbor block and left neighbor block can be used to compensate for the top and left boundaries of the original signal of the current CU. The motion estimation process can be applied (e.g., subsequently) (e.g., in other ways) (e.g., normally).
[0154] Optical flow (PROF) prediction correction can be applied to affine modes. PROF can use optical flow to correct sub-block-based affine motion compensation predictions, for example, to achieve finer-grained motion compensation. For example, a correction can be made to the accuracy prediction sample by adding a difference derived through the optical flow equation (e.g., after sub-block-based affine motion compensation). PROF can include one or more of the following: Sub-block-based affine motion compensation can be performed to generate sub-block predictions. It can compute the spatial gradient of sub-block predictions at one or more sample locations (e.g., each sample location). and For example, a spatial gradient can be computed using one or more pixels that may or may not be partially or completely contiguous. In the example, the gradient of the first pixel can be based on sample values of a second pixel and a third pixel, where the second and third pixels are adjacent to the first pixel. In some examples, the first pixel for calculating the gradient may be adjacent to either or both of the second or third pixels adjacent to the first pixel. In other examples, the first pixel for calculating the gradient may be close to but not adjacent to either or both of the second or third pixels adjacent to the first pixel. The computation can be performed using a 3-tap filter such as [-1, 0, 1], for example, as shown in Equations 15 and 16: (15) (16) For gradient computation, sub-block predictions can be expanded (e.g., by expanding by one pixel on each side). For example, pixels on the expanded boundaries can be copied from the nearest integer pixel position in the reference image. For example, if pixels on the expanded boundaries are copied from the nearest integer pixel position in the reference image, additional interpolation of the filled regions can be avoided. Figure 10 An exemplary sub-block MV V is depicted SB and pixels Illuminance prediction corrections can be calculated using optical flow equations, for example, as shown in Equation 17: (17) in It is for sample location Calculated in pixel MV (by (representation) and pixels The difference between the MV values of the sub-blocks of the same sub-block, such as Figure 10 As shown.
[0155] The affine model parameters and pixel positions relative to the center of the sub-block may not change between sub-blocks. It can be computed for (e.g., the first) sub-block and can be reused for other sub-blocks (e.g., within the same CU). Let... x and y It refers to the horizontal and vertical offsets from the pixel position to the center of the sub-block, which can be derived, for example, from Equation 18. : (18) For the 4-parameter affine model, c and e It can be determined according to Equation 19: (19) For the 6-parameter affine model, c , d , e and f It can be determined according to Equation 20: (20) And among them , , These are the motion vectors of the upper left control point, the upper right control point, and the lower left control point, respectively. w and h These are the width and height of the CU. Add illumination prediction corrections to the sub-block prediction. Final prediction I’ It can be generated, for example, according to Equation 21: (twenty one) DMVR and SbTMVP can be used in different prediction modes to improve the accuracy of predicted MV. The corrected MV following DMVR or SbTMVP can be used (e.g., only) to perform sub-block-based motion compensation afterwards. OBMC can include pixel-level corrections. OBMC can be used to reduce boundary discontinuities at sub-blocks of a CU or sub-CU. OBMC can include multiple motion compensation operations for one or more (e.g., each) sub-blocks. For example, if the MVs of four connected neighboring sub-blocks are available and different from the MV of the current sub-block, they can be used to derive the predicted block for the current sub-block.
[0156] Methods for sub-block / block correction (e.g., pixel-level correction) can be provided. For example, methods can be used to reduce boundary discontinuities. Figure 11 Examples of methods are provided. Figure 11 The method described can be applied to decoders and / or encoders.
[0157] Figure 11 Examples of methods for sub-block / block correction based on one or more of equations (1) to (25) are shown. Examples disclosed herein, as well as other examples, can be found in... Figure 11 The exemplary method 1100 shown operates. Method 1100 includes 1102 and 1104. In 1102, a sample value of a first pixel can be obtained based on, for example, the following: (1) the motion vector (MV) of the current sub-block, (2) the MV of the sub-blocks adjacent to the current sub-block, and (3) the sample value of a second pixel adjacent to the first pixel. In 1104, a block including the current sub-block can be encoded or decoded based on the obtained sample value of the first pixel. When such... Figure 11 When the method is applied to the decoder, Figure 11 1104 in the code can be performed by the decoder, and 1104 may need to decode the block including the current sub-block based on the obtained sample value of the first pixel. When... Figure 11 When the method is applied to an encoder, Figure 11 1104 in the code can be performed by the encoder, and 1104 may need to encode the block including the current sub-block based on the obtained sample value of the first pixel.
[0158] Examples of methods for sub-block / block correction are provided, for example, for encoding and decoding. Examples may refer to "boundaries" that include different types of boundaries, such as the boundaries of blocks, sub-blocks, CUs, and / or PUs. Examples may refer to "neighbors" that may include different types of neighbors, such as spatial and temporal neighbors of blocks, sub-blocks, CUs, and / or PUs. Examples may refer to "adjacent" that includes different types of neighbors, such as adjacent blocks, adjacent sub-blocks, adjacent pixels, and / or pixels adjacent to a boundary. Spatial neighbors can be adjacent in the same frame, while temporal neighbors can be located at the same position in adjacent frames. For example, an adjacent sub-block is a sub-block that can be either a spatial or temporal neighbor. A boundary pixel is a pixel adjacent to a boundary, where the boundary can be any type of boundary. For example, a boundary pixel can be adjacent to the boundary of a block, sub-block, CU, and / or PU.
[0159] Sub-block / block corrections may include sub-block / block boundary corrections. For example, the MV difference between the current block and / or sub-block and neighboring blocks and / or sub-blocks may be calculated and converted into a difference of sample values derived from the optical flow equation. The pixel intensity (e.g., illuminance and / or chromaticity) of the boundary pixels of the current block and / or sub-block may be corrected, for example, by adding the derived difference. The derived difference may be referred to as Block Boundary Optical Flow Prediction Correction (BBPROF). The sample value offset of the boundary pixels may indicate the derived difference. The sample values of the boundary pixels may indicate the pixel intensity of the boundary pixels. BBPROF (e.g., as described herein) may provide pixel-level granularity for sub-block and block boundary corrections. BBPROF (e.g., as described herein) may be applied to any sub-block-based inter-frame prediction mode and / or CU-based inter-frame prediction mode.
[0160] Boundary pixels can include pixels located at the boundaries of blocks and / or sub-blocks. For example, a square sub-block may have four boundaries, including a left boundary, a right boundary, a top boundary, and a bottom boundary. Boundaries can include common boundaries shared between two sub-blocks. Two sub-blocks may be adjacent to each other at their boundaries. A pixel may be located at a boundary when it is near a boundary. For example, a pixel may be located at a boundary when it is in a row of pixels (e.g., 4, 3, 2, or 1) starting from the top boundary of a sub-block, in a row of pixels starting from the bottom boundary of a sub-block, in a column of pixels (e.g., 4, 3, 2, or 1) starting from the left boundary of a sub-block, or in a column of pixels starting from the right boundary of a sub-block. In some examples, when a first pixel is located inside a first sub-block, the first pixel may be a pixel in the first sub-block or located at the common boundary of the first sub-block and adjacent to a pixel inside a second sub-block that shares the common boundary with the first sub-block. In some examples, a pixel may be located at the boundary of a sub-block, but outside that sub-block.
[0161] BBPROF can be applied, for example, in DMVR mode. BBPROF can reduce block boundary discontinuities in DMVR-based sub-block-level motion compensation predictions. Pixel intensity variations can be applied by BBPROF. Pixel intensity variations can be derived, for example, from optical flow equations. Pixel sample value offsets can indicate pixel intensity variations. BBPROF can be used to perform one (e.g., only one) motion compensation operation per sub-block. Motion compensation in DMVR mode can perform one motion compensation operation per sub-block.
[0162] The corrected motion vectors of sub-blocks in the CU can be derived, for example, by performing DMVR (e.g., as described herein). Sub-block-based motion compensation (e.g., as described herein) can be performed to generate sub-block-based predictions.
[0163] The spatial gradient of the sub-block prediction can be computed at one or more (e.g., each) pixel / sample locations (e.g., as described herein). and .
[0164] It can calculate the motion vector difference between the current sub-block and one or more neighboring sub-blocks (considered as neighboring sub-blocks). The MV difference can be calculated at the sub-block level. Each candidate neighboring sub-block that is not far from (e.g., close to or near) the current sub-block can be considered. Sub-blocks with various amounts and / or locations can be selected as neighboring sub-blocks for calculation. In the example, BBPROF in DMVR mode can be calculated using four neighboring sub-blocks (e.g., left, top, right, and bottom neighboring sub-blocks), two neighboring sub-blocks (e.g., top and left neighboring sub-blocks), corner neighboring sub-blocks (e.g., top left and bottom right), or other amounts and positions of neighboring sub-blocks. .
[0165] In the example, four neighboring child blocks can be considered (e.g., the upper, lower, left, and right child blocks adjacent to the current child block), where It can be a set of MV differences that includes four different MV differences. For example, It can be calculated as , where A, B, L, and R can represent the MV difference between the current sub-block and the upper, left, and right sub-blocks, respectively.
[0166] Figure 12 An exemplary MV difference calculation from selected neighboring sub-blocks is depicted, for example, in DMVR mode. Figure 12 As shown, a sub-block can have its own MV difference after DMVR. The current sub-block within the CU (e.g., Figure 12 The current DMVR sub-block in the current DMVR can have four connected neighboring sub-blocks, except for the current sub-block located at the boundary.
[0167] Calculated sub-block level MV difference It can be used to calculate the motion vector offset at one or more (e.g., each) pixel / sample locations within the current sub-block. For example, as shown in Equation 21.
[0168] (twenty one) Where n can be the index of a specific neighboring sub-block, and N can be the total number of neighboring sub-blocks considered. It can be in the vicinity A weighting factor applied to a specific pixel at position (i, j). For example, if left, top, right, and bottom neighboring sub-blocks are considered, N can be equal to 4.
[0169] The set of weighting factors can be, for example, {1 / 4, 1 / 8, 1 / 16, 1 / 32}. Each weighting factor can be used from four rows / columns of pixels on one or more (e.g., each) sides of the current sub-block. The MV difference can be calculated, for example, based on the motion vectors of vertically and / or horizontally neighboring sub-blocks. Pixels in the same row and / or column of the current sub-block can use the same weighting factor. For example, pixels in the first column to the left of the current sub-block can use the same weighting factor (e.g., 1 / 4), and pixels in the second column can use the same weighting factor (e.g., 1 / 8), and so on. The weighting factor can be determined, for example, based on the distance from the current position to the block boundary between the current block and its neighboring blocks. For example, the weighting factor can be smaller when the column and / or row is farther from the block boundary.
[0170] The weighting factor can be adjusted, for example, based on pixel location (e.g., dynamically). In the example, the pixel might be located to the upper left of the current sub-block. The MV differences from the left and upper neighboring sub-blocks can be weighted / combined together to generate the final MV offset at the pixel, for example, if left and upper neighboring sub-blocks exist and if the MV differences from the left and upper neighbors are not zero. For example, if the current sub-block is on the left or top boundary of the current CU, the left or upper neighboring sub-blocks might be unavailable.
[0171] For example, the intensity change of each pixel within the current sub-block can be calculated using optical flow equation 22: (twenty two) in and This can be at one or more (e.g., each) sample locations. The MV offset and spatial gradient at that location can be calculated, for example, in a previous step.
[0172] The prediction of a pixel or sample location can be corrected, for example, by adding the calculated intensity changes (e.g., illuminance or chromaticity) to the sub-block prediction. The corrected prediction of the pixel or sample location can be associated with a list of reference images (e.g., list L0 or list L1). Final prediction I’ It can be generated, for example, according to Equation 23: (twenty three) When applying BBPROF in DMVR mode, up to four neighboring child blocks can be considered. The inner child block can wait for the DMVR process of its neighboring child blocks to complete. In the example, when applying BBPROF in DMVR mode, two neighboring child blocks can be considered. For example, two neighbors can be considered (e.g., only the top and left neighbors), such that the BBPROF of the current child block depends on the DMVR process of the two neighboring child blocks.
[0173] BBPROF can be applied in SbTMVP mode. For example, BBPROF can reduce discontinuities at subblock boundaries in SbTMVP-based subblock-level motion compensation predictions. One or more examples in this paper for applying BBPROF in DMVR mode can be applied to performing BBPROF in SbTMVP mode. Applying BBPROF in SbTMVP mode may include performing one (e.g., only one) motion compensation operation per subblock. SbTMVP motion compensation may perform one motion compensation operation per subblock. Applying BBPROF in SbTMVP mode may include one or more of the following.
[0174] For example, modified motion vectors for one or more (e.g., each) sub-blocks in a CU can be derived by performing SbTMVP (e.g., as described herein). Motion information can be obtained from the juxtaposed sub-CUs. Appropriate time scaling can be applied to the motion information. For example, sub-block-based motion compensation can be performed to generate sub-block-based predictions.
[0175] The spatial gradient of the sub-block prediction can be computed at one or more (e.g., all) pixel / sample locations (e.g., as described in this paper). and .
[0176] The motion vector difference between the current sub-block and, for example, one or more considered neighboring sub-blocks can be calculated. .
[0177] The prediction corrections described in this article can be applied to ATMVP. In the example, the prediction corrections can use... Figure 9 The image shows the motion vectors of the four neighboring sub-blocks of the current sub-block. These motion vectors can be used to derive the MV difference between the current sub-block and its spatial neighbors. .
[0178] Calculated sub-block level MV difference It can be used to calculate the motion vector offset at one or more (e.g., each) pixel / sample locations within the current sub-block. .
[0179] The intensity change of each pixel within the current sub-block can be calculated using optical flow equations (e.g., based on Equation 22).
[0180] The prediction for one or more (e.g., each) reference images in the reference image list can be corrected, for example, by adding intensity variations (e.g., illuminance or chromaticity). Final prediction I’ It can be generated, for example, according to Equation 24: (twenty four) BBPROF can be applied in affine patterns. BBPROF can be applied to affine coding units (e.g., similar to SbTMVP). An affine coding unit can include multiple sub-blocks. The block-level motion vectors (MVs) of one or more (e.g., each) sub-blocks can be derived using an affine motion model (e.g., as described herein). For example, the four parameters of a 4-parameter affine model and / or the six parameters of a 6-parameter affine model can be estimated using two or three control point motion vectors. For example, the block-level motion vectors of sub-blocks within an affine coding unit can be derived using four or six estimated affine model parameters. The sub-block-level MV differences can be calculated, for example, using the motion vectors of different sub-blocks. And / or motion vector offset at one or more (e.g., each) pixel / sample locations within a sub-block. For example, the sub-block level MV difference can be calculated. and / or motion vector offset (For example, BBPROF can be performed as described in this article).
[0181] BBPROF can be applied to pixels adjacent to the boundaries of a CU, for example, as described herein for pixels adjacent to the boundaries of a sub-block. For example, BBPROF can be applied at the CU level. For example, the reference images can be the same or different for adjacent CUs.
[0182] For example, if the selected neighboring CUs (e.g., the upper, lower, left, and right CUs) have the same reference picture as the current CU, then the CU-level MV difference can be calculated (e.g., directly calculated) for the specific CU. The prediction of the boundary pixels of a particular CU can be corrected, for example, by applying BBPROF (e.g., directly).
[0183] For example, if (i) one or more (e.g., all) of the selected neighboring CUs (e.g., top, bottom, left, and right CUs) and (ii) the current CU has different reference images from the reference image list, then temporal motion scaling can be applied to a specific CU. Appropriate temporal motion scaling aligns the reference images of the temporal motion vectors of the selected neighboring CUs with the reference image of the specific CU. CU-level MV differences can be calculated. (e.g., based on the scaled MV of the selected neighboring CUs), and prediction corrections at the CU boundaries can be achieved, for example, by applying BBPROF.
[0184] Several specific implementation variations of BBPROF are provided as additional examples. The amount and location of neighboring (e.g., adjacent) sub-blocks that can be selected for sub-block / block corrections such as sub-block / block boundary corrections (e.g., BBPROF) are not limited to the examples described herein, such as references Figure 12The example described above. Other numbers and / or locations of sub-blocks can be selected. Each candidate neighboring sub-block that is not far from (e.g., close to or near) the current sub-block can be considered. For example, BBPROF can use corner neighboring sub-blocks (e.g., top left, bottom right). Figure 12 The calculation is performed using the four sub-blocks shown or other neighboring sub-blocks of varying numbers and positions. .
[0185] The location of the considered neighboring sub-blocks is not limited to being within the same CU as the current sub-block. The aspect ratio of the considered neighboring sub-blocks from adjacent CUs may or may not be the same as that of the current sub-block. Sub-block / block corrections such as sub-block / block boundary corrections (e.g., BBPROF) can allow for different aspect ratios.
[0186] The MV offset at a pixel / sample location can be derived, for example, based on the sub-block MV difference. The number of rows and / or columns of pixels on one or more (e.g., each) sides of the current sub-block can be configurable and / or dynamically changed, for example, based on one or more (e.g., predefined) criteria. For example, two or more pixel columns on the left side of the current sub-block may include the MV difference with the left neighboring sub-block, for example, instead of the default number of pixel columns, such as, for example, four columns of pixels.
[0187] MV difference can be based on vertically and / or horizontally adjacent sub-blocks. In the example, pixels in the same row and / or column of the current sub-block can use the same weighting factor. In the example, pixels in the same row and / or column of the current sub-block can use different weighting factors.
[0188] The weighting factor can vary. For example, the weighting factor can increase as the spatial distance between the pixel and the vertical or horizontal boundary decreases.
[0189] The intensity difference (e.g., derived from Equation 22) can be multiplied by a weighting factor w, for example, before adding the intensity difference to the prediction, as shown in Equation 25, for example: (25) in w It can be set to a value between 0 and 1, including end values. Signaling can be performed, for example, at the CU level or image level. w For example, a signal can be sent via a weighted index. w Equation 25 can be a variation of Equation 23 and / or Equation 24.
[0190] BBPROF can be used, for example, after combining L0 and L1 predictions based on DMVR with weights. BBPROF can be applied to, for example, a prediction, such as L0 or L1, to reduce complexity. In one example, BBPROF can be applied to a prediction where, for example, the reference image is closer to the current image in the time domain. In another example, BBPROF can be applied to a prediction where, for example, the reference image is farther from the current image in the time domain.
[0191] This document describes numerous embodiments. Features of the embodiments may be provided individually or in any combination across various claim classes and types. Furthermore, embodiments may include one or more of the features, devices, or aspects described individually or in any combination across various claim classes and types, such as, for example, any of the following.
[0192] like Figure 11 The method described herein can be applied to decoders and / or encoders. When, as Figure 11 When the method is applied to an encoder, Figure 11 1104 in the code can be performed by the encoder, and 1104 may need to encode the block including the current sub-block based on the obtained sample value of the first pixel. For example... Figure 11 The method described can be based on one or more of equations (1)-(25). For example, the decoder can decode the current sub-block based on the sample values of the pixel. The pixel may be located at one of the boundaries of the current sub-block. The sample values can be modified sample values obtained based on one or more of equations (1)-(25). As shown in one or more of equations (1)-(25), the decoder can obtain the sample values of the pixel based on, for example, the MV of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and the sample values of the pixels adjacent to the pixel whose sample values are obtained. The decoder can obtain the prediction of the sample values of the pixel, for example, before the decoder modifies the prediction of the sample values. The prediction of the sample values of the pixel may be referred to as For example, as shown in equation (23). As shown in equation (23), the decoder can obtain the sample value of a pixel based on the sample value offset and the prediction of the sample value (e.g., the sum of the sample value offset and the prediction of the sample value). The sample value offset of a pixel can be called... For example, as shown in equation (23). As shown in one or more equations (1)-(25), the sample value offset can be obtained based on, for example, the MV of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and the sample value of the pixel adjacent to the pixel whose sample value was obtained. The decoder can use the MV of the current sub-block and the MV of the sub-blocks adjacent to the current sub-block to obtain the MV difference (e.g., the MV difference in equation (21), as described herein. The MV difference may be referred to as For example, as shown in Equation 21. The decoder can obtain the MV difference using one or more MVs associated with one or more corresponding sub-blocks adjacent to the current sub-block. The decoder can obtain the gradient using the sample values of pixels adjacent to the pixel whose sample value is obtained, for example, as shown in Equations (15) and (16). As in the examples shown in Equations (15) and (16), the decoder can obtain the gradient using one or more sample values of one or more corresponding pixels adjacent to the pixel whose sample value is obtained. The decoder can obtain the sample value of the pixel based on the gradient and the MV difference. In the example, the decoder can obtain the MV offset based on the MV difference, as shown in Equation (21). The decoder can obtain the sample value of the pixel using the MV offset and the gradient, as shown in Equations (22) and (23). The decoder can obtain the sample value offset based on the gradient and the MV difference, and use the sample value offset to obtain the sample value of the pixel. The decoder can determine a weighting factor and use the weighting factor to obtain the sample value of the pixel, for example, as shown in Equation (21). The decoder can decode blocks, including the current sub-block, based on the sample values of the obtained pixels.
[0193] Decoding tools and techniques, including one or more of entropy decoding, inverse quantization, inverse transform, and differential decoding, can be used to implement [the following] in the decoder. Figure 11 The methods described above. These decoding tools and techniques can be used to determine, for example,... Figure 11 The method described above can be used to implement one or more of the sub-block / block modifications; according to, for example Figure 11 The method described for sub-block / block boundary correction; according to, for example Figure 11 The method described includes BBPROF; sub-block / block correction in DMVR mode; sub-block / block correction in SbTMVP mode; sub-block / block correction in affine mode; and according to... Figure 11 The method described above obtains sample values; according to, for example Figure 11 The method described herein obtains sample value offsets; obtains gradients as described herein; obtains MV differences as described herein; obtains predictions of sample values; and other decoder behaviors related to any of the above.
[0194] The encoder can encode the current sub-block based on the sample values of pixels. Pixels can be located at one of the boundaries of the current sub-block. The sample values can be modified sample values obtained based on one or more of equations (1)-(25). As shown in one or more of equations (1)-(25), the encoder can obtain the sample values of pixels based on, for example, the MV of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and the sample values of pixels adjacent to the pixel whose sample values are obtained. The encoder can obtain the prediction of the sample values of pixels, for example, before the encoder modifies the prediction of the sample values. The prediction of the sample values of pixels can be called... For example, as shown in equation (23). As shown in equation (23), the encoder can obtain the sample value of a pixel based on the sample value offset and the prediction of the sample value (e.g., the sum of the sample value offset and the prediction of the sample value). The sample value offset of a pixel can be called... For example, as shown in equation (23). As shown in one or more equations (1)-(25), the sample value offset can be obtained based on, for example, the MV of the current sub-block, the MV of the sub-blocks adjacent to the current sub-block, and the sample value of the pixel adjacent to the pixel whose sample value was obtained. The encoder can use the MV of the current sub-block and the MV of the sub-blocks adjacent to the current sub-block to obtain the MV difference (e.g., the MV difference in equation (21), as described herein. The MV difference may be referred to as For example, as shown in equation (21). The encoder can obtain MV differences using one or more MVs associated with one or more corresponding sub-blocks adjacent to the current sub-block. As shown in equations (15) and (16), the encoder can obtain gradients using sample values of pixels adjacent to the pixel whose sample value is obtained. As in the examples shown in equations (15) and (16), the encoder can obtain gradients using one or more sample values of one or more corresponding pixels adjacent to the pixel whose sample value is obtained. The encoder can obtain sample values of pixels based on gradients and MV differences. In the example, the encoder can obtain MV offsets based on MV differences, as shown in equation (21). The encoder can obtain sample values of pixels using MV offsets and gradients, as shown in equations (22) and (23). The encoder can obtain sample value offsets based on gradients and MV differences, and use the sample value offsets to obtain sample values of pixels. The encoder can determine weighting factors and use the weighting factors to obtain sample values of pixels, for example, as shown in equation (21). The encoder can encode a block, including the current sub-block, based on the sample values of the obtained pixels.
[0195] Encoding tools and techniques, including one or more of quantization, entropy coding, inverse quantization, inverse transform, and differential coding, can be used to implement, in the encoder, such as Figure 11 The methods described above. These encoding tools and techniques can be used according to, for example... Figure 11 The method described herein implements one or more of the sub-block / block modifications; according to, for example Figure 11 The method described for sub-block / block boundary correction; according to, for example Figure 11 The method described includes BBPROF; sub-block / block correction in DMVR mode; sub-block / block correction in SbTMVP mode; sub-block / block correction in affine mode; and according to... Figure 11 The method described above obtains sample values; according to, for example Figure 11The method described herein obtains sample value offsets; obtains gradients as described herein; obtains MV differences as described herein; obtains predictions of sample values; and other encoder behaviors related to any of the above.
[0196] Syntax elements can be inserted into the signaling, for example, to enable the decoder to recognize and execute such... Figure 11 Indications associated with the method or method used. For example, syntax elements may include indications of one or more of BBPROF, DMVR, SbTMVP mode, affine mode, for example, to indicate to the decoder whether one or more of them are enabled or disabled. As an example, syntax elements may include indications of one or more weighting factors as described herein, and / or indications of parameters used by the decoder to perform one or more examples herein.
[0197] For example, the syntax elements applied at the decoder can be selected and / or applied, such as Figure 11 The method described above. For example, the decoder may receive an instruction to enable BBPROF. Based on this instruction, the decoder may perform actions such as... on pixels located at or near the boundaries of a sub-block. Figure 11 The method described.
[0198] The encoder can adjust the prediction residual based on one or more examples presented herein. For example, the residual can be obtained by subtracting the predicted video patch from the original image patch. For example, the encoder can predict the video patch based on sample values of pixels obtained as described herein. The encoder can obtain the original image patch and subtract the predicted video patch from the original image patch to generate the prediction residual.
[0199] A bitstream or signal may include one or more syntax elements or variations thereof. For example, a bitstream or signal may include syntax elements that indicate whether any of the following modes are enabled or disabled: BBPROF, DMVR, SbTMVP, or affine mode.
[0200] Bitstreams or signals may include syntax that conveys information generated according to one or more examples in this document. For example, in the execution of... Figure 11 The example shown generates information or data. The generated information or data can be conveyed in the syntax included in the bitstream or signal.
[0201] This allows the decoder to insert syntax elements into the signal that adapt the residuals in a manner corresponding to those used by the encoder. For example, one or more examples from this paper can be used to generate residuals.
[0202] A method, process, apparatus, medium for storing instructions, medium for storing data, or signal for creating and / or transmitting and / or receiving and / or decoding a bit stream or signal comprising one or more of the said syntax elements or variations thereof.
[0203] A method, process, apparatus, medium for storing instructions, medium for storing data, or signal for creating and / or transmitting and / or receiving and / or decoding according to any of the examples described.
[0204] A method, process, apparatus, medium storing instructions, medium storing data, or signal based on one or more of the following: determining a spatial gradient of a sub-block-based prediction at one or more pixel / sample locations; using MV differences to calculate motion vector offsets at one or more pixel / sample locations; for example, determining the intensity change of each pixel in the current sub-block based on optical flow; for example, correcting the prediction of a list of reference images by adding the calculated intensity changes to the sub-block prediction; determining that a first pixel is adjacent to the boundary of the current sub-block; determining the difference between the MV of the current sub-block and the MV of a sub-block adjacent to the current sub-block; determining the gradient of the first pixel based on sample values of a second pixel adjacent to the first pixel and sample values of a third pixel adjacent to the first pixel; determining a sample value offset based on the determined gradient and the difference between the MV of the current sub-block and the MV of the sub-block adjacent to the current sub-block; obtaining sample values of the first pixel based on the determined sample value offset; for example, determining the gradient based at least on sample values of the second pixel; using the gradient to determine the first pixel. The sample values are obtained by: determining the gradient of the optical flow model based on, for example, at least based on the sample values of the second pixel; using the gradient in the optical flow model to obtain the sample values of the first pixel; using the difference between the MV of the current sub-block and the MV of the sub-block adjacent to the current sub-block to obtain the sample values of the first pixel; further obtaining the sample values of the first pixel based on, for example, the MV of the second sub-block adjacent to the current sub-block; obtaining the sample values of the first pixel based on determining that the first pixel is adjacent to the boundary of the current sub-block; using a weighting factor to obtain the sample values of the first pixel, wherein the weighting factor may or may not vary based on the distance of the first pixel from the corresponding boundary of the current sub-block; determining the sample value offset of the first pixel based on, for example, the MV of the current sub-block, the MV of the sub-block adjacent to the current sub-block, and the sample values of the second pixel adjacent to the first pixel; using the determined sample value offset of the first pixel and the predicted sample values to obtain the sample values of the first pixel; and obtaining the sample values of the first pixel based on determining that the first pixel is adjacent to the boundary of the current sub-block, for example, wherein the boundary of the current sub-block may include the common boundary between the current sub-block and the sub-block adjacent to the current sub-block.
[0205] TVs, set-top boxes, mobile phones, tablets, or other electronic devices that perform block / sub-block / CU corrections according to any of the examples described.
[0206] A TV, set-top box, mobile phone, tablet computer, or other electronic device performs block / sub-block / CU corrections according to any of the examples described and displays the resulting image (e.g., using a monitor, screen, or other type of display).
[0207] A TV, set-top box, mobile phone, tablet computer, or other electronic device selects (e.g., using a tuner) a channel to receive signals including encoded images and perform block / sub-block / CU corrections, according to any of the examples described.
[0208] A TV, set-top box, mobile phone, tablet computer, or other electronic device receives (e.g., using an antenna) an over-the-air signal including an encoded image and performs block / sub-block / CU corrections according to any of the examples described.
[0209] Although features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile optical discs (DVDs)). A processor associated with the software may be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A device for video decoding, comprising: The processor is configured as follows: Obtain the motion vector (MV) difference associated with the first block and the second block, wherein the second block is adjacent to the first block; The sample values of the boundary pixels of the first block are obtained based on the MV difference; as well as The first block is decoded based on the obtained sample values of the boundary pixels.
2. The device of claim 1, wherein the processor is further configured to: Obtain reference images of the first block and the second block; and The MV of the second block is obtained, wherein the MV of the second block is used to obtain the MV difference based on the determination that the reference image of the first block is the same as the reference image of the second block.
3. The device of claim 1, wherein the processor is further configured to: Obtain reference images for the first block and the second block; Determine that the reference image for the first block is different from the reference image for the second block; and Based on the reference image of the first block and the reference image of the second block, the MV of the second block is scaled, wherein the MV difference is determined based on the scaled MV.
4. The device of claim 1, wherein the sample value of the boundary pixel of the first block is further determined based on the sample values of the pixels adjacent to the boundary pixel.
5. The device of claim 1, wherein the processor is further configured to: The gradient is determined at least based on sample values of pixels adjacent to the boundary pixel; and The sample value offset is determined based on the gradient and the MV difference, wherein the sample value of the boundary pixel of the first block is determined based on the sample value offset.
6. The device of claim 5, wherein an optical flow model is used to determine the gradient.
7. The device of claim 1, wherein the boundary pixel is associated with the boundary of the first block, and the processor is further configured to: A weighting factor is determined based on the distance from the boundary of the first block to the boundary pixel, wherein the weighting factor is inversely proportional to the distance; and The weighting factor is applied to the MV difference, wherein the sample value of the boundary pixel of the first block is obtained based on the weighted MV difference.
8. An apparatus for video encoding, comprising: The processor is configured as follows: Obtain the motion vector (MV) difference associated with the first block and the second block, wherein the second block is adjacent to the first block; The sample values of the boundary pixels of the first block are obtained based on the MV difference; as well as The first block is encoded based on the sample values of the obtained boundary pixels.
9. The device of claim 8, wherein the processor is further configured to: Obtain reference images of the first block and the second block; and The MV of the second block is obtained, wherein the MV of the second block is used to obtain the MV difference based on the determination that the reference image of the first block is the same as the reference image of the second block.
10. The device of claim 8, wherein the processor is further configured to: Obtain reference images for the first block and the second block; Determine that the reference image for the first block is different from the reference image for the second block; and Based on the reference image of the first block and the reference image of the second block, the MV of the second block is scaled, wherein the MV difference is determined based on the scaled MV.
11. The device of claim 8, wherein the sample value of the boundary pixel of the first block is further determined based on the sample values of the pixels adjacent to the boundary pixel.
12. The device of claim 8, wherein the processor is further configured to: The gradient is determined at least based on sample values of pixels adjacent to the boundary pixel; and The sample value offset is determined based on the gradient and the MV difference, wherein the sample value of the boundary pixel of the first block is determined based on the sample value offset.
13. A method for video decoding, comprising: Obtain the motion vector (MV) difference associated with the first block and the second block, wherein the second block is adjacent to the first block; The sample values of the boundary pixels of the first block are obtained based on the MV difference; as well as The first block is decoded based on the obtained sample values of the boundary pixels.
14. The method of claim 13, further comprising: Obtain reference images for the first block and the second block; as well as The MV of the second block is obtained, wherein the MV of the second block is used to obtain the MV difference based on the determination that the reference image of the first block is the same as the reference image of the second block.
15. The method of claim 13, further comprising: Obtain reference images for the first block and the second block; It is determined that the reference image for the first block is different from the reference image for the second block; as well as Based on the reference image of the first block and the reference image of the second block, the MV of the second block is scaled, wherein the MV difference is determined based on the scaled MV.
16. The method of claim 13, further comprising: The gradient is determined at least based on sample values of pixels adjacent to the boundary pixel; as well as The sample value offset is determined based on the gradient and the MV difference, wherein the sample value of the boundary pixel of the first block is determined based on the sample value offset.
17. The method of claim 13, wherein the boundary pixel is associated with the boundary of the first block, and the method further comprises: A weighting factor is determined based on the distance from the boundary of the first block to the boundary pixel, wherein the weighting factor is inversely proportional to the distance; as well as The weighting factor is applied to the MV difference, wherein the sample value of the boundary pixel of the first block is obtained based on the weighted MV difference.
18. A method for video encoding, comprising: Obtain the motion vector (MV) difference associated with the first block and the second block, wherein the second block is adjacent to the first block; The sample values of the boundary pixels of the first block are obtained based on the MV difference; as well as The first block is encoded based on the sample values of the obtained boundary pixels.
19. The method of claim 18, further comprising: Obtain reference images for the first block and the second block; as well as The MV of the second block is obtained, wherein the MV of the second block is used to obtain the MV difference based on the determination that the reference image of the first block is the same as the reference image of the second block.
20. The method of claim 18, further comprising: Obtain reference images for the first block and the second block; It is determined that the reference image for the first block is different from the reference image for the second block; as well as Based on the reference image of the first block and the reference image of the second block, the MV of the second block is scaled, wherein the MV difference is determined based on the scaled MV.
21. A communication system, comprising: The core network (CN) is configured to facilitate communication between multiple wireless transmit / receive units (WTRUs) and external networks; The CN is configured to provide the WTRU with access to a circuit-switched network, including the Public Switched Telephone Network (PSTN), in order to enable communication between the WTRU and landline communication equipment; The CN includes or communicates with an Internet Protocol (IP) gateway, the IP gateway including an IP Multimedia Subsystem (IMS) server configured to connect the CN to the PSTN; and The CN is further configured to provide the WTRU with access to additional networks, including wired and / or wireless networks owned or operated by other service providers.