Filling of unavailable intra samples in intra prediction based on block vectors
By using a block vector-based intra-frame prediction method, undecoded samples are filled with weighted averages, medians, or DC/plane prediction processes from neighboring samples. This solves the problem of handling unavailable intra-frame samples in video coding and improves coding efficiency and quality.
Patent Information
- Application Number
- CN202480033068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-07
- Filing Date
- 2024-04-05
- Publication Date
- 2025-12-12
AI Technical Summary
Existing video coding technologies lack effective padding methods when dealing with unavailable intra-frame samples, leading to a decrease in coding efficiency and quality.
An intra-frame prediction method based on block vectors is adopted, which uses the weighted average or median of neighboring samples, or DC/plane prediction process to fill undecoded samples, partially decoded samples, or partially usable samples.
It improves the efficiency and quality of video encoding, especially when dealing with unavailable intra-frame samples, enhancing the performance of the encoder and decoder.
Smart Images

Figure CN121128164A_ABST
Abstract
Description
[0001] Cross Reference to Related Applications This application claims the benefit of European Patent Application No. 23305531.8, filed April 7, 2023, the disclosure of which is incorporated herein in its entirety by reference. BACKGROUND
[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems. SUMMARY
[0003] Systems, methods, and instrumentalities disclosed herein relate to video coding and / or decoding using padding of unavailable intra samples. In an example, a video decoder or video encoder can obtain, according to block vector-based intra prediction, a prediction block for a current block in a current picture. The encoder or decoder can determine that the prediction block includes previously decoded samples and non-decoded samples. Based on the determination that the prediction block includes the previously decoded samples and the non-decoded samples, the non-decoded samples can be padded using neighboring samples. The current block can be encoded or decoded based at least on the padded samples.
[0004] In an example, the non-decoded samples can be padded using a weighted average of the neighboring samples. In an example, the non-decoded samples can be padded using a median calculation of the neighboring samples. In an example, the non-decoded samples can be padded by applying a DC or planar prediction process using the neighboring samples. In an example, the encoder or decoder can determine that the prediction block includes (e.g., further includes) partially decoded samples. Based on the determination that the prediction block includes the partially decoded samples, the partially decoded samples can be padded using the neighboring samples.
[0005] These examples can be performed by a device having a processor. The device can be an encoder or a decoder. These examples can be performed by a computer program product stored on a non-transitory computer readable medium and including program code instructions. These examples can be performed by a computer program including program code instructions.
[0006] The systems, methods, and instrumentalities described herein can relate to a decoder. In some examples, the systems, methods, and instrumentalities described herein can relate to an encoder. In some examples, the systems, methods, and instrumentalities described herein can relate to a signal (e.g., from an encoder and / or received by a decoder). A computer readable medium can include instructions for causing one or more processors to perform the methods described herein. A computer program product can include instructions that, when the program is executed by one or more processors, can cause the one or more processors to perform the methods described herein. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1A is a system diagram illustrating an example communications system, in which one or more disclosed embodiments can be implemented.
[0008] Figure 1B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that can be used within the communications system Figure 1A illustrated in FIG. 1.
[0009] Figure 1C is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that can be used within the communications system Figure 1A illustrated in FIG. 1.
[0010] Figure 1D is a system diagram illustrating a further example RAN and a further example CN that can be used within the communications system Figure 1A illustrated in FIG. 1.
[0011] Figure 2 illustrates an example video encoder.
[0012] Figure 3 illustrates an example video decoder.
[0013] Figure 4 illustrates an example of a system in which various aspects and examples can be implemented.
[0014] Figure 5 illustrates an example of intra block copy (IBC) padding of an un-decoded or un-encoded sample.
[0015] Figure 6 illustrates an example of an intra template matching prediction (intra TMP) search region.
[0016] Figures 7A-7C illustrates an example of an IBC reference region depending on a current block prediction.
[0017] Figure 8 illustrates an example of a reference region for IBC when a coding tree unit (CTU) (m, n) is coded.
[0018] Figure 9 illustrates an example of padding a sample (x) using its neighboring samples (A, L, and AL).
[0019] Figure 10 illustrates an example of padding using (e.g., only using) available samples.
[0020] Figure 11An example is illustrated that uses median calculation of the top and left samples of the current block for padding.
[0021] Figure 12 An example is illustrated that uses available samples (e.g., neighboring samples) to apply padding to unavailable samples by intra prediction (e.g., PC or Planar).
[0022] Figure 13 An example is illustrated that uses available samples (e.g., neighboring samples) to apply padding to unavailable samples by intra prediction (e.g., PC or Planar).
[0023] Figure 14 An example is illustrated that uses available samples (e.g., neighboring samples) to apply padding to unavailable samples by intra prediction (e.g., PC or Planar). DETAILED DESCRIPTION
[0024] A more detailed understanding can be had from the following description, given by way of example in conjunction with the accompanying drawings.
[0025] Figure 1A is a diagram illustrating an example communications system 100 in which one or more disclosed embodiments can be implemented. The communications system 100 can be a multiple access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communications system 100 can enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communications systems 100 can employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique-word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), and the like.
[0026] As Figure 1AAs shown in FIG. 1, the communications system 100 can include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, though it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which can be referred to as a “station” and / or a “STA”)
[0027] The communications system 100 can also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as the CN 106 / 115, the Internet 110, and / or the other networks 112. By way of example, the base stations 114a, 114b can be a base transceiver station (BTS), a Node-B, an eNode B, a Home Node B, a Home eNode B, a gNB, a NR NodeB, a site controller, an access point (AP), a wireless router, and the like. While the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0028] The base stations 114a can be part of the RAN 104 / 113, which can also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base stations 114a and / or the base stations 114b can be configured to transmit and / or receive wireless signals on one or more carrier frequencies (which can be referred to as a cell (not shown)). These frequencies can be in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrums. A cell can provide wireless service to a particular geographic area that can be fixed or it can change in size and shape as network and / or user conditions change. A cell can further be divided into cell sectors each with an associated coverage area. For example, the cell associated with a base station 114a can be divided into three sectors. Thus, in one embodiment, the base station 114a can include three transceivers, one for each sector of the cell. In an embodiment, the base station 114a can employ multiple-input multiple-output (MIMO) techniques and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0029] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).
[0030] More specifically, as indicated above, the communications system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a and the WTRUs 102a, 102b, 102c in the RAN 104 / 113 can implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish the air interface 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0031] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro.
[0032] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as NR Radio Access, which can establish the air interface 116 using New Radio (NR).
[0033] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and WTRUs 102a, 102b, 102c can implement LTE wireless access and NR wireless access together, for instance using dual connectivity (DC) principles. Thus, the air interface utilized by WTRUs 102a, 102b, 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., an eNB and a gNB).
[0034] In other embodiments, the base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 IX, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), and the like.
[0035] Figure 1AThe base station 114b in such embodiments can be, for example, a wireless router, Home Node B, Home eNode B, or access point, and can utilize any suitable RAT for facilitating wireless connectivity access in a localized area, such as a place of business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a roadway, and the like. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. As shown in FIG. 1C, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b can not be required to access the Internet 110 via the CN 106 / 115. Figure 1A
[0036] The RAN 104 / 113 can be in communication with the CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have varying quality of service (QoS) requirements, such as differing throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. The CN 106 / 115 can provide call control, billing services, mobile location-based services, pre-paid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not shown in FIG. 1A, it will be appreciated that the RAN 104 / 113 and / or the CN 106 / 115 can be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which can employ a NR radio technology, the CN 106 / 115 can also be in communication with another RAN (not shown) employing a GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology. Figure 1A
[0037] CN 106 / 115 can also be used as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0038] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a, which can employ cellular-based radio technology, and with base station 114b, which can employ IEEE 802 radio technology.
[0039] Figure 1B This is a system diagram illustrating example WTRU 102. (Example:) Figure 1B As shown, among other things, WTRU 102 may also include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that, while remaining consistent with the embodiments, WTRU 102 may include any sub-combination of the foregoing elements.
[0040] The processor 118 can be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 can also be implemented as a combination of the foregoing. As noted above, the processor 118 can include one or more processors in association with one or more DSP cores. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, which can be coupled to the transmit / receive element 122. While Figure 1B The processor 118 and the transceiver 120 are depicted as separate components, but it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0041] The transmit / receive element 122 can be configured to transmit signals to, or receive signals from, a base station (e.g., the base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF and light signals. It will be appreciated that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0042] Although the transmit / receive element 122 is depicted in the WTRU 102 Figure 1B In one embodiment, the WTRU 102 can include a plurality of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0043] The transceiver 120 can be configured to modulate information to be transmitted by the transmit / receive element 122 and to demodulate information received by the transmit / receive element 122. As indicated above, the WTRU 102 can be a multi-mode device. Thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11, for example.
[0044] The processor 118 of the WTRU 102 can be coupled to, and can receive user input data from, the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can access information from, and store data in, any type of suitable memory, such as the non-removable memory 130 and / or the removable memory 132. The non-removable memory 130 can include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0045] The processor 118 can receive power from the power source 134 and can be configured to distribute and / or control the power to the other components in the WTRU 102. The power source 134 can be any suitable device for powering the WTRU 102. For example, the power source 134 can include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, and the like.
[0046] The processor 118 can also be coupled to the GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or in lieu of, the information from the GPS chipset 136, the WTRU 102 can receive location information over the air interface 116 from a base station (e.g., base stations 114a, 114b) and / or determine its location based on the timing of the signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 can acquire location information by way of any suitable location-determination method while remaining consistent with an embodiment.
[0047] The processor 118 can further be coupled to other peripherals 138 that can include one or more software and / or hardware modules that provide additional features, functionality and / or wired or wireless connectivity. For example, the peripherals 138 can include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a virtual reality and / or an augmented reality (VR / A R) device, an activity tracker, and the like. The peripherals 138 can include one or more sensors, which can be one or more of a gyroscope, an accelerometer, a hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor; a geolocation sensor; an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0048] The WTRU 102 can include a full duplex radio for which transmission and reception of some or all signals can be concurrent and / or simultaneous. The full duplex radio can include an interference management unit to reduce and / or substantially eliminate self-interference and / or cross- interference by separating transmission and / or reception paths. In an embodiment, the WRTU 102 can include a half duplex radio for which transmission and reception cannot be concurrent.
[0049] Figure 1C is a system diagram illustrating the RAN 104 and the CN 106 according to an embodiment. As noted above, the RAN 104 can employ an E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 can also be in communication with the CN 106.
[0050] The RAN 104 can include eNode-Bs 160a, 160b, 160c, though it will be appreciated that the RAN 104 can include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, 160c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, 160c can implement MIMO technology. Thus, the eNode-B 160a, for example, can use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a.
[0051] Each of the eNode-Bs 160a, 160b, 160c can be associated with a particular cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, and the like. As shown, the eNode-Bs 160a, 160b, 160c can communicate with one another over an X2 interface. Figure 1C
[0052] Figure 1C The CN 106 can include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166, as shown. While each of the foregoing elements are depicted as part of the CN 106, it will be appreciated that any of these elements can be owned and / or operated by an entity other than the CN operator.
[0053] The MME 162 can be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and can serve as a control node. For example, the MME 162 can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during an initial attach of the WTRUs 102a, 102b, 102c, and the like. The MME 162 can provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.
[0054] The SGW 164 can be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via the S1 interface. The SGW 164 can generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 can perform other functions, such as anchoring user planes during inter-eNode B handovers, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing contexts of the WTRUs 102a, 102b, 102c, and the like.
[0055] The SGW 164 can be connected to the PGW 166, which can provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0056] The CN 106 can also serve as a gateway for the WTRUs 102a, 102b, 102c to access the PSTN 108, the Internet 110, and / or the other networks 112. The PSTN 108 can include circuit-switched telephone networks that provide
[0057] Although WTRUs are described in Figures 1A-1D representative embodiments as wireless terminals, it is contemplated that in certain representative embodiments such terminals can (e.g., temporarily or permanently) use wired communication interfaces with the communication network.
[0058] In representative embodiments, the other network 112 can be a WLAN.
[0059] A WLAN in Infrastructure Basic Service Set (BSS) mode can have an Access Point (AP) which is connected to Distribution System (DS) or another type of wired / wireless network that carries traffic in to and / or from the BSS. Traffic to STAs that originates from outside the BSS can arrive through the AP and can be delivered to the STAs. Traffic originating from STAs to destinations outside the BSS can be sent to the AP to be delivered to respective destinations. Traffic between STAs within the BSS can be sent through the AP, for example, where the source STA can send traffic to the AP and the AP can deliver the traffic to the destination STA. The traffic between STAs within a BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent between (e.g., directly between) the source and destination STAs with a direct link setup (DLS). In certain representative embodiments, the DLS can use an 802.11e DLS or an 802.11z tunneled DLS (TDLS). A WLAN using Independent BSS (IBSS) mode can not have an AP, and all STAs within or using the IBSS can communicate directly with each other. The IBSS communication mode can sometimes be referred to herein as "ad-hoc" mode of communication.
[0060] When using 802.11 ac infrastructure mode of operation or similar modes of operation, the AP can transmit beacons on a fixed channel, such as a primary channel. The primary channel can be a fixed width (e.g., 20 MHz wide bandwidth) or a width that is dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish a connection with the AP. In certain representative embodiments, carrier sense multiple access with collision avoidance (CSMA / CA) with collision avoidance can be implemented, for example, in 802.11 systems. For CSMA / CA, STAs including the AP (e.g., each STA) can listen to the primary channel. If the primary channel is sensed / detected as busy and / or determined to be busy by a particular STA, the particular STA can back off. Only one STA can transmit in a given BSS at any given time.
[0061] High Throughput (HT) STAs can use 40 MHz wide channels for communication, for example, via a combination of the primary 20 MHz channel with an adjacent or nonadjacent 20 MHz channel to form a 40 MHz wide channel.
[0062] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. The 40 MHz and / or 80 MHz channels can be formed by combining contiguous 20 MHz channels. A 160 MHz channel can be formed by combining 8 contiguous 20 MHz channels, or by combining two non-contiguous 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel coding, the data can be parsed by a segment parser that can separate the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time domain processing can be done on each stream separately. The streams can be mapped on to the two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above described operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Medium Access Control (MAC).
[0063] 802.11af and 802.11ah support sub-1 GHz modes of operation. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah relative to those used in 802.11η and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah can support meter type control / machine type communication, such as MTC devices in a macro coverage area. MTC devices can have certain capabilities, e.g., limited capabilities, including support for (e.g., support only) certain bandwidths and / or limited bandwidth. MTC devices can include a battery with a battery life above a threshold (e.g., to preserve a very long battery life).
[0064] WLAN systems that can support multiple channels and channel bandwidths such as 802.11η, 802.11ac, 802.11af, and 802.11ah include a channel that can be designated as a primary channel. The primary channel can have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by a STA from among all STAs operating in the BSS that supports the smallest bandwidth operating mode. In the example of 802.11ah, for a STA (e.g., MTC-type device) that supports (e.g., only supports) a 1 MHz mode, the primary channel can be 1 MHz wide even if other STAs in the AP and BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, e.g., due to a STA (that only supports a 1 MHz operating mode) transmitting to the AP, then the entire available frequency band can be considered busy even if most of the frequency band remains idle and can be available.
[0065] In the United States, the available frequency bands for 802.11ah use is from 902 MHz to 928 MHz. In Korea, the available frequency bands is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency bands is from 916.5 MHz to 927.5 MHz. Depending on the country code, the total bandwidth available for 802.11ah is 6 MHz to 26 MHz.
[0066] Figure 1D is a system diagram illustrating the RAN 113 and the CN 115 according to an embodiment. As
[0067] The RAN 113 can include gNBs 180a, 180b, 180c, although it will be appreciated that the RAN 113 can include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, 180c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, 180c can implement MIMO technology. For example, gNBs 180a, 108b can utilize beamforming to transmit signals to and / or receive signals from the gNBs 180a, 180b, 180c. Thus, the gNB 180a, for example, can use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a. In an embodiment, the gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, the gNB 180a can transmit multiple component carriers (not shown) to the WTRU 102a. A subset of these component carriers can be on unlicensed spectrum while the remaining component carriers can be on licensed spectrum. In an embodiment, the gNBs 180a, 180b, 180c can implement Coordinated Multi-Point (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0068] The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using transmissions associated with a scalable numerology. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary from transmission to transmission, from cell to cell, and / or from portion of the wireless transmission spectrum to portion of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing varying number of OFDM symbols and / or lasting varying lengths of absolute time).
[0069] The gNBs 180a, 180b, 180c can be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In the standalone configuration, the WTRUs 102a, 102b, 102c can communicate with the gNBs 180a, 180b, 180c without also accessing other RANs (e.g., such as eNode-Bs 160a, 160b, 160c). In the standalone configuration, the WTRUs 102a, 102b, 102c can utilize signals according to one or more standards developed by
[0070] Each of the gNBs 180a, 180b, 180c can be associated with a particular cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support of network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 184a, 184b, routing of control plane information towards access and mobility management functions (AMFs) 182a, 182b, and the like. As shown in FIG. 1C, the gNBs 180a, 180b, 180c can communicate with one another over an Xn interface. Figure 1D As shown in FIG. 1C, the gNBs 180a, 180b, 180c can also be configured to communicate with a core network 180 over an N2 interface.
[0071] Figure 1DThe CN 115, as shown in FIG. 10B, can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements are depicted as part of the CN 115, it will be appreciated that any of these elements can be owned and / or operated by an entity other than the CN operator.
[0072] The AMF 182a, 182b can be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and can serve as a control node. For example, the AMF 182a, 182b can be responsible for authenticating the WTRUs 102a, 102b, 102c, supporting for network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing the registration area for the WTRUs 102a, 102b, 102c, terminating NAS signaling, mobility management, and the like. Network slicing can be used by the AMF 182a, 182b in order to customize CN support for the WTRUs 102a, 102b, 102c based on the type of services utilized by the WTRUs 102a, 102b, 102c. For example, different network slices can be established for different use cases such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services for machine type communication (MTC) access, and / or the like. The AMF 162 can provide control plane functionality for 5G NR access via the N2 interface to the gNBs 180a, 180b, 180c and non-3GPP access technologies (e.g., WiFi) via an N2’ interface to an ePDG (not shown).
[0073] The SMF 183a, 183b can be connected to AMF 182a, 182b in the CN 115 via an N11 interface. The SMF 183a, 183b can also be connected to the UPF 184a, 184b in the CN 115 via an N4 interface. The SMF 183a, 183b can select and control the UPF 184a, 184b and configure the routes for traffic through the UPF 184a, 184b. The SMF 183a, 183b can perform other functions, such as managing and allocating IP address, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, and the like. A PDU session type can be IP-based, non-IP based, Ethernet-based, and the like.
[0074] The UPF 184a, 184b can be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which can provide the WTRUs 102a, 102b, 102c with access to packet- switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPF 184, 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering of downlink packets, providing mobility anchoring, and the like.
[0075] The CN 115 can facilitate communications with other networks. For example, the CN 115 can include, or can communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 can provide the WTRUs 102a, 102b, 102c with access to the other networks 112, which can include other wired and / or wireless networks that are owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c can be connected to a local DN 185a, 185b through the UPF 184a, 184b via the N3 interface between the UPF 184a, 184b and the UPF 184a, 184b and an N6 interface between the UPF 184a, 184b and the DN 185a, 185b.
[0076] In view of Figures 1A-1D And Figures 1A-1D In view of the corresponding description of the above, one or more or all of the functions described herein with reference to one or more of the WTRUs 102a-d, base stations 114a-b, eNode-Bs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN 185a-b, and / or any other device(s) described herein can be performed by one or more emulation devices (not shown). The emulation devices can be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices can be used to test other devices and / or to simulate network and / or WTRU functionality.
[0077] The emulation devices can be designed to implement one or more tests of one or more other devices in a laboratory environment and / or a carrier network environment. For example, one or more emulation devices can perform one or more, or all, of the functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network in order to test other devices within the communication network. One or more emulation devices can perform one or more functions, or all, of the functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices can be directly coupled to the other device(s) for testing and / or can use over-the-air, wireless communication for the purpose of the testing.
[0078] One or more emulation devices can perform one or more, including all, of the functions while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices can be used in a testing laboratory and / or a testing scenario in a non-deployed (e.g., testing) wired and / or wireless communication network in order to implement testing of one or more components. The emulation device(s) can be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which can include one or more antennas) can be used by the emulation devices to transmit and / or receive data.
[0079] Various aspects are described in the present application, including tools, features, examples, models, methods, etc. Many of these aspects are specifically described, and often in a manner that can sound limiting, at least to illustrate the various features. However, this is for purposes of clarity, and does not limit the application or scope of those aspects. In fact, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects described can also be combined with aspects described in earlier filings.
[0080] The aspects described and contemplated in the present application can be implemented in many different forms. The aspects described herein are described in the context of video coding and decoding, and at least one other aspect is generally related to transmitting generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon. Figures 5-14 Some examples can be provided, but other examples are contemplated. Figures 5-14 The discussion of the aspects is not limiting as to the breadth of the implementation. At least one of the aspects is generally related to video encoding and decoding, and at least one other aspect is generally related to transmitting generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.
[0081] In the present application, the terms “reconstruct” and “decode” can be used interchangeably, the terms “pixel” and “sample” can be used interchangeably, and the terms “image,” “picture,” and “frame” can be used interchangeably.
[0082] Various methods are described herein, and each of the methods described can include one or more steps or actions for implementing one or more of the methods described. Unless a specific order of steps or actions is required for correct operation, then the order and / or use of specific steps and / or actions can be modified or combined for each particular method, and the specific order and / or use of each step and / or action can be modified or combined for each particular method. Further, the use of terminology, such as “first,” “second,” etc., in various examples can be used to modify an element, component, step, operation, etc., such as, for example, “first decoding” and “second decoding.” Unless specifically required, the use of such terminology is not intended to limit the order of modified operations. Thus, in this example, the first decoding need not be performed prior to the second decoding, and can occur, for example, prior to, during, or overlapping in time with the second decoding.
[0083] Various methods and other aspects described in this application can be used to modify modules, such as, for example, the decoding modules of the video encoder 200 and decoder 300 shown in Figure 2 and Figure 3 Further, the subject matter disclosed herein can be applied to video coding of any type, format, or version, whether described in a standard or recommendation, whether preexisting or future developed, and extensions of any such standards and recommendations. The aspects described in this application can be used individually or in combination, unless otherwise indicated or technically precluded.
[0084] Various numerical values, such as bits, bit-depths, etc., are used in the examples described in this application. These and other specific values are used for the purpose of describing examples, and the aspects described are not limited to these specific values.
[0085] Figure 2 FIG. 1 is a diagram illustrating an example video encoder. Variations of the example encoder 200 are contemplated, but for clarity the encoder 200 is described below without describing all contemplated variations.
[0086] Before being encoded, the video sequence can undergo pre-encoding processing (201), e.g., applying a color transform on the input color pictures (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a more resilient to compression signal distribution (e.g., using a histogram equalization of one of the color components). Metadata can be associated with the pre-processing and attached to the bitstream.
[0087] In the encoder 200, a picture is encoded by the encoder elements as follows. The picture to be encoded is partitioned (202) and processed in units of, for example, coding units (CU). Each unit is encoded using either intra or inter mode, for example. When a unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of the intra or inter modes to use to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.
[0088] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as the motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and directly apply quantization on the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode the residual without applying the transform or quantization processes.
[0089] The encoder decodes the encoded blocks to provide references for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The image block is reconstructed by combining (255) the decoded prediction residual and the prediction block. A loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored at a reference picture buffer (280).
[0090] Figure 3 FIG. 1 is a diagram illustrating an example of a video encoder. In the example encoder 200, a picture is encoded by the encoder elements as follows. The picture to be encoded is partitioned (202) and processed in units of, for example, coding units (CU). Each unit is encoded using either intra or inter mode, for example. When a unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of the intra or inter modes to use to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block. Figure 2
[0091] In particular, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partitioning information indicates how the pictures are partitioned. Thus, the decoder can divide (335) the pictures according to the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The prediction blocks are obtained (370) from either intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). In-loop filters (365) are applied to the reconstructed pictures. The filtered pictures are stored at the reference picture buffer (380).
[0092] The decoded pictures can further undergo post-decoding processing (385), e.g., inverse color transform (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream. In an example, the decoded pictures (e.g., after the in-loop filters (365) are applied and / or after the post-decoding processing (385) if used) can be sent to a display device for rendering to a user.
[0093] Figure 4 FIG. 4 is a diagram illustrating an example of a system in which various aspects and examples described herein can be implemented. System 400 can be implemented as a device including various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400, singly or in combination, can be implemented in a single integrated circuit (IC), multiple ICs, and / or in discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more of the aspects described in this document.
[0094] The system 400 includes at least one processor 410 configured to execute instructions loaded therein for implementing the various aspects described in this document, for example. Processor 410 can include embedded memory, input output interface, and various other circuitry as known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). The system 400 includes a storage device 440, which can include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. By way of non-limiting example only, the storage device 440 can include an internal storage device, an attached storage device (including a detachable and non-detachable storage device), and / or a network accessible storage device.
[0095] The system 400 includes an encoder / decoder module 430 configured, for example, to process data to provide encoded video or decoded video, and the encoder / decoder module 430 can include its own processor and memory. The encoder / decoder module 430 represents the module(s) that can be included in a device to perform the encoding and / or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, the encoder / decoder module 430 can be implemented as a separate element, or can be incorporated as a combination of hardware and software as known to those skilled in the art, into the processor 410.
[0096] Program code to be loaded onto the processor 410 or the encoder / decoder 430 to perform the various aspects described in this document can be stored in the storage device 440 and then loaded onto the memory 420 for execution by the processor 410. In accordance with various examples, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 can store one or more of the various entries during the execution of the processes described in this document. Such storage entries can include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.
[0097] In some examples, memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions as well as provide working memory for processing needed during encoding or decoding. However, in other examples, memory external to the processing device (e.g., the processing device can be the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory can be the memory 420 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In several examples, the external non-volatile flash memory is used to store, for example, an operating system of the television. In at least one example, fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0098] Input to elements of the system 400 can be provided through various input devices as indicated in block 445. Such input devices include, but are not limited to: (i) a radio frequency (RF) portion that receives RF signals transmitted, for example, over the air by a broadcaster; (ii) a component (COMP) input terminal (or set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 4 Other examples, not shown in FIG. 4, include composite video.
[0099] In various examples, the input devices of block 445 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to one frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band which may, e.g., be referred to as a channel in certain examples, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select a desired stream of data packets. The RF portion of various examples includes one or more elements for performing these functions, e.g., frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various ones of these functions, including, e.g., downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box example, the RF portion and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various examples rearrange the order of the elements described above (and other elements), remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, e.g., inserting amplifiers and analog-to-digital converters. In various examples, the RF portion includes an antenna.
[0100] The USB and / or HDMI terminals can include respective interface processors for connecting the system 400 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented as desired, e.g., within separate input processing ICs or within the processor 410. Similarly, various aspects of USB or HDMI interface processing can be implemented as desired, e.g., within separate interface ICs or within the processor 410. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, e.g., the processor 410 and the encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data stream as desired for presentation on an output device.
[0101] The various elements of the system 400 can be provided within an integrated housing, within which the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement 425 (e.g., internal buses as known in the art, including Inter-IC (I2C) buses, wiring, and printed circuit boards).
[0102] The system 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 can include, without limitation, a transceiver configured to transmit and receive data through the communication channel 460. The communication interface 450 can include, without limitation, a modem or network card, and the communication channel 460 can be implemented, for example, within wired and / or wireless media.
[0103] In various examples, data is streamed or otherwise provided to the system 400 using a wireless network, such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). Wi-Fi signals of these examples are received through the communication channel 460 and the communication interface 450 that are adapted for Wi-Fi communication. The communication channel 460 of these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples provide streamed data to the system 400 using a set-top box that delivers data through an HDMI connection of the input block 445. Still other examples provide streamed data to the system 400 using an RF connection of the input block 445. As indicated above, various examples provide data in a non-streaming manner. Moreover, various examples use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth® network.
[0104] The system 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 of various examples includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 can be used in a television, a tablet computer, a laptop computer, a cellular telephone (mobile phone), or other device. The display 475 can also be integrated with other components (e.g., as in a smartphone), or separate (e.g., an external monitor for a laptop computer). In various examples, the other peripheral devices 495 include one or more of a standalone digital video disc (or digital versatile disc) (DVD, for both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functionality based on the output of the system 400. For example, a disc player performs the functionality of playing the output of the system 400.
[0105] In various examples, control signals are communicated among the system 400 and the display 475, speakers 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 400 via dedicated connections through the respective interfaces 470, 480, and 490. Alternatively, the output devices can be connected to the system 400 using the communication channel 460 via the communication interface 450. In electronic devices such as, for example, televisions, the display 475 and speakers 485 can be integrated in a single unit with the other components of the system 400. In various examples, the display interface 470 includes a display driver such as, for example, a timing controller (T Con) chip.
[0106] For example, if the RF portion of the input 445 is part of a separate set-top box, the display 475 and speakers 485 can instead be separate from one or more of the other components. In various examples in which the display 475 and speakers 485 are external components, the output signals can be provided via dedicated output connections including, for example, HDMI ports, USB ports, or COMP outputs.
[0107] Examples can be performed by computer software implemented by the processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, examples can be implemented by one or more integrated circuits. As a non-limiting example, the memory 420 can be of any type appropriate for the technology environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory and removable memory, as non-limiting examples. As a non-limiting example, the processor 410 can be of any type appropriate for the technology environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0108] Various implementations relate to decoding. As used in this application, “decoding” can encompass all or a portion of a process performed on a received encoded sequence in order to produce a final output suitable for display. In various examples, such a process includes one or more of the processes typically performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transform, and difference decoding. In various examples, such a process can also or instead include processes performed by decoders of various implementations described in this application, e.g., obtaining a prediction block for a current block in a current picture according to block vector-based intra prediction; determining that the prediction block includes previously decoded samples and un-decoded samples; based on the determination that the prediction block includes previously decoded samples and un-decoded samples, filling the un-decoded samples using neighboring samples; decoding the current block based on the filled samples, etc.
[0109] As a further example, in one example, "decoding" refers only to entropy decoding, in another example, "decoding" refers only to difference decoding, and in another example, "decoding" refers to a combination of entropy decoding and difference decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description, and is believed to be well understood by those skilled in the art.
[0110] Various implementations relate to encoding. In a similar manner as discussed above with respect to "decoding," "encoding" as used in this application can encompass all or a portion of a process performed, e.g., on an input video sequence, in order to produce an encoded bitstream. In various examples, such a process includes one or more of the processes typically performed by an encoder, e.g., partitioning, difference encoding, transform, quantization, and entropy encoding. In various examples, such a process can also include or can instead include processes performed by an encoder of the various implementations described in this application, e.g., obtaining a prediction block for a current block in a current picture according to block vector-based intra prediction; determining that the prediction block includes previously decoded samples and un-decoded samples; based on determining that the prediction block includes previously decoded samples and un-decoded samples, filling the un-decoded samples using neighboring samples; encoding the current block based on the filled samples; etc.
[0111] As a further example, in one example, "encoding" refers only to entropy encoding, in another example, "encoding" refers only to difference encoding, and in another example, "encoding" refers to a combination of difference encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will be clear based on the context of the specific description, and is believed to be well understood by those skilled in the art.
[0112] When an accompanying drawing figure is presented as a flow chart, it will be understood that any one or more of the steps, operations, or processes in the figure can be combined with other steps, operations, or processes, or further split into multiple steps, operations, or processes, and each of the steps, operations, or processes can be performed automatically or with the aid of one or more human operators.
[0113] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of the features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0114] Reference to "one example" or "an example" or "one implementation" or "an implementation", as well as other variants, means that a particular feature, structure, characteristic, and so forth being described in connection with an example is included in at least one example. Therefore, the appearance of the phrase "in one example" or "in an example" or "in one implementation" or "in an implementation", as well as any other variations, throughout this application should not necessarily be construed as referring to the same example.
[0115] Furthermore, this application can relate to "determining" pieces of information. Determining information can include one or more of, for example, estimating information, calculating information, predicting information, or retrieving information from memory. Obtaining can include receiving, retrieving, constructing, generating, and / or determining.
[0116] Furthermore, this application can relate to "accessing" pieces of information. Accessing information can include one or more of, for example, receiving information, retrieving information (for example, retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0117] Furthermore, this application can relate to "receiving" pieces of information. As with "accessing", receiving is intended to be a broad term. Receiving information can include one or more of, for example, accessing information or retrieving information (for example, retrieving information from memory). Furthermore, "receiving" is typically involved in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0118] It is to be understood that, in cases where "A / B," "A and / or B," and "at least one of A and B" are used, the usage is intended to cover the selection of one of the first listed options (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in cases where "A, B, and / or C" and "at least one of A, B, and C" are used, such phrasing is intended to cover a selection of at least one of the first listed options (A) only, or the selection of at least one of the second listed options (B) only, or the selection of at least one of the third listed options (C) only, or the selection of combinations of the first, second, and third listed options such as A and B or A and C or B and C, and so forth, or the selection of all of the first, second, and third listed options.
[0119] Also, as used herein, the word "signaling" refers to indicating something to a corresponding decoder, among other things. The encoder signal can include, for example, an encoding function on the input of a block using a precision factor, and so forth. As such, in examples, the same parameters are used at both the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicitly signal) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter along with other parameters, the signaling can be used without transmission (implicitly signaled) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual function, bit savings are achieved in various examples. It is to be understood that the signaling can be implemented in a variety of ways. For example, in various examples, information is signaled to a corresponding decoder using one or more syntax elements, flags, and so forth. While the foregoing involves the verb form of the word "signaling," the word "signal" can also be used (for example) as a noun herein.
[0120] As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals including as process data signals that are formatted to carry information that can be, for example, stored or transmitted. The information can include instructions for performing a method, or data created by one of the described implementations. For example, a signal can be formatted to carry a bitstream of an example described. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium, and can be accessed by a processor.
[0121] Many examples are described herein. Features of examples can be provided alone or in any combination of various claim classes and types. Further, examples can include one or more of the features, devices, or aspects described herein, alone or in any combination of various claim classes and types. For example, features described herein can be implemented in a bitstream or signal that includes information generated as described herein. The information can allow a decoder to decode the bitstream, encoder, bitstream, and / or decoder according to any of the described examples. For example, features described herein can be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein can be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein can be implemented by a TV, set-top box, cell phone, tablet, or other electronic device that performs decoding. The TV, set-top box, cell phone, tablet, or other electronic device can display (for example, using a monitor, screen, or other type of display) resulting images (for example, images reconstructed from residuals from a video bitstream). The TV, set-top box, cell phone, tablet, or other electronic device can receive a signal that includes encoded images and perform decoding.
[0122] Examples can be performed by a device having at least one processor. The device can be an encoder or a decoder. Examples can be performed by a computer program product stored on a non-transitory computer readable medium and including program code instructions. Examples can be performed by a computer program including program code instructions.
[0123] Figure 5An example of intra block copy (IBC) padding of undecoded or uncoded samples is illustrated. IBC can be enhanced with partially decodable prediction candidates. IBC can be an intra mode in which a prediction block is obtained from the current picture (e.g., current frame), and a block vector (similar to a motion vector) can be transmitted to the decoder to locate the coordinates of the prediction block. The block vector can be obtained from the decoded region. In examples, undecoded samples can (e.g., also) be used for prediction. A padding process can be applied to prediction using undecoded or uncoded samples (e.g., as shown in Figure 5 FIG. 1 illustrates an example of a padding process. Padding can be applied to prediction using undecoded or uncoded samples. Padding can be achieved by copying samples from the top-left pixel. The padding mechanism can be improved by other extrapolation techniques, and padding examples can be extended to other template-based coding tools to optimize intra prediction. Examples herein can improve the padding technique, and can integrate padding into template-based tools.
[0124] Figure 6 An example of an intra template matching search region used in intraTMP is illustrated. IntraTMP is an intra prediction mode that can copy the best prediction block from the reconstructed part of the current picture (e.g., current frame), whose L-shaped template can match the current template. For a predefined search range, the encoder can search for the template that is most similar to the current template in the reconstructed part of the current picture, and can use the corresponding block as the prediction block. The encoder can (e.g., then can) signal the use of this mode. The same prediction operation can be performed at the decoder side.
[0125] A prediction signal can be generated by matching the L-shaped causal neighboring block of the current block with another block in the predefined search region in FIG. 7, including: R1 : current CTU R2: top-left CTU R3: top CTU R4: left CTU The sum of absolute differences (SAD) can be used as the cost function. Within the regions (e.g., within each region), the decoder can search for the template that has the minimum SAD with respect to the current template, and can use its corresponding block as the prediction block. The size of the regions (SearchRange_w, SearchRange_h) can be set proportionally to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH where ‘a’ can be a constant that controls the gain / complexity trade-off. For example, ‘a’ can equal 5.
[0126] For CUs having a size less than or equal to 64 in width and height, the intra TMP tool can be enabled. This maximum CU size for intra TMP can be configurable. The intra TMP mode can be signaled at the CU level by a dedicated flag.
[0127] Figures 7A-7C An example of IBC reference region depending on the current block prediction is illustrated. In the example, the encoder or decoder can use template matching in IBC in both IBC merge mode and IBC advanced motion vector prediction (AMVP) mode. The intra block copy template matching (IBC-TM) merge list can be modified (e.g., such that the candidates are selected according to the pruning method with motion distance between candidates as in the regular TM merge mode) compared to the list used by regular IBC merge mode. The ending zero motion realization can be replaced by motion vectors to the left (-W, 0), up (0, -H), and top-left (-W, -H), where W can be the width of the current CU and H can be the height of the current CU.
[0128] In the IBC-TM merge mode, the selected candidate can be refined using template matching before the rate-distortion optimization (RDO) or the decoding process. The IBC-TM merge mode can compete with the regular IBC merge mode and the TM merge flag (which can be signaled).
[0129] In the IBC-TM AMVP mode, up to three candidates can be selected from the IBC-TM merge list. Each of those three selected candidates can be refined using template matching and can be ordered according to their final template matching cost. The first two (e.g., only the first two) can be considered in the motion estimation process.
[0130] For both IBC-TM merge and AMVP modes, the template matching refinement, the IBC motion vector can be constrained to be (i) integer and (ii) within the reference region. In the IBC-TM merge mode, the refinement can be performed in integer precision. In the IBC-TM AMVP mode, the optimization can be performed in integer or 4-pixel precision depending on the AMVR value. In some examples, such refinement accesses samples without interpolation (e.g., only samples are accessed). The refinement motion vector and the template used in a part (e.g., each part) of the refinement process can comply with the constraint of the reference region (e.g., in both IBC-TM merge and IBM-TM AMVP modes).
[0131] Figure 8 An example of the reference region of IBC when a coding tree unit (CTU) (m, n) is coded is illustrated. As Figure 8As shown in the middle, the blue blocks can represent the current CTU, the green blocks can represent the reference region, and the white blocks can represent the invalid reference region. The reference region of IBC can be extended to the above two CTU rows. Figure 8 The reference region for encoding the CTU (m, n) is shown. Specifically, for the CTU (m, n) to be encoded, the reference region can include the CTUs with indices (m - 2, n - 2)... (W, n - 2), (0, n - 1)... (W, n - 1), (0, n)... (m, n), where W can represent the maximum horizontal index within the current tile, slice, or picture. If the CTU size is 256, the reference region can be limited to one CTU row above. This setting can ensure that IBC can not require additional memory in the current ETM platform for CTUs of size 128 or 256. The per-sample block vector search (or referred to as local search) range can be limited to [ - (C « 1), C » 2] horizontally and can be limited to [ - C, C » 2] vertically adaptive reference region extension, where C can represent the CTU size.
[0132] Figure 9 An example of filling a sample (x) using its neighboring samples (A, L, and AL) is illustrated. In the example, an encoder or decoder can obtain a prediction block for a current block in a current picture (e.g., according to block vector based intra prediction). The encoder or decoder can obtain a plurality of samples in the prediction block. The encoder or decoder can determine that the plurality of samples can include previously decoded samples and non-decoded or non-encoded samples. The non-decoded samples or non-encoded samples can be filled using the neighboring samples (e.g., based on determining that the prediction block includes previously decoded samples and non-decoded samples). The current block can be encoded or decoded based on the filled samples.
[0133] It can be optimal to copy the missing samples. This can be because statistics from available samples can not be considered. Examples herein can consider statistics of neighboring (e.g., surrounding) samples. In examples, an average of a median of the neighboring samples can be considered. As Figure 9 As shown in the middle, the blue blocks can represent the current CTU, the green blocks can represent the reference region, and the white blocks can represent the invalid reference region. The reference region of IBC can be extended to the above two CTU rows.
[0134] In an example, a median calculation of neighboring samples can be used to fill in undecoded or unencoded samples. In non-camera captured sequences (e.g., screen share or game sequences), an averaging process can not be used as they can tend to smooth the prediction signal if the characteristics of the original signal are sharp edges. The averaging process can be replaced with a median calculation (e.g., in such examples). The median can be calculated as: x = median(A, L, AL). The median process can select the middle value among A, L, and AL. Similar processes can be used for other unavailable (e.g., undecoded or unencoded) samples.
[0135] Figure 10 An example of filling using available samples is illustrated. In an example, a weighted average of neighboring samples can be used to fill in undecoded samples or unencoded samples. In an example, a weighted average from available samples can be applied without considering the newly filled in sample. This can be to avoid dependencies between samples and / or can be to allow for parallel processing (e.g., as illustrated in Figure 10 The “x” value can be predicted as a weighted average between A and L as follows: x = (wl*A + w2*L) / (wl + w2), where the weights wl and w2 can be scaled according to the distance relative to the current sample “x”. For optimized computation, the weights can be limited to powers of two.
[0136] Figure 11 An example of filling using median calculation of the top and left samples of the current block is illustrated. For using a median based calculation, at least three values can be needed. In addition to the above and left samples, other samples can be used for the median calculation (e.g., as illustrated in Figure 11
[0137] Figure 12 An example of filling of unavailable samples by means of intra prediction mode using neighboring samples (samples directly surrounding the intra prediction region, as illustrated in Figure 12 In an example, the filling process can be performed using the existing default angular prediction process(es). The filling process can be performed by applying a DC or planar prediction process to fill in undecoded samples or unencoded samples using neighboring samples (e.g., samples directly surrounding the intra prediction region, as illustrated in Figure 12
[0138] Figure 13 Examples of padding when reference block portion is available are illustrated. In examples, padding can extend to other areas where samples are not available (e.g., not decoded or not encoded). In examples, multiple samples in a prediction block can be partially available. An encoder or decoder can determine that a prediction block includes partially decoded samples. The partially decoded samples can be padded using neighboring samples (e.g., based on determining that the prediction block includes partially decoded samples). In examples, padding can be extended (e.g., can only be extended) if some of the bottom-right samples are not available. In examples, padding can be applied by copying from the top-left of the current block. In examples, padding can be applied via using an average and a median.
[0139] Figure 14 Examples of padding when right samples are not available are illustrated. In examples, right samples can be missing due to CTU boundaries. In examples, DC or planar prediction using only left samples can be applied. In examples, a weighted average or planar from left samples can be applied. For weighted average or planar from left samples, ‘x’ can be calculated as: x = median(L0, L1, A) or x = (w0*L0 + w1*L1 + w2*A) / (w0+w1+w2). Figure 14 Figure 14
[0140] Examples herein can be extended to intraTMP mode. If an intraTMP search is performed to find the best candidate, partially available blocks can be included as prediction blocks. These blocks can be partially available due to CTU boundaries or picture boundaries. If including prediction candidates with partially decoded (e.g., new) predictions, the distance metric (SAD, SATD, etc. between current block template and reference block template) can be scaled by a factor (>1) to less prefer those prediction blocks. If comparing two candidates, one with complete samples and one with incomplete samples, the complete samples can be better to use. In examples, the scaling factor can be proportional to the number of unavailable samples per region. If half of the samples are not available, the weighting factor can be higher than if only a quarter of the samples are not available.
[0141] Although features and elements are described above in particular combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in combination with others dependent upon the circumstances. The methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer- readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (optical, electrical or electromagnetic) that are transmittable through a wired or wireless connection. Examples of computer-readable media include, but are not limited to, removable tapes, hard disks, floppy disks, memory cards, CD ROMs, and DVD ROMs. The processor associated with a software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. An apparatus for video decoding, the apparatus comprising: a processor configured to: obtain, according to block vector based intra prediction, a prediction block for a current block in a current picture; determine that the prediction block includes previously decoded samples and undecoded samples; based on determining that the prediction block includes previously decoded samples and undecoded samples, fill the undecoded samples using neighboring samples; and decode the current block based at least on the filled samples.
2. The apparatus of claim 1, wherein, fill the undecoded samples using a weighted average of the neighboring samples.
3. The apparatus of claim 1, wherein, fill the undecoded samples using a median calculation of the neighboring samples.
4. The apparatus of claim 1, wherein, fill the undecoded samples by applying a DC or planar prediction process using the neighboring samples.
5. The apparatus of claim 1, wherein, the processor is further configured to: determine that the prediction block further includes partially decoded samples; based on determining that the prediction block includes partially decoded samples, fill the partially decoded samples using neighboring samples.
6. A method for video decoding, the method comprising: obtaining, according to block vector based intra prediction, a prediction block for a current block in a current picture; determining that the prediction block includes previously decoded samples and undecoded samples; based on determining that the prediction block includes previously decoded samples and undecoded samples, filling the undecoded samples using neighboring samples; and decoding the current block based at least on the filled samples.
7. The method of claim 6, wherein, filling the undecoded samples using a weighted average of the neighboring samples.
8. The method of claim 6, wherein, filling the undecoded samples using a median calculation of the neighboring samples.
9. The method of claim 6, wherein, filling the undecoded samples by applying a DC or planar prediction process using the neighboring samples.
10. The method of claim 6, further comprising: determining that the prediction block further includes partially decoded samples; based on determining that the prediction block includes partially decoded samples, filling the partially decoded samples using neighboring samples.
11. An apparatus for video encoding, the apparatus comprising: a processor configured to: obtain, according to block vector based intra prediction, a prediction block for a current block in a current picture; determine that the prediction block includes previously decoded samples and undecoded samples; based on determining that the prediction block includes previously decoded samples and undecoded samples, fill the undecoded samples using neighboring samples; and encode the current block based at least on the filled samples.
12. The apparatus of claim 11, wherein, filling the undecoded samples using a weighted average of the neighboring samples.
13. The apparatus of claim 11, wherein, filling the undecoded samples using a median calculation of the neighboring samples.
14. The apparatus of claim 11, wherein, filling the undecoded samples by applying a DC or planar prediction process using the neighboring samples.
15. The apparatus of claim 11, wherein, the processor is further configured to: determine that the prediction block further includes partially decoded samples; based on determining that the prediction block includes partially decoded samples, fill the partially decoded samples using neighboring samples.
16. A method for video encoding, the method comprising: obtaining, according to block vector based intra prediction, a prediction block for a current block in a current picture; determining that the prediction block includes previously decoded samples and undecoded samples; based on determining that the prediction block includes previously decoded samples and undecoded samples, filling the undecoded samples using neighboring samples; and encoding the current block based at least on the filled samples.
17. The method of claim 16, wherein, filling the undecoded samples using a weighted average of the neighboring samples.
18. The method of claim 16, wherein, filling the undecoded samples using a median calculation of neighboring samples.
19. The method of claim 16, wherein, filling the undecoded samples by applying a DC or planar prediction process using neighboring samples.
20. The method of claim 16, further comprising: determining that the prediction block further includes partially decoded samples; based on determining that the prediction block includes partially decoded samples, filling the partially decoded samples using neighboring samples.
21. A computer program product, stored on a non-transitory computer readable medium, and comprising program code instructions for implementing the steps of the method according to at least one of claims 6 to 10 and 16 to 20 when executed by a processor.
22. A computer readable medium comprising program code instructions for implementing the steps of the method according to at least one of claims 6 to 10 and 16 to 20 when executed by a processor.
23. Video data comprising information representative of an encoded output generated according to one of the methods of any of claims 16 to 20.