Methods and apparatus for a flexible grid region
By partitioning video frames into flexible grid regions using a set of parameters, the method addresses inefficiencies in existing video coding technologies, enhancing data compression and transmission efficiency.
Patent Information
- Application Number
- JP2024119611
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-14
- Filing Date
- 2024-07-25
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2039-09-13
AI Technical Summary
Existing video coding technologies face challenges in efficiently partitioning video frames into flexible grid regions or tiles, which can lead to inefficiencies in data compression and transmission.
A method and apparatus for partitioning video frames into flexible grid regions using a set of parameters, allowing for more dynamic and adaptive tile structures that enhance data compression and transmission efficiency.
The solution enables more efficient data compression and transmission by allowing for more flexible and adaptive tile structures in video frames, improving overall video coding performance.
Smart Images

Figure 0007715890000007 
Figure 0007715890000008 
Figure 0007715890000009
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein generally relate to signaling and processing picture or video information. For example, one or more embodiments disclosed herein relate to methods and apparatuses for using flexible grid regions or tiles in a picture / video frame.
Background Art
[0002] The embodiments disclosed herein generally relate to signaling and processing picture or video information.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
[0005] Provided are a method and an apparatus for using a flexible grid region or tile in a picture / video frame. [Means for Solving the Problems]
[0006] Disclosed are a method and an apparatus for using a flexible grid region in a picture or video frame. In one embodiment, the method includes receiving a set of first parameters that define a plurality of first grid regions that make up a frame. For each of the first grid regions, the method includes receiving a set of second parameters that define a plurality of second grid regions, where the plurality of second grid regions partition the respective first grid region. The method further includes partitioning the frame into the plurality of first grid regions based on the set of first parameters, and partitioning each of the first grid regions into the plurality of second grid regions based on the respective set of second parameters.
[0007] A more detailed understanding can be obtained from the following detailed description, given by way of example, in conjunction with the drawings attached hereto. The figures in the description are by way of example. Therefore, the figures and the detailed description should not be regarded as limiting, and other equally valid examples are possible and may exist. Further, like reference numerals in the figures indicate like elements.
Advantages of the Invention
[0008] Provided are a method and an apparatus for using a flexible grid region or tile in a picture / video frame.
Brief Description of the Drawings
[0009]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 12C
Figure 13
[0010] I. Exemplary Networks and Devices FIG. 1A illustrates an exemplary communication system 100 in which one or more of the disclosed embodiments may be implemented. Communication system 100 can be a multi-connection system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. Communication system 100 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, communication system 100 can utilize one or more channel access methods such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero Tail Unique Word DFT Spread OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, and Filter Bank Multicarrier (FBMC).
[0011] As shown in Figure 1A, the communication system 100 can include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, each of them, which may sometimes be referred to as a "station" and / or "STA", the WTRUs 102a, 102b, 102c, 102d can be configured to transmit and / or receive wireless signals and can include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular telephones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), home appliances, as well as devices operating on commercial and / or industrial wireless networks. Any of the WTRUs 102a, 102b, 102c, 102d may alternatively be referred to as a UE.
[0012] The communication system 100 can also include base station 114a and / or base station 114b. Each of base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as CN106 / 115, the Internet 110, and / or other network 112. By way of example, base stations 114a, 114b can be a base transceiver station (BTS), Node B, eNode B, home Node B, home eNode B, gNB, New Radio (NR) Node B, site controller, access point (AP), and wireless router, among others. Although base stations 114a, 114b are each depicted as a single element, it will be understood that base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0013] The base station 114a can be part of the RAN 104 / 113, and the RAN 104 / 113 can also include other base stations and / or network elements such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes (not shown). The base station 114a and / or the base station 114b can be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies can be in a licensed spectrum, an unlicensed spectrum, or a combination of a licensed spectrum and an unlicensed spectrum. A cell can provide coverage for wireless services in a specific geographic area that can be relatively fixed or can change over time. A cell can further be divided into cell sectors. For example, a cell associated with the base station 114a can be divided into three sectors. Thus, in one embodiment, the base station 114a can include three transceivers, for example, one for each sector of the cell. In an embodiment, the base station 114a can utilize multiple-input multiple-output (MIMO) technology and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0014] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, and the air interface 116 can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).
[0015] More specifically, as mentioned above, the communication system 100 may be a multiple-access system and may utilize one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a and the WTRUs 102a, 102b, and 102c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 116 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0016] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE), and / or LTE Advanced (LTE-A), and / or LTE Advanced Pro (LTE-A Pro).
[0017] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.
[0018] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interface utilized by the WTRUs 102a, 102b, 102c can be characterized by transmissions from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).
[0019] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement wireless technologies such as IEEE 802.11 (e.g., wireless fidelity (WiFi)), IEEE 802.16 (e.g., worldwide interoperability for microwave access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, interim standard 2000 (IS-2000), interim standard 95 (IS-95), interim standard 856 (IS-856), global system for mobile communications (GSM), high speed data rate for GSM evolution (EDGE), and GSM EDGE (GERAN).
[0020] The base station 114b in FIG. 1A can be, for example, a wireless router, a home node B, a home e-node B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a localized area such as an office, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by a drone), and a roadway. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in FIG. 1A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0021] RAN104 / 113 can communicate with CN106 / 115, and CN106 / 115 can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, 102d. The data can have various Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 can provide call control, billing services, mobile location-based services, prepaid originating calls, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN104 / 113 and / or CN106 / 115 can communicate directly or indirectly with other RANs that utilize the same or a different Radio Access Technology (RAT) as RAN104 / 113. For example, in addition to being connected to RAN104 / 113 which may be utilizing New Radio (NR) wireless technology, CN106 / 115 can also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi wireless technology.
[0022] CN106 / 115 can also serve as a gateway for WTRU102a, 102b, 102c, 102d to access the PSTN108, the Internet 110, and / or other networks 112. The PSTN 108 can include a circuit-switched telephone network that provides basic telephone service (POTS). The Internet 110 can include a worldwide system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) within the TCP / IP Internet protocol suite. The network 112 can include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 can include another CN connected to one or more RANs that can utilize the same RAT or a different RAT as the RAN 104 / 113.
[0023] Some or all of the WTRU102a, 102b, 102c, 102d within the communication system 100 can include a multi-mode function (e.g., the WTRU102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU102c shown in Figure 1A can be configured to communicate with a base station 114a that can utilize cellular-based wireless technology and with a base station 114b that can utilize IEEE802 wireless technology.
[0024] Figure 1B is a system diagram illustrating an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 can include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU 102 can include any sub-combination of the above elements while maintaining consistency with the embodiments.
[0025] The processor 118 can be, for example, a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors cooperating with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, and the transceiver 120 can be coupled to the transmit / receive element 122. Although Figure 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0026] The transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmitting / receiving element 122 can be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmitting / receiving element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0027] In FIG. 1B, the transmitting / receiving element 122 is depicted as a single element, but the WTRU 102 can include any number of transmitting / receiving elements 122. More specifically, the WTRU 102 can utilize MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0028] The transceiver 120 can be configured to modulate the signals that will be transmitted by the transmitting / receiving element 122 and demodulate the signals received by the transmitting / receiving element 122. As mentioned above, the WTRU 102 can have a multimode function. Thus, the transceiver 120 can include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs such as, for example, NR and IEEE 802.11.
[0029] The processor 118 of the WTRU 102 can be coupled to and receive user input data from a speaker / microphone 124, keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can obtain information from and store data in any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 can include a random access memory (RAM), read only memory (ROM), hard disk, or any other type of memory storage device. The removable memory 132 can include, for example, a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card. In other embodiments, the processor 118 can obtain information from and store data in a memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0030] The processor 118 can receive power from a power supply 134 and can be configured to distribute power to and / or control the power to other components within the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 can include one or more dry cells (such as nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell.
[0031] Processor 118 can also be coupled to a GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 can receive location information on the air interface 116 from a base station (e.g., base stations 114a, 114b), and / or can determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 can acquire location information using any suitable location determination method while maintaining consistency with the embodiments.
[0032] Processor 118 can further be coupled to other peripheral devices 138, which can include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 can include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a Frequency Modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, and an activity tracker, among others. The peripheral devices 138 can include one or more sensors, which can be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0033] The WTRU 102 can include a full-duplex radio in which some or all of the transmission and reception of signals (associated with certain subframes for both, e.g., UL (for transmission) and downlink (for reception)) can be parallel and / or simultaneous. The full-duplex radio can include an interference management unit 139 to reduce and / or substantially eliminate self-interference, either via hardware (e.g., choke) or via signal processing through a processor (e.g., a separate processor (not shown) or processor 118). In embodiments, the WTRU 102 can include a half-duplex radio for some or all of the transmission and reception of signals (associated with certain subframes for either, e.g., UL (for transmission) or downlink (for reception)).
[0034] Figure 1C is a system diagram illustrating a RAN 104 and a CN 106, according to an embodiment. As mentioned above, the RAN 104 can communicate with the WTRU 102a, 102b, 102c over an air interface 116, using E-UTRA radio technology. The RAN 104 can also communicate with the CN 106.
[0035] The RAN 104 can include eNodeBs 160a, 160b, 160c, although it will be understood that the RAN 104 can include any number of eNodeBs while maintaining consistency with the embodiments. The eNodeBs 160a, 160b, 160c can each include one or more transceivers for communicating with the WTRU 102a, 102b, 102c over the air interface 116. In one embodiment, the eNodeBs 160a, 160b, 160c can implement MIMO technology. Thus, the eNodeB 160a, for example, can transmit a wireless signal to and / or receive a wireless signal from the WTRU 102a using multiple antennas.
[0036] Each of the eNodeBs 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL. As shown in Figure 1C, the eNodeBs 160a, 160b, and 160c can communicate with each other over the X2 interface.
[0037] The CN 106 shown in Figure 1C can include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the above elements is depicted as part of the CN 106, it will be understood that any of these elements can be owned and / or operated by an entity different from the CN operator.
[0038] The MME 162 can be connected to each of the eNodeBs 160a, 160b, and 160c within the RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of the WTRUs 102a, 102b, and 102c, bearer activation / deactivation, and selecting a specific serving gateway during the initial attach of the WTRUs 102a, 102b, and 102c. The MME 162 can provide control plane functions for exchanges between the RAN 104 and other RANs (not shown) that utilize other radio technologies such as GSM and / or WCDMA.
[0039] SGW164 can be connected to each of the eNodeBs 160a, 160b, and 160c within RAN104 via the S1 interface. SGW164 can generally route and transfer user data packets to / from WTRUs 102a, 102b, and 102c. SGW164 can perform other functions such as anchoring the user plane during handover between eNodeBs, triggering paging when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the contexts of WTRUs 102a, 102b, and 102c.
[0040] SGW164 can be connected to PGW166, and PGW166 can provide access to a packet-switched network, such as the Internet 110, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and IP-enabled devices.
[0041] CN106 can facilitate communication with other networks. For example, CN106 can provide access to a circuit-switched network, such as PSTN108, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and traditional fixed-line telephone communication devices. For example, CN106 can include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN106 and PSTN108. Additionally, CN106 can provide access to other networks 112 to WTRUs 102a, 102b, and 102c, and other networks 112 can include other wired and / or wireless networks owned and / or operated by other service providers.
[0042] In FIGS. 1A - 1D, the WTRU is described as a wireless terminal, but in certain representative embodiments, it is contemplated that such a terminal may (e.g., temporarily or permanently) use a wired communication interface to a communication network.
[0043] In some representative embodiments, the other network 112 can be a WLAN.
[0044] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have an access or interface to a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic to an STA originating from outside the BSS can arrive through the AP and be delivered to the STA. Traffic transmitted from an STA to a destination outside the BSS can be sent to the AP for delivery to their respective destinations. Traffic between STAs within the BSS can be sent through the AP. For example, the source STA can send the traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered peer - to - peer traffic and / or sometimes be called peer - to - peer traffic. Peer - to - peer traffic can be sent (e.g., directly) between the source STA and the destination STA using direct link setup (DLS). In certain representative embodiments, the DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) can communicate directly with each other. IBSS mode communication is sometimes referred to herein as "ad - hoc" mode communication.
[0045] When using the operation of 802.11ac infrastructure mode or the operation of a similar mode, the AP can transmit beacons on a fixed channel such as the primary channel. The primary channel can be of a fixed width (e.g., 20 MHz bandwidth), or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In one representative embodiment, for example, in an 802.11 system, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) can be implemented. In the case of CSMA / CA, STAs including the AP (e.g., any STA) can sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA can back off. Within a given BSS, at any given time, one STA (e.g., only one station) can transmit.
[0046] A high throughput (HT) STA can use a 40 MHz wide channel for communication, for example, by combining the primary 20 MHz channel with adjacent or non - adjacent 20 MHz channels to form a 40 MHz wide channel.
[0047] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non - consecutive 80 MHz channels, which may be referred to as an 80 + 80 configuration. In the case of the 80 + 80 configuration, after channel encoding, the data can pass through a segment parser that can split the data into two streams. For each stream separately, an Inverse Fast Fourier Transform (IFFT) process, and time - domain processing can be performed. The streams can be mapped onto two 80 MHz channels and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80 + 80 configuration can be reversed and the combined data can be transmitted to the Medium Access Control (MAC).
[0048] The operation of the sub-1 GHz mode is supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports bandwidths of 5 MHz, 10 MHz, and 20 MHz in the TV white space (TVWS) spectrum, and 802.11ah supports bandwidths of 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz using the non-TVWS spectrum. According to an exemplary embodiment, 802.11ah can support meter type control / machine type communication, such as MTC devices in a macro coverage area. The MTC devices can have limited functionality, including certain functions, for example, support for a certain bandwidth and / or limited bandwidth (e.g., only their support). The MTC devices can include a battery having a battery life above a threshold (e.g., to maintain a very long battery life).
[0049] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that can be designated as primary channels. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs within a BSS. The bandwidth of the primary channel can be set and / or restricted by the STA that supports the minimum bandwidth operating mode among all STAs operating within the BSS. In the example of 802.11ah, for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode, even if the AP and other STAs within the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the primary channel can be 1MHz wide. Carrier sensing and / or network allocation vector (NAV) setting can depend on the status of the primary channel. For example, if the primary channel is busy because an STA (that only supports the 1MHz operating mode) is transmitting to the AP, the entire available frequency band can be considered busy even though most of the frequency band remains idle and available.
[0050] In the United States, the available frequency band that can be used by 802.11ah is from 902MHz to 928MHz. In South Korea, the available frequency band is from 917.5MHz to 923.5MHz. In Japan, the available frequency band is from 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah is from 6MHz to 26MHz, depending on national regulations.
[0051] Figure 1D is a system diagram showing RAN 113 and CN 115 according to an embodiment. As mentioned above, RAN 113 can communicate with WTRUs 102a, 102b, 102c over air interface 116 using NR radio technology. RAN 113 can also communicate with CN 115.
[0052] RAN 113 can include gNBs 180a, 180b, 180c, although it will be understood that RAN 113 can include any number of gNBs while maintaining consistency with the embodiment. gNBs 180a, 180b, 180c can each include one or more transceivers for communicating with WTRUs 102a, 102b, 102c over air interface 116. In one embodiment, gNBs 180a, 180b, 180c can implement MIMO technology. For example, WTRUs 102a, 108b can use beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, 180c. Thus, gNB 180a, for example, can use multiple antennas to transmit wireless signals to and / or receive wireless signals from WTRU 102a. In an embodiment, gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers can be in unlicensed spectrum, while the remaining component carriers can be in licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c can implement multi-point coordinated (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0053] WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol interval, and / or the OFDM sub-carrier interval can be different for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using sub-frames or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or lasting for various lengths of absolute time).
[0054] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (such as eNodeBs 160a, 160b, and 160c). In a stand-alone configuration, WTRUs 102a, 102b, and 102c can utilize one or more of gNBs 180a, 180b, and 180c as mobility anchor points. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals within an unlicensed band. In a non-stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with / connect to gNBs 180a, 180b, and 180c while also communicating with / connecting to another RAN such as eNodeBs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement the DC principle to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNodeBs 160a, 160b, and 160c. In a non-stand-alone configuration, eNodeBs 160a, 160b, and 160c can serve as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0055] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, and routing of control plane information to access and mobility management functions (AMFs) 182a, 182b. As shown in Figure 1D, gNBs 180a, 180b, and 180c can communicate with each other over the Xn interface.
[0056] CN 115 shown in Figure 1D can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and possibly data networks (DNs) 185a, 185b. Although each of the above elements is depicted as part of CN 115, it should be understood that any of these elements can be owned and / or operated by entities different from the CN operator.
[0057] AMF182a and 182b can be connected to one or more of gNBs 180a, 180b, and 180c in RAN113 via the N2 interface and can serve as control nodes. For example, AMF182a and 182b can authenticate users of WTRUs 102a, 102b, and 102c, support network slicing (e.g., handling different PDU sessions with different requirements), select specific SMFs 183a and 183b, manage the registration area, terminate NAS signaling, and perform mobility management, etc. Network slicing can be used by AMF182a and 182b to customize the CN support for WTRUs 102a, 102b, and 102c based on the type of service utilized by WTRUs 102a, 102b, and 102c. For example, different network slices can be established for different use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on high-speed large-capacity mobile broadband (eMBB) access, and / or services for machine type communication (MTC) access. AMF182 can provide control plane functions for exchanges between RAN113 and other RANs (not shown) that utilize other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.
[0058] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b and configure the routing of traffic through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and allocating WTRU or UE IP addresses, managing PDU sessions, enforcing policies and controlling QoS, and providing downlink data notifications. The PDU session type can be IP-based, non-IP-based, and Ethernet-based, etc.
[0059] UPF184a and 184b can be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N3 interface, and they can provide access to a packet-switched network such as the Internet 110 to WTRU102a, 102b, and 102c to facilitate communication between WTRU102a, 102b, and 102c and IP-corresponding devices. UPF184a and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homing PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0060] CN115 can facilitate communication with other networks. For example, CN115 can include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN115 and the PSTN108. In addition, CN115 can provide access to other networks 112 to the WTRU102a, 102b, 102c, and the other networks 112 can include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRU102a, 102b, 102c can be connected to the local data network (DN) 185a, 185b through the UPF184a, 184b via an N3 interface to the UPF184a, 184b and an N6 interface between the UPF184a, 184b and the DN185a, 185b.
[0061] In view of FIGS. 1A-1D, and the corresponding descriptions of FIGS. 1A-1D, one or more of the functions described herein with respect to one or more of the WTRU102a-d, base stations 114a-b, eNode Bs 160a-c, MME162, SGW164, PGW166, gNBs 180a-c, AMF182a-b, UPF184a-b, SMF183a-b, DN185a-b, and / or any other devices described herein can be performed by one or more emulation devices (not shown). An emulation device can be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device can be used to test other devices and / or to simulate network and / or WTRU functions.
[0062] An emulation device can be designed to perform one or more tests on other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices can execute one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices can execute one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device can be directly coupled to another device for the purpose of conducting tests and / or can execute tests using over-the-air wireless communication.
[0063] One or more emulation devices can execute one or more functions including all functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device can be utilized in a test scenario in a test laboratory and / or in a non-deployed (e.g., test) wired and / or wireless communication network to conduct tests on one or more components. One or more emulation devices can be test equipment. Wireless communication via direct RF coupling and / or via an RF circuit (which can include one or more antennas for example) can be used by an emulation device to transmit and / or receive data.
[0064] A video encoding system can be used to compress digital video signals, which can reduce the storage requirements and / or the transmission bandwidth of the video signal over a network, such as either of the networks described above. The video encoding system can include block-based, wavelet-based, and / or object-based systems. The block-based video encoding system can be based on, use, comply with, conform to, etc., one or more standards, such as MPEG-1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, High Efficiency Video Coding (HEVC), and / or Versatile Video Coding (VVC). The block-based video encoding system can include a block-based hybrid video encoding framework.
[0065] In some examples, a video streaming device can include one or more video encoders, each of which can generate a video bitstream at a different resolution, frame rate, or bitrate. The video streaming device can include one or more video decoders, each of which can detect and / or decode an encoded video bitstream. In various embodiments, one or more video encoders and / or one or more decoders can be implemented within a device having a processor, receiver, and / or transmitter communicatively coupled to a memory. The memory can include processor-executable instructions that include instructions for performing any of the various embodiments disclosed herein (e.g., exemplary procedures). In various embodiments, the device can be configured as and / or can use various elements of a Wireless Transmit / Receive Unit (WTRU). Exemplary details of the WTRU and its elements are provided herein in FIGS. 1A-1D and the accompanying disclosure.
[0066] II. HEVC II.1 High Efficiency Video Coding (HEVC) Tiles In some examples, a video frame can be divided into slices and / or tiles. A slice is a sequence of one or more slice segments that starts with an independent slice segment and includes all subsequent dependent slice segments. A tile is rectangular and contains an integer number of coding tree units as specified by HEVC. For each slice and tile, one or both of the following conditions are to be satisfied (see, for example, [1] (Non-Patent Document 1)), that is, 1) all coding tree units within a slice belong to the same tile, and / or 2) all coding tree units within a tile belong to the same slice.
[0067] In some examples, the tile structure in HEVC is signaled in the picture parameter set (PPS) by specifying the height of rows and the width of columns. Individual rows and / or columns can have different sizes, but the partitioning can always span the entire picture from left to right or from top to bottom.
[0068] In some examples, the HEVC tile syntax can be used. In an example, as shown in Table 1, the first flag, tiles_enabled_flag, can be used to specify whether tiles are used. For example, if the first flag (tiles_enabled_flag) is set, the number of tile columns and rows is specified. The second flag, uniform_spacing_flag, can be used to specify whether the tile column boundaries, and also the tile row boundaries, are uniformly distributed across the picture. For example, when uniform_spacing_flag is equal to zero (0), the syntax elements column_width_minus1[i] and row_height_minus1[i] are explicitly signaled to specify the column width and row height. In addition, the third flag, loop_filter_across_tiles_enabled_flag, can be used to specify whether to turn on or off the in-loop filter across tile boundaries for all tile boundaries within a picture.
[0069]
Table 1
[0070] In one implementation, two examples of tile partitions are shown in FIGS. 2A and 2B. In the first example, as shown in FIG. 2A, the tile columns and rows are evenly distributed (into six grid regions) across picture 200. In the second example, as shown in FIG. 2B, the tile columns and rows are not evenly distributed (into six grid regions) across picture 202, and thus the tile column widths and row heights may need to be explicitly specified.
[0071] In some examples, HEVC specifies a special tile set called a motion-constrained tile set (MCTS) via a Supplemental Enhancement Information (SEI) message. The MCTS SEI message indicates that the inter prediction process is constrained such that sample values at fractional sample positions that are derived using sample values outside each identified tile set and / or one or more sample values outside the identified tile set cannot be used for the inter prediction of any samples within the identified tile set [1] (Non-Patent Document 1). In some cases, each MCTS can be independently extracted from the HEVC bitstream and decoded.
[0072] II.2 Padding for Motion Compensation Prediction In some examples, existing video codecs are designed for conventional two-dimensional (2D) video captured on a plane. When motion compensation prediction uses any samples outside the boundary of the reference picture, replicated padding is performed by copying sample values from the boundary of the picture.
[0073] In an example, FIG. 3 illustrates a replicated padding scheme 300. For example, block B0 partially exists outside the reference picture. Portion P0 is filled with the top-left sample of portion P3. Portion P1 has each row filled with the top row of portion P3. Portion P2 has each column filled with the left column of portion P3.
[0074] In some examples, 360-degree video includes video information over the entire sphere and thus 360-degree video inherently has circular characteristics. When considering this circular characteristic, all the information included in the "boundary" wraps around on the spherical surface, so the reference picture of 360-degree video no longer has a "boundary". In some implementations, geometric padding for 360-degree video (e.g., the geometric padding proposed in [5] (Non-Patent Document 5)) can be used.
[0075] In an example, FIG. 4 illustrates a geometry padding process 400 for a 360-degree video having an equirectangular projection (ERP) format. In this example, the geometry padding process for ERP can include padding to be filled at an arrow (e.g., arrow A), which is taken along a corresponding arrow (e.g., arrow A’), and so on, and the alphabet labels indicate the correspondence. For example, at the left and right boundaries of the 360-degree video, samples at A, B, C, D, E, and F are padded with samples at A’, B’, C’, D’, E’, and F’. At the upper boundary, samples at G, H, I, and J are padded with samples at G’, H’, I’, and J’. At the lower boundary, samples at K, L, M, and N are padded with samples at K’, L’, M’, and N’. Compared with the iterative padding method currently used in HEVC, geometry padding can provide meaningful samples and improve the continuity of neighboring samples for areas outside the ERP picture boundary.
[0076] III. Viewport-Dependent Omnidirectional Video Processing The omnidirectional media format (OMAF) is a system standard format developed by the Moving Picture Experts Group (MPEG). OMAF defines a media format that enables omnidirectional media, including 360-degree video, images, audio, and associated time-stamped text. Some viewport-dependent omnidirectional video processing schemes are described, for example, in Appendix D of [2] (Non-Patent Document 2).
[0077] In an example, an equal-resolution MCTS-based viewport-dependent scheme encodes the same omnidirectional video content into several HEVC bitstreams with different picture qualities and bitrates. Each MCTS is included in one region track, and an extractor track is also created. The OMAF player selects the quality at which each subpicture track is received based on the viewing direction.
[0078] FIG. 5 illustrates an exemplary scheme 500 from clause D4.2 of [2] (Non-Patent Document 2). In this example, the OMAF player receives MCTS tracks 1, 2, 5, and 6 at a particular quality and region tracks 3, 4, 7, and 8 at another quality. An extractor track is used to reconstruct the bitstream that can be decoded using a single HEVC decoder. Tiles of the reconstructed HEVC bitstream having different qualities of MCTS can be signaled by the HEVC tile syntax described herein.
[0079] In another example, an MCTS-based viewport-dependent video processing scheme is used to encode the same omnidirectional video source content at several spatial resolutions. Based on the viewing direction, the extractor can select tiles that match the viewing direction at high resolution and other tiles at low resolution. The bitstream disassembled from the extractor track conforms to HEVC and can be decoded by a single HEVC decoder.
[0080] FIG. 6 illustrates an example of a cube map (CMP) partitioning scheme 600 from clause D.6.4 of [2] (Non-Patent Document 2). In this example, preprocessing and encoding for achieving a 6K effective CMP resolution are shown using an HEVC-based viewport-dependent OMAF video profile. The content is encoded at two spatial resolutions having CMP face sizes of 1536×1536 and 768×768, respectively. In both bitstreams, a 6×4 tile grid is used and for each tile position, MCTS is encoded. Each encoded MCTS sequence is stored as a region track. For each different viewport-adaptive MCTS selection, an extractor track is created. This results in the creation of 24 extractor tracks. In each sample of the extractor track, for each MCTS, one extractor is created to extract data from the region track containing one or more selected high-resolution or low-resolution MCTS. Each extractor track uses the same 3×6 tile grid having a tile column width equal to 768, 768, and / or 384 lumasamples (e.g., one or more pixels of luminance), and / or a constant tile row height of 768 lumasamples. Each tile extracted from the low-resolution bitstream contains two slices. The bitstream disassembled from the extractor track has a resolution of 1920×4608, for example, compliant with HEVC level 5.1.
[0081] In some cases, the MCTS of the above-described reconstructed bitstream may not be represented using the HEVC tile syntax as described above (e.g., in Table 1). Instead, slices can be used for each partition. Referring to FIG. 7, in the example, there are two extracted tracks, namely the left extracted track and the right extracted track. The left extracted track has six slice headers represented as slice headers 702, 704, 706, 708, 710, and 712. The right extracted track has twelve slice headers represented as slice headers 714, 716, 718, 720, 722, 724, 726, 728, 730, 732, 734, and 736. In this case, the partitioning of the extractor track in FIG. 6 can ultimately result in twelve slice headers as shown in FIG. 7.
[0082] FIG. 8 illustrates an example of a preprocessing and encoding scheme 800 for achieving a 6K effective ERP resolution (e.g., HEVC-based). Clause D6.3 of Non-Patent Document 2 presents an MCTS-based viewport-dependent scheme for achieving a 6K effective ERP resolution. In the example, an omnidirectional video with a 6K resolution (6144×3072) is resampled to three spatial resolutions, namely 6K (6144×3072), 3K (3072×1536), and 1.5K (1536×768). The 6K sequence and the 3K sequence are each cropped to 6144×2048 (as shown in grid 802) and 3072×1024 (as shown in grid 804), respectively, by excluding the elevation angle range of 30 degrees from the upper and lower sides. The cropped 6K and 3K input sequences are encoded to have an 8×1 tile grid in such a way that each tile becomes an MCTS.
[0083] As shown in grid 806, the upper and lower stripes of size 3072×256 corresponding to the 30-degree elevation range are extracted from the 3K input sequence. The upper and lower stripes are encoded as separate bitstreams with a 4×1 tile grid in such a way that the rows of the tiles form a single MCTS. As shown in grid 808, the upper and lower stripes of size 1536×128 corresponding to the 30-degree elevation range are extracted from the 1.5K input sequence. Each stripe is arranged to form a picture of size 768×256, for example, by placing the left side of the stripe at the upper side of the picture and the right side of the stripe at the lower side of the picture.
[0084] In this example, each MCTS sequence from the cropped 6K and 3K bitstreams can be encapsulated as a separate track. Each bitstream containing the upper or lower stripe of the 3K or 1.5K input sequence can be encapsulated as one track (e.g., track 810).
[0085] For each selection of four adjacent tiles from the cropped 6K bitstream and separately for the viewing directions above and below the equator, an extractor track is prepared. This results in the creation of 16 extractor tracks. Each extractor track uses the same arrangement, for example, as illustrated in FIGS. 9A and 9B. For example, in FIG. 9A, grid 900 includes segmentation by conventional 2×2 tiles (uniform_spacing_flag = 1). In FIG. 9B, grid 902 includes segmentation by flexible tiles (uniform_spacing_flag = 1). The picture size of the bitstream disassembled from the extractor track is 3840×2304, compliant with HEVC level 5.1. In some cases, the tile partitioning of the extractor track may not be specified using the HEVC tile syntax described above (e.g., in Table 1).
[0086] IV. HEVC Tiles In HEVC, tiles are aligned with the Coding Tree Unit (CTU) boundaries. In some examples, the main use of HEVC tiles is to partition a picture into independent segments with a minimal reduction in compression efficiency. In one implementation, HEVC tiles are used to partition a picture for viewport-dependent omnidirectional video processing. In that case, the source video is partitioned and encoded using one or more MCTSs that can be decoded independently of neighboring tile sets. The extractor can select a subset of the tile sets based on the viewport direction and form an HEVC-compliant extractor track for OMAF player consumption.
[0087] For next-generation video compression standards such as Versatile Video Coding (VVC), the CTU size may become larger due to the increasing resolution of images. The granularity of tile segmentation may also become too large and may not align with the frame packing boundaries. It is also difficult to split a picture into equal-sized CTUs for load balancing. Furthermore, the conventional tile structure may not handle the above-mentioned partition structure for OMAF viewport-dependent processing, and moreover, the bit cost of using slices for partitioning is high.
[0088] The flexible tile structure and syntax were proposed by [3] (Non-Patent Document 3) and [4] (Non-Patent Document 4) in MPEG#123. Non-Patent Document 3 proposed that a picture can be split into CTUs of a fixed size like conventional tiles, but to achieve better load balancing and align with the frame packing boundary, the sizes of the rightmost and bottommost CTUs at the tile boundary can be different from the fixed CTU size. The fractional-sized CTUs at the right and bottom edges of each tile are encoded and decoded in the same way as at the picture boundary.
[0089] Non-Patent Document 4 proposed to support flexible tiles that are rectangular but of variable size. Each tile is signaled individually by copying the tile size from the preceding tile size in the decoding order or by a codeword of one tile width and one tile height. Using the proposed syntax, the partitioning structures shown in FIGS. 6 and 8 can be supported. However, such a syntax format may introduce a significant overhead cost for conventional tile structures commonly used, compared to the HEVC tile syntax format.
[0090] Therefore, new or improved methods, schemes, and signaling designs for supporting flexible tiles (e.g., in a video frame) are desired.
[0091] V. Representative Procedures for Flexible Tiles In the present disclosure, the inventors describe numerous embodiments, procedures, methods, architectures, tables, and signal designs for supporting flexible grid regions or tiles, including, for example: 1) geometric padding and loop filter constraints for flexible tiles, 2) signaling to distinguish conventional tiles from flexible tiles in order to reduce overall signaling overhead, 3) grid region-based flexible tile signaling design and scan conversion, and 4) initial quantization parameter (QP) signaling for tile-based video processing.
[0092] In various embodiments, the term "region" as used in the present disclosure can represent a first set of grid regions, and the term "tile" as used in the present disclosure can represent a second set of grid regions. In an example, a picture or video frame can be divided into a first set of grid regions (e.g., regions), and each grid region of the first set of grid regions can be further divided into a second set of grid regions (e.g., tiles). In some cases, the terms "region", "grid regions", and "tile" as used in the present disclosure can be interchangeable and can be represented as either the first set or the second set of grid regions.
[0093] V.1 Padding and Loop Filter Constraints at Tile Boundaries Conventional tile partitioning may not have an integer multiple of CTUs at the right or bottom edge of a picture, and a flexible tile may not have an integer multiple of CTUs at the right or bottom edge of the tile. FIG. 9 illustrates such incompleteness in the case of both conventional tiles and flexible tiles using a conventional method in an example. Incomplete CTUs along the right and bottom edges of each tile can be encoded and decoded as in the case of picture boundaries.
[0094] Geometric padding assumes that all the information contained in a 360-degree video wraps around on the spherical surface, and such cyclic property is maintained regardless of the projection format used to represent the 360-degree video on a 2D plane. Geometric padding can be applied to the 360-degree video picture boundary, but the cyclic property may not be applied to the flexible tile boundary because it depends on the partitioning structure. Based on tile partitioning, the encoder can determine or decide whether it can perform horizontal or vertical geometric padding for each tile, for example, for motion compensation prediction.
[0095] In some embodiments, a padding flag can be signaled (e.g., to a receiver of a WTRU) to indicate whether a padding operation can be performed on a tile edge. If padding_enabled_flag is set, repetitive padding or geometric padding can be performed on the tile edge. In some examples, for a flexible tile syntax structure, each tile can be signaled individually. In some cases, for each tile, a geometry_padding_indicator and a repetitive_padding_indicator can be signaled.
[0096] In some embodiments, in HEVC, loop_filter_across_tiles_enabled_flag is signaled to indicate whether loop filter operations can be performed across tile boundaries in the PPS. When loop_filter_across_tile_enabled_flag is set, loop_filter_indicator can be signaled, for example, to indicate which edges of a tile can be filtered.
[0097] In the example, Table 2 illustrates padding and loop filter syntax formats for tiles or grid regions.
[0098]
Table 2
[0099] In Table 2, padding_enable_flag equal to 1 indicates that padding operations can be used in the current tile, and padding_enable_flag equal to 0 indicates that padding operations are not used in the current tile.
[0100] In Table 2, geometry_padding_indicator is a bitmap that maps each tile edge to a bit. An example of the bitmapping can be that, in clockwise order, the most significant bit is the flag for the upper edge, the second most significant bit is the flag for the right edge, and so on. When the bit value is equal to 1, geometry padding operations can be applied to the corresponding tile edge, and when the bit value is equal to 0, geometry padding operations are not performed on the corresponding tile edge. When it does not exist, the default value of geometry_padding_indicator can be assumed to be equal to 0.
[0101] In Table 2, the repetitive_padding_indicator is a bitmap that maps each tile edge to bits. An example of the bitmapping can be that, in clockwise order, the most significant bit is the flag for the upper edge, the second most significant bit is the flag for the right edge, and so on. When the bit value is equal to 1, the repetitive padding operation can be applied to the corresponding tile edge, and when the bit value is equal to 0, the repetitive padding operation is not executed for the corresponding tile edge. When it does not exist, the default value of the repetitive_padding_indicator can be presumed to be equal to 0.
[0102] In Table 2, the loop_filter_indicator is a bitmap that maps each tile edge to bits. When the bit value is equal to 1, the loop filter operation can be applied across the corresponding tile edge, and when the bit value is equal to 0, the loop filter operation is not executed across the corresponding tile edge. When it does not exist, the default value of the loop_filter_indicator can be presumed to be equal to 0.
[0103] In another embodiment, the padding activation flag padding_on_tile_enabled_flag can be signaled at the PPS level. When padding_on_tile_enabled_flag is equal to 0, the padding_enabled_flag at the tile level is presumed to be 0.
[0104] In another embodiment, when the size of the current tile edge and the size of the corresponding reference boundary are not the same (e.g., different), the geometry padding can be disabled.
[0105] Figure 10 illustrates an example of using flexible tiles in an ERP picture. In this example, the ERP picture 1000 can be divided into a number of tiles, each having a variable size. Geometric padding can be enabled for specific tile edges according to the tile partitioning grid.
[0106] V.2 Signaling for Distinguishing between Conventional Tile Grids and Flexible Tile Grids Conventional tile partitioning restricts all tiles belonging to the same tile row to have the same row height and all tiles belonging to the same tile column to have the same column width. Such restrictions simplify tile signaling and ensure that the tile set is rectangular in shape. Flexible tiles allow individual tiles to have different sizes and enable the characteristics of each tile to be signaled individually. Such signaling supports various partitioning grids but may introduce significant bit overhead. A compromise between overhead bit cost and tile partitioning flexibility can be achieved by including an indicator or flag for distinguishing between conventional and flexible partitioning grids. The indicator or flag can indicate whether the entire picture is partitioned into a normal M×N grid, where M and N are integers. The conventional HEVC tile syntax can be applied to a normal M×N tile grid, while the new flexible tile syntax, as described in Non-Patent Document 4 or this disclosure, can be applied to a flexible tile grid.
[0107] In some examples, the indicator or flag described herein can be signaled in and / or by the sequence parameter set and / or the picture parameter set.
[0108] V.3 Grid Region - Based Signaling for Flexible Tiles In some examples, tile column boundaries, and likewise tile row boundaries, span across the picture. Motivating use cases for flexible tiles are viewport - dependent omnidirectional video processing techniques where multiple MCTS tracks from different picture resolutions are merged into a single HEVC - compliant extractor track. The tile grid of the extractor track can be from different picture resolutions, and thus, as shown in FIGS. 6 and / or 8, tile column and row boundaries may not be continuous across the picture.
[0109] Instead of signaling the size of each tile individually, the signaling scheme / design can be used or configured to signal each grid region, in which a particular tile or region partitioning scheme is utilized. In an example, different regions can have different grid partitionings to enable flexible tiles. In this example, each region can have multiple tiles, and each tile can have the same or different sizes. In some examples, a first tile can have a different size compared to a second tile within the same grid region. In some cases, tiles in each row can share the same height, and tiles in each column can share the same width.
[0110] Table 3 shows an exemplary flexible tile syntax (e.g., multi - level syntax) for use in this exemplary signaling scheme / design.
[0111] [Table 3]
[0112] num_region_columns_minus1 plus 1 specifies the number of region columns that partition the picture. num_region_columns_minus1 shall be in the range from 0 to PicWidthInCtbsY - 1.
[0113] num_region_rows_minus1 plus 1 specifies the number of region rows that partition the picture. num_region_columns_minus1 shall be in the range from 0 to PicHeightInCtbsY - 1.
[0114] Regions can be in raster scanning order from left to right and top to bottom. The total number of regions NumRegions can be derived as follows.
[0115] NumRegions = (num_region_columns_minus1 + 1) × (num_region_rows_minus1 + 1) A uniform_region_flag equal to 1 specifies that the region column boundaries, and also the region row boundaries, are uniformly distributed across the picture. A uniform_spacing_flag equal to 0 specifies that the region column boundaries, and also the region row boundaries, are not uniformly distributed across the picture but are explicitly signaled using the syntax elements region_column_width_minus1, and region_row_height_minus1. When not present, the value of uniform_region_flag is assumed to be equal to 1.
[0116] region_size_unit_idc specifies that the unit size of the region is in coded tree block units. When not present, the default value of region_size_unit_idc is assumed to be equal to 0. The variable RegionUnitInCtbsY is derived as follows.
[0117] RegionUnitInCtbsY = 1 << region_unit_size_idc One plus region_column_width_minus1[i] specifies the width of the i-th region column in terms of the coded tree block unit. When it does not exist, the value of region_column_width_minus1 is presumed to be equal to the picture width PicWidthInCtbsY.
[0118] One plus region_row_height_minus1[i] specifies the height of the i-th region row in terms of the coded tree block unit. When it does not exist, the value of region_row_width_minus1 is presumed to be equal to the picture height PicHeightInCtbsY.
[0119] Figures 11A and 11B illustrate two examples of region-based flexible tile signaling applied to the extractor tracks shown in Figures 6 and 8, respectively.
[0120] Referring to Figure 11A, the extractor track of Figure 6 is reconstructed as track 1100 from two pictures with different resolutions. Two regions are identified, and within each region, tiles are uniformly distributed. The left region of track 1100 is partitioned into a 2×6 grid, and the right region of track 1100 is partitioned into a 1×12 grid.
[0121] Referring to Figure 11B, the extractor track of Figure 8 is reconstructed as track 1110 from four pictures with different resolutions, four regions are identified, and within each region, tiles are uniformly distributed. The partitioning grid of the first region is 4×1, the partitioning grid of the second region is 2×2, the partitioning grid of the third region is 4×1, and the partitioning grid of the fourth region is 1×2.
[0122] In various embodiments, when processing video information (e.g., encoding or decoding a video or picture), the region partitioning and grouping mechanisms described herein can be utilized. In an example, a WTRU (e.g., WTRU 102) can be configured to receive (or identify) a set of first parameters that define a plurality of first grid regions (e.g., tiles) that make up a frame (e.g., a video frame or a picture frame). For each first grid region, the WTRU can be configured to receive (or identify) a set of second parameters that define a plurality of second grid regions, where the plurality of second grid regions can partition their respective first grid regions. The WTRU can be configured to partition the frame into a plurality of first grid regions based on the set of first parameters and to partition each first grid region into a plurality of second grid regions based on their respective sets of second parameters.
[0123] In another example, the WTRU can be configured to receive (or identify) a plurality of sets of parameters or configurations for processing video information. For example, the WTRU can be configured to receive (or identify) a first set of parameters (defining a plurality of first grid regions) and a second set of parameters (defining a plurality of second grid regions). The WTRU can be configured to partition a frame into a plurality of first grid regions based on the first set of parameters and to group (or reorganize) the plurality of first grid regions into a plurality of second grid regions based on the second set of parameters. In some cases, the first grid region or the second grid region can be a tile or a slice and can be used to construct or reconstruct a frame (e.g., a video frame or a picture frame) or to generate one or more bitstreams.
[0124] V.4 Coding Tree Block (CTB) Raster and Flexible Tile Scanning Conversion Process In some embodiments, one or more of the following variables can be derived by initiating a coding tree block raster and flexible tile scanning conversion process.
[0125] a) A list CtbAddrRsToTs[ctbAddrRs] where ctbAddrRs is in the range of 0 or more and PicSizeInCtbsY - 1 or less, specifying the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan. b) A list CtbAddrTsToRs[ctbAddrTs] where ctbAddrTs is in the range of 0 or more and PicSizeInCtbsY - 1 or less, specifying the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture. c) A list TileId[ctbAddrTs] that specifies the conversion from the CTB address in the tile scan to the tile ID, where ctbAddrTs is in the range from 0 to PicSizeInCtbsY - 1. d) A list ColumnWidthInLumaSamples[i][j] that specifies the width of the j-th tile column of the i-th region in luma samples, where j is in the range from 0 to num_tile_columns_minus1[i], and / or e) A list RowHeightInLumaSamples[i][j] that specifies the height of the j-th tile row of the i-th region in luma samples, where j is in the range from 0 to num_tile_rows_minus1[i].
[0126] FIG. 12A illustrates an example of the CTB raster scan of picture frame 1200. FIG. 12B illustrates an example of the CTB raster scan of a conventional tile in picture frame 1210. FIG. 12C illustrates an example of the CTB raster scan of a region-based flexible tile in picture frame 1220. The conversion from the CTB address in the CTB raster scan of a picture to the CTB address in the conventional tile scan is specified in HEVC [1] (Non-Patent Document 1). However, HEVC does not specify how to convert from the CTB address in the CTB raster scan of a picture to the CTB address in the region-based flexible tile scan.
[0127] In some embodiments, the conversion from the CTB address in the CTB raster scan of a picture to the CTB address in the region-based flexible tile scan can be configured as follows, that is, 1) The variables CtbSizeY, PicWidthInCtbsY, and PicHeightInCtbsY are the same as those specified in HEVC [1] (Non-Patent Document 1), and / or 2) Use a new list regionColWidth[i] where i ranges from 0 to num_region_columns_minus1 to specify the width of the i-th region column in CTB units, and the new list can be derived as follows. if( uniform_region_flag ) for( i = 0; i <= num_region_columns_minus1; i++ ) regionColWidth[ i ] =(( i + 1)*PicWidthInCtbsY) / (num_region_columns_minus1+1) - (i * PicWidthInCtbsY ) / (num_region_columns_minus1+1) else { regionColWidth[ num_region_columns_minus1 ] = PicWidthInCtbsY for( i = 0; i < num_region_columns_minus1; i++ ) { regionColWidth[ i ] = RegionUnitInCtbsY * ( region_column_width_minus1[ i ] + 1 ) regionColWidth[ num_region_columns_minus1 ] -= regionColWidth[ i ] } }
[0128] In some embodiments, a new list regionRowHeight[j] where j ranges from 0 to num_region_rows_minus1 to specify the height of the j-th region row in CTB units can be derived as follows. if( uniform_region_flag ) for( j = 0; j <= num_region_rows_minus1; j++ ) regionRowHeight[ j ] = (( j + 1 ) * PicHeightInCtbsY) / (num_region_rows_minus1 + 1) - (j * PicHeightInCtbsY) / (num_region_rows_minus1 + 1) else { regionRowHeight[ num_region_rows_minus1 ] = PicHeightInCtbsY for( j = 0; j < num_region_rows_minus1; j++ ) { regionRowHeight[ j ] = RegionUnitInCtbsY * ( region_row_height_minus1[ j ] + 1 ) regionRowHeight[ num_tile_rows_minus1 ] -= regionRowHeight[ j ] } }
[0129] In some examples, for the i-th region in raster scanning order, the new variables RegionWidthInCtbsY, and RegionHeightInCtbsY can be derived as follows. RegionWidthInCtbsY[ i ] = regionColWidth[ i % (num_region_columns_minus1 + 1) ] RegionRowInCtbsY[ i ] = regionRowHeight[ i / (num_region_row_minus1 + 1) ] RegionSizeInCtbsY[ i ] = RegionWidthInCtbsY[ i ] * RegionRowInCtbsY[ i ]
[0130] In some embodiments, for i in the range from 0 to num_region_columns_minus1 + 1, a new list regionColBd[i] that specifies the location of the column boundary of the i-th region in coding tree block units can be derived as follows. for ( regionColBd[0] = 0; i = 0; i <= num_region_columns_minus1; i++ ) regionColBd[ i + 1 ] = regionColBd[ i ] + regionColWidth[ i ]
[0131] In some embodiments, for j in the range from 0 to num_region_rows_minus1 + 1, a new list regionRowBd[j] that specifies the location of the row boundary of the j-th region in coding tree block units can be derived as follows. for( regionRowBd
[0000] = 0; j = 0; j <= num_region_rows_minus1; j++ ) regionRowBd[ j + 1 ] = regionRowBd[ j ] + regionRowHeight[ j ]
[0132] In some embodiments, for j in the range from 0 to num_tile_columns_minus1[i], a new list colWidth[i][j] that specifies the width of the j-th tile column of the i-th region in CTB units can be derived as follows. if( uniform_spacing_flag ) for( j = 0; j <= num_tile_columns_minus1[ i ]; j++ ) colWidth[i][j] = ((i + 1)*RegionWidthInCtbsY[ i ]) / (num_tile_columns_minus1[ i ]+1)-(i*RegionWidthInCtbsY[i]) / (num_tile_columns_minus1[i]+1) else { colWidth[ num_tile_columns_minus1[i] ] = RegionWidthInCtbsY[i] for( j = 0; j < num_tile_columns_minus1[i]; j++ ) { colWidth[ i ][ j ] = column_width_minus1[ i ][ j ] + 1 colWidth[ i ][ num_tile_columns_minus1 ] -= colWidth[ i ][ j ] } }
[0133] In some embodiments, for j in the range from zero to num_tile_rows_minus1, a new list rowHeight[i][j] specifying the height of the j-th tile row of the i-th region in CTB units can be derived as follows. if( uniform_spacing_flag ) for( j = 0; j <= num_tile_rows_minus1[i]; j++ ) rowHeight[i][j] = ((j+1)*RegionHeightInCtbsY[i]) / (num_tile_rows_minus1[i]+1)-(j*RegionHeightInCtbsY[i]) / (num_tile_rows_minus1[i] + 1) else { rowHeight[i][ num_tile_rows_minus1[i] ] = RegionHeightInCtbsY[i] for( j = 0; j < num_tile_rows_minus1[i]; j++ ) { rowHeight[ i ][ j ] = row_height_minus1[ i ][ j ] + 1 rowHeight[ i ][ num_tile_rows_minus1 ] -= rowHeight[ i ][ j ] } }
[0134] In some examples, the new variables ColumnWidthInLumaSamples[i][j], and RowHeightInLumaSamples[i][j] can be derived as follows. ColumnWidthInLumaSamples[ i ][ j ] = colWidth[ i ][ j ] * CtbSizeY RowHeightInLumaSamples[ i ][ j ] = rowHeight[ i ][ j ] * CtbSizeY
[0135] In some embodiments, a new list colBd[i][j] that specifies the location of the column boundary of the j-th tile of the i-th region in coding tree block units, where j is in the range from 0 to num_tile_columns_minus1[i]+1, can be derived as follows. colBd[i][0] = ( i == 0 )? 0 : colBd[ i -1 ]
[0000] + regionColBd[ i-1 ] colBd[i][0] = ( colBd[i]
[0000] == PicWidthInCtbsY )? 0 : colBd[i][0] for ( j = 0; j <= num_tile_columns_minus1[i]; j++ ) colBd[ i ][ j + 1 ] = colBd[i][j] + colWidth[i][j]
[0136] In some embodiments, for j in the range from 0 to num_tile_rows_minus1[i]+1, a new list rowBd[i][j] that specifies the location of the row boundary of the j-th tile in the i-th region in units of the coded tree block can be derived as follows. rowBd[ i ][0] = ( i == 0 )? 0 : rowBd[ i -1 ][0] + regionRowBd[ i-1 ] rowBd[ i ][0] = (rowBd[ i ][0] == PicHeightInCtbsY)? 0 : rowBd[ i ][0] for( j = 0; j <= num_tile_rows_minus1[ i ]; j++ ) rowBd[ i ][ j + 1 ] = rowBd[ i ][ j ] + rowHeight[ i ][ j ]
[0137] In some embodiments, for ctbAddrRs in the range from 0 to PicSizeInCtbsY-1, a list CtbAddrRsToTs[ctbAddrRs] that specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the region-based tile scan can be derived as follows. for ( ctbAddrRs = 0; ctbAddrRs < PicSizeInCtbsY; ctbAddrRs++ ) { tbX = ctbAddrRs % PicWidthInCtbsY tbY = ctbAddrRs / PicWidthInCtbsY for ( i = 0; i <= num_region_columns_minus1; i++ ) if ( tbX >= regionColBd[ i ]) regionX = i for (j = 0; j <= num_region_rows_minus1; j++) if (tbY >= regionRowBd[j]) regionY = j regionId = regionY * (num_region_columns_minus1 + 1) + regionX for (i = 0; i <= num_tile_columns_minus1[regionId]; i++) if (tbX >= colBd[regionId][i]) tileX = i for (j = 0; j <= num_tile_rows_minus1[regionId]; j++) if (tbY >= rowBd[regionId][j]) tileY = j CtbAddrRsToTs[ctbAddrRs] = 0 for (i = 0; i < regionId; i++) CtbAddrRsToTs[ctbAddrRs] += RegionSizeInCtbsY[i] for (i = 0; i < tileX; i++) CtbAddrRsToTs[ctbAddrRs] += rowHeight[regionId][tileY] * colWidth[regionId][i] for (j = 0; j < tileY; j++) CtbAddrRsToTs[ctbAddrRs] += RegionWidthInCtbsY[regionId] * rowHeight[regionId][j] CtbAddrRsToTs[ctbAddrRs] += (tbY - rowBd[regionId][tileY]) * colWidth[regionId][tileX] + tbX - colBd[regionId][tileX] }
[0138] For ctbAddrTs in the range from 0 to PicSizeInCtbsY - 1, which specifies the conversion from the CTB address in the region-based tile scan to the CTB address in the CTB raster scan of the picture, the list CtbAddrTsToRs[ctbAddrTs] can be derived as follows. for(ctbAddrRs = 0; ctbAddrRs < PicSizeInCtbsY; ctbAddrRs++) CtbAddrTsToRs[CtbAddrRsToTs[ctbAddrRs]] = ctbAddrRs2 For ctbAddrTs in the range from 0 to PicSizeInCtbsY - 1, which specifies the conversion from the CTB address in the tile scan to the tile index or ID, the list TileId[ctbAddrTs] can be derived as follows. for (n = 0, tileIdx = 0, regionIdx = 0; n <= num_region_rows_minus1; n++) for(m = 0; m <= num_region_columns_minus1; m++, regionIdx ++) for(j = 0, j <= num_tile_rows_minus1[regionIdx]; j++) for(i = 0; i <= num_tile_columns_minus1[regionIdx]; i++, tileIdx++) for( y = rowBd[regionIdx][ j ]; y < rowBd[regionIdx][ j + 1 ]; y++ ) for( x = colBd[regionIdx][ i ]; x < colBd[regionIdx][ i + 1 ]; x++ ) TileId[ CtbAddrRsToTs[ y * PicWidthInCtbsY+ x ] ] = tileIdx
[0139] In an alternative embodiment, the tile identifier (ID) for each region-based tile can be represented by a two-dimensional (2D) array. The first index can be the region index, and the second index can be the tile index within the region. FIG. 13 is an example of the TileID representation in picture frame 1300.
[0140] The conversion from the CTB address in the tile scan to the 2D tile IDs TileId0[ctbAddrTs] and TileId1[ctbAddrTs] (e.g., two new lists) where ctbAddrTs is in the range from 0 to PicSizeInCtbsY - 1 can be derived as follows. for ( n = 0, regionIdx = 0; n <= num_region_rows_minus1; n++ ) { for( m = 0; m <= num_tile_columns_minus1; m++; regionIdx++) { TileId0[ CtbAddrRsToTs[ y * PicWidthInCtbsY+ x ] ] = regionIdx for( j = 0, tileIdx = 0; j <= num_tile_rows_minus1[ regionIdx ]; j++ ) for( i = 0; i <= num_tile_columns_minus1[ regionIdx ]; i++, tileIdx++ ) for( y = rowBd[regionIdx][ j ]; y < rowBd[regionIdx][ j + 1 ]; y++ ) for( x = colBd[[regionIdx][ i ]; x < colBd[regionIdx][ i + 1 ]; x++ ) TileId1[ CtbAddrRsToTs[ y * PicWidthInCtbsY+ x ] ] = tileIdx } }
[0141] V.5 Initial quantization parameters for tile coding In some embodiments, HEVC can specify an initial quantization value for each slice. One or more initial quantization parameters (QPs) can be used for the coded blocks within a slice. The initial value of the luma quantization parameter SliceQp for a slice Y is derived as follows.
[0142] SliceQp Y =26+init_qp_minus26+slice_qp_delta where init_qp_minus26 is signaled in the PPS, and slice_qp_delta is signaled in an independent slice segment header.
[0143] The chroma quantization parameter for a slice, and the coded blocks within the slice, are similarly signaled in the PPS and slice header.
[0144] For omnidirectional video processing, a set of tiles can be mapped to a viewport or a face. Each viewport or face can be encoded with a different quality (e.g., resolution) to support viewport-dependent video processing. The quantization parameter of a tile can be inferred from Y SliceQp in the slice header as specified in HEVC, or can be explicitly signaled as a tile characteristic. Y In some cases, the signaling of 360-degree video information [6] (Patent Document 1) can be used. For example, in a case where a specific face is encoded with a higher or lower quality than another face, the QP for each face can be explicitly signaled. Encoded tree blocks belonging to the same face can share the same initial QP signal for the face.
[0145] In some embodiments, QP can be signaled at the region and / or tile level so that all tiles belonging to the same region can share the same initial region QP. Alternatively, each tile can have its own initial QP value based on the initial region QP and the QP offset value of the individual tile. Table 4 shows an exemplary signaling structure according to such an embodiment.
[0146] The region_qp_offset_enabled_flag specifies whether different QPs are used for different regions.
[0147]
Table 4
[0148] The region_qp_offset_enabled_flag specifies whether different QPs are used for different regions.
[0149] region_qp_offset[i] specifies the initial value of QP that is used for the tiles within the region until it is changed by the value of tile_qp_offset within the coding unit layer. Qp for the i-th region Y Quantization parameter RegionQp Y The initial value of [i] can be derived as follows. RegionQp Y [i] = 26 + init_qp_minus26 + region_qp_delta[i]
[0150] tile_qp_offset_enabled_flag specifies whether different QPs are used for different tiles.
[0151] tile_qp_offset[i][m][n] specifies the initial value of QP that is used for the coding blocks within the tile at the position [m][n] of the i-th region. When it does not exist, the value of tile_qp_offset can be presumed to be equal to 0. The value of the quantization parameter TileQpY[i][m][n] can be derived as follows. TileQp Y [ i ][ m ][ n ] = RegionQp Y [ i ] + tile_qp_delta[ i ][ m ][ n ]
[0152] The QP of each tile can be specified in tile index order. The tile index can be derived from the region index, as well as the tile column and row values, as follows. for ( tileIdx = 0, i = 0; i < NumRegions; i++ ) for ( m = 0; m <= num_tile_rows_minus1[i]; m++) for ( n = 0; n <= num_tile_cols_minus1[i]; n++, tileIdx++) TileQpY[tileIdx] = RegionQpY[ i ] + tile_qp_delta[ i ][ m ][ n ]
[0153] In an alternative embodiment, the tile QP offset can be specified in a list, and each tile can derive its initial QP value by referring to the corresponding table index. Table 5 shows an exemplary QP offset list, and Table 6 shows an exemplary tile QP format.
[0154]
Table 5
[0155] tile_qp_offset_list_len_minus1 plus 1 specifies the number of elements of the tile_qp_offset_list syntax. The tile_qp_offset_list specifies a list of QP offset values used in the derivation of the tile QP from the initial QP.
[0156]
Table 6
[0157] tile_qp_offset_idx specifies the index to the tile_qp_offset_list used to determine the value of TileQpOffset Y When it exists, the value of tile_qp_offset_idx shall be in the range of 0 or more and tile_qp_offset_list_len_minus1 or less.
[0158] In some embodiments, the variable TileQpOffset Y [i] of the i-th tile, and TileQp Y [i] can be derived as follows. TileQpOffset Y[i] = tile_qp_offset_list[ tile_qp_offset_idx ] TileQp Y [i] = 26 + init_qp_minus26 + TileQpOffset Y [i]
[0159] The following references, namely, [1] Non-Patent Document 1, [2] Non-Patent Document 2, [3] Non-Patent Document 3, [4] Non-Patent Document 4, [5] Non-Patent Document 5, [6] Patent Document 1, [7] Patent Document 2, and [8] Patent Document 3 are hereby incorporated by reference into this specification.
[0160] VII. Conclusion Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented by a computer program, software, or firmware incorporated within a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software can be used to implement a radio frequency transceiver for use in a WTRU102, UE, terminal, base station, RNC, or any host computer.
[0161] Furthermore, in the embodiments described above, reference was made to a processing platform, a computing system, a controller, and other devices including a processor. These devices can include at least one central processing unit (“CPU”) and a memory. In accordance with the practice of those skilled in the art of computer programming, references to acts and to symbolic representations of operations or instructions can be performed by various CPUs and memories. Such acts and operations or instructions are sometimes said to be “executed,” “executed by a computer,” or “executed by a CPU.”
[0162] Those skilled in the art will understand that acts and symbolic representations of operations or instructions include the manipulation of electrical signals by a CPU. An electrical system represents data bits, which can cause the conversion or reduction of resulting electrical signals and the maintenance of data bits at memory locations within a memory system, thereby restructuring or otherwise changing the operation of the CPU and other processing of signals. The memory locations where data bits are maintained are physical locations having specific electrical, magnetic, optical, or organic characteristics corresponding to or representing the data bits. Representative embodiments are not limited to the platforms or CPUs mentioned above, and it should be understood that other platforms and CPUs can support the provided methods.
[0163] The data bits can also be maintained on a computer-readable medium including a magnetic disk, an optical disk, and any other volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage system readable by a CPU. The computer-readable medium can include cooperative or interconnected computer-readable media, which can be distributed among multiple interconnected processing systems that exist exclusively on the processing system or can be local or remote to the processing system. Representative embodiments are not limited to the memories mentioned above, and it is understood that other platforms and memories can support the methods described.
[0164] In an illustrative embodiment, any of the operations, processes, etc. described herein can be implemented as computer-readable instructions stored on a computer-readable medium. The computer-readable instructions can be executed by a processor of a mobile unit, a network element, and / or any other computing device.
[0165] The differences remaining between the hardware implementation and the software implementation of the system aspects are few. Whether to use hardware or software is generally (for example, but in some situations, the choice between hardware and software may become important, so not always) a design choice representing a cost - effectiveness trade - off. Various means (such as hardware, software, and / or firmware) by which the processes and / or systems, and / or other technologies described herein may be affected can exist, and the preferred means may vary with the circumstances in which the processes and / or systems, and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are of the highest priority, the implementer can primarily select hardware and / or firmware means. If flexibility is of the highest priority, the implementer can primarily select a software implementation. Alternatively, the implementer can select some combination of hardware, software, and / or firmware.
[0166] The foregoing detailed description has described various embodiments of devices and / or processes via the use of block diagrams, flowcharts, and / or examples. As long as such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, each function and / or operation within such block diagrams, flowcharts, or examples can be implemented individually and / or collectively by a wide range of hardware, software, firmware, or substantially any combination thereof, as will be understood by those skilled in the art. Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), application specific standard products (ASSPs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.
[0167] Those skilled in the art will understand that the features and elements were provided above in certain combinations, but each feature or element can be used alone or in any combination with other features and elements. This disclosure should not be limited with respect to the specific embodiments described in this application, which were intended as examples of various aspects. As will be apparent to those skilled in the art, many changes and modifications can be made without departing from its spirit and scope. Elements, acts, or instructions used in the description of this application should not be construed as important or essential to the invention unless expressly provided as such. In addition to those listed herein, functionally equivalent methods and apparatuses within the scope of this disclosure will be apparent to those skilled in the art from the foregoing description. Such changes and modifications are intended to be included within the scope of the appended claims. This disclosure should be limited only by the claims of the appended claims, along with the full scope of equivalents to which such claims are entitled. It should be understood that this disclosure is not limited to a particular method or system.
[0168] It should be understood that the terms used in this specification are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in this specification and when referred to herein, the terms "station" and its abbreviation "STA", "user equipment" and its abbreviation "UE" can mean (i) a wireless transmit and / or receive unit (WTRU) as described below, (ii) any of the various embodiments of a WTRU as described below, (iii) a wireless and / or wireline (e.g., connectable) device configured to use, among other things, some or all of the structure and functionality of a WTRU as described below, (iii) a wireless and / or wireline device configured to use less structure and functionality than all of a WTRU as described below, or (iv) something similar. Details of exemplary WTRUs that can represent (or be interchangeable with) any of the UEs or mobile devices listed herein are provided below with respect to FIGS. 1A-1D.
[0169] In one representative embodiment, some portions of the invention described herein can be implemented via application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, one of ordinary skill in the art will recognize that some aspects of the embodiments disclosed herein can be equivalently implemented, in whole or in part, as one or more computer programs operating on one or more computers (e.g., as one or more programs operating on one or more computer systems), as one or more programs operating on one or more processors (e.g., as one or more programs operating on one or more microprocessors), as firmware, or as substantially any combination thereof, and that designing the circuitry and / or writing code for software and / or firmware is well within the skill of one of ordinary skill in the art in view of the present disclosure. Additionally, one of ordinary skill in the art will understand that the mechanisms of the invention described herein can be distributed in a variety of forms as a program product, and that illustrative embodiments of the invention described herein apply regardless of the particular type of signal bearing medium used to actually carry out the distribution. Examples of signal bearing media include, but are not limited to, recordable type media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, computer memories, etc., and transmission type media such as digital and / or analog communication media (e.g., optical fiber cables, waveguides, wired communication links, wireless communication links, etc.).
[0170] The present invention described herein sometimes exemplifies different components that are included within or connected to other different components. It should be understood that such depicted architectures are merely examples and that in practice many other architectures that achieve the same functionality can be implemented. In a conceptual sense, any arrangement of components for achieving the same functionality is effectively "associated" so as to be able to achieve the desired functionality. Thus, any two components herein that are combined to achieve a particular functionality can be seen as "associated" with each other such that the desired functionality is achieved, regardless of the architecture or intervening components. Similarly, any two components so associated can also be regarded as "operably connected" or "operably coupled" to each other for achieving the desired functionality, and any two components that can be so associated can also be regarded as "operably couplable" to each other for achieving the desired functionality. Particular examples of operably couplable include, but are not limited to, components that can be physically paired and / or physically interact, and / or wirelessly interact and / or wirelessly communicate, and / or logically interact and / or logically communicate with each other.
[0171] Regarding the use of substantially any plural and / or singular terms herein, one of ordinary skill in the art can convert from plural to singular and / or from singular to plural as appropriate to the situation or application. For clarity, various singular / plural substitutions may be explicitly recited herein.
[0172] Generally, in this specification, and in particular in the appended claims (e.g., the body of the appended claims), it will be understood by those skilled in the art that the terms used are generally intended to be "open" terms (e.g., the term "including" should be construed as "including, but not limited to", the term "having" should be construed as "having at least", the term "includes" should be construed as "including, but not limited to", etc.). When a specific number of claim recitations is intended, such intent will be expressly recited in the claim, and it will be further understood by those skilled in the art that when there is no such recitation, no such intent exists. For example, when only one item is intended, the term "single" or similar words can be used. For purposes of understanding, the following appended claims, and / or the description in this specification, can include the use of introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed as implying that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation (e.g., "a" and / or "an" should be construed as meaning "at least one" or "one or more"). The same applies to the use of definite articles used to introduce claim recitations. Additionally, even when a specific number of introduced claim recitations is expressly recited, those skilled in the art will recognize that such recitation should be construed as meaning at least the recited number (e.g., an unmodified recitation of "two recitations" without other modifying phrases means at least two recitations, or two or more recitations).
[0173] Furthermore, when conventional expressions similar to "at least one of A, B, and C" are used, generally, such syntax is intended in the sense understood by one of ordinary skill in the art (e.g., a "system having at least one of A, B, and C" includes, but is not limited to, a system having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together). When conventional expressions similar to "at least one of A, B, or C" are used, generally, such syntax is intended in the sense understood by one of ordinary skill in the art (e.g., a "system having at least one of A, B, or C" includes, but is not limited to, a system having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together). It will be further understood by one of ordinary skill in the art that any substantially disjunctive words and / or phrases presenting two or more alternatives, whether within the description, within the claims, or within the drawings, are to be understood as contemplating the possibility of including one of the items, either of the items, or both of the items. For example, the phrase "A or B" is understood to include the possibilities of "A", or "B", or "A and B". Furthermore, as used herein, the term "any of" followed by a list of multiple items and / or multiple categories of items is intended to include "any of" the items and / or categories of items, "any combination", "any plurality", and / or "any combination of a plurality" of the items and / or categories of items, either individually or in combination with other items and / or other categories of items. Furthermore, as used herein, the term "set" or "group" is intended to include any number of items, including zero. Additionally, as used herein, the term "number" is intended to include any number, including zero.
[0174] In addition, when a feature or aspect of the present disclosure is described with respect to a Markush group, one of ordinary skill in the art will recognize that the present disclosure is thereby also described with respect to any individual member or subgroup of members of the Markush group.
[0175] As will be understood by those of ordinary skill in the art, for all purposes, such as for providing a written description, all ranges disclosed herein include any and all possible sub-ranges, and combinations thereof. Any recited range can be readily recognized as also fully describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range disclosed herein can be readily broken down into a lower third, middle third, and upper third, etc. Also as will be understood by those of ordinary skill in the art, all words such as "up to," "at least," "greater than," and "less than" include the recited number and refer to ranges that can later be divided into sub-ranges as described above. Finally, as will be understood by those of ordinary skill in the art, a range includes each individual member. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on for others.
[0176] Furthermore, the claims should not be read as being limited to the order or elements provided unless stated to that effect in the claims. In addition, the use of the term "means for" in any claim is intended to invoke 35 U.S.C. § 112, paragraph 6, or a means-plus-function claim format, and any claim that does not include the term "means for" is not intended to be read as such.
[0177] A radio frequency transceiver can be implemented for use in a wireless transmit receive unit (WTRU), user equipment (UE), terminal, base station, mobility management entity (MME) or evolved packet core (EPC), or any host computer, using a processor associated with software. The WTRU can be used in conjunction with other components, including hardware, and / or software modules implemented in software including software defined radio (SDR), as well as cameras, video camera modules, videophones, speakerphones, vibration devices, speakers, microphones, television transceivers, hands-free headsets, keyboards, Bluetooth® modules, frequency modulation (FM) radio units, near field communication (NFC) modules, liquid crystal display (LCD) display units, organic light emitting diode (OLED) display units, digital music players, media players, video game player modules, Internet browsers, and / or wireless local area network (WLAN) or ultra wideband (UWB) modules.
[0178] Although the invention has been described in relation to a communication system, it is contemplated that the system can be implemented in software on a microprocessor / general purpose computer (not shown). In certain embodiments, one or more of the functions of the various components can be implemented in software that controls a general purpose computer.
[0179] In addition, although the invention has been illustrated and described herein with reference to specific embodiments, the invention is not intended to be limited to the details shown. Rather, various changes can be made in the details within the scope and range of equivalents of the claims and without departing from the invention.
[0180] Those skilled in the art will understand that throughout this disclosure, certain representative embodiments can be used selectively or in combination with other representative embodiments.
[0181] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented by a computer program, software, or firmware included in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). Using a processor associated with software, a radio frequency transceiver for use in a WRTU, UE, terminal, base station, RNC, or any host computer can be implemented.
[0182]
[0183] One of ordinary skill in the art will understand that acts, and symbolic acts or instructions, include the manipulation of electrical signals by a CPU. An electrical system represents data bits, which can cause the conversion or reduction of the resulting electrical signals and the maintenance of data bits at memory locations within a memory system, thereby reconstructing or otherwise changing the operation of the CPU and other processing of the signals. The memory location where a data bit is maintained is a physical location having specific electrical, magnetic, optical, or organic characteristics corresponding to or representing the data bit.
[0184] Data bits can also be maintained on a computer-readable medium, including magnetic disks, optical disks, and any other volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage systems readable by a CPU. The computer-readable medium can include cooperative or interconnected computer-readable media, which can be distributed among multiple interconnected processing systems that exist exclusively on a processing system or can be local or remote to the processing system. Representative embodiments are not limited to the memories mentioned above, and it is understood that other platforms and memories can support the described methods.
[0185] Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.
[0186] Although the present invention has been described with respect to a communication system, it is contemplated that the system can be implemented in software on a microprocessor / general purpose computer (not shown). In certain embodiments, one or more of the functions of the various components can be implemented in software that controls a general purpose computer.
[0187] In addition, although the present invention has been illustrated and described herein with reference to particular embodiments, the present invention is not intended to be limited to the details shown. Rather, various changes may be made in the details within the scope and range of equivalents of the claims and without departing from the invention.
Industrial Applicability
[0188] The present invention can be used for communication.
Explanation of Signs
[0189] 102, 102a to 102d WTRU 104, 113 RAN 106, 115 Core network 108 PSTN 110 Internet 112 Other network 160a, 160b eNodeB 180a, 180b gNB 118 Processor 120 Transceiver 122 Antenna
Claims
1. A method for encoding video information, comprising: determining a first set of parameters defining a plurality of first grid regions in a frame; for each first grid region, determining a second set of parameters defining a plurality of second grid regions, wherein the plurality of second grid regions partition respective first grid regions; partitioning the frame into the plurality of first grid regions based on the first set of parameters; partitioning each first grid region into the plurality of second grid regions based on respective second sets of parameters; encoding the frame based on the partitioned first and second grid regions; and a method comprising the steps of.
2. The method according to claim 1, wherein the first set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid.
3. The method according to claim 1, wherein for each first grid region, each second set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid.
4. The method according to claim 1, wherein image samples of the frame that cross the boundaries of the first grid regions are not continuous.
5. The method according to claim 1, wherein image samples of the frame mapped to one of the first grid regions are encoded at a first resolution, and image samples of the frame mapped to a second of the first grid regions are encoded at a second resolution.
6. generating a padding flag indicating whether a padding operation is to be performed on each edge of the grid regions in the plurality of first grid regions; encoding the padding flag; The method according to claim 1, further comprising the steps of.
7. A method for decoding video information, comprising: receiving a first set of parameters defining a plurality of first grid regions in a frame; Receiving a second set of parameters that define a plurality of second grid regions for each first grid region, wherein the plurality of second grid regions partition respective first grid regions; Partitioning the frame into the plurality of first grid regions based on the first set of parameters; Partitioning each first grid region into the plurality of second grid regions based on each second set of parameters; Decoding the frame based on the partitioned first and second grid regions; A method comprising the steps of. **Claim 8** The method of claim 7, wherein the first set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid. **Claim 9** The method of claim 7, wherein for each first grid region, each second set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid. **Claim 10** The method of claim 7, wherein the first set of parameters and the second set of parameters are received in any of a sequence parameter set, a picture parameter set, or a slice header. **Claim 11** Receiving a padding flag indicating whether a padding operation is to be performed on each edge of each grid region of the plurality of first grid regions; Performing padding on the edges of the grid regions based on the padding flag; The method of claim 7, further comprising the steps of. **Claim 12** An apparatus for encoding video information, comprising: At least one processor; When executed by the at least one processor, causing the apparatus to: Determine a first set of parameters that define a plurality of first grid regions in a frame; For each first grid region, determine a second set of parameters that define a plurality of second grid regions, each of the plurality of second grid regions partitioning a respective first grid region, Partition the frame into the plurality of first grid regions based on the first set of parameters, Partition each first grid region into the plurality of second grid regions based on each of the second set of parameters, Encode the frame based on the partitioned first grid regions and second grid regions A memory storing instructions to cause the above, An apparatus comprising the above.
13. The apparatus according to claim 12, wherein the first set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid.
14. The apparatus according to claim 12, wherein for each first grid region, each of the second set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid.
15. The apparatus according to claim 12, wherein image samples of the frame that cross the boundary of the first grid region are not continuous.
16. The apparatus according to claim 12, wherein image samples of the frame mapped to one of the first grid regions are encoded at a first resolution, and image samples of the frame mapped to a second one of the first grid regions are encoded at a second resolution.
17. The instructions cause the apparatus to, Generate a padding flag indicating whether a padding operation is to be performed on each of the edges of the grid regions in the plurality of first grid regions, Encode the padding flag The apparatus according to claim 12.
18. An apparatus for decoding video information, comprising At least one processor, and When executed by the at least one processor, cause the apparatus to, Receive, in a frame, a first set of parameters that define a plurality of first grid regions, Receive a second set of parameters that define a plurality of second grid regions for each of the first grid regions, wherein the plurality of second grid regions each partition a respective first grid region, Partition the frame into the plurality of first grid regions based on the first set of parameters, Partition each first grid region into the plurality of second grid regions based on each of the second sets of parameters, Store in a memory instructions to decode the frame based on the partitioned first and second grid regions and an apparatus comprising the same. **Claim 19** The apparatus according to claim 18, wherein the first set of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid. **Claim 20** The apparatus according to claim 18, wherein for each first grid region, each of the second sets of parameters includes a flag indicating that the frame is to be partitioned according to a uniform grid. **Claim 21** The apparatus according to claim 18, wherein the first set of parameters and the second set of parameters are received in any of a sequence parameter set, a picture parameter set, or a slice header. **Claim 22** The instructions cause the apparatus to receive a padding flag indicating whether a padding operation is to be performed on each of the edges of each grid region of the plurality of first grid regions, and further perform padding on the edges of the grid regions based on the padding flag The apparatus according to claim 18.
Citation Information
Patent Citations
Intra prediction from predictive blocks using displacement vectors
JP2016525303A
Method for encoding 360-degree panoramic video, encoding device, and computer program
JP2018534827A
Video encoding / decoding device, method, and computer program
JP2019513320A
Tile Group Partitioning
US62775130P0
Tile group partitioning
US62781749P0