Methods and apparatus for flexible grid regions

By employing flexible tile partitioning and geometric structure filling techniques, the problem of low efficiency in image and video encoding in existing technologies has been solved, enabling more efficient signal notification and processing.

CN112703734BActive Publication Date: 2026-03-13INTERDIGITAL VC HOLDINGS INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-13
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing image and video encoding technologies are inefficient when processing flexible grid areas, making it difficult to achieve efficient signal notification and processing.

Method used

Flexible tile partitioning and geometric structure filling techniques are employed to optimize the encoding and decoding process of image frames through flexible tile division and region-based signaling methods.

Benefits of technology

It improves the encoding efficiency of image and video frames, and enhances the flexibility and efficiency of signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112703734B_ABST
    Figure CN112703734B_ABST
Patent Text Reader

Abstract

Methods and apparatus for using flexible grid regions in image or video frames are disclosed. In one embodiment, a method includes receiving a set of first parameters defining a plurality of first grid regions comprising a frame. For each first grid region, the method includes receiving a set of second parameters defining a plurality of second grid regions, wherein the plurality of second grid regions divide the corresponding first grid regions. The method further includes dividing the frame into the plurality of first grid regions based on the set of first parameters, and dividing each first grid region into the plurality of second regions based on the corresponding set of second parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] The embodiments disclosed herein primarily relate to signaling and processing image or video information. For example, one or more embodiments disclosed herein relate to methods and apparatus for using flexible grid regions or tiles in image / video frames. Attached Figure Description

[0002] A more detailed understanding can be obtained from the following detailed description given by way of example in conjunction with the accompanying drawings. The drawings in the description are examples. Therefore, the drawings and detailed description should not be considered limiting, and other equally valid examples are also possible. Furthermore, the same reference numerals in the drawings denote the same elements, and wherein:

[0003] Figure 1A This is a system diagram illustrating an exemplary communication system in which one or more of the disclosed embodiments may be implemented.

[0004] Figure 1B This illustrates the possibility of implementation according to an embodiment. Figure 1A A system diagram of an exemplary wireless transmit / receive unit (WTRU) used within a communication system shown;

[0005] Figure 1C This illustrates the possibility of implementation according to an embodiment. Figure 1A A system diagram of an exemplary radio access network (RAN) and an exemplary core network (CN) used within the communication system shown;

[0006] Figure 1D This illustrates the possibility of implementation according to an embodiment. Figure 1A A system diagram of another exemplary RAN and another exemplary CN used within the communication system shown;

[0007] Figure 2A This is an example of an HEVC tile partition according to one or more embodiments, the HEVC tile partition having tile columns and tile rows evenly distributed on an image;

[0008] Figure 2B This is an example of an HEVC tile partition according to one or more embodiments, the HEVC tile partition having tile columns and tile rows that are not uniformly distributed on the image;

[0009] Figure 3 This is an illustration of an example of a repeating padding scheme that replicates sample values ​​from image boundaries according to one or more embodiments;

[0010] Figure 4 This is a diagram illustrating an example of a geometry filling process using an equal rectangular projection (ERP) format according to one or more embodiments;

[0011] Figure 5 This is a diagram illustrating an example of merging HEVC MCTS-based regional orbits with the same resolution according to one or more embodiments;

[0012] Figure 6 This is a diagram illustrating an example of cube map (CMP) partitioning according to one or more embodiments;

[0013] Figure 7 This is a diagram illustrating an example of CMP partitioning with a slice header according to one or more embodiments;

[0014] Figure 8 This is a diagram illustrating an example of a preprocessing and encoding scheme for achieving (HEVC-based) 6K effective ERP resolution according to one or more embodiments;

[0015] Figure 9A This is a diagram illustrating an example of partitioning using conventional blocks according to one or more embodiments;

[0016] Figure 9B This is a diagram illustrating an example of partitioning using flexible blocks according to one or more embodiments;

[0017] Figure 10 This is a diagram illustrating an example of geometric filling for flexible blocks according to one or more embodiments;

[0018] Figure 11A This is a diagram illustrating a first example of flexible tile signaling based on a region according to one or more embodiments;

[0019] Figure 11B This is a diagram illustrating a second example of flexible tile signaling based on a region according to one or more embodiments;

[0020] Figure 12A This is an illustration showing an example of a decoded tree block (CTB) raster scan of an image according to one or more embodiments;

[0021] Figure 12B This is a diagram illustrating an example of CTB raster scanning of conventional tiles according to one or more embodiments;

[0022] Figure 12C This is a diagram illustrating an example of CTB raster scanning based on region-based flexible tiles according to one or more embodiments; and

[0023] Figure 13 This is a diagram illustrating an example of using a corresponding tile identifier for each region-based tile according to one or more embodiments. Detailed Implementation

[0024] I. Exemplary Networks and Devices

[0025] Figure 1A This diagram illustrates an exemplary communication system 100 that can implement one or more of the disclosed embodiments. The communication system 100 can be a multiple access system providing voice, data, video, messaging, broadcasting, and other content to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 can use one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtering OFDM, and Filter Bank Multicarrier (FBMC), etc.

[0026] like Figure 1A As shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network components. Each WTRU 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, any WTRU 102a, 102b, 102c, or 102d may be referred to as a “station” and / or “STA”, and may be configured to transmit and / or receive wireless signals. It may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks, etc. Any of WTRU 102a, 102b, 102c, or 102d may be interchangeably referred to as a UE.

[0027] The communication system 100 may further include base station 114a and / or base station 114b. Each base station 114a, 114b may be any type of device configured to enable access to one or more communication networks (e.g., CN 106 / 115, Internet 110, and / or other networks 112) by wirelessly interfacing with at least one of WTRUs 102a, 102b, 102c, 102d. For example, base stations 114a, 114b may be base transceiver stations (BTS), node B, e-node B, home node B, home e-node B, gNB, new radio (NR) node B, site controller, access point (AP), and wireless routers, etc. Although each base station 114a, 114b is described as a single component, it should be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network components.

[0028] Base station 114a may be part of RAN 104 / 113, and the RAN may also include other base stations and / or network components (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies called cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide radio service coverage for a specific geographic area that is relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, that is, each transceiver corresponds to one sector of the cell. In embodiments, base station 114a may use multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, by using beamforming, signals can be transmitted and / or received in a desired spatial direction.

[0029] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, wherein the air interface can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).

[0030] More specifically, as described above, the communication system 100 can be a multiple access system and can use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, and 102c can implement a certain radio technology, such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), wherein the technology can use Wideband CDMA (WCDMA) to establish the air interface 115 / 116. WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0031] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a certain radio technology, such as Evolved UMTS Terrestrial Radio Access (E-UTRA), wherein the technology may use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTA Pro (LTE-A Pro) to establish air interface 116.

[0032] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a certain radio technology, such as NR radio access, wherein the radio technology may use a novel radio (NR) to establish air interface 116.

[0033] In this embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access (e.g., using the dual connectivity (DC) principle). Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or transmissions sent to / from various types of base stations (e.g., eNBs and gNBs).

[0034] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement the following radio technologies, such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (Global Microwave Access Interoperability (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN), etc.

[0035] Figure 1A Base station 114b can be a wireless router, home node B, home e node B, or access point, and can use any suitable RAT to facilitate wireless connectivity in a local area, such as a business premises, residence, vehicle, campus, industrial facility, air corridor (e.g., for use by drones), and road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can establish a wireless local area network (WLAN) by implementing a radio technology such as IEEE 802.11. In another embodiment, base station 114b and WTRUs 102c, 102d can establish a wireless personal area network (WPAN) by implementing a radio technology such as IEEE 802.15. In yet another embodiment, base station 114b and WTRUs 102c, 102d can establish a picocell or femtocell by using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). Figure 1A As shown, base station 114b can be directly connected to the Internet 110. Therefore, base station 114b does not need to access the Internet 110 via CN106 / 115.

[0036] RAN 104 / 113 can communicate with CN 106 / 115, where CN can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more WTRUs 102a, 102b, 102c, 102d. This data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or can perform advanced security functions such as user authentication. Although in Figure 1AWhile not shown, it should be understood that RAN104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs that use the same RAT or a different RAT as RAN 104 / 113. For example, in addition to connecting to RAN 104 / 113 which uses NR radio technology, CN 106 / 115 can also communicate with other RANs (not shown) that use GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technologies.

[0037] CN 106 / 115 may also act as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Simple Old-Style Telephone Service (POTS). The Internet 110 may include a global network of interconnected computer equipment systems using common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, wherein the one or more RANs may use the same RAT or a different RAT as RAN 104 / 113.

[0038] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capability (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers communicating with different wireless networks on different wireless links). For example... Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a, which can use cellular-based radio technology, and with base station 114b, which can use IEEE 802 radio technology.

[0039] Figure 1B This is a system diagram illustrating an example of WTRU 102. (See diagram below.) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive unit 122, a speaker / microphone 124, a keyboard 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and other peripheral devices 138. It should be understood that, while remaining consistent with the embodiments, WTRU 102 may also include any sub-combination of the foregoing components.

[0040] Processor 118 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. Processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables WTRU 102 to operate in a wireless environment. Processor 118 can be coupled to transceiver 120, and transceiver 120 can be coupled to transmitting / receiving unit 122. Although Figure 1B While the processor 118 and transceiver 120 are described as separate components, it should be understood that the processor 118 and transceiver 120 can also be integrated into a single electronic component or chip.

[0041] Transmit / receive component 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmit / receive component 122 may be an antenna configured to transmit and / or receive RF signals. As an example, in an embodiment, transmit / receive component 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In an embodiment, transmit / receive component 122 may be configured to transmit and / or receive RF and optical signals. It should be understood that transmit / receive component 122 may be configured to transmit and / or receive any combination of wireless signals.

[0042] Although Figure 1B The transmit / receive component 122 is described as a single component, but the WTRU 102 may include any number of transmit / receive components 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in an embodiment, the WTRU 102 may include two or more transmit / receive components 122 (e.g., multiple antennas) that transmit and receive radio signals via the air interface 116.

[0043] Transceiver 120 can be configured to modulate signals to be transmitted by transmitter / receiver 122 and demodulate signals received by transmitter / receiver 122. As described above, WTRU 102 can have multimode capability. Therefore, transceiver 120 can include multiple transceivers that allow WTRU 102 to communicate using various RATs (e.g., NR and IEEE 802.11).

[0044] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keyboard 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data from these components. The processor 118 can also output user data to the speaker / microphone 124, keyboard 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 can include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 can include a subscriber identity module (SIM) card, memory stick, secure digital card (SD) memory card, etc. In other embodiments, the processor 118 can access and store information from memory that is not actually located in WTRU 102; for example, such memory could be located in a server or home computer (not shown).

[0045] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power for other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell battery packs (such as nickel-cadmium (Ni-Cd), nickel-zinc (Ni-Zn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, and fuel cells, etc.

[0046] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) related to the current location of the WTRU 102. As a supplement or replacement to the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116, and / or determine its location based on signal timing received from two or more nearby base stations. It should be understood that, while remaining consistent with the embodiments, the WTRU 102 may acquire location information using any suitable positioning method.

[0047] The processor 118 can also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or video), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, etc. Modules, FM radio units, digital music players, media players, video game console modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, and activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0048] WTRU 102 may include a full-duplex wireless device, wherein the reception or transmission of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous for the wireless device. The full-duplex wireless device may include an interference management unit 139 that reduces and / or substantially eliminates self-interference by means of hardware (e.g., choke coils) or by means of a processor (e.g., a separate processor (not shown) or by means of processor 118) for signal processing. In embodiments, WTRU 102 may include a half-duplex wireless device that transmits and receives some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception).

[0049] Figure 1C This is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can communicate with WTRUs 102a, 102b, and 102c using E-UTRA radio technology on air interface 116. RAN 104 can also communicate with CN 106.

[0050] RAN 104 may include eNodeBs 160a, 160b, and 160c; however, it should be understood that RAN 104 may include any number of eNodeBs while remaining consistent with the embodiments. Each eNodeB 160a, 160b, and 160c may include one or more transceivers communicating with WTRUs 102a, 102b, and 102c on air interface 116. In one embodiment, eNodeBs 160a, 160b, and 160c may implement MIMO technology. Thus, for example, eNodeB 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0051] Each eNodeB 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. For example... Figure 1C As shown, nodes B160a, 160b, and 160c can communicate with each other via the X2 interface.

[0052] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing components is described as part of the CN 106, it should be understood that any of these components may be owned and / or operated by an entity other than the CN operator.

[0053] MME 162 can connect to each eNodeB 160a, 160b, and 160c in RAN 104 via the S1 interface and can act as a control node. For example, MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, performing bearer activation / deactivation processes, and selecting a specific serving gateway during the initial attach process of WTRUs 102a, 102b, and 102c, etc. MME 162 can also provide control plane functionality for handover between RAN 104 and other RANs (not shown) using other radio technologies (such as GSM and / or WCDMA).

[0054] The SGW 164 can connect to each eNodeB 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 typically routes and forwards user data packets to / from WTRUs 102a, 102b, and 102c. Furthermore, the SGW 164 can perform other functions, such as anchoring the user plane during handover between eNBs, triggering paging processes when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c, etc.

[0055] SGW 164 can be connected to PGW 166, which can provide packet-switched network (e.g., Internet 110) access for WTRU 102a, 102b, 102c to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.

[0056] CN 106 can facilitate communication with other networks. For example, CN 106 can provide circuit-switched network (e.g., PSTN 108) access for WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and conventional landline communication equipment. For example, CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server), and the IP gateway may act as an interface between CN 106 and PSTN 108. Furthermore, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0057] Although Figure 1A-1D The WTRU is described as a wireless terminal; however, it should be understood that in some typical embodiments, such a terminal may use a wired communication interface (e.g., temporary or permanent) with the communication network.

[0058] In some typical embodiments, the other network 112 may be a WLAN.

[0059] A WLAN employing an Infrastructure Basic Services Set (BSS) model may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may access or interface with a distributed system (DS) or other type of wired / wireless network that sends traffic into and / or out of the BSS. Traffic originating outside the BSS and destined for a STA can be delivered to the STA via the AP. Traffic originating from a STA and destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP; for example, a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as point-to-point traffic. Point-to-point traffic can be sent between the source and destination STAs (e.g., directly therebetween) using Direct Link Establishment (DLS). In some typical embodiments, the DLS may use 802.11e DLS or 802.11z Channelized DLS (TDLS). A WLAN using the Independent BSS (IBSS) mode may not have an access point (AP), and STAs (STAs) within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. Here, the IBSS communication mode is sometimes referred to as a "self-organizing" communication mode.

[0060] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel (e.g., the primary channel). The primary channel can have a fixed width (e.g., a 20 MHz bandwidth) or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish connections with the AP. In some typical embodiments, Carrier-Sensed Multiple Access with Collision Avoidance (CSMA / CA) can be implemented (e.g., in an 802.11 system). For CSMA / CA, STAs, including the AP (e.g., each STA), can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that particular STA can fall back. Within a given BSS, at any given time, only one STA (e.g., only one station) can be transmitting.

[0061] High-throughput (HT) STAs can communicate using a 40MHz wide channel (e.g., by combining a 20MHz wide main channel with adjacent or non-adjacent 20MHz wide channels to form a 40MHz wide channel).

[0062] Very High Throughput (VHT) STAs can support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels or by combining two non-consecutive 80MHz channels (this combination may be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data is transmitted and passed through a segmented parser that splits the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed individually on each stream. The streams can be mapped onto two 80MHz channels, and the data can be transmitted by the STA performing the transmission. On the receiver of the STA performing the reception, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0063] 802.11af and 802.11ah support sub-1 GHz operating modes. Compared to 802.11n and 802.11ac, the channel operating bandwidth and carrier used in 802.11af and 802.11ah are reduced. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to some typical embodiments, 802.11ah can support instrument-type control / machine-type communication (e.g., MTC devices in macro coverage areas). MTCs may have certain capabilities, such as limited capabilities including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include a battery with a battery life exceeding a threshold (e.g., for maintaining a very long battery life).

[0064] For WLAN systems that can support multiple channels and channel bandwidths (e.g., 802.11n, 802.11ac, 802.11af, and 802.11ah), the WLAN system includes a channel that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by a single STA, which is derived from all STAs operating in the BSS that support the minimum bandwidth operating mode. In the example of 802.11ah, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the width of the primary channel can be 1MHz for STAs that support (e.g., only support) the 1MHz mode (e.g., MTC type devices). Carrier sensing and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy (e.g., because an STA (which only supports the 1MHz operating mode) is transmitting to the AP), then the entire available band can be considered busy even if most of the frequency band remains idle and available.

[0065] In the United States, the available frequency band for 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. Depending on the country code, the total bandwidth available for 802.11ah is 6MHz to 26MHz.

[0066] Figure 1DThis is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 can communicate with WTRUs 102a, 102b, and 102c using NR radio technology on air interface 116. RAN 113 can also communicate with CN 115.

[0067] RAN 113 may include gNBs 180a, 180b, and 180c; however, it should be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. Each gNB 180a, 180b, and 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may use beamforming to transmit and / or receive signals to and / or from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In an embodiment, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers (not shown) to WTRU 102a. A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In embodiments, gNBs 180a, 180b, and 180c may implement Cooperative Multipoint (CoMP) technology. For example, WTRU 102a can receive cooperative transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0068] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with scalable digital configurations. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can be different for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., containing different numbers of OFDM symbols and / or continuously varying absolute time lengths).

[0069] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNodeBs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobile anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can use signals in unlicensed frequency bands to communicate with gNBs 180a, 180b, and 180c. In a non-standalone configuration, WTRUs 102a, 102b, and 102c communicate / connect with gNBs 180a, 180b, and 180c simultaneously with other RANs (e.g., eNodeBs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c, as well as one or more eNodeBs 160a, 160b, and 160c, by implementing DC principles. In a non-standalone configuration, eNodeBs 160a, 160b, and 160c can act as mobile anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRUs 102a, 102b, and 102c.

[0070] Each gNB 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support network slicing, implement dual connectivity, implement interoperability processing between NR and E-UTRA, route user plane data to User Plane Functions (UPF) 184a and 184b, and route control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.

[0071] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and may include Data Network (DN) 185a, 185b. While each of the foregoing components is described as part of CN 115, it should be understood that any of these components may be owned and / or operated by entities other than the CN operator.

[0072] AMF 182a and 182b can connect to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different needs), selecting specific SMF 183a and 183b, managing registration areas, terminating NAS signaling, and mobility management, etc. AMF 182a and 1823b can use network slicing to customize the CN support provided to WTRU 102a, 102b, and 102c based on the service types used by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, and / or services for Machine Type Communication (MTC) access, etc. AMF 182 can provide control plane functions for handover between RAN 113 and other RANs (not shown) using other radio technologies (such as LTE, LTE-A, LTE-APro, and / or non-3GPP access technologies such as WiFi).

[0073] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and can configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating WTRU or UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications, etc. PDU session types can be IP-based, non-IP-based, and Ethernet-based, etc.

[0074] UPF 184a and 184b can be connected to one or more gNBs 180a, 180b, and 180c in RAN 113 via the N3 interface, thus providing WTRU 102a, 102b, and 102c with access to a packet-switched network (e.g., Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multihomed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring processing, etc.

[0075] CN 115 can facilitate communication with other networks. For example, CN 115 may include or can communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN 115 and PSTN 108. Furthermore, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to local data networks (DNs) 185a and 185b via the N3 interface connected to UPFs 184a and 184b and the N6 interface between UPFs 184a and 184b and DNs 185a and 185b.

[0076] In view of Figure 1A-1D And about Figure 1A-1D The corresponding descriptions herein refer to one or more of the functions described below, which can be performed by one or more emulation devices (not shown): WTRU 102a-d, Base Station 114a-b, eNodeB 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN 185a-b, and / or any other devices (one or more) described herein. These emulation devices can be one or more devices configured to simulate one or more of the functions described herein. For example, these emulation devices can be used to test other devices and / or simulate network and / or WTRU functions.

[0077] The simulation equipment may be designed to perform one or more tests on other devices in a laboratory environment and / or a carrier network environment. For example, the one or more simulation devices may perform one or more functions while being implemented and / or deployed, wholly or partially, as part of a wired and / or wireless communication network, to test other devices within the communication network. The one or more simulation devices may perform one or more functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation equipment may be directly coupled to other devices to perform tests, and / or may use over-the-air wireless communication to perform tests.

[0078] The one or more simulation devices can perform one or more functions, including all functionalities, without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation devices can be used in test laboratories and / or test scenarios where wired and / or wireless communication networks are not deployed (e.g., under test) to perform tests on one or more components. The one or more simulation devices can be test equipment. The simulation devices can transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (which, as an example, may include one or more antennas).

[0079] Video decoding systems can be used to compress digital video signals, which can reduce storage requirements and / or transmission bandwidth of video signals on networks such as any of the networks mentioned above. Video coding systems can include block-based systems, wavelet-based systems, and / or object-based systems. Block-based video decoding systems can be based on, use, conform to, or comply with one or more standards, such as MPEG-1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, High Efficiency Video Decoding (HEVC), and / or Universal Video Decoding (VVC). Block-based video decoding systems can include block-based hybrid video decoding frameworks.

[0080] In some instances, a video streaming device may include one or more video encoders, each capable of producing video bitstreams at different resolutions, frame rates, or bit rates. The video streaming device may include one or more video decoders, each capable of detecting and / or decoding the encoded video bitstream. In various embodiments, the one or more video encoders and / or one or more decoders may be implemented in a device having a processor communicatively coupled to a memory, a receiver, and / or a transmitter. The memory may include instructions executable by the processor, including instructions for performing any of the various embodiments disclosed herein (e.g., representative processes). In various embodiments, the device may be configured to and / or equipped with various elements of a wireless transmit and receive unit (WTRU). Figures 1A to 1D Detailed examples of WTRUs and their components are provided in the accompanying disclosures.

[0081] II.HEVC

[0082] II.1 High-Efficiency Video Coding (HEVC) Tile

[0083] In some instances, video frames may be divided into slices and / or tiles. A slice is a sequence of one or more segments that begins with an independent segment and contains all subsequent subordinate segments. A tile is rectangular and contains an integer number of decoder tree units as specified by HEVC. For each slice and tile, one or both of the following conditions will be satisfied (e.g., see [1]): 1) all decoder tree units in a slice belong to the same tile; and / or 2) all decoder tree units in a tile belong to the same slice.

[0084] In some examples, the tile structure in HEVC is signaled in the Picture Parameter Set (PPS) by specifying the height of the rows and the width of the columns. Individual rows(s) and / or columns(s) can have different(s) sizes, but the division can always span the entire image from left to right or from top to bottom.

[0085] In some examples, HEVC tile syntax can be used. In one example, as shown in Table 1, the first flag `tile_enabled_flag` can be used to specify whether tiles are used. For example, if this first flag (`tile_enabled_flag`) is set, the number of columns and rows of tiles is specified. The second flag `uniform_spacing_flag` can be used to specify whether tile column boundaries and similar tile row boundaries are evenly distributed across the image. For example, when `uniform_spacing_flag` equals zero (0), the syntax elements `column_width_minus1[i]` and `row_height_minus1[i]` are explicitly signaled to specify the column width and row height. Additionally, the third flag `loop_filter_across_tiles_enabled_flag` can be used to specify whether the loop filter across tile boundaries is enabled or disabled for all tile boundaries in the image.

[0086] Table 1 - HEVC Plot Syntax

[0087]

[0088] In one implementation scheme Figure 2A and Figure 2B The image shows two examples of tile partitioning. In the first example, as shown... Figure 2A As shown, (one or more) tile columns and (one or more) tile rows are evenly distributed across image 200 (in six grid areas). In the second example, as... Figure 2B As shown, the tile columns (one or more) and tile rows (one or more) are not evenly distributed on image 202 (in six grid areas), so it may be necessary to explicitly specify the tile column width and row height.

[0089] In some examples, HEVC specifies a special set of tiles called the Motion-Constrained Tile Set (MCTS) via a Supplemental Enhancement Information (SEI) message. The MCTS SEI message indicates that the inter-frame prediction process is constrained such that sample values ​​outside each identified tile set and / or sample values ​​at partial sample locations derived using one or more sample values ​​outside the identified tile set cannot be used for inter-frame prediction of any sample within the identified tile set [1]. In some cases, each MCTS can be extracted from the HEVC bitstream and decoded independently.

[0090] II.2 Filling for Motion Compensation Prediction

[0091] In some examples, existing video codecs are designed for traditional two-dimensional (2D) video captured on a plane. When motion compensation prediction uses any samples outside the boundaries of the reference image, repeated padding is performed by copying sample values ​​from the image boundaries.

[0092] In one example Figure 3 The repeating fill scheme 300 is shown. For example, block B0 is partially outside the reference image. Part P0 is filled with the top left sample of part P3. Part P1 is filled row by row with the top line of part P3. Part P2 is filled column by column with the left column of part P3.

[0093] In some examples, the 360-degree video includes video information over the entire sphere, thus the 360-degree video is inherently cyclic. When this cyclic property is taken into account, the reference image of the 360-degree video no longer has a “boundary” because the information contained within the “boundary” is entirely surrounded by the sphere. In some implementations, geometrical padding for the 360-degree video can be used (e.g., geometrical padding proposed in JVET-D0075[5]).

[0094] In one example Figure 4 A geometrical fill process 400 for a 360-degree video with an equal rectangular projection format (ERP) is illustrated. In this example, the geometrical fill processing of the ERP may include taking the fill along the corresponding arrow (e.g., arrow A') from the corresponding arrow, and so on, with letter notation indicating this correspondence. For example, at the left and right boundaries of the 360-degree video, samples at A, B, C, D, E, and F are filled with samples at A', B', C', D', E', and F'. At the top boundary, samples at G, H, I, and J are filled with samples at G', H', I', and J'. At the bottom boundary, samples at K, L, M, and N are filled with samples at K', L', M', and N'. Compared to the repetitive fill method currently used in HEVC, this geometrical fill can provide meaningful samples and improve the continuity of adjacent samples in blocks outside the boundaries of the ERP image.

[0095] III. Windows-related omnidirectional video processing

[0096] Omnidirectional Media Format (OMAF) is a system standard format developed by the Moving Picture Experts Group (MPEG). OMAF defines a media format that allows omnidirectional media to include 360-degree video, images, audio, and associated timed text. For example, several window-dependent omnidirectional video processing schemes are described in Appendix D of the OMAF specification [2].

[0097] In the example, a window-related scheme based on equal-resolution MCTS encodes the same omnidirectional video content into several HEVC bitstreams with different picture qualities and bit rates. Each MCTS is included in a region track, and an extractor track is also created. The OMAF player selects the quality received for each sub-picture track based on the viewing direction.

[0098] Figure 5 Example scheme 500 from clause D4.2 of OMAF[2] is shown. In this example, the OMAF player receives MCTS tracks 1, 2, 5, and 6 at a specific quality and regional tracks 3, 4, 7, and 8 at another quality. The extractor tracks are used to reconstruct a bitstream that can be decoded by a single HEVC decoder. The tiles of the reconstructed HEVC bitstream of MCTS with different qualities can be signaled by the HEVC tile syntax discussed herein.

[0099] In another example, a window-related video processing scheme based on MCTS is used to encode the same omnidirectional video source content (one or more) into several spatial resolutions. Based on the viewing direction, the extractor can select those high-resolution tiles that match the viewing direction and other low-resolution tiles. The bitstream parsed from the extractor track conforms to HEVC and can be decoded by a single HEVC decoder.

[0100] Figure 6 An example of a cubemap (CMP) partitioning scheme 600 from OMAF [2] Clause D.6.4 is shown. In this example, preprocessing and encoding are shown to achieve an effective CMP resolution of 6K using a window-dependent OMAF video profile based on HEVC. Content is encoded at two spatial resolutions with CMP face sizes of 1536×1536 and 768×768, respectively. In both bitstreams, a 6×4 tile grid is used, and MCTSs are decoded for each tile location. Each decoded MCTS sequence is stored as a region track. Extractor tracks are created for each different window-adaptive MCTS selection. This results in the creation of 24 extractor tracks. In each sample of the extractor tracks, an extractor is created for each MCTS to extract data from a region track containing one or more selected high-resolution or low-resolution MCTSs. Each extractor track uses the same 3×6 tile grid with a tile column width equal to 768, 768, and / or 384 luma samples (e.g., one or more luma pixels) and / or a constant tile row height of 768 luma samples. Each tile extracted from the low-resolution bitstream contains two slices. The bitstream parsed from the extractor track has a resolution of 1920×4608, conforming to, for example, HEVC level 5.1.

[0101] In some cases, the HEVC tile syntax discussed above (e.g., in Table 1) may not be used to represent the MCTS of the reconstructed bitstream(s) described above. Instead, slices can be used for each partition. (See reference) Figure 7 In one example, there are two extraction tracks: a left extraction track and a right extraction track. The left extraction track has six slice headers, denoted as slice headers 702, 704, 706, 708, 710, and 712. The right extraction track has twelve slice headers, denoted as slice headers 714, 716, 718, 720, 722, 724, 726, 728, 730, 732, 734, and 736. In this case, Figure 6 The division of the extractor orbitals (one or more) can end with 12 slice headers, such as... Figure 7 As shown.

[0102] Figure 8 An example of a preprocessing and encoding scheme 800 for achieving (e.g., HEVC-based) 6K effective ERP resolution is shown. OMAF Clause D6.3 provides a window-dependent scheme based on MCTS for achieving 6K effective ERP resolution. In one example, omnidirectional video at 6K resolution (6144x3072) is resampled to three spatial resolutions: 6K (6144x3072), 3K (3072x1536), and 1.5K (1536x768). The 6K and 3K sequences are cropped to 6144×2048 (as shown in grid 802) and 3072×1024 (as shown in grid 804), respectively, by excluding a 30-degree elevation range from the top and bottom. The cropped 6K and 3K input sequences are encoded using an 8×1 tile grid, with each tile being an MCTS.

[0103] As shown in grid 806, a top stripe and a bottom stripe of size 3072×256 corresponding to a 30-degree elevation angle range are extracted from the 3K input sequence. These top and bottom stripes are encoded as separate bitstreams with a 4×1 tile grid, with each stripe acting as a single MCTS. As shown in grid 808, a top stripe and a bottom stripe of size 1536×128 corresponding to a 30-degree elevation angle range are extracted from the 1.5K input sequence. Each stripe can be placed within an image of size 768×256, for example, by placing the left side of the stripe at the top of the image and the right side of the stripe at the bottom of the image.

[0104] In this example, each MCTS sequence from the pruned 6K and 3K bitstreams can be encapsulated as a separate track. Each bitstream containing the top or bottom stripe of the 3K or 1.5K input sequence can be encapsulated as a track (e.g., track 810).

[0105] Extractor tracks were prepared for each selection of four adjacent tiles from the cropped 6K bitstream, and for viewing directions above and below the mid-latitude line, respectively. This resulted in the creation of 16 extractor tracks. Each extractor track used the same arrangement, e.g., as... Figure 9A and Figure 9B As shown. For example, in Figure 9A In this context, grid 900 comprises segments created using traditional 2×2 tiles (uniform_spacing_flag = -1). Figure 9B In this context, grid 902 includes segmentation via flexible tiles (uniform_spacing_flag = -1). The image size of the bitstream parsed from the extractor track is 3840 × 2304, conforming to HEVC level 5.1. In some cases, the tile division of the extractor track may not need to be specified using the HEVC tile syntax discussed above (e.g., in Table 1).

[0106] IV. HEVC Plot

[0107] In HEVC, tiles are aligned with the boundaries of the Code Tree Unit (CTU). In some examples, the primary purpose of HEVC tiles is to divide an image into independent segments with minimal compression efficiency loss. In one implementation, HEVC tiles are used to divide an image for window-dependent omnidirectional video processing. The source video is then divided and encoded using one or more MCTSs, which can be decoded independently of adjacent tile sets. The extractor can select a subset of the tile set based on the window orientation and form an HEVC-compliant extractor track for consumption by the OMAF player.

[0108] For next-generation video compression standards (one or more), such as Universal Video Decoding (VVC), the size of the CTU may become larger due to increased image resolution. The granularity of tile segmentation may also become too large to align with frame packing boundaries. It will also be difficult to divide images into equally sized CTUs for load balancing. Furthermore, traditional tile structures may not handle the partitioning structures used for OMAF window-related processing, and the bit cost of segmentation using slicing is high.

[0109] In MPEG#123, flexible tile structures and syntax were proposed by JVET-K0155[3] and JVET-K0260[4]. JVET-K0155 proposed that a picture can be divided into CTUs of constant size as in traditional tiles, while the rightmost and bottommost CTUs in the tile boundaries can be different in size from the constant CTU size to achieve better load balancing and alignment with frame packing boundaries. The scattered CTUs in the right and bottom edges of each tile are encoded and decoded in the same way as in the picture boundaries.

[0110] JVET-K0260 proposes support for flexible tiles with rectangular shapes but different sizes. Each tile can be individually signaled by either copying the tile size from the previous tile size in decoding order or by using a tile width and a tile height codeword. Utilizing the proposed syntax, it is possible to support… Figure 6 and Figure 8 The partitioning structure is shown in the figure. However, compared to the HEVC tile syntax format, this syntax format may result in significant overhead costs for commonly used traditional tile structures.

[0111] Therefore, new or improved methods, schemes, and signal designs are needed to support (e.g., in video frames) flexible tiles.

[0112] V. Representative Process of Flexible Blocks

[0113] In this disclosure, we describe various embodiments, processes, methods, architectures, tables, and signal designs that support flexible grid regions or tiles, including, for example: 1) constraints on the geometry filling and loop filtering of flexible tiles; 2) signaling that distinguishes between conventional tiles and flexible tiles to reduce total signaling overhead; 3) flexible tile signaling design and scan conversion based on grid regions; and 4) initial quantization parameter (QP) signaling for tile-based video processing.

[0114] In various embodiments, the term "region" used in this invention may refer to a first set of grid regions, and the term "tile" used in this invention may refer to a second set of grid regions. In one example, an image or video frame may be divided into a first set of grid regions (e.g., regions), and each grid region of the first set of grid regions may be further divided into a second set of grid regions (e.g., tiles). In some cases, the terms "region," "grid region," and "tile" used in this disclosure may be interchangeable and may be represented as either a first set of grid regions or a second set of grid regions.

[0115] V.1 Fill and Loop Filter Constraints on Tile Boundaries

[0116] Traditional tile partitioning may not have an integer multiple of CTU at the right or bottom edge of the image, and flexible tiles may also not have an integer multiple of CTU at the right or bottom edge of the tile. Figure 9 illustrates this incompleteness in both traditional and flexible tile cases using the conventional method in one example. The incomplete CTUs along the right and bottom edges of each tile can be encoded and decoded in the same way as within the image boundaries.

[0117] The geometrical fill assumes that all information contained in the 360-degree video is wrapped around a sphere, and this cyclic property holds regardless of the projection format used to represent the 360-degree video on a 2D plane. Geometrical fill can be applied to the boundaries of 360-degree video images, but not to flexible tile boundaries, because the cyclic property depends on the partitioning structure. Based on the tile partitioning, the encoder can determine whether, for example, horizontal or vertical geometrical fill can be deployed for motion compensation prediction.

[0118] In some embodiments, a padding flag (e.g., notifying the receiver of the WTRU) can be signaled to indicate whether a padding operation can be performed on the edges of one or more tiles. If the padding_enabled_flag is set, repetitive padding or geometric padding can be performed on the edges of the one or more tiles. In some examples, for flexible tile syntax structures, each tile can be signaled individually. In some cases, the geometry_padding_indicator and repetitive_padding_indicator can be signaled for each tile.

[0119] In some embodiments, the `loop_filter_across_tiles_enabled_flag` is signaled in HEVC to indicate whether loop filter operations can be performed on tile boundaries in PPS. For example, if `loop_filter_across_tile_enabled_flag` is set, the `loop_filter_indicator` can be signaled to indicate which edge of the tile can be filtered.

[0120] In one example, Table 2 shows the syntax format for fill and loop filters for (one or more) tile or (one or more) grid regions.

[0121] Table 2 - Fill and Loop Filter Syntax

[0122]

[0123] In Table 2, a padding_enable_flag value of 1 indicates that padding operations can be used in the current block, while a padding_enable_flag value of 0 indicates that padding operations are not used in the current block.

[0124] In Table 2, `geometry_padding_indicator` is a bitmap that maps each tile edge to one bit. An example of this bit mapping could be that the most significant bit is a flag for the top edge, the second most significant bit is a flag for the right edge, and so on in clockwise order. When the bit value is 1, geometry padding can be applied to the corresponding tile edge; when the bit value is 0, no geometry padding is performed on the corresponding tile edge. If it does not exist, it can be inferred that the default value of `geometry_padding_indicator` is 0.

[0125] In Table 2, `repetitive_padding_indicator` is a bitmap that maps each tile edge to one bit. An example of this bit mapping could be that the most significant bit is a flag for the top edge, the second most significant bit is a flag for the right edge, and so on in clockwise order. When the bit value is 1, repetitive padding is applied to the corresponding tile edge; when the bit value is 0, no repetitive padding is performed on the corresponding tile edge. If it does not exist, it can be inferred that the default value of `repetitive_padding_indicator` is 0.

[0126] In Table 2, `loop_filter_indicator` is a bitmap that maps each tile edge to a one-bit value. When the bit value is 1, a loop filter operation can be performed on the corresponding tile edge; when the bit value is 0, a loop filter operation is not performed on the corresponding tile edge. If it does not exist, it can be inferred that the default value of `loop_filter_indicator` is 0.

[0127] In another embodiment, the enable flag `padding_on_tile_enabled_flag` can be populated with a signal at the PPS level. When `padding_on_tile_enabled_flag` equals 0, `padding_enabled_flag` at the tile level is inferred to be 0.

[0128] In another embodiment, the geometry filling can be deactivated when the size of the current tile edge is not the same as (e.g., different) the size of the corresponding reference boundary.

[0129] Figure 10This example demonstrates the use of flexible tiles in an ERP image. In this example, the ERP image 1000 can be divided into multiple tiles, each with a different size. Geometric fill can be enabled for specific tile edges based on the tile grid.

[0130] V.2 Signaling used to distinguish between traditional tile grids and flexible tile grids

[0131] Traditional tile partitioning restricts all tiles belonging to the same row to have the same row height and all tiles belonging to the same column to have the same column width. This restriction simplifies tile signaling and ensures that the tile set is rectangular in shape. Flexible tiles allow individual tiles to have different sizes and allow each tile's attributes to be signaled individually. This signaling supports various partitioning grids but can introduce significant bit overhead. A trade-off between the overhead bit cost and the flexibility of tile partitioning can be achieved by including indicators or flags to distinguish between traditional partitioning grids and flexible partitioning grids. The indicators or flags can indicate whether the entire image is partitioned into a regular M×N grid, where M and N are integers. The traditional HEVC tile syntax can be applied to regular M×N tile grids, while new flexible tile syntaxes, such as those discussed in JVET-K0260 or in this disclosure, can be applied to flexible tile grids.

[0132] In some examples, the indicators or flags discussed herein can be signaled at or within the sequence parameter set and / or picture parameter set.

[0133] V.3 Signaling for Grid Regions in Flexible Tiles

[0134] In some examples, tile column boundaries and similar tile row boundaries can span across the image. A prime example of the use of flexible tiles is window-dependent omnidirectional video processing methods, where multiple MCTS tracks from different image resolutions are merged into a single HEVC-compliant extractor track. The tile grid of this extractor track can originate from different image resolutions, and therefore the tile column and row boundaries can be discontinuous across the image, such as... Figure 6 and / or Figure 8 As shown.

[0135] Instead of individually signaling the size of each tile, a signaling scheme / design can be used or configured to signal each grid region, where a specific tile or region partitioning scheme is employed. In one example, different regions can have different grid partitions to achieve one or more flexible tiles. In this example, a corresponding region can have multiple tiles, and each tile can have the same or different sizes. In some examples, a first tile can have a different size compared to a second tile within the same grid region. In some cases, tiles in each row can share the same height, and tiles in each column can share the same width.

[0136] Table 3 shows an exemplary flexible block syntax (e.g., multi-level syntax) used in this exemplary signaling scheme / design.

[0137] Table 3 - Flexible Block Syntax

[0138]

[0139] The increment of 1 in `num_region_columns_minus1` specifies the number of columns that divide the image into regions. `num_region_columns_minus1` should be in the range of 0 to `PicWidthInCtbsY-1`, which includes 0 and `PicWidthInCtbsY-1`.

[0140] Increasing 1 to num_region_rows_minus1 specifies the number of rows that divide the image into regions. num_region_columns_minus1 should be in the range of 0 to PicHeightInCtbsY-1, which includes 0 and PicHeightInCtbsY-1.

[0141] The regions can be defined in a raster scan order from left to right and from top to bottom. The total number of regions, NumRegion, can be derived as follows:

[0142] NumRegions=(num_region_columns_minus1+1)*(num_region_rows_minus1+1)

[0143] A uniform_region_flag value of 1 indicates that region column boundaries and similar region row boundaries are evenly distributed across the image. A uniform_spacing_flag value of 0 indicates that region column boundaries and similar region row boundaries are not evenly distributed across the image, but are explicitly signaled using the syntax elements region_column_width_minus1 and region_row_height_minus1. When these are not present, the value of uniform_region_flag is inferred to be 1.

[0144] `region_size_unit_idc` specifies the unit size of the region in decode tree blocks. If `region_size_unit_idc` does not exist, its default value is assumed to be 0. The variable `RegionUnitInCtbsY` can be derived as follows:

[0145] RegionUnitInCtbsY=1< <region_unit_size_idc

[0146] The increment of region_column_width_minus1[i] by 1 specifies the width of the i-th region column in units of the decode tree block. When region_column_width_minus1 does not exist, it is inferred that the value of region_column_width_minus1 is equal to the image width PicWidthInCtbsY.

[0147] region_row_height_minus1[i] incremented by 1 specifies the height of the i-th region row in units of the decode tree block. When region_row_width_minus1 does not exist, it is inferred that the value of region_row_width_minus1 is equal to the image height PicHeightInCtbsY.

[0148] Figure 11A and Figure 11B The applications are shown respectively. Figure 6 and Figure 8 Two examples of region-based flexible tile signaling for the extractor track are shown.

[0149] Reference Figure 11A , Figure 6The extractor track was reconstructed into track 1100 from two images with different resolutions. Two regions were identified, with tiles evenly distributed within each region. The left region of track 1100 was divided into a 2×6 grid, and the right region of track 1100 was divided into a 1×12 grid.

[0150] refer to Figure 11B , Figure 8 The extractor track was reconstructed into tracks 1110 from four images of different resolutions, and four regions were identified, with tiles evenly distributed within each region. The first region was divided into a 4x1 grid, the second region into a 2x2 grid, the third region into a 4x1 grid, and the fourth region into a 1x2 grid.

[0151] In various embodiments, when processing video information (e.g., encoding or decoding video or images), the region partitioning and grouping mechanisms discussed herein may be employed. In one example, a WTRU (e.g., WTRU 102) may be configured to receive (or identify) a set of first parameters that define a plurality of first grid regions (e.g., tiles) comprising frames (e.g., video frames or image frames). For each first grid region, the WTRU may be configured to receive (or identify) a set of second parameters that define a plurality of second grid regions, and the plurality of second grid regions may divide the corresponding first grid regions. The WTRU may be configured to divide the frame into the plurality of first grid regions based on the set of first parameters, and to divide each first grid region into the plurality of second grid regions based on the corresponding set of second parameters.

[0152] In another example, the WTRU can be configured to receive (or identify) multiple sets of parameters or configurations for processing video information. For example, the WTRU can be configured to receive (or identify) a first set of parameters (which defines multiple first grid regions) and a second set of parameters (which defines multiple second grid regions). The WTRU can be configured to divide a frame into the multiple first grid regions based on the first set of parameters, and to group (or reconstruct) the multiple first grid regions into the multiple second grid regions based on the second set of parameters (one or more sets). In some cases, the first grid regions or the second grid regions can be tiles or slices, and can be used to construct or reconstruct frames (e.g., video frames or picture frames) or generate one or more bitstreams.

[0153] V.4 Decoding Tree Block (CTB) Raster and Flexible Tile Scan Conversion Process

[0154] In some embodiments, one or more of the following variables can be derived by invoking the decode tree block raster and flexible tile scan conversion process:

[0155] a) For ctbAddrRs in the range from 0 to PicSizeInCtbsY-1 (inclusive), the list CtbAddrRsToTs[ctbAddrRs] specifies the conversion from CTB address in the CTB raster scan of the image to CTB address in the tile scan.

[0156] b) For ctbAddrTs in the range from 0 to PicSizeInCtbsY-1 (inclusive), the list CtbAddrtStRs[ctbAddrTs] specifies the conversion from the CTB address in the tile scan of the image to the CTB address in the CTB raster scan.

[0157] c) For ctbAddrTs in the range from 0 to PicSizeInCtbsY-1 (inclusive), the list TileId[ctbAddrTs] specifies the conversion from CTB address to tile ID in tile scan;

[0158] d) For j in the range from 0 to num_tile_columns_minus1[i] (inclusive), the list ColumnWidthInLumaSamples[i][j] specifies the width of the j-th tile column of the i-th region in units of luminance samples; and / or

[0159] e) For j in the range from 0 to num_tile_rows_minus1[i] (inclusive), the list RowHeightInLumaSamples[i][j] specifies the height of the j-th tile row of the i-th region in units of luminance samples.

[0160] Figure 12A An example of a CTB raster scan of image frame 1200 is shown. Figure 12B An example of a CTB raster scan of a conventional tile in image frame 1210 is shown. Figure 12C An example of a region-based flexible tile scan is shown in picture frame 1220. HEVC[1] specifies the conversion from CTB addresses in a CTB raster scan of a picture to CTB addresses in a conventional tile scan. However, HEVC does not specify how to convert from CTB addresses in a CTB raster scan of a picture to CTB addresses in a region-based flexible tile scan.

[0161] In some embodiments, the conversion from CTB addresses in a CTB raster scan of an image to CTB addresses in a region-based flexible tile scan can be configured as follows:

[0162] 1) The variables CtbSizeY, PicWidthInCtbsY, and PicHeightInCtbsY are the same as those specified in HEVC[1]; and / or 2) For i in the range from 0 to num_region_columns_minus1 (inclusive), use a new list region_ColWidth[i] to specify the width of the i-th region column in CTB units, and the new list can be derived as follows:

[0163] In some embodiments, for a value j in the range from 0 to num_region_rows_minus1 (inclusive), a new list region_RowHeight[j] specifies the height of the j-th region row in CTB units, and this new list can be derived as follows:

[0164]

[0165] In some examples, the new variables RegionWidthInCtbsY and RegionHeightInCtbsY for the i-th region in the raster scan sequence can be derived as follows: RegionWidthInCtbsY[i] = regionColWidth[i%(num_region_columns_minus1+1)] RegionRowInCtbsY[i] = regionRowHeight[i / (num_region_row_minus1+1)] RegionSizeInCtbsY[i] = RegionWidthInCtbsY[i] * RegionRowInCtbsY[i]

[0166] In some embodiments, for i in the range from 0 to num_region_columns_minus1+1 (inclusive), a new list region_ColBd[i] specifies the position of the boundary of the i-th region column in units of the decode tree block. This new list can be derived as follows: for(regionColBd[0]=0;i=0;i<=num_region_columns_minus1;i++)

[0167] regionColBd[i+1]=regionColBd[i]+regionColWidth[i]

[0168] In some embodiments, for a j in the range from 0 to num_region_rows_minus1+1 (inclusive), a new list region_RowBd[j] specifies the position of the j-th region row boundary in units of decode tree blocks. This new list can be derived as follows: for(regionRowBd[0]=0;j=0;j<=num_region_rows_minus1;j++)

[0169] regionRowBd[j+1]=regionRowBd[j]+regionRowHeight[j]

[0170] In some embodiments, for j in the range from 0 to num_tile_columns_minus1[i] (inclusive), a new list colWidth[i][j] specifies the width of the j-th tile column of the i-th region in CTB units, and this new list can be derived as follows:

[0171]

[0172] In some embodiments, for a value j ranging from 0 to num_tile_rows_minus1 (inclusive), a new list rowHeight[i][j] specifies the height of the j-th tile row of the i-th region in CTB units. This new list can be derived as follows:

[0173]

[0174]

[0175] In some examples, the new variables ColumnWidthInLumaSamples[i][j] and RowHeightInLumaSamples[i][j] can be derived as follows:

[0176] ColumnWidthInLumaSamples[i][j]=colWidth[i][j]*CtbSizeY

[0177] RowHeightInLumaSamples[i][j]=rowHeight[i][j]*CtbSizeY

[0178] In some embodiments, for j ranging from 0 to num_tile_columns_minus1[i]+1 (inclusive), a new list colBd[i][j] specifies the position of the j-th tile column boundary of the i-th region in units of decoding tree blocks. This new list can be derived as follows:

[0179] colBd[i][0]=(i==0)? 0:colBd[i-1][0]+regionColBd[i-1]

[0180] colBd[i][0]=(colBd[i][0]==PicWidthInCtbsY)? 0:colBd[i][0]

[0181] for(j=0;j<=num_tile_columns_minus1[i];j++)

[0182] colBd[i][j+1]=colBd[i][j]+colWidth[i][j]

[0183] In some embodiments, for a value j ranging from 0 to num_tile_rows_minus1[i]+1 (inclusive), a new list rowBd[i][j] specifies the position of the j-th tile row boundary of the i-th region in units of decoding tree blocks. This new list can be derived as follows:

[0184] rowBd[i][0]=(i==0)? 0:rowBd[i-1][0]+regionRowBd[i-1]

[0185] rowBd[i][0]=(rowBd[i][0]==PicHeightInCtbsY)? 0:rowBd[i][0]

[0186] for(j=0;j<=num_tile_rows_minus1[i];j++)

[0187] rowBd[i][j+1]=rowBd[i][j]+rowHeight[i][j]

[0188] In some embodiments, for ctbAddrRs ranging from 0 to PicSizeInCtbsY-1 (inclusive), the list CtbAddrRsToTs[ctbAddrRs] specifies the conversion from CTB addresses in a CTB raster scan of an image to CTB addresses in a region-based tile scan, and this list can be derived as follows:

[0189]

[0190]

[0191] For ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 (inclusive), the list CtbAddrTsToRs[ctbAddrTs] specifies the conversion from CTB addresses in a region-based tile scan of an image to CTB addresses in a CTB raster scan. This list can be exported as follows:

[0192] for(ctbAddrRs=0;ctbAddrRs <PicSizeInCtbsY;ctbAddrRs++)

[0193] CtbAddrTsToRs[CtbAddrRsToTs[ctbAddrRs]]=ctbAddrRs2

[0194] For ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 (inclusive), the list TileId[ctbAddrTs] specifies the conversion from CTB addresses in a tile scan to tile indices or IDs, and this list can be exported as follows:

[0195]

[0196] In an alternative embodiment, the tile identifier (ID) for each region-based tile can be represented by a two-dimensional (2D) array. The first index can be a region index, and the second index can be a tile index within that region. Figure 13 This is an example of a tile ID representation in image frame 1300.

[0197] For ctbAddrTs in the range from 0 to PicSizeInCtbsY-1 (inclusive), the transformation from CTB addresses in a tile scan to 2D tile IDs (e.g., two new lists) TileId0[ctbAddrTs] and TileId1[ctbAddrTs] can be derived as follows:

[0198]

[0199] V.5 Initial Quantization Parameters for Patch Decoding

[0200] In some embodiments, HEVC can specify an initial quantization value for each slice. One or more initial quantization parameters (QP) can be used in the decoding block within the slice. The initial value of the luminance quantization parameter SliceQpY for a slice is derived as follows:

[0201] SliceQpY=26+init_qp_minus26+slice_qp_delta

[0202] In this context, init_qp_minus26 is signaled in PPS, and slice_qp_delta is signaled in the individual slice header.

[0203] The colorimetric parameters used for the slice and the decoding blocks within the slice are also signaled in the PPS and the slice header.

[0204] For omnidirectional video processing, a set of tiles can be mapped to windows or faces. Each window or face can be decoded to a different quality (e.g., resolution) to support window-dependent video processing. The quantization parameters of the tiles can be inferred from the slice header SliceQpY as specified in HEVC, or can be explicitly signaled as characteristics of the tiles.

[0205] In some examples, signaling of 360-degree video information [6] can be used. For example, in cases where a particular face is encoded at a higher or lower quality than another face, the QP for each face can be explicitly signaled. Decoder blocks belonging to the same face can share the same initial QP signal for that face.

[0206] In some embodiments, the QP can be signaled at the region and / or tile level so that all tiles belonging to the same region can share the same initial region QP. Alternatively, each tile may have its own initial QP value based on the initial region QP and individual tile QP offset values. Table 4 shows an exemplary signaling structure according to this embodiment.

[0207] Table 4 - QP Signaling for Region-Based Flexible Tiles

[0208]

[0209] The `region_qp_offset_enabled_flag` specifies whether different QPs should be used for different regions (one or more).

[0210] `region_QP_offset[i]` specifies the initial value of QP for tiles in the region, until modified by the value of `tile_QP_offset` in the decoding unit layer. The initial value of the QpY quantization parameter for the i-th region is `RegionQp`. Y [i] can be derived as follows:

[0211] RegionQp Y [i]=26+init_qp_minus26+region_qp_delta[i]

[0212] The tile_qp_offset_enabled_flag specifies whether different QPs are used for different tiles.

[0213] `tile_QP_offset[i][m][n]` specifies the initial value of the QP to be used for the decoded block in the tile at position [m][n] of the i-th region. If it does not exist, the value of `tile_qp_offset` can be inferred to be 0. The value of the quantization parameter `TileQpY[i][m][n]` can be derived as follows:

[0214] TileQp Y [i][m][n] = RegionQp Y [i]+tile_qp_delta[i][m][n]

[0215] The QP for each tile can be specified in tile index order. The tile index can be derived from the region index and the values ​​of the tile columns and rows as follows:

[0216]

[0217] In an alternative embodiment, tile QP offsets can be specified in a list, and each tile can derive its initial QP value by referring to a corresponding table index. Table 5 shows an exemplary list of QP offsets, while Table 6 shows an exemplary tile QP format.

[0218] Table 5 - QP Table

[0219]

[0220] Incrementing `tile_qp_offset_list_len_minus1` by 1 specifies the number of elements in the `tile_qp_offset_list` syntax. `tile_QP_offset_list` specifies a list of one or more QP offset values ​​used when deriving a tile QP from an initial QP.

[0221] Table 6 - Initial QP Signaling for Blocks

[0222]

[0223] tile_qp_offset_idx specifies the index in tile_qp_offset_list, which is used to determine TileQPOffset. Y The value of `tile_qp_offset_idx` should be in the range of 0 to `tile_qp_offset_list_len_minus1` when it exists, including both 0 and `tile_qp_offset_list_len_minus1`.

[0224] In some embodiments, the variable TileQpOffset of the i-th tile Y [i] and TileQp Y [i] can be derived as follows:

[0225] TileQpOffset Y [i]=tile_qp_offset_list[tile_qp_offset_idx]

[0226] TileQp Y [i]=26+init_qp_minus26+TileQpOffset Y [i]

[0227] Each of the following references is incorporated herein by reference: [1] JCTVC-R1013_v6, “Draft high efficiency video coding (HEVC) version 2”, June 2014; [2] ISO / IEC JTC1 / SC29 / WG11N17827, “WD2 of ISO / IEC 23090-2 OMAF 2”. nd[3] JVET-K0155, “AHG12: Flexible Tile Partitioning”, July 2018; [4] JVET-K0260, “Flexible Tile”, July 2018; [5] JVET-D0075, “AHG8: Geometry padding for 360 video coding”, October 2016; [6] PCT Patent Application Publication No. WO2018 / 045108; [7] U.S. Patent Application No. 62 / 775,130; and [8] U.S. Patent Application No. 62 / 781,749.

[0228] VII. Conclusion

[0229] Although the features and elements are described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in computer programs, software, or firmware embedded in a computer-readable medium and executed by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multipurpose discs (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU 102, UE, terminal, base station, RNC, or any host computer.

[0230] Furthermore, in the above embodiments, references to processing platforms, computing systems, controllers, and other devices including processors are mentioned. These devices may include at least one central processing unit (“CPU”) and memory. According to the practice of those skilled in the art of computer programming, references to symbolic descriptions of actions and operations or instructions can be executed by various CPUs and memories. These actions and operations or instructions may be referred to as “executed,” “computer-executed,” or “CPU-executed.”

[0231] Those skilled in the art will understand that the operations or instructions described by the actions and symbols include the CPU's manipulation of electrical signals. The electrical system represents the identification of data bits, causing electrical signals to be transformed or restored, and maintaining the storage location of data bits in the memory system, thereby reconfiguring or otherwise altering the CPU's operation and other signal processing. Maintaining the storage location of data bits involves having specific electrical, magnetic, optical, or organic properties corresponding to or representing the data bits. It should be understood that exemplary embodiments are not limited to the platforms or CPUs described above, and other platforms and CPUs may support the provided methods.

[0232] Data bits can also be stored on computer-readable media, including disks, optical disks, and any other large CPU-readable storage system, whether volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)). The computer-readable media can include cooperative or interconnected computer-readable media that reside exclusively on the processor system or are distributed among multiple interconnected processing systems, which may be local to the processing system or remote. It is understood that representative implementations are not limited to the memories described above, and other platforms and memories may support the described methods.

[0233] In the illustrated embodiments, any of the operations, processes, etc., described herein can be implemented as computer-readable instructions stored on a computer-readable medium. These computer-readable instructions can be executed by a processor of a mobile unit, network element, and / or any other computing device.

[0234] There is a distinction between hardware and software implementations in a system. The use of hardware or software is generally (but not always, as the choice between hardware and software can be critical in certain environments) a design choice that considers a trade-off between cost and efficiency. Various tools (e.g., hardware, software, and / or firmware) can influence the processes and / or systems and / or other technologies described herein, and the preferred tools can vary depending on the context of the deployed processes and / or systems and / or other technologies. For example, if the implementer determines that speed and accuracy are paramount, they may choose primarily hardware and / or firmware tools. If flexibility is paramount, they may choose primarily software implementation. Alternatively, the implementer may choose some combination of hardware, software, and / or firmware.

[0235] The foregoing detailed description has presented various implementations of the apparatus and / or process using block diagrams, flowcharts, and / or examples. Within the scope of one or more functions and / or operations contained in these block diagrams, flowcharts, and / or examples, those skilled in the art will understand that each function and / or operation within these block diagrams, flowcharts, or examples can be implemented individually and / or together in a wide range of hardware, software, or firmware, or substantially any combination thereof. Suitable processors include, for example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs); field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0236] While features and elements are provided above in specific combinations, it will be understood by those skilled in the art that each feature or element can be used alone or in any combination with other features and elements. This disclosure is not limited to the specific embodiments described herein, which are intended as examples of various aspects. Many modifications and variations can be made without departing from their essence and scope, as is known to those skilled in the art. Elements, actions, or instructions used in the description of this application should not be construed as critical or essential to the invention unless explicitly stated otherwise. In addition to the methods and apparatuses listed herein, those skilled in the art will recognize functionally equivalent methods and apparatuses within the scope of this disclosure based on the above description. These modifications and variations should also fall within the scope of the appended claims. This disclosure is defined solely by the appended claims, including their full equivalents. It should be understood that this disclosure is not limited to specific methods or systems.

[0237] It should also be understood that the terminology used herein is for describing particular implementations only and is not restrictive. The terms “station” and its abbreviation “STA”, “user equipment” and its abbreviation “UE” as used herein can refer to (i) a wireless transmitting and / or receiving unit (WTRU), as described below; (ii) an implementation of any number of WTRUs, as described below; (iii) a device with wireless and / or wired capabilities (e.g., wired), configured with some or all of the structure and functions of a WTRU (e.g., as described above); (iii) a device with wireless and / or wired capabilities, configured with fewer than all the structure and functions of a WTRU, as described below; and / or (iv) others. Details of example WTRUs that can represent (or be used interchangeably with) any UE or mobile device described herein have been referenced above. Figures 1A to 1D Provided.

[0238] In some representative embodiments, portions of the subject matter described herein may be implemented via application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, those skilled in the art will understand that some aspects of the embodiments disclosed herein, in whole or in part, may be equivalently implemented by integrated circuits as one or more computer programs running on one or more computers (e.g., one or more programs running on one or more computer systems), one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), firmware, or substantially any combination of these, and that designing circuitry and / or writing code for such software and / or firmware according to this disclosure is known to those skilled in the art. Furthermore, those skilled in the art will understand that the mechanisms of the subject matter described herein can be distributed as various forms of program products, and that exemplary embodiments of the subject matter described herein are applicable regardless of the specific type of signal-bearing medium used to actually perform that distribution. Examples of signal-bearing media include, but are not limited to, the following: recordable media, such as floppy disks, hard disks, CDs, DVDs, digital tapes, computer memory, etc., and transmission media, such as digital and / or analog communication media (e.g., optical fibers, waveguides, wired communication links, wireless communication links, etc.).

[0239] The topics described herein sometimes show different components that are contained in or connected to different other components. It is understood that the architectures depicted are merely examples, and many other architectures implementing the same functionality can be implemented in practice. Conceptually, any arrangement of components implementing the same functionality is effectively “associated” to thus implement the desired functionality. Therefore, any two components combined here to implement a particular function can be considered “associated” with each other to implement the desired functionality, regardless of the architecture or intermediate components. Similarly, any two associated components can also be considered “operationally connected” or “operationally coupled” to each other to implement the desired functionality, and any two components that can be associated in this way can also be considered “operationally coupled” to each other to implement the desired functionality. Specific examples of operationally coupled components include, but are not limited to, physically pairable and / or physically interactive components and / or wirelessly interactive components and / or logically interactive and / or logically interactive components.

[0240] Regarding the use of virtually any plural and / or singular terms herein, those skilled in the art can escape from plural to singular and / or from singular to plural as appropriate in context and / or application. For clarity, various singular / plural substitutions may be explicitly proposed herein.

[0241] Those skilled in the art will understand that the terminology used herein, and especially in the claims (e.g., the body of the claims), is generally “open-ended” (e.g., the term “comprising” should be understood as “including but not limited to,” the term “having” should be understood as “at least having,” the term “comprising” should be understood as “including but not limited to,” etc.). Those skilled in the art will also understand that if a claim describes a particular quantity, it will be explicitly stated in the claim, and without such a description, there is no such meaning. For example, the term “single” or similar language may be used to indicate only one item. To aid understanding, the following claims and / or the description herein may contain the use of the prepositional phrases “at least one” or “one or more” to introduce the claim description. However, the use of these phrases should not be construed as implying that a claim description introduced by the indefinite article “a” limits any particular claim containing such an introduced claim description to an embodiment containing only one such description, even if the same claim includes the prepositional phrases “one or more” or “at least one” and the indefinite article (e.g., “a”) (e.g., “a” should be understood as meaning “at least one” or “one or more”). The same applies to the use of definite articles used to introduce the claim description. Furthermore, even if a specific quantity described in the derived claim is explicitly stated, those skilled in the art will understand that such a description should be interpreted as indicating at least the quantity described (e.g., the simple description of "two descriptions" without any other modifiers indicates at least two descriptions, or two or more descriptions).

[0242] Furthermore, in these instances where the convention of "at least one of A, B, and C" is used, this convention is generally understood by those skilled in the art (e.g., "the system has at least one of A, B, and C" can include, but is not limited to, the system having only A, only B, only C, A and B, A and C, B and C, and / or A, B, and C, etc.). In these instances where the convention of "at least one of A, B, or C" is used, this convention is generally understood by those skilled in the art (e.g., "the system has at least one of A, B, or C" can include, but is not limited to, the system having only A, only B, only C, A and B, A and C, B and C, and / or A, B, and C, etc.). Those skilled in the art will also understand that any substantially separating word and / or phrase indicating two or more alternatives, whether in the specification, claims, or drawings, should be understood to include the possibility of including one of two items, either one or both items. For example, the phrase "A or B" is understood to include the possibility of "A" or "B" or "A" and "B". Furthermore, the term "any" as used herein, followed by a list of multiple items and / or various items, is intended to include "any," "any combination," "any number," and / or "any combination of multiple" items, alone or in combination with other items and / or other types of items. Furthermore, the terms "set" or "group" as used herein are intended to include any number of items, including zero. Furthermore, the term "quantity" as used herein is intended to include any quantity, including zero.

[0243] Furthermore, if the features or aspects of this disclosure are described in accordance with the Markush Group, those skilled in the art will understand that this disclosure is also described in accordance with any individual member or subgroup of members of the Markush Group.

[0244] Those skilled in the art will understand that, for any and all purposes, such as for providing a written description, all scopes disclosed herein also include any and all possible subscopes and combinations thereof. Any scope listed herein can be readily understood as sufficient to describe and implement the same scope divided into at least two, three, four, five, ten, etc., equal parts. As a non-limiting example, each scope described herein can be readily divided into a lower third, a middle third, and an upper third, etc. Those skilled in the art will also understand that all language such as “more than,” “at least,” “greater than,” “less than,” etc., includes the described numbers and scopes that can subsequently be divided into the aforementioned subscopes. Finally, those skilled in the art will understand that a scope includes each individual member. Thus, for example, a group and / or set with 1-3 cells refers to a group / set with 1, 2, or 3 cells. Similarly, a group / set with 1-5 cells refers to a group / set with 1, 2, 3, 4, or 5 cells, and so on.

[0245] Furthermore, the claims should not be construed as limiting to the provided order or elements unless the description has such an effect. Additionally, the use of the term "means for..." in any claim is intended to invoke 35 U.S.SC §112. The claim format of 6 or device + function, and any claim without the term "device for..." does not have this intention.

[0246] The software-associated processor can be used to implement radio frequency transceivers in a Wireless Transmit / Receive Unit (WTRU), User Equipment (UE), terminal, base station, Mobility Management Entity (MME), or Evolved Packet Core (EPC), or any host computer. The WTRU can incorporate hardware and / or software-implemented modules (including Software-Defined Radio (SDR)) and other components, such as cameras, video camera modules, video phones, walkie-talkies, vibration devices, speakers, microphones, television transceivers, hands-free headsets, and keyboards. Modules, FM radio units, Near Field Communication (NFC) modules, Liquid Crystal Display (LCD) units, Organic Light Emitting Diode (OLED) units, digital music players, media players, video game console modules, Internet browsers and / or any Wireless Local Area Network (WLAN) or Ultra Wideband (UWB) modules.

[0247] Although the invention has been described in relation to a communication system, it will be understood that the system can be implemented in software on a microprocessor / general-purpose computer (not shown). In some embodiments, the functionality of one or more of the various components can be implemented in software that controls the general-purpose computer.

[0248] Furthermore, although the invention has been shown and described with reference to specific embodiments, it is not intended to be limited to the details shown. Rather, various modifications to the details may be made within the scope of the claims and without departing from the invention.

[0249] Throughout this disclosure, those skilled in the art will understand that certain representative embodiments may be used in place of or in combination with other representative embodiments.

[0250] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in computer programs, software, or firmware embedded in a computer-readable medium and executed by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multipurpose discs (DVDs). The processor associated with the software can be used to implement radio frequency transceivers used in WTRUs, UEs, terminals, base stations, RNCs, or any host computer.

[0251] Furthermore, in the above embodiments, references to processing platforms, computing systems, controllers, and other devices including processors are mentioned. These devices may include at least one central processing unit (“CPU”) and memory. According to the practice of those skilled in the art of computer programming, references to symbolic descriptions of actions and operations or instructions can be executed by various CPUs and memories. These actions and operations or instructions may be referred to as “executed,” “computer-executed,” or “CPU-executed.”

[0252] Those skilled in the art will understand that the operations or instructions described by the actions and symbols include the CPU's manipulation of electrical signals. An electrical system representation can identify data bits, causing electrical signals to be transformed or restored, and maintaining the storage location of data bits in the memory system, thereby reconfiguring or otherwise altering the CPU's operation and other signal processing. Maintaining the storage location of data bits involves having specific electrical, magnetic, optical, or organic properties corresponding to or representing the data bits.

[0253] Data bits can also be stored on computer-readable media, including disks, optical disks, and any other large storage system that is volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) and readable by the CPU. Computer-readable media can include cooperative or interconnected computer-readable media that reside exclusively on the processor system or are distributed among multiple interconnected processing systems, which may be local to the processing system or remote. It is understood that representative implementations are not limited to the memories described above, and other platforms and memories may support the described methods.

[0254] As examples, suitable processors include general-purpose processors, special-purpose processors, traditional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC) and / or state machine.

[0255] Although the invention has been described with respect to a communication system, it is contemplated that the system can be implemented in software on a microprocessor / general-purpose computer (not shown). In some embodiments, the functionality of one or more of the various components can be implemented using software that controls the general-purpose computer.

[0256] Furthermore, although the invention has been described and illustrated herein with reference to specific embodiments, the invention is not intended to be limited to the details shown. Rather, various modifications to the details may be made within the scope of the claims and without departing from the invention.

Claims

1. A method for encoding video information, further comprising: Determine the first set of parameters, which defines multiple first grid regions within the frame; For each first grid region, a second set of parameters is determined, which defines a plurality of second grid regions, wherein the plurality of second grid regions divide the corresponding first grid regions; Based on the first set of parameters, the frame is divided into the plurality of first grid regions; Based on the corresponding second set of parameters, each first grid region is divided into the plurality of second grid regions; Based on the first and second grid regions of the frame, an image is encoded into a bitstream, wherein the image is encoded according to a first resolution level and a second resolution level, wherein the encoded image region of the image is mapped to a corresponding second grid region of the frame, and wherein the encoded image region of one of the images mapped to the first grid region of the frame is encoded according to the first resolution level, and the encoded image region of the other image mapped to the first grid region of the frame is encoded according to the second resolution level; as well as The first set of parameters and the corresponding second set of parameters for each of the first grid regions are encoded into the bitstream.

2. The method according to claim 1, wherein, The first set of parameters includes a flag that indicates that the frame will be divided according to a uniform grid.

3. The method according to claim 1, wherein, For the first grid region, the corresponding second set of parameters includes a flag indicating that the first grid region will be divided according to a uniform grid.

4. The method according to claim 1, wherein, The mapped image region that crosses the boundary of the first grid region is not a continuous region of the image being encoded.

5. The method according to claim 1, further comprising: Generate a fill flag, the fill flag indicating whether to perform a fill operation for each edge of the corresponding grid region of the plurality of first grid regions; as well as The padding flag is encoded into the bitstream.

6. A method for decoding video information, comprising: Decoding from the bitstream defines the first set of parameters for multiple first grid regions within the frame; For each first grid region, a second set of parameters is decoded from the bitstream, the second set of parameters defining a plurality of second grid regions, wherein the plurality of second grid regions divide the corresponding first grid regions; Based on the first set of parameters, the frame is divided into the plurality of first grid regions; Based on the corresponding second set of parameters, each first grid region is divided into the plurality of second grid regions; as well as Based on the first and second grid regions of the frame, an image is decoded from the bitstream, wherein the image is decoded according to a first resolution level and a second resolution level, wherein the decoded image region of the image is mapped to a corresponding second grid region of the frame, and wherein the decoded image region of one of the images mapped to the first grid region of the frame is decoded according to the first resolution level, and the decoded image region of the other image mapped to the first grid region of the frame is decoded according to the second resolution level.

7. The method according to claim 6, wherein, The first set of parameters includes a flag that indicates that the frame will be divided according to a uniform grid.

8. The method according to claim 6, wherein, For the first grid region, the corresponding second set of parameters includes a flag indicating that the first grid region will be divided according to a uniform grid.

9. The method according to claim 6, wherein, The first set of parameters and the second set of parameters are decoded from the bitstream in any of the following ways: sequence parameter set, image parameter set, or slice header.

10. The method of claim 6, further comprising: Decode the padding flag from the bitstream, the padding flag indicating whether to perform a padding operation for each edge of the corresponding grid region of the plurality of first grid regions; as well as Based on the fill flag, the edges of the grid region are filled.

11. An apparatus for video encoding, comprising: One or more processors are configured as follows: The first set of parameters defines the multiple first grid regions within the frame; For each first grid region, a second set of parameters is determined, which defines a plurality of second grid regions, wherein the plurality of second grid regions divide the corresponding first grid regions; Based on the first set of parameters, the frame is divided into the plurality of first grid regions; Based on the corresponding second set of parameters, each first grid region is divided into the plurality of second grid regions; as well as Based on the first and second grid regions of the frame, an image is encoded into a bitstream, wherein the image is encoded according to a first resolution level and a second resolution level, wherein the encoded image region of the image is mapped to a corresponding second grid region of the frame, and wherein the encoded image region of one of the images mapped to the first grid region of the frame is encoded according to the first resolution level, and the encoded image region of the other image mapped to the first grid region of the frame is encoded according to the second resolution level; as well as The first set of parameters and the corresponding second set of parameters for each of the first grid regions are encoded into the bitstream.

12. The apparatus according to claim 11, wherein, The first set of parameters includes a flag that indicates that the frame will be divided according to a uniform grid.

13. The apparatus according to claim 11, wherein, For the first grid region, the corresponding second set of parameters includes a flag indicating that the first grid region will be divided according to a uniform grid.

14. The apparatus according to claim 11, wherein, The mapped image region that crosses the boundary of the first grid region is not a continuous region of the image being encoded.

15. The apparatus of claim 11, further comprising: Generate a fill flag, the fill flag indicating whether to perform a fill operation on each edge of the corresponding grid region of the plurality of first grid regions; as well as The padding flag is encoded into the bitstream.

16. An apparatus for video decoding, comprising: One or more processors are configured as follows: Decoding from the bitstream defines the first set of parameters for multiple first grid regions within the frame; For each first grid region, a second set of parameters is decoded from the bitstream, the second set of parameters defining a plurality of second grid regions, wherein the plurality of second grid regions divide the corresponding first grid regions; Based on the first set of parameters, the frame is divided into the plurality of first grid regions; Based on the corresponding second set of parameters, each first grid region is divided into the plurality of second grid regions; as well as Based on the first and second grid regions of the frame, an image is decoded from the bitstream, wherein the image is decoded according to a first resolution level and a second resolution level, wherein the decoded image region of the image is mapped to a corresponding second grid region of the frame, and wherein the decoded image region of one of the images mapped to the first grid region of the frame is decoded according to the first resolution level, and the decoded image region of the other image mapped to the first grid region of the frame is decoded according to the second resolution level.

17. The apparatus according to claim 16, wherein, The first set of parameters includes a flag that indicates that the frame will be divided according to a uniform grid.

18. The apparatus according to claim 16, wherein, For the first grid region, the corresponding second set of parameters includes a flag indicating that the first grid region will be divided according to a uniform grid.

19. The apparatus of claim 16, wherein the first set of parameters and the second set of parameters are decoded from the bitstream in any of the following: a sequence parameter set, an image parameter set, or a slice header.

20. The apparatus of claim 16, wherein the one or more processors are further configured to: Decode the padding flag from the bitstream, the padding flag indicating whether to perform a padding operation for each edge of the corresponding grid region of the plurality of first grid regions; and Based on the fill flag, the edges of the grid region are filled.

Citation Information

Patent Citations

  • Method and system for signaling of 360-degree video information

    WO2018045108A1

  • Video encoding and decoding method and device

    CN108028933A

  • Sub-Pictures for Pixel Rate Balancing on Multi-Core Platforms

    US20130202051A1

  • Tile grouping in HEVC and l-HEVC file formats

    US20170289556A1

  • System and method for improving efficiency in encoding / decoding a curved view video

    WO2018035721A1