Methods and apparatus for point cloud compression bitstream format

The method addresses the challenge of efficiently representing and transmitting large-scale 3D point clouds by employing a video-based point cloud compression bitstream format based on ISOBMFF, ensuring effective decoding and reconstruction for enhanced 3D rendering applications.

JP7691541B2Active Publication Date: 2025-06-11INTERDIGITAL VC HOLDINGS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024045278
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-18
Filing Date
2024-03-21
Publication Date
2025-06-11
Estimated Expiration
2039-09-11

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently representing, compressing, and transmitting large-scale 3D point clouds over communication networks, particularly in ensuring effective decoding and reconstruction of 3D spaces.

Method used

A method and apparatus for a point cloud compression bitstream format, utilizing a video-based point cloud compression (V-PCC) bitstream structure based on the ISOBMFF standard, which enables flexible storage and extraction of point cloud components, and supports efficient decoding and reconstruction.

Benefits of technology

The proposed solution enables efficient compression, transmission, and decoding of 3D point clouds, improving storage and rendering capabilities in applications like telepresence, virtual reality, and dynamic 3D maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691541000025
    Figure 0007691541000025
  • Figure 0007691541000026
    Figure 0007691541000026
  • Figure 0007691541000027
    Figure 0007691541000027
Patent Text Reader

Abstract

To provide methods, apparatus, systems, architectures and interfaces for encoding and / or decoding point cloud bitstreams.SOLUTION: Included in methods, apparatuses, systems, architectures and interfaces for encoding and / or decoding point cloud bitstreams including a coded point cloud sequence is an apparatus that may include a processor and memory. A method may include: mapping components of a point cloud bitstream into tracks; generating information for identifying any of geometry streams and texture streams according to the mapping of the components; generating information associated with layers corresponding to respective geometry component streams; and generating information indicating operation points associated with the point cloud bitstream.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following generally relates to communication networks, wireless and / or wired. For example, one or more embodiments disclosed herein relate to methods and apparatuses for decoding information associated with a three-dimensional (3D) point cloud that can be transmitted and / or received using a wireless communication network and / or a wired communication network.

Background Art

[0002] A 3D point cloud can provide a representation of a physical space, a virtual space, and / or an immersive media. For example, a point cloud can be a set of points representing a 3D space using coordinates indicating the position of each point along one or more attributes such as color, transparency, acquisition time, laser reflectivity, or material properties, among others, associated with one or more of the points. Point clouds can be captured in several ways. Point clouds can be captured using a plurality of cameras and / or any of depth sensors such as, for example, a light detection and ranging (LiDAR) laser scanner. To represent a 3D space, the number of points used to reconstruct objects and scenes (e.g., realistically) using a point cloud can be in the millions or billions of dimensions and may be further increased. Such a large number of points in a point cloud may require efficient representation and compression for storing and transmitting point cloud data and may be applied beforehand, for example, when capturing and rendering 3D points used in areas such as telepresence, virtual reality, and large-scale dynamic 3D maps.

Summary of the Invention

[0003] A more detailed understanding can be obtained from the following detailed description together with the accompanying drawings given by way of example. In and of itself, the drawings and the detailed description are not considered limiting, and other equally effective embodiments are possible. Further, the same reference numerals in the drawings indicate the same elements.

Advantages of the Invention

[0004] Provided are a method and an apparatus for a point cloud compression bitstream format.

Brief Description of the Drawings

[0005]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Best Mode for Carrying Out the Invention

[0006] Exemplary Networks and Devices FIG. 1A is a diagram illustrating an exemplary communication system 100 that can implement one or more of the disclosed embodiments. The communication system 100 may be a multi-connectivity system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word discrete Fourier transform spread OFDM (ZT UW DTS-S-OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multicarrier (FBMC).

[0007] As shown in FIG. 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it should be recognized that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, any of them may be referred to as a "station" and / or "STA", and the WTRUs 102a, 102b, 102c, 102d may be configured to transmit and / or receive wireless signals and may be a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (e.g., remote surgery), an industrial device and application (e.g., a robot and / or other wireless device operating in an industrial and / or automated processing chain situation), a home appliance device, and a device operating on a commercial and / or industrial wireless network, among others. Any of the WTRUs 102a, 102b, 102c, 102d may also be interchangeably referred to as a UE.

[0008] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as CN 106, Internet 110, and / or other network 112. By way of example, base stations 114a, 114b may be a base transceiver station (BTS), NodeB, eNodeB, home NodeB, home eNodeB, gNB, NR NodeB, site controller, access point (AP), and wireless router, among others. Although base stations 114a, 114b are each represented as a single element, it will be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0009] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements such as a base station controller (BSC), a radio network controller (RNC), relay nodes (not shown). The base station 114a and / or the base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be in the licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. The cell may provide coverage for wireless services in a particular geographic area that may be relatively fixed or may change over time. The cell may further be divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one for each sector of the cell. In an embodiment, the base station 114a may utilize multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0010] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0011] More specifically, as described above, the communication system 100 may be a multi-connection system and may employ one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a within RAN104 / 113 and the WTRUs 102a, 102b, 102c may establish the air interface 116 using wideband CDMA (WCDMA) and may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed Uplink (UL) Packet Access (HSUPA).

[0012] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may establish the air interface 116 using Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) and may implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA).

[0013] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may establish the air interface 116 using NR and may implement radio technologies such as NR radio access.

[0014] In an embodiment, the base station 114a, and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a, and the WTRUs 102a, 102b, 102c may implement both LTE radio access and NR radio access using, for example, the dual connectivity (DC) principle. Accordingly, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by transmissions to / from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).

[0015] In other embodiments, the base station 114a, and the WTRUs 102a, 102b, 102c may implement wireless technologies such as IEEE802.11 (i.e., Wireless Fidelity (WiFi)), IEEE802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM) for mobile communications, High-Speed Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0016] The base station 114b in FIG. 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a localized area such as an office, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by a drone), and a roadway. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in FIG. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.

[0017] RAN 104 / 113 may communicate with CN 106 / 115, and CN 106 / 115 may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, 102d. The data may have various Quality of Service (QoS) requirements such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid originating calls, Internet connectivity, video distribution, etc., and / or may perform high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 may communicate directly or indirectly with other RANs that utilize the same RAT or a different RAT as RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113 which may be utilizing New Radio (NR) radio technology, CN 106 / 115 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0018] CN106 / 115 may also serve as a gateway for WTRU102a, 102b, 102c, 102d to access the PSTN108, the Internet 110, and / or other networks 112. The PSTN108 may include a circuit-switched telephone network that provides basic telephone service (POTS). The Internet 110 may include a worldwide system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) within the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs that may utilize the same RAT or a different RAT as the RAN104 / 113.

[0019] Some or all of the WTRU102a, 102b, 102c, 102d within the communication system 100 may include a multimode function (e.g., the WTRU102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU102c shown in Figure 1A may be configured to communicate with a base station 114a that may employ a cellular-based wireless technology and with a base station 114b that may utilize IEEE802 wireless technology.

[0020] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU 102 may include any sub-combination of the above elements while maintaining consistency with the embodiments.

[0021] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors cooperating with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, among others. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, and the transceiver 120 may be coupled to the transmit / receive element 122. Although Figure 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0022] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 may be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0023] In FIG. 1B, the transmit / receive element 122 is shown as a single element, but the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0024] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have a multimode function. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.

[0025] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may obtain information from and may store data in any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 may include a random access memory (RAM), a read only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include, for example, a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card. In other embodiments, the processor 118 may access information from and may store data in a memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).

[0026] The processor 118 may receive power from a power source 134 and may be configured to distribute power to and / or control power for other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cells (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell.

[0027] Processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of WTRU 102. In addition to, or instead of, information from the GPS chipset 136, WTRU 102 may receive location information on air interface 116 from a base station (e.g., base stations 114a, 114b), and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that WTRU 102 may obtain location information using any suitable positioning method while maintaining consistency with the embodiments.

[0028] Processor 118 may also be coupled to other peripheral devices 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripheral devices 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or video), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a Frequency Modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, and an activity tracker, among others. Peripheral devices 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0029] WTRU102 may include a full-duplex radio in which some or all of the transmission and reception of signals associated with a particular subframe for both the uplink (e.g., for transmission) and the downlink (e.g., for reception) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit 139 to reduce and / or substantially eliminate self-interference, either via hardware (e.g., a choke) or via signal processing through a processor (e.g., a separate processor (not shown) or processor 118). In an embodiment, WTRU102 may include a half-duplex radio for some or all of the transmission and reception of signals associated with a particular subframe for either the uplink (e.g., for transmission) or the downlink (e.g., for reception).

[0030] FIG. 1C is a system diagram illustrating RAN104 and CN106 in accordance with an embodiment. As described above, RAN104 may employ E-UTRA radio technology to communicate with WTRU102a, 102b, 102c through air interface 116. RAN104 may also communicate with CN106.

[0031] RAN104 may include eNodeBs 160a, 160b, 160c, although it will be understood that RAN104 may include any number of eNodeBs while maintaining consistency with the embodiment. Each of eNodeBs 160a, 160b, 160c may include one or more transceivers for communicating with WTRU102a, 102b, 102c over air interface 116. In one embodiment, eNodeBs 160a, 160b, 160c may implement MIMO technology. Thus, eNodeB160a, for example, may transmit a wireless signal to and / or receive a wireless signal from WTRU102a using multiple antennas.

[0032] Each of eNodeBs 160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, and user scheduling in the UL and / or DL. As shown in Figure 1C, eNodeBs 160a, 160b, and 160c may communicate with each other over the X2 interface.

[0033] CN106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the above elements is depicted as part of CN106, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.

[0034] MME 162 may be connected to each of eNodeBs 160a, 160b, and 160c within RAN 104 via the S1 interface and may act as a control node. For example, MME 162 may be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, and selecting a specific serving gateway during the initial attach of WTRUs 102a, 102b, and 102c. MME 162 may provide control plane functions for exchanges between RAN 104 and other RANs (not shown) that utilize other radio technologies such as GSM and / or WCDMA.

[0035] SGW164 may be connected to each of eNodeBs 160a, 160b, and 160c within RAN104 via the S1 interface. SGW164 may generally route and transfer user data packets to / from WTRUs 102a, 102b, and 102c. SGW164 may perform other functions such as anchoring the user plane during eNodeB handover, triggering paging when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.

[0036] SGW164 may be connected to PGW166, and PGW166 may provide access to a packet switched network, such as the Internet 110, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and IP-enabled devices.

[0037] CN106 may facilitate communication with other networks. For example, CN106 may provide access to a circuit switched network, such as PSTN 108, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and conventional fixed line communication devices. For example, CN106 may include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN106 and PSTN 108. Additionally, CN106 may provide access to other network 112 to WTRUs 102a, 102b, and 102c, and other network 112 may include other wired and / or wireless networks owned and / or operated by other service providers.

[0038] In FIGS. 1A through 1D, the WTRU is described as a wireless terminal, but in some representative embodiments, it is contemplated that such a terminal may (e.g., temporarily or permanently) use a wired communication interface to a communication network.

[0039] In some representative embodiments, the other network 112 may be a WLAN.

[0040] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have an access or interface to a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic to an STA originating from outside the BSS may arrive and be delivered to the STA through the AP. Traffic transmitted from an STA to a destination outside the BSS may be transmitted to the AP for delivery to each destination. Traffic between STAs within the BSS may be transmitted through the AP; for example, the source STA may transmit the traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered peer-to-peer traffic and / or may be referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted (e.g., directly) between the source STA and the destination STA using direct link setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) may communicate directly with each other. Communication in IBSS mode may sometimes be referred to herein as "ad hoc" mode communication.

[0041] When using the operation of 802.11ac infrastructure mode or the operation of a similar mode, the AP may transmit beacons on a fixed channel such as the primary channel. The primary channel may have a fixed width (e.g., 20 megahertz bandwidth), or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In one representative embodiment, for example, in an 802.11 system, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) may be implemented. In the case of CSMA / CA, STAs including the AP (e.g., any STA) may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. Within a given BSS, at any given time, one STA (e.g., only one station) may transmit.

[0042] A high throughput (HT) STA may use a 40 megahertz width channel for communication, for example, by combining the primary 20 megahertz channel with adjacent or non - adjacent 20 megahertz channels to form a 40 megahertz width channel.

[0043] Very High Throughput (VHT) STAs can support channels with widths of 20 megahertz, 40 megahertz, 80 megahertz, and / or 160 megahertz. 40 megahertz and / or 80 megahertz channels may be formed by combining consecutive 20 megahertz channels. A 160 megahertz channel may be formed by combining eight consecutive 20 megahertz channels, or may be formed by combining two non - consecutive 80 megahertz channels, which may be referred to as an 80 + 80 configuration. In the case of the 80 + 80 configuration, after channel encoding, the data may be passed through a segment parser that can split the data into two streams. For each stream separately, an Inverse Fast Fourier Transform (IFFT) process and time - domain processing may be performed. The streams may be mapped onto two 80 megahertz channels and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80 + 80 configuration may be reversed, and the combined data may be transmitted to the Media Access Control (MAC).

[0044] Operation in the sub - 1 - GHz mode is supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 - MHz, 10 - MHz, and 20 - MHz bandwidths in the TV white - space (TVWS) spectrum, and 802.11ah supports 1 - MHz, 2 - MHz, 4 - MHz, 8 - MHz, and 16 - MHz bandwidths using the non - TVWS spectrum. According to an exemplary embodiment, 802.11ah may support meter - type control / machine - type communication, such as MTC devices in a macro - coverage area. The MTC devices may have limited functionality including certain functions, for example, support for a certain bandwidth and / or limited bandwidth support (e.g., only their support). The MTC devices may include a battery having a battery life above a threshold (e.g., to maintain a very long battery life).

[0045] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that may be designated as primary channels. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs within a BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the minimum bandwidth operating mode among all STAs operating within the BSS. In the example of 802.11ah, for an STA (e.g., an MTC type device) that supports the 1 megahertz mode (e.g., supports only that), the primary channel may be 1 megahertz wide even if the AP and other STAs within the BSS support 2 megahertz, 4 megahertz, 8 megahertz, 16 megahertz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. For example, if the primary channel is busy because an STA (supporting only the 1 megahertz operating mode) is transmitting to the AP, the entire available frequency band may be considered busy even if most of the frequency band remains idle and available.

[0046] In the United States, the available frequency band that may be used by 802.11ah is from 902 megahertz to 928 megahertz. In Korea, the available frequency band is from 917.5 megahertz to 923.5 megahertz. In Japan, the available frequency band is from 916.5 megahertz to 927.5 megahertz. The total bandwidth available for 802.11ah is from 6 megahertz to 26 megahertz, depending on national regulations.

[0047] Figure 1D is a system diagram illustrating RAN 113 and CN 115 according to an embodiment. As described above, RAN 113 may communicate with WTRUs 102a, 102b, 102c over air interface 116 using NR radio technology. RAN 113 may also communicate with CN 115.

[0048] RAN 113 may include gNBs 180a, 180b, 180c, although it will be understood that RAN 113 may include any number of gNBs while maintaining consistency with the embodiment. Each of gNBs 180a, 180b, 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, 102c over air interface 116. In one embodiment, gNBs 180a, 180b, 180c may implement MIMO technology. For example, WTRUs 102a, 108b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, 180c. Thus, gNB 180a, for example, may transmit and / or receive radio signals from WTRU 102a using multiple antennas. In an embodiment, gNBs 180a, 180b, 180c may implement carrier aggregation technology. For example, gNB 180a may transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c may implement multi-site coordinated (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).

[0049] WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using transmissions associated with scalable numerology. For example, the OFDM symbol interval and / or the OFDM sub-carrier interval may vary for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using sub-frames or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or lasting for various lengths of absolute time).

[0050] gNBs 180a, 180b, and 180c may be configured to communicate with WTRUs 102a, 102b, and 102c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 102a, 102b, and 102c may communicate with gNBs 180a, 180b, and 180c without accessing other RANs (such as eNodeBs 160a, 160b, and 160c for example). In a stand-alone configuration, WTRUs 102a, 102b, and 102c may utilize one or more of gNBs 180a, 180b, and 180c as a mobility anchor point. In a stand-alone configuration, WTRUs 102a, 102b, and 102c may communicate with gNBs 180a, 180b, and 180c using signals within an unlicensed band. In a non-stand-alone configuration, WTRUs 102a, 102b, and 102c may communicate with and / or connect to gNBs 180a, 180b, and 180c while also communicating with / connecting to another RAN such as eNodeBs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c may implement the DC principle to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNodeBs 160a, 160b, and 160c. In a non-stand-alone configuration, eNodeBs 160a, 160b, and 160c may serve as a mobility anchor for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c may provide additional coverage and / or throughput to serve WTRUs 102a, 102b, and 102c.

[0051] Each of gNBs 180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, and routing of control plane information to access and mobility management functions (AMFs) 182a, 182b. As shown in FIG. 1D, gNBs 180a, 180b, and 180c may communicate with each other over the Xn interface.

[0052] CN 115 shown in FIG. 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and perhaps data networks (DNs) 185a, 185b. Although each of the above elements is depicted as part of CN 115, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.

[0053] AMF182a and 182b may be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N2 interface and may serve as control nodes. For example, AMF182a and 182b may authenticate users of WTRU102a, 102b, and 102c, support network slicing (e.g., handling different PDU sessions with different requirements), select specific SMF183a and 183b, manage the registration area, terminate NAS signaling, and perform mobility management. Network slicing may be used by AMF182a and 182b to customize the CN support for WTRU102a, 102b, and 102c based on the type of services utilized by WTRU102a, 102b, and 102c. For example, different network slices may be established for different use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on high-speed large-capacity mobile broadband (eMBB) access, and / or services for machine-type communication (MTC) access. AMF162 may provide control plane functions for exchanges between RAN113 and other RANs (not shown) that utilize other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.

[0054] SMF183a and 183b may be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b may also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b may select and control UPF184a and 184b and configure the routing of traffic through UPF184a and 184b. SMF183a and 183b may perform other functions such as managing and allocating UE IP addresses, managing PDU sessions, enforcing policies and controlling QoS, and providing downlink data notifications. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0055] UPF184a and 184b may be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N3 interface, and they can provide access to a packet-switched network such as the Internet 110 to WTRU102a, 102b, and 102c to facilitate communication between WTRU102a, 102b, and 102c and IP-corresponding devices. UPF184a and 184b may perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homing PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0056] CN115 can facilitate communication with other networks. For example, CN115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN115 and the PSTN108. Additionally, CN115 may provide access to other networks 112 to the WTRU102a, 102b, 102c, and the other networks 112 may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRU102a, 102b, 102c may be connected to the local data network (DN) 185a, 185b through the UPF184a, 184b via an N3 interface to the UPF184a, 184b and an N6 interface between the UPF184a, 184b and the DN185a, 185b.

[0057] With reference to FIGS. 1A through 1D, and the corresponding descriptions of FIGS. 1A through 1D, one or more of the functions described herein with respect to one or more of the WTRU102a through d, base stations 114a through b, eNodeB160a through c, MME162, SGW164, PGW166, gNB180a through c, AMF182a through b, UPF184a through b, SMF183a through b, DN185a through b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation device may be used to test other devices and / or to simulate network and / or WTRU functionality.

[0058] An emulation device may be designed to perform one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices may perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices may perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device may be directly coupled to another device for testing purposes and / or may perform tests using over-the-air wireless communication.

[0059] One or more emulation devices may perform one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be utilized in a test scenario in a test laboratory and / or in a non-deployed (e.g., test) wired and / or wireless communication network to perform tests on one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via an RF circuit (which may include, for example, one or more antennas) may be used by an emulation device to transmit and / or receive data.

[0060] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, video gaming devices, video game consoles, and cellular, satellite, or other wireless telephones. Many digital video devices implement video compression techniques such as those described in standards defined by Moving Picture Experts Group (MPEG)-2, MPEG-4, etc., of MPEG, International Telecommunications Union (ITU)-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), and extended versions of such standards, to more efficiently transmit and receive digital video information including information associated with three-dimensional (3D) point clouds.

[0061] FIG. 2 is a block diagram showing an exemplary video decoding and decoding system 10 that can implement and / or implement one or more embodiments. The system 10 may include a source device 12 that transmits encoded video information to a destination device 14 via a communication channel 16.

[0062] The source device 12 and the destination device 14 may be any of a wide range of devices. In some embodiments, the source device 12 and the destination device 14 may include a wireless handset or, in that case, any wireless device such as a wireless transmit and / or receive unit (WTRU) that can communicate video information through a communication channel 16 that includes a wireless link. However, explicitly, implicitly, and / or inherently, the methods, apparatuses, and systems described, disclosed, or otherwise provided herein (collectively "provided") are not necessarily limited to wireless applications or settings. For example, those techniques may be applied to over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet video transmissions, encoded digital video encoded on a storage medium, or other scenarios. Accordingly, the communication channel 16 may include any combination of wireless or wired media suitable for the transmission of encoded video data and / or may be either of them.

[0063] The source device 12 may include a video encoder unit 18, a transmit and / or receive (Tx / Rx) unit 20, and a Tx / Rx element 22. As shown, the source device may optionally include a video source 24. The destination device 14 may include a Tx / RX element 26, a Tx / Rx unit 28, and a video decoder unit 30. As shown, the destination device 14 may optionally include a display device 32. Each of the Tx / Rx units 20 and 28 may be a transmitter, a receiver, or a combination of a transmitter and a receiver (e.g., a transceiver or a transmitter-receiver) or may include them. Each of the Tx / Rx elements 22 and 26 may be, for example, an antenna. In accordance with the present disclosure, the video encoder unit 18 of the source device 12 and / or the video decoder unit 30 of the destination device may be configured and / or adapted to apply the coding techniques provided herein (collectively "adapted").

[0064] The source device 12 and the destination device 14 may include other elements / components or arrangements. For example, the source device 12 may be adapted to receive video data from an external video source. Also, the destination device 14 may include and / or interface with an external display device (not shown) instead of including and / or using a (e.g., integrated) display device 32. In some embodiments, the data stream generated by the video encoder unit 18 may be transmitted to other devices without the need to modulate the data on a carrier signal, such as direct digital transfer, and the other devices may or may not modulate the data for transmission.

[0065] The system 10 shown in FIG. 2 is only one example. The techniques provided herein may be performed by any digital video encoding and / or digital video decoding device. The techniques provided herein are generally performed by separate video encoding and / or video decoding devices, but the techniques may also typically be performed by a combined video encoder / decoder, referred to as a "CODEC". Moreover, the techniques provided herein may also be performed by a video pre-processor or the like. The source device 12 and the destination device 14 are only examples of such coding devices that generate (and receive and generate) video information encoded for transmission from the source device 12 to the destination device 14. In some embodiments, the devices 12 and 14 may operate in a substantially symmetric manner such that each of the devices 12 and 14 includes both a video encoding component and / or components (collectively "components") and a video decoding component and / or components. Thus, the system 10 can support either unidirectional video or bidirectional video between the devices 12 and 14 for any of, for example, video streaming, video playback, video broadcast, video telephony, and video conferencing. In some embodiments, the source device 12 may be, for example, a video streaming server adapted to generate (and / or receive and generate) encoded video information for one or more destination devices, and the destination devices may communicate with the source device 12 through a wired communication system and / or a wireless communication system.

[0066] The external video source and / or video source 24 may be, and / or may include, a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed from a video content provider. Alternatively, the external video source and / or video source 24 may generate data based on computer graphics as a combination of source video, or live video, archived video, and computer-generated video. In some embodiments, when the video source 24 is a video camera, the source device 12 and the destination device 14 may be, or may embody, a camera phone or a video phone. However, as mentioned above, the techniques provided herein may be applicable generally to video coding and may be applied to wireless and / or wired applications. In either case, the captured video, previously captured video, computer-generated video, video feed, or other type of video data (collectively "uncoded video") may be encoded by the video encoder unit 18 to form encoded video information.

[0067] The Tx / Rx unit 20 may modulate the encoded video information, e.g., according to a communication standard, to form one or more modulated signals that carry the encoded video information. The Tx / Rx unit 20 may also pass the modulated signal to its transmitter for transmission. The transmitter may transmit the modulated signal to the destination device 14 via the Tx / Rx element 22.

[0068] At the destination device 14, the Tx / Rx unit 28 may receive the modulated signal via the Tx / Rx element 26 through the channel 16. The Tx / Rx unit 28 may demodulate the modulated signal to obtain the encoded video information. The Tx / RX unit 28 may pass the encoded video information to the video decoder unit 30.

[0069] The video decoder unit 30 may decode the encoded video information so as to obtain the decoded video data. The encoded video information may include syntax information defined by the video encoder unit 18. This syntax information may include one or more elements ("syntax elements") that can be useful for decoding some or all of the encoded video information. The syntax elements may include, for example, characteristics of the encoded video information. The syntax elements may also include characteristics of the non-encoded video used to form the encoded video information and / or may describe the processing of the non-encoded video.

[0070] The video decoder unit 30 may output the decoded video data for later storage and / or display on an external display (not shown). Alternatively, the video decoder unit 30 may output the decoded video data to the display device 32. The display device 32 may be any individual various display devices, a plurality of various display devices, a combination of various display devices adapted to display the decoded video data to the user, and / or may include them. Examples of such display devices include liquid crystal displays (LCDs), plasma displays, organic light emitting diode (OLED) displays, cathode ray tubes (CRTs), and the like.

[0071] Communication channel 16 may be any wireless communication medium or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines, or any combination of wireless and wired media. Communication channel 16 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. Communication channel 16 generally represents any suitable communication medium or set of different communication media for transmitting video data from source device 12 to destination device 14, including any suitable combination of wired or wireless media. Communication channel 16 may include a router, a switch, a base station, or any other device that can be useful for facilitating communication from source device 12 to destination device 14. Details of an exemplary communication system that can facilitate such communication between devices 12 and 14 are provided below with reference to FIGS. 8, 9A-9E. Details of the devices that can represent devices 12 and 14 are also provided below.

[0072] Video encoder unit 18 and video decoder unit 30 may operate according to one or more standards and / or specifications, such as H.264 (``H.264 / SVC'') as extended according to, for example, MPEG-2, H.261, H.263, H.264, H.264 / AVC, SVC extensions. However, it will be understood that the methods, apparatuses, and systems provided herein are applicable to other video encoders, decoders, and / or CODECs implemented according to (and / or compliant with) different standards or proprietary specifications including video encoders, decoders, and / or CODECs that have not yet been developed. However, further, the techniques provided herein are not limited to any particular coding standard.

[0073] The relevant portions of H.264 / AVC described above are incorporated herein by reference and are available from the International Telecommunications Union as ITU-T Recommendation H.264, which may also be referred to herein as the H.264 standard and H.264 specification, or the H.264 / AVC standard or specification, or more specifically, from "ITU-T Rec. H.264 and ISO / IEC 14496-10 (MPEG4-AVC), 'Advanced Video Coding for Generic Audiovisual Service', v5, March, 2010". The H.264 / AVC standard has been developed by the ITU-T Video Coding Experts Group (VCEG) together with ISO / IEC MPEG as a collective partnership product known as the Joint Video Team (JVT). In some aspects, the techniques provided herein may generally be applicable to devices compliant with the H.264 standard. The JVT continues to consider extensions to the H.264 / AVC standard.

[0074] Work to advance the H.264 / AVC standard has been undertaken in various ITU-T forums such as the Key techniques Area (KTA) forum. At least some of the forums are partially exploring advances in coding techniques that exhibit higher coding efficiency than those indicated by the H.264 / AVC standard. For example, ISO / IEC MPEG and ITU-T VCEG established the Joint Collaborative Team on Video Coding (JCT-VC) to begin developing the next generation video coding and / or compression standard, namely, the High Efficiency Video Coding (HEVC) standard. In some aspects, the techniques provided herein can result in coding improvements for the H.264 / AVC and / or HEVC (currently under drafting) standards, and / or coding improvements in accordance with them.

[0075] Although not shown in FIG. 2, each of the video encoder unit 18 and the video decoder unit 30 may include and / or may be integrated with an audio encoder and / or an audio decoder (if necessary). The video encoder unit 18 and the video decoder unit 30 may include an appropriate MUX-DEMUX unit, or other hardware and / or software, to handle the encoding of both audio and video in a common data stream or alternatively in separate data streams. If applicable, the MUX-DEMUX unit may comply with, for example, the ITU-T Recommendation H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0076] The plurality of video encoder units 18 and video decoder units 30, or each of them, may be included in one or more encoders or decoders. Any one of the one or more encoders or decoders may be integrated as part of a CODEC and may be integrated with or otherwise combined with each respective camera, computer, mobile device, subscriber device, broadcast device, set-top box, and server, etc. Further, the video encoder units 18 and video decoder units 30 may be implemented as any one of various suitable encoder circuits and decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Alternatively, either or both of the video encoder unit 18 and video decoder unit 30 may be substantially implemented in software, and thus, the operations of the elements of the video encoder unit 18 and / or video decoder unit 30 are performed by appropriate software instructions executed by one or more processors (not shown). Again, such embodiments may also include off-chip components such as an external storage device (in the form of non-volatile memory), an input / output interface, etc., in addition to the processor.

[0077] In other embodiments, some of the elements of video encoder unit 18 and video decoder unit 30 may be implemented as hardware, and others may be implemented using appropriate software instructions executed by one or more processors (not shown). In any embodiment where the operations of elements of video encoder unit 18 and / or video decoder unit 30 may be executed by software instructions executed by one or more processors, such software instructions may be maintained on a magnetic disk, optical disk, and any other volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read only memory (“ROM”)) mass storage device system readable by a CPU. The computer-readable media may include cooperative computer-readable media or interconnected computer-readable media that exist exclusively on a processing system or are distributed among a plurality of interconnected processing systems that can be local or remote to the processing system.

[0078] The 3D Graphics Subgroup of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Joint Technical Committee 1 / SC29 / Working Group 11 (JTC1 / SC29 / WG11) Moving Picture Experts Group (MPEG) has developed a 3D Point Cloud Compression (PCC) standard that includes (1) a geometry-based compression standard for static point clouds, and (2) a video-based compression standard for dynamic point clouds. Those standards can enable the storage and transmission of 3D point clouds. Those standards can also support both irreversible and reversible coding of point cloud geometry coordinates and attributes.

[0079] FIG. 3 is a diagram showing the structure of a bitstream for video-based point cloud compression (V-PCC).

[0080] Referring to FIG. 3, a bitstream, e.g., a generated video bitstream, and metadata may be multiplexed together to generate a V-PCC bitstream. The bitstream syntax, e.g., the bitstream syntax of the V-PCC standard associated with MPEG, may be defined as shown in Table 1.

[0081] [Table 1]

[0082] Table 1 V-PCC bitstream syntax

[0083] Referring to the bitstream syntax of FIG. 3, the bitstream may start with a global header, e.g., a global header applicable to the entire PCC bitstream, followed by a series of group-of-frame (GOF) units. A GOF, e.g., one GOF unit, can result in a representation of any number of PCC frames (e.g., a concatenated representation) that share characteristics definable in the GOF header (e.g., the header at the beginning of the GOF unit and / or the header at the start of the GOF unit). That is, a GOF unit may include a GOF header and a series of component streams following it.

[0084] The component stream may include one or more video streams (e.g., a video stream for texture, one or two video streams for geometry), and a metadata stream. However, the present disclosure is not limited thereto, and the component stream may include any number of metadata streams. The metadata stream may include sub-streams such as, for example, a sub-stream for an occupancy map and a sub-stream for auxiliary information. The information in the metadata stream may be associated with a geometry frame and may be used to reconstruct a point cloud. The streams within the GOF unit may (1) be in order and (2) not be interleaved per frame.

[0085] FIG. 4 is a diagram showing the structure of a V-PCC bitstream as a series of V-PCC units.

[0086] In the version of the V-PCC community draft, the bitstream may be composed of a set of V-PCC units, as shown in FIG. 4. For example, as defined in the V-PCC CD, the syntax of the V-PCC units is shown in Table 2 below. In such a case, each V-PCC unit has a V-PCC unit header and a V-PCC unit payload. The V-PCC unit header describes the V-PCC unit type, as shown in Table 3 below. The V-PCC units with unit types 2, 3, and 4 may be defined as occupancy data unit, geometry data unit, and attribute data unit, respectively (e.g., in the V-PCC CD). Those data units represent (e.g., the three (e.g., main) components required) for reconstructing the point cloud. In addition to the V-PCC unit type, the V-PCC attribute unit header also specifies the attribute type and its index, which enables multiple instances of the same attribute type to be supported.

[0087] The payloads of the occupancy V-PCC unit, geometry V-PCC unit, and attribute V-PCC unit correspond to video data units (e.g., HEVC NAL units) that can be decoded by the video decoder specified in the corresponding occupancy V-PCC unit, geometry V-PCC unit, and attribute V-PCC unit.

[0088] [Table 2]

[0089] Table 2 V-PCC Unit Syntax

[0090] [Table 3]

[0091] Table 3 V-PCC Unit Header Syntax

[0092] [Table 4]

[0093] Table 4 V-PCC Unit Payload Syntax

[0094] The V-PCC CD specifies a V-PCC bitstream as a set of V-PCC units, and the set of V-PCC units consists of five types of V-PCC units: VPCC_SPS, VPCC_PSD, VPCC_OVD, VPCC_GVD, and VPCC_AVD. VPCC_SPS is referenced by other unit types via vpcc_sequence_parameter_set_id in the unit header.

[0095] Figure 5 shows the V-PCC unit data types, unit header syntax, and references to the active sequence parameter set (SPS). The SPS includes sequence level syntax elements such as sps_frame_width, sps_frame_height, sps_layer_count_minus1, and configuration flags. The SPS also includes syntax structures such as profile_tier_level, occupancy_parameter_set, geometry_parameter_set, and one or more attribute_parameter_sets.

[0096] VPCC_PSD also includes a plurality of PSD parameter set unit types, such as PSD_SPS, PSD_GFPS, PSD_GPPS, PSD_AFPS, PSD_APPS, PSD_FPS, and PSD_PFLU. Each parameter set may refer to a different sequence level parameter set or PSD level parameter set, and each parameter set includes, for example, a plurality of override flags, enable flags, or present flags in order to reduce overhead.

[0097] Figure 6 is a diagram showing the SPS parameter set and the PSD parameter set. The parameter sets included in SPS and PSD, as well as the reference links between the parameter sets and the upper level parameter sets, are shown in Figure 6. The dashed lines in Figure 6 indicate that the parameters in the upper level parameter set may be overwritten by the lower level parameter set.

[0098] <ISO Base Media File Format> The file format for time-based media may include several parts in accordance with the MPEG standard, for example, the ISO / IEC 14496 (MPEG-4) standard. For example, those parts may be based on, included in, and / or derived from the ISO Base Media File Format (ISOBMFF), which is a structural media-independent definition.

[0099] The file format according to ISOBMFF can support (e.g., can include, etc.) structural information and / or media data information about the timed presentation of media data, such as audio, video, virtual / augmented reality, etc. ISOBMFF can also support un-timed data, such as metadata at different levels within the file structure. According to ISOBMFF, a file may have a logical structure of a video such that the video can include a set of tracks that are temporally parallel. According to ISOBMFF, a file may have a time structure such that a track can include, for example, a sequence of samples over time. The sequence of samples may be mapped to the timeline of the entire video. ISOBMFF is based on the concept of a box-structured file. A box-structured file may include a series of boxes having sizes and types (e.g., the boxes may be referred to as atoms). According to ISOBMFF, a type may be identified according to a 32-bit value represented by four printable characters, also known as a 4-character code (4CC). According to ISOBMFF, un-timed data may be included, for example, in a metadata box at the file level, or may be added to a video box or a stream of timed data, such as a track within a video.

[0100] An ISOBMFF container includes boxes, which may be referred to as MovieBox (moov) and can contain metadata about a (e.g., continuous) media stream included in a file (e.g., a container). The metadata may be signaled within the MovieBox, e.g., within the hierarchy of boxes within a TrackBox (trak). A track can represent a continuous media stream included in a file. The media stream may be a sequence of samples, such as audio access units or video access units of an elementary media stream, and may be enclosed within a MediaDataBox (mdat) present at the topmost level of the file (e.g., a container). The metadata for each track may include, for example, a list of sample description entries, each of which provides (1) the coding and / or encapsulation format used in the track, and (2) initialization data for processing the format. Each sample may be associated with a sample description entry for the track. An explicit timeline map (e.g., for each track) may be defined using tools, such as an edit list. The edit list may be signaled using an EditListBox, and each entry may define a portion of the track timeline either by (1) mapping a portion of the composition timeline, or (2) indicating an empty time (e.g., an "empty" edit in cases where a portion of the presentation timeline map is not mapped to media). The EditListBox may have the following syntax.

[0101]

Number

[0102] Media files may be incrementally generated using tools such as fragmentation, may be downloaded progressively, and / or may be adaptively streamed. In accordance with ISOBMFF, a fragmented container may include a MovieBox followed by a series of fragments, such as video fragments. Each video fragment may include (1) a MovieFragmentBox (moof) that can include a subset of the sample table, and (2) a MediaDataBox (mdat) that can include samples of a subset of the sample table. The MovieBox may include only non-sample-specific information, such as track description information and / or sample description information. Within a video fragment, a set of track fragments may be represented by several TrackFragmentBox (traf) instances. A track fragment may have zero or more track runs, and a track run can record (e.g., represent) a consecutive run of samples for that track. The MovieFragmentBox may include a MovieFragmentHeaderBox (mfhd) that can include a sequence number (e.g., a number starting from 1 and sequentially changing in value for each video fragment in the file).

[0103] <3D point cloud> To enable new forms of interaction and communication with VR and / or new media, 3D point clouds may be used for new media such as VR and immersive 3D graphics. MPEG has developed standards for defining bitstreams for compressed dynamic point clouds through its 3D working group. The bitstream defined in the MPEG standard is organized into a series of group of frames (GOF) units, and each GOF unit contains a series of component streams for several frames. In the case of the MPEG standard bitstream, the PCC decoder may need to analyze the entire bitstream, for example starting from the first bit, to search for a particular GOF and / or synchronize GOF boundaries. In such cases, since the PCC frames are not internally interleaved within the GOF unit, the entire GOF unit needs to be accessed (e.g., read, stored, etc.) for safe decoding and reconstruction. Further, in such cases, the presentation timing information is specific to the frame timing information of the video-coded component bitstream. Also, in such cases, the video codec used for the component stream may not be signaled at a higher level within the PCC bitstream, and the PCC bitstream does not provide support for media profiles, tiers, and / or levels that are specific to PCC.

[0104] According to an embodiment, for example, a bitstream such as a PCC bitstream may be based on (e.g., conform to, be similar to, etc.) the ISOBMFF. For example, the file format for a V-PCC bitstream may be based on the ISOBMFF. According to an embodiment, the V-PCC bitstream can provide flexible storage and extraction of components of the PCC stream (e.g., different components, a plurality of components, a set of components, etc.). According to an embodiment, the V-PCC bitstream may be reconstructed as an ISOBMFF bitstream (e.g., in the manner of the ISOBMFF, according to the ISOBMFF, similar to the ISOBMFF, conforming to the ISOBMFF).

[0105] FIG. 7 is a diagram showing the mapping of the GOF stream to the video fragment.

[0106] Fragments, for example, ISOBMFF fragments, may be used to define (e.g., identify, delineate, demarcate, etc.) the V-PCC bitstream. Referring to FIG. 7, each video fragment, for example, may be defined by (1) mapping the GOF header data to the MovieFragmentBox and (2) mapping the GOF video stream and / or GOF metadata (e.g., auxiliary information, occupancy map, etc.) to the MediaDataBox of the video fragment. In the case of FIG. 7, each GOF unit may be mapped to an ISOBMFF fragment, or in other words, a one-to-one mapping between the GOF unit and the video fragment is shown.

[0107] In addition, the parameter set reference structure design for the VPCC patch sequence data unit (VPCC_PSD) can be problematic in certain cases. That is, the case where patch_frame_parameter_set refers to the active patch sequence parameter set via pfps_patch_sequence_parameter_set_id, the active patch parameter set via pfps_geometry_patch_frame_parameter_set_id, and the active attribute patch parameter set via pfps_attribute_patch_frame_parameter_set_id is a problematic case. Each active geometry patch parameter set refers to the active geometry frame parameter set via gpps_geometry_frame_parameter_set_id, and each active geometry frame parameter set refers to the active patch sequence parameter set via gfps_patch_sequence_parameter_set_id. Further, each active attribute patch parameter set refers to the active attribute frame parameter set via apps_attribute_frame_parameter_set_id, and each active attribute frame parameter set refers to the active patch sequence parameter set via afps_patch_sequence_parameter_set_id.

[0108] In the problematic cases described above, when the values of pfps_patch_sequence_parameter_set_id, gfps_patch_sequence_parameter_set_id, and afps_patch_sequence_parameter_set_id are different, the patch frame parameter set may end up referring to three different active patch sequence parameter sets, which can be problematic when the different active patch sequence parameter sets contain different parameter values.

[0109] <ISOBMFF-based V-PCC bitstream> FIG. 8 is a diagram showing a V-PCC bitstream structure according to an embodiment.

[0110] According to an embodiment, the V-PCC bitstream structure may be based on the ISOBMFF bitstream structure. According to an embodiment, for example, the items and / or elements shown in FIG. 8 may be mapped to (for example, corresponding) ISOBMFF boxes. According to an embodiment, the component stream may be mapped to, for example, individual tracks within a container file. According to an embodiment, the component stream for the V-PCC stream may include either (1) one or more (for example, two or three) video streams for either geometry information or texture information, and (2) one or more temporal metadata streams for either occupancy maps or auxiliary information.

[0111] According to an embodiment, other component streams (for example, other than the type of component stream discussed above) may be included in the V-PCC stream. For example, the other stream may include a stream for any number or type of attributes associated with points of a point cloud, for example, a 3D point cloud. According to an embodiment, for example, a (for example, additional) temporal metadata track may be included in the container file to provide GOF header information. According to an embodiment, metadata may be signaled. According to an embodiment, metadata such as information describing the characteristics of the component stream and / or the relationship between different tracks within the file may be signaled using, for example, tools provided according to the MPEG standard.

[0112] According to an embodiment, samples for media and / or timed metadata tracks may be included in a MediaDataBox (mdat). According to an embodiment, samples of a stream may be sequentially stored in a MediaDataBox. For example, in the case of media storage, samples of each stream may be stored together in a MediaDataBox with the continuous stream, so that there can be a continuation that includes all samples of a first stream followed by another continuation that includes all samples of a second stream.

[0113] According to an embodiment, samples of a component (e.g., a component stream) may be divided into chunks. For example, samples of a component stream may be divided into chunks according to any of the sizes of GOF units. According to an embodiment, chunks may be interleaved. Chunks may be interleaved within a MediaDataBox, for example, to support progressive download of a V-PCC bitstream. According to an embodiment, chunks may be of different sizes (i.e., may have different sizes), and samples within a chunk may be of different sizes (i.e., may have different sizes).

[0114] SampleToChunkBox (stsc) may be included in the SampleTableBox (stbl) of a track, and the SampleToChunkBox (stsc) may contain a table. According to an embodiment, the SampleToChunkBox may be used to discover (e.g., indicate them, or use them to determine) any of a chunk containing samples, a position associated with a chunk (e.g., one or more samples), or information describing samples associated with a chunk. According to an embodiment, the ChunkOffsetBox (stco or co64) may be included in the SampleTableBox (stbl) of a track and may indicate (e.g., provide) the index of each chunk within the containing file (e.g., within a container).

[0115] <Geometry and Texture Track> According to an embodiment, a component video stream of a PCC bitstream may be mapped to a track within an ISOBMFF container file. For example, each component video stream (e.g., each of a texture stream and a geometry stream) within a PCC bitstream may be mapped to a track within an ISOBMFF container file. In such cases, an access unit (AU) of a component stream may be mapped to a sample for the corresponding track. There may be cases where a component stream, e.g., a texture stream and a geometry stream, is not directly rendered.

[0116] According to embodiments, a constrained video scheme may be used to signal post-decoder requirements associated with a track of a component stream. For example, a constrained video scheme, such as defined according to ISOBMFF, may be used to signal post-decoder requirements associated with tracks of a texture stream and a geometry stream. According to embodiments, signaling post-decoder requirements associated with a track of a component stream can enable a player / decoder to inspect a file (e.g., a container) and identify requirements for rendering a bitstream. According to embodiments, signaling post-decoder requirements associated with a track of a component stream may not enable a legacy player / decoder to decode and / or render the component stream. According to embodiments, a constrained scheme (e.g., a constrained video scheme) may be applied to either the geometry track and the texture track of a PCC bitstream.

[0117] According to embodiments, either the geometry track and the texture track may be a constrained video scheme track (e.g., may be converted to, labeled as, and considered as such). According to embodiments, for either the geometry track and the texture track, the respective sample entry code may be set to the four-character code (4CC) "resv", and a RestrictedSchemeInfoBox may be added to the respective sample description while, for example, leaving all other boxes unmodified. According to embodiments, the original sample entry type, which can be based on the video codec used to encode the stream, may be stored in an OriginalFormatBox within the RestrictedSchemeInfoBox.

[0118] The properties of the constraints (e.g., scheme type) may be defined in the SchemeTypeBox, and the information associated with the scheme (e.g., the data required for it) may be stored in the SchemeInformationBox, for example, as defined by ISOBMFF. The SchemeTypeBox and SchemeInformationBox may be stored within the RestrictedSchemeInfoBox. According to an embodiment, a scheme_type field (e.g., included in the SchemeTypeBox) may be used to indicate a restricted scheme for point cloud geometry. For example, in the case of a geometry video stream track, the scheme_type field included in the SchemeTypeBox may be set to "pctx" indicating that the nature of the constraint is a restricted scheme for point cloud geometry. As another example, in the case of a texture video stream track, the scheme_type field may be set to "pctx" indicating a restricted scheme for point cloud texture. The PCCDepthPlaneInfoBox may be included in the SchemeInformationBox of each track. According to an embodiment, in the case where two or more geometry tracks are present in a file (e.g., a container), the PCCDepthPlaneInfoBox can indicate the depth image plane information for each track (e.g., identify it, include information indicating it, etc.). For example, in the case where two geometry tracks are present, the depth image plane information can indicate which track contains the video stream of depth image plane 0 and which track contains the video stream of depth image plane 1. According to an embodiment, the PCCDepthPlaneInfoBox may include a depth_image_layer which can be a field containing the depth image plane information. For example, the depth_image_layer may be the index of the depth image plane (e.g., information indicating it), where the value 0 indicates depth image plane 0, the value 1 indicates depth image plane 1, and other values are reserved for later use.According to the embodiment, the PCCDepthPlaneInfoBox including the depth_image_layer is,

[0119]

Number

[0120] According to the embodiment, in the case where (1) multiple layers are available for either the geometric component or the texture component, and (2) any number of components are carried in the component track, those layers may be signaled in the PCCComponentLayerInfoBox within the SchemeInformationBox of the track. The PCCComponentLayerInfoBox is,

[0121]

Number

[0122] According to the embodiment, the semantics for the PCCComponentLayerInfoBox may include (1) min_layer can indicate the index of the minimum layer for the V-PCC component carried by the track, and (2) max_layer can indicate the index of the maximum layer for the V-PCC component carried by the track.

[0123] According to the embodiment, the V-PCC texture component may be a subtype of a (e.g., more) general video-coded component type, which may also be referred to as a V-PCC attribute component (e.g., may be considered as such). Further, a set of attribute tracks may be present in a container where a subset of those tracks can carry information about texture attributes. The attribute track may be, for example, a constrained video scheme track having a scheme_type field of a SchemeTypeBox set to the 4CC "pcat". The PCCAttributeInfoBox within the SchemeInformationBox can identify the type of the attribute, and the value of attribute_type can indicate the type of the attribute, for example, as defined in the V-PCC CD. The PCCAttributeInfoBox is

[0124]

Number

[0125] Video coders for encoding texture video streams and geometry video streams are not restricted. Further, the texture video stream and the geometry video stream may be encoded using different video codecs. According to an embodiment, a decoder (e.g., a PCC decoder / player) may identify the codec (e.g., the type of codec) used for a component video stream. For example, the PCC decoder / player may identify the type of codec used for a particular component video stream by checking the sample entry of its track within an ISOBMFF container file. The header of each GOF within the V-PCC stream may include a flag, such as absolute_d1_flag, indicating how geometry layers other than the layer closest to the projection plane are coded. In cases where absolute_d1_flag is set, two geometry streams may be used to reconstruct the 3D point cloud, and in cases where absolute_d1_flag is not set, only one geometry stream may be used to reconstruct the 3D point cloud.

[0126] According to an embodiment, the value of absolute_d1_flag may vary across GOF units. For example, during one or more periods within the presentation time, there are no samples in the second geometry track. According to an embodiment, the value of absolute_d1_flag that varies across GOF units may be signaled using an EditListBox in the second geometry track. According to an embodiment, a parser (e.g., included in a PCC decoder / player) may determine whether it is possible to reconstruct the second geometry track based on the information in the edit list. For example, the PCC decoder / player may determine whether it is possible to reconstruct the second geometry track by checking the edit list of the second geometry track for sample availability at a given timestamp.

[0127] <Occupancy Map and Auxiliary Information Track> According to an embodiment, the decoder may use either an occupancy map or auxiliary information to reconstruct the 3D point cloud. For example, on the decoder side, the point cloud may be reconstructed from the geometry stream using the occupancy map and auxiliary information. The occupancy map and auxiliary information may be part of a stream other than the geometry stream within each GOF unit. According to an embodiment, the occupancy map and auxiliary information may be included in a (e.g., separate) time-limited metadata track, which may be referred to as an occupancy map track. According to an embodiment, a sample for the occupancy map track may include either the occupancy map or the auxiliary information for a single frame. According to an embodiment, the occupancy map track may be identified by the following sample entry in the sample description of the track.

[0128]

Number

[0129] According to an embodiment, two time-limited metadata tracks may be used to convey the occupancy map information and the auxiliary information separately. According to an embodiment, the occupancy map track may have a sample entry as shown above for the case of a single combined occupancy map and auxiliary information track. According to an embodiment, the time-limited metadata track for the auxiliary information may have the following sample entry in its sample description.

[0130]

Number

[0131] According to an embodiment, auxiliary information such as patch data may be conveyed in the samples of the point cloud metadata track, and for example, a separate auxiliary information track may not be required.

[0132] According to an embodiment, the occupancy map may be coded using a video coder, and the generated video stream may be placed in a constrained video scheme track. According to an embodiment, the scheme_type field of the SchemeTypeBox of the constrained video scheme track may be set to "pomv", for example, to indicate the constrained video scheme of the point cloud occupancy map.

[0133] <Point Cloud Metadata Track> Metadata for the PCC bitstream may appear at different levels within the bitstream, for example, within the global header and within the headers of GOF units. Further, the metadata may be applicable at either the frame level or the patch level for the occupancy map. According to an embodiment, the point cloud metadata track may include metadata associated with either the global header or the GOF header. According to an embodiment, the point cloud metadata track may be a time-limited metadata track (e.g., separate, single, etc.), and the metadata information may be organized as described below.

[0134] The global header information may be applied to all GOF units within the stream. According to an embodiment, the global header information may be stored in the sample description of the timing metadata track, which may be considered as an entry point when parsing the PCC file. According to an embodiment, the PCC decoder / player that decodes / plays the PCC stream may search for this timing metadata track within the container. According to an embodiment, this timing metadata track may be identified by a PointCloudSampleEntry within the sample description of the track. According to an embodiment, the PointCloudSampleEntry may include a PCCDecoderConfigurationRecord to provide, for example, either (1) information regarding the PCC profile of the bitstream, and (2) information regarding a video codec that the player may need to support in order to decode the component stream. According to an embodiment, the PointCloudSampleEntry may also include a PCCHeaderBox to include, for example, information signaled in the global bitstream header (e.g., of MPEGV-PCC).

[0135] According to an embodiment, the syntax of the PointCloudSampleEntry may be as follows.

[0136]

Number

[0137] According to the embodiment, the semantics for the fields of PCCHeaderStruct may be that (1) pcc_category2_container_version indicates the version of the PCC bitstream, (2) gof_metadata_enabled_flag indicates whether PCC metadata is enabled at the GOF-level, (3) gof_scale_enabled_flag indicates whether scaling is enabled at the GOF-level, (4) gof_offset_enabled_flag indicates whether offsetting is enabled at the GOF-level, (5) gof_rotation_enabled_flag indicates whether rotation is enabled at the GOF-level, (6) gof_point_size_enabled_flag indicates whether point size is enabled at the GOF-level, and (7) gof_point_shape_enabled_flag indicates whether point shape is enabled at the GOF-level. According to the embodiment, the semantics for the fields of PCCDecoderConfigurationRecord may be that (1) configurationVersion is the version field, and changes that are not compatible with the record are indicated by changes in the version number within the version field, (2) general_profile_space specifies the context for the interpretation of general_profile_idc, (3) general_tier_flag specifies the tier context for the interpretation of general_level_idc, (4) general_profile_idc indicates the profile to which the coded point cloud sequence conforms when general_profile_space is equal to 0, and (5) general_level_idc indicates the level to which the coded point cloud sequence conforms.

[0138] According to an embodiment, information applied to a GOF unit (e.g., any information applied to all GOF units) may be stored in the sample description of the timing metadata track. According to an embodiment, a field of the PCCDecoderConfigurationRecord may be part of the PCCHeaderStruct. According to an embodiment, the PCCHeaderBox may be the top-level box within the MovieBox. According to an embodiment, the PCC decoder / player can (e.g., easily) identify whether it can decode and play a file, for example, without having to parse all the tracks in the file to discover the PCC metadata track, and can determine whether the listed profiles are supported. According to an embodiment, each sample in the point cloud metadata track may include GOF header information, for example, as defined according to MPEG V-PCC. According to an embodiment, the syntax of the GOFHeaderSample and the GOFHeaderStruct, which is a data structure including all fields defined in the GOF header, is shown as follows.

[0139]

Number

[0140] According to an embodiment, a parser (e.g., the PCC decoder / player) may identify how many frames are in a GOF unit by parsing the GOF metadata sample. For example, the parser may identify how many frames are in a GOF unit so that, for example, an exact number of samples can be read from the geometry video track and the texture video track. According to an embodiment, the point cloud metadata track may be linked to a component video track. For example, the track reference tool of the ISOBMFF standard may be used to link the point cloud metadata track to a component video track.

[0141] According to an embodiment, the content description reference "cdsc" may be used to link a PCC metadata track to a component track. Or in other words, a content description reference "cdsc" from a PCC metadata track to a component track may be generated. According to an embodiment, the link may be formed by (1) adding a TrackReferenceBox to a TrackBox (e.g., inside), and (2) placing a TrackReferenceTypeBox of type "cdsc" inside the TrackReferenceBox. According to an embodiment, the TrackReferenceTypeBox may include any number of track_IDs that specify the component video track(s) that the PCC metadata refers to. According to an embodiment, for example, instead of "cdsc", a new track reference type for the PCC bitstream may be defined. According to an embodiment, the chain of track references may be used by (1) adding a "cdsc" track reference from a PCC metadata track to a geometry video track(s), and (2) adding an "auxl" track reference from the geometry video track(s) to an occupancy map and texture track(s).

[0142] According to an embodiment, for example, instead of a timing metadata track, a point cloud parameter set track may be used. According to an embodiment, the point cloud parameter set track may be similar to an AVC parameter set track, for example, as defined by ISO / IEC. According to an embodiment, the sample entry for this track may also be defined as follows.

[0143]

Number

[0144] According to an embodiment, a PCC parameter stream sample entry may include a PCC parameter stream configuration box, and the PCC parameter stream configuration box may be defined as follows.

[0145]

Number

[0146] According to an embodiment, samples in a PCC parameter set track may have equal decoding times, for example, when the first frame of the corresponding GOF is decoded / when the parameter set(s) become valid at that time (e.g., in that instance).

[0147] According to an embodiment, in a case where a bitstream is structured as a series of V-PCC units as described, for example, in a V-PCC CD, a parameter set V-PCC unit may be identified, for example, by a media handler type 4CC "vpcc" and carried in a (e.g., new type of) track having a sample entry of type "vpc1". According to an embodiment, a track (e.g., of a new type) identified by a media handler type 4CC "vpcc" may be

[0148]

Number

[0149] According to an embodiment, the vpcc_unit_payload array may include the payload of the sequence level parameter set (e.g., only that). According to an embodiment, for example, in a case where the sequence parameter set is defined to include any one of an occupancy parameter set, a geometry parameter set, or an attribute parameter set as defined in the V-PCC CD, the vpcc_unit_payload array may include the sequence parameter set V-PCC unit (e.g., only that). According to an embodiment, in a case where a plurality of sequence level parameter sets are defined, the vpcc_unit_payload may be, for example, the payload of one of the sequence level parameter sets (e.g., any one of a sequence parameter set, a geometry parameter set, an occupancy parameter set, or an attribute parameter set) by separating the sequence parameter set from other component parameter sets (e.g., a geometry parameter set, an occupancy parameter set, and an attribute parameter set) (e.g., as necessary). According to an embodiment, in a case where the patch unit sequence parameter set (e.g., PSD_SPS as defined in the V-PCC CD) includes information applicable to the entire sequence, the PSD_SPS payload may also be stored (e.g., also that) in the vpcc_unit_payload array of the VPCCSampleEntry. According to an embodiment, for example, as an alternative to directly extending the SampleEntry, the VPCCSampleEntry may be defined to extend the SampleEntry and may be defined to extend the (e.g., newly defined) VolumentricSampleEntry that can provide a basic sample entry type for volumetric media. The samples in this track may correspond to point cloud frames. Each V-PCC sample may include any number of vpcc_unit_payload instances, for example, with the constraint of including only the patch_sequence_data V-PCC unit payload.Samples corresponding to the same frame across a component track may have the same composition time as the corresponding samples for that frame within the V-PCC track.

[0150] According to an embodiment, the VPCCSampleEntry may be such that the vpcc_unit_payload array contains the payload of a sequence level parameter set, e.g., a sequence parameter set (e.g., only that), and, if separate, the payloads of a geometry parameter set, an occupancy parameter set, and an attribute parameter set (e.g., only that). According to an embodiment, the VPCCSampleEntry may

[0151]

Number

[0152] <Component track having multiple layers> A component track can carry more than one layer of components, and a playback device can be capable of (e.g., needs to) identifying and extracting samples belonging to a particular layer. According to an embodiment, a sample grouping function (e.g., of ISO / IEC 14496-12) may be utilized. According to an embodiment, for example, a new sample group description having a grouping type set to the 4CC "vpld" for grouping component layer samples may

[0153]

Number

[0154] According to an embodiment, the semantics for VPCCLayerSampleGroupEntry can be that (1) the layer_index can be the index of the layer to which the samples of the group belong, and (2) the absolute_coding_flag can indicate whether the samples of the layer associated with the sample group depend on samples from another layer sample group. In the case where the absolute_coding_flag is set to 1, the samples may not depend on samples from another layer. In the case where the absolute_coding_flag is set to 0, the samples can depend on samples from another layer, and (3) the predictor_layer_index can be the index of the layer on which the samples of the group depend.

[0155] According to an embodiment, the mapping of samples to the corresponding layer group may be performed using SampleToGroupBox, for example, as defined in ISO / IEC14496-12. SampleToGroupBox may include, for example, several entries, and each entry associates several consecutive samples with one of the group entries in SampleGroupDescriptionBox.

[0156] A single point of an entry for point cloud data in a container file According to an embodiment, information about (e.g., all) the tracks that make up a single V-PCC content may be signaled at a single location within the container file. For example, a player can identify those tracks and their types as early as possible without having to parse the sample descriptions of each track. According to an embodiment, such early identification may be achieved, for example, by signaling the track information in one box at the top level of the container file or within a MetaBox ("meta") that exists at the top level of the file.

[0157] According to an embodiment, such a box may be a (e.g., newly defined) box that has a new box type or inherits from and extends EntityToGroupBox, as defined in ISO / IEC 14496-12. According to an embodiment, the information to be signaled may include a list of trackIDs of (e.g., all) tracks belonging to the V-PCC content. For each track to be signaled, along with the track type (e.g., metadata, occupancy map, geometry, etc.), the component layer carried by the track, if applicable, may be signaled (e.g., only signaled) in such a box. According to an embodiment, such a box may also include information regarding the profile and level of the content (e.g., that too). According to an embodiment, such a box (e.g., a newly defined box) for carrying the above-mentioned information may be

[0158]

Number

[0159] According to an embodiment, the semantics for the fields of the VPCCContentBox may be that (1) the content_id is a unique id for the V-PCC content among all the V-PCC contents stored in the container, and the num_tracks indicates the total number of tracks that are part of the V-PCC content, (2) the track_id is a trackID of one of the tracks stored in the container, (3) the track_type indicates the type of the component track (e.g., texture, geometry, metadata, etc.), (4) the min_layer indicates the index of the minimum layer for the V-PCC component carried by the track, and (5) the max_layer indicates the index of the maximum layer for the V-PCC component carried by the track.

[0160] According to an embodiment, another example for the definition of the V-PCC content information box may be, for example, when extending the EntityToGroupBox as defined by ISO / IEC. That is, according to an embodiment, the V-PCC content information box

[0161]

Number

[0162] According to an embodiment, the semantics of any of track_type, min_layer, and max_layer may be the same as the semantics of the corresponding fields for the VPCCContentBox defined above.

[0163] <Alternative Versions of Point Cloud Content and Component Signaling> According to an embodiment, in a case where more than one version of the same point cloud is available in an ISOBMFF container (e.g., the same point cloud at different resolutions), each version may have a separate point cloud metadata track.

[0164] According to an embodiment, alternative track mechanisms defined in ISO / IEC 14496-12 may be used to signal that those tracks are alternatives to each other. According to an embodiment, point cloud metadata tracks that are alternatives to each other may have the same value for the alternate_group field in their respective TrackHeaderBox(es) within the ISOBMFF container (e.g., if required).

[0165] Similarly, when multiple versions (e.g., bitrates) of a point cloud component (e.g., any of a geometry component, an occupancy component, or an attribute component) are available, the alternate_group fields within the TrackHeaderBox(es) for components of different versions may have the same value (e.g., if required).

[0166] According to an embodiment, a single point cloud metadata track carrying metadata for different versions of the same point cloud may be available in the ISOBMFF container. According to an embodiment, the sequence parameter set for each version may be signaled in a separate sample entry within the SampleDescriptionBox for the track's sample table. The type of those sample entries may be VPCCSampleEntry. According to an embodiment, a sample grouping function (e.g., of ISO / IEC 14496-12) may be used to group samples within the point cloud metadata track belonging to each version.

[0167] <Fragmented ISOBMFF Container for V-PCC Bitstream> FIG. 9 is a diagram showing a fragmented ISOBMFF container for a V-PCC bitstream according to an embodiment.

[0168] According to an embodiment, a GOF unit may be mapped to an ISOBMFF video fragment. Referring to FIG. 9, each video fragment may correspond to one or more GOF units in a (e.g., elementary) V-PCC bitstream. According to an embodiment, a video fragment may include only samples for the corresponding GOF unit. According to an embodiment, for example, the entire bitstream such as a global stream header, and metadata regarding the number of tracks present in (e.g., included in) the container may be stored in a MovieBox. According to an embodiment, the MovieBox may include a (e.g., one) TrackBox for each component stream and an (e.g., additional) TrackBox for a GOF header timing metadata track.

[0169] According to an embodiment, a case of one-to-one mapping may exist, in which case each video fragment includes only one GOF unit. In such a case, there may be no need for a GOF header timing metadata track. According to an embodiment, the GOF header may be stored in a MovieFragmentHeaderBox. According to an embodiment, the MovieFragmentHeaderBox may include an optional box including a PCCGOFHeaderBox. According to an embodiment, the PCCGOFHeaderBox may be defined as follows.

[0170]

Number

[0171] In an embodiment, in a case where a V-PCC elementary stream is composed of a set of V-PCC units, the V-PCC sequence parameter set information may be included in a VPCCSampleEntry for the point cloud metadata track in the MovieBox.

[0172] <Multiple point cloud streams> According to an embodiment, the ISOBMFF container may include more than one V-PCC stream. According to an embodiment, each stream may be represented by a set of tracks. According to an embodiment, track grouping (e.g., a track grouping tool) may be used to identify the stream to which a track belongs. According to an embodiment, for example, for one PCC stream, a TrackGroupBox ("trgr") may be added to (1) the TrackBox of all component streams, and (2) the PCC metadata track. According to an embodiment, the syntax for the PCCGroupBox can define a (e.g., new) type of track grouping, the TrackGroupTypeBox may be defined according to ISOBMFF, and may include a single track_group_id field. According to an embodiment, the syntax for the PCCGroupBox may be as follows.

[0173]

Number

[0174] According to an embodiment, tracks belonging to the same PCC stream may have the same track_group_id (e.g., the same value for it) for track_group_type "pccs", and tracks belonging to different PCC streams may have different / respective track_group_ids. According to an embodiment, a PCC stream may be identified according to the track_group_id in a TrackGroupTypeBox having a track_group_type equal to "pccs".

[0175] According to an embodiment, for example, in a case where a plurality of point cloud streams are included (e.g., allowed) in a single container, the PCCHeaderBox may be used to indicate the operation point and the global header of each PCC stream. According to an embodiment, the syntax of the PCCHeaderBox may be as follows.

[0176]

Number

[0177] According to an embodiment, the semantics of the identified fields may be that (1) number_of_pcc_streams can indicate how many point cloud streams can be stored in the file, and (2) pcc_stream_id can be a unique identifier for each point cloud stream corresponding to the track_group_id for the track of the component stream.

[0178] <Signaling of PCC Profile> To implement a media coding standard in an interoperable manner, for example, among various applications having similar functional requirements, profiles, tiers, and levels may be (e.g., may be specified to be) used as conformance points. A profile can define a set of coding tools and / or algorithms used in generating (e.g., conforming) bitstreams, and a level can define (e.g., can impose) constraints on (e.g., specific, important, etc.) parameters of a bitstream, such as parameters corresponding to either decoder processing load or memory capabilities, etc.

[0179] According to an embodiment, for example, a brand may be used to indicate conformance to a V-PCC profile by indicating the brand in a track-specific manner. ISOBMFF includes the concept of a brand, which can be indicated using the compatible_brands list within the FileTypeBox. Each brand is a 4-character code registered with ISO that identifies a concise specification. The presence of a brand within the compatible_brands list of the FileTypeBox may be used to indicate that the file conforms to the brand requirements. Similarly, a TrackTypeBox (e.g., within a TrackBox) may be used to indicate the conformance of an individual track to a specific brand. According to an embodiment, for example, a brand may be used to indicate conformance to a V-PCC profile because a TrackTypeBox can have a syntax similar or identical to the syntax of the FileTypeBox and can be used to indicate a brand in a track-specific manner. According to an embodiment, a V-PCC profile may also be signaled as part of a PCCHeaderBox. According to an embodiment, a V-PCC profile may also be signaled in a VPCCContentBox, for example, by referring to a single point of an entry for point cloud data within a container file as defined above.

[0180] <VPCC Parameter Set Reference> As discussed above, the specific parameter set reference structure designed for VPCC_PSD can be problematic

[0181] FIG. 10 is a diagram showing a PSD parameter set reference structure according to an embodiment.

[0182] According to an embodiment, for example, in contrast to the problematic structure, the parameters of the frame-level geometry parameter set and the attribute parameter set may be integrated into a single component parameter set. According to an embodiment, such a single component parameter set may refer to a single active patch sequence parameter set, and the parameters of the geometry patch parameter set and the attribute patch parameter set may be integrated into a single component patch parameter set, and the single component patch parameter set refers to an active geometry attribute frame parameter set. According to an embodiment, the patch frame parameter set may refer to a single active geometry attribute patch parameter set. The processed PSD parameter set reference structure is shown in FIG. 10.

[0183] FIG. 11 is a diagram showing another PSD parameter set reference structure according to an embodiment.

[0184] According to an embodiment, any of the geometry frame parameter set and the attribute frame parameter set may be included in the patch sequence parameter set. That is, the parameters of the geometry patch parameter set and the attribute patch parameter set may be combined to form a component_patch_parameter_set. According to an embodiment, the component_patch_parameter_set may refer to an active patch sequence parameter set. According to an embodiment, as shown in FIG. 11, the patch frame parameter set may refer to an active component patch parameter set.

[0185] <Support for Spatial Access and Signaling of the Target Region>

[0186] A target region (RoI) within the point cloud may be defined by a 3D bounding box. According to an embodiment, for example, patches resulting from the projection of points within the RoI may be packed into a set of tiles within a 2D frame of any of the geometry component, occupancy component, and attribute component. According to an embodiment, the tiles (e.g., a set of tiles within a 2D frame) may be encoded with high quality / resolution, and the tiles may be (e.g., subsequently) coded independently. For example, the tiles may be coded independently as HEVC MCTS tiles, and their respective samples may be stored in separate ISOBMFF tracks. This can enable (e.g., facilitate) spatial random access to the RoI without the need to decode the entire 2D frame.

[0187] According to an embodiment, for example, corresponding 2D tile tracks across components of a point cloud (e.g., within it) may be grouped together using, for example, a track grouping tool (e.g., as discussed above). According to an embodiment, a TrackGroupBox ("trgr") may be added to a TrackBox associated with component tracks (e.g., all of them). A new type of track grouping for 2D tile tracks of V-PCC component tracks may have a TrackGroupTypeBox (as defined according to ISO / IEC) and may include a single track_group_id field according to an embodiment. The new type of track grouping may be defined as

[0188] [Number] as follows.

[0189] According to an embodiment, tracks belonging to the same point cloud 2D tile may have the same value of track_group_id for track_group_type "p2dt". According to an embodiment, the track_group_id of tracks associated with a point cloud 2D tile may be different from the track_group_id of tracks associated with another (e.g., any other) point cloud 2D tile. The track_group_id within a TrackGroupTypeBox having a track_group_type equal to "p2dt" may be used as an identifier for the point cloud 2D tile.

[0190] According to an embodiment, a 3D RoI within a point cloud may be associated with any number of point cloud 2D tiles using, for example, a VPCCRegionsOfInterestBox. According to an embodiment, the VPCCRegionsOfInterestBox

[0191]

Number

[0192] According to an embodiment, the semantics for the fields of 3DRegionBox and / or VPCCRegionsOfInterestBox may include any of the following: (1) region_x can be the x-coordinate of the reference point of the bounding box; (2) region_y can be the y-coordinate of the reference point of the bounding box; (3) region_z can be the z-coordinate of the reference point of the bounding box; (4) region_width can indicate the length of the bounding box along the x-axis; (5) region_height can indicate the length of the bounding box along the y-axis; (6) region_depth can indicate the length of the bounding box along the z-axis; (7) roi_count can indicate the number of RoIs in the point cloud; (8) 2d_tile_count can indicate the number of point cloud 2D tiles associated with the RoI; and (9) track_group_ids can be an array of track group identifiers for track groups of type "p2dt" (e.g., corresponding to point cloud 2D tiles).

[0193] According to an embodiment, in the case where the RoI in the point cloud sequence is static (e.g., does not change), the VPCCRegionsOfInterestBox may be included in either the VPCCSampleEntry in the PCC metadata track or the VPCCContentGroupingBox in the MetaBox. According to an embodiment, in the case where the RoI in the point cloud sequence is dynamic, the VPCCRegionsOfInterestBox may be signaled in the samples of the PCC metadata track.

[0194] <Conclusion> Although the features and elements have been described above in specific combinations, those skilled in the art will recognize that each feature or element may be used alone or in any combination with other features and elements. Additionally, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electrical signals (transmitted through wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). The processor associated with the software may be used to implement a radio frequency transceiver for use in a TRU, UE, terminal, base station, RNC, or any host computer.

[0195] Moreover, in the above-described embodiments, mention has been made of processing platforms, computing systems, controllers, and other devices including processors. Those devices may include at least one central processing unit (“CPU”) and memory. In accordance with the convention of those skilled in computer programming, references to the operations of algorithms or instructions and symbolic representations may be performed by various CPUs and memories. Such operations and algorithms or instructions may be referred to as “executed,” “executed by a computer,” or “executed by a CPU.”

[0196] One of ordinary skill in the art will recognize that operations and symbolically represented computations or instructions involve the manipulation of electrical signals by a CPU. An electronic system represents data bits whose conversions or deformations resulting from electrical signals and the maintenance of data bits at memory locations within a memory system can reconfigure or alter other processing of the signals along with the operations of the CPU. The memory location where the data bits are maintained is a physical location corresponding to the data bits or having specific electrical, magnetic, optical, or organic characteristics representing the data bits. It should be understood that representative embodiments are not limited to the platforms or CPUs mentioned above, and that other platforms and CPUs can support the provided methods.

[0197] Data bits may also be maintained on a computer-readable medium including magnetic disks, optical disks, and any other volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage device systems readable by a CPU. The computer-readable medium may exist exclusively on a processing system or may be distributed across a plurality of interconnected processing systems local or remote to the processing system. It may include cooperating or interconnected computer-readable media. It will be understood that representative embodiments are not limited to the memories mentioned above, and that other platforms and memories can support the described methods.

[0198] In an exemplary embodiment, any of the computations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium. The computer-readable instructions may be executed by a processor of a mobile unit, network element, and / or any other computing device.

[0199] There is little distinction between the hardware implementation and the software implementation of the system aspects. The use of hardware or software is generally a design choice that represents a cost - effectiveness trade - off (e.g., although not always, in certain contexts, the choice between hardware and software may become important). There may thus be various vehicles (e.g., hardware, software, and / or firmware) that can act on the processes, systems, and / or other technologies described herein, and the preferred vehicle may vary with the context in which the processes, systems, and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are paramount, the implementer may choose a vehicle mainly of hardware and / or firmware. If flexibility is paramount, the implementer may choose some combination of hardware, software, and / or firmware.

[0200] The foregoing detailed description has shown various embodiments of devices and / or processes via the use of block diagrams, flowcharts, and / or examples. As long as such block diagrams, flowcharts, and examples include one or more functions and / or operations, those skilled in the art will understand that each function and / or operation within such block diagrams, flowcharts, or examples may be implemented individually and / or collectively by a wide range of hardware, software, firmware, or any combination thereof. Suitable processors include, by way of example, general - purpose processors, special - purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application - specific integrated circuits (ASICs), application - specific standard products (ASSPs), field - programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0201] Features and elements were provided above in certain combinations, but one of ordinary skill in the art will recognize that each feature or element may be used alone or in any combination with other features and elements. This disclosure is not limited to the specific embodiments described in this application, which are intended as examples of various aspects. As will be apparent to those of ordinary skill in the art, many modifications and variations may be made without departing from its spirit and scope. Elements, operations, or instructions used in the description of this application should not be construed as important or essential to the invention unless expressly provided as such. In addition to the methods and apparatuses listed herein, functionally equivalent methods and apparatuses within the scope of the disclosure will be apparent to those of ordinary skill in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims. This disclosure will thus be limited only by such claims, according to the full scope of equivalents to which the appended claims are entitled. It will be understood that this disclosure is not limited to a particular method or system.

[0202] Also, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, when referred to herein, the terms "base station" and its abbreviation "STA", "user equipment" and its abbreviation "UE" may mean (i) a wireless transmit and / or receive unit (WTRU) as described hereinafter, (ii) any of several embodiments of a WTRU as described hereinafter, (iii) in particular, a wireless-enabled device and / or a wired-enabled device (e.g., a tetherable one) constituted by some or all of the structures and functionalities of a WTRU as described hereinafter, (iii) a wireless-enabled device and / or a wired-enabled device constituted by structures and functionalities of a WTRU that are not all of those as described hereinafter, or (iv) something similar. Details of an exemplary WTRU that can represent (or be interchangeable with) any UE or mobile device described herein are provided below with respect to FIGS. 1A through 1D.

[0203] In certain representative embodiments, some portions of the subject matter described herein may be implemented via application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, one of ordinary skill in the art will recognize that some aspects of the embodiments disclosed herein may be equivalently implemented in an integrated circuit as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as any combination thereof, and that designing the circuitry and / or writing code for the software and / or firmware is well within the skill of one of ordinary skill in the art in view of the present disclosure. Additionally, one of ordinary skill in the art will recognize that the mechanisms of the subject matter described herein may be distributed as a program product in a variety of forms and that the illustrative embodiments of the subject matter described herein apply regardless of the particular type of signal bearing medium actually used to carry out the distribution. Examples of signal bearing media include, but are not limited to, recordable types of media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, computer memories, etc., and transmission types of media such as digital and / or analog communication media (e.g., optical fiber cables, waveguides, wired communication links, wireless communication links, etc.).

[0204] The subject matter described in this specification may refer to different components that are included within or connected to different other components. It should be understood that such described architectures are merely examples, and in fact, many other architectures that achieve the same functionality may be implemented. In a conceptual sense, any arrangement of components for achieving the same functionality may be effectively associated so as to be able to achieve the desired functionality. Thus, any two components combined herein to achieve a particular functionality may be considered to be "associated" with each other to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components so associated may also be considered to be "operably connected" or "operably coupled" to each other to achieve the desired functionality, and any two components that can be so associated may also be considered to be "operably couplable" to each other to achieve the desired functionality. Specific examples of operably couplable include, but are not limited to, components that can physically engage and / or physically interact, components that wirelessly interact and / or wirelessly interact with each other, and / or components that logically interact and / or logically interact with each other.

[0205] Regarding the use of substantially any plural and / or singular terms herein, one of ordinary skill in the art can translate from plural to singular and / or from singular to plural depending on the context and / or application. Various singular / plural substitutions may be explicitly shown herein for simplicity.

[0206] Generally, the terms used herein, and especially those used in the appended claims (e.g., the body of the claims), will be understood by those of ordinary skill in the art to be interpreted generally as “open terms” (e.g., the term “having” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “including but not limited to,” etc.). Further, where a specific number of introduced claim recitations is intended, such intent will be clearly recited in the claim, and where no such recitation is present, it will be understood by those of ordinary skill in the art that no such intent exists. For example, if only one item is intended, the term “single” or similar language may be used. By way of aid in understanding, the following appended claims and / or the description herein may include the use of introductory phrases “at least one” and “one or more” to introduce claim recitations. However, even when the same claim uses an introductory phrase “one or more” or “at least one” and an indefinite article such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”), the use of such phrases should not be interpreted as suggesting that the introduction of a claim recitation by the indefinite article “a” or “an” limits any particular claim that includes the recited claim to an embodiment that includes only one such recitation. This also applies to the use of definite articles used to introduce claim recitations. Additionally, even where a specific number of introduced claim recitations is clearly recited, those of ordinary skill in the art will recognize that such recitation should be interpreted as intending to mean at least the recited number (e.g., a mere recitation of “two” recitations without other modifiers means at least two recitations or two or more recitations).

[0207] Furthermore, in those examples where conventions such as "at least one of A, B, and C" are used, generally, such a structure is intended in the sense that one of ordinary skill in the art would understand the structure (e.g., a system having "at least one of A, B, and C" includes, but is not limited to, A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those examples where conventions such as "at least one of A, B, or C" are used, generally, such a structure is intended in the sense that one of ordinary skill in the art would understand the structure (e.g., a system having "at least one of A, B, or C" includes, but is not limited to, A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Furthermore, any virtual disjunctive words and / or phrases that exist in two or more alternative terms should be understood to contemplate including one of the terms, any of the terms, or both terms, regardless of whether they are in the description, claims, or drawings. For example, the phrase "A or B" would be understood to include the possibility of including "A" or "B" or "A and B". Furthermore, as used herein, the term "any of a plurality of items and / or a list of categories of a plurality of items that follows" is intended to include "any combination", "any plurality of combinations and / or any combination" of the "plurality of items and / or a plurality of items" together with individual and / or other items and / or other categories of items. Moreover, as used herein, the term "set" or "group" is intended to include any number of items including zero. Additionally, as used herein, the term "number" is intended to include any number including zero.

[0208] In addition, although the disclosed features or aspects have been described in terms of Markush groups, one of ordinary skill in the art will recognize that the disclosure thereby describes any number of individual Markush groups or any number of subgroups of Markush groups.

[0209] As will be understood by those skilled in the art, for any and all purposes, such as providing the described description, all ranges disclosed herein include all possible sub-ranges and combinations of those sub-ranges. Any recited range will be readily recognized as being fully described and enabled to be decomposed into at least equal halves, thirds, quarters, fifths, tenths, etc. of the same range. By way of non-limiting example, each range disclosed herein may be readily decomposed into lower thirds, middle thirds, upper thirds, etc. Also, as will be understood by those skilled in the art, all language such as "up to", "at least", "above", and "below" includes the recited number and then refers to a range that can be decomposed into sub-ranges as discussed above. Finally, as will be understood by those skilled in the art, ranges include each individual member. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on.

[0210] Moreover, the claims should not be read as being limited to the order or elements provided unless so indicated. In addition, the use of the term "means" in any claim is not intended to invoke 35 U.S.C. § 112, paragraph 6, or the format of a means-plus-function claim, and any claim without the term "means" is not so intended.

[0211] Processors associated with software may be used to implement radio frequency transceivers for use in a wireless transmit / receive unit (WTRU), user equipment (UE), terminal, base station, mobility management entity (MME) or evolved packet core (EPC), or any host computer. The WTRU may include hardware and / or software implemented modules, along with other components such as a software defined radio (SDR), as well as a camera, video camera module, video phone, speakerphone, vibrating device, speaker, microphone, television transceiver, hands-free headset, keyboard, Bluetooth® module, frequency modulation (FM) radio unit, near field communication (NFC) module, liquid crystal display (LCD) display unit, organic light emitting diode (OLED) display unit, digital music player, media player, video game player module, Internet browser, and / or any wireless local area network (WLAN) or ultra wideband (UWB) module, and may be used with modules.

[0212] Although the invention has been described with respect to a communication system, it is contemplated that the system may be implemented in software on a microprocessor / general purpose computer (not shown). In certain embodiments, one or more of the functions of the various components may be implemented in software that controls the general purpose computer.

[0213] In addition, although the invention has been illustrated and described with reference to specific embodiments, the invention is not intended to be limited to the details shown. Rather, various modifications may be made in the equivalent scope of the claims and without departing from the invention.

[0214] Throughout the disclosure, those skilled in the art will understand that specific exemplary embodiments may be used in alternative embodiments or in combination with other exemplary embodiments.

[0215] Although the features and elements have been described above in specific combinations, those skilled in the art will recognize that each feature or element may be used alone or in any combination with other features and elements. Additionally, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks, and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). The processor associated with the software may be used to implement a radio frequency transceiver for use in a WRTU, UE, terminal, base station, RNC, or any host computer.

[0216] Furthermore, in the embodiments described above, mention has been made of processing platforms, computing systems, controllers, and other devices including processors. Those devices may include at least one central processing unit (“CPU”) and memory. In accordance with the convention of those skilled in computer programming, references to the operations of operations or instructions and symbolic representations may be performed by various CPUs and memories. Such operations and operations or instructions may be referred to as “executed,” “executed by a computer,” or “executed by a CPU.”

[0217] One skilled in the art will recognize that operations and symbolically represented operations or instructions involve the manipulation of electrical signals by a CPU. An electronic system represents data bits whose conversions or deformations resulting from electrical signals and the maintenance of data bits at memory locations within a memory system can reconfigure or alter other processing of the signals along with the operation of the CPU. The memory location where a data bit is maintained is a physical location corresponding to the data bit or having specific electrical, magnetic, optical, or organic characteristics representing the data bit.

[0218] Data bits may also be maintained on a computer-readable medium including magnetic disks, optical disks, and any other volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage device systems readable by a CPU. The computer-readable medium may exist exclusively on a processing system or be distributed across a plurality of interconnected processing systems local or remote to the processing system. It may include cooperating or interconnected computer-readable media. It will be understood that representative embodiments are not limited to the memory mentioned above and that other platforms and memories can support the methods described.

[0219] Suitable processors include, by way of example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), application specific standard products (ASSPs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0220] Although the invention has been described in connection with a communication system, it is contemplated that the system may be implemented in software on a microprocessor / general purpose computer (not shown). In certain embodiments, one or more of the functions of the various components may be implemented in software that controls a general purpose computer.

[0221] In addition, although the invention has been illustrated and described above with reference to particular embodiments, the invention is not intended to be limited to the details shown. Rather, various modifications may be made in detail within the scope of equivalents of the claims and without departing from the invention.

Claims

1. 1. A method for communicating decoding information for a point cloud (PC) bitstream of a coded PC sequence, comprising: mapping a metadata component bitstream of the PC bitstream to a first constrained video track of an International Organization for Standardization / International Electrotechnical Commission Base Media File Format (ISOBMFF), the first constrained video track of the ISOBMFF including one or more access units of the metadata component bitstream; mapping a geometry sub-bitstream of the PC bitstream to a second constrained video track of the ISOBMFF, the second constrained video track of the ISOBMFF including one or more access units of the geometry sub-bitstream; mapping a dedicated sub-bitstream of the PC bitstream to a third constrained video track of the ISOBMFF, the third constrained video track of the ISOBMFF including one or more access units of the dedicated sub-bitstream; linking the metadata component bitstream to the geometry sub-bitstream and the occupation sub-bitstream using a track referencing tool such that the first constrained video track contains indications of a plurality of component bitstreams of the PC bitstream encoded in the ISOBMFF, the plurality of component bitstreams including the geometry sub-bitstream and the occupation sub-bitstream; generating an ISOBMFF container to transmit the mapping of the metadata component bitstream, the mapping of the geometry sub-bitstream, and the mapping of the occupation sub-bitstream to the PC bitstream; A method for providing the above.

2. adding a constrained scheme information box to the first constrained video track, the second constrained video track and the third constrained video track, the constrained scheme information box indicating that each video track is a constrained video track; The method of claim 1 further comprising:

3. 2. The method of claim 1, wherein the second constrained video track and the third constrained video track include an indication of one or more layers of the PC bitstream.

4. The method of claim 3 , wherein each of the one or more layers is associated with a respective depth relative to a depth image plane.

5. 2. The method of claim 1, wherein the first constrained video track includes a plurality of samples associated with the metadata component bitstream of the PC bitstream, the second constrained video track includes a plurality of second samples associated with the geometry sub-bitstream of the PC bitstream, and the third constrained video track includes a plurality of third samples associated with the occupancy sub-bitstream of the PC bitstream.

6. 2. The method of claim 1, wherein the first constrained video track, the second constrained video track and the third constrained video track in one coded PC sequence are signaled in a single location of the ISOBMFF.

7. mapping an attribute substream of the PC bitstream to a fourth constrained video track of the ISOBMFF, the fourth constrained video track of the ISOBMFF including one or more access units of the attribute substream; The method of claim 1 further comprising:

8. The method of claim 7 , wherein the attribute substreams are associated with attribute types including color, transparency, time of acquisition, and another material property.

9. The method of claim 1 , further comprising generating one or more timed metadata component bitstreams, wherein samples associated with the timed metadata component bitstreams are contained in a MediaDataBox.

10. 1. An apparatus comprising: a circuit for communicating decoding information for a PC bitstream of a coded point cloud (PC) sequence, the apparatus comprising: the circuitry includes any one of a transmitter, a receiver, a processor, and a memory; mapping a metadata component bitstream of the PC bitstream to a first constrained video track of an International Organization for Standardization / International Electrotechnical Commission Base Media File Format (ISOBMFF), the first constrained video track of the ISOBMFF including one or more access units of the metadata component bitstream; Mapping a geometry sub-bitstream of the PC bitstream to a second constrained video track of the ISOBMFF, the second constrained video track of the ISOBMFF including one or more access units of the geometry sub-bitstream; mapping a dedicated sub-bitstream of the PC bitstream into a third constrained video track of the ISOBMFF, the third constrained video track of the ISOBMFF including one or more access units of the dedicated sub-bitstream; linking the metadata component bitstream to the geometry sub-bitstream and the occupation sub-bitstream using a track referencing tool such that the first constrained video track includes indications of a plurality of component bitstreams of the PC bitstream encoded in the ISOBMFF, the plurality of component bitstreams including the geometry sub-bitstream and the occupation sub-bitstream; Generate an ISOBMFF container to transmit the mapping of the metadata component bitstream, the mapping of the geometry sub-bitstream, and the mapping of the occupation sub-bitstream to the PC bitstream. The apparatus is configured to:

11. The circuit comprises:

11. The apparatus of claim 10, further configured to add a constrained scheme information box to the first constrained video track, the second constrained video track and the third constrained video track, the constrained scheme information box indicating that each video track is a constrained video track.

12. the first constrained video track includes a plurality of samples associated with the metadata component bitstream of the PC bitstream, the second constrained video track includes a plurality of second samples associated with the geometry sub-bitstream of the PC bitstream, and the third constrained video track includes a plurality of third samples associated with the occupancy sub-bitstream of the PC bitstream; 11. The apparatus of claim 10, wherein the second constrained video track and the third constrained video track include an indication of one or more layers of the PC bitstream.

13. The apparatus of claim 12 , wherein each of the one or more layers is associated with a respective depth relative to a depth image plane.

14. The apparatus of claim 10 , wherein the metadata component bitstream includes a first reference to the geometry sub-bitstream and a second reference to the occupancy sub-bitstream.

15. 11. The apparatus of claim 10, wherein the first constrained video track, the second constrained video track, and the third constrained video track in one coded PC sequence are signaled in a single location of the ISOBMFF.

16. 11. The apparatus of claim 10, wherein the circuitry is further configured to map an attribute substream of the PC bitstream to a fourth constrained video track of the ISOBMFF, the fourth constrained video track of the ISOBMFF including one or more access units of the attribute substream.

17. The apparatus of claim 16 , wherein the attribute substreams are associated with attribute types including color, transparency, and capture.

18. 11. The apparatus of claim 10, wherein the circuitry is further configured to generate one or more timed metadata component bitstreams, samples associated with the timed metadata component bitstreams being included in a MediaDataBox.

19. The method of claim 1 , wherein the metadata component bitstream includes a first reference to the geometry sub-bitstream and a second reference to the occupancy sub-bitstream.

20. 1. An apparatus comprising: a circuit for decoding a PC bitstream of an encoded point cloud (PC) sequence, the apparatus comprising: the circuitry includes any one of a transmitter, a receiver, a processor, and a memory; a first constrained video track of an International Organization for Standardization / International Electrotechnical Commission Base Media File Format (ISOBMFF) that includes one or more access units of a metadata component bitstream of said PC bitstream; a second constrained video track of the ISOBMFF that includes one or more access units of a geometry sub-bitstream of the PC bitstream; and a third constrained video track of the ISOBMFF that includes one or more access units of a dedicated sub-bitstream of the PC bitstream; the first constrained video track includes indications of a plurality of component bitstreams of the PC bitstream encoded in the ISOBMFF, the plurality of component bitstreams including the geometry sub-bitstream and the occupancy sub-bitstream; Decoding the first constrained video track, the second constrained video track, and the third constrained video track of the ISOBMFF container to display the PC bitstream. The apparatus is configured to:

Citation Information

Patent Citations

  • Separate track storage of texture and depth views for multiview coding plus depth

    CN104919801A

  • Method, apparatus and stream for immersive video format

    EP3349182A1

  • Design of track and operating point signaling in layered hevc file format

    JP2018524891A