Video-based point cloudstream

The video decoder efficiently compresses and delivers 3D point cloud data by parsing region identifiers and track groups within a media container file, addressing the challenge of large data sets in 3D reconstruction.

JP2026062770APending Publication Date: 2026-04-10INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2025-12-22
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently compressing and delivering 3D point cloud data due to the large number of points required to realistically reconstruct objects and scenes, necessitating improved representation and delivery techniques.

Method used

A video decoder processes a media container file containing a video-based point cloud compression (V-PCC) bitstream, parsing region identifiers and track groups to decode video tracks, utilizing timed metadata and sample entries for efficient compression and decoding of 3D regions.

Benefits of technology

Enables efficient compression and delivery of 3D point cloud data by accurately decoding and representing 3D regions, reducing storage and transmission bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062770000001_ABST
    Figure 2026062770000001_ABST
Patent Text Reader

Abstract

The present invention provides a video decoding apparatus and method, as well as a video encoding apparatus, for processing video data associated with a three-dimensional (3D) space. [Solution] In a video-based point cloud compression (V-PCC) container structure 600 used to enable spatial access to a specific region in 3D space, the 3D space 602 of the intocloud is divided into 3D cubes 602a, 602b, 602c, etc., representing multiple regions and / or objects in the 3D space, and the points belonging to each of the regions and / or objects in the 3D space are clustered, and bounding boxes are used to represent the regions or objects. The resulting patches from the projection of points within each bounding box are packed together into one or more tile groups in a 2D frame of the track, and are multiple V-PCC component streams or, for example, occupancy, geometry, attribute streams.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 852,046, filed May 23, 2019, and U.S. Provisional Patent Application No. 62 / 907,249, filed September 27, 2019, the entire disclosures of which are hereby incorporated by reference herein for all purposes.

Background Art

[0002] Background

[0002] Video coding systems can be used to compress and / or decompress digital video signals (e.g., to reduce the storage and / or transmission bandwidth required for such signals). 3 - dimensional (3D) point clouds have emerged as an advanced representation of immersive media. These point clouds can be captured in many ways, for example, using multiple cameras, depth sensors, and / or light detection and ranging (LiDAR) laser scanners. The number of points required to realistically reconstruct an object and / or scene in 3D space can be in the millions or billions. Therefore, efficient representation, compression, and / or delivery techniques for point cloud data are desirable.

Summary of the Invention

[0003] Summary

[0003] Systems, methods, and means for processing video data associated with a three-dimensional (3D) space are disclosed. A video decoder described herein may include a processor configured to receive a media container file (e.g., an ISO-based media file format (ISOBMFF) container file) containing a video-based point cloud compression (V-PCC) bitstream. The processor may parse the media container file and / or the V-PCC bitstream contained therein to determine the region identifier (ID) of a 3D region in the 3D space and the track group ID of each of the one or more track groups associated with the 3D space. The processor may determine that one or more track groups are associated with a 3D region based on the determination that the track group ID of each of the one or more track groups is linked to the region ID of the 3D region. The processor may decode video tracks belonging to one or more track groups (e.g., corresponding to one or more tiles in a 2D frame) to draw a visual representation of the 3D region in the 3D space. One or more track groups described herein may share a common track group type, and it may be determined further on the track group type that one or more track groups are associated with a 3D region. A medial container file may include one or more structures that define the number of regions associated with the 3D space and the number of track groups associated with each of these regions, and the processor may be configured to determine, based on the information contained in the structures, that the track group ID of each of the one or more track groups is linked to the region ID of the 3D region.

[0004]

[0004] The medial container file may include timed metadata containing information associated with a subset of the updated region, which may indicate updates to this subset of region (e.g., location, dimensions, etc.). Furthermore, the video track may include one or more sample entries, each of which may include an indication of the length of a data field indicating the network abstraction layer (NAL) unit size. The sample entries may further include an indication of the number of V-PCC parameter sets associated with the sample entry or the number of arrays of atlas NAL units associated with the sample entry. [Brief explanation of the drawing]

[0005] Brief explanation of the drawing [Figure 1A]

[0005] This is a system diagram showing an exemplary communication system in which one or more disclosed embodiments may be implemented. [Figure 1B]

[0006] This is a system diagram showing an exemplary wireless transmit / receive unit (WTRU) that may be used in the communication system shown in Figure 1A according to one embodiment. [Figure 1C]

[0007] This is a system diagram showing an exemplary radio access network (RAN) and an exemplary core network (CN) that may be used in the communication system shown in Figure 1A according to one embodiment. [Figure 1D]

[0008] This is a system diagram showing another exemplary RAN and another exemplary CN that may be used in the communication system shown in Figure 1A according to one embodiment. [Figure 2]

[0009] This shows an exemplary video-based point cloud compression (V-PCC) bitstream structure containing multiple V-PCC units. [Figure 3]

[0010] An example media container structure is shown. [Figure 4]

[0011] This section provides illustrative constraints for the alignment of intra-random access point (IRAP) samples for a component. [Figure 5]

[0012] An example is shown where the least common multiple of the IRAP periods is used to demonstrate V-PCC IRAP. [Figure 6]

[0013] This document illustrates an exemplary media container structure that can be used to enable spatial access to a specific region within a 3D space. [Modes for carrying out the invention]

[0006] Detailed explanation

[0014] A more detailed understanding can be obtained from the following explanation, which is provided as an example in conjunction with the attached drawings.

[0007]

[0015] Figure 1A is a diagram showing an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast communications to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content by sharing system resources (including wireless bandwidth). For example, the communication system 100 may employ one or more channel access methods (such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique-word OFDM (UW-OFDM), resource block filter OFDM, filter bank multicarrier (FBMC), etc.).

[0008]

[0016] As shown in Figure 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRU102a, 102b, 102c, and 102d (all also referred to as "stations" and / or "STAs") may be configured to transmit and / or receive radio signals and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, radio sensors, hotspots or Mi-Fi devices, "Internet of Things (IoT)" devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots, and / or other radio devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, and devices operating on commercial and / or industrial radio networks. Any of WTRU102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0009]

[0017] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks (such as CN106 / 115, the Internet 110, and / or other networks 112). As an example, base stations 114a and 114b may be base transceiver stations (BTS), node B, eNode B, Home Node B, Home eNode B, gNB, NR Node B, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are each described as single elements, it will be understood that they may include any number of interconnected base stations and / or network elements.

[0010]

[0018] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc.) (not shown). Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, and may also be called cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. Cells may provide coverage of radio services to a specific geographic area, which may be relatively fixed or change over time. Cells may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers (i.e., one for each sector of the cell). In one embodiment, the base station 114a may employ multiple-input multiple output (MIMO) technology and utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0011]

[0019] Base stations 114a and 114b may communicate with one or more WTRUs 102a, 102b, 102c, and 102d over a radio interface 116 which may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The radio interface 116 may be established using any suitable radio access technology (RAT).

[0012]

[0020] Specifically, as noted above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a and WTRUs 102a, 102b, and 102c within RAN 104 / 113 may implement radio technologies (such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA)) that can establish radio interfaces 115 / 116 / 117 using broadband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or HSPA+. HSPA may include High-Speed ​​Downlink Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0013]

[0020] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies (such as Evolved UMTS Terrestrial Radio Access (E-UTRA)) that can establish a radio interface 116 by using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).

[0014]

[0022] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as NR Radio Access, which can establish a radio interface 116 by using New Radio (NR).

[0015]

[0023] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together, for example, by using the dual connectivity (DC) principle. Accordingly, the radio interfaces utilized by WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions (e.g., eNB and gNB) transmitted to / from multiple types of base stations.

[0016]

[0024] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.

[0017]

[0025] For example, base station 114b in Figure 1A may be a wireless router, Home Node B, Home eNode B, or access point, and may utilize any suitable RAT to facilitate wireless connectivity in local areas such as offices, homes, vehicles, premises, industrial facilities, air corridors (e.g., for use by drones), and roadways. In one embodiment, base stations 114b and WTRUs 102c, 102d may implement wireless technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base stations 114b and WTRUs 102c, 102d may implement wireless technologies such as IEEE 802.15 to establish a wireless personal network (WPAN). In yet another embodiment, base stations 114b and WTRUs 102c, 102d may utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not be required to access the internet 110 via CN106 / 115.

[0018]

[0026] RAN104 / 113 may communicate with CN106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more WTRU102a, 102b, 102c, and 102d. The data may have various Quality of Service (QoS) requirements, including various throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in Figure 1A, it is understood that RAN104 / 113 and / or CN106 / 115 may communicate directly or indirectly with other RANs employing the same RAT as RAN104 / 113 or different RATs. For example, in addition to connecting to RAN104 / 113 which may utilize NR radio technology, CN106 / 115 may also communicate with another RAN (not shown) employing GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0019]

[0027] CN106 / 115 can also act as a gateway for WTRU102a, 102b, 102c, 102d to access the PSTN108, the Internet 110, and / or other networks 112. The PSTN108 can include a circuit-switched telephone network that provides traditional telephone service (POTS: plain old telephone service). The Internet 110 can include a global system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) within the TCP / IP Internet protocol suite. The network 112 can include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 can include another CN connected to one or more RANs that may employ the same RAT or a different RAT as the RAN104 / 113.

[0020]

[0028] Some or all of the WTRU102a, 102b, 102c, 102d within the communication system 100 may include multimode capabilities (e.g., the WTRU102a, 102b, 102c, 102d may include multiple transceivers for communicating with various wireless networks over various wireless links). For example, the WTRU102c shown in FIG. 1A can be configured to communicate with a base station 114a that may employ cellular-based wireless technology and a base station 114b that may employ IEEE 802 wireless technology.

[0021]

[0029] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, among many others, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the aforementioned elements, while conforming to the embodiment.

[0022]

[0030] The processor 118 could be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or other functionalities that enable the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to a transceiver 120 which can be coupled to a transmit / receive element 122. Figure 1B depicts the processor 118 and the transceiver 120 as separate components, but it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.

[0023]

[0031] The transmit / receive element 122 may be configured to transmit and / or receive signals over the radio interface 116 to or from a base station (e.g., base station 114a). For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be a emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of radio signals.

[0024]

[0032] Although the transmit / receive element 122 is depicted as a single element in Figure 1B, the WTRU 102 may include any number of transmit / receive elements 122. Specifically, the WTRU 102 may employ MIMO technology. Therefore, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals on the radio interface 116.

[0025]

[0033] The transceiver 120 may be configured to modulate the signal transmitted by the transmit / receive element 122 and demodulate the signal received by the transmit / receive element 122. As noted above, the WTRU 102 may have multimode capability. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0026]

[0034] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input data from there. The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132, and store data therein. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information from memory that is not physically located on the WTRU 102 (such as on a server or home computer (not shown)) and store data therein.

[0027]

[0035] The processor 118 may be configured to receive power from the power supply 134 and distribute and / or control this power to other components within the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, a fuel cell, etc.

[0028]

[0036] The processor 118 may also be coupled to a GPS chipset 136 which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) on the radio interface 116 and / or determine its position based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information by any preferred location determination method, while conforming to the embodiments.

[0029]

[0037] The processor 118 is further coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, e-compass, satellite transceiver, digital camera (for photography and / or video), Universal Serial Bus (USB) port, vibration device, television transceiver, hands-free headset, Bluetooth® module, frequency modulation (FM) radio unit, digital music player, media player, video game player module, internet browser, virtual reality and / or augmented reality (VR / AR) device, activity tracker, and the like. Peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, orientation sensor, proximity sensor, temperature sensor, time sensor, geographic position sensor, altimeter, light sensor, contact sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.

[0030]

[0038] WTRU102 may include a full-duplex radio where some or all of the transmission and reception of signals (e.g., associated with specific subframes of both UL (e.g., for transmission) and downlink (e.g., for reception) links may occur simultaneously. The full-duplex radio may include an interference management unit to reduce and / or nearly eliminate either self-interference via hardware (e.g., chokes) or self-interference via signal processing via a processor (e.g., a separate processor (not shown) or processor 118). In one embodiment, WRTU102 may include a half-duplex radio for the transmission and reception of some or all of the signals (e.g., associated with specific subframes of both UL (e.g., for transmission) and downlink (e.g., for reception) links).

[0031]

[0039] Figure 1C is a system diagram showing RAN104 and CN106 according to one embodiment. As noted above, RAN104 may employ E-UTRA radio technology to communicate with WTRU102a, 102b, and 102c on the radio interface 116. RAN104 may also communicate with CN106.

[0032]

[0040] RAN104 may include eNode-B 160a, 160b, and 160c. However, it should be understood that RAN104 may include any number of eNode-B, while conforming to the embodiment. Each of the eNode-B 160a, 160b, and 160c may include one or more transceivers for communicating with WTRU 102a, 102b, and 102c on the radio interface 116. In one embodiment, the eNode-B 160a, 160b, and 160c may implement MIMO technology. Thus, for example, the eNode-B 160a may use multiple antennas to transmit and / or receive radio signals to and from the WTRU 102a.

[0033]

[0041] Each of the eNode-B 160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling within UL and / or DL, etc. As shown in Figure 1C, the eNode-B 160a, 160b, and 160c may communicate with each other over the X2 interface.

[0034]

[0042] The CN106 shown in Figure 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. Although each of the aforementioned elements is depicted as part of CN106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0035]

[0043] The MME162 can be connected to each of the eNode-B 162a, 162b, and 162c within RAN104 via the S1 interface and can function as a control node. For example, the MME162 may be responsible for authenticating users of WTRU102a, 102b, and 102c, activating / deactivating bearers, and selecting specific serving gateways during initial task generation of WTRU102a, 102b, and 102c. The MME162 may also provide control plane functionality for switching between RAN104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0036]

[0044] The SGW164 can be connected to eNode B 160a, 160b, and 160c within RAN104 via the S1 interface. The SGW164 can typically route and forward user data packets to and from WTRU102a, 102b, and 102c. The SGW164 can also perform other functions, such as fixing the user plane during inter-eNode B handovers, triggering paging when DL data is available to WTRU102a, 102b, and 102c, and managing and storing the context of WTRU102a, 102b, and 102c.

[0037]

[0045] SGW164 may be connected to PGW166, which can provide WTRU102a, 102b, and 102c with access to a packet-switched network such as the Internet 110, in order to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices.

[0038]

[0046] CN106 can facilitate communication with other networks. For example, CN106 can provide WTRU102a, 102b, and 102c with access to a circuit-switched network such as PSTN108 to facilitate communication between WTRU102a, 102b, and 102c and traditional terrestrial circuit communication equipment. For example, CN106 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between CN106 and PSTN108. In addition, CN106 can provide WTRU102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0039]

[0047] Although the WTRU is described as a wireless terminal in Figures 1A to 1D, in some representative embodiments, it is intended that such a terminal may use a wired communication interface with a communication network (for example, temporarily or permanently).

[0040]

[0048] In a typical embodiment, the other network 112 may be a WLAN.

[0041]

[0049] In Infrastructure Basic Service Set (BSS) mode, a WLAN may have access points (APs) of the BSS and one or more stations (STAs) associated with the APs. APs may have access to or interfaces with a Distribution System (DS) or another type of wired / wireless network that carries traffic into and from the BSS. Traffic originating outside the BSS to an STA may reach the STA via an AP and be delivered to the STA. Traffic originating from an STA to a destination outside the BSS may be sent to an AP for delivery to its respective destination. Traffic between STAs within the BSS may be transmitted, for example, via an AP, where the source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted (e.g., directly) between the source and destination STAs via a direct link setup (DLS). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using Independent BSS (IBSS) mode may not have APs, and STAs within or using IBSS (e.g., all STAs) may communicate directly with each other. The IBSS mode of communication may occasionally be referred to herein as the “ad hoc” mode of communication.

[0042]

[0050] When using the 802.11ac infrastructure operating mode or a similar operating mode, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., 20 MHz wideband) or a dynamically set width via signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish association with the AP. In some typical embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. With respect to CSMA / CA, an STA, including the AP (e.g., any STA), may sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, that STA may withdraw. One STA (e.g., just one station) may transmit at a given time in a given BSS.

[0043]

[0051] A high-throughput (HT) STA may use a 40MHz wideband channel for communication via a combination of a primary 20MHz channel and adjacent or non-adjacent 20MHz channels to form a 40MHz wideband channel, for example.

[0044]

[0052] Very High Throughput (VHT) STAs can support 20MHz, 40MHz, 80MHz, and / or 160MHz wideband channels. 40MHz and / or 80MHz channels can be formed by combining adjacent 20MHz channels. A 160MHz channel, sometimes called an 80+80 configuration, can be formed by combining eight adjacent 20MHz channels or two discontinuous 80MHz channels. For 80+80 configurations, the channel-coded data can be passed through a segment analyzer that can split the data into two streams. Inverse Fast Fourier Transform (IFFT) and time-domain processing can be performed separately on each stream. The streams can be mapped to two 80MHz channels, and the data can be transmitted by a transmitting STA. At a receiver receiving the STA, the above operation for the 80+80 configuration can be reversed, and the combined data can be sent to Medium Access Control (MAC).

[0045]

[0503] Sub-1GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced to those used in 802.11n and 802.11ac for 802.11af and 802.11ah. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths by using the non-TVWS spectrum. According to a typical embodiment, 802.11ah may support Meter Type Control / Machine-Type communications such as MTC devices in a macro coverage area. MTC devices may have several capabilities (e.g., limited capabilities including support for several bandwidths and / or limited bandwidths (e.g., support only)). MTC devices may include batteries with battery life exceeding a threshold (e.g., to maintain very long battery life).

[0046]

[0054] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs within the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the minimum bandwidth operating mode among all STAs when operating within the BSS. In the 802.11ah example, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports 1MHz mode (e.g., only supports 1MHz mode), even if the AP and other STAs within the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes. Carrier sense and / or Network Allocation Vector (NAV) settings may depend on the state of the primary channel. If the primary channel is busy transmitting to the AP, for example due to an STA (which only supports 1MHz operation mode), then the entire available frequency band can be considered busy, even if the majority of the frequency band remains idle and may be available.

[0047]

[0055] In the United States, the usable frequency band for 802.11ah is 902MHz to 928MHz. In South Korea, the usable frequency band is 917.5MHz to 923.5MHz. In Japan, the usable frequency band is 916.5MHz to 927.5MHz. The total usable frequency band for 802.11ah is 6MHz to 26MHz, depending on the country identification code.

[0048]

[0056] Figure 1D is a system diagram showing RAN113 and CN115 according to one embodiment. As noted above, RAN113 may employ NR radio technology to communicate with WTRU102a, 102b, and 102c on the radio interface 116. RAN113 may also communicate with CN115.

[0049]

[0057] RAN113 may include gNB180a, 180b, and 180c, however, it should be understood that RAN113 may include any number of gNBs, while conforming to the embodiment. Each of gNB180a, 180b, and 180c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c on the radio interface 116. In one embodiment, gNB180a, 180b, and 180c may implement MIMO technology. For example, gNB180a and 108b may utilize beamforming to transmit and / or receive signals from gNB180a, 180b, and 180c. Thus, gNB180a may, for example, use multiple antennas to transmit and / or receive radio signals from WTRU102a. In one embodiment, gNB180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB180a may transmit multiple component carriers to WTRU102a (not shown). A subset of these component carriers may be on the unlicensed spectrum, while the remaining component carriers may be on the licensed spectrum. In one embodiment, gNB180a, 180b, and 180c may implement Coordinated Multi-Point (CoMP) technology. For example, WTRU102a may receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).

[0050]

[0058] WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c by using transmissions associated with scalable numerology. For example, OFDM symbol intervals and / or OFDM subcarrier intervals can vary with respect to various transmissions, various cells, and / or various parts of the radio transmission spectrum. WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c by using subframes or transmission time intervals (TTI) of varying or scalable lengths (for example, by including various numbers of OFDM symbols and / or by lasting for various lengths of absolute time).

[0051]

[0509] The gNB180a, 180b, and 180c can be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, the WTRU102a, 102b, and 102c can communicate with the gNB180a, 180b, and 180c without accessing other RANs (e.g., eNode-B 160a, 160b, and 160c). In a standalone configuration, the WTRU102a, 102b, and 102c can use one or more of the gNB180a, 180b, and 180c as mobility anchor points. In a standalone configuration, the WTRU102a, 102b, and 102c can communicate with the gNB180a, 180b, and 180c by using signals within the unlicensed bandwidth. In a non-standalone configuration, WTRU102a, 102b, and 102c communicate / connect with gNB180a, 180b, and 180c, while also communicating / connecting with other RANs such as eNode-B 160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c may implement DC principles to communicate with one or more gNB180a, 180b, and 180c and one or more eNode-B 160a, 160b, and 160c almost simultaneously. In a non-standalone configuration, eNode-B 160a, 160b, and 160c can act as mobility anchors for WTRU102a, 102b, and 102c, while gNB180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.

[0052]

[0060] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling within UL and / or DL, support for network slicing, dual connectivity, interaction between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a, 184b, and routing of control plane information to Access and Mobility Management Functions (AMF) 182a, 182b. As shown in Figure 1D, the gNB180a, 180b, and 180c may communicate with each other over the Xn interface.

[0053]

[0061] The CN115 shown in Figure 1D may include at least one AMF182a, 182b, at least one UPF184a, 184b, at least one Session Management Function (SMF)183a, 183b, and possibly a Data Network (DN)185a, 185b. Although each of the aforementioned elements is depicted as part of the CN115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0054]

[0062] AMF182a and 182b may be connected to one or more gNB180a, 180b, and 180c within RAN113 via the N2 interface and may function as control nodes. For example, AMF182a and 182b may be responsible for authenticating users of WTRU102a, 102b, and 102c, supporting network slicing (e.g., handling various PDU sessions with different requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, and mobility management. Network slicing may be used by AMF182a and 182b to customize CN support for WTRU102a, 102b, and 102c based on the type of services utilized by WTRU102a, 102b, and 102c. For example, various network slices can be established for various use cases, such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services using machine-type communication (MTC) access, and so on. The AMF162 may provide control plane functionality for switching between RAN113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies (such as WiFi).

[0055]

[0063] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b, and configure traffic routing through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0056]

[0064] UPF184a and 184b may be connected via an N3 interface to one or more gNB180a, 180b, and 180c in RAN113 (which may provide WTRU102a, 102b, and 102c with access to packet-switched networks such as the Internet 110) to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices. UPF184 and 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0057]

[0065] CN115 can facilitate communication with other networks. For example, CN115 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between CN115 and PSTN108. In addition, CN115 may provide WTRU102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a, 102b, 102c may be connected to local data networks (DN) 185a, 185b via UPF184a, 184b through an N3 interface to UPF184a, 184b and an N6 interface between UPF184a, 184b and DN185a, 185b.

[0058]

[0606] In terms of the correspondence between Figures 1A-1D and Figures 1A-1D, one or more or all of the functions described herein with respect to one or more of the WTRU102a-d, base stations 114a-b, eNode-B 160a-c, MME162, SGW164, PGW166, gNB180a-c, AMF182a-b, UPF184a-b, SMF183a-b, DN185a-b, and / or other devices described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.

[0059]

[0067] Emulation devices may be designed to perform tests on one or more other devices in a laboratory environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all of the functions of other devices in a communication network while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network. One or more emulation devices may perform one or more or all of the functions of other devices while temporarily implemented / deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for testing purposes and / or may perform testing by using wireless communication.

[0060]

[0068] One or more emulation devices may perform one or more functions (including all functions) while not being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test scenario in a test chamber and / or undeployed (e.g., test) wired and / or wireless communication network to perform testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via an RF circuit configuration (e.g., including one or more antennas) may be used by the emulation device to transmit and / or receive data.

[0061]

[0069] 3D point clouds (e.g., high-quality 3D point clouds) can be used to represent immersive media. A point cloud can contain one or more points (e.g., a set) that can be represented in 3D space using coordinates that indicate the location and / or one or more attributes of each point. For example, attributes may include one or more of the following associated with each point: color, transparency, time of acquisition, laser reflection, or material properties, etc. Point clouds can be captured in many ways. For example, multiple cameras and depth sensors can be used to capture a point cloud. Light detection and ranging (LiDAR) laser scanners can be used to capture a point cloud. The number of points contained in a point cloud to realistically reconstruct objects and / or scenes in 3D space can be approximately millions or billions. Efficient representation and compression can facilitate the storage and / or transmission of point cloud data.

[0062]

[0070] Figure 2 shows an exemplary structure 200 of a bitstream for video-based point cloud compression (V-PCC) that is transmitted (e.g., signaled) by an encoding device and can be analyzed and decoded by a decoding device. The V-PCC bitstream 200 may contain a set of one or more V-PCC units 202, and Table 1 contains exemplary syntax for signaling V-PCC units. Each V-PCC unit 202 may contain a V-PCC unit header 204 and a V-PCC unit payload 206, the V-PCC unit payload 206 may contain one or more sequence parameter sets 208, occupied video data 210, various types of patch data groups 212, geometry video data 214, or attribute video data 216. The V-PCC unit header 204 may define the V-PCC unit type of a V-PCC unit (for example, as shown by the vpcc_unit_type field in Table 2), which can be one of several values, including VPCC_OVD, VPCC_GVD and VPCC_AVD, VPCC_PDG, and VPCC_SPS, which may correspond to, for example, occupation, geometry, attribute, patch data group, and sequence parameter set data units, respectively. V-PCC units of some or all of these unit types may be used to reconstruct the point cloud. The V-PCC attribute unit header may specify attribute types and their indices. The V-PCC attribute unit header may allow for support of multiple instances of the same attribute type.As shown, vpcc_unit_type may indicate the type of V-PCC unit, vpcc_sequence_parameter_set_id may indicate the identifier of the V-PCC sequence parameter set, vpcc_attribute_index may indicate an index of the V-PCC attribute, vpcc_attribute_dimension_index may indicate an index of the dimension partition of the V-PCC attribute, sps_multiple_layer_streams_present_flag may indicate whether the sequence parameter set (SPS) is associated with multiple layers or views, vpcc_layer_index may indicate an index of one of the multiple layers, pcm_separate_video_data may indicate parameters associated with pulse code modulation (PCM) video data (e.g., in a separate video stream) and / or the encoding of the PCM data, and vpcc_reserved_zero_23bits or vpcc_reserved_zero_27bits may indicate the number of reserved zero bits.

[0063]

[0071] The occupancy payload, geometry, and / or attribute V-PCC unit may correspond to video data units that can be decoded by a video decoder (e.g., HEVC network abstraction layer (NAL) units) (e.g., defined by the corresponding occupancy, geometry, and attribute parameter set V-PCC unit). Table 3 shows exemplary V-PCC unit payload syntax.

[0064] [Table 1]

[0065] [Table 2]

[0066] [Table 3]

[0067]

[0072] In some cases (for example, when lossless coding is used in V-PCC), the encoder may generate a patch of missing points containing information about points that may be lost after reconstruction from the compressed V-PCC bitstream. Missing points are sometimes called missing pulse code modulation (PCM) points. PCM points can be coded directly, for example, without utilizing a patch projection process. The patch of missing points may allow the decoder to reconstruct (e.g., completely reconstruct) the original point cloud that may be provided as input to the V-PCC encoder. The patch containing information related to missing points may be packed into the same video (e.g., the same video stream as the stream carrying the other points) or into a separate video (e.g., a separate video stream from the stream carrying the other points).

[0068]

[0073] A patch data group (PDG) may be replaced by a patch NAL (PNAL) unit (e.g., an atlas NAL unit). A PNAL unit may be equivalent to a network extraction layer (NAL) unit used in a video stream. Each PNAL unit may include a header containing the unit type and / or additional information (e.g., layer identification). A PNAL unit may be defined in one or more formats. One or more formats may include a simple PNAL unit stream format and / or a sample stream format. In the sample stream format, an additional header may precede the PNAL unit. The additional header may indicate the size of the PNAL unit (e.g., the exact size).

[0069]

[0074] The International Organization for Standardization (ISO) Based Media File Format (ISOBMFF) can define a structured media-independent file format. ISOBMFF (e.g., an ISOBMFF container file) can contain structured and / or media data information (e.g., for the timed presentation of media content such as audio and video). An ISOBMFF container file can also include support for non-timed data (e.g., metadata at various levels within the file structure). The logical structure of the file can be that of a movie (e.g., it can mimic a movie) and may contain a set of time-parallel tracks. The time structure of the file means that tracks may contain a sequence of samples in time. This sequence of samples can be mapped to the timeline of the entire movie. ISOBMFF can be based on the concept of a box-structured file. A box-structured file can contain a set of boxes (e.g., atoms) which may each have a size and / or type (e.g., each box may be associated with a size and type). This type can be a 32-bit value and can be represented by four printable characters (also known as a four-character code (4CC)). Non-timed data can, for example, be included in a metadata box at the file level and / or attached to a movie box within a movie or to one of the streams (e.g., tracks) of timed data.

[0070]

[0075] An ISOBMFF container (e.g., an ISOBMFF container file) may contain a MovieBox ("moov"). A MovieBox may contain metadata for media streams (e.g., sequential media streams) present in the file. Metadata can be transmitted within the box hierarchy of the MovieBox (e.g., within a TrackBox ("trak")). A track may represent a media stream (e.g., a sequential media stream) present in the file. A media stream may contain a series of samples (e.g., sample entries), such as audio or video access units of the base media stream, and may be encapsulated within a MediaDataBox ("mdat") (e.g., which may be at the top level of the container). The metadata for each track may contain a list of sample description entries, each sample description entry providing the coding or encapsulation format used in the track and initialization data for processing this format. Each sample may be associated with one of the sample description entries of the track. Tools may be used to define an explicit timeline map for each track. For example, an edit list may define an explicit timeline map for each track. The edit list may be signaled by using an EditListBox (or similar entity) having the exemplary syntax shown in Table 4, where each entry defines a portion of the track timeline, for example by mapping a composite timeline or by indicating "empty" time or "empty" edits (e.g., some portion of the presentation timeline that does not map any media).

[0071] [Table 4]

[0072]

[0076] ISOBMFF can be used to handle situations where a file author (e.g., an encoding device) may indicate a specific action to be performed on a player or renderer. In the case of a video stream, the file author may indicate such an action by using a restricted video scheme track. If the video track is a restricted video scheme track (e.g., as defined in subsection 8.15 of the ISO / IEC 14496-12 standard), post-decoder requirements may be signaled on the track. A track may be converted to a restricted video scheme track by setting its sample entry code to the 4-character code (4CC) "resv" and adding a RestrictedSchemeInfoBox (or similar entity) to its sample description. One or more (e.g., all other) boxes may be left unmodified. The original sample entry type (based on the video codec used to encode the stream) may be stored in an OriginalFormatBox (or similar entity) within the RestrictedSchemeInfoBox. A RestrictedSchemeInfoBox may contain three boxes: OriginalFormatBox, SchemeTypeBox, and SchemeInformationBox. OriginalFormatBox may store the original sample entry type based on the video codec used to encode the component stream. The nature of the constraints (e.g., features) may be defined within SchemeTypeBox.

[0073]

[0077] Figure 3 shows an exemplary structure of an ISOBMFF V-PCC container 300. Based on this exemplary structure, a V-PCC ISOBMFF container may include one or more of the following: A V-PCC ISOBMFF container may include a V-PCC track 302. A V-PCC track 302 may include a sample carrying one or more sequence parameter sets and / or one or more non-video encoded information V-PCC unit payloads (e.g., V-PCC unit types VPCC_SPS and / or VPCC_PDG). A V-PCC track 302 may provide a track reference to other tracks containing a sample carrying one or more video compression V-PCC unit payloads (e.g., V-PCC unit types VPCC_GVD, VPCC_AVD, and / or VPCC_OVD). A V-PCC ISOBMFF container may contain one or more limited video scheme tracks 304 (e.g., payloads for V-PCC units of type VPCC_GVD) whose samples may contain NAL units of the video-coded base stream of geometry data. A V-PCC ISOBMFF container may contain one or more limited video scheme tracks 306 (e.g., payloads for V-PCC units of type VPCC_AVD) whose samples may contain NAL units of the video-coded base stream of attribute data. A V-PCC ISOBMFF container may contain one or more limited video scheme tracks 308 (e.g., payloads for V-PCC units of type VPCC_OVD) whose samples may contain NAL units of the video-coded base stream of occupied map data.

[0074]

[0078] A container format for point cloud data may be provided. Transmission of missing PCM point information may be supported by the container format for point cloud data. Signaling for V-PCC tile groups and / or spatial access may be provided. The sample format for V-PCC tracks may support PNAL units. Many file format structures that enable flexible access to various components, layers, and / or spatial regions within a V-PCC bitstream may be provided to support and / or provide signaling for PCM information. A track (e.g., a single track) may be used to store information for layers of V-PCC components (e.g., all layers) if, for example, the layers of V-PCC components constitute a single video stream. A sample grouping mechanism may be used to group samples belonging to each layer.

[0075]

[0079] If layers are stored in separate tracks, the track grouping tool can be used to signal that the separate tracks belong to layers that are part of the same V-PCC component. For example, a track group type (e.g., VPCCComponentGroupBox or a similar entity) can be defined by extending, for example, TrackGroupTypeBox. TrackGroupTypeBox may include a track_group_id field, which is the group identifier, and a track_group_type field, which stores a four-character code that identifies the group type. The pair of track_group_id and track_group_type can identify a track group in a container file. VPCCComponentGroupBox can be defined as follows: VPCCComponentGroupBox can be a box type of "vplg" and can be placed within a TrackGroupBox container. In some examples, VPCCComponentGroupBox can be optional (e.g., not mandatory). In some examples, there can be multiple VPCCComponentGroupBox boxes within a TrackGroupBox.

[0076]

[0080] Table 5 shows an example of the VPCCComponentGroupBox syntax.

[0077] [Table 5]

[0078]

[0081] Tracks belonging to the same component layer (e.g., all tracks) may have a VPCCComponentGroupBox within a TrackGroupBox. Each VPCCComponentGroupBox may have the same track_group_id value. The V-PCC media player can identify tracks belonging to the same V-PCC component by analyzing each track in the container and / or by identifying those that have VPCCComponentGroupBoxes with the same track_group_id value.

[0079]

[0082] To refer collectively to the base tracks belonging to the same component (e.g., all tracks), the track reference corresponding to the V-PCC component within the main V-PCC track may use the track_group_id of the component's track group (e.g., to identify one or more track groups associated with the component). For example, the TrackReferenceTypeBox corresponding to a component may have an entry in its track_IDs array that uses the track_group_id to identify the component's track group.

[0080]

[0083] In some examples, TrackGroupTypeBox may contain a flag field, and a bit (e.g., bit 0 of the field: bit 0 is the least significant bit) may be used to indicate the uniqueness of track_group_id. Tracks carrying geometry and / or attribute information of the same layer may be grouped. Grouped tracks carrying geometry and / or attribute information may enable a media player to have scalable access to V-PCC content. VPCCLayerGroupBox or similar entities may be defined and given the box type "vplg". VPCCLayerGroupBox may be placed inside a TrackGroupBox container. In some examples, VPCCLayerGroupBox may be optional (e.g., not mandatory). In some examples, multiple VPCCLayerGroupBox boxes may exist within a TrackGroupBox.

[0081]

[0084] Table 6 shows an example of the VPCCLyerGroupBox syntax.

[0082] [Table 6]

[0083]

[0085] As shown in the illustrative syntax, the VPCCLayerGroupBox field may contain one or more of the following fields: The layer_index field may indicate the index of the layer to which one or more tracks in the group belong. The absolute_coding_flag field may indicate whether the geometry tracks in this track group depend on geometry tracks in another layer. If absolute_coding_flag is set to 1, the track does not need to depend on another layer. If absolute_coding_flag is set to 0, the track may depend on another layer. The predictor_layer_index field may indicate the index of the layer to which the geometry tracks in this group depend.

[0084]

[0086] V-PCC component tracks may be provided. In some cases (for example, when one or more V-PCC stream components such as occupation, geometry, and / or attribute components are video-coded), one or more tracks carrying information related to the V-PCC stream components (e.g., any of the occupation, geometry, and / or attribute components) may be signaled as limited video scheme tracks. Limited video scheme tracks may not be for direct rendering. The scheme_type field in SchemeTypeBox may be set to 4CC with respect to the components of the V-PCC content (e.g., "pccv"). Data associated with V-PCC schemes may be stored in SchemeInformationBox. For example, data associated with a V-PCC scheme, which may be carried in SchemeInformationBox and defined as follows, may be signaled within VPCCComponentInfoBox (or a similar entity).

[0085]

[0087] Table 7 shows an example of VPCCComponentInfoBox syntax.

[0086] [Table 7]

[0087]

[0088] As shown in the illustrative semantics, a VPCCComponentInfoBox may contain one or more of the following fields: The component_type field may indicate the type of component. For example, a value of 0 for component_type may be reserved. A value of 1 for component_type may indicate an occupied map component. A value of 2 for component_type may indicate a geometry component. A value of 3 for component_type may indicate an attribute component. It should be noted that these numbers are provided herein as examples and other numbers may be used to indicate various component types. The is_pcm_flag field may indicate whether the information carried in the track is of a PCM point. If is_pcm_flag is set (for example to a value of 1), the track may carry PCM information of the component indicated by component_type. The all_layers_present_flag field may indicate whether the track carries information for all layers of the component. Encoded data for the component's layers (for example, all layers) may be present in the track if all_layers_present_flag is set (for example to a value of 1). Otherwise (for example, if all_layers_present_flag is not set or is set to a value of 0), the track may carry coded data from a single layer of a component. The layer_index field may indicate the component layer to which the data carried by the track belongs.

[0088]

[0089] A SchemeInformationBox may contain additional VPCCAttributeInfoBoxes (or similar entities) that provide additional descriptions of attribute components, for example, if the component track carries attribute information (for example, if component_type is set to 3). A VPCCAttributeInfoBox can be defined as shown in Table 8.

[0089] [Table 8]

[0090]

[0090] As shown in the exemplary semantics of Table 8, the VPCCAttributeInfoBox may include one or more of the following fields: The attr_index field may indicate an index of an attribute in the list of attributes. The attr_type field may indicate an attribute of an attribute type. The attr_dimensions field may indicate the number of dimensions of the attribute (e.g., total). The attr_first_dim_index field may indicate an index of the first attribute dimension carried by the track (e.g., a zero-based index).

[0091]

[0091] VPCCAttributeInfoBox may include a segmentation index. Table 9 shows another example of VPCCAttributeInfoBox syntax.

[0092] [Table 9]

[0093]

[0092] As shown in the exemplary syntax of Table 9, a VPCCAttributeInfoBox may contain one or more of the following fields: The attr_index field may indicate an index of the attribute in the list of attributes. The attr_type field may indicate the type of attribute. The attr_dimensions field may indicate the number of dimensions of the attribute (e.g., total). The attr_dim_partition_index field may represent an index of the dimensionality partitioning carried by the track (e.g., zero-based).

[0094]

[0093] In some cases, the VPCCComponentBox may carry a vpcc_unit_header()HLS struct of vpcc_unit_type corresponding to the component information carried by the truck (for example, it may carry it directly). If vpcc_unit_type is VPCC_AVD, the presence of a VPCCAttributeInfoBox in the SchemeInformationBox may be optional.

[0095]

[0094] Information regarding missing PCM points may include geometry data and / or attribute data. PCM point information may be packed into the video stream of the associated component and / or available as a separate video stream (e.g., one video stream per component). PCM point information may be carried in a separate track (e.g., one track of information relating to a component) if the PCM point information is available separately. A separate track may be signaled as a restricted video scheme track with the is_pcm_flag field in VPCCComponentBox set to 1 (e.g., as described herein). Each track carrying PCM point information may be included in a track group of the associated component. A track reference from the main track to the track_group_id of the V-PCC component may refer to tracks of PCM and / or non-PCM points (e.g., referencing them collectively).

[0096]

[0095] Geometry and attribute information associated with PCM can be grouped, for example, to enable easy identification and access to missing points. Track grouping can be defined, for example, using a VPCCPCMTrackGroupBox (or similar entity) as shown in Table 10 to identify tracks having PCM point information.

[0097] [Table 10]

[0098]

[0096] A track that carries information related to the PCM points of a V-PCC content may contain a VPCCPCMTrackGroupBox (or a similar entity) within a TrackGroupBox. Each VPCCPCMTrackGroupBox may have the same track_group_id value.

[0099]

[0097] A 4CC value (e.g., "pccp") may be defined in the reference_type field of the TrackReferenceTypeBox. The 4CC value may be used to signal a track reference that carries PCM point data (e.g., including geometry and / or attribute data).

[0100]

[0098] In some cases, there may be no constraints on the predictive structure used to encode the various components of the V-PCC bitstream. Thus, it may be possible to encode various components and / or (if the components are not in the same video stream, for example) various layers of the same component by an encoding configuration that would result in non-aligned intra-refresh periods across various component substreams. Such non-aligned intra-refresh periods across various component substreams can make random access difficult, as an intra-coded sample in a patch stream within the main V-PCC track at a given decoding time may not have a corresponding intra-coded sample in other component tracks at the same decoding time. Without additional information indicating the position of a sync sample in a component track relative to the main track, a media player may have to rely on scanning the component track of the nearest sync sample.

[0101]

[0099] Synchronization samples within one V-PCC component may not be properly aligned with synchronization samples within other components. Synchronization samples within a main track may have corresponding synchronization samples in other (e.g., all other) component tracks. For example, if the intra-refresh period of a patch sequence stream is once every 30 frames, a geometry component may have an intra-refresh period once every 60 frames and / or a texture attribute may have an intra-refresh period once every 30 frames. For example, intra-refresh frames may occur every 30 seconds in the main track and other (e.g., all other) components. Intra-refresh frames may have the same decoding time.

[0102]

[0100] Constraints on coded intra-random access point (IRAP) periods across components may be defined so that IRAP samples are aligned across tracks (e.g., to support random access). For example, an encoder may be constrained to generate a substream with synchronized samples aligned at regular intervals. This constraint may lead the decoder and / or client to assume that: IRAP samples are available in one or more other components (e.g., all other) at the same time that an IRAP sample is detected in one component track (e.g., any component). The IRAP of each component may represent the IRAP of the VPCC bitstream. Figure 4 shows an example of a constraint on the alignment of component IRAP samples.

[0103]

[0101] The constraints described herein may eliminate the need for additional information to signal the matching of synchronization samples across the entire track. Once a synchronization sample reaches the main track, a corresponding synchronization sample with the same decoding time may be found in other (e.g., all other) component tracks.

[0104]

[0102] In some examples, the IRAP periods and main patch sequence tracks of various components may be selected so that time-aligned (e.g., synchronized) intra-samples appear at regular intervals. When time-aligned (e.g., synchronized) intra-samples appear at regular intervals, a V-PCC media player accessing the synchronized samples in the main V-PCC track may find corresponding synchronized samples with the same decoding time in other component tracks. Each component may have different IRAP periods. The IRAP period of the main V-PCC track may be the least common multiple of the IRAP periods of the other (e.g., all) component tracks. The IRAP of the main V-PCC track may represent the IRAP of the V-PCC bitstream. Figure 5 shows an example of using the least common multiple of IRAP periods to represent V-PCC IRAP.

[0105]

[0103] In some cases, there may be no constraints on the IRAP duration of the V-PCC components. The IRAP of the main V-PCC track may represent the IRAP of the V-PCC bitstream. With respect to other components, the nearest IRAP may be placed given the decoding and / or presentation time of the IRAP in the main V-PCC.

[0106]

[0104] The V-PCC high-level syntax (HLS) may support tile groups. In video encoding standards (e.g., HEVC), a 2D frame may be divided into a grid of tiles. One or more tile groups may correspond to a rectangular area within a 2D frame containing many tiles. Motion-constrained tile sets (MCTS) may be decoded (e.g., independently) and enable the extraction of specific areas within a frame. In V-PCC, patches corresponding to points belonging to a region in space (e.g., a 3D region or cube) may be packed into one or more MCTS. Tile groups and MCTS may be used interchangeably as described herein.

[0107]

[0105] Figure 6 shows an example of a V-PCC container structure 600 that can be used to enable spatial access to a specific region in 3D space. As shown, the 3D space 602 of the point cloud (e.g., the bounding box corresponding to the 3D space) can be divided into a 3D cubic grid (e.g., cubes 602a, 602b, 602c, etc.) representing multiple regions and / or objects in 3D space. Points belonging to each of the regions and / or objects in 3D space can be clustered, and bounding boxes can be used to represent those regions or objects. Points belonging to different parts of the same object can be grouped together, and each of those parts can be represented by its own bounding box.

[0108]

[0106] Patches resulting from the projection of points within each bounding box of the resulting bounding box can be packed together into one or more tile groups within a 2D frame of multiple V-PCC component streams or tracks (e.g., occupancy, geometry, and / or attribute streams or tracks). Patches can be encoded using an encoding configuration that generates tile groups (e.g., independently decodeable tile groups). These tile groups can be carried in separate tracks within an ISOBMFF container, and therefore the term “tile group” can be used interchangeably with “track group” herein (e.g., a tile group can be an instance of a track group). Carrying tile groups in separate tracks can enable a decoder (e.g., a media player) to access and / or download tracks that carry information related to a particular region or object in 3D space. For example, if tile groups are carried in separate tracks of a V-PCC bitstream, a media player can access and / or download only the tracks related to a particular region in 3D space when decoding a particular region (e.g., when rendering a visual representation of the region).

[0109]

[0107] Tracks that have corresponding tile groups across the entire V-PCC component (for example, carrying information about points in bounding boxes representing regions or objects) can be grouped together using the Track Grouping tool. A TrackGroupBox ("trgr") or similar entity can be added to the TrackBox of each of these tracks, and the track grouping type of the V-PCC tile group can be defined by extending TrackGroupTypeBox as shown below (for example by using the track_group_id or tile_group_id fields).

[0110]

[0108] Table 11 shows an example of the VPCCTileGroupBox syntax.

[0111] [Table 11]

[0112]

[0109] As shown in the illustrative semantics of Table 11, a VPCCTileGroupBox may contain a tile_group_id field (or a similar field) that identifies a V-PCC tile group (for example, as an identifier for a V-PCC tile group). In some examples, tile_group_id may correspond to (or be identical to) a tile group address (for example, a field such as ptgh_address that may be included in the tile group header of a V-PCC bitstream). Tracks belonging to the same point cloud tile group may have the same value for track_group_id with track_group_type "vptg". The track_group_id of a track from one point cloud tile group may be different from the track_group_id of a track from another point cloud tile group. For example, as shown in Figure 6, a first tile group corresponding to 3D region 602a may have a track group ID of 1, and a second tile group corresponding to 3D region 602b may have a track group ID of 2. Therefore, a track_group_id in a TrackGroupTypeBox with a track_group_type equal to "vptg" (or a similar 4CC value) can be used as an identifier for a point cloud tile group in an ISOBMFF container file.

[0113]

[0110] For example, sample grouping may be used to signal which sample belongs to which V-PCC tile group. For example, sample grouping may be used when information relating to two or more V-PCC tile groups for one V-PCC component is carried in a track (for example, for a set of V-PCC tile groups, there is a set of sample groups in the track, and each group of samples is associated with its respective V-PCC tile group). Sample group entries may be defined (for example, as shown in Table 12 below), where the semantics (e.g., definition) of tile_group_id may be the same as that of tile_group_id defined in VPCCTileGroupBox as described herein. The group type may be "vpge" or a similar 4CC value. The container may be SampleGroupDescriptionBox ("sgpd") or a similar entity. VPCCTileGroupBox may not be mandatory (e.g., it may be optional), and each track may have multiple VPCCTileGroupBoxes (e.g., may be associated). Table 12 shows an example syntax for VPCCTileGroupEntry.

[0114] [Table 12]

[0115]

[0111] In some examples, subtracks carrying one or more V-PCC tile groups may be defined within a component track. One or more V-PCC tile groups may be defined by using a SubTrackSampleGroupBox (or similar entity) and enumerating VPCCTileGroupEntry instances (or similar entities) corresponding to the V-PCC tile groups carried in each subtrack within the corresponding SubTrackSampleGroupBox (for example, by referring to their group_description_index). One or more V-PCC tile groups may be defined by defining a V-PCC-specific VPCCTileGroupSubTrackBox (for example, as shown in Table 13). The box type may be set to "vpst" or a similar 4CC value. The container may be a SubTrackDefinitionBox ("strd") or a similar entity. The VPCCTileGroupSubTrackBox may not be mandatory (for example, it may be optional), and each track may have multiple VPCCTileGroupSubTrackBoxes.

[0116] [Table 13]

[0117]

[0112] The union (e.g., set) of tile_group_ids within a VPCCTileGroupSubTrackBox can describe a subtrack defined by a box (e.g., it can describe them collectively). The semantics of a VPCCTileGroupSubTrackBox may include one or more of the following fields: The item_count field may represent the count of the number of tile groups enumerated in the VPCCTileGroupSubTrackBox. The tile_group_id field may represent the identifier of the V-PCC tile group included in this subtrack. The tile_group_id field within a VPCCTileGroupSubTrackBox may correspond to the tile_group_id defined in a VPCCTileGroupEntry.

[0118]

[0113] For example, a mapping may be provided between regions or objects in 3D space (e.g., each of the 3D bounding boxes) and their respective tile groups to enable a client (e.g., a media player or decoder) to identify which track to access / download in order to render a region in 3D space (e.g., represented by a bounding box). It should be noted that while the position of tile groups in a 2D frame may not change, the position and possibly size (e.g., dimensions) of bounding boxes (e.g., regions) in 3D space may change over time due to the movement of objects represented by points within the bounding box. 3D regions within a point cloud are defined using the exemplary 3D region structure shown in Table 14.

[0119] [Table 14]

[0120]

[0114] As shown in the illustrative semantics of Table 14, a 3DRegionStuct may contain one or more of the following fields: The region_id field may represent a unique identifier of the 3D region. The region_x field may represent the x-coordinate of a reference point associated with the 3D region (e.g., the bounding box associated with the region). The region_y field may represent the y-coordinate of the reference point. The region_z field may represent the z-coordinate of the reference point. The region_width field may indicate the length of the 3D region (e.g., the bounding box associated with the region) along the x-axis. The region_height field may indicate the length of the 3D region (e.g., the bounding box associated with the region) along the y-axis. The region_depth field may indicate the length of the 3D region (e.g., the bounding box associated with the region) along the z-axis. The dimensions_included_flag field may indicate whether the dimensions of the 3D region (e.g., the bounding box associated with the region) are signaled in the same instance of the struct. For example, if dimensions_included_flag has a value of 0, this may indicate that the dimension is not transmitted, or that the dimension of the same region may have already been transmitted (for example, a previous instance of VPCC3DRegionStruct with the same region_id transmitted the dimension). If dimensions_included_flag has a value of 1, this may indicate that the dimension is transmitted.

[0121]

[0115] A 3D region or object within a point cloud can be associated with one or more point cloud tile groups (e.g., instances of track groups) by using VPCCRegionToTileGroupBox or a similar entity. Table 15 shows exemplary VPCCRegionToTileGroupBox syntax.

[0122] [Table 15]

[0123]

[0116] As shown in the illustrative semantics of Table 15, VPCCRegionToTileGroupBox may represent a mapping relationship between a region (or object) in 3D space and one or more tile groups (e.g., track groups). VPCCRegionToTileGroupBox may include one or more of the following fields: The num_regions field may indicate the number of 3D regions in the point cloud associated with the 3D space. The region_id field may identify a 3D region (e.g., may include an identifier). The num_tile_groups field may indicate the number of V-PCC tile groups associated with the 3D region. The tile_group_id field may identify a V-PCC tile group. Thus, VPCCRegionToTileGroupBox may link one or more tiles to a 3D region via at least the tile_group_id field and the region_id field.

[0124]

[0117] The VPCCRegionToTileGroupBox may be signaled in a sample entry of the main V-PCC track 604 or in a sample entry of a separate time-limited metadata track 606 associated with the main V-PCC track, as shown in Figure 6. The time-limited metadata track 606 (which may be separate from the main V-PCC track, for example) may be contained within an ISOBMFF container and may be used to update one or more properties (e.g., location and / or dimensions) of a predefined 3D region of a point cloud over time. This time-limited metadata track 606 may contain a defined sample entry (e.g., VPCC3DRegionSampleEntry) having the 4CC (or similar 4CC value) of "vp3r", and the defined sample entry may extend MetadataSampleEntry or a similar entity as shown by the exemplary syntax (e.g., VPCC3DRegionInfoBox or a similar entity) in Table 16.

[0125] [Table 16]

[0126]

[0118] As shown in the illustrative semantics of Table 16, the VPCC3DRegionInfoBox may include a num_regions field indicating the total number of 3D regions in 3D space. The time-limited metadata track 606 may be linked to the main V-PCC track 604 by using, for example, the 4CC (or similar 4CC value) of "cdsc" as a track reference. Each (e.g.) sample within this time-limited metadata track may define a 3D region by using, for example, the illustrative syntax shown in Table 17 below. The VPCC3DRegionSample structure (e.g., or similar entities) may be extended with the derived track format.

[0127] [Table 17]

[0128]

[0119] As shown in the illustrative semantics of Table 17, VPCC3DRegionSample may include a num_regions field which may indicate the number of 3D regions being signaled within the sample. The number of 3D regions being signaled within the sample may be equal to or different from the total number of available regions. For example, the number of 3D regions being signaled within the sample may indicate 3D regions whose properties (e.g., location and / or dimensions) are updated within the sample.

[0129]

[0120] Patch information may be carried in a V-PCC track. The VPCCDecoderConfigurationRecord (or similar entity) and sample format syntax of the V-PCC track may be formatted to support the transmission of a patch information substream structured, for example, as a series of Patch Network Extraction Layer (PNAL) units. The VPCCDecoderConfigurationRecord may provide configuration information to the decoder (for example, at the beginning of the decoding process). The VPCCDecoderConfigurationRecord may contain one or more parameter sets and / or one or more supplemental enhancement information (SEI) messages. The VPCCDecoderConfigurationRecord may contain a lengthSizeMinusOne field. Exemplary VPCCDecoderConfigurationRecord syntax may be shown in Table 18 below.

[0130] [Table 18]

[0131]

[0121] As shown in the illustrative semantics of Table 18, the VPCCDecoderConfigurationRecord may include a configurationVersion field indicating the current version of the configuration record. In some examples, non-conforming changes to the decoder configuration record may be indicated by a change in the configuration version number. The decoder may be configured not to attempt to decode the applicable configuration record or stream if the configuration version number is not recognized. The VPCCDecoderConfigurationRecord may include a lengthSizeMinusOne field, where the value of lengthSizeMinusOne plus 1 may indicate the length (e.g., in bytes) of the PNALUnitLength field in the V-PCC sample (e.g., in the stream to which this configuration record applies). For example, a PNALUnitLength field length of 1 byte may be indicated by a lengthSizeMinusOne value of 0. The value of the lengthSizeMinusOne field may be 0, 1, or 3, which may correspond to lengths (e.g., PNALUnitLength) encoded by 1, 2, or 4 bytes, respectively.

[0132]

[0122] In some examples, the decoder configuration record may include one or more setup unit arrays, such as the first setup unit array of the V-PCC parameter set (e.g., the Vsequence parameter set, and a second setup unit array of other setup units in the patch information substream). Table 19 below shows examples of one or more setup unit arrays.

[0133] [Table 19]

[0134]

[0123] As shown in the semantics example in Table 19, a VPCCDecoderConfigurationRecord (or similar entity) may contain one or more of the following fields: The configurationVersion field (or a field with a similar name) may indicate the current version of the configuration record. In some examples, non-conforming changes to the decoder configuration record may be indicated by a change in the configuration version number. The decoder may be configured not to attempt to decode the applicable configuration record or stream if the configuration version number is not recognized. The numOfSequenceParameterSets field (or a field with a similar name) may indicate the number of signed (e.g., defined) V-PCC parameter sets (e.g., arrays) in the decoder configuration record (e.g., for a stream to which the decoder configuration record applies). The numOfSetupUnitArrays field may indicate the number of arrays of signed (e.g., defined) PNAL units of a specified type (e.g., indicated by PNAL_unit_type) in the decoder configuration record (e.g., for a stream to which the decoder configuration record applies). The array_completeness field may indicate whether all PNAL units are included in the array. For example, if the array_completeness field is equal to 1, this may indicate that the subsequent array contains PNAL units of a given type (e.g., all PNAL units) (e.g., nothing in the stream). If the array_completeness field is equal to 0, this may indicate that additional PNAL units of the indicated type may exist in the stream. The default value and / or allowed value of array_completeness may be constrained by the sample entry name or sample entry type of the corresponding sample entry. For example, VPCCDecoderConfigurationRecord may be used in various sample entries.The container for VPCCDecoderConfigurationRecord can be a VPCCDecoderConfigurationBox (or a similar entity), which may be a box contained within a VPCCSampleEntry. VPCCSampleEntry can be of different types, and the type of sample entry may set constraints on the allowed and / or default values ​​of the array_completeness field within the encapsulated VPCCDecoderConfigurationRecord.

[0135]

[0124] A VPCCDecoderConfigurationRecord (or similar entity) may include a PNAL_unit_type field indicating the type of PNAL unit in a subsequent array (for example, all PNAL units in the array may be of the indicated type). The PNAL_unit_type field may have one of the following values ​​(for example, it may be restricted to take) that indicates a PUP_PSPS, PUP_PREFIX_SEI, or PUP_SUFFIX_SEI PNAL unit. A VPCCDecoderConfigurationRecord (or similar entity) may include a numPNALUnits field indicating the number of PNAL units of the indicated type included in the configuration record (for example, of the stream to which this configuration record applies). A Supplemental Enhancement Information (SEI) array may include a Declaration SEI message (for example, it may include only this). A Declaration SEI message may include an SEI message that indicates information about the stream in general. For example, User Data SEI may be a Declaration SEI message.

[0136]

[0125] A VPCCDecoderConfigurationRecord (or similar entity) may contain a pnalUnitLength field (e.g., in bytes) indicating the length of the PNAL unit. A VPCCDecoderConfigurationRecord (or similar entity) may contain a pnalUnit field which may be used to hold a PUP_PSPS or declared SEI PNAL unit.

[0137]

[0126] Based on the exemplary VPCCDecoderConfigurationRecord syntax shown herein, the sample format of a sample within a V-PCC track (represented, for example, as VPCCSample) is shown in Table 20 below.

[0138] [Table 20]

[0139]

[0127] As shown in the illustrative semantics of Table 20, the VPCCDecoderConfigurationRecord field may indicate the decoder configuration record in the corresponding V-PCC sample entry. The PNALUnitLength field may indicate the size of the PNAL unit (e.g., measured in bytes). In some examples, the PNALUnitLength field may include the sizes of both the PNAL unit header and the PNAL unit payload. In some examples, the PNALUnitLength field may not include the size of the PNALUnitLength field itself. Furthermore, the PNALUnit field may be included to represent a PNAL unit (e.g., a single atlas NAL unit).

[0140]

[0128] In some examples, patch information within a V-PCC track sample may be formatted based on a patch information sample stream (e.g., an atlas sample stream). A VPCCDecoderConfigurationRecord (or similar entity) may include a lengthSizeMinusOne field. Table 21 shows another exemplary VPCCDecoderConfigurationRecord syntax.

[0141] [Table 21]

[0142]

[0129] Fields (e.g., variables) in the exemplary syntax of Table 21 can be defined in the same way as those in Table 19. For example, the value of lengthSizeMinusOne plus 1 may indicate the length (e.g., in bytes) of the PNALUnitLength field (e.g., in the V-PCC sample in the stream to which this configuration record applies). Thus, the size of one byte in the PNALUnitLength field may be indicated by the lengthSizeMinusOne field, which has a value of 0. In the exemplary syntax of Table 21, the lengthSizeMinusOne field may be defined as an unsigned int(3), and therefore the value of the lengthSizeMinusOne field can be in the range of 0 to 7.

[0143]

[0130] Although features and elements have been described above in particular in combination, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. In addition, the methods described herein may be implemented in computer programs, software, or firmware incorporated into computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, optical media such as CD-ROM disks, and digital multipurpose disks (DVDs). Processors related to software may be used to implement radio frequency transceivers used in WTRUs, UEs, terminals, base stations, RNCs, or any host computer.

Claims

1. A video decoding device configured to process video data associated with a three-dimensional (3D) space, Receive media container files, The medial container file is parsed in order to determine the region identifier (ID) of the 3D region in the 3D space and the respective track group IDs of one or more track groups associated with the 3D space. Based on the determination that each of the one or more track groups' track group IDs is linked to the area ID of the 3D area, it is determined that the one or more track groups are associated with the 3D area. A video decoding device including a processor configured to decode video tracks belonging to one or more track groups in order to render a visual representation of the 3D region of the 3D space.

2. The video decoding apparatus according to claim 1, wherein the one or more track groups share a common track group type, and based further on the determination that the one or more track groups share the common track group type, it is determined that the one or more track groups are associated with the 3D region.

3. The video decoding apparatus according to claim 1, wherein the media container file includes a structure that defines the number of regions associated with the 3D space and the number of track groups associated with each of the regions, and the processor is configured to determine, based on the information contained in the structure, that each of the track group IDs of the one or more track groups is linked to the region ID of the 3D region.

4. The video decoding apparatus according to claim 1, wherein the medial container file includes time-limited metadata indicating an update to at least one characteristic of the 3D region.

5. The video decoding apparatus according to claim 4, wherein the processor is configured to determine, based on the time-limited metadata, that each of the track group IDs of the one or more track groups is linked to the area ID of the 3D area.

6. The video decoding apparatus according to claim 4, wherein the 3D space includes a plurality of regions, and the time-limited metadata includes information associated with an updated subset of the regions.

7. The video decoding apparatus according to claim 1, wherein the processor is further configured to determine the 3D region and reference points associated with the dimensions of the 3D region based on the medial container file.

8. The video decoding apparatus according to claim 1, wherein the video tracks belonging to the one or more track groups correspond to one or more tiles in a two-dimensional (2D) frame.

9. The video decoding apparatus according to claim 1, wherein the video track comprises one or more sample entries, each of which comprises an indication of the length of a data field indicating the network extraction layer (NAL) unit size.

10. The video decoding apparatus according to claim 9, wherein each of the one or more sample entries further includes an indication of the number of V-PCC parameter sets associated with the sample entry or the number of arrays of Atlas NAL units associated with the sample entry.

11. A method for decoding video data associated with a three-dimensional (3D) space, Receiving media container files and, The medial container file is parsed in order to determine the region identifier (ID) of the 3D region in the 3D space and the respective track group IDs of one or more track groups associated with the 3D space. Based on the determination that each of the one or more track groups' track group IDs is linked to the area ID of the 3D area, it is determined that the one or more track groups are associated with the 3D area, In order to render a visual representation of the 3D region of the 3D space, the video tracks belonging to one or more track groups are decoded. A method that includes this.

12. The method according to claim 11, wherein the one or more track groups share a common track group type, and based further on the determination that the one or more track groups share the common track group type, it is determined that the one or more track groups are associated with the 3D region.

13. The method according to claim 11, wherein the media container file includes a structure that defines the number of regions associated with the 3D space and the number of track groups associated with each of the regions, and based on the information contained in the structure, it is determined that each of the track group IDs of the one or more track groups is linked to the region ID of the 3D region.

14. The method according to claim 11, wherein the media container file includes time-limited metadata indicating an update to at least one characteristic of the 3D region.

15. The method according to claim 14, wherein the 3D space includes a plurality of regions, and the time-limited metadata includes information associated with an updated subset of the regions.

16. The method according to claim 11, wherein the video track belonging to the one or more track groups corresponds to one or more tiles in a two-dimensional (2D) frame.

17. The method according to claim 11, wherein the video track comprises one or more sample entries, each of which comprises an indication of the length of a data field indicating the network extraction layer (NAL) unit size.

18. The method according to claim 17, wherein each of the one or more sample entries further includes an indication of the number of V-PCC parameter sets associated with the sample entry or the number of arrays of Atlas NAL units associated with the sample entry.

19. A video encoding device configured to encode and transmit information associated with a three-dimensional (3D) space, The aforementioned 3D space is divided into one or more 3D regions, each of which is assigned its own region identifier (ID). The video data associated with at least one of the one or more 3D regions is encoded into multiple video-based point cloud compression (V-PCC) component tracks. The aforementioned multiple V-PCC component tracks are organized into a track group, and a track group ID is assigned to the track group. The fact that the 3D space includes the one or more 3D regions and that the track group is linked to at least one of the one or more 3D regions is indicated in a separate V-PCC track, and the track group is linked to at least one of the one or more 3D regions via the track group ID of the track group and the region ID of at least one of the one or more 3D regions. A video encoding device including a processor configured to transmit a media container file to a receiving device, wherein the media container file includes the one or more V-PCC component tracks and the separate V-PCC tracks.

20. The video encoding apparatus according to claim 19, wherein the media container file further includes time-limited metadata indicating updates associated with the one or more 3D regions.