Partial access support in isobmff containers for video-based point cloud streams

By dividing 3D space into spatial regions and mapping them to V-PCC tile sets within an ISOBMFF container, the solution addresses the limitations of current V-PCC signaling, enabling efficient and flexible partial access to video-based point cloud streams.

JP2025170343APending Publication Date: 2025-11-18INTERDIGITAL PATENT HOLDINGS INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025138461
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-10-05
Filing Date
2025-08-21
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Current video-based point cloud compression (V-PCC) signaling is insufficient for partial access of V-PCC sequences, hindering flexible access to different parts of a coded point cloud sequence.

Method used

The video encoding device divides a 3D space into spatial regions, maps them to V-PCC tile sets, and sends associated tracks and metadata in a timed metadata V-PCC bitstream within an ISOBMFF container, allowing independent decoding and updating of these regions.

Benefits of technology

Enables flexible partial access and efficient compression of video-based point cloud streams by allowing independent decoding and updating of spatial regions within the ISOBMFF container.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170343000001_ABST
    Figure 2025170343000001_ABST
Patent Text Reader

Abstract

To provide systems, devices, and methods for partial access support in ISOBMFF containers for video-based point cloud streams.SOLUTION: A video encoding device may partition a 3D space into a first spatial region and a second spatial region. The video encoding device may map the first spatial region to a first set of V-PCC tiles and the second spatial region to a second set of V-PCC tiles. The video encoding device may determine a first track to carry first mapping information associated with the first spatial region that is mapped to the first set of V-PCC tiles. The video encoding device may determine a second track to carry second mapping information associated with the second spatial region that is mapped to the second set of V-PCC files. The video encoding device may send, in a timed-metadata V-PCC bit stream, the first track and the second track.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 009,931, filed April 14, 2020, U.S. Provisional Patent Application No. 63 / 042,892, filed June 23, 2020, U.S. Provisional Patent Application No. 63 / 062,983, filed August 7, 2020, and U.S. Provisional Patent Application No. 63 / 087,425, filed October 5, 2020, the disclosures of which are incorporated herein by reference in their entireties. [Background technology]

[0002] A point cloud may include a set of points represented in 3D space using coordinates that indicate the location and attributes of each point. Reconstructing objects and scenes based on a point cloud may require processing millions of points. Efficient compression may be essential for storing and transmitting point cloud data.

[0003] A video-based point cloud compression (V-PCC) bitstream may include a sequence of V-PCC units. Each V-PCC unit may include a V-PCC header and a V-PCC payload. The V-PCC header may describe the V-PCC unit type, and the V-PCC payload may provide data associated with the V-PCC unit type. The sequence of V-PCC units may be signaled to a video decoder in the V-PCC bitstream. Current V-PCC signaling may not be sufficient for certain types of access (e.g., partial access) of a V-PCC sequence. Summary of the Invention

[0004] Systems, devices, and methods are described herein for partial access support in International Organization for Standardization Base Media File Format (ISOBMFF) containers for video-based point cloud streams. The file format structure may allow flexible partial access to different parts of a coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0005] The video encoding device may divide a 3D space into a first spatial region and a second spatial region. The video encoding device may map the first spatial region to a first video-based point cloud compression (V-PCC) tile set and the second spatial region to a second V-PCC tile set. Each of the first V-PCC tile set and the second V-PCC tile set may be associated with an atlas frame. Each of the first V-PCC tile set and the second V-PCC tile set may be independently decodable. Mapping the first spatial region to the first V-PCC tile set and the second spatial region to the second V-PCC tile set may be based on tile identification and / or track identification. The first V-PCC tile set may be associated with a first patch set, and the second V-PCC tile set may be associated with a second patch set. The video encoding device may determine a first track carrying first mapping information associated with a first spatial region mapped to a first V-PCC tile set. The video encoding device may determine a second track carrying second mapping information associated with a second spatial region mapped to a second V-PCC tile set. The video encoding device may send the first track and the second track in a timed metadata V-PCC bitstream. The first track and the second track may be sent within a media container file.

[0006] The video encoding device may determine an update dimension flag. The update dimension flag may indicate an update to one or more dimensions of the first spatial region or an update to one or more dimensions of the second spatial region. The video encoding device may send the update dimension flag in a timed metadata V-PCC bitstream.

[0007] The first spatial region may be associated with a first object. The second spatial region may be associated with a second object. The video encoding device may determine one or more object flags. The video encoding device may send the object flags in a timed metadata V-PCC bitstream. The video encoding device may determine an object dependent flag indicating that a first object associated with the first spatial region depends on a second object associated with the second spatial region and may send the object dependent flag in the timed metadata V-PCC bitstream. The video encoding device may determine an update object flag indicating an update to the first object associated with the first spatial region or an update to the second spatial region related to the second object and may send the update object flag in the timed metadata V-PCC bitstream. [Brief explanation of the drawings]

[0008] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 1C] 1A is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 1D]1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 2] FIG. 1 illustrates an embodiment of a block-based video encoder. [Figure 3] FIG. 1 illustrates an embodiment of a video decoder. [Figure 4] FIG. 1 illustrates an example system in which various aspects and embodiments may be implemented. [Figure 5] FIG. 2 illustrates an exemplary interface between a server and a client. [Figure 6] FIG. 2 illustrates an exemplary interface between a server and a client. [Figure 7] FIG. 1 illustrates an example of requesting content by a client (e.g., a head-mounted display (HMD)). [Figure 8] FIG. 1 illustrates an example of a video-based point cloud compression (V-PCC) bitstream structure as a sequence of V-PCC units. [Figure 9] FIG. 10 illustrates an example of tile and tile group division of an atlas frame. [Figure 10] FIG. 1 illustrates an exemplary structure of a multi-track ISOBMFF V-PCC container. [Figure 11] 1 illustrates an example of tile mapping of an atlas frame associated with a three dimensional (3D) space. DETAILED DESCRIPTION OF THE INVENTION

[0009] A detailed description of illustrative embodiments will now be described with reference to various figures. While the description provides detailed examples of possible implementations, it should be noted that the details are intended to be illustrative and in no way limit the scope of the present application.

[0010] 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. Communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. Communication system 100 may enable multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0011] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which may be referred to as a “station” and / or “STA,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain contexts), consumer electronics devices, devices operating in commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.

[0012] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNodeB, a Home Node B, a Home eNodeB, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0013] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple output (MIMO) technology and may utilize multiple transceivers for each sector of the cell, for example, using beamforming to transmit and / or receive signals in desired spatial directions.

[0014] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0015] More specifically, as noted above, the communications system 100 may be a multiple-access system and may use one or more channel access schemes, such as, for example, CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114 a and the WTRUs 102 a, 102 b, 102 c in the RAN 104 / 113 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA may include communications protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​Uplink Packet Access (HSUPA).

[0016] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).

[0017] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may establish the air interface 116 using New Radio (NR).

[0018] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using dual connectivity (DC) principles. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0019] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity, WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access, WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.

[0020] 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area such as a location such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may establish a picocell or a femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in FIG. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 through the CN 106 / 115.

[0021] The RAN 104 / 113 may communicate with the CN 106 / 115, which may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. 1A, it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 may communicate, directly or indirectly, with other RANs employing the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also communicate with another RAN (not shown) employing GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0022] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a public switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as the transmission control protocol (TCP), user datagram protocol (UDP), and / or the internet protocol (IP) of the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may employ the same RAT as the RAN 104 / 113 or a different RAT.

[0023] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that may use a cellular-based wireless technology and a base station 114b that may use an IEEE 802 wireless technology.

[0024] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.

[0025] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0026] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0027] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0028] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0029] The processor 118 of the WTRU 102 may be coupled to and may receive user-entered data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).

[0030] The processor 118 may receive power from the power source 134, but may be configured to distribute and / or control the power to other components in the WTRU 102. The power source 134 may be any suitable device for providing power to the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0031] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 may obtain location information by way of any suitable location-determination method while remaining consistent with an embodiment.

[0032] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripheral device 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0033] The WTRU 102 may include a full-duplex radio where transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and or substantially eliminate self-interference through hardware (e.g., chokes) or processor-based signal processing (e.g., via a separate processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmission and reception of either some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)).

[0034] 1C is a system diagram illustrating the RAN 104 and the CN 106 according to one embodiment. As mentioned above, the RAN 104 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 using E-UTRA radio technology. The RAN 104 may also communicate with the CN 106.

[0035] The RAN 104 may include eNode-Bs 160a, 160b, and 160c, although it will be understood that the RAN 104 may include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In an embodiment, the eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Thus, the eNode-B 160a may, for example, use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.

[0036] Each of the eNode-Bs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling, etc. in the UL and / or DL. As shown in FIG. 1C, the eNode-Bs 160a, 160b, 160c may communicate with each other via an X2 interface.

[0037] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. Although each of the foregoing elements is shown as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0038] The MME 162 may be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and may function as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies such as GSM and / or WCDMA.

[0039] The SGW 164 may be connected to each of the eNode-Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to and from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions, such as anchoring the user plane during inter-eNode-B handovers, triggering paging when DL data is available to the WTRUs 102a, 102b, 102c, and managing and storing the context of the WTRUs 102a, 102b, 102c.

[0040] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0041] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional landline communications devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. Furthermore, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0042] Although the WTRU is depicted in FIGS. 1A-1D as a wireless terminal, it is contemplated that in certain representative embodiments, such a terminal may use a wired communication interface (e.g., temporarily or permanently) with the communication network.

[0043] In a representative embodiment, the other network 112 may be a WLAN.

[0044] A WLAN in infrastructure Basic Service Set (BSS) mode may have an access point (AP) of the BSS and one or more stations (STAs) associated with the AP. The AP may have access or interface to a Distribution System (DS) or another type of wired / wireless network that carries traffic into and / or out of the BSS. Traffic originating from outside the BSS to a STA may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP and transmitted to the respective destination. Traffic between STAs within the BSS may be transmitted, for example, through the AP; the source STA may send traffic to the AP, which may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted between a source STA and a destination STA (e.g., directly between them) in a direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all of the STAs) may communicate directly with each other. The IBSS mode of communication may be referred to herein as an "ad hoc" communication mode.

[0045] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a width that is dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, for example, in an 802.11 system, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) with collision avoidance may be implemented. With CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit at any given time in a given BSS.

[0046] High Throughput (HT) STAs may use 40 MHz wide channels for communication, which may be formed, for example, through a combination of a primary 20 MHz channel and adjacent or non-adjacent 20 MHz channels.

[0047] A Very High Throughput (VHT) STA may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. The 40 MHz and / or 80 MHz wide channels may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining eight contiguous 20 MHz channels or by combining two non-contiguous 80 MHz channels, which may be referred to as an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data may pass through a segment parser that may split the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing may be performed separately on each stream. The streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration may be reversed and the combined data may be transmitted to the Medium Access Control (MAC).

[0048] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah may support meter-type control / machine-type communications, such as MTC devices within macro coverage areas. MTC devices may have specific capabilities, including, for example, support for (e.g., only for) specific and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).

[0049] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as a primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be configured and / or limited by the STAs among all STAs operating in the BSS that support the minimum bandwidth operating mode. In an 802.11ah example, the primary channel can be 1 MHz wide for STAs (e.g., MTC-type devices) that support (e.g., only) the 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) configuration can depend on the condition of the primary channel. For example, if the primary channel is busy due to a STA (that only supports 1 MHz mode of operation) transmitting to the AP, the entire available frequency band may be considered busy, even though most of the frequency band may remain idle and be available for use.

[0050] In the United States, the available frequency band that can be used by 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total bandwidth available for 802.11ah is 6MHz to 26MHz depending on the country code.

[0051] 1D is a system diagram illustrating the RAN 113 and the CN 115 according to one embodiment. As mentioned above, the RAN 113 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 using NR radio technology. The RAN 113 may also communicate with the CN 115.

[0052] The RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that the RAN 113 may include any number of gNBs while remaining consistent with the embodiments. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a, 180b may utilize beamforming to transmit and / or receive signals to the gNBs 180a, 180b, and 180c. Thus, the gNB 180a may, for example, transmit wireless signals to and / or receive wireless signals from the WTRU 102a using multiple antennas. In one embodiment, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on an unlicensed spectrum, and the remaining component carriers may be on a licensed spectrum. In one embodiment, the gNBs 180a, 180b, and 180c may implement coordinated multi-point (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).

[0053] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., including different numbers of OFDM symbols and / or lasting different lengths of absolute time).

[0054] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNode-Bs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate with and connect to a gNB 180a, 180b, 180c while also communicating with and connecting to another RAN, such as an eNode-B 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNode-Bs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, while the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.

[0055] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a, 184b, routing of control plane information to Access and Mobility Management Functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.

[0056] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements is shown as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0057] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may function as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize the CN support of the WTRUs 102a, 102b, 102c based on the type of service utilizing the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases, such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services for machine type communication (MTC) access, and / or the like. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.

[0058] The SMFs 183a and 183b may be connected to the AMFs 182a and 182b in the CN 115 via an N11 interface. The SMFs 183a and 183b may also be connected to the UPFs 184a and 184b in the CN 115 via an N4 interface. The SMFs 183a and 183b may select and control the UPFs 184a and 184b and configure the routing of traffic through the UPFs 184a and 184b. The SMFs 183a and 183b may perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0059] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.

[0060] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. Additionally, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to local data networks (DNs) 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.

[0061] 1A-1D and their corresponding descriptions, one or more or all of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, eNode-Bs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or simulate network and / or WTRU functions.

[0062] The emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices may be directly coupled to another device for testing purposes and / or may perform testing using terrestrial wireless communication.

[0063] One or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in an undeployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, e.g., one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0064] This application describes various aspects, including tools, features, examples or embodiments, models, approaches, and the like. Many of these aspects are described with specificity and, at least to illustrate their individual characteristics, are described in what may sometimes sound definitive terms. However, this is for purposes of clarity of description and does not limit the applicability or scope of the aspects. In fact, all of the different aspects may be combined and interchanged to provide further aspects. Moreover, aspects may similarly be combined and interchanged with aspects described in prior applications.

[0065] Aspects described and contemplated in this application may be implemented in many different forms. While Figures 1-10 described herein may provide some examples, other embodiments are also contemplated. The discussion of Figures 1-10 is not intended to limit the breadth of implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects may be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.

[0066] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably.

[0067] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," etc. may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." The use of such terms does not imply a modified ordering of operations unless specifically required. Thus, in this embodiment, the first decoding does not need to be performed before the second decoding, but may occur, for example, before the second decoding, during the second decoding, or during a time that overlaps with the second decoding.

[0068] Various methods and other aspects described herein may be used to modify, for example, modules of video encoder 200 and decoder 300, as shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein presents aspects not limited to VVC or HEVC and may apply to any type, format, or version of video coding, whether described in a standard or recommendation, whether existing, or developed in the future, and to extensions of any such standard and recommendation (including, for example, VVC and HEVC). Unless otherwise indicated or technically excluded, aspects described herein may be used individually or in combination.

[0069] In the embodiments described herein, various numerical values ​​are used, such as the remaining byte count as 013, and nal_unit_type values ​​in the ranges of 0 to 5 and 10 to 21. These and other specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0070] 2 illustrates an exemplary video encoder. While variations of exemplary encoder 200 are contemplated, encoder 200 is described below for clarity without describing all possible variations.

[0071] Before being encoded, a video sequence may undergo encoding pre-processing (201), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with that pre-processing and attached to the bitstream.

[0072] In the encoder 200, a picture is coded by the encoder elements, as described below. The picture to be coded is divided (202) and processed, for example, in units of coding units (CUs). Each unit is coded, for example, using either intra mode or inter mode. When a unit is coded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) whether to use intra mode or inter mode to code the unit, and indicates the intra / inter decision, for example, via a prediction mode flag. A prediction residual is calculated (210), for example, by subtracting the predicted block from the original image block.

[0073] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.

[0074] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (255) to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (Sample Adaptive Offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).

[0075] Figure 3 illustrates an example video decoder. In the exemplary decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs a decoding pass that is the inverse of the encoding pass, as described in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.

[0076] In particular, the decoder's input includes a video bitstream, which may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coding information. Picture partition information indicates how the picture is partitioned. The decoder may then partition the picture according to the decoded picture partition information (335). The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined with a predicted block (355) to reconstruct an image block. The predicted block may be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0077] The decoded picture may further undergo post-decoding processing (385), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing may use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0078] FIG. 4 illustrates an example system in which various aspects and embodiments described herein may be implemented. System 400 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400, singly or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or separate components. In various embodiments, system 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 400 is configured to implement one or more of the aspects described herein.

[0079] The system 400 includes at least one processor 410 configured to execute instructions loaded therein, for example, to implement various aspects described herein. The processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). The system 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Flash, magnetic disk drives, and / or optical disk drives. Storage devices 440 may include, by way of non-limiting example, internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0080] System 400 includes an encoder / decoder module 430 configured to process data to provide, for example, encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 430 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software, as known to those skilled in the art.

[0081] Program code loaded into the processor 410 or the encoder / decoder 430 to perform various aspects described herein may be stored in the storage device 440 and then loaded onto the memory 420 for execution by the processor 410. According to various embodiments, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 may store one or more of various items during the execution of the storage processes described herein. Such stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and operational logic.

[0082] In some embodiments, memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be the memory 420 and / or the storage device 440, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store, for example, the television's operating system. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC, High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).

[0083] Input to the elements of system 400 may be provided through a variety of input devices, as shown in block 445. Such input devices may include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, terrestrially transmitted RF signals by a broadcast station, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples include composite video, not shown in FIG. 4.

[0084] In various embodiments, the input devices of block 445 have associated respective input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to as a channel (for example), (iv) demodulating the downconverted, bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, downconverting received signals to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0085] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 400 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 410, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, as desired, within a separate interface IC or within processor 410. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and an encoder / decoder 430, which operates in combination with memory and storage elements to process the data stream as desired for presentation on an output device.

[0086] The various elements of system 400 may be provided within a unitary housing in which the various elements may be interconnected and transmit data between them using a suitable connection arrangement 425, such as internal buses known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards.

[0087] System 400 includes a communication interface 450 that enables communication with other devices over a communication channel 460. Communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 460. Communication interface 450 may include, but is not limited to, a modem or a network card, and communication channel 460 may be implemented in a wired and / or wireless medium, for example.

[0088] In various embodiments, data is streamed or otherwise provided to system 400 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE, the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these examples is received via communication channel 460 and communication interface 450 adapted for Wi-Fi communication. Communication channel 460 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, streaming data is provided to system 400 using a set-top box that delivers data via an HDMI connection in input block 445. In yet other embodiments, streaming data is provided to system 400 using an RF connection in input block 445. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0089] System 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 can be for a television, a tablet, a laptop, a mobile phone, or other device. The display 475 can also be integrated into other components (e.g., as in a smartphone) or be separate (e.g., an external monitor for a laptop). Other peripheral devices 495, in various example embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (DVR, or digital versatile disc, as an abbreviation for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 495 to provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0090] In various embodiments, control signals are communicated between system 400 and display 475, speakers 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow control between devices with or without user intervention. Output devices may be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speakers 485 may be integrated into a single unit with other components of system 400 within an electronic device such as a television. In various embodiments, display interface 470 includes a display driver, such as a timing controller (TCon) chip.

[0091] Display 475 and speakers 485 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various embodiments where display 475 and speakers 485 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0092] The embodiments may be executed by the processor 410, or by computer software implemented by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 420 may be of any type appropriate to the technology environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 may be of any type appropriate to the technology environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0093] Various implementations include decoding. As used herein, "decoding" can encompass all or part of the processes performed on a received encoded sequence to produce a final output suitable for, for example, a display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of various implementations described herein, such as decoding a portion of a coded point cloud sequence (e.g., encapsulated in an ISOBMFF container using one or more file format structures, e.g., as disclosed herein) to provide partial access to the coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0094] As a further embodiment, in one example, "decoding" refers to entropy decoding only, while in another embodiment, "decoding" refers to differential decoding only, while in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader decoding process will be clear based on the context of the particular description and is believed to be well understood by those skilled in the art.

[0095] Various implementations include encoding. Similar to the above discussion regarding "decoding," "encoding," as used herein, can encompass, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transforming, quantizing, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by the encoders of various embodiments described herein, such as, for example, encoding a video-based point cloud bitstream that includes one or more file format structures (e.g., as disclosed herein) to provide partial access support to different portions of the coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0096] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader encoding process will be clear based on the context of the particular description and is believed to be well understood by those skilled in the art.

[0097] It should be noted that syntax elements as used herein, e.g., atlas_tile_group_layer_rbsp(), VPCCTileGroupSampleEntry, VolumetricSampleEntry, TrackGroupTypeBox, SpatialRegionGroupBox, TrackGroupTypeBox, DynamicVolumetricMetadataSampleEntry, 3DSpatialRegionStruct, VPCCVolumetricMetadataSample, VPCCAtlasSampleEntry, etc., are descriptive terms and therefore do not preclude the use of other syntax element names.

[0098] Where a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.

[0099] Implementations and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the discussed feature implementation may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in, for example, a processor, which generally refers to a processing device and includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, handheld / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0100] References to "one embodiment," "an embodiment," "an example," "one implementation," or "an implementation," and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment," "in an example," "in one implementation," or "in an implementation" in various places throughout this specification, as well as appearances of any other variations, do not necessarily all refer to the same embodiment or example.

[0101] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0102] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0103] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves in some way, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0104] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of only the third listed alternative (C), or selection of only the first and second listed alternatives (A and B), or selection of only the first and third listed alternatives (A and C), or selection of only the second and third listed alternatives (B and C), or selection of all three alternatives (A, B, and C). This can be expanded as many times as the number of listed items, as would be apparent to one of ordinary skill in this and related arts.

[0105] Also, as used herein, the term "signaling" means, in particular, indicating something to a corresponding decoder. In some embodiments, an encoder may signal (e.g., in the encoded bitstream and / or within an encapsulation file such as an ISOBMFF container), for example, a V-PCC parameter set, an SEI message, metadata, edit lists, post-decoder requirements, signals enabling flexible partial access to different parts of a coded point cloud sequence encapsulated in an ISOBMFF container, a dependency list for each signaled object, mapping to spatial domains, 3D bounding box information, etc. In this way, in embodiments, the same parameters are used on both the encoder and decoder sides. Thus, for example, an encoder may send specific parameters to a decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling may be used to simply enable the decoder to know and select specific parameters without transmission (implicit signaling). By avoiding the transmission of any actual functionality, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to a corresponding decoder. While the above refers to the verb form of the word "signal," the word "signal" can also be used as a noun herein.

[0106] As will be apparent to one skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. Information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The signal it carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0107] Capturing and rendering three-dimensional (3D) images (e.g., using 3D point clouds) can have many applications (e.g., telepresence, virtual reality, and large-scale dynamic 3D maps). 3D point clouds can be used to represent immersive media. A 3D point cloud can include a set of points represented in 3D space. The (e.g., each) point can include coordinates and / or one or more attributes. The coordinates can indicate the location of the (e.g., each) point. The attributes can include, for example, one or more of color, transparency, acquisition time, laser reflectivity, or material properties associated with each point. Point clouds can be captured or developed in several ways. Point clouds can be captured or developed (e.g., to sample 3D space) using, for example, multiple cameras and depth sensors, Light Detection and Ranging (LiDAR) laser scanners, etc. Points (e.g., represented by coordinates and / or attributes) can be generated, for example, by sampling objects in 3D space. A point cloud may include multiple points, each of which may be represented by a set of coordinates (e.g., x, y, z coordinates) that map to 3D space. In an embodiment, a 3D object or scene may be represented or reconstructed as a point cloud containing millions or billions of sampled points. A 3D point cloud can represent a static and / or dynamic (moving) 3D scene.

[0108] Point cloud data may be represented and / or compressed (e.g., point cloud compression (PCC)), for example, to store and / or transmit the point cloud data (e.g., efficiently). For example, to support efficient and interoperable storage and transmission of 3D point clouds, geometry-based compression may be used to encode and decode static point clouds, and video-based compression may be used to encode and decode dynamic point clouds. Point cloud sampling, representation, compression, and / or rendering may support lossy and / or lossless coding (e.g., encoding or decoding) of the geometric coordinates and / or attributes of the point cloud.

[0109] FIG. 5 illustrates a system interface 500 for a server 502 and a client 510. The server 502 may be a point cloud server connected to the Internet 504 and other networks 506. The client 510 is also connected to the Internet 504 and other networks 506, enabling communication between nodes (e.g., the server 502 and the client 510). Each node includes a processor, a non-transitory computer-readable memory storage medium, and executable instructions stored in the storage medium that are executable by the processor to implement methods or portions of methods disclosed herein. One or more of the nodes may further include one or more sensors. The client 510 may include (e.g., may also include) a graphics processor 512 for rendering 3D video for a display, such as a head-mounted display (HMD) 508. Any or all of the nodes may comprise a WTRU and communicate over a network, as described above with respect to FIGS. 1A-1D.

[0110] FIG. 6 illustrates a system interface 600 for a server 602 and a client 604. The server 602 may be a point cloud content server 602 and may include a database of point cloud content, logic for processing level of detail, and server management functions. In some examples, processing detail may reduce the resolution for transmission to a client 604 (e.g., a viewing client 604) due to bandwidth limitations or as permitted because the viewing distance is sufficient to allow the reduction. The point cloud content server 602 may communicate with the client 604, and point cloud data and / or point cloud metadata may be exchanged. In some examples, the point cloud data rendered for the viewer may undergo a data structuring process to reduce and / or increase the level of detail, such as from the point cloud data and / or point cloud metadata (e.g., streamed from the point cloud server 602 to the viewing client 604). The point cloud server 602 may stream the point cloud data at the resolution provided for spatial capture, or in some embodiments, may downsample to comply with, for example, bandwidth constraints or viewing distance tolerances. The point cloud server 602 may dynamically reduce the level of detail. In some examples, the point cloud server 602 may (e.g., further) segment the point cloud data and identify objects within the point cloud. In some examples, points in the point cloud data that correspond to selected objects may be replaced with lower resolution data.

[0111] A client 604 (e.g., a client 604 with an HMD) may request a portion and / or tile of a point cloud from the point cloud content server 602 via a bitstream, e.g., a video-based point cloud compression (V-PCC) coded bitstream. For example, the portion and / or tile of the point cloud may be retrieved based on the location and / or orientation of the HMD.

[0112] FIG. 7 illustrates an example 700 of requesting content by a client (e.g., an HMD). It is understood that HMD and client are used interchangeably, such that one or more steps described as being performed by an HMD may be performed by a client (e.g., instead of the HMD). At 702, a position of the HMD may be determined. At 702, an orientation of the HMD may be determined. A viewport from received viewports may be selected. At 704, a timed metadata track indicating one or more 6DoF viewports may be received by the HMD and / or client from a point cloud server. At 706, one or more tile group tracks may be requested from the point cloud server of FIG. 5 or 6. At 708, the requested tile group tracks may be received (e.g., at the HMD). The received tile group track set may convey information for rendering a spatial region or object within a point cloud scene, for example, as described herein. Systems based on FIGS. 1A-6 may be implemented based on the disclosure herein.

[0113] FIG. 8 illustrates an example of a video-based point cloud compression (V-PCC) bitstream structure as a sequence of V-PCC units. A V-PCC bitstream may include a sequence of V-PCC units (e.g., as shown in the example of FIG. 8). A V-PCC unit (e.g., each V-PCC unit) may have a V-PCC unit header and / or a V-PCC unit payload. The V-PCC unit header may describe a V-PCC unit type. Table 1 illustrates an example of a V-PCC unit type. An attribute video data V-PCC unit header may specify one or more attribute types and / or indices, which may enable supporting multiple instances of the same attribute type. Table 2 illustrates an example of a V-PCC attribute type. The occupancy, geometry, and / or attribute video data unit payload (e.g., as shown by the example of FIG. 8) may correspond to a video data unit (e.g., a network abstraction layer (NAL) unit) that can be decoded by a video decoder. The video decoder corresponding video coding component sub-bitstreams (eg, each video coding component sub-bitstream, such as occupancy, geometry, and / or attribute substreams) may be signaled in a V-PCC parameter set.

[0114] [Table 1]

[0115] [Table 2]

[0116] The V-PCC bitstream high-level syntax (HLS) may support, for example, tile groups (e.g., tile sets) within one or more atlas frames. An atlas frame may be divided into tiles and / or tile groups (e.g., tile sets). An atlas frame may be divided into, for example, one or more tile rows and / or one or more tile columns. A tile may be, for example, a rectangular region of an atlas frame. A tile group (e.g., tile set) may contain one or more tiles of an atlas frame. The tiles of a tile group (e.g., tile set) may be independently decodable. The number of tiles in a tile group may vary.

[0117] FIG. 9 is a diagram illustrating an example of tile and tile group division of an atlas frame (e.g., into 24 tiles and 9 tile groups). FIG. 9 is shown with alternating shading to distinguish the nine tile groups. In the example, rectangular tile group division (e.g., only rectangular tile group division) may be supported. A tile group may, for example, include several tiles of an atlas frame that collectively form a rectangular region of the atlas frame (e.g., two or four tiles per tile group, as shown in the example of FIG. 9). A tile group may include a V-PCC tile set associated with the atlas frame.

[0118] Supplemental enhancement information (SEI) messages may be signaled in the V-PCC bitstream, for example, to associate patches and / or volumetric shapes (e.g., rectangles) in an atlas frame with objects in a scene represented by a point cloud. SEI messages may enable and / or support annotation, labeling, and / or adding properties to one or more objects. Objects may correspond to real objects (e.g., physical objects in a scene) and / or conceptual objects (e.g., objects that may be related to physical or other properties). Objects may be associated with parameters and / or properties (e.g., different parameters and / or properties), which may correspond to information (e.g., information provided), for example, during creation and / or editing of a point cloud or a scene graph. Dependency relationships may be defined between different objects. For example, an object may be part of one or more other objects.

[0119] Objects in a point cloud may be persistent in time or may be updated (e.g., at any time and / or frame). Related information (e.g., information associated with an object) may persist, for example, until updated or replaced (e.g., by update / association signaling) or until the end of the bitstream. One or more patches and / or 2D volumetric rectangles may be associated with one or more objects. A 2D volumetric rectangle may include one or more patches, for example, as shown in FIG. 11 herein.

[0120] Time-based media may be stored in one or more file formats, such as the ISO Base Media File Format (ISOBMFF). Files within a media file format (e.g., ISOBMFF) may contain structure and / or media data information for timing presentation of media data, such as audio, video, etc. The file format (e.g., ISOBMFF) may support non-timed data, such as metadata, at different levels within the file structure. The logical structure of the file may be, for example, a video including a set of time-parallel tracks. The temporal structure of the file may be, for example, tracks including a sequence of samples in time. The sequence may be mapped to a timeline of the video (e.g., the entire video). ISOBMFF may be based, for example, on a box-structured file. A box-structured file may include a series of boxes (e.g., atoms) that may have a size and a type. A type (e.g., among multiple types) may be, for example, a 32-bit value. A type may be selected or chosen to be, for example, four printable characters, which may be referred to as a four-character code (4CC). Non-timed data may be included, for example, within a metadata box (eg, at the file level or attached to a stream of timed data, which may be referred to as a video box or track within a video).

[0121] An ISOBMFF container may contain multiple top-level boxes. For example, a video box ("moov") may be the top-level box in an ISOBMFF container. The video box ("moov") may contain metadata for continuous media streams that may be present in the file. The metadata may be signaled within a hierarchy of boxes within the video box, for example, within a track box ("trak"). A track may represent a media stream (e.g., a continuous media stream present in a file). A media stream may contain a sequence of samples (e.g., audio and / or video access units of an elementary media stream). The samples may be encapsulated within a MediaDataBox ("mdat"), which may be present at the top level of the container. Track metadata (e.g., for each track) may include, for example, a list of sample description entries. The sample description entries (e.g., for each sample description entry) may, for example, provide the encoding format and / or encapsulation format that may be used in the track and / or provide initialization data for processing the encoding format and / or encapsulation format. A sample (e.g., each sample) may be associated with one or more sample description entries of a track. An explicit timeline map may be defined for a track (e.g., each track), which may be referred to as an edit list. An edit list may be signaled, for example, using an EditListBox, which may have the following syntax: A sample description entry (e.g., each sample description entry) may define a portion of the track timeline, for example, by mapping a portion of the composition timeline and / or by indicating "empty" time (e.g., a portion of the presentation timeline that maps without media, resulting in an "empty" edit).

[0122] An example syntax for an EditListBox may be provided as follows:

[0123]

number

[0124] The ISOBMFF may support imposing one or more actions on a player and / or renderer. In an embodiment (e.g., in the case of a video stream), a restricted video scheme track may be used to impose one or more actions. For example, post-decoder requirements may be signaled on a video track that is a restricted video scheme track. A track may be converted to a restricted video scheme track, for example, by setting the track's sample entry code to a four-character code (4CC) (e.g., "resv") and adding a RestrictedSchemeInfoBox to the track's sample description (e.g., without modifying other boxes). The original sample entry type, which may be based on the video codec used to encode the stream, may be stored in an OriginalFormatBox within the RestrictedSchemeInfoBox. The RestrictedSchemeInfoBox may include one or more boxes (e.g., three boxes, such as OriginalFormatBox, SchemeTypeBox, and SchemeInformationBox). The OriginalFormatBox may store the original sample entry type, which may be based on the video codec used to encode the component stream. The nature of the restriction can be defined in a SchemeTypeBox.

[0125] 10 illustrates an exemplary structure of a multi-track ISOBMFF V-PCC container. In an embodiment, the multi-track V-PCC container may include, for example, one or more of the following: The multi-track V-PCC container may include a V-PCC track 10002 containing samples that may, for example, carry V-PCC parameter sets and / or atlas sub-bitstream parameter sets and / or atlas sub-bitstream NAL units (e.g., in sample entries). V-PCC and VPCC are used interchangeably herein. A track may include track references to other tracks that may, for example, carry payloads of video compression V-PCC units (e.g., unit types VPCC_OVD, VPCC_GVD, and / or VPCC_AVD). The multi-track V-PCC container may include, for example, a constrained video scheme track, where samples may include access units of a video coded elementary stream of occupancy map data (e.g., payloads of V-PCC units of type VPCC_OVD). A multi-track V-PCC container may, for example, include one or more restricted video scheme tracks, where samples may include access units of video coded elementary streams of geometry data (e.g., payloads of V-PCC units of type VPCC_GVD). A multi-track V-PCC container may, for example, include zero or more restricted video scheme tracks, where samples may include access units of video coded elementary streams of attribute data (e.g., payloads of V-PCC units of type VPCC_AVD).

[0126] There is growing interest in new media (e.g., VR and / or immersive 3D graphics). 3D point clouds may represent immersive media. Immersive media may enable new forms of interaction and communication with virtual worlds. 3D point clouds may be represented by large amounts of information. Efficient coding (e.g., efficient coding algorithms) may reduce storage and / or transmission resources and time involved in storing and transmitting 3D point cloud data (e.g., dynamic 3D point cloud data).

[0127] A point cloud sequence may represent a scene having multiple objects. In embodiments, individual objects (e.g., represented in a point cloud sequence) may be accessed (e.g., streamed and / or rendered), for example, without decoding other portions of the scene. Similarly, one or more portions of an object (e.g., a single object) represented by a point cloud may be accessed without decoding the entire point cloud.

[0128] The SEI message may, for example, annotate, label, and / or add properties to patches and / or volumetric rectangles. One or more SEI messages may, for example, enable partial access and rendering of a V-PCC sequence. Atlas sub-bitstream data may be carried in tracks (e.g., a single track). Carrying sub-bitstream data in a single track may lead to a streaming application downloading and decoding excessive atlas information, even if, for example, a user may be interested in (e.g., only interested in) a particular region / object within the V-PCC content or a subset of atlases within the V-PCC content, which may, for example, lead to excessive consumption of time and computing resources and degrade the user experience. Tracks (e.g., and associated signaling) may impose restrictions (e.g., excessive restrictions) on viewport signaling and / or may be incompatible with camera parameter and / or viewport position SEI messages.

[0129] The file format structure may allow flexible partial access to different parts of the coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0130] A V-PCC atlas tile group track may be provided. A tile group (e.g., each tile set), or a group of tile groups, may be encapsulated in a separate track (e.g., called an atlas tile group track), for example, if an atlas substream of a V-PCC bitstream contains multiple tile groups. The atlas tile group track may carry NAL units with atlas_tile_group_layer_rbsp() payloads of one or more atlas tile groups, for example, to allow access to the tile groups (e.g., direct access to the tile groups).

[0131] Patches in an atlas frame that may correspond to spatial regions and / or objects in a point cloud scene may be mapped to atlas tile groups, for example, to support partial access in an ISOBMFF container of a V-PCC coded stream. Tile groups may be carried in separate atlas tile group tracks in the container. For example, when tile groups are carried in separate atlas tile group tracks in the container, a player, streaming client, etc. may be able to identify and retrieve tile group tracks (e.g., only a set of tile group tracks) that carry information for rendering selected spatial regions or objects in the point cloud scene.

[0132] A V-PCC track 10002 may be linked to one or more attratile group tracks based on track references having a track reference type defined, for example, using a four-character code (4CC) (e.g., "pcct"). Track references of a defined track reference type may be used, for example, to link a V-PCC track 10002 to one or more attratile group tracks (e.g., to each attratile group track). An attratile group track (e.g., each attratile group track) may be grouped with one or more other video coded V-PCC component tracks that may carry component information for tile groups (e.g., tile sets) within the attratile group track (e.g., using ISO / IEC 14496-12 track groups). A track group definition may include, for example, addresses of tile groups that may be associated with tracks within the track group.

[0133] A V-PCC tile group track may be identified, for example, by a sample description (e.g., VPCCTileGroupSampleEntry). The sample entry type of a V-PCC atlas tile group track may be, for example, "vpt1". The definition of a VPCCTileGroupSampleEntry may be, for example, as follows:

[0134] [Table 3]

[0135]

number

[0136] A sample entry may describe a media sample of a V-PCC tile group track. In an embodiment, a VPCCTileGroupSampleEntry may not include a VPCCConfigurationBox. A VPCCConfigurationBox may be included in the sample description of the main V-PCC track 10002. Other boxes (e.g., other optional boxes) may be included.

[0137] The semantics of the fields of a VPCCTileGroupSampleEntry may be, for example, as follows: The parameter compressorname (e.g., in the base class VolumetricSampleEntry) may indicate the name of the compressor used (e.g., the value "\013VPCC Coding"). The first byte may indicate the count of the remaining bytes, which may be represented, for example, by \013 (e.g., octal 13, which is decimal 11) as the number of bytes remaining in the string.

[0138] Samples in atlas tile group tracks may, for example, have a sample format (e.g., the same sample format) defined for samples of V-PCC track 10002 (e.g., as provided in ISO / IEC 23090-10). NAL units carried in atlas tile group track samples may, for example, have nal_unit_type values ​​within multiple ranges (e.g., an inclusive range of 0 to 5 and an inclusive range of 10 to 21).

[0139] In (e.g., additional or alternative) embodiments, the number and / or layout of tile groups (e.g., tile sets) within an atlas frame may be fixed (e.g., over the duration of the coded point cloud sequence), e.g., to avoid an increase (e.g., explosion) in the number of tracks within the container file.

[0140] In (e.g., additional or alternative) embodiments, an atlas tile group track may include a track reference to a V-PCC track 10002 of the atlas to which the atlas tile group (e.g., carried by the atlas tile group track) belongs. The track reference may enable a parser to identify the V-PCC track 10002 associated with the atlas tile group track. For example, the parser may identify the V-PCC track 10002 associated with the atlas tile group track based on the track identification (ID) of the atlas tile group track.

[0141] Atlas style group tracks and component tracks may be grouped. V-PCC component tracks associated with atlas style group tracks (e.g., tracks that may carry video coded occupancy 10004, geometry 10006, and / or attribute information 10008) may be grouped with the track using a track group with a "vptg" TrackGroupTypeBox, for example, as follows:

[0142]

number

[0143] The semantics of the fields of a VPCCTileGroupBox may be, for example: num_tile_groups_minus1+1 may indicate the number of V-PCC tile groups or V-PCC tile sets associated with the track group. The tile_group_id may indicate the ID of the V-PCC tile group or tile set, and may be the same as the atgh_address (e.g., in ISO / IEC23090-5).

[0144] In (e.g., additional or alternative) embodiments, a SpatialRegionGroupBox may be used to group atlas tile group tracks and corresponding component tracks, for example, based on updates to the syntax of the SpatialRegionGroupBox to include a list of associated tile group identifiers (e.g., similar to embodiments described herein).

[0145] In (e.g., additional or alternative) embodiments, a single track reference from V-PCC tracks 10002, which may use the track_group_id of a VPCCTileGroupBox, may be used to reference (e.g., collectively reference) one or more tracks (e.g., all tracks) that may be associated with a V-PCC tile group (e.g., a V-PCC set of tiles) or a V-PCC tile group set. In an example embodiment, the TrackReferenceTypeBox of the track reference may have an entry in its track_ID array with the track_group_id of the track group of the V-PCC tile group or V-PCC tile group set. A bit of a flag of the TrackGroupTypeBox (e.g., bit 0 or the least significant bit) may be used, for example, to indicate the uniqueness of the track_group_id. The semantics of the flag may be defined, for example, as follows: Bit 0 of the flag of the TrackGroupTypeBox (e.g., bit 0 is the least significant bit) may be used, for example, to indicate the uniqueness of the track_group_id. In an embodiment, (flags&1) equal to 1 in a TrackGroupTypeBox of a particular track_group_type may indicate that the track_group_id in that TrackGroupTypeBox is not equal to the track_ID value and is not equal to the track_group_id of a TrackGroupTypeBox with a different track_group_type. (flags&1) may, for example, be equal to 1 in (e.g., all) TrackGroupTypeBoxes of (e.g., the same) values ​​of track_group_type and track_group_id if (flags&1) is equal to 1 in TrackGroupTypeBoxes with particular values ​​of track_group_type and track_group_id.

[0146] In (e.g., additional or alternative) embodiments, the VPCCTileGroupBox may contain the track ID of the atlas track to which the tile group track belongs. The VPCCTileGroupBox may, for example, extend the TrackGroupTypeBox "vptg" as follows:

[0147]

number

[0148] In this case, the various fields (eg, the semantics of the fields) of the VPCCTileGroupBox may include: atls_track_ID can be the track ID of the atlas track to which the tile group represented by the VPCCTileGroupBox belongs. num_tile_groups_minus1+1 may be the number of V-PCC tile groups or V-PCC tile sets associated with the track group. The tile_group_id may be the ID of the V-PCC tile group (for example, additionally provided as atgh_address in ISO / IEC23090-5).

[0149] A VPCCTileGroupBox may, for example, use an atlas ID as an alternative to using a track ID. A VPCCTileGroupBox may use the atlas ID of the atlas sub-bitstream to which the tile group represented by the VPCCTileGroupBox belongs. In this case, for example, a VPCCTileGroupBox may extend the TrackGroupTypeBox "vptg" as follows:

[0150]

number

[0151] In this case, the various fields (eg, the semantics of the fields) of the VPCCTileGroupBox may include: atlas_id, which may be equal to the atlas ID of the atlas to which the tile group represented by the VPCCTileGroupBox belongs. atlas_id may be equal to one of the vps_atlas_id values ​​that may be signaled in the V-PCC parameter set (VPS), for example. num_tile_groups_minus1+1 may be the number of V-PCC tile groups or V-PCC tile sets associated with the track group. The tile_group_id may be the ID of the V-PCC tile group (for example, additionally provided as atgh_address in ISO / IEC23090-5).

[0152] A volumetric metadata track may be a timed metadata track that may carry information about one or more objects (e.g., one or more different objects) within the point cloud scene and / or 3D spatial division. The object information may be carried in the samples of the track. A timed metadata track may have a defined sample entry (e.g., DynamicVolumetricMetadataSampleEntry) with the 4CC "dyvm" that may extend MetadataSampleEntry, for example, as follows:

[0153]

number

[0154] The volumetric metadata track may, for example, include a "cdsc" track reference to the V-PCC track 10002.

[0155] One or more samples of a volumetric metadata track may include, for example, a table that may map object identifiers to one or more track groups carrying V-PCC tile groups (e.g., V-PCC tile sets) mapped to one or more corresponding objects. One or more samples may include a dependency list for a signaled object (e.g., each signaled object), which may include identifiers of other objects on which the signaled object depends. A sample of a volumetric metadata track may be defined, for example, as follows:

[0156]

number

[0157] where 3DSpatialRegionStruct may be defined, for example, as follows:

[0158]

number

[0159] The semantics of the fields of the VPCCVolumetricMetadataSample may include, for example, one or more of the following: The region_updates_flag may indicate, for example, whether the sample contains updates to a 3D spatial region. The object_updates_flag may indicate, for example, whether the sample contains updates to point cloud scene objects. num_obj_updates may, for example, indicate the number of point cloud scene objects updated in the sample. obj_index_length[i] may, for example, indicate the object index length (eg, number of bytes) of the ith object in the sample's object update list. object_index[i] may, for example, indicate the index of the ith object in the sample's object update list. obj_cancel_flag[i] may indicate, for example, whether the ith object in the sample's object update list is to be cancelled. obj_spatial_region_mapping_flag[i] may indicate, for example, whether mapping to the spatial domain may be signaled for the i-th object in the sample's object update list. obj_depdendencies_present_flag[i] may, for example, indicate whether object dependency information may be available for the i-th object in the sample's object update list (e.g., where a value of 0 may indicate that the object does not depend on other objects, and a value of 1 may indicate that the object depends on one or more objects in the point cloud scene). obj_bounding_box_present_flag[i] may, for example, indicate whether 3D bounding box information may be available for the i-th object in the object update list of the sample (e.g., where a value of 0 may indicate that no bounding box information is given, and a value of 1 may indicate that 3D bounding box information for the i-th object may be signaled in the sample). num_spatial_regions[i] may, for example, indicate the number of 3D spatial regions with which the ith object in the sample's object update list may be associated. region_id[j][i] may indicate, for example, the identifier of the jth spatial region with which the ith object in the sample's object update list may be associated. num_track_groups[i] may, for example, indicate the number of track groups that the ith object in the sample's object update list may be associated with. track_group_id[j][i] may, for example, indicate the identifier of the jth track group (eg, the jth tile set) with which the ith object in the sample object update list may be associated. num_obj_depedencies[i] may indicate, for example, the number of objects on which the ith object in the sample's object update list may depend. obj_dep_index_length[j][i] may indicate, for example, the length in bytes of the index of the jth object on which the ith object in the sample's object update list may depend; or obj_index[j][i] may, for example, indicate the index of the jth object on which the ith object may depend in the sample's object update list.

[0160] In (e.g., additional or alternative) embodiments, updated objects in a sample of a volumetric metadata track may be mapped (e.g., directly mapped) to a V-PCC tile group (e.g., a V-PCC tile set) that includes, for example, patches associated with one or more objects. A corresponding sample format syntax (e.g., in this embodiment) may be, for example, as follows:

[0161]

number

[0162] The semantics of the fields in the sample format syntax may be similar to the semantics of the fields in the sample format in the embodiments described herein, with the exception of, for example, one or more of the following fields: num_tile_groups[i] may indicate, for example, the number of V-PCC tile groups or V-PCC tile sets that the i-th object in the sample's object update list may be associated with; or tile_group_id[j][i] may, for example, indicate the identifier of the jth V-PCC tile group (e.g., the jth V-PCC tile set) with which the ith object in the sample's object update list may be associated. For example, the identifier may be identical to the value of atgh_address in the atlas tile group header of the V-PCC tile group (e.g., where atgh_address may specify the tile group address of the tile group). The value of atgh_address may, for example, be inferred to be equal to 0 if not present.

[0163] A sample (e.g., any sample) in a volumetric metadata track may be marked as a synchronization sample. For a sample in a volumetric metadata track, if at least one of the referenced Visual Volumetric Video-based Coding (V3C) track and V3C and atlas style track media sample references having the same decoding time is a synchronization sample, the sample may be marked as a synchronization sample. A sample that does not have the same decoding time as the synchronization sample may (e.g., or may not) be marked as a synchronization sample. A synchronization sample in a timed metadata track may carry information about spatial regions and / or objects (e.g., all available spatial regions and / or objects) available at the timestamp of the synchronization sample. An asynchronous sample in a timed metadata track may carry updates (e.g., updates only) to spatial region and / or 3D object information relative to previous samples up to and including the first preceding synchronization sample.

[0164] In an embodiment, for example, if track grouping is not used to group tracks belonging to the same attratile group, and the attratile group tracks are linked to associated component tracks (e.g., using track references), then updated objects in the sample of the volumetric metadata track may be mapped to track IDs associated with the attratile group tracks that carry information related to the attratile group. The V-PCC component tracks associated with the tile group tracks may be identified, for example, by following the track references from the attratile group tracks.

[0165] In an embodiment (eg, additional or alternative), a sample of a volumetric metadata track may carry a volumetric annotation SEI message.

[0166] In (e.g., additional or alternative) embodiments, the volumetric metadata track may replace (e.g., or may be used together with) a dynamic spatial domain timed metadata track (e.g., as specified in ISO / IEC CD23090-10), e.g., as a general track that may carry metadata for 3D spatial regions and / or objects in a point cloud scene.

[0167] A V-PCC atlas track may be provided. Atlas sub-bitstreams (e.g., each atlas sub-bitstream) may be carried in a separate track called an atlas track, e.g., if a V-PCC bitstream has two or more atlas sub-bitstreams. An atlas track may carry (e.g., carry only) atlas NAL units belonging to the atlas sub-bitstream associated with the track. NAL units associated with one or more tile groups (e.g., one or more tile sets) may be carried in a separate atlas tile group track, e.g., if an atlas sub-bitstream associated with the atlas track includes multiple atlas tile groups (e.g., multiple atlas tile sets).

[0168] Atlas sub-bitstreams of a V-PCC bitstream may be carried in separate atlas tracks. V-PCC track 10002 may contain track references (e.g., of a certain type defined using 4CC) to each atlas track, which may link the main track to the atlas track.

[0169] A V-PCC atlas track may be identified, for example, by a VPCCAtlasSampleEntry sample description. The sample entry type of a V-PCC atlas track may be, for example, "vpa1" or "vpag." The definition of a VPCCAtlasSampleEntry may be, for example, as follows:

[0170] [Table 4]

[0171]

number

[0172] A sample entry (e.g., as shown in the examples herein) may describe a media sample of a V-PCC atlas track. In an example, a VPCCAtlasSampleEntry may not include a VPCCConfigurationBox. For example, a VPCCConfigurationBox may be included in the sample description of a main V-PCC track. Other boxes (e.g., other optional boxes) may be included.

[0173] The semantics of the fields of a VPCCAtlasSampleEntry may include, for example, one or more of the following: compressorname (e.g., base class VolumetricSampleEntry) may indicate, for example, the name of the compressor to be used along with a value (e.g., "\013VPCC Coding"), where, for example, the first byte is a count of the remaining bytes (e.g., represented by \13, where 13 (e.g., octal 13) is 11 (e.g., decimal 11)), and the number of bytes is the remainder of the string. lengthSizeMinusOne+1 may, for example, indicate the length (e.g., in bytes) of the NALUnitLength field in the sample in the atlas stream to which the configuration record applies (e.g., a size of 1 byte may be indicated by a value of 0), and the value of the field may be equal to ssnh_unit_size_precision_bytes_minus1 in the sample_stream_nal_header() of the atlas substream. numOfSetupUnitArrays may, for example, indicate the number of arrays of atlas NAL units of the indicated type. array_completeness may, for example, indicate (e.g., when equal to 1) that atlas NAL units of a given type (e.g., all atlas NAL units) are possible in the subsequent array and not in the stream, or (e.g., when equal to 0) that additional atlas NAL units of the indicated type are possible in the stream (e.g., where default and allowed values ​​may be constrained by the sample entry name). NAL_unit_type may, for example, indicate the type of atlas NAL unit in the subsequent array (which may, for example, have an indicated type), where NAL_unit_type may have a value (e.g., as defined in ISO / IEC 23090-5) and / or where NAL_unit_type may be restricted to one or more values ​​indicating, for example, NAL_ASPS, NAL_AFPS, NAL_PREFIX_SEI, and / or NAL_SUFFIX_SEI atlas NAL units. numNALUnits may, for example, indicate the number of atlas NAL units of the indicated type that may be included in the configuration record of the stream to which the configuration record applies, where the SEI array may include (e.g., may only include) SEI messages of a declarative nature (e.g., SEI messages that provide information about the entire stream, such as user data SEIs). SetupUnitLength may indicate the size (e.g., in bytes) of the setupUnit field, where the length field may include, e.g., the size of the NAL unit header and / or the NAL unit payload, or may not include, e.g., the length field; or The setupUnit may store NAL units of type NAL_ASPS, NAL_AFPS, NAL_PREFIX_SEI, or NAL_SUFFIX_SEI (e.g., as defined in ISO / IEC 23090-5), where NAL_PREFIX_SEI or NAL_SUFFIX_SEI (e.g., if present in the setupUnit) may store SEI messages of a "declarative" nature.

[0174] Tracks of an atlas (e.g., the same atlas) may be grouped. In an embodiment, tracks (e.g., all tracks) carrying information belonging to an atlas sub-bitstream (e.g., the same atlas sub-bitstream) may be grouped together, for example, using track grouping (e.g., as described in ISO / IEC 14496-12) and / or a defined track group type. A track group of an atlas may include, for example, an atlas track, an atlas tile group track, and a V-PCC component track that may be associated with the atlas. The track group type may be defined, for example, as follows, using, for example, a "vpsg" TrackGroupTypeBox (e.g., the TrackGroupTypeBox may have a track_group_id field defined in ISO / IEC 14496-12):

[0175]

number

[0176] Static spatial regions may be signaled. Static 3D spatial regions may be defined for V-PCC content. Atlas tile groups may be carried in separate tracks. VPCCSpatialRegionsBox (e.g., as provided in ISO / IEC 23090-10) may be extended to indicate (e.g., using a flag such as all_tiles_in_single_track_flag) whether a tile group (e.g., all tilesets) is carried in a single atlas track or whether each tile group (e.g., each tile set) is carried separately in an atlas tile group track. As provided in the exemplary syntax herein, a flag-based VPCCSpatialRegionsBox may associate track group IDs of track groups of various tracks (e.g., all tracks) corresponding to atlas tile group tracks with 3D spatial regions (e.g., 3D spatial regions signaled in VPCCSpatialRegionsBox).

[0177] An example syntax for VPCCSpatialRegionsBox may be provided as follows:

[0178]

number

[0179] The various fields of the VPCCSpatialRegionsBox may include: all_tiles_in_single_track_flag, which may indicate whether tiles (e.g., all tiles) are carried in the V3C track of the corresponding atlas, or whether tiles (e.g., all tiles) are carried separately in atlas tile tracks. A value of 1 may, for example, indicate that tiles (e.g., all tiles) are carried in the V3C track. A value of 0 may, for example, indicate that tiles are carried in separate atlas tile tracks. component_track_group_id, which may identify the track group of tracks carrying the V3C component of the associated 3D spatial region, or tile_track_group_id, which may identify the track group of the atlas tile track for the associated 3D spatial region.

[0180] In an embodiment, for example, when track groups are not used to group tracks that belong to an atlas style group (e.g., the same atlas style group), and the atlas style group tracks are linked to associated component tracks using track references, the track ID (e.g., track ID only) of the atlas style group track associated with the 3D spatial region may be signaled, and the component tracks for the atlas style track may be identified (e.g., by following the track references from the atlas style group track to the component track).

[0181] An example syntax for VPCCSpatialRegionsBox may be as follows:

[0182]

number

[0183] In an embodiment, tile_track_id may represent the track ID of an atlas tile group track associated with the 3D spatial region.

[0184] A system, method, and / or means for signaling viewport information may be implemented. In an embodiment, one or more camera parameters may be signaled. A six degrees of freedom (6DoF) viewport may be defined by two types of camera parameters, e.g., extrinsic camera parameters and intrinsic camera parameters. The extrinsic camera parameters may be signaled (e.g., using an ExtCameraInfoStruct data structure).

[0185] An example syntax for the ExtCameraInfoStruct data structure may be as follows:

[0186]

number

[0187] The semantics of the fields defined in ExtCameraInfoStruct may be as follows: pos_x, pos_y, and pos_z may respectively indicate the x, y, and z coordinates (e.g., in meters) of the viewport's position in the global reference coordinate system. -16 It may be in meters. quat_x, quat_y, and quat_z may indicate the x, y, and z components of the rotation of the viewport area, respectively, using quaternion representation. The coordinate values ​​may be floating-point values ​​in the inclusive range of -1 to 1. The values ​​may specify the x, y, and z components of the rotation, i.e., qX, qY, and qZ, to be applied to transform the global coordinate axes into the camera's local coordinate axes, using quaternion representation. The fourth component of the quaternion qW may be calculated as follows: qW=sqrt(1-(qX 2 +qY 2 +qZ 2 )) The point (w, x, y, z) is at an angle 2 about the axis directed by the vector (x, y, z). * cos^{-1}(w)=2 * It can represent rotation by sin^{-1}(sqrt(x^{2}y^{2}z^{2}).

[0188] The internal camera parameters may be signaled, for example, using the IntCameraInfoStruct data structure.

[0189] Viewport information (eg, using a ViewportInfoStruct data structure) may be signaled based on, for example, external and internal camera parameters.

[0190] An example syntax for the ViewportInfoStruct data structure may be as follows:

[0191]

number

[0192] The semantics of the fields defined in ViewportInfoStruct may be as follows: center_view_flag may be a flag that indicates whether the signaled viewport position corresponds to the center of the viewport and / or to one of the two stereo positions of the viewport. A value of 1 may indicate that the signaled viewport position corresponds to the center of the viewport. A value of 0 may indicate that the signaled viewport position corresponds to one of the two stereo positions of the viewport. left_view_flag may be a flag that indicates whether the signaled viewport information is for a left stereo position of the viewport or a right stereo position of the viewport. A value of 1 may indicate that the signaled viewport information corresponds to a left stereo position of the viewport. A value of 0 may indicate that the signaled viewport information corresponds to a right stereo position of the viewport. extCamInfo can be an instance of ExtCameraInfoStruct that defines the external camera parameters of the viewport. intCamInfo can be an IntCameraInfoStruct instance that defines the internal camera parameters of the viewport.

[0193] A viewport-timed metadata track may be implemented. In an embodiment, a general timed metadata track for indicating a 6DoF viewport may include a ViewportInfoSampleEntry in a SampleDescriptionBox. The purpose of the timed metadata track may be indicated by the track sample entry type. An exemplary ViewportInfoSampleEntry data structure may include a ViewportConfigurationBox data structure (e.g., one ViewportConfigurationBox data structure).

[0194] An example syntax for the ViewportConfigurationBox data structure may be as follows:

[0195]

number

[0196] The semantics of the fields defined in the ViewportConfigurationBox data structure may be as follows: dynamic_int_camera_flag equal to 0 may indicate that the internal camera parameters are fixed for all samples that reference the sample entry. If dynamic_ext_camera_flag is equal to 0, dynamic_int_camera_flag may be equal to 0. dynamic_ext_camera_flag equal to 0 may indicate that the external camera parameters are fixed for all samples that reference the sample entry.

[0197] A sample format for a viewport metadata track (e.g., all viewport metadata tracks) may start with a common portion followed by an extension portion that may be specific to a sample entry in the viewport metadata track. A sample format for a viewport metadata track may be implemented.

[0198] An example syntax for the ViewportInfoSample data structure may be as follows:

[0199]

number

[0200] The semantics of the fields defined in ViewportInfoSample may be as follows: num_viewports may indicate the number of viewports signaled in the sample. viewport_id[i] may be an identifier number that may be used to identify the ith viewport. viewport_cancel_flag[i] equal to 1 may indicate that the viewport with ID, viewport_id[i], may have been canceled. Indicates that viewport information for the i-th viewport follows (e.g., may be conditional on the flag value being 0). int_camera_flag[i] equal to 1 may indicate that the internal camera parameters are present in the i-th viewport camera parameter set of the current sample. int_camera_flag[i] may be equal to 0, for example, if dynamic_int_camera_flag is equal to 0. Furthermore, int_camera_flag[i] may be set as 0, for example, if ext_camera_flag is equal to 0. ext_camera_flag[i] equal to 1 may indicate that the external camera parameters are present in the i-th viewport camera parameter set of the current sample. ext_camera_flag[i] may be equal to 0, for example, if dynamic_camera_flag[i] is equal to 0.

[0201] If a viewport-timed metadata track is present, the external camera parameters represented by ExtCameraInfoStruct() may be present, for example, at the sample entry or sample level. The following may be prohibited from occurring simultaneously: dynamic_ext_camera_flag equals 0 for all samples and ext_cam_flag[i] equals 0 for all samples.

[0202] If a timed metadata track links to one or more media tracks with a "cdsc" track reference, the timed metadata track may describe one or more media tracks individually (eg, each media track).

[0203] A recommended viewport may be implemented. The recommended viewport metadata track may include a RecommendedViewportSampleEntry data structure. The RecommendedViewportSampleEntry data structure may extend the ViewportInfoSampleEntry data structure and include an additional RecommendedViewportInfoBox that may identify the type of recommended viewport signaled in the recommended viewport metadata track.

[0204] An example syntax for the RecommendedViewportSampleEntry data structure may be as follows:

[0205]

number

[0206] The semantics of the fields defined in RecommendedViewportInfoBox may be as follows: viewport_type may specify the viewport type, as listed in Table 3, for all samples that reference a sample entry containing a RecommendedViewportInfoBox. viewport_description can be a null-terminated UTF-8 string that provides a text description of the viewport type.

[0207] Table 3 shows examples of viewport types.

[0208] [Table 5]

[0209] The samples in the viewport metadata track may have the same format as the ViewportInfoSample.

[0210] An initial viewport may be implemented. In an embodiment, the metadata may indicate, for example, an initial viewport that should be used when playing the associated media track.

[0211] When playing a file (e.g., when the file includes an initial viewport metadata track), the player can be expected to parse the initial viewport metadata track associated with the media track and to parse the initial viewport metadata track when rendering the media track.

[0212] The data structure, ViewportInfoSampleEntry, may be implemented with a sample entry type "6inv" that may be used, for example, for the initial viewport metadata track.

[0213] A sample of the initial viewport track may be implemented.

[0214] An exemplary syntax for the InitialViewportSample data structure may have the following format:

[0215]

number

[0216] The semantics of the fields defined in InitialViewportSample may be as follows: A refresh_flag equal to 0 may, for example, specify that the signaled viewport should be used when starting playback from a time-parallel sample in the associated media track. A refresh_flag equal to 1 may, for example, specify that the signaled viewport should always be used when rendering time-parallel samples of each associated media track, both when playing continuously and when starting playback from a time-parallel sample.

[0217] Spatial scalability may be supported. In an embodiment, patches (e.g., in V3C) may support a feature that allows for subsampling of patches across different dimensions before encoding the associated information of the patch. The feature may be referred to as a level of detail (LoD) patch mode. Atlas tiles may allow for dividing the atlas into independently decodable rectangular regions. In one embodiment, patches within an independently decodable rectangular region may not be allowed to use information from patches in other independently decodable rectangular regions. Combining atlas tiles and patch LoD modes may enable various scalability features for use in different applications.

[0218] The LoD (Level of Detail) of a static spatial region may be signaled. To signal the LoD of a static spatial region, the syntax of V3CSpatialRegionsBox may be extended by introducing an additional spatial_scalability_enabled_flag. The spatial_scalability_enabled_flag may signal whether multiple LoDs are supported for the carried V3C content. If the flag is set, the 3D spatial region (e.g., each 3D spatial region) signaled in the V3CSpatialRegionsBox may include an additional num_lods field indicating the number of LoDs available for the 3D spatial region. For each LoD associated with a spatial region, the characteristics of the LoD may be signaled. In one embodiment, a mapping of the tiles storing the patch of the LoD to the corresponding tile ID may be signaled.

[0219] An example syntax for the V3CSpatialRegionsBox data structure (eg, extended to support multiple LoDs) may have the following format:

[0220]

number

[0221] The semantics for the fields defined above may be as follows: lod_scale_min_x and lod_scale_min_y may indicate the minimum LoD scaling factors for the local x and y coordinates of one or more patches in one or more tiles associated with the LoD (e.g., the minimum pdu_lod_scale_x_minus1 value and the minimum pdu_lod_scale_y_idc value, respectively, across a batch (e.g., all patches) in the LoD). lod_scale_max_x and lod_scale_max_y may indicate the values ​​of the maximum LoD scaling factors for the local x and y coordinates of one or more patches in one or more tiles associated with the LoD (e.g., the maximum pdu_lod_scale_x_minus1 value and the maximum pdu_lod_scale_y_idc value, respectively, across the patches in the LoD (e.g., all patches)).

[0222] The LoD of a dynamic spatial region may be signaled. To signal the LoD of a dynamic spatial region, the sample format of one or more samples of a volumetric metadata track may support signaling the LoD of a spatial region (e.g., each spatial region) listed in the sample. A mapping between the LoD and an atlas tile that stores a patch of the LoD (e.g., each LoD) may be signaled.

[0223] An example syntax for the VPCCVolumetricMetadataSample data structure may have the following format:

[0224]

number

[0225] In an embodiment, an object_updates_flag may be associated with one or more of the added object and / or the removed object.

[0226] Player behavior may be implemented based on the adaptive LoD. In an embodiment, a player may identify the presence of dynamic volumetric metadata, for example, when parsing a file and finding a timed metadata track with a DynamicVolumetricMetadatasampleEntry and a "cdsc" track reference to a V3C track. For example, if a dynamic volumetric metadata track is not associated with the main track of V3C content and a V3CSpatialRegionsBox is present in the main track, a static 3D spatial region set may be associated with the V3C content. At some point during playback (e.g., any point), the player may identify a target 3D spatial region set based on the current viewport and characteristics of one or more spatial regions signaled in the V3CSpatialRegionsBox (for static spatial regions) or in the samples of the dynamic volumetric metadata track (e.g., for dynamic spatial regions). For example, if scalability is enabled for 3D spatial regions and / or objects signaled in a V3CSpatialRegionsBox or for samples in a dynamic volumetric metadata track, the player may determine the desired LoD for each of the target spatial regions based on one or more constraints (e.g., the current viewport and / or available network bandwidth). For each target LoD for each target spatial region, the player may identify the tile ID of the tile associated with the LoD based on the mapping in the V3CSpatialRegionsBox or the samples in the dynamic volumetric metadata track. The player may identify the attra tile track (e.g., by inspecting the tile ID in the attra tile track's sample entry) that carries the tile associated with the target LoD. Corresponding component tracks may be identified by following the track criteria from the selected attra tile track to the component track.

[0227] The LoD information may be signaled in an atlas tile track. To facilitate efficient access to the LoD, the tiles carried by the atlas tile track may be limited to tiles associated with the same LoD. For streaming applications, this may allow players and / or streaming clients to download data from the tile track that provides the target LoD.

[0228] The example syntax of AltasTileSampleEntry may allow for signaling LoD information for tiles carried by an AltasTile track.

[0229]

number

[0230] The semantics for the fields defined above may be as follows: The spatial_scalability_enabled_flag may indicate whether LoD mode is enabled for the tile track. lod_id may be an identifier of the LoD. LoDInfoStruct() may be an instance of LoDInfoStruct, which carries the information of the LoD.

[0231] In an embodiment, an atlas tile may include tile tracks associated with different LoDs.

[0232] An example syntax for AtlasTileSampleEntry may be provided to support two use cases (e.g., a single LoD in an atlas tile track and multiple LoDs per tile), for example, as follows:

[0233]

number

[0234] The semantics of the flags disclosed above may be as follows: The single_lod_flag may indicate whether all tiles carried by the atlas tile truck belong to the same LoD. A value of 1 may indicate that all tiles belong to the same LoD. Otherwise, each tile may be associated with a different LoD.

[0235] 11 shows an example of tile mapping of an atlas frame associated with a 3D space. The 3D space may be divided into one or more spatial regions, shown in FIG. 11 as V0, V1, V2, V3, and V4. Each of the spatial regions may be mapped to a V-PCC tile set (e.g., a V-PCC tile group) associated with the atlas frame. V0, V1, V2, V3, and V4 may be mapped to tile groups 0, 1, 2, 3, and 4, respectively. Mapping each of the spatial regions to a tile set may be based on tile identification information (e.g., tile_group_id), as described with respect to FIG. 10.

[0236] Mapping information associated with the mapping of each spatial region to a tile set may be carried in multiple tracks. For example, mapping information associated with mapping spatial region V0 to tile group 0 may be carried in track 0, and mapping information associated with mapping spatial region V1 to tile group 1 may be carried in track 1. Track identification information (e.g., track_group_id) may be used to adjust the mapping information, as described with respect to FIG. 10. Tracking identification information and / or tile identification may be signaled in the timed metadata V-PCC bitstream. In such a case, the track associated with the signaled tracking identification information may be decoded to present mapping information to the associated tile set.

[0237] An object 11000 may be associated with one or more spatial regions. The object may be an area and / or item that may be of interest to a user. One or more flags (e.g., obj_spatial_region_mapping_flag[i]) may be used to indicate that the object is associated with one or more spatial regions, as described with respect to Figure 10. The flags may be signaled in the timed metadata V-PCC bitstream.

[0238] One or more flags may be used to indicate changes (e.g., updates) associated with a spatial region (e.g., region_updates_flag) and / or an object (e.g., object_updates_flag), as described in Figure 10. The flags may be carried in tracks associated with tile sets. Tracks containing flags may be decoded, and the mapping information may be used to access tile sets associated with the updated spatial region, while tracks without flags, for example, do not need to be decoded.

[0239] One or more patches may be associated with a tile set. In embodiments, patches may be mapped to tile sets (e.g., tile groups). As shown by way of example in FIG. 11, tile group 0 may include patches P0, P1, P2, P3, and P4. tile group 1 may include patches P0 and P1. tile group 2 may include patches P0, P1, and P2, and tile group 3 may include patch P0. tile group 4 may include patches P1, P2, and P3. The patches may indicate an orientation associated with the object represented by the spatial region.

[0240] Systems, devices, and methods are described herein for partial access support in International Organization for Standardization Base Media File Format (ISOBMFF) containers for video-based point cloud streams. The file format structure may enable flexible partial access to different parts of a coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0241] The video encoding device may divide a 3D space into a first spatial region and a second spatial region. The video encoding device may map the first spatial region to a first V-PCC tile set and map the second spatial region to a second V-PCC tile set. Each of the first V-PCC tile set and the second V-PCC tile set may be associated with an atlas frame. Each of the first V-PCC tile set and the second V-PCC tile set may be independently decodable. The mapping of the first spatial region to the first V-PCC tile set and the second spatial region to the second V-PCC tile set may be based on tile identification and / or track identification. The first V-PCC tile set may be associated with a first patch set, and the second V-PCC tile set may be associated with a second patch set. The video encoding device may determine a first track carrying first mapping information associated with a first spatial region mapped to a first V-PCC tile set. The video encoding device may determine a second track carrying second mapping information associated with a second spatial region mapped to a second V-PCC tile set. The video encoding device may send the first track and the second track in a timed metadata V-PCC bitstream. The first track and the second track may be sent within a media container file.

[0242] The video encoding device may determine an update dimension flag. The update dimension flag may indicate an update to one or more dimensions of the first spatial region or an update to one or more dimensions of the second spatial region. The video encoding device may send the update dimension flag in a timed metadata V-PCC bitstream.

[0243] The first spatial region may be associated with a first object. The second spatial region may be associated with a second object. The video encoding device may determine one or more object flags indicating that the first spatial region is associated with the first object and the second spatial region is associated with the second object. The video encoding device may send the object flags in a timed metadata V-PCC bitstream. The video encoding device may determine an object dependent flag indicating that a first object associated with the first spatial region depends on a second object associated with the second spatial region and may send the object dependent flag in the timed metadata V-PCC bitstream. The video encoding device may determine an update object flag indicating an update to the first object associated with the first spatial region or an update to the second object associated with the second spatial region and may send the update object flag in the timed metadata V-PCC bitstream.

[0244] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. 1. A video decoding device, comprising: determining a plurality of atlas tile tracks associated with an atlas track of the volumetric content; determining a first set of tile identifiers (IDs) associated with a first attratile track of the plurality of attratile tracks and a second set of tile IDs associated with a second attratile track of the plurality of attratile tracks; identifying a first atlas tile track sample associated with the first set of tile IDs, the first atlas tile track sample including a first set of atlas Network Abstraction Layer (NAL) units associated with the first set of tile IDs; identifying a second atlas tile track sample associated with the second set of tile IDs, the second atlas tile track sample including a second set of atlas NAL units associated with the second set of tile IDs; obtaining a first set of atlas NAL units associated with the first set of tile IDs and a second set of atlas NAL units associated with the second set of tile IDs; decoding the first set of atlas NAL units and the second set of atlas NAL units to reconstruct the volumetric content; 1. A video decoding device comprising: a processor configured to execute:

2. The video decoding device of claim 1 , wherein the first set of tile IDs or the second set of tile IDs are indicated in a tile group configuration.

3. The video decoding device of claim 1 , wherein the first atlas tile track sample or the second atlas tile track sample is identified using a tile group sample entry.

4. The processor: obtaining a first set of component tracks associated with a first set of tiles and a second set of component tracks associated with a second set of tiles, the first set of component tracks indicating component information associated with the first set of tiles, and the second set of component tracks indicating component information associated with the second set of tiles; decoding the first set of atlas NAL units and the second set of atlas NAL units further based on component information associated with the first set of tiles and component information associated with the second set of tiles; The video decoding device of claim 1 , further configured to perform:

5. 5. The video decoding device of claim 4, wherein component information associated with the first set of tiles and component information associated with the second set of tiles indicate at least one of occupancy information, geometry information, or attribute information.

6. The video decoding device of claim 1 , wherein the processor is further configured to receive a timed metadata track associated with the atlas track.

7. 7. The video decoding device of claim 6, wherein the timed metadata track indicates viewport information associated with the atlas track, the viewport information indicating external or internal camera parameters associated with a viewport.

8. 7. The video decoding device of claim 6, wherein the timed metadata track indicates spatial domain information associated with the atlas track, the spatial domain information indicating a mapping of a spatial domain to a set of tiles associated with the first atlas tile track or the second atlas tile track.

9. The video decoding device of claim 1 , wherein the reconstructed volumetric content comprises a point cloud sequence associated with a spatial domain, the spatial domain associated with a user's viewport.

10. 1. A video decoding method comprising: determining a plurality of atlas tile tracks associated with an atlas track of the volumetric content; determining a first set of tile identifiers (IDs) associated with a first attratile track of the plurality of attratile tracks and a second set of tile IDs associated with a second attratile track of the plurality of attratile tracks; identifying a first atlas tile track sample associated with the first set of tile IDs, the first atlas tile track sample including a first set of atlas Network Abstraction Layer (NAL) units associated with the first set of tile IDs; identifying a second atlas tile track sample associated with the second set of tile IDs, the second atlas tile track sample including a second set of atlas NAL units associated with the second set of tile IDs; obtaining a first set of atlas NAL units associated with the first set of tile IDs and a second set of atlas NAL units associated with the second set of tile IDs; decoding the first set of atlas NAL units and the second set of atlas NAL units to reconstruct the volumetric content; 1. A video decoding method comprising:

11. The video decoding method of claim 10 , wherein the first set of tile IDs or the second set of tile IDs are indicated in a tile group configuration.

12. The video decoding method of claim 10 , wherein the first attra-tile track sample or the second attra-tile track sample is identified using a tile group sample entry.

13. The method comprises: obtaining a first set of component tracks associated with a first set of tiles and a second set of component tracks associated with a second set of tiles, the first set of component tracks indicating component information associated with the first set of tiles, and the second set of component tracks indicating component information associated with the second set of tiles; decoding the first set of atlas NAL units and the second set of atlas NAL units further based on component information associated with the first set of tiles and component information associated with the second set of tiles; The video decoding method of claim 10, further comprising:

14. 14. The video decoding method of claim 13, wherein component information associated with the first set of tiles and component information associated with the second set of tiles indicates at least one of occupancy information, geometry information, or attribute information.

15. The video decoding method of claim 10 , the method further comprising receiving a timed metadata track associated with the atlas track.

16. 16. The video decoding method of claim 15, wherein the timed metadata track indicates viewport information associated with the atlas track, the viewport information indicating external or internal camera parameters associated with a viewport.

17. 16. The video decoding method of claim 15, wherein the timed metadata track indicates spatial domain information associated with the atlas track, the spatial domain information indicating a mapping of a spatial domain to a set of tiles associated with the first atlas tile track or the second atlas tile track.

18. The video decoding method of claim 10 , wherein the reconstructed volumetric content comprises a point cloud sequence associated with a spatial domain, the spatial domain associated with a user's viewport.

19. 1. A video encoding device, comprising: determining a plurality of atlas tile tracks associated with an atlas track of the volumetric content; determining a first set of tile identifiers (IDs) associated with a first attratile track of the plurality of attratile tracks and a second set of tile IDs associated with a second attratile track of the plurality of attratile tracks; identifying a first atlas tile track sample associated with the first set of tile IDs, the first atlas tile track sample including a first set of atlas Network Abstraction Layer (NAL) units associated with the first set of tile IDs; identifying a second atlas tile track sample associated with the second set of tile IDs, the second atlas tile track sample including a second set of atlas NAL units associated with the second set of tile IDs; obtaining a first set of atlas NAL units associated with the first set of tile IDs and a second set of atlas NAL units associated with the second set of tile IDs; transmitting the first set of atlas NAL units and the second set of atlas NAL units to a decoding device, wherein the first set of atlas NAL units and the second set of atlas NAL units are to be used to reconstruct the volumetric content; and 1. A video encoding device comprising: a processor configured to execute:

20. The video encoding device of claim 19 , wherein the first set of tile IDs or the second set of tile IDs are indicated in a tile group configuration.

21. The video encoding device of claim 19 , wherein the first atlas tile track sample or the second atlas tile track sample is identified using a tile group sample entry.

22. The processor: obtaining a first set of component tracks associated with a first set of tiles and a second set of component tracks associated with a second set of tiles, the first set of component tracks indicating component information associated with the first set of tiles, and the second set of component tracks indicating component information associated with the second set of tiles; transmitting the first set of atlas NAL units and the second set of atlas NAL units further based on component information associated with the first set of tiles and component information associated with the second set of tiles; 20. The video encoding device of claim 19, further configured to perform:

23. 23. The video encoding device of claim 22, wherein component information associated with the first set of tiles and component information associated with the second set of tiles indicate at least one of occupancy information, geometry information, or attribute information.

24. The video encoding device of claim 19 , wherein the processor is further configured to obtain a timed metadata track associated with the atlas track.

25. 25. The video encoding device of claim 24, wherein the timed metadata track indicates viewport information associated with the atlas track, the viewport information indicating external or internal camera parameters associated with a viewport.

26. 25. The video encoding device of claim 24, wherein the timed metadata track indicates spatial domain information associated with the atlas track, the spatial domain information indicating a mapping of a spatial domain to a set of tiles associated with the first atlas tile track or the second atlas tile track.

27. 20. The video encoding device of claim 19, wherein the reconstructed volumetric content comprises a point cloud sequence associated with a spatial domain, the spatial domain associated with a user's viewport.

28. 1. A video encoding method comprising: determining a plurality of atlas tile tracks associated with an atlas track of the volumetric content; determining a first set of tile identifiers (IDs) associated with a first attratile track of the plurality of attratile tracks and a second set of tile IDs associated with a second attratile track of the plurality of attratile tracks; identifying a first atlas tile track sample associated with the first set of tile IDs, the first atlas tile track sample including a first set of atlas Network Abstraction Layer (NAL) units associated with the first set of tile IDs; identifying a second atlas tile track sample associated with the second set of tile IDs, the second atlas tile track sample including a second set of atlas NAL units associated with the second set of tile IDs; obtaining a first set of atlas NAL units associated with the first set of tile IDs and a second set of atlas NAL units associated with the second set of tile IDs; transmitting the first set of atlas NAL units and the second set of atlas NAL units to a decoding device, wherein the first set of atlas NAL units and the second set of atlas NAL units are to be used to reconstruct the volumetric content; and 1. A video encoding method comprising:

29. 29. The video encoding method of claim 28, wherein the first set of tile IDs or the second set of tile IDs are indicated in a tile group configuration.

30. 30. The video encoding method of claim 28, wherein the first atlas tile track sample or the second atlas tile track sample is identified using a tile group sample entry.

31. The method comprises: obtaining a first set of component tracks associated with a first set of tiles and a second set of component tracks associated with a second set of tiles, the first set of component tracks indicating component information associated with the first set of tiles, and the second set of component tracks indicating component information associated with the second set of tiles; transmitting the first set of atlas NAL units and the second set of atlas NAL units further based on component information associated with the first set of tiles and component information associated with the second set of tiles; 30. The video encoding method of claim 28, further comprising:

32. 32. The video encoding method of claim 31 , wherein component information associated with the first set of tiles and component information associated with the second set of tiles indicates at least one of occupancy information, geometry information, or attribute information.

33. 29. The video encoding method of claim 28, wherein the method further comprises obtaining a timed metadata track associated with the atlas track.

34. 34. The video encoding method of claim 33, wherein the timed metadata track indicates viewport information associated with the atlas track, the viewport information indicating external or internal camera parameters associated with a viewport.

35. 34. The video encoding method of claim 33, wherein the timed metadata track indicates spatial domain information associated with the atlas track, the spatial domain information indicating a mapping of a spatial domain to a set of tiles associated with the first atlas tile track or the second atlas tile track.

36. 29. The video encoding method of claim 28, wherein the reconstructed volumetric content comprises a point cloud sequence associated with a spatial domain, the spatial domain being associated with a user's viewport.

Citation Information

Patent Citations

  • Video-based point cloud stream

    JP2022533225A

  • File format for point cloud data

    JP2022550150A

  • Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method

    JP2022551690A

  • Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method

    JP2023509092A

  • Image processing device and method

    JP7487742B2