MMT Signaling for Streaming Visual Volumetric Video-Based (V3C) and Geometry-Based Point Cloud (G-PCC) Media

The method of streaming visual volumetric video-based coding (V3C) and geometry-based point cloud coding (G-PCC) media addresses the challenge of efficiently compressing and streaming high-quality three-dimensional point clouds and immersive video content, enabling lossy and lossless coding and supporting applications like telepresence and virtual reality.

JP7822390B2Active Publication Date: 2026-03-02INTERDIGITAL PATENT HOLDINGS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023540501
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-05
Filing Date
2022-01-05
Publication Date
2026-03-02
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently compressing and streaming high-quality three-dimensional point clouds and immersive video content, particularly in supporting lossy and lossless coding of point cloud geometry and attributes, as well as immersive video with constrained six degrees of freedom, which are essential for applications like telepresence and virtual reality.

Method used

The implementation of methods and systems for streaming visual volumetric video-based coding (V3C) and geometry-based point cloud coding (G-PCC) media, utilizing MPEG Media Transport Protocol (MMTP) packets to recover requested media assets based on a viewport, and employing MMT signaling for efficient transmission and compression.

Benefits of technology

Enables efficient and interoperable streaming of high-quality three-dimensional point clouds and immersive video content, supporting lossy and lossless coding, and facilitating applications such as telepresence and virtual reality with improved motion parallax and six degrees of freedom.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007822390000003
    Figure 0007822390000003
  • Figure 0007822390000004
    Figure 0007822390000004
  • Figure 0007822390000005
    Figure 0007822390000005
Patent Text Reader

Abstract

Methods, systems, and apparatuses for streaming visual volumetric video-based coding (V3C) media and geometry-based point cloud coding (G-PCC) media are described herein. The method implemented at a receiving device may include receiving one or more of a first message including a list of media assets available for streaming from a sending device, or one or more messages each describing the media assets. The method may further include sending a second message indicating a request for a subset of the media assets to be streamed from the sending device. The requested subset of the media assets may be determined based on a viewport of the receiving device. The method may further include receiving Motion Picture Experts Group (MPEG) Media Transport Protocol (MMTP) packets and processing the packets to recover at least a portion of the requested subset of the media assets.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Application No. 63 / 134,038, filed January 5, 2021, and U.S. Provisional Application No. 63 / 134,143, filed January 5, 2021, the contents of which are incorporated herein by reference. [Background technology]

[0002] High-quality three-dimensional (3D) point clouds and other visual volumetric media have recently emerged as advanced representations of immersive media, such as immersive video content in which a real or virtual 3D scene is captured by multiple real or virtual cameras.

[0003] Recent advances in 3D point capture and rendering technologies may enable novel applications in the fields of telepresence, virtual reality, and large-scale dynamic 3D maps. The 3D Graphics Subgroup of the ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG) is currently working on the development of two 3D point cloud compression (PCC) standards: a geometry-based compression standard for static point clouds and a video-based compression standard for dynamic point clouds. The goal of these standards may be to support efficient and interoperable storage and transmission of 3D point clouds. One of the requirements for these standards may be to support lossy and / or lossless coding of point cloud geometry coordinates and attributes. MPEG-I Visual is another MPEG subgroup working on the development of standards for the compression of immersive video content to support 6DoF virtual walkthroughs with true motion parallax within bounded volumes. Because both video-based point cloud compression and immersive video with constrained six degrees of freedom (6DoF) may rely on video-coded components, these codings of these two types of immersive media may be collectively referred to as visual volumetric video-based coding (V3C), and the same bitstream format may be used to represent their coded information. Summary of the Invention

[0004] Methods, systems, and apparatuses for streaming visual volumetric video-based coding (V3C) media and geometry-based point cloud coding (G-PCC) media are described herein. The method, implemented at a receiving device, may include receiving one or more of a first message including a list of media assets available for streaming from a sending device, or one or more messages each describing the media assets. The method may further include sending a second message indicating a request for a subset of the media assets to be streamed from the sending device. The requested subset of media assets may be determined based on a viewport of the receiving device. The method may further include receiving Motion Picture Experts Group (MPEG) Media Transport Protocol (MMTP) packets and processing the packets to recover at least a portion of the requested subset of media assets. [Brief explanation of the drawings]

[0005] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which like reference numerals indicate similar elements and in which: [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communications system shown in FIG. 1A, according to one embodiment. [Figure 1C] 1B is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1D]1B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 2] FIG. 1 illustrates an example of a video encoder. [Figure 3] FIG. 1 illustrates an example of a video encoder. [Figure 4] FIG. 1 illustrates an example of an exemplary system in which various aspects and embodiments described herein may be implemented. [Figure 5] FIG. 2 illustrates an exemplary system interface between a server and a client. [Figure 6] FIG. 10 illustrates another exemplary system interface between a server and a client. [Figure 7] FIG. 1 illustrates the structure of an exemplary V3C bitstream. [Figure 8] 1 is a table illustrating examples of supported V3C attribute types. [Figure 9] FIG. 1 illustrates an exemplary structure of a V3C container that may be implemented in accordance with the ISOBMFF standard. [Figure 10] FIG. 1 illustrates an exemplary multi-track container with two or more atlases and multiple atlas tiles. [Figure 11] FIG. 1 illustrates an example of a bitstream structure. [Figure 12] 1 is a table providing an example syntax structure of a G-PCC TLV encapsulation unit. [Figure 13] 1 is a table providing possible values ​​of the TLV type parameter and corresponding descriptions. [Figure 14] 1 is a table providing an example syntax structure of a G-PCC TLV unit payload. [Figure 15] A diagram illustrating an exemplary sample structure in which a bitstream providing G-PCC geometry information and attribute information is stored in a single track. [Figure 16]FIG. 1 illustrates an exemplary structure of a multi-track ISOBMFF G-PCC container. [Figure 17] FIG. 1 illustrates an example end-to-end architecture of a system in which MMT signaling is performed. [Figure 18] 1A-1C are exemplary diagrams of package structures according to some embodiments. [Figure 19] 1 is a table providing a list of defined application message types. [Figure 20] 1 is a table providing an example syntax structure of a V3C asset descriptor. [Figure 21] 10 is a table illustrating an example syntax of a V3CAssetGroupMessage. [Figure 22] 1 is a table illustrating exemplary V3C data type values ​​as may be used in the Data_type field. [Figure 23] 10 is a table illustrating an example syntax of a V3CSelectionMessage. [Figure 24] 10 is a table providing a definition of the switching_mode field. [Figure 25] 10 is a table illustrating an example syntax of a V3CViewChangeFeedbackMessage. [Figure 26] 1 is a table providing an example syntax structure of a G-PCC asset descriptor. [Figure 27] 1 is a table illustrating examples of defined G-PCC application message types. [Figure 28] 10 is a table illustrating an example syntax of a group message. [Figure 29] 10 is a table illustrating exemplary G-PCC data type values ​​as may be used in the Data_type field. [Figure 30] 10 is a table illustrating an example syntax of a GPCC selection feedback message. [Figure 31]10 is a table providing a definition of the switching_mode field. [Figure 32] 10 is a table illustrating an example syntax of a G-PCC view change feedback message (e.g., "GPCCViewChangeFeedback"); DETAILED DESCRIPTION OF THE INVENTION

[0006] 1A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. Communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. Communication system 100 may enable multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 100 may use one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word discrete Fourier transform spread OFDM (ZT-UW-DFT-S-OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0007] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a radio access network (RAN) 104, a core network (CN) 106, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which may be referred to as a station (STA), may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a mobile phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (e.g., remote surgery), an industrial device and application (e.g., robots and / or other wireless devices operating in an industrial and / or automated processing chain context), a consumer electronic device, a device operating in a commercial and / or industrial wireless network, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.

[0008] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNode B (eNB), a Home Node B, a Home eNode B, a next generation Node B such as a gNode B (gNB), a new radio (NR) Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0009] The base station 114a may be part of the RAN 104, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple output (MIMO) technology and may utilize multiple transceivers per sector of the cell, for example, using beamforming to transmit and / or receive signals in desired spatial directions.

[0010] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0011] More specifically, as noted above, the communications system 100 may be a multiple-access system and may use one or more channel access schemes, such as, for example, CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a of the RAN 104 and the WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 116 using wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​Uplink (UL) Packet Access (HSUPA).

[0012] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-Advanced, LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).

[0013] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may establish the air interface 116 using NR.

[0014] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using dual connectivity (DC) principles. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions transmitted to / from multiple types of base stations (e.g., eNBs and gNBs).

[0015] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (i.e., Wireless Fidelity, WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access, WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.

[0016] 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area such as a location such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may establish a picocell or a femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in FIG. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 through the CN 106.

[0017] The RAN 104 may communicate with the CN 106, which may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput, latency, error tolerance, reliability, data throughput, mobility, etc. The CN 106 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 and / or CN 106 may communicate directly or indirectly with other RANs that use the same RAT as the RAN 104 or a different RAT. For example, in addition to being connected to the RAN 104, which may utilize NR radio technology, the CN 106 may also communicate with another RAN (not shown) using GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0018] The CN 106 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a public switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as the transmission control protocol (TCP), the user datagram protocol (UDP), and / or the internet protocol (IP) of the TCP / IP Internet protocol suite. The networks 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may use the same RAT as the RAN 104 or a different RAT.

[0019] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that may use a cellular-based wireless technology and a base station 114b that may use an IEEE 802 wireless technology.

[0020] 1B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.

[0021] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0022] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0023] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0024] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.

[0025] The processor 118 of the WTRU 102 may be coupled to and may receive user-entered data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).

[0026] The processor 118 may receive power from the power source 134, but may also be configured to distribute and / or control the power to other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0027] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 may obtain location information by way of any suitable location-determination method while remaining consistent with an embodiment.

[0028] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors. The sensor may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, a humidity sensor, and the like.

[0029] The WTRU 102 may include a full-duplex radio for transmitting and receiving some or all of the signals (e.g., associated with a particular subframe on both the UL (e.g., for transmission) and DL (e.g., for reception)) simultaneously and / or together. The full-duplex radio may include an interference management unit for reducing and or substantially eliminating self-interference through hardware (e.g., chokes) or signal processing via a processor (e.g., via a separate processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmitting and receiving some or all of the signals (e.g., associated with a particular subframe on either the UL (e.g., for transmission) or DL ​​(e.g., for reception)).

[0030] 1C is a system diagram illustrating the RAN 104 and the CN 106 according to one embodiment. As mentioned above, the RAN 104 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 using E-UTRA radio technology. The RAN 104 may also communicate with the CN 106.

[0031] The RAN 104 may include eNodeBs 160a, 160b, and 160c, although it will be understood that the RAN 104 may include any number of eNodeBs while remaining consistent with an embodiment. The eNodeBs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In an embodiment, the eNodeBs 160a, 160b, and 160c may implement MIMO technology. Thus, the eNodeB 160a may, for example, use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.

[0032] Each of the eNodeBs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling, etc. in the UL and / or DL. As shown in FIG. 1C, the eNodeBs 160a, 160b, 160c may communicate with each other via an X2 interface.

[0033] 1C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (PGW) 166. Although the foregoing elements are shown as part of the CN 106, it will be understood that any of these elements may also be owned and / or operated by an entity other than the CN operator.

[0034] The MME 162 may be connected to each of the eNodeBs 162a, 162b, 162c in the RAN 104 via an S1 interface and may function as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies such as GSM and / or WCDMA.

[0035] The SGW 164 may be connected to each of the eNode-Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to and from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions, such as anchoring the user plane during inter-eNode-B handovers, triggering paging when DL data is available to the WTRUs 102a, 102b, 102c, and managing and storing the context of the WTRUs 102a, 102b, 102c.

[0036] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0037] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional landline communications devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0038] Although the WTRU is depicted in FIGS. 1A-1D as a wireless terminal, it is contemplated that in certain representative embodiments, such a terminal may use a wired communication interface (e.g., temporarily or permanently) with the communication network.

[0039] In a representative embodiment, the other network 112 may be a WLAN.

[0040] A WLAN in infrastructure Basic Service Set (BSS) mode may have an access point (AP) of the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic originating from outside the BSS to a STA may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP and transmitted to the respective destination. Traffic between STAs within the BSS may be transmitted, for example, through the AP; the source STA may send traffic to the AP, which may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be viewed and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted between a source STA and a destination STA (e.g., directly between them) in a direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all of the STAs) may communicate directly with each other. The IBSS mode of communication may be referred to herein as an "ad hoc" communication mode.

[0041] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a dynamically configured width. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. With CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit at any given time in a given BSS.

[0042] High Throughput (HT) STAs may use 40 MHz wide channels for communication, which may be formed, for example, through a combination of a primary 20 MHz channel and adjacent or non-adjacent 20 MHz channels.

[0043] A Very High Throughput (VHT) STA may support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. The 40 MHz and / or 80 MHz wide channels may be formed by combining contiguous 20 MHz channels. A 160 MHz channel may be formed by combining eight contiguous 20 MHz channels or by combining two non-contiguous 80 MHz channels, which may be referred to as an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data may pass through a segment parser that may split the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing may be performed separately on each stream. The streams may be mapped to two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration may be reversed, and the combined data may be transmitted to the Medium Access Control (MAC).

[0044] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah may support meter-type control / machine-type communications (MTC), such as MTC devices in macro coverage areas. MTC devices may have specific capabilities, including, for example, support for (e.g., only support for) specific and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).

[0045] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as a primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be configured and / or limited by the STAs among all STAs operating in the BSS that support the minimum bandwidth operating mode. In an 802.11ah example, the primary channel can be 1 MHz wide for STAs (e.g., MTC-type devices) that support (e.g., only) the 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or Network Allocation Vector (NAV) configuration can depend on the condition of the primary channel. For example, if the primary channel is busy, a STA (that only supports 1 MHz mode of operation) transmitting to the AP may cause all of the available frequency bands to be considered busy, even if most of the available frequency bands are idle.

[0046] In the United States, the available frequency band that can be used by 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total bandwidth available for 802.11ah is 6MHz to 26MHz depending on the country code.

[0047] 1D is a system diagram illustrating the RAN 104 and the CN 106 according to one embodiment. As mentioned above, the RAN 104 may communicate with the WTRUs 102a, 102b, 102c over the air interface 116 using NR radio technology. The RAN 104 may also communicate with the CN 106.

[0048] The RAN 104 may include gNBs 180a, 180b, and 180c, although it will be understood that the RAN 104 may include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In an embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a, 180b may utilize beamforming to transmit and / or receive signals to the gNBs 180a, 180b, and 180c. Thus, the gNB 180a may, for example, transmit wireless signals to and / or receive wireless signals from the WTRU 102a using multiple antennas. In one embodiment, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers may be on an unlicensed spectrum, and the remaining component carriers may be on a licensed spectrum. In one embodiment, the gNBs 180a, 180b, and 180c may implement Coordinated Multi-Point (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).

[0049] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of varying or scalable lengths (e.g., including varying numbers of OFDM symbols and / or varying lengths of absolute time).

[0050] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNodeBs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate with and connect to gNBs 180a, 180b, 180c while also communicating with and connecting to another RAN, such as eNodeBs 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNodeBs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNodeBs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, while the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.

[0051] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, DC, interworking between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a, 184b, routing of control plane information to Access and Mobility Management Functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D , the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.

[0052] 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While the foregoing elements are shown as part of the CN 106, it will be understood that any of these elements may also be owned and / or operated by an entity other than the CN operator.

[0053] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 104 via an N2 interface and may function as a control node. For example, the AMF 182a, 182b may be responsible for user authentication of the WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selection of the SMF 183a, 183b for registration, management of registration areas, termination of non-access stratum (NAS) signaling, mobility management, etc. The network slicing may be used by the AMF 182a, 182b to customize the CN support for the WTRUs 102a, 102b, 102c based on the type of service utilizing the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases, such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services for MTC access, etc. The AMFs 182a, 182b may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.

[0054] The SMFs 183a and 183b may be connected to the AMFs 182a and 182b in the CN 106 via an N11 interface. The SMFs 183a and 183b may also be connected to the UPFs 184a and 184b in the CN 106 via an N4 interface. The SMFs 183a and 183b may select and control the UPFs 184a and 184b and configure the routing of traffic through the UPFs 184a and 184b. The SMFs 183a and 183b may perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing DL data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0055] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 104 via an N3 interface, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks such as the Internet 110 to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as packet routing and forwarding, user plane policy enforcement, support for multi-homed PDU sessions, handling user plane QoS, DL packet buffering, mobility anchoring, etc.

[0056] The CN 106 may facilitate communication with other networks. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to the local DNs 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.

[0057] 1A-1D and the corresponding description thereof, one or more or all of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or simulate network and / or WTRU functions.

[0058] The emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices may be directly coupled to another device for testing and / or conducting tests using over-the-air wireless communication.

[0059] One or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in an undeployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, e.g., one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0060] Various methods and other aspects described herein may be used to modify, for example, modules of video encoder 200 and decoder 300, as shown in Figures 2 and 3. Furthermore, the subject matter disclosed herein presents aspects not limited to V3C, G-PCC, and may apply to any type, format, or version of video coding, whether described in a standard or recommendation, whether existing, or developed in the future, and to extensions of any such standard and recommendation (including, for example, V3C and G-PCC). Unless otherwise indicated or technically excluded, aspects described herein may be used individually or in combination.

[0061] In the examples described herein, various numerical values ​​are used, such as the number of bits reserved for fields in a V3C application message or a G-PCC application message. These and other specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0062] 2 illustrates an example of a video encoder. While variations of the exemplary encoder 200 are contemplated, the encoder 200 is described below for clarity without describing all possible variations.

[0063] Before being encoded, a video sequence may undergo encoding preprocessing (201), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the preprocessing, and such metadata may be attached to the bitstream.

[0064] In the encoder 200, a picture may be coded by encoder elements, as described below. The picture to be coded may be divided (202) and processed, for example, in units of coding units (CUs). Each unit may be coded, for example, using either intra mode or inter mode. If the unit is coded in intra mode, the unit performs intra prediction (260); in inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder may determine (205) whether to use intra mode or inter mode to code the unit, and indicate the intra / inter decision, for example, via a prediction mode flag. A prediction residual may be calculated (210), for example, by subtracting the predicted block from the original image block.

[0065] The prediction residual may then be transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, may be entropy coded (245) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process.

[0066] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. Combining the decoded prediction residual with the prediction block (255) reconstructs an image block. A ln-loop filter (265) is applied to the reconstructed picture to reduce coding artifacts, for example, to perform deblocking / sample adaptive offset (SAO) filtering. The filtered image is stored in a reference picture buffer (280).

[0067] Figure 3 illustrates an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by a decoder element as described below, and the video decoder 300 generally performs a decoded pass that is the inverse of the encoding pass as described in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data. In particular, the decoder's input may include a video bitstream that may be generated by the video encoder 200. The bitstream may first be entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned, and the decoder may then divide the picture according to the decoded picture partition information (335). The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (355) to reconstruct an image block, which may be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0068] The decoded picture may further undergo post-decoding processing (385), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing may use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0069] FIG. 4 illustrates an example system in which various aspects and embodiments described herein may be implemented. System 400 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400, singly or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or separate components, and in various embodiments, system 400 is communicatively coupled to one or more other systems or other electronic devices, e.g., via a communication bus or via dedicated input and / or output ports. In various embodiments, system 400 is configured to implement one or more of the aspects described herein.

[0070] The system 400 includes at least one processor 410 configured to execute instructions loaded therein to implement various aspects described herein, for example, and may include embedded memory, input / output interfaces, and various other circuits as known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a nonvolatile memory device). The system 400 includes a storage device 440, which may include nonvolatile and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. The storage device 440 may include, by way of non-limiting example, an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.

[0071] System 400 includes, for example, an encoder / decoder module 430 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 430 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 430 may be implemented as a separate element of system 400 or may be incorporated within processor 410 as a combination of hardware and software, as is known to those skilled in the art.

[0072] Program code to be loaded into the processor 410 or the encoder / decoder 430 to perform various aspects described herein may be stored in the storage device 440 and then loaded into the memory 420 for execution by the processor 410, and according to various embodiments, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 may store one or more of various items during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and operational logic.

[0073] In some embodiments, memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 410 or the encoder / decoder module 430) may be used for one or more of these functions. The external memory may be the memory 420 and / or the storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store, for example, the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations, such as MPEG-2. MPEG refers to the Moving Picture Experts Group, and MPEG-2 may also be referred to as ISO / IEC 13818. ISO / IEC 13818-1 is also known as H.222, and 13818-2 is sometimes known as H.262), HEVG (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).

[0074] Inputs to the elements of system 400 may be provided through various input devices, as shown in block 445. Such input devices may include, but are not limited to, (i) a Radio Frequency (RF) section that receives, for example, RF signals transmitted throughout a broadcast by a broadcaster, (ii) a Component (COMP) input terminal (or set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other embodiments may include composite video, although not shown in FIG. 4.

[0075] In various embodiments, the input devices of block 445 may have associated respective input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which in particular embodiments may be referred to (for example) as a channel, (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner to perform various of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and refiltering to a desired frequency band. Various embodiments may rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include, for example, inserting elements between existing elements, such as an amplifier and an analog-to-digital converter, and in various embodiments, the RF section includes an antenna.

[0076] Additionally, it will be appreciated that the USB and / or HDMI terminals may include respective interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections, and that various aspects of input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or, if desired, within processor 410. Similarly, aspects of USB or HDMI interface processing may be implemented, if desired, within a separate interface IC or within processor 410. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and an encoder / decoder 430 operating in combination with memory and storage elements to process the data stream as desired for presentation on an output device.

[0077] The various elements of system 400 may be provided within a unitary housing in which the various elements may be interconnected and transmit data between them using a suitable connection arrangement 425, such as internal buses known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards.

[0078] System 400 includes a communication interface 450 that enables communication with other devices over a communication channel 460. Communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 460. Communication interface 450 may include, but is not limited to, a modem or a network card, and communication channel 460 may be implemented in a wired and / or wireless medium, for example.

[0079] Data may, in various embodiments, be streamed or otherwise provided to system 400 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these examples is received via communication channel 460 and communication interface 450 adapted for Wi-Fi communication. Communication channel 460 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, streaming data is provided to system 400 using a set-top box that delivers data via an HDMI connection in input block 445. Still other embodiments provide streamed data to system 400 using an RF connection in input block 445. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments may use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0080] System 400 can provide output signals to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, a tablet, a laptop, a mobile phone, or other device. The display 475 may also be integrated with other components (e.g., as in the case of a smartphone) or may be separate (e.g., an external monitor for a laptop). The other peripheral devices 495, in various example embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (DVR, in both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0081] In various embodiments, control signals may be communicated between system 400 and display 475, speakers 485, or other peripheral devices 495 using signaling such as AVLink, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to system 400 using communication channel 460 via communication interface 450. Display 475 and speakers 485 may be integrated into a single unit with other components of system 400 in an electronic device such as a television, and in various embodiments, display interface 470 includes a display driver, such as a timing controller (T Con) chip.

[0082] Display 475 and speakers 485 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various embodiments where display 475 and speakers 485 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0083] The embodiments may be executed by the processor 410, or by computer software implemented by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 420 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 may be of any type suitable for the technical environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0084] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of the processes performed on a received encoded sequence, e.g., to generate a final output suitable for display; in various embodiments, such processes include processes typically performed by a decoder, e.g., one or more of entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders in various implementations described herein, e.g., decoding a portion of a coded point cloud sequence (e.g., encapsulated in an ISOBMFF container using one or more file format structures, e.g., as disclosed herein) to provide partial access to the coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0085] As a further embodiment, in some instances, "decoding" may refer only to entropy decoding, in other embodiments, "decoding" may refer only to differential decoding, and in other embodiments, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally will be clear based on the context of the particular description and will be well understood by one of ordinary skill in the art.

[0086] Various implementations involve encoding. Similar to the above discussion regarding "decoding," "encoding" as used herein can encompass, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transforming, quantizing, and entropy coding. In various embodiments, such processes also or alternatively include processes performed by the encoders of various embodiments described herein, such as, for example, encoding a video-based point cloud bitstream that includes one or more file format structures (e.g., as disclosed herein) to provide partial access support for different portions of the coded point cloud sequence (e.g., encapsulated in an ISOBMFF container).

[0087] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase encoding process is intended to refer specifically to a subset of operations or to a broader encoding process in general will be clear based on the context of the particular description and is believed to be well understood by those skilled in the art.

[0088] It should be noted that the syntax elements used herein, such as V3CSelectionMessage, V3CAssetGroupMessage, and V3CViewChangeFeedbackMessage, are descriptive terms and therefore do not preclude the use of other syntax element names.

[0089] Where a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.

[0090] Implementations and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the discussed feature may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be performed in, for example, a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, personal digital assistants (PDAs), and other devices that facilitate communication of information between end users.

[0091] "One embodiment," "embodiment," "example," "one implementation," or "implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment," "in an example," "in one implementation," or "in an implementation," as well as any other variations thereof, appearing in various places throughout this application are not necessarily all referring to the same embodiment or example.

[0092] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0093] Additionally, the application may refer to accessing various portions of information, which may include, for example, one or more of retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, inferring information, or deducing information.

[0094] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves, in some way or another, operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0095] It should be understood that the use of "and / or" below and at least one in the cases of "e.g., "A / B," "A and / or B," and "at least one of A and B" is intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of both alternatives (A and B). As a further example, for "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of only the third listed alternative (C), or selection of only the first and second listed alternatives (A and B), or selection of only the first and third listed alternatives (A and C), or selection of only the second and third listed alternatives (B and C), or selection of all three alternatives (A and B and C). This can be expanded as many times as the number of listed items, as would be apparent to one of ordinary skill in this and related arts.

[0096] Also, as used herein, the word "signaling" refers to, among other things, indicating something to a corresponding decoder. In some embodiments, an encoder may signal (e.g., in an encoded bitstream and / or an encapsulation file such as an ISOBMFF container), for example, a parameter set, an SEI message, metadata, an edit list, post-decoder requirements, a signal enabling flexible partial access to different portions of a coded point cloud sequence encapsulated in an ISOBMFF container, a dependency list for each signaled object, a mapping to a spatial domain, 3D bounding box information, etc. In this way, in one embodiment, the same parameters are used on both the encoder and decoder sides. Thus, for example, an encoder may send specific parameters to a decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling may be used to simply enable the decoder to know and select specific parameters without transmission (implicit signaling). It should be understood that by avoiding transmission of any actual functionality, bit savings are realized in various embodiments, and signaling may be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to a corresponding decoder. Although the above relates to the verb form of the word "signal," the word "signal" may also be used herein as a noun.

[0097] As will be apparent to one skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. Information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The signal it carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0098] Capturing and rendering three-dimensional (3D) images (e.g., using 3D point clouds) may have many applications, such as telepresence, virtual reality, and large-scale dynamic 3D maps. 3D point clouds may be used to represent immersive media. A 3D point cloud may include a set of points represented in 3D space. The (e.g., each) point may include coordinates and / or one or more attributes. The coordinates may indicate the location of the (e.g., each) point. The attributes may include, for example, one or more of color, transparency, acquisition time, laser or material properties, etc. associated with each point. Point clouds may be captured or developed in several ways. Point clouds may be captured or developed using, for example, multiple cameras and depth sensors, light detection and ranging (LIDAR) laser scanners, etc. (e.g., to sample 3D space). The points (e.g., represented by coordinates and / or attributes) may be generated, for example, by sampling objects in 3D space. A point cloud can include multiple points, each of which may be represented by a set of coordinates (e.g., x, y, z coordinates) that map to 3D space, and in one example, a 3D object or scene may be represented or reconstructed with a point cloud containing millions or billions of sampled points. 3D point clouds can represent static and / or dynamic (moving) 3D scenes.

[0099] Point cloud data may be represented and / or compressed (e.g., point cloud compression (PCC)), for example, to store and / or transmit the point cloud data (e.g., efficiently). For example, to support efficient and interoperable storage and transmission of 3D point clouds, geometry-based compression may be used to encode and decode static point clouds, and video-based compression may be used to encode and decode dynamic point clouds. Point cloud sampling, representation, compression, and / or rendering may support lossy and / or lossless coding (e.g., encoding or decoding) of geometric coordinates and / or attributes of the point cloud.

[0100] FIG. 5 illustrates a system interface 500 for a server 502 and a client 510. The server 502 may be a point cloud server connected to the Internet 504 and other networks 506. The client 510 may also be connected to the Internet 504 and other networks 506 to enable communication between nodes (e.g., the server 502 and the client 510). Each node may include a processor, a non-transitory computer-readable memory storage medium, and executable instructions stored in the storage medium that are executable by the processor to implement methods or portions of methods disclosed herein. One or more nodes may further include one or more sensors. The client 510 may include (e.g., may also include) a graphics processor 512 for rendering 3D video for a display, such as a head-mounted display (HMD) 508. Any or all of the nodes may be equipped with a WTRU and communicate over a network, as described above with respect to FIGS. 1A-1D.

[0101] FIG. 6 illustrates a system interface 600 for a server 602 and a client 604. The server 602 may be a point cloud content server 602 and may include a database of point cloud content, logic for processing level of detail, and server management functions. In some examples, processing detail may reduce the resolution for transmission to a client 604 (e.g., a viewing client 604) as permitted due to bandwidth limitations or because the viewing distance is sufficient to allow the reduction. The point cloud content server 602 may communicate with the client 604 and may exchange point cloud data and / or point cloud metadata. In some examples, the point cloud data rendered for a viewer may undergo a data structuring process to reduce and / or increase the level of detail, such as from the point cloud data and / or point cloud metadata (e.g., streamed from the point cloud server 602 to the viewing client 604). The point cloud server 602 may stream the point cloud data at the resolution at which spatial capture was provided, or in some embodiments, may downsample to comply with bandwidth constraints or viewing distance tolerances, for example. The point cloud server 602 may dynamically reduce the level of detail, and in some examples, the point cloud server 602 may segment (e.g., similarly) the point cloud data and identify objects within the point cloud. In some examples, points in the point cloud data that correspond to selected objects may be replaced with lower resolution data.

[0102] A client 604 (e.g., a client 604 with an HMD) may request a portion and / or tile of a point cloud from the point cloud content server 602 via a bitstream, e.g., a video-based point cloud compression (V-PCC) coded bitstream. For example, the portion and / or tile of the point cloud may be retrieved based on the location and / or orientation of the HMD.

[0103] A point cloud may consist of a set of points represented in 3D space using coordinates indicating the location of each point along with one or more attributes associated with each point, such as color, transparency, acquisition time, laser reflectivity, or material properties. Point clouds may be captured in several ways. For example, one technique for capturing point clouds may use multiple cameras and depth sensors. Light detection and ranging (LiDAR) laser scanners may also be used to capture point clouds. The number of points required to realistically reconstruct objects and scenes using point clouds may be in the millions (or even billions). Therefore, efficient representation and compression may be essential for storing and transmitting point cloud data. Similar to point clouds, some immersive video types may also be capable of representing visual volumetric content, e.g., with six degrees of freedom (6DoF), providing support for playback of 3D scenes within a limited range of viewing positions and orientations.

[0104] As substantially described in the above paragraphs, at least two 3D point cloud compression (PCC) standards are proposed: a geometry-based compression standard for static point clouds, and a video-based compression standard for dynamic point clouds. With regard to the video-based compression standard for dynamic point clouds, visual volumetric video-based coding (V3C) is an example, and various aspects of a V3C-based implementation may be described as follows.

[0105] 7 illustrates an example of the structure of an exemplary V3C bitstream. As shown in FIG. 7, the bitstream may include a V3C sample stream 701, which may include a set of V3C units, each having a V3C unit header and a V3C unit payload. The V3C unit header may describe the V3C unit type. For example, the V3C unit type may include V3C_OVD, V3C_GVD, and / or V3C_AVD. V3C units with unit types V3C_OVD, V3C_GVD, and V3C_AVD may be dedicated video data units, geometry attribute video data units, and attribute video data units, respectively. These data units may represent three major components required to reconstruct visual volumetric media content. The dedicated V3C unit payload, geometry V3C unit payload, and attribute V3C unit payload may correspond to video data units (e.g., NAL units) that can be decoded by an appropriate video decoder. A V3C bitstream may also include one or more V3C_VPS units, which may provide parameter sets that define syntax elements that can be used in V3C unit headers. A V3C bitstream may further include an atlas sub-bitstream (e.g., indicated by a V3C unit header V3C_AD), which may carry a network abstraction layer (NAL) sample stream 702 that includes at least a unit including a NAL unit header and a unit that encapsulates data that defines (or partially defines) the coded atlas. For example, as shown in Figure 7, a NAL unit may include a payload of an atlas tile group layer 703 (e.g., a raw byte sequence payload (RBSP)) that corresponds to an atlas tile group, which may include a header and data that describe a patch (i.e., a region in the atlas associated with volumetric information).

[0106] Figure 8 is a table illustrating examples of supported V3C attribute types. The V3C attribute unit header may specify an attribute type in addition to the V3C unit type. The V3C attribute unit header may also specify an index, allowing multiple instances of the same attribute type to be supported. For example, supported attribute types may include texture, material, transparency, reflectance, or surface normal.

[0107] This document describes the V3C container file format.

[0108] Figure 9 shows an example structure of a V3C container as may be implemented in accordance with the ISOBMFF standard. In general, a V3C container may contain volumetric video data 900 further defined by atlas data, geometry data, attribute data, and occupancy data. More specifically, the container may include a V3C atlas track 910 that contains V3C parameter sets and atlas parameter sets within sample entries and atlas component bitstream NAL units within samples. The V3C atlas track may also contain track references to other tracks 920, 930, and 940, or V3C atlas style tracks, that carry payloads of video compressed V3C units (i.e., V3C unit types equal to V3C_OVD, V3C_GVD, and V3C_AVD).

[0109] The container may include one or more V3C video component tracks whose samples contain access units of a video coded elementary stream for geometry data (i.e., payloads of V3C units of type equal to V3C_GVD), as illustrated at 920 in Figure 9. The container may include zero or more V3C video component tracks whose samples contain access units of a video coded elementary stream for attribute data (i.e., payloads of V3C units of type equal to V3C_AVD), as illustrated at 930 in Figure 9. The container may include zero or more V3C video component tracks whose samples contain access units of a video coded elementary stream for occupancy data (i.e., payloads of V3C units of type equal to V3C_OVD), as illustrated at 940 in Figure 9.

[0110] FIG. 10 illustrates an exemplary multi-track container having two or more atlases and multiple atlas styles. When multiple atlases are present in V3C media, these atlases may be carried in separate atlas tracks with track references to associated V3C component tracks (i.e., tracks carrying associated occupancy maps, geometry, and attribute information). When the atlas data includes two or more atlas styles, these atlas styles may be stored in separate atlas style tracks referenced by the atlas track, and additional track references are stored from the atlas style tracks to tracks carrying the atlas style's associated V3C video component information carried by the atlas style tracks. This may be illustrated, for example, in FIG. 10. As illustrated in 1001, the V3C track "v3cb" may contain multiple atlases. The atlases may be stored in separate V3C tracks 1010 and 1020, for example, with sample entries for "v3a1" or "v3ag." The V3C tracks 1010 and 1020 may each include multiple atlas tile tracks 1011 and 1012, and each of the atlas tile tracks 1011 and 1012 may include a V3C component track 1013 and 1014, respectively.

[0111] As substantially described in the paragraphs above, a geometry-based compression for static point clouds (G-PCC) standard may also be defined to support efficient and interoperable storage and transmission of 3D point clouds. Methods, apparatus, and systems are proposed herein that may be performed and / or implemented in accordance with such a geometry-based compression standard.

[0112] FIG. 11 illustrates an example of the structure of a bitstream encoded in accordance with the G-PCC standard. As shown in FIG. 11, a G-PCC bitstream 1100 may carry a set of G-PCC units, also known as a type-length-value (TLV) encapsulation structure. As shown in 1110, (i.e., each) G-PCC TLV unit may include information indicating a TLV type 1111 and a G-PCC TLV unit payload 1112. Although not shown in FIG. 11, the GPCC TLV unit may further include information indicating a G-PCC TLV unit payload length, which may be expressed in terms of bytes or bits, for example. The G-PCC TLV unit payload 1112 may include information of a given type. For example, the G-PCC TLV unit payload may carry information of a given type, which may be, for example, a sequence parameter set, a geometry parameter set, a geometry data unit, an attribute parameter set, an attribute data unit, a tile inventory, a frame boundary marker, or a default attribute data unit.

[0113] 12 is a table providing an example syntax structure of a G-PCC TLV encapsulation unit, which may be defined, for example, according to the MPEG standard. As shown in FIG. 12, the TLV encapsulation unit may indicate a payload type using a first number of bits (or bytes), for example, 8 bits. The TLV encapsulation unit payload length may be represented by a second number of bits (e.g., 32 bits). The G-PCC TLV encapsulation unit may include a payload having the indicated payload type and payload length.

[0114] Figure 13 is a table providing possible values ​​of the TLV type parameter and a corresponding description of each of the possible values. As shown in Figure 13, the TLV payload type may be a sequence parameter set, a geometry parameter set, a geometry data unit, an attribute parameter set, an attribute data unit, a tile inventory, a frame boundary marker, or a default attribute data unit. G-PCC TLV units with unit types "2" and "4" may be a geometry data unit and an attribute data unit, respectively.

[0115] Figure 14 is a table providing an example syntax structure of a G-PCC TLV unit payload. The example syntax shown in Figure 14 may correspond to the syntax structure defined in MPEG-I Part 9 (ISO / IEC 23090-9), for example. The payload information of the geometry G-PCC unit and the attribute G-PCC unit may be decoded by a G-PCC decoder and may correspond to the media data units (e.g., TLV units) specified in the corresponding geometry G-PCC unit and the attribute parameter set G-PCC unit.

[0116] The high-level syntax (HLS) of a G-PCC file may support the concepts of slices and tile groups in geometry data and attribute data. A frame may be divided into multiple tiles and slices. A slice may be understood as a set of points that can be coded or decoded independently. A slice may, for example, contain one geometry data unit and zero or more attribute data units. Information in an attribute data unit may depend on corresponding information in geometry data units within the same slice. Within a slice, a geometry data unit may necessarily appear before its associated attribute unit. Data units of a slice may be consecutive. The ordering of slices within a frame need not necessarily be specified.

[0117] In some schemes, groups of slices may be identified by a common tile identifier. Consistent with some standards, a tile inventory may be provided that describes the bounding box of each tile. A tile may overlap another tile within the bounding box. Each slice may contain an index that identifies the tile to which the slice belongs.

[0118] This specification describes a G-PCC container file format. When a G-PCC bitstream is carried in a single track, it may require that the G-PCC encoded bitstream be represented by a single track declaration. Single-track encapsulation of G-PCC data may, in some cases, utilize simple ISOBMFF encapsulation, where the G-PCC bitstream is stored in a single track without further processing. Each sample within such a track may contain one or more G-PCC components. In other words, each sample may contain one or more TLV encapsulation structures.

[0119] 15 illustrates an exemplary sample structure in which bitstreams providing G-PCC geometry information and attribute information are stored in a single track. As shown in FIG. 15, a sample 1500 of a track carrying a G-PCC bitstream may include at least one of a first TLV 1510 providing a parameter set, a second TLV 1520 providing geometry data, and a third TLV 1530 providing attribute data corresponding to the geometry data of the second TLV 1520.

[0120] When one or more encoded G-PCC geometry bitstreams and one or more encoded G-PCC attribute bitstreams are stored in separate tracks, each sample in the track may contain at least one TLV encapsulation structure carrying a single G-PCC component data.

[0121] FIG. 16 shows an example structure of a multi-track ISOBMFF G-PCC container that may be implemented according to some standards, such as MPEG-I Part 18 (ISO / IEC 23090-18). A multi-track G-PCC container may contain information units known as "boxes," represented in FIG. 16 by ftyp, moov, and mdat structures 1610, 1620, and 1630, respectively, which may correspond to the base media file format defined in ISO / IEC 14496-12. The ftyp box 1610 may provide, for example, file type description information and common data structures used in media files. The moov box 1620 and the mdat box 1630 may contain G-PCC tracks 1621 and 1631 that together contain geometry bitstream samples carrying geometry parameter sets, sequence parameter sets, and geometry data TLV units. The tracks may also contain track references to other tracks carrying G-PCC attribute component payloads. The moov box 1620 and the mdat box 1630 may collectively contain G-PCC tracks 1622 and 1632 that may contain attribute parameter sets for the respective attributes, and attribute bitstream samples that carry attribute data TLV units.

[0122] When a G-PCC bitstream is carried in multiple tracks, the G-PCC component tracks may be linked using a track reference tool, which may be implemented according to several standards (e.g., ISO / IEC 14496-12). In some cases, one or more TrackReferenceTypeBoxes may be added to a TrackReferenceBox within a TrackBox of a G-PCC track. The TrackReferenceTypeBox may contain an array of track_IDs that specify the tracks to which the G-PCC track references. To link a G-PCC geometry track to a G-PCC attribute track, the reference_type of the TrackReferenceTypeBox of the G-PCC geometry track may identify the associated attribute track. The four-character code (4CC) associated with these track reference types may be "gpca," which may indicate that the referenced track contains an encoded bitstream of G-PCC attribute data.

[0123] When the geometry stream of a G-PCC bitstream contains multiple tiles, each tile or group of tiles may be encapsulated in a separate track, which may be called a geometry tile track. A geometry tile track may carry the TLV units of one or more geometry tiles, thus allowing direct access to these tiles. Similarly, an attribute stream of a G-PCC bitstream containing multiple tiles may also be carried in multiple attribute tile tracks.

[0124] Data for one or more G-PCC tiles may be carried in separate geometry and attribute tile tracks of the container. To support partial access in ISOBMFF containers for G-PCC coded streams, tiles corresponding to spatial regions within a point cloud scene may be signaled in samples of timed metadata tracks, such as a track with a Dynamic3DSpatialRegionSampleEntry, which may be defined consistent with some MPEG standards, or in a GPCCSpatialRegionInfoBox box, as may also be defined in some MPEG standards. This may enable players and streaming clients to retrieve a set of tile tracks that carry the information needed to render a particular spatial region or tile within a point cloud scene.

[0125] The G-PCC base track may carry, for example, a TLV encapsulation structure containing only SPS, GPS, APS, and tile inventory information, as described in ISO / IEC 23090-9. To link the G-PCC base track to the geometry tile track, a track reference with a new track reference type may be defined using 4CC "gpbt". The new type of track reference may be used to link the G-PCC base track to each geometry tile track.

[0126] Each geometry tile track can be linked to one or more other attributes of a G-PCC tile track that carries attribute information for the respective tile or tile group using a track reference tool, such as may be implemented in accordance with ISO / IEC 14496-12. These track reference types 4CC can be, for example, "gpca", as may be defined in accordance with the MPEG standard.

[0127] A point cloud scene may be coded in alternative forms. In such cases, the alternative forms of the coded G-PCC data may be indicated by an alternative track mechanism, such as may be implemented in accordance with ISO / IEC 14496-12. For example, the alternate_group field of the TrackHeaderBox may be used to indicate the alternatives for the coded G-PCC data. When each alternative G-PCC bitstream is stored in a single track, G-PCC tracks containing coded G-PCC bitstreams that may be alternatives to one another may have the same alternate_group value in their TrackHeaderBox. When each alternative G-PCC bitstream is stored in a multi-track container, i.e., when the different component bitstreams of each alternative G-PCC bitstream are carried in separate tracks, the G-PCC geometry tracks of the alternative G-PCC bitstreams may have the same alternate_group value in their TrackHeaderBox.

[0128] Methods, procedures, apparatus, and systems for MPEG Media Transport (MMT) are described herein. Generally speaking, a set of tools can be used to enable advanced media transport and distribution services. The tools may be distributed across three different functional areas: media processing unit (MPU) formatting, distribution, and signaling. While such tools may be designed to be used efficiently together, they may also be used independently.

[0129] The Media Processing Unit (MPU) functional area may define the logical structure of media content, the packages and formats of data units processed by MMT entities, and their instantiation using, for example, the ISO Base Media File Format. Packages may specify components containing media content and their relationships to provide the information necessary for advanced distribution. Data formats may be defined to encapsulate media data encoded for either storage or distribution, and to enable easy conversion between data to be stored and data to be distributed.

[0130] The distribution functional area may define an application layer transport protocol called MMT Protocol (MMTP) and a payload format. The application layer transport protocol may provide enhanced capabilities for the distribution of multimedia data, such as multiplexing and support for mixed use of streaming and download distribution in a single packet flow. The payload format may enable the transport of coded media data independent of media type and encoding method.

[0131] The signaling functional area may define the format of signaling messages for managing the distribution and consumption of media data. Signaling messages for consumption management may be used to signal the structure of packages, and signaling messages for distribution management may be used to signal the structure of payload formats and protocol configurations.

[0132] The MMT protocol may support multiplexing of different media data from various assets, such as media processing units (MPUs), via a single MMTP packet flow. It may deliver multiple types of data to a receiving entity in order of consumption to aid synchronization between different types of media data without introducing large delays or requiring large buffers. MMTP may also support multiplexing of media data and signaling messages within a single packet flow.

[0133] In some embodiments, an MMTP payload may be carried in only one MMTP packet. Fragmentation and aggregation may be provided by the payload format or may not be provided by the MMTP itself. MMTP may define two packetization modes: Generic File Delivery (GFD) mode and MPU mode. GFD mode may identify data units using their byte position within a transport object. MPU mode may identify data units using their role and media position within the MPU. The MMT protocol may support mixed use of packets with two different modes in a single delivery session. A single packet flow of MMT packets may optionally consist of two types of payload.

[0134] Figure 17 depicts an example end-to-end architecture of a system in which MMT signaling is implemented. The architecture may include at least, but not limited to, a package provider 1710, one or more asset providers 1721 and 1722, an MMT transmitting entity 1730, and an MMT receiving entity 1740. As shown in Figure 17, the MMT transmitting entity 1730 may receive a package from the package provider 1710. The MMT transmitting entity 1730 may be responsible for transmitting the package to the MMT receiving entity 1740 as an MMTP packet flow. The MMT transmitting entity 1730 may be requested to collect media content from content providers based on package presentation information provided by the package provider 1710. The media content may be provided as assets that are segmented into a series of encapsulated MMT processing units that form an MMTP packet flow. Thus, the MMT transmitting entity 1730 may collect asset information from one or more of the asset providers 1721 and / or 1722.

[0135] Signaling messages may be used to manage the delivery and consumption of packages. The interface between the MMT transmitting entity 1730 and the MMT receiving entity 1740, as well as their operation, may be standardized. The MMT protocol (MMTP) may be used by the MMT receiving entity 1740 to receive and demultiplex the streamed media based on packet_id and payload type. The decapsulation procedure performed by the MMT receiving entity 1740 may depend on the type of payload carried and may be processed separately, for example, in the scenario shown in FIG. 17.

[0136] Various aspects of the MMT data model are described herein. The MMT protocol may provide for both streaming and download delivery of coded media data. For streaming delivery, the MMT protocol may assume a specific data model that includes MPUs, assets, and packages. The MMT protocol may preserve the data model during delivery by using signaling messages to indicate the structural relationships between MPUs, assets, and packages.

[0137] A collection of encoded media data and its associated metadata may constitute a package. The package may be distributed from one or more MMT transmitting entities to one or more MMT receiving entities. One or more portions of the encoded media data of a package, such as a portion of audio or video content, may constitute an asset.

[0138] Assets may be associated with identifiers that may not depend on the actual physical location or service provider providing the asset, so that assets can be globally uniquely identified. Assets with different identifiers may not be interchangeable. For example, two different assets may carry two different encodings of the same content, but they may not be interchangeable. MMT cannot specify a specific identification mechanism, but may allow for the use of URIs or UUIDs for this purpose. Each asset may have its own timeline, which may have a different duration than the timeline of the overall presentation created by the package.

[0139] Each MPU may constitute a non-overlapping portion of an asset, i.e., two consecutive MPUs of the same asset may not contain the same media samples. Each MPU may be consumed independently by the presentation engine of an MMT receiving entity.

[0140] Figure 18 is an illustration of a package structure according to some embodiments. As shown in Figure 18, a package 1800 may be a logical entity. The package 1800 may contain one or more presentation information documents 1810, one or more assets 1820, and associated asset delivery characteristics (ADCs) for each asset. Each of the assets 1820 may contain one or more MPUs 1830. Processing of the package may be performed on a per-MPU basis, and each MPU may share the same asset ID.

[0141] MMT assets are further described herein according to some embodiments. An asset may be any multimedia data used to build a multimedia presentation. An asset may be a logical grouping of MPUs that share the same asset ID for carrying encoded media data. The encoded media data of an asset may be timed or untimed. Timed data may include encoded media data with an inherent timeline and may require synchronized decoding and presentation of data units at a specified time. Untimed data may include any other type of data that does not have an inherent timeline for the decoding and presentation of its media content. The decoding and presentation times of each item of untimed data may not necessarily be related to the decoding and presentation times of other items of the same untimed data. For example, they may be determined by user interaction or presentation information.

[0142] Two MPUs of the same asset carrying timed media data may have no overlap in their presentation time. Any type of data referenced by the presentation information may be considered an asset. Examples of types of media data that may be considered individual assets may include audio data, video data, or web page data.

[0143] The features and characteristics of a Media Processing Unit (MPU) are described herein. A Media Processing Unit (MPU) can be a media data item that can be processed by an MMT entity and consumed by a Presentation Engine independently of other MPUs.

[0144] Processing of an MPU by an MMT entity may include encapsulation / decapsulation and packetization / depacketization. An MPU may include an MMT hint track that indicates the boundaries of the MFU for media-aware packetization. Consumption of an MPU may include media processing (e.g., encoding / decoding) and presentation.

[0145] For packetization purposes, an MPU may be fragmented into data units that may be smaller than an access unit (AU). The syntax and semantics of an MPU may be independent of the type of media data carried in the MPU. An MPU of a single asset may have either timed or untimed media. An MPU may contain portions of data formatted according to one or more of several standards, such as MPEG-4 AVC (ISO / IEC 14496-10) or MPEG-2 TS.

[0146] A single MPU may contain an integer number of AUs or untimed data. In the case of timed data, a single AU may not be fragmented into multiple MPUs. In the case of untimed data, a single MPU may contain one or more untimed data items to be consumed by the Presentation Engine. MPUs may be identified by an associated asset identification (asset_id) and / or sequence number.

[0147] Aspects of MMTP payloads are described herein. The MMTP payload may be a generic payload used to packetize and convey media data such as MPUs, generic objects, and other information for consuming packages via the MMT protocol. An appropriate MMTP payload format may be used to packetize the MPUs, generic objects, and signaling messages.

[0148] An MMTP payload may carry a complete MPU or a fragment of an MPU, a signaling message, a generic object, repair symbols for the AL-FEC scheme, or other data units or structures. The type of payload may be indicated by a type field in the MMT protocol packet header. For each payload type, one or more data units for delivery and, additionally or alternatively, a type-specific payload header may be defined. For example, when an MMTP payload carries a fragment of an MPU, the fragment of the MPU (e.g., an MFU) may be considered as a single data unit. The MMT protocol may aggregate multiple data units of the same data type into a single MMTP payload. It may also fragment a single data unit into multiple MMTP packets.

[0149] An MFU may be a sample or sub-sample of timed data, or an item of untimed data. An MFU may contain media data that may be smaller than an AU for timed data, and the contained media data may be processed by a media decoder. An MFU may include an MFU header that contains information about the boundaries of the media data being carried. An MFU may contain an identifier to uniquely distinguish an MFU within an MPU. It may also provide dependency and priority information relative to other MFUs within the same MPU.

[0150] An MMTP payload may include a payload header and payload data. Some data types may allow fragmentation and aggregation, in which case a single data unit may be split into multiple pieces, or a set of data units may be delivered in a single MMTP packet.

[0151] In recent years, there has been considerable interest in new and emerging media types such as virtual reality (VR) and immersive video and 3D graphics. High-quality 3D point clouds have emerged in recent years as an advanced representation of immersive media, enabling new forms of interaction and communication with virtual worlds. The large amount of information required to represent such point clouds may require efficient coding algorithms. New standards for video-based point cloud compression are currently under development and will form the basis for visual volumetric video-based coding (V3C). Standards for geometry-based point cloud compression are also being developed and may define bitstreams for compressed static point clouds. In parallel, standards are also under development defining the transport of V3C media and geometry-based point cloud data.

[0152] While discussions surrounding V3C carriage and point cloud standards may address storage and signaling aspects of V3C data and point cloud data, such discussions may be limited in that they may relate only to signaling for dynamic adaptive streaming over HTTP based on the MPEG-DASH standard, for example. Another important candidate standard for enabling different streaming and distribution applications is MPEG Media Transport (MMT). However, the MMT standard may not currently provide a signaling mechanism for V3C media. Therefore, new signaling elements are desired that allow streaming clients to identify V3C streams and their component substreams. In addition, it may also be necessary to signal different types of metadata associated with V3C components to enable streaming clients to select the optimal version(s) of V3C content or its components that they can support or deliver given specific network constraints or a user's viewport at any given time.

[0153] Furthermore, it is assumed that practical point cloud applications will require streaming point cloud data over a network. Such applications may perform either live streaming or on-demand streaming of point cloud content, depending on how the content was generated. Due to the large amount of information required to represent a point cloud, such applications may need to support adaptive streaming techniques to avoid network overload and provide an optimal viewing experience at any given moment, e.g., with respect to the network capacity at that moment. Components of point cloud content may be divided into tiles. One or more streaming clients may (e.g., only) desire (e.g., decide or select) to stream specific tile portions of a geometry component (e.g., instead of the entire point cloud data), e.g., based on bandwidth availability. G-PCC component tile data may be encapsulated in different G-PCC tile tracks. (E.g., each) tile track may represent a set of G-PCC component tiles or a set of all G-PCC component tiles.

[0154] Currently, MMT does not provide a signaling mechanism for point cloud media, including point cloud streams based on the MPEG G-PCC standard. Therefore, it is important to define new signaling elements that allow streaming clients to identify point cloud streams and their component substreams. It is also necessary to signal different kinds of metadata associated with point cloud components, to allow streaming clients to select the optimal version of a point cloud or its components that they can support.

[0155] The solutions described herein may provide new signaling elements that allow an MMT streaming client to identify different components and metadata associated with V3C and GPCC media content and to select the media data the client needs to retrieve from the content server at any point during a streaming session. Additionally, the solutions described herein may provide various methods for encapsulation of G-PCC data for MMT streaming and the MMT signaling messages required to support delivery of G-PCC data via MMT.

[0156] MMT delivery of V3C content is further described herein. V3C content may assist MMT transmission entities in the streaming process. For example, presentation information may contain information describing a V3C-compliant MPU to enable proper processing by an application.

[0157] The player may receive information about the current viewing direction, the current viewport, and the display characteristics of the device on which the player is operating. Based on this information, view-dependent streaming may be used to reduce the bandwidth required in a streaming session. In the case of MMT, view-dependent streaming may be achieved by one or more techniques.

[0158] In some client-based streaming approaches, an MMT receiving entity may be instructed by a player to select a subset of assets that carry the V3C information needed to render the portion of the V3C content contained within (or intersecting with) the current viewport. MMT session control procedures may be used to request the selected set of assets from an MMT transmitting entity. The player may use V3C application-specific signaling messages from the server to select the appropriate assets to switch to for view-dependent streaming.

[0159] In some server-based approaches, the MMT receiving entity may rely on the MMT transmitting entity to select the correct subset of assets that provide V3C information for rendering the portion of the V3C content that covers the current viewport. The receiving entity may use V3C application-specific signaling to send information about the current viewport to the transmitting entity.

[0160] Methods and procedures for mapping V3C containers to MMT assets are described herein. To support the delivery of V3C content using MMT, each track in a multi-track ISOBMFF V3C container may be encapsulated as a separate asset. Thus, the number of assets may be equal to the number of tracks in the container. Assets that belong to the same V3C component may be logically grouped into asset groups. These asset groups may be signaled to a receiving entity to allow a streaming client to determine which asset group to request. V3C application-specific MMT signaling is described herein.

[0161] For the purpose of streaming V3C encoded data using MMT, several V3C-specific MMT messages are defined. For example, V3C application-specific signaling can include the transmission of group messages such as V3CAssetGroupMessage, selection messages such as V3CSelectionMessage, or view change feedback messages such as V3CViewChangeFeedbackMessage. In some embodiments, these messages may include an application identifier with, for example, the Uniform Resource Name (URN) "urn:mpeg:mmt:app:v3c:2020", which may allow the transmitting entity to associate the signaling with a V3C application.

[0162] Figure 19 is a table providing a list of defined application message types. In the proposed MMT V3C signaling, a set of application message types may be defined, and each message type in the set may be associated with an application message name, as shown in Figure 19. Via the V3C AssetGroupMessage, a sending entity may inform a client about a set of assets available at the server and provide a list of assets being streamed to a receiving entity. In the V3CSelectionMessage, a client may request that a set of assets be streamed by the sending entity to the receiving entity. In the V3CViewChangeFeedbackMessage, a client may send an indication of the user's current viewing direction and viewport to the server in a server-based view-dependent streaming session.

[0163] When transmitting V3C content over MMT, in some embodiments, the V3CAssetGroupMessage may be mandatory and may provide the receiving entity with a list of assets available at the server associated with the V3C content. This message may also be used to inform the receiving entity about which of these assets are currently being streamed to the receiving entity. From this list, a client running on the receiving entity may request a unique subset of these V3C assets using a V3CSelectionMessage message.

[0164] For view-dependent delivery of V3C content via MMT, a client may send its current viewport information to the server using a V3CViewChangeFeedbackMessage message, which may then select and deliver assets corresponding to that viewport to the client. The V3CAssetGroupMessage may also be used to update the client about a selected subset of assets. Figure 20 is a table providing an example syntax structure of a V3C asset descriptor. The asset descriptor may be used to inform receiving entities and consuming applications about the content of an asset carrying V3C content. Semantics of the V3C asset descriptor are provided herein. The descriptor tag, e.g., "Descriptor_tag," may indicate the type of the descriptor. The descriptor length, e.g., "Descriptor_length," may specify the length in bytes counting from the next byte after this field to the last byte of the descriptor. The data type, e.g., "Data_type," may indicate the type of V3C data present in this asset. Values ​​of this field are further shown in Figure 22 and may be introduced and substantially described in the following paragraphs. A dependency flag, e.g., "Dependency_flag," may indicate whether a V3C asset depends on data in another V3C asset for decoding. A value of 0 may indicate that this V3C component asset group data can be decoded independently. A value of 1 may indicate that this V3C asset depends on other V3C asset data for decoding. An alternate group flag, e.g., "Alternate_group_flag," may indicate whether this V3C asset has alternate versions. A value of 0 may indicate that this V3C component asset does not have any alternate assets. A value of 1 may indicate that this V3C asset has one or more alternates. An alternate group ID, e.g., "Alternate_group_id," may indicate an ID that identifies a group of alternate assets. Different encoded versions of the same V3C asset may have the same value for this field.Dependent asset ID, e.g., "Dep_asset_id", may indicate the value of the asset ID on which decoding of this asset depends. In some cases, this value may only be present when dependency_flag is set to 1. For example, a V3C video component asset may use the corresponding V3C atlas component asset ID for this field. "Num_tiles" may indicate the number of tiles carried in this asset. "tile_id" may indicate a unique identifier for a particular atlas tile.

[0165] FIG. 21 is a table illustrating an example of exemplary syntax for a V3CAssetGroupMessage. Consistent with the table of FIG. 21, the semantics of the V3CAssetGroupMessage may be described as follows: "Message_id" may indicate an identifier of the V3C application message. "Version" may indicate the version of the V3C application message. "Length" may indicate the length of the V3C application message in bytes, counting from the start of the next field to the last byte of the message. The value of this field may not be equal to 0. An application identifier, e.g., "Application_identifier," may indicate an application identifier as a URN that uniquely identifies the application that consumes the content of this message. "App_message_type" may indicate an application-specific message type, substantially as described above with respect to FIG. 19. "Num_V3C_asset_groups" may indicate the number of V3C asset groups, each group containing assets associated with a V3C component. "asset_group_id" may indicate an identifier of an asset group associated with a V3C component. "Num_assets" may indicate the number of assets in the asset group associated with the V3C component. "Start_time" may indicate the presentation time of the V3C component for which the asset states listed in this message are applicable. "Data_type" may indicate the type of V3C data present in this asset group. Example values ​​for this field may be described in the context of FIG. 22 and introduced and substantially described in the following paragraphs. "Pending_flag" may indicate whether all data components are ready for rendering of the asset group. For example, if set to "1", it may indicate that the data is ready; otherwise, the flag may be "0". "asset_id" may provide an asset identifier for the asset. "state_flag" may indicate the delivery state of the asset.When set to one ('1'), this may indicate that the transmitting entity is actively transmitting the asset to the receiving entity. When set to zero ('0'), this may indicate that the transmitting entity is not actively transmitting the asset to the receiving entity. The 'Sending_time_flag' may indicate the presence of a 'sending_time' for the first MMTP packet containing the first MPU of the asset stream. The default value may be '0'. The 'alternate_group_flag' may indicate whether this V3C component asset has an alternate version. A value of 0 may indicate that this V3C asset does not have any alternate assets. A value of 1 may indicate that this V3C asset has alternate assets. A dependency flag, for example, 'Dependency_flag', may indicate whether this V3C component asset depends on data in other V3C assets for decoding. A value of 0 may indicate that this V3C component asset group data can be decoded independently. A value of 1 may indicate that this V3C asset depends on other V3C asset data for decoding. The sending time, e.g., "Sending_time," may indicate the sending time for the first MMTP packet containing the first MPU of the asset stream. Using this information, the client may prepare a new packet processing pipeline for the new asset stream. "alternate_group_id" may indicate an identifier for an alternate V3C component asset. Different encoded versions of the same V3C asset may have the same value for this field. "Dep_asset_group_id" may indicate an ID for an asset on which decoding of this asset depends. In some cases, this value may only be present, for example, when dependency_flag is set to 1. For example, a V3C attribute component asset may use the V3C atlas component asset ID that corresponds to this field. "all_tiles_present_flag" may indicate whether all tiles of the atlas component are part of the asset.A value of 1 may indicate that data for all atlas tiles is available in the asset. A value of 0 may indicate that data for a subset of atlas tiles is available within the asset. "Num_tiles" may indicate the number of tiles carried in this asset. "tile_id" may provide a unique identifier for a particular atlas tile.

[0166]

[00130] Figure 22 is a table illustrating exemplary V3C data type values ​​that may be used in the Data_type field. As shown in Figure 22, the values ​​of the Data_type field may indicate all V3C component data, atlas component data, occupancy component data, geometry component data, attribute component data, codec initialization data, dynamic volumetric timed metadata information, or viewport timed metadata information.

[0167] FIG. 23 is a table illustrating an example syntax of the V3CSelectionMessage. Consistent with the table of FIG. 23, the semantics of the V3CSelectionMessage may be described as follows: "Message_id" may indicate an identifier of the V3C application message. "Version" may indicate the version of the V3C application message. "Length" may indicate the length of the V3C application message in bytes, e.g., counting from the beginning of the next field to the last byte of the message. The value of this field may not be equal to 0. "Application_identifier" may indicate an application identifier as a URN that uniquely identifies the application that consumes the content of this message. "App_message_type" may indicate an application-specific message type, substantially as described above in the paragraph above with respect to FIG. 19. "Num_selected_asset_groups" may indicate the number of asset groups for which there is an associated state change request by the receiving entity. "asset_group_id" may indicate the identifier of the asset group associated with the V3C content. "switching_mode" may indicate the switching mode used for asset selection requested by the receiving entity. A list of values ​​for "switching_mode" may be defined, for example, consistent with the following paragraphs introducing and describing Figure 23. "Num_assets" may indicate the number of assets signaled for state change due to the specified switching mode. "Asset_id" may indicate the identifier of the asset for state change due to the specified switching mode.

[0168] FIG. 24 is a table providing a definition of the switching_mode field. As shown in FIG. 24, the "switching_mode" field may indicate the switching mode used for asset selection. For example, if the switching mode is set to refresh, for each asset listed in the V3CSelectionMessage, the State_flag of each asset is set to "1", and the State_flag of all assets not listed in the V3CSelectionMessage is set to "0". If the switching mode is set to toggle, for each asset listed in the V3CSelectionMessage, the State_flag of each asset is changed, for example, if originally "0" then to "1", or if originally "1" then to "0", but the State_flag of all assets not listed in the V3CSelectionMessage remains unchanged. If the switching mode is set to send all for all assets in the asset group specified in the V3CSelectionMessage, the State_flag of each asset is set to "1".

[0169] FIG. 25 is a table illustrating an example syntax for the V3CViewChangeFeedbackMessage. Consistent with the table of FIG. 25, the semantics of the V3CViewChangeFeedbackMessage may be described as follows: "Message_id" may indicate an identifier for the V3C application message. "Version" may indicate the version of the V3C application message. "Length" may indicate the length of the V3C application message in bytes, counting from the start of the next field to the last byte of the message. The value of this field shall not be equal to 0. "Application_identifier" may indicate an application identifier as a URN that uniquely identifies the application that consumes the content of this message. "App_message_type" may indicate an application-specific message type, substantially as described above in the paragraph above with respect to FIG. 19. "Vp_pos_x", "vp_pos_y", and "vp_pos_z" may indicate the x, y, and z coordinates, respectively, of the position of the viewport in the global reference coordinate system in meters. The values ​​may be, for example, 2 -16 The coordinates may be provided in units of meters. "Vp_quat_x", "vp_quat_y", and "vp_quat_z" may indicate the x, y, and z components of the rotation of the viewport area, respectively, using quaternion representation. The coordinate values ​​may be floating-point values ​​in the range of -1 to 1, inclusive. These values ​​may specify the x, y, and z components, i.e., qX, qY, and qZ, of the rotation applied to transform the global coordinate axes into the camera's local coordinate axes, using quaternion representation. The fourth component of the quaternion qW may be generated according to Equation 1:

[0170]

number

[0171]

number

[0172] "clipping_near_plane" and "clipping_far_plane" may indicate the near and far depth (or distance) based on the viewport's near and far clipping planes in meters. "Horizontal_fov" may specify a longitude range corresponding to the horizontal size of the viewport area, for example, in radians. This value may be in the range of 0 to 2π. "vertical_fov" may specify a latitude range corresponding to the vertical size of the viewport area, for example, in radians. This value may be in the range of 0 to π.

[0173] Methods and apparatus for streaming client behavior are described herein. MMT clients may be guided by information provided in application-specific signaling messages. The following are examples of client behavior for streaming V3C content using the MMT signaling presented herein:

[0174] In some methods, an MMT transmitting entity may send a "V3CAssetGroupMessage" application message to interested clients. The receiving client may parse the "V3CAssetGroupMessage" application message to identify V3C media assets present at the MMT content transmitting entity. To identify available V3C media content, the streaming client may check the "application_identifier" field in the "V3CAssetGroupMessage" application message, which is set to "urn:mpeg:mmt:app:v3c:2020". All or part of the V3C assets available in the V3C content may be identified by checking the asset IDs signaled in the "V3CAssetGroupMessage" application message. The client may select the required assets to be streamed based on the user's current viewport. The MMT client may send a "V3CSelectionMessage" application message to the transmitting entity to request V3C assets of interest from the list of available V3C assets. The MMT transmitting entity may form MMTP packets using MTP and send the MTTP packets to the client.

[0175] In several ways, an MMT client may receive an MMTP packet and depacketize an MPU or MFU. The MPU / MFU may contain timed media content or non-timed V3C media content. If an MMT client receives an MMTP packet with the asset group "data_type" set to "0x05," the V3C asset data represents initialization information such as VPS, ASPS, AAPS, AFPS, and SEI messages. If an MMT client receives an MMTP packet with the asset group "data_type" set to "0x06," the V3C asset data may represent 3D spatial domain timed metadata information. The information in this asset may be used for partial access of V3C content. If an MMT client receives an MMTP packet with the asset group "data_type" set to "0x07," the V3C asset data may indicate initial or recommended viewport information. This information can be used to enable automatic viewport changes based on different criteria. The MMT client may select the required V3C asset based on, for example, the user's viewport or a recommended viewport and one or more corresponding 3D spatial regions. The MMT client may send a "V3CSelectionMessage" application message to the sending entity requesting the target V3C asset.

[0176] In some methods, when a user's viewport changes in a client-based streaming approach, the MMT client may request a different set of V3C assets using a "V3CSelectionMessage" application message. When a user's viewport changes in a server-based streaming approach, the MMT client may send a "V3CViewChangeFeedbackMessage" message to the transmitting entity to signal the user's current viewport. Upon receiving this message, the MMT transmitting entity selects a new set of V3C assets based on the user's new viewport information and sends a "V3CAssetGroupMessage" application message with the corresponding V3C assets to the MMT client. The MMT transmitting entity may stream the V3C asset data as MMTP packets. The MMT client may start receiving MMTP packets for all requested V3C assets and extract the MPUs and MFUs from the MMTP payload. The MPUs and MFUs may directly contain media samples or may contain media segments. The MMT client may begin parsing the media segment container (e.g., ISOBMFF) to extract elementary stream information and structure the V3C bitstream according to the V3C standard. The bitstream may be passed to a V3C decoder. If the MMTP payload contains V3C media samples, the elementary stream data is extracted and structured according to the V3C bitstream standard. The bitstream may be passed to a V3C decoder.

[0177]

[0003] Embodiments directed to encapsulating and signaling G-PCC data in MMT are described herein. Unlike traditional media content, G-PCC media content may include multiple components, such as geometry and attributes. Each component may be encoded separately as a substream of a G-PCC bitstream. Components such as geometry and attributes may be encoded using, for example, a G-GPCC encoder. However, these substreams may need to be decoded together with additional metadata to render a point cloud.

[0178] G-PCC encoded content may be distributed over a network using MMT. When a G-PCC component in ISOBMFF is signaled using multiple tracks, each track may be proposed to be encapsulated in a separate asset, which may then be packetized into MMTP packets in the usual way. A G-PCC-defined application message is also proposed to allow servers and clients to identify groups of multiple assets for a particular G-PCC component.

[0179] G-PCC media content may include one or more (e.g., multiple) components, such as geometry and attributes. The (e.g., each) component may be encoded separately as a substream of a G-PCC bitstream. Components such as geometry and attributes may be encoded using, for example, a G-GPCC encoder. The substreams may be decoded together with additional metadata, for example, to render a point cloud.

[0180] G-PCC data may be encapsulated and signaled in MMT. G-PCC encoded content may be distributed over a network using MMT. G-PCC data may be encapsulated for MMT streaming using various encapsulation methods (e.g., as described herein). MMT signaling messages may support (e.g., be generated and transmitted) distribution of G-PCC data via MMT.

[0181] A G-PCC component within an ISOBMFF may be signaled using multiple tracks. (E.g., each) track (e.g., of multiple tracks) may be encapsulated in a separate asset, and the separate asset may (e.g., then) be packetized into an MMTP packet. G-PCC-defined application messages may (e.g., also) be configured / deployed, for example, for servers and clients to identify groups of multiple assets to or for a particular G-PCC component.

[0182] In some examples (e.g., to support distribution of G-PCC content using MMT), (e.g., each) track in a multi-track ISOBMFF G-PCC container may be encapsulated in a separate asset. The number of assets may be equal to the number of tracks in the multi-track ISOBMFF G-PCC container. In some examples, multiple assets corresponding to a (e.g., single) G-PCC component may be grouped and signaled as an asset group in a message (e.g., a "GPCCAssetGroupMessage" application message). Alternative component tracks may be exposed in a message (e.g., using a "GPCCAssetGroupMessage" message), for example, to enable server and client selection decisions (e.g., efficiently) (e.g., without first parsing the ISOBMFF file in the MMTP packet).

[0183] MMT may define application-specific signaling messages, which may support (e.g., allow) the delivery of application-specific information. G-PCC-specific signaling messages may be defined (e.g., configured) to stream G-PCC-encoded data using MMT. The G-PCC-specific signaling messages may have an application identifier with a Uniform Resource Name (URN) value (e.g., a URN value of "urn:mpeg:mmt:app:gpcc:2020").

[0184] FIG. 26 is a table providing an example syntax structure of a G-PCC asset descriptor. The asset descriptor may be used to inform receiving entities and consuming applications about the content of an asset carrying G-PCC content. Semantics of the G-PCC asset descriptor are provided herein. 'descriptor_tag' may indicate the type of descriptor. 'Descriptor_length' may specify the length in bytes counting from the next byte after this field to the last byte of the descriptor. 'Data_type' may indicate the type of G-PCC data present in this asset group. Values ​​for this field are further shown in FIG. 29 and may be introduced and substantially described in the following paragraphs. 'Dependency_flag' may indicate whether the G-PCC asset depends on data in another G-PCC asset for decoding. A value of 0 may indicate that this G-PCC component asset group data can be independently decoded. A value of 1 may indicate that this G-PCC asset depends on other G-PCC asset data for decoding. "alternate_group_flag" may indicate whether this G-PCC asset has alternate versions. A value of 0 may indicate that this G-PCC component asset does not have any alternate assets. A value of 1 may indicate that this G-PCC asset has one or more alternates. "alternate_group_id" may indicate an ID that identifies a group of alternate assets. Different encoded versions of the same G-PCC asset may have the same value for this field. "Dep_asset_id" may indicate the value of the asset ID on which decoding of this asset depends. In some cases, this value may only be present when dependency_flag is set to 1. For example, a G-PCC attribute component asset may use the corresponding G-PCC geometry component asset ID for this field. "Num_tiles" may indicate the number of tiles carried in this asset. tile_id indicates a unique identifier for a specific tile in the tile inventory.When dynamic_tile_id_flag is set to the value 0, tile_id may represent one of the tile id values ​​present in the tile inventory.

[0185] MMT G-PCC signaling may include one or more of a set of (e.g., defined) application message types, such as group messages like GPCCAssetGroupMessage, selection feedback messages like GPCCSelectionMessageFeedback, and / or change view feedback messages like GPCCViewChangeFeedback.

[0186]

[00137] Figure 27 is a table illustrating example defined G-PCC application message types. As shown in Figure 27, the application message type may indicate that the message is a GPCCCAssetGroupMessage, a GPCCSelectionMessageFeedback message, or a GPCCViewChangeFeedback message. In an example of a GPCCAssetGroupMessage message type, a sending entity may send a group message (e.g., a GPCCCAssetGroupMessage message) to inform a client about a set of assets available at the server and / or a list of assets that can be streamed (e.g., are being streamed) to a receiving entity. In an example of a selection feedback message type (e.g., a GPCCSelectionMessageFeedback message type), a client may use the selection feedback message to request that a set of assets be streamed by the sending entity to a receiving entity. In an example of a change view feedback message (e.g., a GPCCViewChangeFeedback message), a client may use the view change feedback message to send an indication of the user's current viewing space to the server.

[0187] A group message (e.g., a GPCCCAssetGroupMessage message) may be used to transmit G-PCC encoded content via MMT. The group message (e.g., a GPCCCAssetGroupMessage message) may provide a client with a list of G-PCC data type assets available at the server and / or may inform the client about which of the assets can be streamed (e.g., are currently being streamed) to a receiving entity. A client may request a unique subset of G-PCC data type assets (e.g., from a list). The request may be made, for example, using a GPCCSelectionFeedback message.

[0188] The client may, for example, send current viewing space (e.g., frustum) information to the server using the GPCCViewChangeFeedback message (e.g., for view-dependent delivery of G-PCC content via MMT). The server may select and deliver to the client an asset corresponding to the viewing space. The GPCCAssetGroupMessage may (e.g., also) be updated and sent to the client. Table 4 provides examples of defined G-PCC application message types.

[0189] FIG. 28 is a table illustrating an example syntax for a group message such as GPCCCAssetGroupMessage. Consistent with the table of FIG. 28, the semantics of the GPCCCAssetGroupMessage may be as follows: "Message_id" may indicate an identifier of the G-PCC application message. "Version" may indicate the version of the G-PCC application message. "Length" may indicate the length of the G-PCC application message (e.g., in bytes counting from the beginning of the next field to the last byte of the message). The value of the length field may not be equal to zero (0). The application identifier (e.g., "application_identifier") may indicate an application identifier, for example, as a URN that (e.g., uniquely) identifies the type of application that consumes the content of the message. The application message type (e.g., "app_message_type") may define an application-specific message type (e.g., as given by the example in Table 4). The length of the application message type field may be, for example, 8 bits. The number of G-PCC asset groups (e.g., "num_gpcc_asset_groups") may indicate the number of G-PCC asset groups. An asset group (e.g., each) may contain assets associated with a G-PCC component. The asset group identifier (e.g., "asset_group_id") may indicate an identifier for an asset group associated with a G-PCC component. The number of assets (e.g., "num_assets") may indicate the number of assets in an asset group associated with a G-PCC component. The start time (e.g., "start_time") may indicate the presentation time of the G-PCC component at which the asset states listed in the message may be applicable. The data type (e.g., "data_type") may indicate the type of G-PCC point cloud data present in the asset group, further described in the following paragraph with respect to FIG. 29.A pending flag (e.g., "pending_flag") may indicate, for example, whether (e.g., all) data components are ready to render for an asset group. A pending flag set to "1" may indicate that the data is ready. A pending flag set to zero ("0") may indicate that the data is not ready. A dependency flag (e.g., "dependency_flag") may indicate whether a G-PCC component asset group depends on other G-PCC component asset group data for decoding. A value of zero ("0") may indicate that the G-PCC component asset group data can be decoded independently. A value of one ("1") may indicate that the G-PCC component asset group depends on other G-PCC component asset group data for decoding. A dependency asset group ID (e.g., "dep_asset_group_id") may indicate the value of the asset group ID on which asset group content decoding depends. A value may be present, for example, if / only if dependency_flag is set to 1. For example, a G-PCC attribute component asset group may use the corresponding G-PCC geometry component asset group ID for the dependent asset group ID field. The asset ID (e.g., "asset_id") may provide an asset identifier for the asset. The alternate asset group flag (e.g., "alternate_asset_group_flag") may indicate whether the G-PCC component asset has alternate versions. A value of 0 ("0") may indicate that the G-PCC component asset does not have alternate versions. A value of 1 ("1") may indicate that the G-PCC component asset has alternate versions. The value of the alternate group flag field may be set to 1 ("1"), for example, if / when different encoded versions of the same G-PCC component and / or asset are available in the bitstream.The value of the alternate group flag field may be set to zero ('0'), for example, if / when different encoded versions of the same G-PCC component and / or asset are not available in the bitstream. The alternate asset group ID (e.g., 'alternate_asset_group_id') may indicate an alternate G-PCC component asset value (e.g., a unique value). Different encoded versions of a G-PCC component or asset may represent the same value for the alternate asset group ID field. The state flag (e.g., 'state_flag') may indicate the delivery state of the asset. A state flag set to one ('1') may indicate that the transmitting entity is actively sending the asset to the receiving entity. A state flag set to zero ('0') may indicate that the transmitting entity is not actively sending the asset to the receiving entity. The sending time flag (e.g., 'sending_time_flag') may indicate the presence of a sending time (e.g., sending_time) for the first MMTP packet containing the first MPU of the asset stream. The default value may, for example, be zero ('0'). The sending time (e.g., "sending_time") may indicate the sending time of the first MMTP packet containing the first MPU of the asset stream. The client may prepare a new packet processing pipeline for the new asset stream (e.g., using the sending time information). The dynamic tile flag (e.g., "dynamic_tile_flag") may indicate whether the number of tiles and / or tile identifiers can change dynamically within the asset. A value of 0 ("0") may indicate that the number of tiles and tile identifiers in the asset do not change throughout the bitstream and / or that the number of tiles (e.g., "num_tiles") and tile IDs (e.g., "tile_id") are signaled. A value of 1 ("1") may indicate the number of tiles, and the tile identifiers may change within the asset. A value of 1 ("1") may indicate that the tile IDs present in the tile track are changing dynamically over time in the bitstream.The number of tiles (e.g., "num_tiles") may indicate the number of tiles carried in the asset. The tile ID (e.g., "tile_id") may indicate a (e.g., unique) identifier for a particular tile in the tile inventory. The tile ID (e.g., "tile_id") may represent a tile id value (e.g., one of the tile id values) present in the tile inventory, for example, if / when a dynamic tile flag (e.g., "dynamic_tile_flag") is set to a value of 0 ("0").

[0190] Figure 29 is a table showing exemplary G-PCC data type values ​​that may be used in the Data_type field. As shown in Figure 24, the values ​​of the Data_type field may indicate all G-PCC component data, geometry data, attribute data, SPS, GPS, APS, and tile inventory data, or 3D spatial domain timed metadata information.

[0191] FIG. 30 is a table illustrating an example syntax of a GPCC selection feedback message (e.g., "GPCCSelectionFeedback"). Consistent with the table of FIG. 30, the semantics of the GPCCSelectionFeedback message may be as follows: The message ID (e.g., "message_id") may indicate an identifier of the G-PCC application message. The version (e.g., "version") may indicate the version of the G-PCC application message. The length (e.g., "length") may indicate the length of the G-PCC application message (e.g., in bytes counting from the beginning of the next field to the last byte of the message). The value of the length field may not be equal to 0. The application identifier (e.g., "application_identifier") may indicate an application identifier, e.g., as a URN that (e.g., uniquely) identifies the type of application that consumes the content of the message. The application message type (e.g., "app_message_type") may define an application-specific message type (e.g., substantially as described in the paragraph above with respect to FIG. 27). The length of the application message type field may be, e.g., 8 bits. The number of selected asset groups (e.g., "num_selected_asset_groups") may indicate the number of asset groups for which there is an associated state change request by the receiving entity. The asset group ID (e.g., "asset_group_id") may indicate an identifier for an asset group associated with the G-PCC content. The switching mode (e.g., "switching_mode") may indicate the switching mode used for asset selection (e.g., as requested by the receiving entity). The number of assets (e.g., "num_assets") may indicate the number of assets signaled for state change (e.g., according to the specified switching mode).The asset ID (eg, "asset_id") may indicate an identifier of the asset for the state change (eg, according to a specified switching mode).

[0192] FIG. 31 is a table illustrating a definition of the switching_mode field. As shown in FIG. 31, the "switching_mode" field may indicate the switching mode used for asset selection. For example, if the switching mode is set to refresh, for each asset listed in GPCCSelectionMessageFeedback, the State_flag of each asset is set to "1," and the State_flag of all assets not listed in GPCCSelectionMessageFeedback is set to "0." If the switching mode is set to toggle, for each asset listed in GPCCSelectionMessageFeedback, the State_flag of each asset is changed, for example, if originally "0" then to "1," and if originally "1" then to "0," but the State_flag of all assets not listed in GPCCSelectionMessageFeedback remains unchanged. If the switching mode is set to send all for all assets in the asset group specified in GPCCSelectionMessageFeedback, the State_flag of each asset is set to "1."

[0193] FIG. 32 is a table illustrating an example syntax of a G-PCC view change feedback message (e.g., "GPCCViewChangeFeedback"). Consistent with the table of FIG. 32, the semantics of the GPCCViewChangeFeedback message may be as follows: The message ID (e.g., "message_id") may indicate an identifier of the G-PCC application message. The version may indicate the version of the G-PCC application message. The length may indicate the length of the G-PCC application message (e.g., in bytes counting from the beginning of the next field to the last byte of the message). The value of the length field may not be equal to 0. The application identifier (e.g., "application_identifier") may indicate an application identifier, for example, as a URN that (e.g., uniquely) identifies the type of application that consumes the content of the message. The application message type (e.g., "app_message_type") may define an application-specific message type (e.g., as given by the example in Table 4). The length of the application message type field may be, for example, 8 bits. The viewport position coordinates (e.g., vp_pos_x, vp_pos_y, vp_pos_z) may indicate the x, y, and z coordinates of the viewport's position in the global reference coordinate system in meters. Values ​​may be, for example, 2 -16The units may be meters. The viewport rotation (e.g., vp_quat_x, vp_quat_y, vp_quat_z) may indicate the x, y, and z components of the rotation of the viewport area (e.g., using a quaternion representation). The values ​​may be, for example, floating-point values ​​in the range of −1 to 1, inclusive. The values ​​may specify the x, y, and z components (e.g., qX, qY, and qZ) of a rotation to be applied to transform the global coordinate axes into the local coordinate axes of the camera (e.g., using a quaternion representation). The fourth component of the quaternion qW may be calculated, for example, according to Equation 1, substantially as set forth in the paragraph above. The point (w, x, y, z) may represent a rotation about the axis directed by the vector (x, y, z) by an angle determined according to Equation 2, also substantially as set forth in the paragraph above.

[0194] Clipping at the near plane (e.g., clipping_near_plane) and clipping at the far plane (e.g., clipping_far_plane) may indicate near and far depths or distances, for example, based on the near and far clipping planes of the viewport (e.g., in meters).

[0195] The horizontal field of view (FOV) (e.g., horizontal_fov) may specify the longitude range (e.g., in radians) that corresponds to the horizontal size of the viewport area. This value may be in the range 0 to 2π.

[0196] The vertical FOV (e.g., vertical_fov) may specify a latitudinal range (e.g., in radians) that corresponds to the vertical size of the viewport area. This value may be in the range 0 to π.

[0197] Streaming client behavior may be provided (e.g., defined or configured). MMT clients may be guided, for example, by information provided in application-specific signaling messages. Example client behaviors are provided for streaming geometry-based point cloud compressed content (e.g., using the example MMT signaling disclosed herein).

[0198] The MMT transmitting entity may send a "GPCCAssetGroupMessage" application message to interested clients. The receiving client may parse the "GPCCAssetGroupMessage" application message to identify the G-PCC media assets present at the MMT content transmitting entity. The streaming client may, for example, check the "application_identifier" field (e.g., set to "urn:mpeg:mmt:app:gpcc:2020") in the "GPCCAssetGroupMessage" application message to identify available G-PCC media content. Available G-PCC assets (e.g., all G-PCC assets) in the G-PCC point cloud content may be identified, for example, by checking the asset_id present in the "GPCCAssetGroupMessage" application message. The client may choose (e.g., select) the asset_id to be streamed, for example, based on the user's current viewport. The MMT client may send a "GPCCSelectionFeedback" application message to the transmitting entity requesting G-PCC assets of interest from the list of available G-PCC assets. The MMT transmitting entity may form MMTP packets using MTP. The MMT transmitting entity may send the MTTP packets to a client. The MMT client may receive the MMTP packets. The MMT client may depacketize the MPU or MFU. The MPU / MFU may contain timed or untimed G-PCC media content.

[0199] The G-PCC asset data may represent, for example, initialization information (e.g., SPS, GPS, APS, and / or tile inventory) when an MMT client receives an MMTP packet with the asset group "data_type" set to "3". The G-PCC asset data may represent, for example, 3D spatial domain timed metadata information when an MMT client receives an MMTP packet with the asset group "data_type" set to "4". The G-PCC asset information may be used for partial access of the G-PCC data.

[0200] The MMT client may select G-PCC assets based on the user viewport and the corresponding 3D spatial region. The MMT client may send a "GPCCSelectionFeedback" application message to the sending entity requesting the G-PCC assets of interest. The MMT client may request a different set of G-PCC assets (e.g., using the "GPCCSelectionFeedback" application message) if / when, for example, the user viewport changes.

[0201] For example, the MMT client may send a "GPCCViewChangeFeedback" message to the transmitting entity (e.g., to signal the user's current viewport) if / when the user viewport changes. The MMT transmitting entity (e.g., upon receiving the message from the MMT client) may select a G-PCC asset (e.g., based on the user's new viewport information). The MMT transmitting entity may send a "GPCCAssetGroupMessage" application message with the corresponding G-PCC asset to the MMT client. The MMT transmitting entity may stream the G-PCC asset data as MMTP packets.

[0202] The MMT client may start receiving MMTP packets for the requested G-PCC assets (e.g., all). The MMT client may extract the MPU and MFU from the MMTP payload. The MPU and MFU may contain media samples (e.g., directly) or media segments.

[0203] The MMT client may start parsing the media segment container (e.g., ISOBMFF) to extract elementary stream information, structure the G-PCC bitstream, and pass the bitstream to the G-PCC decoder. For example, if / when the MMTP payload contains G-PCC media samples, the elementary stream data may be extracted and structured, and the bitstream may be passed to the G-PCC decoder.

[0204] Systems, methods, and apparatus for MPEG Media Transport (MMT) streaming of geometry-based point clouds (G-PCCs) are described herein. G-PCC-encoded content may be distributed over a network using MMT. G-PCC data may be encapsulated for MMT streaming. MMT signaling messages may support the distribution of G-PCC data via MMT. (e.g., each) track may be encapsulated into a separate asset that can be packetized into an MMTP packet, for example, when a G-PCC component in the International Standards Organization Base Media File Format (ISOBMFF) is signaled using multiple tracks. G-PCC definition application messages may enable servers and clients to identify groups of multiple assets for a G-PCC component.

[0205] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. 1. A method implemented in a receiving device for streaming Motion Picture Experts Group (MPEG) Media Transport Protocol (MMTP) media content, comprising: receiving an asset group message from a transmitting device including asset descriptor data describing one or more asset groups available to be streamed, the asset descriptor data including a field indicating a respective data type associated with each asset of the one or more asset groups; transmitting an asset selection message to the transmitting device, the asset selection message including a request for at least a subset of assets of the one or more asset groups that are available to be streamed, the asset selection message including at least one identifier associated with at least one of the requested subset of assets; receiving one or more MMTP packets from the transmitting device in response to the asset selection message; and processing the one or more MMTP packets to recover at least a portion of the requested subset of assets of the one or more asset groups.

2. 10. The method of claim 1, further comprising: sending a viewport change message to the transmitting device including an indication of a current viewport of the receiving device; and receiving another asset group message including updated asset descriptor data describing one or more asset groups that are available to be streamed based on the current viewport of the receiving device.

3. 10. The method of claim 1, wherein the requested at least the subset of assets of the one or more asset groups that are available to be streamed is selected by the receiving device based on a current viewport of the receiving device.

4. The method of claim 1 , wherein the asset descriptor data includes a unique identifier associated with each of the assets of the one or more asset groups.

5. The method of claim 1 , wherein the transmitted asset selection message includes information identifying an application intended to consume the requested subset of assets.

6. The method of claim 1 , wherein the asset descriptor data describing the one or more asset groups available to be streamed describes Volumetric Video Based Coding (V3C) data.

7. 7. The method of claim 6, wherein the respective data type associated with each asset in the one or more asset groups is one of atlas component data, occupancy component data, geometry component data, attribute component data, dynamic volumetric timed metadata information, or viewport timed metadata information.

8. The method of claim 1 , wherein the asset descriptor data describing the one or more asset groups available to be streamed describes geometry-based point cloud compressed (G-PCC) data.

9. 9. The method of claim 8, wherein the respective data type associated with each asset of the one or more asset groups is one of geometry data, attribute data, attribute parameter set, sequence parameter set, geometry parameter set, tile inventory data, frame boundary marker data, default data, or three-dimensional spatial domain timed metadata information.

10. 10. The method of claim 1, wherein the asset group message includes information indicating one or more of: a dependency of an asset on another asset for decoding; an indication of the other asset on which the asset depends; whether the asset has an alternative version; and an identification of the alternative version of the asset.

11. A receiving device configured to stream MPEG (Motion Picture Experts Group) MMTP (Media Transport Protocol) media content. a processor; a communication interface; the processor and the communication interface are configured to receive from a transmitting device an asset group message including asset descriptor data describing one or more asset groups available to be streamed, the asset descriptor data including a field indicating a respective data type associated with each asset of the one or more asset groups; the processor and the communication interface are configured to transmit to the transmitting device an asset selection message including a request for at least a subset of assets of the one or more asset groups that are available to be streamed, the asset selection message including at least one identifier associated with at least one of the requested subset of assets; the processor and the communication interface are configured to receive one or more MMTP packets from the transmitting device in response to the asset selection message; The processor processes the one or more MMTP packets to recover at least a portion of the requested subset of assets of the one or more asset groups.

12. 12. The receiving device of claim 11, wherein the processor and the communication interface are configured to send a viewport change message to the transmitting device including an indication of the current viewport of the receiving device and to receive another asset group message including updated asset descriptor data describing one or more asset groups that are available to be streamed based on the current viewport of the receiving device.

13. 12. The receiving device of claim 11, wherein the requested at least the subset of assets of the one or more asset groups that are available to be streamed is selected by the receiving device based on a current viewport of the receiving device.

14. The receiving device of claim 11 , wherein the asset descriptor data includes a unique identifier associated with each of the assets of the one or more asset groups.

15. The receiving device of claim 11 , wherein the transmitted asset selection message includes information identifying an application intended to consume the requested subset of assets.

16. The receiving device of claim 11 , wherein the asset descriptor data describing the one or more asset groups available to be streamed describes Volumetric Video Based Coding (V3C) data.

17. 17. The receiving device of claim 16, wherein the respective data type associated with each asset of the one or more asset groups is one of atlas component data, occupancy component data, geometry component data, attribute component data, dynamic volumetric timed metadata information, or viewport timed metadata information.

18. The receiving device of claim 11 , wherein the asset descriptor data describing the one or more asset groups available to be streamed describes geometry-based point cloud compression (G-PCC) data.

19. 20. The receiving device of claim 18, wherein the respective data type associated with each asset of the one or more asset groups is one of geometry data, attribute data, attribute parameter set data, sequence parameter set data, geometry parameter set data, tile inventory data, frame boundary marker data, default data, or three-dimensional spatial domain timed metadata information.

20. 12. The receiving device of claim 11, wherein the asset group message includes one or more of an asset's dependency on another asset for decoding, an indication of the other asset on which the asset depends, whether the asset has an alternative version, and an identification of the alternative version of the asset.

Citation Information

Patent Citations

  • Systems and methods for network-based media processing

    US20190028691A1

  • An apparatus, a method and a computer program for video coding and decoding

    WO2020008106A1

  • Information processing device and information processing method

    WO2020137642A1

  • Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

    WO2020189895A1

  • IMMERSIVE VIDEO CODING TECHNIQUES FOR THREE DEGREE OF FREEDOM PLUS / METADATA FOR IMMERSIVE VIDEO (3DoF+ / MIV) AND VIDEO-POINT CLOUD CODING (V-PCC)

    WO2020232281A1