Signaling parameter sets for geometry-based point cloud streams
By receiving and processing the point cloud compressed data in the ISOBMFF file in the video decoding system and utilizing the sample group description information and sample-to-group box information to group the G-PCC samples into multiple tracks, the problem of low point cloud data compression efficiency is solved and a more efficient video decoding system is achieved.
Patent Information
- Application Number
- CN202510870310.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-28
- Filing Date
- 2023-06-27
- Publication Date
- 2025-09-12
AI Technical Summary
Existing video decoding systems are inefficient in compressing and transmitting point cloud data, especially the representation and compression mechanism of geometry-based point cloud compression (G-PCC) data.
By receiving and processing the geometry-based point cloud compression data in the ISOBMFF file, the G-PCC samples are grouped into multiple tracks using the sample group description information and the sample-to-group box information, and grouped according to the parameter set information to ensure that the samples in each track have the same parameter set information.
It improves the compression and transmission efficiency of point cloud data, optimizes the performance of the video decoding system, and adapts to the changing requirements of parameter sets in different time periods.
Smart Images

Figure CN120640010A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202380054229.0, filed on June 27, 2023, and entitled “Parameter Set for Geometry-Based Point Cloud Streams Using Signals.”
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 356,335, filed on June 28, 2022, the disclosure of which is incorporated herein by reference in its entirety. Background Art
[0004] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, wavelet-based systems, object-based systems, and / or block-based systems (such as hybrid block-based video coding systems). Representation and / or compression mechanisms used to store and / or transmit point cloud data may be inefficient. Summary of the Invention
[0005] Systems, methods, and tools for signaling parameter sets for geometry-based point cloud streaming are disclosed.
[0006] An example device may receive an International Organization for Standardization Base Media File Format (ISOBMFF) file. The ISOBMFF file may include geometry-based point cloud compression (G-PCC) data carried using one or more tracks, the one or more tracks having sample group description information and sample to group box information. The sample group description information may indicate grouping type information and multiple sample group description entries, the multiple sample group description entries indicating geometry-based volume or point cloud parameter set information. The sample to group box information may include one or more sample to group box entries, each sample to group box entry having multiple entry parameters, the multiple entry parameters including: the grouping type information, the grouping type parameter, and a sample group description index. The entry parameters may also include an entry count and / or a sample count.
[0007] In an example, a device may receive an indication of a geometry-based point cloud compression (G-PCC) sample and G-PCC information. The G-PCC sample may be associated with a plurality of tracks. The G-PCC information may include parameter set information, and the parameter set information may include information associated with a sample set. The device may determine that the parameter set information has changed from a first time to a second time. Based on the change in the parameter set information, the device may group the G-PCC sample into each track of the plurality of tracks. The grouped G-PCC samples may include the same parameter set information.
[0008] Each feature disclosed anywhere herein is described and may be implemented separately / individually and in any combination with any other feature disclosed herein and / or with any feature disclosed elsewhere that may be implicitly or explicitly mentioned herein or that may otherwise fall within the scope of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1A is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented.
[0010] Figure 1B is an example of a method that can be used according to an embodiment of the present invention. Figure 1A A system diagram of an example wireless transmit / receive unit (WTRU) for use within an example communication system.
[0011] Figure 1C is an example of a method that can be used according to an embodiment of the present invention. Figure 1A System diagram of an example radio access network (RAN) and an example core network (CN) for use within an illustrated communication system.
[0012] Figure 1D is an example of a method that can be used according to an embodiment of the present invention. Figure 1A System diagram of another example RAN and another example CN for use within the illustrated communication system.
[0013] Figure 2 is a diagram illustrating an example video encoder.
[0014] Figure 3 is a diagram illustrating an example of a video decoder.
[0015] Figure 4 is a diagram illustrating an example of a system in which various aspects and examples may be implemented.
[0016] Figure 5 An example of a geometry-based point cloud compression (G-PCC) bitstream structure is shown.
[0017] Figure 6 An example of a sample structure when a decoded G-PCC bitstream is stored in a single track is shown.
[0018] Figure 7 An example of a multi-track G-PCC bitstream container structure is shown.
[0019] Figure 8 An example of a G-PCC bitstream container structure with G-PCC entries and tile entries is shown.
[0020] Figure 9An example of a G-PCC bitstream container structure with a G-PCC item, a tile item, and a spatial region item is shown.
[0021] Figure 10 An example of temporal levels in a G-PCC sequence is shown.
[0022] Figure 11 An example of using the 'gpsg' sample group in multi-track and time-level tracks is shown.
[0023] Figure 12 An example of using the 'gpsg' sample group in a multi-component track is shown. DETAILED DESCRIPTION
[0024] A detailed description of exemplary embodiments will now be described with reference to the various drawings.While this specification provides detailed examples of possible implementations, it should be noted that the details are intended to be illustrative and in no way limit the scope of the application.
[0025] Figure 1A is a diagram illustrating an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multi-carrier (FBMC), etc.
[0026] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112. However, it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in an industrial and / or automated process chain environment), a consumer electronic device, a device operating on a commercial and / or industrial wireless network, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0027] The communication system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as the CNs 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be base transceiver stations (BTSs), Node-Bs, eNode-Bs, Home Node-Bs, Home eNode-Bs, gNBs, NR Node-Bs, site controllers, access points (APs), wireless routers, and the like. While the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0028] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage to a specific geographic area, which may be relatively fixed or may vary over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, one for each sector of the cell. In an embodiment, base station 114a may employ multiple-input, multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0029] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0030] More specifically, as noted above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a in the RAN 104 / 113 and the WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may use Wideband CDMA (WCDMA) to establish the air interface 115 / 116 / 117. WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0031] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0032] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may establish the air interface 116 using New Radio (NR).
[0033] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, for example, using the principle of dual connectivity (DC). Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions transmitted to / from multiple types of base stations (e.g., eNBs and gNBs).
[0034] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.
[0035] Figure 1AThe base station 114b in may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, and the like. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or a femtocell. As Figure 1A As shown, base station 114b may have a direct connection to the Internet 110. Therefore, base station 114b may not need to access the Internet 110 via CN 106 / 115.
[0036] The RAN 104 / 113 may be in communication with the CN 106 / 115, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have different quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not described in detail in the text, the CN 106 / 115 may be used to provide a variety of services, such as voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. Figure 1A Although not shown in the figures, it will be appreciated that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, the CN 106 / 115 may also be in communication with another RAN (not shown) that employs GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0037] The CN 106 / 115 may also act as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet protocol suite. The networks 112 may include wired communication networks and / or wireless communication networks owned and / or operated by other service providers. For example, the networks 112 may include another CN connected to one or more RANs, which may employ the same RAT as the RAN 104 / 113 or a different RAT.
[0038] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communication system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). Figure 1A The illustrated WTRU 102c may be configured to communicate with the base station 114a, which may employ a cellular-based radio technology, and with the base station 114b, which may employ an IEEE 802 radio technology.
[0039] Figure 1B is a system diagram illustrating an example WTRU 102. Figure 1B As shown, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be appreciated that the WTRU 102 may include any subcombination of the foregoing elements while remaining consistent with an embodiment.
[0040] The processor 118 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal decoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. Although Figure 1B The processor 118 and the transceiver 120 are depicted as separate components, but it is understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0041] The transmit / receive element 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via an air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It should be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0042] Although the transmit / receive element 122 Figure 1B Although depicted as a single element in FIG. 1 , the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0043] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As noted above, the WTRU 102 may have multi-mode capabilities. For example, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0044] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, or the like. In other embodiments, the processor 118 may access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0045] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control power to the other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0046] The processor 118 may also be coupled to the GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or in lieu of the information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may acquire location information by any suitable location-determination method while remaining consistent with an embodiment.
[0047] The processor 118 may also be coupled to other peripherals 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripherals 138 may include one or more sensors, which may be one or more of the following: a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor; a geo-location sensor; an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0048] The WTRU 102 may include a full-duplex radio for which transmission and reception of some or all of the signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference via signal processing performed via hardware (e.g., a choke) or via a processor (e.g., a separate processor (not shown) or via the processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio for which transmission and reception of some or all of the signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.
[0049] Figure 1C 1 is a system diagram illustrating the RAN 104 and the CN 106 in accordance with an embodiment. As noted above, the RAN 104 may employ E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 may also be in communication with the CN 106.
[0050] The RAN 104 may include eNode-Bs 160a, 160b, 160c, though it will be appreciated that the RAN 104 may include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, 160c may implement MIMO technology. Thus, the eNode-B 160a, for example, may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 102a.
[0051] Each of the eNodeBs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. Figure 1C As shown, the eNode-Bs 160a, 160b, 160c may communicate with one another via an X2 interface.
[0052] Figure 1C The illustrated CN 106 may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the foregoing elements is depicted as part of the CN 106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0053] The MME 162 may be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and may serve as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during an initial attach of the WTRUs 102a, 102b, 102c, and the like. The MME 162 may also provide control plane functions for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.
[0054] The SGW 164 may be connected to each of the eNode-Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. The SGW 164 may generally route and forward user data packets to and from the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions, such as anchoring the user plane during inter-eNode-B handovers, triggering paging when downlink data is available for the WTRUs 102a, 102b, 102c, managing and storing the context of the WTRUs 102a, 102b, 102c, and the like.
[0055] The SGW 164 may be connected to the PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0056] The CN 106 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional land-line communications devices. For example, the CN 106 may include, or may be in communication with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 106 and the PSTN 108. The CN 106 may also provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired networks and / or wireless networks owned and / or operated by other service providers.
[0057] Even though the WTRU Figures 1A to 1D Although described as wireless terminals, it is contemplated that in certain representative embodiments such terminals may (eg, temporarily or permanently) employ a wired communications interface with a communications network.
[0058] In a representative embodiment, the other network 112 may be a WLAN.
[0059] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or out of the BSS. Traffic originating from outside the BSS and destined for a STA can reach the AP and be delivered to the STA. Traffic originating from a STA and destined for a destination outside the BSS can be transferred to the AP for delivery to the destination. Traffic between STAs within the BSS can be transferred through the AP, for example, where a source STA can transfer traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as point-to-point traffic. Point-to-point traffic can be transferred between a source STA and a destination STA (e.g., directly between them) using direct link setup (DLS). In certain representative embodiments, the DLS can use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (eg, all STAs in each STA) may communicate directly with each other. The IBSS communication mode may sometimes be referred to herein as an "ad hoc" communication mode.
[0060] When using the 802.11ac infrastructure operating mode or a similar operating mode, the AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, for example, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented in an 802.11 system. With CSMA / CA, STAs (e.g., each STA) (including the AP) may sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or determined to be busy, the particular STA may back off. One STA (e.g., only one station) may transmit in a given BSS at any given time.
[0061] High throughput (HT) STAs may communicate using a 40 MHz wide channel, for example, via a primary 20 MHz channel combined with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.
[0062] Very high throughput (VHT) STAs can support 20MHz, 40MHz, 80MHz, and / or 160MHz wide channels. 40MHz channels and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels, or by combining two non-contiguous 80MHz channels (this may be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, the data can pass through a segment parser that can separate the data into two streams. Each stream can be individually processed using an inverse fast Fourier transform (IFFT) and time domain processing. These streams can be mapped to two 80MHz channels, and the data can be sent by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration can be reversed, and the combined data can be transmitted to the medium access control (MAC).
[0063] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah relative to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah may support meter type control / machine type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support for (e.g., only support for) certain bandwidths and / or limited bandwidths. MTC devices may include batteries with battery life above a threshold (e.g., to maintain very long battery life).
[0064] WLAN systems that support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as a primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by a STA (that supports the minimum bandwidth operating mode) from among all STAs operating in the BSS. In the example of 802.11ah, for a STA (e.g., an MTC-type device) that supports (e.g., only supports) a 1 MHz mode, the primary channel may be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) settings may depend on the status of the primary channel. If the primary channel is busy, for example, because a STA (that only supports 1 MHz operating mode) is transmitting to the AP, the entire available frequency band may be considered busy, even if most of the frequency band remains idle and potentially available.
[0065] In the United States, the available frequency band for 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. The total bandwidth available for 802.11ah ranges from 6MHz to 26MHz, depending on the country code.
[0066] Figure 1D 1 is a system diagram illustrating the RAN 113 and the CN 115 according to an embodiment. As noted above, the RAN 113 may employ NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 may also be in communication with the CN 115.
[0067] The RAN 113 may include gNBs 180a, 180b, and 180c, though it will be appreciated that the RAN 113 may include any number of gNBs while remaining consistent with an embodiment. Each of the gNBs 180a, 180b, and 180c may include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, and 180c may implement MIMO technology. For example, the gNBs 180a and 180b may utilize beamforming to transmit and / or receive signals to and from the gNBs 180a, 180b, and 180c. Thus, the gNB 180a may, for example, use multiple antennas to transmit and / or receive wireless signals to and from the WTRU 102a. In an embodiment, the gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a may transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In an embodiment, gNBs 180a, 180b, and 180c may implement coordinated multi-point (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0068] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or Transmission Time Intervals (TTIs) of varying or scalable lengths (e.g., including varying numbers of OFDM symbols and / or continuously varying absolute time lengths).
[0069] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c while not accessing other RANs (e.g., such as the eNodeBs 160a, 160b, 160c). In a standalone configuration, the WTRUs 102a, 102b, 102c may use one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in an unlicensed band. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect with the gNB 180a, 180b, 180c while also communicating / connecting with another RAN, such as the eNode-B 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNode-B 160a, 160b, 160c may serve as a mobility anchor for the WTRUs 102a, 102b, 102c, and the gNB 180a, 180b, 180c may provide additional coverage and / or throughput for serving the WTRUs 102a, 102b, 102c.
[0070] Each of the gNBs 180a, 180b, 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in UL and / or DL, support of network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 184a, 184b, routing of control plane information towards access and mobility management functions (AMFs) 182a, 182b, etc. Figure 1D As shown, gNBs 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0071] Figure 1DThe illustrated CN 115 may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and possibly data networks (DNs) 185a, 185b. While each of the aforementioned elements is depicted as part of the CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0072] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via the N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling of different PDU sessions with different requirements), selecting a specific SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. The AMF 182a, 182b may use network slicing to customize CN support for the WTRUs 102a, 102b, 102c based on the type of services utilized by the WTRUs 102a, 102b, 102c. For example, different network slices may be established for different use cases, such as services relying on Ultra Reliable Low Latency (URLLC) access, services relying on enhanced Mobile Broadband (eMBB) access, and / or services for Machine Type Communication (MTC) access. The AMF 162 may provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0073] The SMFs 183a and 183b can connect to the AMFs 182a and 182b in the CN 115 via the N11 interface. The SMFs 183a and 183b can also connect to the UPFs 184a and 184b in the CN 115 via the N4 interface. The SMFs 183a and 183b can select and control the UPFs 184a and 184b and configure traffic routing through the UPFs 184a and 184b. The SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, and so on.
[0074] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via the N3 interface. These gNBs may provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, and the like.
[0075] The CN 115 may facilitate communications with other networks. For example, the CN 115 may include, or may communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. Furthermore, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired networks and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to local data networks (DNs) 185a, 185b through UPFs 184a, 184b via the N3 interface to the UPFs 184a, 184b and the N6 interface between the UPFs 184a, 184b and the local data networks (DNs) 185a, 185b.
[0076] Given that Figures 1A to 1D as well as Figures 1A to 1D As described herein, one or more or all of the functions described herein with reference to one or more of the following may be performed by one or more emulated devices (not shown): the WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MMEs 162, SGWs 164, PGWs 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other devices described herein. An emulated device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulated device may be used to test other devices and / or simulate network and / or WTRU functions.
[0077] Emulated devices can be designed to implement one or more tests of other devices in a laboratory environment and / or in a carrier network environment. For example, one or more emulated devices can perform one or more functions or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network in order to test other devices within the communication network. One or more emulated devices can perform one or more functions or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulated device can be directly coupled to another device for testing purposes and / or can use over-the-air wireless communications to perform testing.
[0078] One or more emulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation device can be used in a test scenario in a test lab and / or in a non-deployed (e.g., testing) wired and / or wireless communication network to enable testing of one or more components. The one or more emulation devices can be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which can include one or more antennas) can be used by the emulation device to send and / or receive data.
[0079] This application describes a number of aspects, including tools, features, examples or embodiments, models, methods, etc. Many of these aspects are described in a specific manner and, at least to illustrate individual features, are typically described in a manner that sounds potentially restrictive. However, this is for clarity of description and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with the various aspects described in earlier submissions.
[0080] The various aspects described and contemplated in this application can be implemented in many different forms. Figures 5 to 8 Some embodiments may be provided, but others are also contemplated. Figures 5 to 8 The discussion does not limit the breadth of the specific implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.
[0081] In this application, the terms "reconstruction" and "decoding" are used interchangeably, the terms "pixel" and "sample" are used interchangeably, and the terms "image", "picture" and "frame" are used interchangeably.
[0082] Various methods are described herein, and each method in the various methods includes one or more steps or actions for realizing the described method.Unless the correct operation method requires the steps or actions of a specific order, the order and / or purposes of specific steps and / or actions can be modified or combined.Additionally, terms such as "first", "second" etc. can be used to modify elements, parts, steps, operations, etc. in various embodiments, such as, for example, "first decoding" and "second decoding".Unless specifically required, the use of such terms does not imply the sequencing of the modification operation.Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can, for example, occur before, during, or in the overlapping time period of the second decoding.
[0083] Various methods and other aspects described herein can be used to modify modules, such as Figure 2 and Figure 3 The pre-encoding process 201, intra-frame prediction 260, entropy decoding 245 and / or entropy decoding module 330, intra-frame prediction 360, and post-decoding process 385 of the video encoder 200 and video decoder 300 are shown respectively. In addition, the subject matter disclosed herein presents aspects that are not limited to VVC or HEVC and can be applied, for example, to any type, format, or version of video coding (whether described in a standard or described in a recommendation, whether pre-existing or developed in the future), as well as extensions of any such standards and recommendations (e.g., including VVC and HEVC). Unless otherwise indicated or technically excluded, the various aspects described in this application can be used alone or in combination.
[0084] Various numerical values are used in the examples described herein, such as minimum and maximum value ranges (e.g., 0 to 1, 0 to N, or 0 to 255), bit values for indication or determination, default values, ID numbers (e.g., for adaptation IDs), etc. These and other specific values are for the purpose of describing the examples, and the described aspects are not limited to these specific values.
[0085] Figure 2 is a diagram illustrating an example video encoder. Variations of the example encoder 200 are contemplated, but the following describes encoder 200 for clarity without describing all contemplated variations.
[0086] Before being encoded, the video sequence may undergo pre-encoding processing (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.
[0087] In encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of, for example, coding units (CUs). Each unit is encoded, for example, using intra mode or inter mode. When a unit is encoded in intra mode, the unit performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) whether to use intra mode or inter mode to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the predicted block from the original image block.
[0088] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly code the residual without applying the transform or quantization process.
[0089] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The image block is reconstructed by combining the decoded prediction residual and the predicted block (255). An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).
[0090] Figure 3 is a diagram illustrating an example of a video decoder. In the example decoder 300, the bitstream is decoded by the decoder elements, as described below. The video decoder 300 generally performs the same operations as described above. Figure 2 The encoding process described herein is a decoding process that is the reverse of the encoding process described herein. The encoder 200 may also typically perform video decoding as part of encoding the video data. For example, the encoder 200 may perform one or more of the video decoding steps presented herein. The encoder may, for example, reconstruct the decoded image to maintain synchronization with the decoder with respect to one or more of the following: reference pictures, entropy coding context, and other decoder-related state variables.
[0091] Specifically, the input to the decoder includes a video bitstream, which may be generated by the video encoder 200. First, the bitstream is entropy decoded (330) to obtain transform coefficients, motion vectors, and other decoded information. Picture partition information indicates how the picture is partitioned. Therefore, the decoder may partition (335) the picture according to the decoded picture partition information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The image block is reconstructed by combining (355) the decoded prediction residual and the predicted block. The predicted block may be obtained (370) from intra-frame prediction (360) or motion compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380).
[0092] The decoded picture may also undergo post-decoding processing (385), such as an inverse color transform (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping performed in the pre-encoding process (201). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0093] Figure 4 4 is a diagram illustrating an example of a system in which various aspects and embodiments described herein can be implemented. System 400 may be embodied as a device comprising various components described below and configured to perform one or more aspects of the various aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various embodiments, system 400 is coupled to one or more other systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 400 is configured to implement one or more aspects of the various aspects described in this document.
[0094] The system 400 includes at least one processor 410 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. The processor 410 may include embedded memory, input / output interfaces, and various other circuit systems as known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). The system 400 includes a storage device 440, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drive, and / or optical disk drive. As non-limiting examples, the storage device 440 may include an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.
[0095] System 400 includes an encoder / decoder module 430, which is configured to process data to provide encoded video or decoded video, for example, and may include its own processor and memory. Encoder / decoder module 430 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 430 may be implemented as a separate element of system 400, or may be incorporated into processor 410 as a combination of hardware and software known to those skilled in the art.
[0096] Program code to be loaded onto the processor 410 or encoder / decoder 430 to perform various aspects described in this document may be stored in the storage device 440 and subsequently loaded onto the memory 420 for execution by the processor 410. According to various embodiments, one or more of the processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more of the various items during execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results of processing equations, formulas, operations, and operational logic.
[0097] In some embodiments, memory internal to the processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be memory 420 and / or storage device 440, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, the operating system of the television. In at least one embodiment, fast external dynamic volatile memory (such as RAM) is used as working memory for video coding and decoding operations, such as, for example, MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Experts Team (JVET)).
[0098] Inputs to the elements of system 400 may be provided through various input devices as indicated in block 445. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives an RF signal, such as that transmitted over the air by a broadcaster; (ii) a component (COMP) input terminal (or a set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples ( Figure 4 ) includes composite video.
[0099] In various embodiments, the input device of block 445 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band); (ii) down-converting the selected signal; (iii) again band-limiting a narrower frequency band to select a signal band that may be referred to as a channel in certain embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select the desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal that for example sends by wired (, cable) medium, and by filtering, down-conversion and filtering to the frequency band of expectation again to carry out frequency selection.Various embodiments rearrange the order of (and other) element described above, remove some elements in these elements, and / or add other elements that perform similar or different functions.Adding element can be included in and inserts element between existing element, such as for example inserts amplifier and analog to digital converter.In various embodiments, the RF part comprises antenna.
[0100] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 400 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 410, as needed. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 410, as needed. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 410 and the encoder / decoder 430, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.
[0101] The various components of system 400 can be disposed within an integrated housing. Within the integrated housing, the various components can interconnect and transmit data between the components using a suitable connection arrangement 425 (e.g., an internal bus as known in the art, including an inter-IC (I2C) bus, wiring, and a printed circuit board).
[0102] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. Communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data over communication channel 460. Communication interface 450 may include, but is not limited to, a modem or a network card, and communication channel 460 may be implemented, for example, within a wired and / or wireless medium.
[0103] In various embodiments, data is streamed or otherwise provided to the system 400 using a wireless network, such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these examples are received via a communication channel 460 and a communication interface 450 suitable for Wi-Fi communication. The communication channel 460 of these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments provide streaming data to the system 400 using a set-top box that delivers data via an HDMI connection of an input box 445. Still other embodiments provide streaming data to the system 400 using an RF connection of an input box 445. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use a wireless network other than Wi-Fi, such as a cellular network or a Bluetooth network.
[0104] System 400 can provide output signals to various output devices, including display 475, speaker 485 and other peripheral devices 495. The display 475 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display and / or a foldable display. Display 475 can be used for a television, a tablet computer, a laptop computer, a cellular phone (mobile phone) or other device. Display 475 can also be integrated with other components (for example, as in a smart phone), or be separate (for example, an external monitor for a laptop computer). In various examples of the embodiments, other peripheral devices 495 include one or more of an independent digital video disc (or digital versatile disc) (DVR, a term for both), a CD player, a stereo system and / or a lighting system. Various embodiments use one or more peripheral devices 495 that provide functions based on the output of system 400. For example, a CD player performs the function of playing the output of system 400.
[0105] In various embodiments, control signals are communicated between the system 400 and the display 475, speakers 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to the system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, output devices can be connected to the system 400 via communication interface 450 using communication channel 460. The display 475 and speakers 485 can be integrated into a single unit with other components of the system 400 in an electronic device such as, for example, a television. In various embodiments, the display interface 470 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0106] For example, if the RF portion of input 445 is part of a separate set-top box, the display 475 and speaker 485 may alternatively be separate from one or more of the other components. In various embodiments in which the display 475 and speaker 485 are external components, the output signal may be provided via a dedicated output connection (including, for example, an HDMI port, a USB port, or a COMP output).
[0107] These embodiments may be implemented by computer software implemented by processor 410, or by hardware, or by a combination of hardware and software. As a non-limiting example, these embodiments may be implemented by one or more integrated circuits. Memory 420 may be of any type suitable for the technical environment and, as a non-limiting example, may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 410 may be of any type suitable for the technical environment and, as a non-limiting example, may include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0108] Various implementations involve decoding. As used herein, "decoding" may encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include (or alternatively include) processes performed by a decoder of the various implementations described herein, such as receiving an International Organization for Standardization Base Media File Format (ISOBMFF) file. The ISOBMFF file may include geometry-based point cloud compression (G-PCC) data carried using one or more tracks having sample group description information and sample-to-group box information. The sample group description information may indicate grouping type information and multiple sample group description entries, each of which indicates geometry-based volume or point cloud parameter set information. The sample-to-group box information may include one or more sample-to-group box entries, each of which has multiple entry parameters, the multiple entry parameters including: the grouping type information, the grouping type parameter, and a sample group description index; and so on.
[0109] As a further embodiment, in one example, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or broadly to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0110] Various implementations relate to encoding. In a manner similar to the discussion above regarding "decoding," "encoding," as used herein, may encompass, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential encoding, transforms, quantization, and entropy encoding. In various embodiments, such processes also include (or alternatively include) processes performed by the encoder of the various implementations described herein, such as receiving an International Organization for Standardization Base Media File Format (ISOBMFF) file. The ISOBMFF file may include geometry-based point cloud compression (G-PCC) data carried using one or more tracks having sample group description information and sample-to-group box information. The sample group description information may indicate grouping type information and multiple sample group description entries, each of which indicates geometry-based volume or point cloud parameter set information. The sample-to-group box information may include one or more sample-to-group box entries, each of which has multiple entry parameters, including: the grouping type information, the grouping type parameter, and a sample group description index; among other things.
[0111] As a further example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or broadly to a broader encoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0112] It should be noted that the syntax elements used herein (such as those indicated in Tables 1 to 23 and otherwise indicated in the discussions or figures presented herein) are descriptive terms. Therefore, they do not preclude the use of other syntax element names.
[0113] When the figures are presented as flow charts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the figures are presented as block diagrams, it should be understood that they also provide flow charts of the corresponding methods / processes.
[0114] During the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate-distortion optimization is often formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these approaches may be based on extensive testing of all coding options (including all considered modes or coding parameter values) and a complete evaluation of their coding costs and the associated distortion of the reconstructed signal after decoding and encoding. Faster approaches can also be used to reduce coding complexity, particularly for the calculation of approximate distortion based on predictions or prediction residual signals rather than reconstructed residual signals. A hybrid of these two approaches may also be used, such as by using approximate distortion for only some of the possible coding options and full distortion for others. Other approaches only evaluate a subset of the possible coding options. More generally, many approaches employ any of a variety of techniques to perform optimization, but optimization does not necessarily involve a complete evaluation of both coding costs and associated distortion.
[0115] The specific implementations and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of specific implementation (e.g., discussed only as a method), the specific implementation of the features discussed can also be implemented in other forms (e.g., a device or program). The device can be implemented in, for example, appropriate hardware, software, and firmware. These methods can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication equipment, such as, for example, a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other equipment that facilitates information communication between end users.
[0116] Reference to "one embodiment," "an embodiment," "an example," or "an implementation," or "an implementation," and variations thereof, means that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment," "in an embodiment," "in an example," or "in one implementation," or "in an implementation," and any other variations thereof in various places throughout this application are not necessarily all referring to the same embodiment or example.
[0117] Additionally, the present application may refer to "determining" various information. Determining information may include, for example, estimating information, calculating information, predicting information, or retrieving information from a memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.
[0118] Furthermore, the present application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0119] Additionally, this application may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information can include, for example, accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves, in one way or another, an operation such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0120] It should be understood that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one" is intended to encompass selecting only the first-listed option (A), or only the second-listed option (B), or both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass selecting only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first-listed option and the second-listed option (A and B), or only the first-listed option and the third-listed option (A and C), or only the second-listed option and the third-listed option (B and C), or all three options (A, B, and C). As will be apparent to one of ordinary skill in this and related arts, this can be extended to as many items as listed.
[0121] In addition, as used herein, the word "signaling" refers in particular to indicating something to the corresponding decoder. For example, in some embodiments, the encoder (e.g., to the decoder) signals MPD, adaptation set, representation, preselection, G-PCC component, G-PCCComponent descriptor, G-PCC descriptor or basic attribute descriptor, supplementary attribute descriptor, G-PCC tile inventory descriptor, G-PCC static space region descriptor, GPCCTileId descriptor, GPCC3DRegionID descriptor, other descriptors, elements and attributes, metadata, mode, etc. (e.g., as disclosed herein, included in Tables 1 to 23). In this way, in an embodiment, the same parameters are used at both the encoder side and the decoder side. Therefore, for example, the encoder can send specific parameters (explicit signaling) to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding sending any actual function, bit saving is achieved in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the above refers to the verb form of the word "signal", the word "signal" can also be used as a noun in this document.
[0122] It will be apparent to one of ordinary skill in the art that a specific implementation may generate a variety of signals formatted to carry, for example, storable or transmittable information. The information may include, for example, instructions for performing a method or data generated by one of the described specific implementations. For example, a signal may be formatted to carry a bit stream of the described embodiment. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0123] High-quality three-dimensional (3D) point clouds can be used to represent immersive media. A point cloud can include a set of points represented in 3D space using coordinates indicating the location of the point (e.g., each point) and one or more attributes, such as color, transparency, laser reflectivity, or material properties associated with the point (e.g., each point). Point clouds can be captured in a variety of ways. For example, point clouds can be captured using multiple cameras and depth sensors, light detection and ranging (LiDAR) laser scanners, etc. The number of points used to realistically reconstruct objects and scenes using point clouds can be in the millions (e.g., or even billions). Therefore, efficient representation and compression can be used to store and transmit point cloud data.
[0124] Technologies for capturing and rendering 3D points have applications in areas such as telepresence, virtual reality, and / or large-scale dynamic 3D mapping. One or more of the following 3D point cloud compression (PCC) technologies can be employed: geometry-based compression standards for static point clouds and video-based compression standards for dynamic point clouds. These technologies can be used to support efficient and interoperable storage and transmission of 3D point clouds. Based on these technologies, lossy and / or lossless decoding of point cloud geometry and attributes can be supported.
[0125] Figure 5 An example of a geometry-based point cloud compression (G-PCC) bitstream structure is shown. Features associated with G-PCC are provided. Figure 5 As shown, a bitstream structure for geometry-based point cloud compression (G-PCC) may be shown. A G-PCC bitstream may include a set of G-PCC units, which may be referred to as a type-length-value (TLV) encapsulation structure, such as Figure 5 As shown. The syntax of the G-PCC TLV unit may be given in Table 1, where the G-PCC TLV unit (e.g., each G-PCC TLV unit) has a TLV type (e.g., tlv_type), a G-PCC TLV unit payload length (e.g., tlv_num_payload_bytes), and a G-PCCTLV unit payload (e.g., tlv_payload_byte[i]). tlv_type may describe the G-PCC unit type as shown in Table 2. G-PCC TLV units with unit types 2 and 4 may be geometry data units and attribute data units, respectively. The data units may represent the main components used to reconstruct the point cloud. The payload of the geometry and attribute G-PCC units may correspond to a media data unit (e.g., a TLV unit), which may be decoded by a G-PCC decoder specified in the corresponding geometry and attribute parameter set G-PCC unit.
[0126] Table 1G-PCC TLV Unit Syntax
[0127]
[0128] Table 2 tlv_type and associated data unit description
[0129]
[0130]
[0131] Table 3 G-PCC attribute types by known_attribute_label
[0132] known_attribute_label Attribute Type 0 color 1 Reflectivity 2 Frame Index 3 Material ID 4 transparency 5 Normal
[0133] Table 4G-PCC TLV Encapsulation Unit Payload Syntax
[0134]
[0135]
[0136] The G-PCC bitstream high-level syntax (HLS) may include slices and tile groups in geometry and attribute data. A frame may be divided into multiple tiles and slices. A slice may be a set of points that may be independently encoded and / or decoded. A slice may include one geometry data unit and zero or more attribute data units. An attribute data unit may depend on a corresponding geometry data unit within the same slice. Within a slice, a geometry data unit may appear before an associated attribute unit (e.g., any associated attribute units). The data units of a slice may be consecutive. The ordering of slices within a frame may not be specified.
[0137] A tile group can be identified by a common tile identifier. A tile inventory can describe a bounding box of a tile (e.g., each tile). A tile can overlap another tile within the bounding box. A slice (e.g., each slice) can include an index that identifies the tile to which the slice belongs.
[0138] This document provides features associated with the ISO Base Media File Format (ISOBMFF). ISOBMFF may include G-PCC data carried using one or more tracks. ISOBMFF may have multiple sections that may indicate a file format for storing time-based media. These sections may be based on and / or derived from ISOBMFF, which may be structural, media-independent definitions. ISOBMFF may include structures and media data information for various presentations (e.g., timed presentations) of media data (such as audio, video, etc.). There may be support for non-timed data (such as metadata at various levels within the file structure). The logical structure of a file may be a movie that may include a set of time-parallel tracks. The temporal structure of a file may be such that a track includes a temporal sequence of samples, and these sequences are mapped into the timeline of the entire movie. ISO BMFF may be based on a box-structured file. A box-structured file may include a series of boxes (e.g., which may be referred to as atoms) of a certain size and type. The type may be a 32-bit value and may be selected as four printable characters (e.g., a four-character code (4CC)). The non-timed data may be included in a metadata box (at the file level), or attached to a movie box or one of the timed data streams (called tracks) within a movie.
[0139] Among the top-level boxes within the ISOBMFF container may be a MovieBox ('moov'). The MovieBox may include metadata for the continuous media streams present in the file. The metadata may be signaled within the hierarchy of boxes within the MovieBox (e.g., within the TrackBox ('trak')). A track may represent a continuous media stream present in a file. A media stream may include a sequence of samples (such as audio or video units of an elementary media stream) and may be encapsulated within a MediaDataBox ('mdat') present at the top level of the container. The metadata for a track (e.g., each track) may include a list of sample description entries, where an entry (e.g., each entry) provides the decoding or packaging format used in the track and initialization data for processing the samples in the track. A sample (e.g., each sample) may be associated with a sample description entry for the track. Tools may define an explicit timeline mapping for a track (e.g., each track). This may be referred to as an edit list and may be signaled using an EditListBox using the following syntax, where an entry (e.g., each entry) defines a portion of a track timeline by mapping a portion of a component timeline or by indicating an empty time (e.g., portions of a presentation timeline that are not mapped to any media, which may be referred to as an empty edit).
[0140]
[0141] Figure 6An example of a sample structure when a decoded G-PCC bitstream is stored in a single track is shown. Features associated with the G-PCC container file format are provided herein. If a G-PCC bitstream is carried in a single track, a G-PCC encoded bitstream may be represented by a single track declaration. Single track encapsulation of G-PCC data may utilize simple ISOBMFF encapsulation by storing the G-PCC bitstream in a single track (e.g., without further processing). A sample in the track (e.g., each sample) may include one or more G-PCC components. A sample (e.g., each sample) may include one or more TLV encapsulation structures. As Figure 6 As shown, the sample structure can be shown when the G-PCC geometry and attribute bitstreams are stored in a single track.
[0142] Figure 7 An example of a multi-track G-PCC bitstream container structure is shown. If the decoded G-PCC geometry bitstream and the decoded G-PCC attribute bitstream are stored in separate tracks, the samples in the track (e.g., each sample) may include at least one TLV encapsulation structure carrying data for a single G-PCC component. Figure 7 The structure of a multi-track G-PCC container is shown. These boxes can be mapped to corresponding boxes.
[0143] Based on this structure, a multi-track G-PCC ISOBMFF container may include one or more of the following: a G-PCC track including a geometry parameter set, a sequence parameter set, and a geometry bitstream sample carrying a geometry data TLV unit (e.g., the track may include track references to other tracks carrying payloads of G-PCC attribute components); or zero or more G-PCC tracks, where a track (e.g., a track) includes attribute parameter sets for corresponding attributes and attribute bitstream samples carrying attribute data TLV units.
[0144] If the G-PCC bitstream is carried in multiple tracks, the track reference tool can be used to link between G-PCC component tracks. A TrackReferenceTypeBoxes can be added to the TrackReferenceBox inside the TrackBox of the G-PCC track. The TrackReferenceTypeBox may include an array of track_IDs specifying the tracks referenced by the G-PCC track. To link a G-PCC geometry track to a G-PCC attribute track, the reference_type of the TrackReferenceTypeBox in the G-PCC geometry track may identify the associated attribute track. The 4CC of these track reference types may be 'gpca', which may indicate that the referenced track includes the decoded bitstream of the G-PCC attribute component.
[0145] If the 3D spatial region information and the associated G-PCC tiles within the 3D spatial region in the G-PCC bitstream are dynamically changing, the timed metadata track may carry the dynamically changing 3D spatial region information. The 3D spatial region information timed metadata track may provide an association over time between the 3D spatial region information and the corresponding G-PCC tiles of the 3D spatial region (e.g., each 3D spatial region).
[0146] The timed metadata track may include a 'cdsc' track reference to a G-PCC base track. The G-PCC base track may include a track reference type defined using 4CC 'gbsr' to the timed metadata track.
[0147] Non-timed G-PCC data can be encapsulated into ISOBMFF files using entries. An entry can be a box that carries data (not sample data) that does not use timed processing.
[0148] A single item or multiple items with G-PCC tiles may be used to support the carrying of non-timed G-PCC data. For multiple items with G-PCC tiles, items of type 'gpt1' may be described herein along with property items and item references to support partial access.
[0149] Figure 8 An example of a G-PCC bitstream container structure having a G-PCC item and a tile item is shown. Figure 8 As shown, non-timed G-PCC data including three G-PCC tiles can be carried in multiple items by storing the G-PCC tiles (e.g., each G-PCC tile) in a separate item. This can enable a decoding device (e.g., a player) to identify the item including the appropriate G-PCC tile by interpreting the associated spatial region item properties.
[0150] Figure 9 An example of a G-PCC bitstream container structure with a G-PCC item, a tile item, and a spatial region item is shown. To support a finer-grained indication of a G-PCC tile, subsample information may be used (e.g., even if a G-PCC tile item may include multiple G-PCC tiles, as described with respect to Figure 9 For example, the subsample information may be adapted to indicate an identifier of a tile included in the G-PCC tile entry.
[0151] A G-PCC temporal level may be a subset of frames in a G-PCC bitstream. A G-PCC temporal level may include a subsequence whose frame rate is lower than the frame rate of the actual bitstream sequence. A G-PCC frame (e.g., each G-PCC frame) may be associated with a temporal level (e.g., a particular temporal level). A temporal level (e.g., each temporal level) may be identified by a temporal level identifier (e.g., a unique temporal identifier), where the temporal level identifier (ID) of the first temporal level is 0.
[0152] A G-PCC bitstream (e.g., G-PCC data) may be carried and / or stored in one or more temporal level tracks. The G-PCC data may include multiple G-PCC samples. Information describing the temporal level tracks and the mapping between samples and temporal levels may be obtained in a file. A G-PCC sample belonging to a temporal level may not have decoding dependencies (e.g., any decoding dependencies) on a G-PCC sample (e.g., any G-PCC sample) present in a higher temporal level. Prior to the decoding (e.g., decoding) process, samples may be extracted from the temporal level tracks and combined into a single consistent bitstream. When extracting a G-PCC bitstream for a target temporal level having an ID greater than 0 and a target tile ID, data from lower temporal level samples (e.g., all lower temporal level samples) may be included in the resulting bitstream, and a track may be selected accordingly during the extraction process.
[0153] Figure 10 An example of the time level in a G-PCC sequence is shown. Figure 10 As shown, if the G-PCC bitstream is divided into three temporal levels, playback of the G-PCC bitstream at 30fps, 45fps and 60fps can be enabled: temporal level 0 can represent a 30fps subsequence, and temporal level 1 and temporal level 2 can represent a 15fps subsequence, respectively.
[0154] A G-PCC track that includes a GPCCScalabilityInfoBox in its sample entry may be referred to as a temporal level track that carries a bitstream subset. This box may signal scalability information for the G-PCC track. When this box is present in a track with a sample entry of type 'gpe1', 'gpeg', 'gpc1', 'gpcg', 'gpcb', and / or 'gpeb', it may indicate support for temporal scalability and provide information about the temporal levels present in the G-PCC track.
[0155] This article provides features related to media such as VR and immersive 3D graphics. High-quality 3D point clouds can provide representations of immersive media, enabling various forms of interaction and communication with virtual worlds. The large amount of information used to represent such point clouds can require efficient decoding algorithms. This article describes example techniques for geometry-based point cloud compression. This article describes example techniques for temporal scalability. Temporal scalability can provide support for temporal partial access to G-PCC data encapsulated in a container.
[0156] A point cloud sequence may represent a scene with multiple tiles. In some examples, a decoding device may be able to access (e.g., stream and / or render) individual tiles without decoding the rest of the scene. Similarly, a point cloud may represent a single object. Parts of the object may be accessed (e.g., stream and / or render) without decoding the entire point cloud.
[0157] Multiple temporal tracks can be used to support G-PCC data within a file. A container (e.g., an ISOBMFF container) can signal dynamically changing G-PCC parameter sets, including, for example, a geometry parameter set (GPS), a sequence parameter set (SPS), and / or an attribute parameter set (APS). Frame-Specific Attribute Properties (FSAP) parameter sets can specify attribute properties. These attribute properties can be used for specific attribute frames. When multiple tracks, multiple temporal tracks, or multiple temporal tile tracks are present, coding devices (e.g., encoders and / or decoders) may not be aware of how such parameter sets are carried within ISOBMFF.
[0158] As described herein, signaling techniques can be used to carry dynamically changing G-PCC parameter sets in single-track, multi-track, G-PCC temporal track, and temporal tile track scenarios. Constraints on temporal track and temporal tile track sample entries can be updated as described herein.
[0159] Features associated with G-PCC parameter set sample groups are provided herein. Features associated with time-level track cases are provided herein. G-PCC data may be carried using one or more tracks. A track (e.g., each of the one or more tracks) may include sample group description information (e.g., in a SampleGroupDescriptionBox) and sample-to-group box information (e.g., in a SampleToGroupBox). In an example, if G-PCC data is carried using multiple time-level tracks and the parameter set information changes over time, a G-PCC parameter set information sample group with grouping_type equal to 'gpsg' may be used to signal parameter set information related to samples present in the time-level track. The grouping type may indicate that G-PCC samples are grouped together, and the G-PCC samples use the same sample group description entry (e.g., the same parameter set information). For example, a G-PCC parameter set information sample group with grouping type 'gpsg' may be used to group G-PCC samples that use the same parameter set information in a time-level track. Using 'gpsg' for the grouping_type in a sample group may indicate that samples in the temporal track are assigned to the corresponding parameter set carried in the sample group. If there is a SampleToGroupBox with grouping_type equal to 'gpsg', there may be a companion SampleGroupDescriptionBox with the same grouping type. The SampleToGroupBox may include the index of the sample group description entry to which each sample belongs.
[0160] If there is a SampleToGroupBox with grouping_type equal to 'gpsg', the SampleGroupDescriptionEntry at index 1 of the associated SampleGroupDescriptionBox may carry the parameter set for decoding the first sample in the temporal track. SampleGroupDescriptionBoxes at index 2 and above may carry the changed parameter set associated with the sample or consecutive sample sets present in the temporal track (e.g., only the changed parameter set). In an example, the SampleGroupDescriptionEntry may carry the setting element of one of the SPS, GPS, APS, and FSAP parameter sets (e.g., only the setting element).
[0161] In an example, if 'gpc1', 'gpcg', 'gpe1' or 'gpeg' sample entries are used in a temporal level track, the sample group may include a G-PCC parameter set for decoding samples present in the temporal level track.
[0162] In this example, the SampleGroupDescriptionBox may not carry the FSAP. The FSAP may be carried in the samples of the attribute track.
[0163] Features associated with non-temporal tracks (e.g., component tracks and single component tracks) are provided herein. If a single G-PCC track or multiple G-PCC tracks are used to carry G-PCC data and the parameter set information changes over time, a G-PCC parameter set information sample group with grouping_type 'gpsg' may be used to signal the parameter set information related to the samples present in the G-PCC track. The G-PCC parameter set information sample group with grouping type 'gpsg' may be used to group G-PCC samples that use the same parameter set information in the G-PCC track. The SampleToGroupBox may include grouping type information (e.g., grouping_type). The SampleGroupDescriptionBox may include grouping type information (e.g., grouping_type). Using 'gpsg' for the grouping_type in the sample group may indicate that the samples in the G-PCC track are assigned to the corresponding parameter sets carried in the sample group. The sample may include a sample group description entry (e.g., one or more SampleGroupDescriptionEntry). The sample group description entry may indicate parameter set information (e.g., based on geometry or point cloud parameter set information). If there is a SampleToGroupBox with grouping_type equal to 'gpsg', a companion SampleGroupDescriptionBox with the same grouping type may be present and may include the index of the group to which the sample belongs. If there is a SampleToGroupBox with grouping_type equal to 'gpsg', the SampleGroupDescriptionEntry at index 1 of the associated SampleGroupDescriptionBox may carry the parameter set for decoding the first sample in the track. SampleGroupDescriptionEntries at index 2 and above may carry updated parameter sets (e.g., only the updated parameter sets) associated with samples or consecutive sample sets present in the track.
[0164] In an example, the SampleGroupDescriptionEntry may carry a setting element of one of the SPS, GPS, APS, and FSAP parameter sets (eg, only the setting element).
[0165] If 'gpc1', 'gpcg', 'gpe1', or 'gpeg' sample entries are used in a G-PCC track, and a sample group with grouping_type 'gpsg' exists, the sample group may include a G-PCC parameter set for decoding samples present in the track.
[0166] A G-PCC temporal level track or a G-PCC track may include one or more SampleToGroupBox boxes with grouping_type equal to 'gpsg'. If there is more than one SampleToGroupBox box with grouping_type equal to 'gpsg', the SampleToGroupBox boxes may have unique grouping_type_parameter values and the version of each of the SampleToGroupBox boxes may be set to 1.
[0167] In this example, the SampleGroupDescriptionBox may not carry FSAP. FSAP may be carried in samples of a specific attribute track.
[0168] Features associated with component track cases (e.g., normal and temporal level track cases) are provided herein. In an example, if 'gpe1' or 'gpeg' sample entries are used in a track, a G-PCC temporal level track or a G-PCC track may include multiple SampleToGroupBox boxes with grouping_type equal to 'gpsg' but with different grouping_type_parameter parameters. A sample group with a parameter set setting unit of a particular type may be identified using the 'gpsg' grouping_type and a unique grouping_type_parameter value. If there are multiple SampleToGroupBox boxes with different grouping_type_parameter values, the SampleGroupDescriptionEntry present in the sample group description box with grouping type 'gpsg' (e.g., each SampleGroupDescriptionEntry) may include one (e.g., only one) of SPS, GPS, APS, or frame-specific attribute parameters (e.g., but not including a combination of parameter sets). _setupUnits may be 1 in the SampleGroupDescriptionEntry representing one of the SPS, GPS, or FSAP setup units. The numOfSetupUnits present in the SampleGroupDescriptionEntry representing an APS may be equal to the number of attributes present in the G-PCC bitstream or an updated APS set of one or more attributes present in the bitstream.
[0169] The group type may be 'gpsg', which may be associated with a container sample group description box (e.g., 'snpd'). The group type 'gpsg' may not be used (e.g., it may be associated with a quantity of 0). If used, it may be associated with a quantity of 1 or greater.
[0170] The G-PCC parameter set information sample group entry may define parameter set information for samples using the same G-PCC parameter set.If there are multiple instances of the SampleToGroupBox box with grouping_type equal to 'gpsg', the version of the SampleToGroupBox box may be set to 1.
[0171] An example syntax for a G-PCC parameter set sample group is provided below:
[0172]
[0173]
[0174] Example semantics may include one or more of the following: numOfSetupUnits or setupUnit. The parameter numOfSetupUnits may specify the number of G-PCC setup units signaled in the sample group description entry. The parameter setupUnit may include a G-PCC unit that carries one of the SPS, GPS, APS, or FSAP parameters.
[0175] G-PCC parameter sets may not be carried in temporal tile tracks. In an example, when multiple temporal tile tracks are used to carry G-PCC data, sample groups with grouping_type 'gpsg' may not exist in tracks with sample entries 'gpcb', 'gpeb', or 'gpt1'.
[0176] In an example, if multiple temporal tile tracks are used to carry G-PCC data, samples in a G-PCC tile base track with a sample entry type of 'gpcb' or 'gpeb' may carry a G-PCC unit that includes one or more of: SPS, GPS, APS, tile inventory, or FSAP information.
[0177] The presentation time of a sample can be used to identify a G-PCC tile base track sample for decoding a G-PCC temporal tile track sample. The presentation time of the corresponding tile base track sample can be equal to or less than the presentation time of the temporal tile track sample. If the presentation times of a tile base track sample and the corresponding temporal tile track sample are not equal, a tile base track sample with a previous presentation time that is closer to the presentation time of the temporal tile track sample can be used to decode the tile base track sample or identify tile inventory information.
[0178] The presentation time of the attribute G-PCC tile track sample can be used to identify the G-PCC tile base track sample with FSAP parameter information, which is used to decode the G-PCC attribute temporal tile track sample. The presentation time of the corresponding tile base track sample with FSAP information can be the same as the presentation time of the temporal tile track sample carrying the attribute data.
[0179] Figure 11 Examples of using the 'gpsg' sample group in multi-track and time-level tracks are shown. Features associated with the extraction process are provided. Figure 11 As shown, the GPCC file structure may have two single-component tracks. For a track, there may be a sample group description box (e.g., an 'sgpd' box) with grouping type 'gpsg' and a sample to group box (e.g., one or more 'sgpd' boxes). Track 1 may be a geometry track (e.g., a geometry component track). Track 1 may include two sample to group boxes and a sample group description box with grouping_type equal to 'gpsg', including SPS and GPS parameter sets. Track 2 may be an attribute track. The sample to group box may include a sample to group box entry, which may include multiple parameters. For example, Figure 11 As illustrated, a sample-to-group box entry may include a grouping type; a grouping type parameter (eg, grouping_type_parameter); an entry count (eg, entry_count); a sample count (eg, sample_count); and / or a sample group description index (eg, sample_description_index).
[0180] The grouping_type_parameter may include a value indicating a parameter set setting unit of a sample group description entry. For example, a sample-to-group box with a grouping_type_parameter of 0 may include associated sample group description entries (e.g., SPS and GPS) including a setting unit for an SPS parameter set. A sample-to-group box with a grouping_type_parameter of 1 may include associated sample group description entries (e.g., SPS and GPS) including a setting unit for a GPS parameter set. As another example, a sample-to-group box with a grouping_type_parameter of 2 may include associated sample group description entries (e.g., APS and FSAP) including a setting unit for an APS parameter set. A sample-to-group box with a grouping_type_parameter of 3 may include associated sample group description entries (e.g., APS and FSAP) including a setting unit for an FSAP parameter set.
[0181] If the GPS information changes over time, the GPS information (e.g., new GPS information) can be carried in the sample group description entry. Samples using a changing GPS parameter set can be indicated in a sample-to-group box with grouping_type_parameter equal to 1. sample_count can indicate how many samples in a sample set use a particular parameter set. Figure 11 As shown, samples 1 to 100 in track 1 (e.g., as indicated by sample_count[1] being set to 100) may use GPS parameter set data from the sample group entry description at index 2; and the following 200 samples, i.e., samples 101 to 300 (e.g., as indicated by sample_count[2] being set to 200), may use GPS parameter set data from the sample group entry description at index 3. The sample group entry description at index 2 and the sample group entry description at index 3 may be present in the sample group description box. Similarly, a sample to group box with grouping_type_parameter equal to 2 may indicate that samples 1 to 200 in track 2 (e.g., as indicated by sample_count[1] being set to 200) use APS parameter set data from the sample group entry description at index 1; and the following 100 samples, i.e., samples 201 to 300 (e.g., as indicated by sample_count[2] being set to 100), use APS parameter set data from the sample group entry description at index 2.
[0182] entry_count may indicate the number of sample sets. Sample sets may each use a different sample group description entry. For example, Figure 11 As shown in the First Sample to Group box of , if entry_count is 1, there may be 1 sample set that all use the same sample group description entry (e.g., the SPS indicated by sample_description_index1 in this case). As another example, Figure 11 As shown in the Second Sample to Group box, if entry_count is 2, there may be 2 sample sets, where the samples in one set all use the same sample group description entry (e.g., in this case, the first 100 samples use GPS, as indicated by sample_description_index2; and the next 200 samples use GPS New, as indicated by sample_description_index3).
[0183] Figure 12 An example of using the 'gpsg' sample group in a multi-component track is shown. Figure 12 As shown, the GPCC file structure may have a single multi-component track. This track may include multiple instances of a sample group description box and a sample to group box with a grouping type of 'gpsg'. Figure 12 As shown, the first track (e.g., track 1) may include four sample group boxes with grouping_type equal to 'gpsg' and different grouping_type_parameter values (representing SPS, GPS, APS, and FSAP parameter sets). The sample entry to the group box may be associated with a sample group description entry (including a setting element for one of the SPS, GPS, APS, and FSAP parameter sets).
[0184] like Figure 12As shown, the samples may use the SPS in the SampleGroupDescriptionEntry present at position 1. As indicated in the SampleToGroupBox where grouping_type is equal to 'gpsg' and the grouping_type_parameter value is 1, samples from 1 to 100 may use the GPS parameter set data from the sample group entry description located at sample group description index 2, and samples 101 to 300 may use the GPS parameter set data from the sample group entry description located at sample group description index 4 present in the sample group description box. As indicated in the SampleToGroupBox where 'gpsg' is equal to grouping_type and the grouping_type_parameter value is 2, samples from 1 to 200 may use the APS parameter set data from the sample group entry description located at sample group description index 3, and samples 201 to 300 may use the APS parameter set data from the sample group entry description located at sample group description index 5.
[0185] This document provides example constraints for temporal tracks and temporal tile tracks. A G-PCC track that includes a GPCCScalabilityInfoBox in its sample entry may be referred to as a temporal track that carries a bitstream subset. In the example, if multiple temporal tracks are used to carry G-PCC data, the temporal tracks that carry the G-PCC bitstream may use the same sample entry type. In the example, if multiple temporal tile tracks are used to carry G-PCC data, the temporal tile tracks that carry the G-PCC bitstream may use the same sample entry type 'gpt1'.
[0186] Although features and elements are described above in particular combinations, it will be understood by one of ordinary skill in the art that each feature or element may be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (sent via a wired or wireless connection) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as built-in hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile disks (DVDs)). A processor associated with software may be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method for signaling geometry-based volume or point cloud parameter set information, the method comprising: A file comprising geometry-based point cloud compression (G-PCC) data associated with a track is received, wherein the track comprises sample group description information and sample-to-group box information, wherein the sample group description information indicates grouping type information and a plurality of sample group description entries, wherein the plurality of sample group description entries comprises: a first sample group description entry, the first sample group description entry being associated with a first sample group description index, the first sample group description index being associated with a first parameter set type; and a second sample group description entry, the second sample group description entry being associated with a second sample group description index, the second sample group description index being associated with a second parameter set type; and The sample-to-box information includes: a first sample-to-group box, the first sample-to-group box including a first grouping type parameter, the first grouping type parameter indicating that a first sample set uses a first parameter set, the first parameter set being associated with the first sample group description index; and A second sample-to-group box includes a second grouping type parameter, the second grouping type parameter indicates that a second sample set uses a second parameter set, the second parameter set is associated with the second sample group description index.
2. The method of claim 1 , wherein the first sample-to-group box further comprises an entry count, wherein the entry count indicates a number of sample sets in a plurality of sample sets including the first sample set, wherein each sample set in the plurality of sample sets is associated with a different sample group description entry in the plurality of sample group description entries.
3. The method of claim 1 , wherein the plurality of sample group description entries further comprises a third sample group description entry, the third sample group description entry being associated with a third sample group description index, the third sample group description index being associated with a third parameter set type; and wherein the sample-to-group box information further comprises a third sample-to-group box, the third sample-to-group box comprising a third grouping type parameter, the third grouping type parameter indicating that a third sample set uses a third parameter set, the third parameter set being associated with the third sample group description index.
4. The method of claim 1 , wherein the track is a first track, the sample group description information is first sample group description information, the sample-to-group box information is first sample-to-group box information, the plurality of sample group description entries is a first plurality of sample group description entries, and the geometry-based volume or point cloud parameter set information is a first geometry-based volume or point cloud parameter set information; and wherein the G-PCC data is further associated with a second track, wherein the second track includes second sample group description information and second sample-to-group box information, wherein the second sample group description information indicates a second plurality of sample group description entries, the second plurality of sample group description entries indicating second geometry-based volume or point cloud parameter set information; and wherein the second plurality of sample group description entries includes: a third sample group description entry associated with a third sample group description index associated with a third parameter set type; and a fourth sample group description entry, the fourth sample group description entry being associated with a fourth sample group description index, the fourth sample group description index being associated with a fourth parameter set type; and The second sample to group box information includes: a third sample-to-group box, the third sample-to-group box including a third grouping type parameter, the third grouping type parameter indicating that a third sample set uses a third parameter set, the third parameter set being associated with the third sample group description index; and A fourth sample-to-group box includes a fourth grouping type parameter, the fourth grouping type parameter indicating that a fourth sample set uses a fourth parameter set, the fourth parameter set being associated with the third sample group description index.
5. The method of claim 1, wherein the first parameter set is of the first parameter set type, the second parameter set is of the first parameter set type, and the first parameter set is a sequence parameter set, a geometry parameter set, an attribute parameter set, or a frame-specific attribute characteristic parameter set. The method according to claim 1 , wherein the grouping type information indicates that samples are grouped together, the samples using a same sample group description entry among the plurality of sample group description entries.
7. The method of claim 1 , wherein the first sample-to-group box further comprises a first sample count indicating the number of samples in the first sample set; and the second sample-to-group box further comprises a second sample count indicating the number of samples in the second sample set.
8. The method of claim 1, wherein the file is an ISO Base Media File Format (ISOBMFF) file, and at least one sample in the second sample set is located in the first sample set.
9. The method of claim 1, wherein the track is a geometric track, an attribute track, or a multi-component track.
10. The method of claim 1, wherein the tracks comprise one or more time-level tracks.
11. An apparatus for signaling geometry-based volume or point cloud parameter set information, the apparatus comprising: A processor configured to: A file comprising geometry-based point cloud compression (G-PCC) data associated with a track is received, wherein the track comprises sample group description information and sample-to-group box information, wherein the sample group description information indicates grouping type information and a plurality of sample group description entries, wherein the plurality of sample group description entries comprises: a first sample group description entry, the first sample group description entry being associated with a first sample group description index, the first sample group description index being associated with a first parameter set type; and a second sample group description entry, the second sample group description entry being associated with a second sample group description index, the second sample group description index being associated with a second parameter set type; and The sample-to-box information includes: a first sample-to-group box, the first sample-to-group box including a first grouping type parameter, the first grouping type parameter indicating that a first sample set uses a first parameter set, the first parameter set being associated with the first sample group description index; and A second sample-to-group box includes a second grouping type parameter, the second grouping type parameter indicates that a second sample set uses a second parameter set, the second parameter set is associated with the second sample group description index.
12. The apparatus of claim 11 , wherein the first sample-to-group box further comprises an entry count, wherein the entry count indicates a number of sample sets in a plurality of sample sets including the first sample set, wherein each sample set in the plurality of sample sets is associated with a different sample group description entry in the plurality of sample group description entries.
13. The apparatus of claim 11 , wherein the plurality of sample group description entries further comprises a third sample group description entry, the third sample group description entry being associated with a third sample group description index, the third sample group description index being associated with a third parameter set type; and The sample-to-group box information further includes a third sample-to-group box, the third sample-to-group box includes a third grouping type parameter, the third grouping type parameter indicates that a third sample set uses a third parameter set, and the third parameter set is associated with the third sample group description index.
14. The apparatus of claim 11 , wherein the track is a first track, the sample group description information is first sample group description information, the sample-to-group box information is first sample-to-group box information, the plurality of sample group description entries is a first plurality of sample group description entries, and the geometry-based volume or point cloud parameter set information is a first geometry-based volume or point cloud parameter set information; and wherein the G-PCC data is further associated with a second track, wherein the second track includes second sample group description information and second sample-to-group box information, wherein the second sample group description information indicates a second plurality of sample group description entries, the second plurality of sample group description entries indicating second geometry-based volume or point cloud parameter set information; and wherein the second plurality of sample group description entries includes: a third sample group description entry associated with a third sample group description index associated with a third parameter set type; and a fourth sample group description entry, the fourth sample group description entry being associated with a fourth sample group description index, the fourth sample group description index being associated with a fourth parameter set type; and The second sample to group box information includes: a third sample-to-group box, the third sample-to-group box including a third grouping type parameter, the third grouping type parameter indicating that a third sample set uses a third parameter set, the third parameter set being associated with the third sample group description index; and A fourth sample-to-group box includes a fourth grouping type parameter, the fourth grouping type parameter indicating that a fourth sample set uses a fourth parameter set, the fourth parameter set being associated with the third sample group description index.
15. The apparatus of claim 11, wherein the first parameter set is of the first parameter set type, the second parameter set is of the first parameter set type, and the first parameter set is a sequence parameter set, a geometry parameter set, an attribute parameter set, or a frame-specific attribute characteristic parameter set. 16 . The apparatus according to claim 11 , wherein the grouping type information indicates that samples are grouped together, the samples using a same sample group description entry among the plurality of sample group description entries.
17. The apparatus of claim 11, wherein the first sample-to-group box further comprises a first sample count indicating the number of samples in the first sample set; and the second sample-to-group box further comprises a second sample count indicating the number of samples in the second sample set.
18. The apparatus of claim 11, wherein the file is an ISO Base Media File Format (ISOBMFF) file, and at least one sample in the second sample set is located in the first sample set.
19. The apparatus of claim 11, wherein the track is a geometric track, an attribute track, or a multi-component track.
20. The apparatus of claim 11, wherein the tracks comprise one or more time-level tracks.
21. A method comprising: receiving an indication of a common tile identifier; identifying a slice based on the common tile identifier; as well as The slice is decoded.