Dynamic Adaptation of Sub-bitstreams of Components of Volumetric Content in a Streaming Service

The dynamic adaptation of visual volumetric content in video coding systems, facilitated by messages and parameter sets indicating changes in attributes and codecs, addresses the challenge of inefficient resource utilization and improves user experience by optimizing decoding processes in real-time.

JP7689086B2Active Publication Date: 2025-06-05INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021577997
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-02
Filing Date
2020-07-02
Publication Date
2025-06-05
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

Existing video coding systems struggle to dynamically adapt to changes in visual volumetric content, such as point cloud component substreams, in real-time, especially in terms of bitrate adaptation and codec changes, which can lead to inefficient resource utilization and poor user experience.

Method used

The system employs dynamic adaptation of visual volumetric content by using messages and parameter sets to indicate changes in active attributes and codecs. This includes processing messages like Component Codec Change (CCC), Active Attribute (AA), Component Change Parameter Set (CCPS), and Parameter Set Activation (PSA) to dynamically adjust the decoding process based on the operating environment, resource availability, and client capabilities.

Benefits of technology

This approach enables efficient bitrate adaptation and codec changes, optimizing resource utilization and improving user experience by ensuring that the visual volumetric content is decoded using the most appropriate attributes and codecs, even in varying network conditions and client capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689086000009
    Figure 0007689086000009
  • Figure 0007689086000010
    Figure 0007689086000010
  • Figure 0007689086000011
    Figure 0007689086000011
Patent Text Reader

Abstract

A media content processing device can decode visual volumetric content based on one or more messages that may indicate which attribute sub-bitstreams of one or more attribute sub-bitstreams are active, as indicated in a parameter set. The parameter set may include a visual volumetric video-based parameter set. The messages may be received by a decoder to indicate the one or more active attribute sub-bitstreams. The decoder can perform decoding, such as determining which attribute sub-bitstreams to use to decode the visual media content, based on the one or more messages. For example, one or more messages may be generated and sent to the decoder to indicate deactivation of one or more attribute sub-bitstreams. The decoder can determine the inactive attribute sub-bitstreams and skip the inactive attribute sub-bitstreams to decode the visual media content based on the one or more messages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 869,705, filed on Jul. 2, 2019, entitled “Dynamic Adaptation of Point Cloud Component Substreams in a Point Cloud Streaming Service”, the entire disclosure of which is incorporated herein by reference as if fully set forth herein.

Background Art

[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based systems such as wavelet-based systems, object-based systems, and / or block-based hybrid video coding systems.

Summary of the Invention

[0003] Systems, methods, and means are disclosed for the dynamic adaptation of visual volumetric content, such as a point cloud component sub-bitstream in a point cloud streaming service. The dynamic adaptation of visual volumetric content can be based on one or more messages and / or parameter sets that indicate one or more changes in the visual volumetric content. The changes can include, for example, changes in active attributes and / or codecs. For example, bitrate adaptation processing in a streaming session can add / drop one or more attributes for a video component and / or change the representation of the video component to a representation encoded with a different codec (based on an operating environment such as, for example, resource availability, bandwidth attributes, client coding capabilities, and / or client rendering capabilities). One or more messages and / or parameter sets are generated and sent to a decoder, which can indicate visual volumetric content processing information that can include, for example, change indications. The decoder can perform decoding, such as determining which attribute sub-bitstreams and / or codecs to use for decoding, based on the one or more messages and / or parameter sets.

[0004] One or more messages and / or parameter sets can include, for example, a Component Codec Change (CCC) message, an Active Attribute (AA) message, a Component Change Parameter Set (CCPS), and / or a Parameter Set Activation (PSA) message. For example, a Supplemental Enhancement Information (SEI) message and / or parameter set of a visual volumetric content bitstream can support adaptive streaming of visual volumetric content. A message (e.g., a CCC SEI message) can cause a decoder (e.g., a visual volumetric content decoder) to determine which video codec to use for a reference component by notifying the decoder of a codec change for one or more components of visual volumetric content. A message (e.g., an AA SEI message) can cause a decoder (e.g., a visual volumetric content decoder) to determine which active attributes to use and which inactive attributes to ignore for a reference component by notifying the decoder of an attribute change for one or more components of visual volumetric content. A parameter set (e.g., a CCPS) can include information about changes (e.g., attributes or codecs) made to one or more components of visual volumetric content with respect to a parameter set (e.g., a Sequence Parameter Set (SPS)), thereby causing a decoder to determine which video codec and active attributes to use for a reference component. A message (e.g., a PSA SEI message) can cause a decoder to determine which parameter set is active for one or more components of visual volumetric content.The decoder can determine, for example, based on a message and / or parameter set, which of a plurality of attribute sub-bitstreams indicated by a parameter set associated with the visual volumetric content is active, and can decode the visual volumetric content using the active attribute sub-bitstream.

[0005] In one example, the message can indicate a set of attribute sub-bitstreams that are active for use in decoding a bitstream of visual volumetric content (e.g., after receipt of the message). One or more messages and / or parameter sets can signal the activation of a subset of attribute component substreams in the bitstream of the visual volumetric content. One or more messages and / or parameter sets can signal codec changes to the component substreams of the visual volumetric content. The visual volumetric content decoder can be configured to perform one or more of, for example, obtaining an indication indicating whether at least one attribute signaled by a reference parameter set is inactive, obtaining an active attribute indication (e.g., the number of active attributes and their indices) if the indication indicates that at least one attribute in the reference parameter set is inactive, identifying inactive attributes based on the active attribute indication, or skipping inactive attributes in the reference parameter set during decoding.

[0006] In one example, the method may be implemented to perform dynamic adaptation of a point cloud component sub-bitstream in a point cloud streaming service. The method can include determining to deactivate an attribute sub-bitstream of a plurality of attribute sub-bitstreams indicated by a parameter set associated with the visual volumetric content, and generating a message as described herein to indicate deactivation of the attribute sub-bitstream. The method may be implemented by a device such as, for example, a visual media content processing or coding device. The visual media content processing or coding device can include a DASH client or a streaming client, such as, for example, a video-based point cloud compression (V-PCC) client.

[0007] A method for decoding visual media content can include determining which attribute sub-bitstream of a plurality of attribute sub-bitstreams indicated by a parameter set associated with the visual volumetric content to use for decoding the visual volumetric content based on a message indicating which of the attribute sub-bitstreams are active, and decoding the visual volumetric content using the active attribute sub-bitstreams based on the message. The attributes indicated by the parameter set can characterize the visual media content.

[0008] A method for processing media content can include determining to deactivate an attribute sub-bitstream of a plurality of attribute sub-bitstreams indicated by a parameter set associated with the visual volumetric content, and generating a message indicating deactivation of the attribute sub-bitstream.

[0009] A method for decrypting visual media content includes, for example, obtaining a parameter set associated with the visual volumetric content, receiving a message indicating which of a plurality of attribute sub-bitstreams indicated in the parameter set is active, determining active and non-active attribute sub-bitstreams based on the message, decrypting the visual volumetric content using the active attribute sub-bitstreams, and skipping the non-active attribute sub-bitstreams.

[0010] The message can be signaled in the bitstream. The message can include a supplementary enhancement information (SEI) message. The message can have a persistence scope that lasts until the end of the bitstream. The message can have a persistence scope that lasts until another different message is received. The message can include an indicator indicating the number of active attribute sub-bitstreams among a plurality of attribute sub-bitstreams indicated in the parameter set associated with the visual volumetric content.

[0011] The parameter set can include a visual volumetric parameter set containing attribute information. The message can refer to a part of the attribute information of the active attribute sub-bitstream. The parameter set can indicate a plurality of attributes. The message can include an indicator indicating that a plurality of attribute sub-bitstreams indicated in the parameter set are active.

[0012] The parameter set may indicate map information associated with each of a plurality of attribute sub-bitstreams. The message may indicate which map information is active, for example, by indicating which of the plurality of attribute sub-bitstreams indicated by the parameter set is active. The visual volumetric content may be decoded using the active map information, for example, the map information associated with the active attribute sub-bitstream.

[0013] The plurality of attribute sub-bitstreams may indicate, for example, texture information, material information, transparency information, and / or reflectance information associated with the visual volumetric content. The active attribute sub-bitstream may be determined based on the message. The non-active attribute sub-bitstream may be skipped for decoding the visual volumetric content. An attribute sub-bitstream that is indicated by the parameter set but not by the message may be determined to be a non-active attribute sub-bitstream. The non-active attribute sub-bitstream may be skipped for decoding the visual volumetric content.

[0014] For example, deactivation of an attribute sub-bitstream may be indicated by the message by not referring to an indicator associated with the attribute sub-bitstream indicated by the parameter set. Deactivation of an attribute sub-bitstream may be determined, for example, based on bitrate adaptation.

[0015] One or more methods may be implemented by an apparatus comprising one or more processors configured to execute computer-executable instructions stored on a computer-readable medium or a computer program product that, when executed by the one or more processors, perform the one or more methods. Accordingly, the apparatus comprises one or more processors configured to perform the one or more methods. The computer-readable medium or computer program product comprises instructions for causing the one or more processors to execute the one or more methods by executing the instructions. The computer-readable medium may include data content generated according to one or more methods. The signal may include a message according to one or more methods. The device may include a device such as a visual media content processing or coding device. The device may include a television, a mobile phone, a tablet, or a set-top box. The device may include at least one of (i) an antenna configured to receive a signal including data representing an image, (ii) a band limiter configured to limit the received signal to a band of frequencies including data representing an image, and (iii) a display configured to display the image.

[0016] Each feature disclosed in any part of this specification, while described separately / individually, may be implemented in any combination with any other feature disclosed in this specification and / or any other feature disclosed elsewhere that may be implicitly or explicitly referred to in this specification or that may be included within the scope of the disclosed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0017]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0018] The details of the exemplary embodiments will be described with reference to the various figures. It should be understood that this description provides detailed examples of possible implementations, but the details are for illustrative purposes only and do not limit the scope of the present application.

[0019] FIG. 1A is a diagram showing an exemplary communication system 100 that can implement one or more of the disclosed embodiments. The communication system 100 may be a multi-connection system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 may enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT spread OFDM (ZT-UW-DFT-S-OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multi-carrier (FBMC).

[0020] As shown in FIG. 1A, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, the WTRUs 102a, 102b, 102c, 102d may each be referred to as a “station” and / or “STA,” which may be configured to transmit and / or receive wireless signals, and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular telephones, personal digital assistants (PDAs), smartphones, laptops, notebooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, IoT devices, wristwatches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. The WTRUs 102a, 102b, 102c, and 102d may each be interchangeably referred to as a UE.

[0021] The communication system 100 can also include base station 114a and / or base station 114b. Each of base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as CN 106 / 115, Internet 110, and / or other network 112. For example, base stations 114a, 114b can be a base transceiver station (BTS), Node B, eNode B, home Node B, home eNode B, gNB, New Radio (NR) Node B, site controller, access point (AP), wireless router, etc. Although base stations 114a, 114b are each depicted as a single element, it will be understood that base stations 114a, 114b can include any number of interconnected base stations and / or network elements.

[0022] Base station 114a may be part of RAN 104 / 113 which may also include other base stations and / or network elements such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. (not shown). Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies which may be referred to as a cell (not shown). These frequencies may be in an authorized spectrum, an unlicensed spectrum, or a combination of an authorized spectrum and an unlicensed spectrum. A cell can provide coverage for wireless services in a particular geographic area which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0023] Base stations 114a, 114b can communicate with one or more of WTRUs 102a, 102b, 102c, 102d via air interface 116 which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 may be established using any suitable radio access technology (RAT).

[0024] More specifically, as described above, the communication system 100 may be a multi-connection system and may employ one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base stations 114a and the WTRUs 102a, 102b, 102c in the RAN 104 / 113 may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) that can establish the air interfaces 115 / 116 / 117 using Wideband CDMA (WCDMA). WCDMA (registered trademark) may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).

[0025] In one embodiment, the base stations 114a and the WTRUs 102a, 102b, 102c may implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish the air interface 116 using Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).

[0026] In one embodiment, the base stations 114a and the WTRUs 102a, 102b, 102c may implement radio technologies such as NR radio access that can establish the air interface 116 using New Radio (NR).

[0027] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interface utilized by WTRUs 102a, 102b, 102c may be characterized by transmissions sent between multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).

[0028] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement wireless technologies such as IEEE802.11 (i.e., WiFi (Wireless Fidelity)), IEEE802.16 (i.e., WiMAX (Worldwide Interoperability for Microwave Access)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), GSM (registered trademark) (Global System for Mobile communications), GSM Evolution Enhanced Data Rate (EDGE), GSM EDGE (GERAN), etc.

[0029] The base station 114b in Fig. 1A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a local area such as a workplace, home, vehicle, campus, industrial facility, aerial corridor (e.g., for use by drones), road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. As shown in Fig. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.

[0030] RAN 104 / 113 may communicate with CN 106 / 115, which may be any type of network configured to provide voice, data, applications, and / or VoIP services to one or more of WTRUs 102a, 102b, 102c, 102d. The data may have varying quality of service (QoS) requirements such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 may provide call control, billing services, mobile location information services, prepaid originating calls, Internet connectivity, video distribution, etc., and / or may perform high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 may communicate directly or indirectly with other RANs that employ the same or a different RAT than RAN 104 / 113. For example, in addition to being connected to a RAN 104 / 113 that may utilize NR radio technology, CN 106 / 115 may also communicate with another RAN (not shown) that employs GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0031] CN 106 / 115 may also act as a gateway for WTRUs 102a, 102b, 102c, 102d to access PSTN 108, Internet 110, and / or other network 112. PSTN 108 may include a circuit-switched telephone network that provides basic telephone service (POTS). Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as TCP, UDP, and / or IP in the TCP / IP Internet protocol suite. Network 112 may include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, network 112 may include another CN that is connected to one or more RANs that may employ the same or a different RAT than RAN 104 / 113.

[0032] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communication system 100 may include multimode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that may employ a cellular-based radio technology and with a base station 114b that may employ IEEE 802 radio technology.

[0033] FIG. 1B is a system diagram showing an exemplary WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transceiver element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a GPS chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU 102 may comprise any sub-combination of the foregoing elements while maintaining consistency with embodiments.

[0034] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, other types of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120 which may be coupled to the transceiver element 122. Although the processor 118 and the transceiver 120 are shown in FIG. 1B as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0035] The transmitting / receiving element 122 may be configured to transmit and receive signals to / from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmitting / receiving element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmitting / receiving element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0036] Although the transmitting / receiving element 122 is depicted as a single element in FIG. 1B, the WTRU 102 may include any number of transmitting / receiving elements 122. More particularly, the WTRU 102 may employ MIMO technology. Accordingly, in one embodiment, the WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.

[0037] The transceiver 120 may be configured to modulate signals to be transmitted by the transmitting / receiving element 122 and demodulate signals received by the transmitting / receiving element 122. As described above, the WTRU 102 may have multi-mode capabilities. Accordingly, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs such as, for example, NR and IEEE 802.11.

[0038] The processor 118 of the WTRU 102 may be coupled to the speaker / microphone 124, keypad 126, and / or display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and may receive user input data therefrom. The processor 118 may also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Further, the processor 118 may access information in any suitable type of memory, such as the non-removable memory 130 and / or removable memory 132, and store data therein. The non-removable memory 130 may include RAM, ROM, a hard disk, or any other type of memory storage device. The removable memory 132 may include a SIM card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information in a memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown), and store data therein.

[0039] The processor 118 may receive power from the power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 may be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), a solar cell, a fuel cell, and the like.

[0040] Processor 118 may also be coupled to a GPS chipset 136 configured to provide location information (e.g., longitude and latitude) regarding the current location of WTRU 102. In addition to, or instead of, information from GPS chipset 136, WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that WTRU 102 may obtain location information by any suitable location determination method while maintaining consistency with the embodiments.

[0041] Processor 118 may further be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripheral devices 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biosensor, and / or a humidity sensor.

[0042] The WTRU 102 may include a full-duplex radio in which some or all of the transmission and reception of signals (associated with a particular subframe for both, e.g., UL (for transmission) and downlink (for reception, e.g.)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference either via hardware (e.g., a choke) or via signal processing through a processor (e.g., a separate processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for some or all of the transmission and reception of signals (associated with a particular subframe for either, e.g., UL (for transmission) or downlink (for reception, e.g.)).

[0043] Figure 1C is a system diagram showing the RAN 104 and the CN 106, according to one embodiment. As described above, the RAN 104 may employ E-UTRA radio technology and communicate with the WTRU 102a, 102b, 102c via the air interface 116. The RAN 104 may also communicate with the CN 106.

[0044] The RAN 104 may include eNodeBs 160a, 160b, 160c, although it will be understood that the RAN 104 may include any number of eNodeBs while maintaining consistency with the embodiment. Each of the eNodeBs 160a, 160b, 160c may be equipped with one or more transceivers for communicating with the WTRU 102a, 102b, 102c via the air interface 116. In one embodiment, the eNodeBs 160a, 160b, 160c may implement MIMO technology. Thus, the eNodeB 160a, for example, may use multiple antennas to transmit and / or receive radio signals from the WTRU 102a.

[0045] Each of the eNodeBs 160a, 160b, and 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in the UL and / or DL, etc. As shown in Figure 1C, the eNodeBs 160a, 160b, and 160c may communicate with each other via the X2 interface.

[0046] The CN 106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the above elements is shown as part of the CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0047] The MME 162 may be connected to each of the eNodeBs 160a, 160b, and 160c within the RAN 104 via the S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during the initial attach of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide control plane functions for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies such as GSM and / or WCDMA.

[0048] The SGW 164 may be connected to each of the eNodeBs 160a, 160b, 160c within the RAN 104 via the S1 interface. The SGW 164 can generally route and transfer user data packets between the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as anchoring the user plane during handovers between eNodeBs, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, and managing and storing the contexts of the WTRUs 102a, 102b, 102c.

[0049] The SGW 164 may be connected to a PGW 166 that can provide the WTRUs 102a, 102b, 102c access to a packet switched network such as the Internet 110 to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0050] The CN 106 can facilitate communication with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c access to a circuit switched network such as the PSTN 108 to facilitate communication between the WTRUs 102a, 102b, 102c and conventional fixed communication devices. For example, the CN 106 may include, or be able to communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. Further, the CN 106 can provide the WTRUs 102a, 102b, 102c access to other networks 112 that may include other wired and / or wireless networks owned and / or operated by other service providers.

[0051] Although the WTRU is described as a wireless terminal in FIGS. 1A - 1D, in some representative embodiments, it is contemplated that such a terminal may use a wired communication interface to the communication network (e.g., temporarily or permanently).

[0052] In a representative embodiment, the other network 112 may be a WLAN.

[0053] In an infrastructure basic service set (BSS) mode WLAN, there is an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have access or an interface to a distribution system (DS) that carries traffic to and / or from the BSS or another type of wired / wireless network. Traffic destined for the resulting STA may arrive and be sent through the AP. Traffic originating from an STA to a destination outside the BSS may be sent to the AP and delivered to each destination. Traffic between STAs within the BSS can be sent via the AP. For example, the source STA may send traffic to the AP and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent (e.g., directly) between the source STA and the destination STA using direct link setup (DLS). In certain representative embodiments, the DLS can use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode may be referred to herein as the "ad hoc" communication mode.

[0054] When operating in 802.11ac infrastructure mode or using a similar operating mode, the AP can transmit beacons on a fixed channel such as the primary channel. The primary channel can be of a fixed width (e.g., a 20 MHz wide bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In certain representative embodiments, the Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) scheme can be implemented, for example, in an 802.11 system. In the case of CSMA / CA, STAs including the AP (e.g., all STAs) can sense the primary channel. If the primary channel is sensed / detected by a particular STA and / or is determined to be busy, the particular STA can back off. One STA (e.g., just one station) can transmit at any time in a particular BSS.

[0055] A High Throughput (HT) STA can use a 40 MHz wide channel for communication, for example, by combining the primary 20 MHz channel with an adjacent or non - adjacent 20 MHz channel to form a 40 MHz wide channel.

[0056] A Very High Throughput (VHT) STA can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel may be formed by combining eight consecutive 20 MHz channels or by combining two non - consecutive 80 MHz channels, the latter of which may be referred to as an 80 + 80 configuration. In the case of the 80 + 80 configuration, data may be passed through a segment parser after channel encoding, and the segment parser can split the data into two streams. The Inverse Fast Fourier Transform (IFFT) process and time - domain processing can be performed separately on each stream. The streams may be mapped onto two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80 + 80 configuration can be reversed, and the combined data can be transmitted to the Medium Access Control (MAC).

[0057] The sub - 1 GHz operation mode is supported by 802.11af and 802.11ah. In 802.11af and 802.11ah, the channel operation bandwidth and carriers are relatively reduced compared to those used in 802.11n and 802.11ac. 802.11af supports bandwidths of 5 MHz, 10 MHz, and 20 MHz in the TV White Space (TVWS) spectrum, and 802.11ah supports bandwidths of 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz using the non - TVWS spectrum. According to an exemplary embodiment, 802.11ah can support Meter Type Control / Machine Type Communication, such as MTC devices in a macro - coverage area. An MTC device may have limited capabilities, including specific capabilities, e.g., support for a specific and / or limited bandwidth (e.g., only support). An MTC device may include a battery with a battery life exceeding a threshold (e.g., maintaining a very long battery life).

[0058] The WLAN system can support multiple channels and channel bandwidths such as 802.11n, 802.11ac, 802.11af, and 802.11ah, and this WLAN system includes channels that can be designated as primary channels. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or restricted by the STA among all STAs operating in a BSS that supports the minimum bandwidth operating mode. In the example of 802.11ah, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. For example, if the primary channel is busy because an STA (supporting only the 1MHz operating mode) is transmitting to the AP, the entire available frequency band may be considered busy even if most of the frequency band remains idle and available.

[0059] In the United States, the available frequency band that can be used with 802.11ah is 902 MHz to 928 MHz. In Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total bandwidth available for use with 802.11ah is 6 MHz to 26 MHz depending on the country code.

[0060] Figure 1D is a system diagram showing RAN 113 and CN 115 according to one embodiment. As described above, RAN 113 can employ NR radio technology to communicate with WTRUs 102a, 102b, 102c via air interface 116. RAN 113 may also communicate with CN 115.

[0061] RAN 113 may include gNBs 180a, 180b, 180c, although it will be understood that RAN 113 can include any number of gNBs while remaining consistent with one embodiment. Each of gNBs 180a, 180b, 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, 102c via air interface 116. In one embodiment, gNBs 180a, 180b, 180c can implement MIMO technology. For example, gNBs 180a, 108b can transmit and / or receive signals with gNBs 180a, 180b, 180c using beamforming. Thus, for example, gNB 180a can transmit and / or receive wireless signals with WTRU 102a using multiple antennas. In one embodiment, gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers can be on unlicensed spectrum and the remaining component carriers can be on licensed spectrum. In one embodiment, gNBs 180a, 180b, 180c can implement multi-site coordinated (CoMP) technology. For example, WTRU 102a can receive coordinated transmission from gNB 180a and gNB 180b (and / or gNB 180c).

[0062] The WTRUs 102a, 102b, and 102c can communicate with the gNBs 180a, 180b, and 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or the OFDM subcarrier spacing can vary depending on different transmissions, different cells, and / or different portions of the radio transmission spectrum. The WTRUs 102a, 102b, and 102c can communicate with the gNBs 180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing various numbers of OFDM symbols and / or lasting for various lengths of absolute time).

[0063] gNBs 180a, 180b, and 180c may be configured to communicate with WTRUs 102a, 102b, and 102c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNodeBs 160a, 160b, 160c, etc.). In a stand-alone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobility anchor points. In a stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in an unlicensed band. In a non-stand-alone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c while also communicating / connecting with another RAN such as eNodeBs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c can implement the DC principle to communicate with one or more of gNBs 180a, 180b, and 180c and one or more of eNodeBs 160a, 160b, and 160c almost simultaneously. In a non-stand-alone configuration, eNodeBs 160a, 160b, and 160c can act as mobility anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.

[0064] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown), and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, access to control plane information, and routing to mobility management functions (AMFs) 182a, 182b, etc. As shown in FIG. 1D, gNBs 180a, 180b, and 180c can communicate with each other via the Xn interface.

[0065] CN 115 shown in FIG. 1D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and optionally data networks (DNs) 185a, 185b. Although each of the above elements is shown as part of CN 115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0066] AMF182a and 182b are connected to one or more gNBs 180a, 180b, 180c in RAN113 via the N2 interface and can function as control nodes. For example, AMF182a, 182b can be responsible for user authentication of WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling different PDU sessions with different requirements), selection of specific SMFs 183a, 183b, management of the registration area, termination of NAS signaling, mobility management, etc. Network slicing can be used by AMF182a, 182b to customize the CN support for WTRUs 102a, 102b, 102c based on the type of services being used by WTRUs 102a, 102b, 102c. For example, different network slices can be established for various use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on enhanced massive mobile broadband (eMBB) access, services for machine type communication (MTC) access, etc. AMF182 can provide control plane functions for switching between RAN113 and other RANs (not shown) that use other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.

[0067] SMF183a, 183b can be connected to AMF182a, 182b in CN115 via the N11 interface. SMF183a, 183b can also be connected to UPF184a, 184b in CN115 via the N4 interface. SMF183a, 183b can select and control UPF184a, 184b and configure the routing of traffic through UPF184a, 184b. SMF183a, 183b can perform other functions such as management and allocation of the IP address of the WTRU or UE, management of the PDU session, policy enforcement and QoS control, provision of downlink data notifications, etc. The PDU session type can be IP-based, non-IP-based, Ethernet (R) - based, etc.

[0068] UPF 184a and 184b can be connected to one or more gNBs 180a, 180b, 180c within RAN 113 via the N3 interface, which provides access for WTRUs 102a, 102b, 102c to a packet switched network such as the Internet 110 and facilitates communication between the WTRUs 102a, 102b, 102c and IP-enabled devices. UPF 184, 184b can perform other functions such as packet routing and forwarding, user plane policy enforcement, support for multi-home PDU sessions, user plane QoS handling, buffering of downlink packets, and providing a mobility anchor.

[0069] CN 115 can facilitate communication with other networks. For example, CN 115 may include or communicate with an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between CN 115 and the PSTN 108. Additionally, CN 115 can provide access for WTRUs 102a, 102b, 102c to other networks 112 that may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c can be connected to the DNs 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the local data networks (DNs) 185a, 185b.

[0070] In view of FIGS. 1A-1D and the corresponding descriptions of FIGS. 1A-1D, one or more of the functions described herein with respect to one or more of WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b and / or any other devices described herein may be implemented by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate a network and / or WTRU functionality.

[0071] An emulation device can be designed to implement one or more tests of other devices in a laboratory environment and / or an operator network environment. For example, one or more emulation devices can be fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network and perform one or more or all of the functions. One or more emulation devices can be temporarily implemented / deployed as part of a wired and / or wireless communication network to perform one or more or all of the functions. An emulation device can be directly coupled to another device for testing purposes and / or can use wireless communication to execute the test.

[0072] One or more emulation devices can implement one or more functions, including all functions, without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device can be used in a test scenario in a test laboratory and / or in a test scenario in a non-deployed (e.g., for testing) wired and / or wireless communication network to implement tests for one or more components. One or more emulation devices can be test equipment. Wireless communication via direct RF coupling and / or via an RF circuit (which may comprise one or more antennas) can be used by the emulation device such that data can be transmitted and / or received.

[0073] This application describes various aspects, including tools, functions, examples and embodiments, models, approaches, etc. Many of these aspects are specifically described and are often described in a way that may sound restrictive in order to at least indicate individual characteristics. However, this is for clarity of explanation and does not limit the application or scope of these aspects. In fact, all of the various aspects can be combined and exchanged to provide further aspects. Additionally, these aspects can also be combined and exchanged with aspects described in previous applications.

[0074] The aspects described and contemplated in this application can be implemented in various forms. The FIGS. 5-8 described herein may provide some embodiments, but other embodiments are contemplated. The discussions of FIGS. 5 and 8 do not limit the scope of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmission of a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods described herein, and / or a computer-readable storage medium storing a bitstream generated according to any of the methods described herein.

[0075] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture" and "frame" can be used interchangeably.

[0076] Various methods are described herein, and each method includes one or more steps or actions for achieving the described method. Unless a particular order of steps or actions is required for proper operation of the method, the order and / or use of particular steps and / or actions can be changed or combined. Further, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, actions, etc., such as "first decoding" and "second decoding". The use of such terms does not, except where particularly required, imply a changed order of operations. Thus, in this example, the first decoding need not be performed before the second decoding and can occur, for example, before, during, or overlapping with the second decoding.

[0077] The various methods and other aspects described in this application can be used to modify (for use in) modules, such as the intra prediction modules, entropy encoding modules, and / or decoding modules (260, 360, 245, 330) of an encoder 200 and a decoder 300 as shown in FIGS. 2 and 3. Further, the subject matter disclosed herein presents aspects that are not limited to VVC or HEVC, and can be applied to any type, format, or version of video coding, and extensions of such standards and recommendations (including VVC and HEVE), whether existing or to be developed in the future, regardless of whether they are described in a standard or recommendation. Unless otherwise specified or technically excluded, the aspects described in this application can be used individually or in combination.

[0078] Various numerical values are used in the examples described in this application, such as bit value logic (e.g., logic with values of 0 or 1), numerical ranges (e.g., 0 to 255), a payload type value of 12 for a CCC SEI message, and a payload type value of 13 for an AA SEI message. These and other specific values are for illustrative purposes, and the aspects described are not limited to these specific values.

[0079] FIG. 2 is a diagram showing an exemplary video encoder. Although variations of the exemplary encoder 200 are contemplated, the encoder 200 is described below for clarity without explaining all the variations that may be expected.

[0080] Before being encoded, a video sequence can undergo pre-encoding processing (101), for example, applying a color conversion to the input color picture (e.g., conversion from RGB4:4:4 to YCbCr4:2:0), or performing remapping of the input picture components (e.g., using histogram equalization of one of the color components) to make the signal distribution more flexible for compression. Metadata can be attached to the bitstream in association with the pre-processing.

[0081] In encoder 200, pictures are encoded by encoder elements as described below. The pictures to be encoded are divided (202) and processed, for example, in units of coding units (CUs). Each unit is encoded using, for example, either an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction is performed (260). In the inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) whether to use the intra mode or the inter mode to encode a unit, and indicates the intra mode / inter mode decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.

[0082] The prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, that is, the residual is encoded directly without applying the transformation or quantization process.

[0083] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).

[0084] FIG. 3 is a diagram showing an example of a video decoder. In an exemplary decoder 300, a bitstream is decoded by decoder elements as described below. Video decoder 300 generally performs a decoding path that is the reverse of the encoding path, as described in FIG. 2. Encoder 200 can also generally perform video decoding as part of the encoding of video data. For example, encoder 200 can perform one or more of the video decoding steps presented in this application. The encoder reconstructs the decoded image and maintains synchronization with the decoder with respect to, for example, one or more of a reference picture, an entropy encoding context, and other decoder-related state variables.

[0085] In particular, the input to the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Picture partitioning information indicates how a picture is partitioned. Thus, the decoder can partition the picture (335) according to the decoded picture partitioning information. The transform coefficients are inverse quantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (355) to reconstruct the image block. The prediction block can be obtained from intra prediction (360) or motion compensated prediction (i.e., inter prediction) (375) (370). A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0086] The decoded image can further undergo post-decoding processing (385), such as inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping that performs the reverse of the remapping process performed in pre-encoding processing (201). In post-decoding processing, metadata derived in pre-encoding processing and signaled in the bitstream can be used.

[0087] Figure 4 is a diagram showing an example of a system in which various aspects and embodiments described in the present application are implemented. System 400 can be embodied as a device that includes various components described below and is configured to execute one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 400 can be embodied alone or in combination as a single integrated circuit (IC), multiple ICs, and / or individual components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of System 400 are distributed across multiple ICs and / or individual components. In various embodiments, System 400 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, System 400 is configured to implement one or more of the aspects described in the present application.

[0088] System 400 includes, for example, at least one processor 410 configured to execute instructions loaded therein to implement various aspects described herein. The processor 410 can include embedded memory, input / output interfaces, and various other circuits known in the art. System 400 includes at least one memory 420 (e.g., volatile memory devices and / or non-volatile memory devices). System 400 includes a storage device 440 that can include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM (Electrically Erasable Programmable Read-Only Memory), ROM (Read-Only Memory), PROM (Programmable Read-Only Memory), RAM (Random Access Memory), DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), flash, magnetic disk drives, and / or optical disk drives. The storage device 440 can include, by way of non-limiting example, an embedded storage device, a connected storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.

[0089] System 400 includes, for example, an encoder / decoder module 430 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 430 can include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device can include one or both of an encoding module and a decoding module. Further, the encoder / decoder module 430 can be implemented as a separate element of the system 400 or, as is known to those skilled in the art, can be incorporated within the processor 410 as a combination of hardware and software.

[0090] The program code loaded into the processor 410 or the encoder / decoder 430 to execute the various aspects described herein is stored in the storage device 440 and can then be loaded into the memory 420 for execution by the processor 410. According to various embodiments, one or more of the processor 410, the memory 420, the storage device 440, and the encoder / decoder module 430 can store one or more of the various items during the execution of the processes described herein. Such stored items can include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and logical operations.

[0091] In some embodiments, the memory internal to the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide a working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory can be the memory 420 and / or the storage device 440, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM is used as a working memory for video encoding and decoding operations. Video encoding and decoding operations include, for example, MEPG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding and is also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the JVET (Joint Video Experts Team)).

[0092] Inputs to the elements of system 400 can be provided via various input devices, as shown in block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, an RF signal wirelessly transmitted by a broadcast station, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) an HDMI (High Definition Multimedia Interface) input terminal. Other examples not shown in FIG. 4 include composite video.

[0093] In various embodiments, the input device of block 445 has respective input processing elements known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band of frequencies (e.g.) to select a signal frequency band that can be called a channel in certain embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF portion can include a tuner that performs these various functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments can rearrange the order of the above (and other) elements, delete some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting an amplifier or an analog-to-digital converter, for example. In various embodiments, the RF portion includes an antenna.

[0094] Furthermore, the USB and / or HDMI terminals can each include an interface processor for connecting the system 400 to other electronic devices via a USB and / or HDMI connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within the processor 410 as needed. Similarly, aspects of USB or HDMI interface processing can be implemented, as needed, within a separate interface IC or within the processor 410. The demodulated, error-corrected, and de-multiplexed stream is provided to various processing elements, including, for example, the processor 410 and an encoder / decoder 430 operating in combination with memory and storage elements, to process the data stream required for presentation on the output device.

[0095] The various elements of the system 400 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and data can be transmitted between them using an appropriate connection configuration 425, such as an internal bus known in the art that includes an Inter-Integrated Circuit (I2C) bus, wiring, and a printed circuit board.

[0096] The system 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 can include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 can include, but is not limited to, a modem or a network card, and the communication channel 460 can be implemented, for example, within a wired and / or wireless medium.

[0097] In various embodiments, data is streamed or otherwise provided to system 400 using a Wi-Fi network, such as a wireless network like IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via a communication channel 460 and a communication interface 450 that are adapted for Wi-Fi communication. The communication channel 460 of these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 400 using a set-top box that distributes data via the HDMI connection of input block 445. Still other embodiments provide streamed data to system 400 using the RF connection of input block 445. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0098] System 400 can provide output signals to various output devices including display 475, speaker 485, and other peripheral devices 495. The display 475 in various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 can be for a television, a tablet, a laptop, a mobile phone, or other devices. The display 475 can also be integrated with other components (such as in a smartphone) or separately (such as an external monitor for a laptop). Other peripheral devices 495 include, in various examples of embodiments, one or more stand-alone digital video discs (or digital versatile discs) (both terms DVR), disc players, stereo systems, and / or lighting devices. Various embodiments use one or more peripheral devices 495 that provide functions based on the output of system 400. For example, a disc player performs the function of playing back the output of system 400.

[0099] In various embodiments, control signals are communicated between system 400 and display 475, speaker 485, or other peripheral devices 495 using signaling such as an AV link, CEC (Consumer Electronics Control), or other communication protocols that enable device - to - device control regardless of the presence or absence of user intervention. The output devices can be communicatively coupled to system 400 via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, the output devices can be connected to system 400 using communication channel 460 via communication interface 450. The display 475 and speaker 485 can be integrated into a single unit with other components of system 400 within an electronic device such as a television. In various embodiments, the display interface 470 includes a display driver such as a timing controller (T Con) chip.

[0100] Alternatively, the display 475 and the speaker 485 can be separated from one or more of the other components, for example, if the RF portion of the input 445 is part of a separate set-top box. In various embodiments where the display 475 and the speaker 485 are external components, the output signal can be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.

[0101] Embodiments can be implemented by computer software executed by a processor 410, or by hardware, or by a combination of hardware and software. By way of non-limiting example, embodiments can be implemented by one or more integrated circuits. The memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 can be of any type suitable for the technical environment and can include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.

[0102] Decoding is included in various implementations. As used herein, "decoding" can include all or part of a process that is performed, for example, on a received encoded sequence to generate a final output suitable for display. In various embodiments, such a process typically includes one or more processes performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process can also, or instead, include processes performed by the decoders of the various implementations described herein, such as determining which attribute sub-bitstreams to use to decode visual volumetric content based on a message indicating which of the attribute sub-bitstreams indicated in a parameter set associated with the visual volumetric content is active, decoding visual volumetric content using the active attribute sub-bitstreams, receiving a message indicating which of the plurality of attribute sub-bitstreams indicated in a parameter set associated with the visual volumetric content is active, determining, based on the message, the active and non-active attribute sub-bitstreams, receiving the active attribute sub-bitstreams, decoding visual volumetric content using the active attribute sub-bitstreams and skipping the non-active attribute sub-bitstreams, determining the non-active attribute sub-bitstreams based on the message, determining that attribute sub-bitstreams shown in the parameter set but not shown in the message are non-active attribute sub-bitstreams, skipping the non-active attribute sub-bitstreams for decoding visual volumetric content, etc.

[0103] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the term "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will become clear based on the specific context of the description and will be well understood by those skilled in the art.

[0104] Various implementations include encoding. Similar to the above description regarding "decoding", "encoding" as used in the present application can include all or part of the process performed on an input video sequence, for example, to generate an encoded bitstream. In various embodiments, such a process typically includes one or more processes performed by an encoder, such as splitting, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such a process can also or alternatively include processes performed by the encoders of the various implementations described in the present application, such as determining to deactivate an attribute sub-bitstream among a plurality of attribute sub-bitstreams indicated in a parameter set associated with visual volumetric content, determining to deactivate an attribute sub-bitstream among a plurality of attribute sub-bitstreams indicated in a parameter set associated with visual volumetric content based on bitrate adaptation, generating and transmitting a message indicating deactivation of the attribute sub-bitstream, and the like.

[0105] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the term "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the context of a particular description and will be well understood by those skilled in the art.

[0106] Note that the syntactic elements used herein, such as Tables 1 to 7 and the syntactic elements shown in the descriptions and drawings presented herein, are illustrative terms. As such, they do not preclude the use of other syntactic element names.

[0107] It should be understood that when a figure is presented as a flowchart, it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0108] During the quantization process, a balance or trade-off between rate and distortion is typically considered in view of computational complexity constraints. Rate-distortion optimization is usually formulated as minimizing a rate-distortion function that is a weighted sum of the rate and a weighted sum of the distortion. There are various approaches to solving the rate-distortion optimization problem. For example, an approach is based on an extensive test of all encoding options including all modes or coding parameter values to be considered, and fully evaluates the coding cost and the associated distortion of the reconstructed signal after coding and decoding. In particular, a faster approach can also be used to save encoding complexity in the calculation of an approximate distortion based on a prediction or prediction residual signal rather than the reconstructed signal. These two approaches can also be used in combination, such as using the approximate distortion for only some of the possible encoding options and the full distortion for other encoding options. In other approaches, only a subset of the possible encoding options is evaluated. More generally, many approaches employ one of various techniques to perform the optimization, but the optimization is not necessarily a full evaluation of both the coding cost and the associated distortion.

[0109] The implementations and aspects described herein can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of implementation (e.g., described only as a method), the implementation of the described functions can also be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. These methods can be implemented on a processor that generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor includes, for example, a computer, a communication device such as a mobile phone, a personal digital assistant (PDA), and other devices that facilitate the communication of information between end users.

[0110] References to "one embodiment", "an embodiment", "an example", "an implementation" or "implementations", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in connection with the embodiments are included in at least one embodiment. Thus, the appearances of the descriptions of "in one embodiment", "in an embodiment", "in an example", "in an implementation" or "in implementations" and any other variations thereof that appear in various places in this application do not necessarily all refer to the same embodiment or example.

[0111] Furthermore, this application may refer to the "determination" of various information. The determination of information can include, for example, one or more of the estimation of information, the calculation of information, the prediction of information, or the retrieval of information from memory. Acquisition may include reception, retrieval, construction, generation, and / or determination.

[0112] Furthermore, this application may refer to "access" to various information. Access to information can include, for example, one or more of the reception of information, the retrieval of information (e.g., from memory), the storage of information, the transfer of information, the copying of information, the calculation of information, the determination of information, the prediction of information, or the estimation of information.

[0113] Furthermore, this application may refer to the "reception" of various information. Reception, like "access", means a broad term. The reception of information can include, for example, one or more of the access to information or the retrieval of information (e.g., from memory). Furthermore, "reception" is typically included in some form during operations such as, for example, the storage of information, the processing of information, the transmission of information, the transfer of information, the copying of information, the deletion of information, the calculation of information, the determination of information, the prediction of information, or the estimation of information.

[0114] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any of " / ", "and / or", and "at least one" is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first- and second-listed options (A and B), or only the first- and third-listed options (A and C), or only the second- and third-listed options (B and C), or the selection of all three options (A, B, and C). As will be apparent to those skilled in the art and those in the relevant fields, this can be extended for any number of items described.

[0115] Also, as used herein, the term "signal" refers, among other things, to indicating something to the corresponding decoder. For example, in some embodiments, an encoder signal (e.g., to a decoder) changes to a video component sub-bitstream for decoding, a change or active / inactive codec of a video component codec for decoding, a change or active / inactive attribute of a video component attribute for decoding, a change or active / inactive parameter set of an attribute set of a video component for decoding, etc. Thus, in one embodiment, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used (implicit signaling) without sending the specific parameters so that the decoder can recognize and select the specific parameters. By avoiding the transmission of actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the word "signal", but the word "signal" can also be used as a noun in this specification.

[0116] As will be apparent to those skilled in the art, implementations can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry a bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding of the data stream and modulation of a carrier having the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. The signal can be transmitted via various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0117] The examples described herein may be applicable to 3D content. The examples described herein may be applicable to visual volumetric content. Visual volumetric content may be captured in a point cloud. Visual volumetric content may be captured in immersive video content. The examples described herein may be applicable to video-based point cloud compression (V-PCC) or visual volumetric video-based coding (V3C). For example, a particular example may be described from the perspective of V-PCC, but the example is equally applicable to V3C. Thus, in a sense, these terms can be used interchangeably, and an example described from the perspective of V-PCC is equally applicable to V3C.

[0118] 3D point clouds can be used to represent 3D content (e.g., immersive media). A point cloud can include a set of points represented in three-dimensional (3D) space. Each point can be associated with coordinates indicating the position of the point and / or one or more attributes (e.g., the color, transparency, acquisition time, laser reflectance, or material properties of the point). A point cloud can be captured or deployed using, for example, one or more cameras, depth sensors, and / or light detection and ranging (LiDAR) laser scanners. A point cloud may contain multiple points. Each point can be represented by a series of coordinates (e.g., x, y, z coordinates) mapped to 3D space. Points can be generated based on sampling of an object. In an example, the number of points in a point cloud can be on the order of millions or billions. One or more objects and / or scenes can be reconstructed using the point cloud.

[0119] 3D content can be processed using V3C, including decoding and encoding. Volumetric content can be represented and / or compressed, for example, to be efficiently stored and transmitted. Volumetric content can include visual volumetric content. Visual volumetric content can be processed based on visual volumetric video-based coding (V3C) and / or video-based point cloud compression (V-PCC). Visual volumetric content can include V3C content or V-PCC content.

[0120] V3C may be based on point cloud compression. Point cloud compression may support lossy and / or lossless coding (e.g., encoding or decoding) of the geometric coordinates and / or attributes of the point cloud. The point cloud may be deployed to support various applications (e.g., telepresence, virtual reality (VR), large-scale dynamic 3D maps).

[0121] FIG. 5 shows an example of the bitstream structure of V-PCC (video-based point cloud compression). As shown in FIG. 5, a video bitstream may be generated. The V-PCC bitstream may be generated, for example, by multiplexing the generated video bitstream and its respective metadata together.

[0122] The V-PCC bitstream may include a set of V-PCC units, for example, as shown in FIG. 5. Table 1 shows a syntax example of the V-PCC unit, which may include a V-PCC unit header and a V-PCC unit payload. Table 2 shows a syntax example of the V-PCC unit header. Table 3 shows a syntax example of the V-PCC unit payload. The V-PCC unit header may include an indication of the V-PCC unit type (e.g., as shown in Table 2). In the example, a V-PCC unit whose unit type is VPCC_OVD, VPCC_GVD, or VPCC_AVD may include, for example, occupancy, geometry, and / or attribute data units. The three components of the occupancy, geometry, and / or attribute data units can be used to reconstruct visual volumetric content (e.g., represented by a point cloud). The V-PCC attribute unit header may indicate an attribute type and an index that may be associated with the attribute type. There may be multiple instances of the attribute type.

[0123] A payload including occupancy, geometry, and attribute V-PCC units may correspond to a video data unit (e.g., a Network Abstraction Layer (NAL) unit) that can be decoded by a video coder (e.g., a decoder), and may be specified in a corresponding occupancy, geometry, and attribute parameter set V-PCC unit.

[0124] [Table 1]

[0125] [Table 2]

[0126] [Table 3]

[0127] One or more messages (e.g., supplementary additional information messages) can be used to assist one or more processes related to the coding (e.g., encoding or decoding), reconstruction, display, etc. of media content.

[0128] Changes in network delivery conditions can be adapted dynamically. For example, MPEG Dynamic Adaptive Streaming over HTTP (MPEG-DASH) is a delivery format that can adapt dynamically to changes in network delivery conditions, e.g., to provide an end-user experience.

[0129] Dynamic HTTP streaming can deliver multimedia content at one or more bitrates available at the server. The multimedia content may include multiple media components (e.g., audio, video, text). If the media components are different, their characteristics may also be different. The characteristics of the media components can be described, for example, by a media presentation description (MPD).

[0130] FIG. 6 shows an exemplary MPD hierarchical data model. As shown in FIG. 6, the MPD can describe a series of periods (e.g., time intervals). For example, a set of encoded versions of media content components may not change during a period. A period can be associated with a start time and a duration. A period can be composed of one or more adaptation sets (e.g., the adaptation set 1 shown in FIG. 6). A DASH streaming client can be, for example, a WTRU as described herein with respect to FIGS. 1A-1D.

[0131] An adaptation set can represent a set of encoded versions of one or more (e.g., several) media content components that share one or more (e.g., the same) properties such as language, media type, image aspect ratio, role, accessibility, viewpoint, evaluation property, etc. For example, an adaptation set can include various bitrates of the visual components of the multimedia content. An adaptation set may include different bitrates of the audio components (e.g., low-quality stereo and / or high-quality surround sound) of (e.g., the same) multimedia content. An adaptation set may include multiple representations.

[0132] A representation may describe an encodable version of one or more (e.g., several) media components. A representation may differ from other representations, for example, by bitrate, resolution, number of channels, and / or other characteristics. A representation can include one or more segments. Attributes of representation elements (e.g., @id, @bandwidth, @qualityRanking, and @dependencyId) can be used to specify one or more properties of the representation (for example).

[0133] Segments can be retrieved in a request (e.g., an HTTP request). Each segment may include a URL (e.g., an addressable location on a server). Segments can be downloaded, for example, using HTTP GET or HTTP GET with byte ranges.

[0134] The DASH client can parse the MPD XML document. The DASH client can select a collection of adaptation sets (e.g., those suitable for its environment), for example, based on the elements of the adaptation sets. The client can select a (e.g., one) representation within an adaptation set (e.g., within each adaptation set). The client can select a representation based on, for example, the value of the @bandwidth attribute, the client coding capabilities, and / or the client rendering capabilities. The client can download the initialization segment of the selected representation. The client can access the content (e.g., by requesting the entire segment or a byte range of the segment). The client may continue to consume the media content, for example, after the presentation has started or during the presentation. The client can request (e.g., continuously request) media segments and / or parts of media segments during the presentation. The client can play the content according to the timeline of the media presentation. The client can switch from a first representation to a second representation, for example, based on updated information from the client environment. The client can continuously play the content over, for example, two or more periods. The media presentation (e.g., consumed by the client in segments) ends, a period starts, and / or the MPD can be refetched, for example, towards the announced end of the media in the representation.

[0135] Changes to the volumetric content component substream (e.g., changes to active attributes and / or codecs) may be signaled to the decoder. The changes can be caused, for example, by the bitrate adaptation process of the streaming session.

[0136] Multimedia applications such as virtual reality (VR) and immersive 3D graphics can be implemented in or represented by 3D point clouds, enabling an updated form of interaction and / or communication with the virtual world. Dynamic volumetric content (e.g., represented by a point cloud) can generate a large amount of information. Efficient coding algorithms can be used to compress the volumetric content, for example, to reduce the use of storage and / or transmission resources by the volumetric content. For example, the bitstream of compressed dynamic volumetric content may use fewer transmission resources than the bitstream of uncompressed information.

[0137] DASH can be used to carry (e.g., stream) volumetric content information / data. DASH signaling can support DASH volumetric content data streaming. The volumetric content bitstream may include multiple component sub-streams and may increase the number of dimensions for bitrate adaptation. A computing device (e.g., a DASH client such as a smartphone) may have multiple video codec instances per codec. A DASH client (e.g., a WTRU) can select (e.g., dynamically) from among the encoded media versions of the volumetric content components, for example, based on or to manage media processing capacity or capabilities. The DASH client can be a WTRU.

[0138] The dynamic selection of the coded media version of the volumetric content component can be based on, for example, the number of instances supported by the client, the resolution, and / or the codec type. Support for dynamic selection can be implemented, for example, in a high-level syntax. For example, a V-PCC track can include a V-PCC sequence parameter set (SPS) such as VPCC_SPS, which can include a codec identifier (ID) for the video-coded component sub-stream. Switching between different codecs can be implemented, for example, by using multiple VPCC_SPS units with different combinations of codec IDs for the component sub-stream. The content creator may or may not generate various V-PCC tracks that include various combinations of codec and resolution.

[0139] Certain attribute components may not be available for use in the reconstruction of volumetric content. Bitrate adaptation (e.g., available to a streaming client) can be included or implemented by dropping (e.g., not downloading or deactivating) attribute sub-bitstreams that are not available for use in the reconstruction of volumetric content (e.g., point cloud). The number of attributes and their respective indices used in the reconstruction of volumetric content can be stored in the VPCC_SPS unit. The volumetric content decoder can be configured to handle (e.g., process) a change in the number of attributes without re-initializing with a new SPS, based on, for example, a supplementary enhancement information (SEI) message, such as a component codec change (CCC) SEI message.

[0140] A visual media content processing device (e.g., an encoder or a decoder) can drop attributes, which may include, for example, not downloading an attribute or information about or associated with the attribute. A decoder can skip or ignore inactive attributes, which may include, for example, not decoding information about inactive attributes in a video bitstream or information associated with inactive attributes (e.g., an inactive attribute sub-bitstream), not allocating a decoder, not expecting or receiving an update of data, and not using it.

[0141] Sub-streams, sub - streams, and sub-bitstreams can be used interchangeably to refer to a part of a bitstream. A sub - stream is associated with an attribute and may refer to an attribute sub-bitstream.

[0142] An attribute (e.g., as indicated by a parameter set) may characterize volumetric content. An attribute can include scalar or vector properties that can be associated with (e.g., each) point in a volumetric frame, such as color, transparency, reflectivity, texture, surface normal, timestamp, material ID, etc.

[0143] Attribute information can include information about an attribute, such as the number of attributes, attribute index, attribute type, attribute ID, codec ID, persistence of the attribute, number of dimensions or channels of the attribute, attribute partition, etc.

[0144] The parameter set can include attribute information of attributes that characterize volumetric content. The parameter set can include one or more parameters such as V-PCC components (e.g., V-PCC units such as occupancy, geometry, or attribute data units), codecs (e.g., occupancy, geometry, or attribute), geometry (e.g., coordinates of points in a point cloud), attributes (e.g., having a status such as modified, active, inactive, etc.), and attribute information. The parameter set can be of any type or format, such as a sequence parameter set (SPS), a component change parameter set (CCPS), a video-based point cloud compression parameter set (VPS), etc.

[0145] The SEI message can support adaptive streaming of volumetric content. For example, a component codec change (CCC) SEI message can notify a volumetric content decoder of a codec change to one or more visual volumetric content components. The CCC SEI message can reference the volumetric content sequence parameter set. The change in codec ID signaled in the CCC SEI message may be related to (e.g., associated with) the corresponding component signaled in the referenced SPS unit. The volumetric content decoder can instantiate a video coder (e.g., a new video decoder) for each component and codec ID signaled in the received CCC SEI message. Table 4 shows an example syntax of the CCC SEI message.

[0146] A message (e.g., a CCC SEI message) can include a prefix SEI message and may be transmitted in a patch data group (PDG) unit of the PDG_PREFIX_SEI type. For example, the payload type value of a CCC SEI message can be set to value 12. The persistence scope of a CCC SEI message can include the remainder of the bitstream. For example, a change in the codec of a signaled component may persist until the stream ends or another CCC SEI message is detected.

[0147] Bitrate adaptation can include a decision to switch (e.g., change or update) the representation of volumetric content components to another representation. A volumetric content streaming client can decide to switch the representation of visual volumetric content components to another representation. The representations available for switching can be coded using a different codec. For example, a first representation can be encoded using AVC, and a second representation can be encoded using HEVC. A volumetric content streaming client can insert a VPCC_PDG unit that includes a CCC SEI message containing an indication, selection, or codec ID of the codec of the selected representation (e.g., based on an indication to switch to another available representation). The VPCC_PDG unit can be inserted, for example, in front of the volumetric content bitstream that is sent to a coding device (e.g., a decoder). For example, a streaming client can decide to change or update the representation of a volumetric content component to a representation encoded using a codec different from the representation specified in the SPS (e.g., V-PCC SPS). As described herein, a change or update to the representation of a volumetric content component can be performed, for example, in response to a limited network bandwidth during a streaming session. The representation of a volumetric content component can be changed or updated to a lower bitrate representation, which can be coded using a different codec, for example, when the available bandwidth during a streaming session is low (e.g., below a threshold).

[0148] [Table 4]

[0149] The semantics of the fields of the CCC SEI message syntax (such as the example of the CCC SEI message syntax shown in Table 4) may include, for example, one or more of the following.

[0150] Variables such as sps_id can indicate the ID of the visual volumetric content sequence parameter set (e.g., V-PCC sequence parameter set) to which the CCC SEI message is related.

[0151] Variables such as occupancy_codec_change_flag can indicate whether the codec of the occupancy component has been changed. For example, when the value of occupancy_codec_change_flag is 1, it may indicate that the codec has been changed.

[0152] Variables such as Geometry_codec_change_flag can indicate whether the codec of the geometry component has been changed. For example, when the value of geometry_codec_change_flag is 1, it may indicate that the codec has been changed.

[0153] Variables such as attributes_codecs_change_flag can indicate whether the codec of one or more attribute components has been changed. For example, when the value of attributes_codecs_change_flag is 1, it indicates that the codec of at least one attribute has been changed, and when the value of attributes_codecs_change_flag is 0, it may indicate that no codec change has occurred for any attribute.

[0154] Variables such as occupancy_codec_id can indicate the identifier of the new or updated codec of the occupancy map information. For example, occupancy_codec_id can be set to a value in the range from 0 to 255.

[0155] Variables such as Geometry_codec_id can indicate the identifier of a new or updated codec for geometry information. For example, Geometry_codec_id can be set to a value within the range from 0 to 255.

[0156] Variables such as pcm_geometry_codec_change_flag can indicate whether the codec of the geometry data of the PCM-coded point has been changed. For example, when the value of pcm_geometry_codec_change_flag is 1, it may indicate that the codec has been changed. PCM may represent pulse code modulation. In some examples, PCM may represent a point cloud map.

[0157] For example, variables such as pcm_geometry_codec_id may indicate the identifier of the new codec for the geometry data of the PCM-coded point (e.g., when the PCM-coded point is encoded in another stream). For example, the value of pcm_geometry_codec_id can be set to a value within the range from 0 to 255.

[0158] Variables such as attribute_count_minus1 and attribute_count_plus1 may indicate the number of attributes for which codec changes are signaled in the CCC SEI message.

[0159] Variables such as attribute_idx[i] may indicate the attribute index of the i-th attribute of the volumetric content sequence parameter set (e.g., the V-PCC sequence parameter set) of the related SEI message.

[0160] Variables such as attribute_codec_change_flag[i] may indicate whether the codec of the i-th attribute of the SEI message has been changed. For example, when the value of attribute_codec_change_flag[i] is 1, it may indicate that the codec has been changed.

[0161] Variables such as attribute_codec_id[i] may indicate the identifier of the new codec of the attribute video data with index i in the SEI message. For example, attribute_codec_id[i] can be in the range from 0 to 255.

[0162] Variables such as pcm_attribute_codec_change_flag[i] may indicate whether the codec of the attribute video data of the PCM coding point of attribute i of the SEI message has been changed. For example, when the value of pcm_attribute_codec_change_flag[i] is 1, it may indicate that the codec has been changed.

[0163] Variables such as pcm_attribute_codec_id[i] may indicate, for example, if it exists, the identifier of the new codec of the attribute data of the PCM coding point of attribute i in the SEI message (e.g., the PCM coding point is encoded in another stream). For example, the value of pcm_attribute_codec_id[i] can be in the range from 0 to 255.

[0164] The Active Attribute (AA) SEI message can notify (e.g., indicate) to the visual volumetric content decoder that one or more attributes signaled (e.g., in the V-PCC of the codec) have been changed for one or more dual volumetric content components. The AA SEI message may refer to a specific visual volumetric content sequence parameter set. The AA SEI message can include, for example, the attribute indices of the active attributes when it is necessary to activate a subset of the attributes (e.g., only that subset). The sequence parameter set decoder may ignore (e.g., consider non-active) other attributes in the referenced SPS (e.g., attributes not listed in the AA SEI message). Table 5 shows an example syntax of the AA SEI message.

[0165] The AA SEI message may be a prefix SEI message. The AA SEI message may be transmitted in a patch data group unit of type PDG_PREFIX_SEI. The payload type value of the AA SEI message may be, for example, 13. The persistence scope of the AA SEI message can include, for example, the remainder of the bitstream. For example, the signaled active attributes may persist until the end of the stream or until a subsequent AA SEI message is received.

[0166] Bitrate adaptation may include (e.g., may be implemented by) a volumetric content streaming client that determines to drop one or more of the attributes of volumetric content. The volumetric content streaming client can determine to drop one or more attributes of volumetric content, for example, because the network bandwidth is limited. Volumetric content streaming can insert a VPCC_PDG unit including an AA SEI message containing a list of active attributes, for example, when the bitrate adaptation process of the V-PCC streaming client determines to drop one or more attributes of the V-PCC content. Volumetric content streaming can insert a VPCC_PDG unit including an AA SEI message containing a list of active attributes, for example, before the volumetric content bitstream is sent to the decoder. A decoder (e.g., a decoder that receives the VPCC_PDG unit) may skip inactive attributes or not request volumetric content units in the bitstream associated with the inactive attributes.

[0167] The message can communicate information, for example, to assist in coding, reconstructing, and displaying volumetric content. The information can include one or more attributes, attribute information, a display indicating whether an attribute sub-bitstream shown in a parameter set associated with the volumetric content is active or inactive, a display indicating whether an attribute associated with the attribute information of the parameter set is active, a display of the persistence of the message (e.g., until the end of the bitstream or until another message is received), a display of the number of active attributes identified in the parameter set, a display indicating that multiple attributes within the parameter set are active, and / or a display indicating that an attribute is deactivated when the message does not refer to the attribute shown in the parameter set. The message can be of any type or format, such as an SEI message (e.g., a Component Coding Change (CCC) message, an Active Attribute (AA) message, a Parameter Set Activation (PSA) message, etc.). The message can refer to, relate to, or be associated with a parameter set (e.g., a V3C sequence parameter set). The message may be included in a bitstream.

[0168]

Table 5

[0169] The semantics of the fields of the AA SEI message syntax (such as the example of the AA SEI message syntax shown in Table 5) can include, for example, one or more of the following.

[0170] A display such as sps_id may indicate the ID of the visual volumetric content content (e.g., V-PCC) sequence parameter set to which the AA SEI message relates.

[0171] Displays such as the all_attributes_active_flag may indicate whether the attributes signaled in the referenced SPS are active. For example, when the value of the all_attributes_active_flag is 1, it may indicate that all attributes are active. When the value of the all_attributes_active_flag is 0, it may indicate that a subset of the attributes is active.

[0172] Variables such as attribute_count_minus1 and attribute_count_plus1 may indicate the number of active attributes signaled in the AA SEI message.

[0173] Variables such as attribute_idx[i] may indicate the attribute index of the V-PCC SPS of the active attribute at index i of the relevant AA SEI message.

[0174] The semantics of the fields of the AA SEI message syntax in Table 6 are merely examples. A message (e.g., an AA SEI message) may contain less information. For example, the display sps_id may not be included in the message. A message (e.g., an AA SEI message) may contain more information (e.g., the active layer or map information indicated in the parameter set).

[0175] The component change parameter set (CCPS) may contain information about changes applied to the component substream, for example, regarding the volumetric sequence parameter set (e.g., the V-PCC SPS). The CCPS may contain information about codec changes and / or active attributes. The CCPS may be carried, for example, in a V-PCC unit (e.g., a specific type of V-PCC unit). Table 6 shows an example of the CCPS syntax.

[0176]

Table 6-1

[0177]

Table 6-2

[0178] The semantics of the fields of the CCPS syntax (such as the example of the CCPS message syntax shown in Table 6) may include, for example, one or more of the following.

[0179] Variables such as ccps_component_change_parameter_set_id may provide a CCPS identifier so that they can be referenced by other syntax elements. For example, the value of ccps_component_change_parameter_set_id can be in the range from 0 to 255.

[0180] Variables such as ccps_sps_id may indicate the ID of the volumetric content (e.g., V-PCC) sequence parameter set related to CCPS.

[0181] Variables such as ccps_component_codec_change_flag may indicate whether a codec change has occurred for one or more volumetric content (e.g., V-PCC) components. For example, when the value of ccps_component_codec_change_flag is 1, it may indicate that a codec change has occurred.

[0182] Variables such as ccps_active_attributes_change_flag may indicate whether the set of active attributes has been changed. For example, when the value of ccps_active_attributes_change_flag is 1, it may indicate that a change has occurred in the set of active attributes.

[0183] Variables such as the ccps_occupancy_codec_change_flag can indicate whether the codec of the occupancy component has changed. For example, when the value of the ccps_occupancy_codec_change_flag is 1, it may indicate that the codec has been changed.

[0184] Variables such as the ccps_geometry_codec_change_flag can indicate whether the codec of the geometry component has changed. For example, when the value of the ccps_geometry_codec_change_flag is 1, it may indicate that the codec has been changed.

[0185] Variables such as the ccps_attributes_codecs_change_flag can indicate whether the codec of one or more attribute components has changed. For example, when the value of the ccps_attributes_codecs_change_flag is 1, it may indicate that the codec of at least one attribute has been changed. When the value of the ccps_attributes_codecs_change_flag is 0, it may indicate that no codec change has occurred for any attribute.

[0186] Variables such as the ccps_occupancy_codec_id can indicate the identifier of the new codec for occupancy map information. For example, the value of the ccps_occupancy_codec_id can be within the range from 0 to 255.

[0187] Variables such as the ccps_geometry_codec_id can indicate the identifier of the new codec for geometry information. For example, the value of the ccps_geometry_codec_id can be within the range from 0 to 255.

[0188] Variables such as the `ccps_pcm_geometry_codec_change_flag` can indicate whether the codec for the geometry video data of the PCM coding point has been changed. For example, when the value of the `ccps_pcm_geometry_codec_change_flag` is 1, it may indicate that the codec has been changed.

[0189] For example, variables such as the `ccps_pcm_geometry_codec_id` may indicate the identifier of the new codec for the geometry data of the PCM coding point, for example, when the PCM coding point is encoded in another video stream if it exists. For example, the value of the `ccps_pcm_geometry_codec_id` can be in the range from 0 to 255.

[0190] Variables such as `ccps_codec_change_attribute_count_minus1` and `ccps_codec_change_attribute_count_plus1` may indicate the number of attributes for which codec changes are signaled in CCPS.

[0191] Variables such as `ccps_codec_change_attribute_idx[i]` may indicate the attribute index of the V-PCC SPS for the attribute at index `i` in the `codec_change_information()` structure of CCPS.

[0192] Variables such as `ccps_attribute_codec_change_flag[i]` may indicate whether the codec for the `i`-th attribute in the `codec_change_information()` structure of CCPS has been changed. For example, when the value of `ccps_attribute_codec_change_flag[i]` is 1, it may indicate that the codec for the `i`-th attribute has been changed.

[0193] Variables such as ccps_attribute_codec_id[i] may indicate the identifier of the new codec for the attribute data with index i of the CCPS's codec_change_information() structure. For example, the value of ccps_attribute_codec_id[i] can be in the range from 0 to 255.

[0194] Variables such as ccps_pcm_attribute_codec_change_flag[i] may indicate whether the codec of the attribute video data at the PCM coding point of attribute i in the CCPS's codec_change_information() structure has changed. For example, if the value of ccps_pcm_attribute_codec_change_flag[i] is 1, it may indicate that the codec has changed.

[0195] For example, variables such as ccps_pcm_attribute_codec_id[i], if they exist, may indicate the identifier of the new codec for the attribute data at the PCM coding point of attribute i in the CCPS's codec_change_information() structure, for example, when the PCM coding point is encoded in a different stream. For example, the value of ccps_pcm_attribute_codec_id[i] can be in the range from 0 to 255.

[0196] Variables such as ccps_all_attributes_active_flag may indicate whether the attributes signaled in the referenced SPS are active. For example, if the value of ccps_all_attributes_active_flag is 1, it may indicate that all attributes are active. If the value of ccps_all_attributes_active_flag is 0, it may indicate that a subset of the attributes is active.

[0197] Variables such as ccps_active_attribute_count_minus1 and ccps_active_attribute_count_plus1 may indicate the number of active attributes signaled by CCPS.

[0198] Variables such as ccps_active_attribute_idx[i] may indicate the volumetric content (e.g., V-PCC) SPS attribute index of the active attribute at index i of the CCPS's active_attribute_information() structure.

[0199] The parameter set activation (PSA) SEI message may indicate the active parameter set for the volumetric content (e.g., V-PCC) units that may follow the PSA SEI message. Table 7 shows an example syntax of the PSA SEI message.

[0200]

Table 7

[0201] The semantics of the fields of the PSA SEI message syntax (such as the example of the PSA SEI message syntax shown in Table 7) may include, for example, one or more of the following.

[0202] Variables such as active_sequence_parameter_set_id may indicate or may be the same as the value of sps_sequence_parameter_set_id of the active volumetric content SPS (e.g., the V-PCC SPS that needs to be activated). For example, the value of active_sequence_parameter_set_id may be in the range from 0 to 15.

[0203] Variables such as component_change_flag may indicate whether to activate the component change parameter set. For example, when the value of component_change_flag is 1, it may indicate that it is necessary to activate the component change parameter set.

[0204] Variables such as active_component_change_parameter_set_id may indicate or be equal to the value of ccps_component_change_parameter_set_id of the activated CCPS. The value of active_component_change_parameter_set_id can be in the range from 0 to 255.

[0205] The PSA SEI message (e.g., as shown in Table 7, which may be generated by the exemplary encoder 200 and received by the exemplary decoder 300) may indicate which parameter set associated with the volumetric content (e.g., V-PCC) is active for the V-PCC component. The CCC SEI message (e.g., generated by the exemplary encoder 200 and received by the exemplary decoder 300) may indicate which codec should be used to decode the volumetric content component (e.g., the V-PCC unit shown in FIGS. 5, 1, 2, and 3) in the volumetric content (e.g., the V-PCC bitstream shown in FIG. 5).

[0206] FIG. 7 shows an example of a method for processing visual volumetric content based on one or more messages. The method described in FIG. 7 can be performed by an exemplary encoder or an exemplary client device. The examples disclosed herein and other examples can operate according to the exemplary method 700 shown in FIG. 7. Method 700 includes reference numerals 702 and 704. At reference numeral 702, a determination can be made to deactivate an attribute sub-bitstream of one or more attribute sub-bitstreams shown in a parameter set associated with the visual volumetric content. At reference numeral 704, a message indicating the deactivation of the attribute sub-bitstream can be generated and transmitted. The exemplary method 700 can be implemented by an encoder, such as, for example, the exemplary encoder 200. The exemplary method 700 can be implemented by a streaming client, such as, for example, a visual volumetric content streaming client. The exemplary method 700 can be implemented according to the exemplary syntax and semantics described herein for parameter sets including, for example, visual volumetric bitstream SEI messages and / or, for example, CCC messages, AA messages, CCPS, and / or PSA messages.

[0207] FIG. 8 shows an example of a method for processing visual volumetric content based on one or more messages. The method described in FIG. 8 can be applied to a decoder. The examples disclosed herein and other examples can operate according to the exemplary method 800 shown in FIG. 8. Method 800 includes reference numerals 802 and 804. At reference numeral 802, a determination can be made as to which attribute sub-bitstream of one or more attribute sub-bitstreams shown in a parameter set associated with visual volumetrics is to be used to decode the visual volumetric content based on a message indicating which attribute sub-bitstream is active. At reference numeral 804, the visual volumetric content can be decoded using the active attribute sub-bitstream (e.g., the active attribute sub-bitstream determined at 802) based on the message. The exemplary method 800 can be implemented by a decoder such as, for example, the exemplary decoder 300. The exemplary method 800 can be implemented by a media content decoder such as, for example, a visual volumetric content decoder. The exemplary method 800 can be implemented according to the exemplary syntax and semantics described herein for, for example, a visual volumetric bitstream SEI message and / or a parameter set including, for example, a CCC message, an AA message, a CCPS, and / or a PSA message.

[0208] Many embodiments are described herein. The features of the embodiments can be provided singly or in any combination across various claim categories and types. Further, embodiments can include one or more of the features, devices, or aspects described herein, singly or in any combination, across various claim categories and types such as, for example, any of the following.

[0209] The decoder can decode media content based on one or more messages, which can indicate which of one or more attribute sub-bitstreams shown in a parameter set are active. The media content can include, for example, visual volumetric content. The parameter set can be associated with the visual volumetric content. For example, the parameter set can include a visual volumetric video-based parameter set. The decoder can perform decoding such as determining which attribute sub-bitstream to use to decode the visual volumetric content based on one or more messages. A media content decoder, such as an exemplary decoder 300 operating according to the exemplary method shown in FIG. 8, can determine which attribute sub-bitstream to use to decode the visual volumetric content based on a message indicating which of one or more attribute sub-bitstreams shown in a parameter set associated with the media content are active. A media content decoder, such as an exemplary decoder 300 operating according to the exemplary method shown in FIG. 8, can use the active attribute sub-bitstream based on the message to decode the visual volumetric content.

[0210] The attribute sub-bitstreams shown in the parameter set may indicate information regarding the visual volumetric content. The attribute sub-bitstream can include a sub-bitstream associated with an attribute that characterizes the visual volumetric content. The attribute sub-bitstream can include attributes that characterize the media content, such as color, transparency, reflectivity, texture, etc.

[0211] The message can include, for example, a visual volumetric SEI message such as an AA SEI message (e.g., as shown in Table 5), which can refer to a visual volumetric SPS and indicate an attribute index of an active attribute. A message indicating a list of active attribute sub-bitstreams can be received by the decoder. The message can be generated and sent to the decoder, for example, to indicate the inactivation of one or more attribute sub-bitstreams. The decoder can determine the inactive attribute sub-bitstreams based on the message. The decoder can decode the visual volumetric content using the attribute sub-bitstreams indicated as active in the message (e.g., the AA SEI message). The decoder can skip the attribute sub-bitstreams indicated as inactive in the message for decoding the visual volumetric content.

[0212] Decoder tools and techniques including one or more of entropy decoding, inverse quantization, inverse transformation, and differential decoding can be used to enable the method described in FIG. 8 in a decoder. Using these decoding tools and techniques, receiving media content such as visual volumetric content according to the methods described in FIG. 8 or other parts of this specification; decoding the media content; determining which attribute sub-bitstreams to use to decode media content such as visual volumetric content; receiving and analyzing messages, such as SEI messages such as AA SEI messages, CCC SEI messages, and PSA SEI messages; receiving and analyzing parameter sets such as SPS and CCPS that may indicate map information associated with the attribute sub-bitstream, wherein the active map information is indicated by the indicated active attribute sub-bitstream; decoding visual volumetric content using the active attribute sub-bitstream determined based on the message and / or parameter set or the active map information associated with the active attribute sub-bitstream; determining the persistence of the message, such as until the end of the bitstream or until another message arrives; determining the number of active attribute sub-bitstreams in one or more attribute sub-bitstreams indicated in a parameter set such as a parameter set including a VPS having attribute information referenced by a message indicating an active attribute sub-bitstream associated with the visual volumetric content; determining that one or more attribute sub-bitstreams of a parameter set associated with the visual volumetric content are active based on an indicator in the message;Determining that an attribute sub-bitstream is inactive (e.g., based on the attribute sub-bitstream being indicated in a parameter set rather than a message, or, for example, based on a message that does not reference the attribute sub-bitstream or an indicator of the attribute sub-bitstream, or, for example, based on bitrate adaptation), and skipping the inactive attribute sub-bitstream for decoding of visual volumetric content; and enabling one or more of the other decoder operations related to any of the above.;

[0213] An encoder (e.g., the encoder described herein) can encode media content to generate visual volumetric content that can be transmitted as a bitstream, for example, in a streaming service. The encoder or a client device (e.g., an application) can generate and transmit one or more messages that can indicate which of one or more attribute sub-bitstreams indicated in a parameter set are active. The parameter set can be associated with the visual volumetric content. For example, the parameter set can include a visual volumetric video-based parameter set (VPS).

[0214] An encoder or a client device (e.g., an application) can determine which of one or more attribute sub-bitstreams indicated in a parameter set are active or inactive, based on, for example, an evaluation of an attribute sub-bitstream, an operating environment, resource availability, bandwidth attributes, client coding capabilities, and / or client rendering capabilities. In one example, the encoder or client device can determine which of one or more attribute sub-bitstreams indicated in the parameter set to deactivate based on bitrate adaptation in a streaming session. The encoder or client device can indicate the deactivation of one or more attributes based on a determination of which of one or more attribute sub-bitstreams indicated in the parameter set to deactivate. A media content processor (e.g., an encoder such as exemplary encoder 200) operating according to the exemplary method shown in FIG. 7 can determine to deactivate an attribute sub-bitstream of one or more attribute sub-bitstreams indicated in a parameter set associated with visual volumetric content. The media content processor can include a streaming client device as described herein.

[0215] A media content processor (e.g., an encoder such as exemplary encoder 200) operating according to the exemplary method shown in FIG. 7 can generate a message indicating deactivation of one or more attribute sub-bitstreams and can transmit it, for example, to a decoder. The message can include, for example, a visual volumetric SEI message such as an AA SEI message (such as shown in Table 5), which can refer to the VPS and indicate the attribute index of the active attributes. A message including one or more inactive and / or active attribute sub-bitstreams can be transmitted by the encoder to the decoder.

[0216] Encoding tools and techniques including one or more of quantization, entropy encoding, inverse quantization, inverse transform, and differential encoding can be used to enable the method described in FIG. 7 in an encoder. Using these encoding tools and techniques, generating or transmitting media content such as visual volumetric content according to the methods described in FIG. 7 or other parts of this specification; encoding media content; determining which attribute sub-bitstreams to use for encoding media content such as visual volumetric content; generating and transmitting messages, such as SEI messages such as AA SEI messages, CCC SEI messages, and PSA SEI messages; generating and transmitting parameter sets such as SPS and CCPS that may indicate map information associated with the attribute sub-bitstream, wherein the active map information is indicated by the indicated active attribute sub-bitstream; encoding visual volumetric content using the active attribute sub-bitstream or the active map information associated with the active attribute sub-bitstream; indicating the persistence of a message, such as until the end of the bitstream or until another message arrives; indicating the number of active attribute sub-bitstreams in one or more attribute sub-bitstreams indicated in a parameter set (such as a parameter set including a VPS having attribute information referenced by a message indicating the active attribute sub-bitstream) associated with the visual volumetric content; indicating in a message that one or more attribute sub-bitstreams of a parameter set associated with the visual volumetric content are active;Indicating that the attribute sub-bitstream is inactive (e.g., based on the attribute sub-bitstream being indicated in a parameter set rather than a message, or, for example, based on a message that does not refer to the attribute sub-bitstream or an indicator of the attribute sub-bitstream, or, for example, based on bitrate adaptation), and skipping an inactive attribute sub-bitstream for encoding visual volumetric content; and enabling one or more of the other decoder operations associated with any of the above.;

[0217] Syntax elements such as the syntax elements shown in Tables 1 to 7 can be inserted into the signaling, for example, to enable the decoder to identify an active attribute sub-bitstream and / or an indication of the codec and perform a decoding method as shown in FIG. 8. For example, the syntax element can include one or more of the following indications for the decoder to indicate whether one or more of an attribute sub-bitstream, an attribute sub-bitstream ID, a parameter set, an attribute sub-bitstream change indication, an attribute codec change indication, an attribute codec ID, an attribute active indication, an attribute sub-bitstream inactive indication, an active attribute count indication, an active parameter set indication, an active parameter set ID, an active parameter set change indication are active or inactive for use in decoding. As an example, the syntax element can include one or more indications of an attribute index, an attribute state such as active, inactive or a change flag, and / or an attribute count, as described herein, and / or an indication of a parameter used by the decoder to perform one or more of the examples described herein.;

[0218] The method described in FIG. 8 can be selected and / or applied, for example, based on the syntax elements applied in the decoder. For example, the decoder can receive an indication (e.g., in a message or parameter set) indicating a change in the attribute sub-bitstream and / or the active / inactive status of the codec used to decode visual volumetric content. Based on the indication, the decoder can select the attribute sub-bitstream and / or codec used to decode the visual volumetric content, as described in FIG. 8.

[0219] The bitstream or signal can include one or more of the described syntax elements or variations thereof. For example, the bitstream or signal can include a syntax element indicating an active attribute sub-bitstream and / or codec for performing the decoding method as described in FIG. 8.

[0220] The bitstream or signal can include a syntax for transmitting information generated according to one or more examples herein. For example, the information or data can be generated when performing an example as shown in FIGS. 7 and 8, including any example described herein within the scope of the examples shown in FIGS. 7 and 8. The generated information or data can be transmitted in the syntax included in the bitstream or signal.

[0221] Syntax elements that enable the decoder to use an active attribute sub-bitstream and codec to decode visual volumetric components in a manner corresponding to the method used by the encoder can be inserted into the signal. For example, one or more messages and / or parameter sets indicating the attribute sub-bitstream and codec to be used for decoding can be generated using one or more examples herein.

[0222] A method, process, apparatus, media storage instructions, media storage data, or signal for generating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof.

[0223] A method, process, apparatus, media storage instructions, media storage data, or signal for generating and / or transmitting and / or receiving and / or decoding according to any of the described examples.

[0224] A method, process, apparatus, media storage instructions, media storage data, or signal that, although not necessarily limited, conforms to any number or combination of one or more of the following: processing or decrypting media content; performing dynamic adaptation of a point cloud component sub-bitstream in a point cloud streaming service; obtaining an indication of whether at least one attribute signaled in a reference parameter set is inactive; determining to deactivate an attribute sub-bitstream of one or more attribute sub-bitstreams indicated in a parameter set associated with visual volumetric content: generating and transmitting a message indicating deactivation of the attribute sub-bitstream; obtaining an indication of active attributes (e.g., the number of active attributes and their respective attribute indices) if the indication indicates that at least one attribute of the reference parameter set is inactive; identifying inactive attributes, for example, based on the indication of active attributes; skipping inactive attributes of the reference parameter set during decoding; obtaining a parameter set associated with visual volumetric content; receiving a message indicating which attribute sub-bitstreams of one or more attribute sub-bitstreams indicated in the parameter set are active; determining active and inactive attribute sub-bitstreams based on the message; determining which attribute sub-bitstreams to use to decrypt visual volumetric content based on a message indicating which attribute sub-bitstreams of one or more attribute sub-bitstreams indicated in a parameter set associated with visual volumetric content are active; decrypting visual volumetric content using the active attribute sub-bitstreams based on the message;Decoding visual volumetric content using an active attribute sub-bitstream and skipping non-active attribute sub-bitstreams; Signaling or receiving a message such as an SEI message in a bitstream; Signaling or receiving such a message having a persistence scope that lasts until the bitstream ends or another message different from the message is received; Signaling or receiving a message comprising an indicator indicating the number of active attribute sub-bitstreams among a plurality of attribute sub-bitstreams indicated in a parameter set associated with visual volumetric content; Signaling or receiving a parameter set including a VPS containing attribute information; Signaling or receiving a message referring to a part of the attribute information within the VPS of the active attribute sub-bitstream; Signaling or receiving a message including an indicator indicating that a plurality of attribute sub-bitstreams indicated in the parameter set are active; Signaling or receiving a parameter set indicating map information associated with each attribute sub-bitstream of a plurality of attributes; Signaling or receiving a message indicating which map information is active by indicating which of the plurality of attribute sub-bitstreams indicated in the parameter set is active; Decoding visual volumetric content using the active map information associated with the active attribute sub-bitstream; Signaling and / or receiving one or more attribute sub-bitstreams indicating, for example, texture information, material information, transparency information, and / or reflectance information related to or characterizing visual volumetric content; Determining non-active attribute sub-bitstreams based on a message; Determining as non-active sub-bitstreams attribute sub-bitstreams indicated in the parameter set but not indicated in the message;Skipping an inactive attribute sub-bitstream for decoding visual volumetric content; indicating or receiving an indication of inactivation of an attribute sub-bitstream in a message; determining that an attribute sub-bitstream is inactivated when the attribute sub-bitstream or an indicator of the attribute sub-bitstream is not referenced in the message; and / or determining that an attribute sub-bitstream is inactivated based on bitrate adaptation.

[0225] A television, set-top box, mobile phone, tablet, or other electronic device that performs dynamic adaptation of visual volumetric content, such as a visual volumetric component sub-bitstream in a visual volumetric streaming service, according to any of the examples described herein.

[0226] A television, set-top box, mobile phone, tablet, or other electronic device that performs dynamic adaptation of visual volumetric content, such as a visual volumetric component sub-bitstream in a visual volumetric streaming service, according to any of the examples described herein, and displays the resulting visual representation (e.g., using a monitor, screen, or other type of display).

[0227] A television, set-top box, mobile phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal including an encoded volumetric frame and performs dynamic adaptation of visual volumetric content, such as a visual volumetric component sub-bitstream in a visual volumetric streaming service, according to any of the examples described herein.

[0228] Wirelessly receive (e.g., using an antenna) a signal comprising an encoded volumetric frame according to any of the examples described herein, and perform dynamic adaptation of visual volumetric content such as a visual volumetric component sub-bitstream in a visual volumetric streaming service, in a television, set-top box, mobile phone, tablet, or other electronic device.

[0229] Although features and elements are described above in specific combinations, one of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Also, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile disks (DVDs)). A processor in conjunction with software may be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. An apparatus for decrypting media content, comprising one or more processors, wherein the one or more processors obtain a parameter set associated with visual volumetric content, receive a SEI (supplemental enhancement information) message indicating which of a plurality of attribute sub-bitstreams indicated in the parameter set is active, determine an active attribute sub-bitstream and an inactive attribute sub-bitstream based on the SEI message, decrypt the visual volumetric content using the active attribute sub-bitstream and skip the inactive attribute sub-bitstream The apparatus is configured as follows.

2. The apparatus according to claim 1, wherein the SEI message is signaled in a bitstream.

3. The apparatus according to claim 1, wherein the SEI message has a persistence scope that lasts until the end of the bitstream.

4. The apparatus according to claim 1, wherein the SEI message has a persistence scope that lasts until another SEI message different from the SEI message is received.

5. The apparatus according to claim 1, wherein the SEI message includes an indicator indicating the number of active attribute sub-bitstreams among a plurality of attribute sub-bitstreams indicated in a parameter set associated with the visual volumetric content.

6. The parameter set includes a visual volumetric video-based parameter set (VPS) including attribute information about the plurality of attribute sub-bitstreams, The apparatus according to claim 1, wherein the SEI message refers to attribute information about a subset of the plurality of attribute sub-bitstreams.

7. A method for decrypting media content, comprising obtaining a parameter set associated with visual volumetric content, receiving a SEI (supplemental enhancement information) message indicating which of a plurality of attribute sub-bitstreams indicated in the parameter set is active, Determining an active attribute sub-bitstream and an inactive attribute sub-bitstream based on the SEI message; Decoding the visual volumetric content using the active attribute sub-bitstream and skipping the inactive attribute sub-bitstream; A method comprising the steps of.

8. The method according to claim 7, wherein the SEI message is signaled in a bitstream.

9. The method according to claim 7, wherein the SEI message has a persistence scope that lasts until the end of the bitstream.

10. The method according to claim 7, wherein the SEI message has a persistence scope that lasts until another SEI message different from the SEI message is received.

11. The method according to claim 7, wherein the SEI message includes an indicator indicating the number of active attribute sub-bitstreams among a plurality of attribute sub-bitstreams indicated in a parameter set associated with the visual volumetric content.

12. The parameter set includes a visual volumetric video-based parameter set (VPS) that includes attribute information about the plurality of attribute sub-bitstreams, The method according to claim 7, wherein the SEI message refers to attribute information about a subset of the plurality of attribute sub-bitstreams.

Citation Information

Patent Citations

  • Instructions and activation of parameter sets for video coding.

    JP2015529436A

  • Color conversion encoding method, decoding method and corresponding equipment

    JP2016529787A