Low-complexity multiplane image profiles
The low-complexity MPI profile for MPEG immersive video encoding addresses the challenge of high computational demands in volumetric video decoding, allowing efficient processing on devices with limited power by defining parameters like MPI layers and reference cameras.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-07-11
- Publication Date
- 2026-07-29
AI Technical Summary
Existing video coding systems are not suitable for certain types of decoding/rendering devices, particularly those with limited computational power, due to high complexity in processing volumetric video content.
A low-complexity multi-plane image (MPI) profile is introduced for MPEG immersive video (MIV) encoding, which includes parameters such as the number of MPI layers, layer arrangement, and reference cameras, allowing devices to encode and decode volumetric video efficiently, even with limited computing power.
The MPI profile reduces rendering complexity, enabling efficient decoding and rendering of volumetric video on less powerful devices by limiting computational requirements.
Smart Images

Figure 2026525331000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of European Patent Application No. 23306246.2, filed on July 19, 2023, the content of which is incorporated herein by reference.
Background Art
[0002] Video coding systems are used to compress digital video signals, for example, to reduce the storage capacity and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block - based systems, wavelet - based systems, and / or object - based systems. Some video coding mechanisms may not be suitable for certain types of decoding / rendering devices.
Summary of the Invention
[0003] Systems, methods, and means for providing and / or using a low - complexity multi - plane image (MPI) profile (e.g., for MPEG immersive video (MIV) encoded content) are disclosed.
[0004] For example, a video encoding and / or decoding device may include a processor. The device can be configured to receive a representation of an MIV-encoded volumetric video, which can be encoded using the MPI format. The representation of the MIV-encoded volumetric video may be received via a single video bitstream. The device can receive indications of a set of MPI parameters and / or a set of MIV parameters associated with the MIV-encoded volumetric video. The device can determine that the set of MPI parameters and / or the set of MIV parameters is associated with a low-complexity configuration. A low-complexity configuration may limit the complexity of rendering the MIV-encoded volumetric video. Based on the set of MPI parameters and / or the set of MIV parameters, the device can encode and / or decode the MIV-encoded volumetric video.
[0005] One or more functions may be included. In the example, the device may receive at least one constraint on a set of MPI parameters and / or a set of MIV parameters. Based on at least one constraint, the device may determine that the set of MPI parameters and / or a set of MIV parameters is associated with a low-complexity configuration. The device may receive an indication of an MIV profile. An MIV profile may be associated with a low-complexity configuration. Based on the indication of the MIV profile, the device may determine at least one parameter from the set of MPI parameters and / or a set of MIV parameters. Indications of a set of MPI parameters and / or a set of MIV parameters can be received via SEI (supplemental enhancement information) messages.
[0006] The set of MPI parameters associated with volumetric video encoded in MIV may include at least one of the following: the number of MPI layers, the layer arrangement, and the number of reference cameras. The number of MPI layers may be associated with the maximum number of layers based on a low-complexity configuration. The device can determine that the set of MPI parameters and the set of MIV parameters are associated with a low-complexity configuration based on the number of MPI layers. The number of MPI layers is 16 or 32. The set of MIV parameters associated with volumetric video encoded in MIV may include at least one of the following: the total number of layers, the frame packing indication, and the level limit. The level limit may be associated with a single atlas, pixel rate, or bitrate.
[0007] A video encoding and / or decoding method may include receiving a representation of a volumetric video encoded in MIV, which can be encoded using the MPI format. The representation of the volumetric video encoded in MIV may be received via a single video bitstream. The method may include receiving an indication of a set of MPI parameters and / or a set of MIV parameters associated with the volumetric video encoded in MIV. The method may include encoding that determines that the set of MPI parameters and / or the set of MIV parameters are associated with a low-complexity configuration. A low-complexity configuration may limit the rendering complexity of the volumetric video encoded in MIV. The method may include encoding and / or decoding the representation of the volumetric video encoded in MIV based on the set of MPI parameters and / or the set of MIV parameters.
[0008] One or more functions may be included. For example, this method may include receiving at least one constraint on a set of MPI parameters and / or a set of MIV parameters. This method may include determining, based on at least one constraint, that the set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration. This method may include receiving an indication of an MIV profile. An MIV profile may be associated with a low-complexity configuration. This method may include determining at least one parameter from the set of MPI parameters and / or a set of MIV parameters based on the indication of the MIV profile. Indications of a set of MPI parameters and / or a set of MIV parameters can be received via SEI (supplemental enhancement information) messages.
[0009] The set of MPI parameters associated with volumetric video encoded in MIV may include at least one of the following: the number of MPI layers, the layer arrangement, and the number of reference cameras. The number of MPI layers may be associated with the maximum number of layers, based on a low-complexity configuration. This method may include determining, based on the number of MPI layers, that the set of MPI parameters and the set of MIV parameters are associated with a low-complexity configuration. The number of MPI layers may be 16 or 32.
[0010] The set of MIV parameters associated with volumetric video encoded in MIV can include the number of complete layers, frame packing indications, and at least one level limit. The level limit can be associated with a single atlas, pixel rate, or bitrate.
[0011] This method may include receiving a patch atlas using volumetric video encoded in MIV. The patch atlas may include one or more patches. One or more patches may be associated with texture information or transparency information of at least one layer of the volumetric video encoded in MIV. Each patch of one or more patches may represent an entire layer. This method may further include encoding and / or decoding a representation of the volumetric video encoded in MIV based on the patch atlas.
[0012] For example, the device can generate a volumetric video encoded in MIV, which is then encoded using the MPI format. The device can determine a set of MPI parameters associated with the MIV-encoded volumetric video and a set of MIV parameters associated with the MIV-encoded volumetric video.
[0013] The device can determine that a set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration and can be transmitted to a less computationally powerful device. The device can transmit (for example, via a bitstream) a representation of volumetric video encoded in MIV, along with indications of the set of MPI parameters and / or the set of MIV parameters. The MPI parameters may include the number of MPI layers, indications of the arrangement of the MPI layers, and / or the number of reference cameras. The MIV parameters may include the total number of layers, indications of frame packing, and / or multiple level limits.
[0014] In the example, the device can send an indication (e.g., a tag) that a set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration.
[0015] Indications for a set of MPI parameters and / or a set of MIV parameters may indicate the level of the MIV profile.
[0016] In the example, the device can send indications of a set of MPI parameters and / or a set of MIV parameters using SEI (supplemental enhancement information) messages. In the example, the set of MPI parameters and / or a set of MIV parameters can indicate that those sets of parameters are associated with a low-complexity configuration.
[0017] In the example, the decoder can receive (for example, in a bitstream) a representation of a volumetric video encoded in MIV using the MPI format, and / or an indication of a set of MPI parameters and a set of MIV parameters associated with the MIV-encoded volumetric video. In the example, the decoder may have limited computing power to decode MPI content encoded in MIV.
[0018] In the example, the decoder can determine that the MPI parameters and / or MIV parameters are associated with a low-complexity configuration.
[0019] In the example, the decoder may receive an indication that a set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration.
[0020] The decoder can decode and / or render the representation. The representation can be decoded and / or rendered using a set of MPI parameters and / or a set of MIV parameters associated with the low-complexity configuration.
[0021] The systems, methods, and means described herein may include a decoder. In some examples, the systems, methods, and means described herein may include an encoder. In some examples, the systems, methods, and means described herein may include a signal (e.g., a signal from an encoder and / or a signal received by a decoder). The computer-readable medium may include instructions that cause one or more processors to execute the methods described herein. The computer program product may include instructions that cause one or more processors to execute the methods described herein when the program is executed by one or more processors.
Brief Description of the Drawings
[0022] [Figure 1A] FIG. is a system diagram showing an exemplary communication system in which one or more of the disclosed embodiments may be implemented. [Figure 1B] FIG. is a system diagram showing an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1C] FIG. is a system diagram showing an exemplary radio access network (RAN) and an exemplary core network (CN) that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 1D] FIG. is a system diagram showing a further exemplary RAN and a further exemplary CN that may be used within the communication system shown in FIG. 1A, according to one embodiment. [Figure 2] FIG. is a diagram showing an exemplary video encoder. [Figure 3] FIG. is a diagram showing an exemplary video decoder. [Figure 4] FIG. is a diagram showing an example of a system in which various aspects and examples may be implemented. [Figure 5] FIG. is a diagram showing an example of a multi-plane image (MPI) slice using a perspective camera projection model. [Figure 6]This is a diagram showing an example of a multi-spherical image (MSI) slice using an orthographic cylindrical projection method.
Embodiments for Carrying Out the Invention
[0023] A more detailed understanding will be obtained from the following description, which is shown by way of example together with the accompanying drawings.
[0024] FIG. 1A is a diagram showing an exemplary communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 can be a multi-connection system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 can employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT-spread OFDM (ZT UW DFT-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multi-carrier (FBMC).
[0025] As shown in Figure 1A, the communication system 100 may include radio transceiver units (WTRUs) 102a, 102b, 102c, and 102d, RAN 104 / 113, CN 106 / 115, the public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it should be understood that the disclosed embodiments intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a radio environment. For example, WTRU102a, 102b, 102c, and 102d, any of which may be called “stations” and / or “STAs,” may be configured to transmit and / or receive radio signals and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, radio sensors, hotspots or Mi-Fi devices, IoT (Internet of Things) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other radio devices operating in an industrial and / or automated processing chain context), consumer electronics devices, and devices operating on commercial and / or industrial radio networks. Any of WTRU102a, 102b, 102c, and 102d may interchangeably be called UEs.
[0026] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as CN 106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be transceiver base stations (BTS), node B, enode B, home node B, home enode B, gNB, NR node B, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are shown as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0027] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown) such as base station controllers (BSCs), radio network controllers (RNCs), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be called cells (not shown). These frequencies may be in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrum. Cells may provide coverage for radio services to a particular geographic area, which may be relatively fixed or may change over time. Cells may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology, which may utilize multiple transceivers per sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0028] Base stations 114a, 114b can communicate with one or more of WTRUs 102a, 102b, 102c, 102d via an air interface 116 which may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0029] More specifically, as described above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a and WTRU 102a, 102b, and 102c in RAN 104 / 113 may implement radio technologies such as Universal Mobile Communications System (UMTS) Terrestrial Radio Access (UTRA) which can establish air interfaces 115 / 116 / 117 using broadband CDMA (WCDMA®). WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0030] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as Advanced UMTS Terrestrial Radio Access (E-UTRA) that can establish an air interface 116 using Long-Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).
[0031] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using New Radio (NR).
[0032] In one embodiment, base station 114a and WTRU 102a, 102b, 102c can implement multiple radio access technologies. For example, base station 114a and WTRU 102a, 102b, 102c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Thus, the air interface utilized by WTRU 102a, 102b, 102c may be characterized by transmissions made to and from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNB and gNB).
[0033] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c can implement wireless technologies such as IEEE 802.11 (i.e., WiFi (Wireless Fidelity)), IEEE 802.16 (i.e., WiMAX (Worldwide Interoperability for Microwave Access)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), GSM (Registered Trademark) (Global System for Mobile communications), GSM Advanced High Speed Data Rate (EDGE), and GSM EDGE (GERAN).
[0034] In Figure 1A, base station 114b may be, for example, a wireless router, home node B, home enode B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in local areas such as workplaces, homes, vehicles, premises, industrial facilities, aerial corridors (for use by drones, for example), and roads. In one embodiment, base station 114b and WTRU 102c, 102d can implement wireless technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 114b and WTRU 102c, 102d can implement wireless technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRU 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not need to access the internet 110 via CN 106 / 115.
[0035] RAN104 / 113 may communicate with CN106 / 115, which may be any type of network configured to provide voice, data, application, and / or VoIP (Voice over Internet Protocol) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have varying quality of service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location services, prepaid calling, internet connectivity, video distribution, and / or implement high-level security functions such as user authentication. Although not shown in Figure 1A, it should be understood that RAN104 / 113 and / or CN106 / 115 may communicate directly or indirectly with other RANs employing the same or different RATs as RAN104 / 113. For example, in addition to connecting to RAN104 / 113, which may utilize NR radio technology, CN106 / 115 may also communicate with another RAN (not shown) employing GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0036] CN106 / 115 can also act as a gateway for WTRU102a, 102b, 102c, 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing simple telephone services (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as TCP, UDP, and / or IP within the TCP / IP Internet Protocol Suite. Network 112 may include wired and / or wireless networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that can employ the same RAT as RAN104 / 113 or a different RAT.
[0037] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 can include multimode capability (for example, WTRUs 102a, 102b, 102c, and 102d can include multiple transceivers for communicating with different radio networks via different radio links). For example, WTRU 102c shown in Figure 1A may be configured to communicate with base station 114a which can employ cellular-based radio technology and with base station 114b which can employ IEEE 802 radio technology.
[0038] Figure 1B is a system diagram showing an exemplary WTRU 102. As shown in Figure 1B, the WTRU 102 may include, in particular, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a GPS chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any partial combination of the above elements while remaining consistent with the embodiment.
[0039] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, another type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to a transceiver 120 that can be coupled to the transmit / receive element 122. Although the processor 118 and transceiver 120 are shown as separate components in Figure 1B, it should be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0040] The transmitting / receiving element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmitting / receiving element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In another embodiment, the transmitting / receiving element 122 may be configured to transmit and / or receive both RF signals and optical signals. It should be understood that the transmitting / receiving element 122 may be configured to transmit and / or receive any combination of radio signals.
[0041] Although the transmit / receive element 122 is shown as a single element in Figure 1B, the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals via the air interface 116.
[0042] The transceiver 120 may be configured to modulate the signal to be transmitted by the transmitting / receiving element 122 and to demodulate the signal received by the transmitting / receiving element 122. As described above, the WTRU 102 may have multimode capability. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0043] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input data from them. The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and store data therein. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or other types of memory storage devices. Removable memory 132 may include a subscriber identification module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown), and store data therein.
[0044] The processor 118 can receive power from the power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 can be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0045] The processor 118 may also be coupled to a GPS chipset 136 which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may obtain location information by any preferred location determination method while remaining consistent with the embodiment.
[0046] The processor 118 may be further coupled to other peripherals 138 which may include one or more software and / or hardware modules that provide additional features, functions and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, e-compass, satellite transceiver, digital camera (for photos and / or videos), USB port, vibration device, television transceiver, hands-free headset, Bluetooth® module, frequency modulation (FM) radio unit, digital music player, media player, video game player module, internet browser, virtual reality and / or augmented reality (VR / AR) device, activity tracker, etc. Peripherals 138 may include one or more sensors which may be one or more of a gyroscope, accelerometer, Hall effect sensor, orientation sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biosensor, and / or humidity sensor.
[0047] WTRU102 may include a full-duplex radio where the transmission and reception of some or all of the signal (related to a particular subframe for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference through signal processing either through hardware (e.g., chokes) or a processor (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU102 may include a half-duplex radio for the transmission and reception of some or all of the signal (related to a particular subframe for either UL (e.g., for transmission) or downlink (e.g., for reception)).
[0048] Figure 1C is a system diagram showing RAN104 and CN106 according to one embodiment. As described above, RAN104 can employ E-UTRA radio technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. RAN104 may also communicate with CN106.
[0049] RAN104 may include enodes B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of enodes B while remaining consistent with the embodiment. Each of enodes B160a, 160b, and 160c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, enodes B160a, 160b, and 160c can implement MIMO technology. Thus, enode B160a may, for example, use multiple antennas to transmit and / or receive radio signals from WTRU102a.
[0050] Each of the e-nodes B160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. As shown in Figure 1C, the e-nodes B160a, 160b, and 160c can communicate with each other via the X2 interface.
[0051] The CN106 shown in Figure 1C may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. Although each of the above elements is shown as part of CN106, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0052] The MME162 can be connected to each of the e-nodes B160a, 160b, and 160c in RAN104 via the S1 interface and can act as a control node. For example, the MME162 may be responsible for authenticating users of WTRU102a, 102b, and 102c, activating / deactivating bearers, and selecting a specific serving gateway for the initial attachment of WTRU102a, 102b, and 102c. The MME162 can provide control plane functionality for switching between RAN104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.
[0053] The SGW164 can be connected to each of the e-nodes B160a, 160b, and 160c in RAN104 via the S1 interface. The SGW164 can generally route and forward user data packets to and from WTRU102a, 102b, and 102c. The SGW164 can perform other functions such as anchoring the user plane during handover between e-nodes B, triggering paging when DL data is available for WTRU102a, 102b, and 102c, and managing and remembering the context of WTRU102a, 102b, and 102c.
[0054] SGW164 may be connected to PGW166, which can provide WTRU102a, 102b, and 102c with access to a packet-switched network such as the Internet 110 to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0055] CN106 can facilitate communication with other networks. For example, CN106 can give WTRU102a, 102b, and 102c access to circuit-switched networks such as PSTN108 to facilitate communication between WTRU102a, 102b, and 102c and conventional fixed communication devices. For example, CN106 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN106 and PSTN108. Furthermore, CN106 can give WTRU102a, 102b, and 102c access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0056] Although the WTRU is shown as a wireless terminal in Figures 1A to 1D, in some typical embodiments where such a terminal can be used (for example, temporarily or permanently), wired communication is considered to interface with the communication network.
[0057] In a typical embodiment, the other network 112 may be a WLAN.
[0058] A WLAN in Infrastructure Basic Service Set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interfaces with a distribution system (DS) or another type of wired / wireless network that carries traffic entering and leaving the BSS. Traffic originating from outside the BSS to an STA may arrive through the AP and be sent to the STA. Traffic originating from an STA to a destination outside the BSS may be sent to the AP for distribution to its respective destination. Traffic between STAs within the BSS may be sent through the AP, for example, here, a source STA may send traffic to the AP, and the AP may send traffic to the destination STA. Traffic between STAs within the BSS is considered and / or sometimes referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between a source STA and a destination STA (e.g., directly between them) using a Direct Link Setup (DLS). In some typical embodiments, the DLS may be an 802.11e DLS or an 802.11z Tunneled DLS (TDLS). A WLAN using Independent BSS (IBSS) mode may not have access points (APs), and STAs within or using IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode is sometimes referred to as the “ad-hoc” communication mode in this specification.
[0059] When using the 802.11ac infrastructure operating mode or a similar operating mode, an AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel may be of a fixed width (e.g., a 20 MHz bandwidth) or a dynamically set width via signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In some typical embodiments, Carrier Sensitivity Multiple Access / Collision Avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. In CSMA / CA, an STA, including the AP (e.g., any STA), can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that particular STA can backoff. One STA (e.g., just one station) can transmit at a given time in a given BSS.
[0060] High-throughput (HT) STAs can use a 40MHz wide channel for communication via a combination of primary 20MHz channels, for example, with adjacent or non-adjacent 20MHz channels, to form a 40MHz wide channel.
[0061] Extremely high throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. 160 MHz channels can be formed by combining eight consecutive 20 MHz channels, or by combining two discontinuous 80 MHz channels, sometimes referred to as an 80+80 configuration. In an 80+80 configuration, data can be passed through a segment parser that can split the data into two streams after channel coding. Inverse fast Fourier transform (IFFT) processing and time-domain processing can be performed separately for each stream. Streams can be mapped onto two 80 MHz channels, and data can be transmitted by a transmitting STA. At the receiver of a receiving STA, the operation described above for the 80+80 configuration can be reversed, and the combined data can be sent to a media access control (MAC).
[0062] Sub-1GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths using the non-TVWS spectrum. According to a typical embodiment, 802.11ah can support meter-type control / machine-type communications, such as MTC devices in a macro coverage area. MTC devices may have limited capabilities, including support for some and / or limited bandwidths (e.g., support only that much). MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).
[0063] WLAN systems that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only) the 1MHz mode. Carrier detection and / or network allocation vector (NAV) settings may depend on the status of the primary channel. For example, if the primary channel is busy for an STA (which only supports 1MHz operating mode), a large portion of the frequency band remains idle, and transmitting the entire available frequency band to the AP may be considered busy, even if it could be available.
[0064] In the United States, the available frequency band that can be used by 802.11ah ranges from 902 MHz to 928 MHz. In South Korea, the available frequency band ranges from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band ranges from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah ranges from 6 MHz to 26 MHz, depending on the country code.
[0065] Figure 1D is a system diagram showing RAN113 and CN115 according to one embodiment. As described above, RAN113 can employ NR radio technology to communicate with WTRU102a, 102b, and 102c via air interface 116. RAN113 may also communicate with CN115.
[0066] RAN113 may include gNB180a, 180b, and 180c, but it will be understood that RAN113 may include any number of gNBs while remaining consistent with the embodiment. Each of the gNB180a, 180b, and 180c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, gNB180a, 180b, and 180c can implement MIMO technology. For example, gNB180a and 180b may utilize beamforming to transmit and / or receive signals from WTRU102a, 102b, and 102c. Thus, gNB180a may, for example, use multiple antennas to transmit and / or receive radio signals from WTRU102a. In one embodiment, gNB180a, 180b, and 180c can implement carrier aggregation technology. For example, gNB180a can transmit multiple component carriers to WTRU102a (not shown). A subset of these component carriers may be on the unlicensed spectrum, while the remaining component carriers may be on the licensed spectrum. In one embodiment, gNB180a, 180b, and 180c can implement coordinated multipoint (CoMP) technology. For example, WTRU102a can receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).
[0067] WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using transmissions related to scalable numerology. For example, OFDM symbol intervals and / or OFDM subcarrier intervals can vary for different transmissions, cells, and / or parts of the radio transmission spectrum. WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using subframes or transmit time intervals (TTIs) of varying or scalable lengths (e.g., containing a variety of OFDM symbols and / or lasting for a variable absolute time).
[0068] gNB180a, 180b, and 180c can be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c without accessing other RANs (e.g., e-nodes B160a, 160b, and 160c). In a standalone configuration, WTRU102a, 102b, and 102c can utilize one or more of gNB180a, 180b, and 180c as mobility anchor points. In a standalone configuration, WTRU102a, 102b, and 102c can communicate with gNB180a, 180b, and 180c using signals in unlicensed bands. In a non-standalone configuration, WTRU102a, 102b, and 102c can communicate with / connect to gNB180a, 180b, and 180c while also communicating with / connecting to other RANs such as enodes B160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c can implement the DC principle to communicate substantially simultaneously with one or more gNB180a, 180b, and 180c and one or more enodes B160a, 160b, and 160c. In a non-standalone configuration, enodes B160a, 160b, and 160c can act as mobility anchors for WTRU102a, 102b, and 102c, and gNB180a, 180b, and 180c can provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.
[0069] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interconnection between NR and E-UTRA, routing of user plane data to user plane functions (UPF) 184a and 184b, and routing of control plane information to access and mobility management functions (AMF) 182a and 182b. As shown in Figure 1D, the gNB180a, 180b, and 180c can communicate with each other via the Xn interface.
[0070] The CN115 shown in Figure 1D may include at least one AMF182a, 182b, at least one UPF184a, 184b, at least one Session Management Function (SMF)183a, 183b, and optionally a Data Network (DN)185a, 185b. While each of the above elements is shown as part of the CN115, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0071] AMF182a and 182b can be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N2 interface and can act as control nodes. For example, AMF182a and 182b may be responsible for authenticating users of WTRU102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, and mobility management. Network slicing may be used by AMF182a and 182b to customize the support for WTRU102a, 102b, and 102c CNs based on the type of service being utilized. For example, different network slices may be established for different use cases, such as services relying on high-reliability, low-latency (URLLC) access, services relying on extended large-scale mobile broadband (eMBB) access, and services using machine-type communications (MTC) access. AMF182a, 182b can provide control plane functionality for switching between RAN113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.
[0072] SMF183a and 183b can be connected to AMF182a and 182b in CN115 via the N11 interface. SMF183a and 183b can also be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b can select and control UPF184a and 184b and configure traffic routing through UPF184a and 184b. SMF183a and 183b can perform other functions such as managing and allocating IP addresses for UEs, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.
[0073] UPF184a, 184b may be connected to one or more of the gNB180a, 180b, 180c in RAN113 via an N3 interface that can give WTRU102a, 102b, 102c access to a packet-switched network such as the Internet 110 to facilitate communication between WTRU102a, 102b, 102c and IP-enabled devices. UPF184a, 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0074] CN115 can facilitate communication with other networks. For example, CN115 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN115 and PSTN108. Furthermore, CN115 can grant WTRU102a,102b,102c access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a,102b,102c may be connected to the Local Data Network (DN) 185a,185b through UPF184a,184b via an N3 interface to UPF184a,184b and an N6 interface between UPF184a,184b and DN185a,185b.
[0075] In view of Figures 1A to 1D and the corresponding descriptions of Figures 1A to 1D, one or more or all of the functions described herein with respect to one or more of the WTRU102a to d, base stations 114a to b, e-nodes B160a to c, MME162, SGW164, PGW166, gNB180a to c, AMF182a to b, UPF184a to b, SMF183a to b, DN185a to b, and / or any other devices described herein may be implemented by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.
[0076] Emulation devices may be designed to perform one or more tests on other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one, more, or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless network to test other devices in a communications network. One or more emulation devices may perform one, more, or all functions while being temporarily implemented / deployed as part of a wired and / or wireless network. Emulation devices may be directly coupled to another device so that they can be tested and / or performed using over-the-air radio communication.
[0077] One or more emulation devices can perform one or more functions, including all of them, without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test scenario in a test laboratory and / or an undeployed (e.g., for testing) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. To transmit and / or receive data, the emulation device may use direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).
[0078] This application describes various embodiments, including tools, features, examples, models, and methods. Many of these embodiments are described with specific characteristics and are often described in a manner that may sound limiting, at least to illustrate individual features. However, this is for the sake of clarity of description and does not limit the applications or scope of those embodiments. In fact, all different embodiments can be combined and interchangeable to give further embodiments. Furthermore, embodiments can also be combined and interchangeable with embodiments described in prior applications.
[0079] The embodiments described and intended in this application can be implemented in many different forms. Figures 5 and 6 described herein may give some examples, but other examples are intended. The description of Figures 5 and 6 is not limited to the scope of implementation. At least one of the embodiments relates generally to video encoding and decoding, and at least one other embodiment relates generally to transmitting a generated or encoded bitstream. These and other embodiments can be implemented as a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods, apparatus, or described methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.
[0080] In this application, the terms “reconstructed” and “decoded” may be used interchangeably; the terms “pixel” and “sample” may be used interchangeably; and the terms “image,” “picture,” and “frame” may be used interchangeably.
[0081] This specification describes various methods, each comprising one or more steps or actions for achieving the described method. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc., may be used in various examples to modify elements, components, steps, actions, etc., such as “first decryption” and “second decryption.” The use of such terms does not imply any ordering of modified actions unless specifically required. Thus, in this example, the first decryption does not need to be performed before the second decryption, but may, for example, be performed before, during, or over a period of overlapping time with the second decryption.
[0082] Various methods and other embodiments described herein may be used to modify modules of the video encoder 200 and decoder 300 shown in Figures 2 and 3, for example, the decoding module. Furthermore, the subject matter disclosed herein may be applied to any type, format, or version of video coding, whether described in standards or recommendations, whether existing or future development, or extensions of such standards and recommendations. Unless otherwise specified or technically hindered, the embodiments described herein may be used individually or in combination.
[0083] Various numerical values are used in the examples described in this application. These numerical values and other specific values are used to illustrate various examples, and the embodiments described are not limited to these specific numerical values.
[0084] Figure 2 shows an exemplary video encoder. While variations of the exemplary encoder 200 are intended, encoder 200 is described below for clarity without describing all expected variations.
[0085] Before encoding, the video sequence may undergo pre-encoding (201), for example, applying color conversions to the input color picture (e.g., converting from RGB4:4:4 to YCbCr4:2:0), or performing remapping of input picture components to make signal delivery more compression-elastic (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.
[0086] In encoder 200, the picture is encoded by encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of encoding units (CUs), for example. Each unit is encoded using either intra-mode or inter-mode, for example. When a unit is encoded in intra-mode, it performs intra-prediction (260). In inter-mode, motion estimation (275) and motion compensation (270) are performed. The encoder decides whether to use intra-mode or inter-mode to encode a unit (205), and indicates the intra / inter decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting the prediction block from the original image block (210).
[0087] The predicted residual is then transformed (225) and quantized (230). The quantization transformation coefficients, as well as the motion vector and other syntax elements, are entropy encoded (245) and output as a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is coded directly without the application of any transformation or quantization process.
[0088] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantization transformation coefficients are inversely quantized (240), inversely transformed (250), and the prediction residuals are decoded. Combining the decoded prediction residuals with the prediction blocks (255) restores the image blocks. An in-loop filter (265) is applied to the restored picture to reduce encoding artifacts, for example, by performing deblocking / SAO (Sample Adaptive Offset) filtering. The filtered image is stored in a reference picture buffer (280).
[0089] Figure 3 shows an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by the decoder elements as described below. The video decoder 300 generally performs a decoding path that is the reverse of the encoding path described in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.
[0090] In particular, the decoder input includes a video bitstream that may be generated by the video encoder 200. The bitstream is first entropy-decoded (330) to obtain transformation coefficients, motion vectors, and other coded information. Picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (335). The transformation coefficients are inversely quantized (340) and inversely transformed (350) to decode the prediction residuals. Combining the decoded prediction residuals with the prediction blocks (355) restores the image blocks. The prediction blocks may be obtained from intra-predictions (360) or motion-compensated predictions (i.e., inter-predictions) (375) (370). An in-loop filter (365) is applied to the restored image. The filtered image is stored in a reference picture buffer (380).
[0091] The decoded picture may further undergo post-decoding (385), such as reverse color conversion (e.g., conversion from YCbCr4:2:0 to RGB4:4:4) or reverse remapping, which is the reverse of the remapping process performed in pre-encoding (201). The post-decoding process may use metadata derived in pre-encoding and signaled in the bitstream. For example, the decoded image (e.g., after the application of the in-loop filter (365) and / or after post-decoding (385) if post-decoding is used) may be sent to a display device for rendering to the user.
[0092] Figure 4 shows an example of a system in which various embodiments and examples described herein may be implemented. System 400 may be implemented as a device comprising various components described below and configured to implement one or more of the embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital TV receivers, personal video recording systems, connected household electrical appliances, and servers. The elements of System 400 may be implemented individually or in combination as a single integrated circuit (IC), multiple ICs, and / or individual components. For example, in at least one example, the processing and encoder / decoder elements of System 400 are distributed across multiple ICs and / or individual components. In various examples, System 400 is communicably coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various examples, System 400 is configured to implement one or more of the embodiments described herein.
[0093] System 400 includes, for example, at least one processor 410 configured to execute instructions loaded therein to implement various embodiments described herein. The processor 410 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 400 includes at least one memory 420 (e.g., volatile memory devices and / or non-volatile memory devices). System 400 includes a storage device 440 which may include, but is not limited to, electrically erasable programmable read-only memory (EEPROM), ROM, programmable read-only memory (PROM), RAM, dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives, as well as non-volatile and / or volatile memory. The storage device 440 may, in non-limiting examples, include internal storage devices, accessory storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0094] System 400 includes, for example, an encoder / decoder module 430 configured to process data to give encoded or decoded video, the encoder / decoder module 430 of which may include its own processor and memory. The encoder / decoder module 430 represents a module that may be included in the device to perform encoding and / or decoding functions. As is known, the device may include one or both of the encoding and decoding modules. Furthermore, the encoder / decoder module 430 may be implemented as a separate element of System 400 or may be incorporated into the processor 410 as a combination of hardware and software, as is known to those skilled in the art.
[0095] Program code to be loaded onto a processor 410 or encoder / decoder 430 that implements the various embodiments described herein may be stored in a storage device 440 and subsequently loaded into memory 420 for execution by the processor 410. According to various examples, one or more of the processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more of various items during the implementation of the processing described herein. Such stored items include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of expressions, formulas, operations, and operational logic.
[0096] In some examples, the internal memory of the processor 410 and / or the encoder / decoder module 430 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other examples, memory outside the processing device (for example, the processing device may be either the processor 410 or the encoder / decoder module 430) is used for one or more of these functions. The external memory may be memory 420 and / or storage device 440, for example, dynamic volatile memory and / or non-volatile flash memory. In some examples, external non-volatile flash memory is used, for example, to store the television operating system. In at least one example, light-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations.
[0097] Inputs to the elements of System 400 can be provided through various input devices, as shown in Block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion for receiving RF signals transmitted wirelessly by a broadcaster, for example, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI®) input terminal. Other examples not shown in Figure 4 include composite video.
[0098] In various examples, the input device of block 445 has the respective relevant input processing elements known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band of frequencies to select a signal frequency band (e.g., sometimes called a channel in some examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) multiplexing to select a desired stream of data packets. The RF portion of various examples includes one or more elements for performing these functions, e.g., frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner for performing various functions of these functions, e.g., down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near the baseband) or to the baseband. In one example of a set-top box, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium, filter it, down-convert it, and filter it again to a desired frequency band to perform frequency selection. Various examples may rearrange the order of the elements described above (and others), remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements in between existing elements, such as inserting amplifiers and analog-to-digital converters. In various examples, the RF section may include an antenna.
[0099] USB and / or HDMI terminals may include their respective interface processors for connecting the system 400 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, if necessary, in a separate input processing IC or within the processor 410. Similarly, aspects of USB or HDMI interface processing may be implemented, if necessary, in a separate interface IC or within the processor 410. The demodulated, error-corrected, and multiplexed streams are fed to various processing elements, including, for example, the processor 410 and encoder / decoder 430, which work in conjunction with memory and storage elements to process the data streams as necessary for presentation on an output device.
[0100] Various elements of system 400 may be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected with each other using a suitable connection configuration 425, for example, an internal bus known in the art, including an inter-IC (I2C) bus, wiring, and a printed circuit board, and can transmit data between them.
[0101] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 460. The communication interface 450 may also include, but is not limited to, a modem or network card and the communication channel 460, which may be implemented in a wired and / or wireless medium, for example.
[0102] In various examples, data is streamed to or provided to system 400 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. In these examples, communication channel 460 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other examples, data is provided to system 400 using a set-top box that distributes data via the HDMI connection of input block 445. In yet another example, data is provided to system 400 using the RF connection of input block 445. As described above, various examples provide data in a non-streaming manner. Furthermore, various examples use wireless networks other than Wi-Fi, e.g., cellular networks or Bluetooth® networks.
[0103] System 400 can provide output signals to various output devices, including a display 475, a speaker 485, and other peripheral devices 495. Various examples of the display 475 include, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 475 may be for a television, tablet, laptop, cell phone (mobile phone), or other device. The display 475 may also be integrated into other components (for example, in the case of a smartphone) or separate (for example, an external monitor for a laptop). Various examples of the other peripheral devices 495 include, in various examples, one or more of a standalone digital video disc (or digital multipurpose disc) (DVD for both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 495 that provide functionality based on the output of System 400. For example, a disc player performs the function of playing the output of System 400.
[0104] In various examples, control signals are communicated between the system 400 and the display 475, speaker 485, or other peripheral devices 495 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices may be coupled to the system 400 via dedicated connections through their respective interfaces 470, 480, and 490. Alternatively, output devices may be connected to the system 400 using communication channels 460 via communication interface 450. The display 475 and speaker 485 may be integrated into a single unit with other components of the system 400 in an electronic device such as a television. In various examples, the display interface 470 may include a display driver, such as a timing controller (TCon) chip.
[0105] The display 475 and speaker 485 may, alternatively, be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various examples where the display 475 and speaker 485 are external components, the output signal may be provided via a dedicated output connection, such as an HDMI port, a USB port, or a COMP output.
[0106] An example may be executed by computer software implemented by the processor 410 or by hardware, or by a combination of hardware and software. In a non-limiting example, an example may be implemented by one or more integrated circuits. The memory 420 may be of any type appropriate for the technical environment, and in a non-limiting example, may be implemented using any suitable data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 410 may be of any type appropriate for the technical environment, and in a non-limiting example, may include one or more of microprocessors, general-purpose computers, dedicated computers, and processors based on multi-core architectures.
[0107] Various implementations are involved in decoding. As used in this application, “decoding” may include all or part of processing performed on a received encoded sequence, for example, to produce a final output suitable for display. In various examples, such processing may include one or more of the processing commonly performed by decoders, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such processing may also, or alternatively, include processing performed by the various implementations of decoders described in this application, such as receiving, decoding, and rendering a representation of volumetric video encoded in MPEG immersive video (MIV).
[0108] For further examples, in one instance, "decoding" refers only to entropy decoding; in another instance, "decoding" refers only to differential decoding; and in yet another instance, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" refers specifically to a subset of operations or, in general, to a broader decoding process will become clear from the context of the particular description and will be well understood by those skilled in the art.
[0109] Various implementations are involved in encoding. In a manner similar to the above description of "decoding," "encoding" as used in this application may include all or part of the processing performed on, for example, an input video sequence to generate an encoded bitstream. In various examples, such processing includes one or more of the processing commonly performed by an encoder, such as partitioning, differential encoding, transformation, quantization, and entropy coding. In various examples, such processing also includes, or alternatively, processing performed by the encoder of the various embodiments described in this application, such as generating volumetric video encoded in MPEG immersive video (MIV), and determining a set of MPI parameters associated with the MIV-encoded volumetric video and a set of MIV parameters associated with the MIV-encoded volumetric video.
[0110] As further examples, in one example, “encoding” refers only to entropy coding; in another example, “encoding” refers only to differential coding; and in yet another example, “encoding” refers to a combination of differential coding and entropy coding. Whether the phrase “encoding process” refers specifically to a subset of operations or, in general, to a broader encoding process will become clear from the context of the particular description and will be well understood by those skilled in the art.
[0111] It should be noted that the MIV coding syntax elements used herein, such as MaxNumLayers, mvp_num_views_minus1, ave_view_complete_in_atlas_flag[atlasID][0], VpsPackingInformationPresentFlag, and vps_packed_video_present_flag[atlasID], are descriptive terms. Therefore, they do not preclude the use of other syntax element names.
[0112] Please understand that when a diagram is presented as a flowchart, it also provides a block diagram of the corresponding device. Similarly, please understand that when a diagram is presented as a block diagram, it also provides a flowchart of the corresponding method / process.
[0113] For example, the implementations and embodiments described herein may be implemented as methods or processes, apparatus, software programs, data streams, or signals. Even when only the content of a single form of implementation is discussed (e.g., only as a method), the implementation of the feature discussed may also be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, with appropriate hardware, software, and firmware. Methods may be implemented, for example, as processors, which generally refer to processing devices including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.
[0114] References to “one example,” “one implementation,” or “one implementation,” as well as other variations thereof, mean that the specific features, structures, characteristics, etc. described in relation to the example are included in at least one example. Therefore, references to “one example,” “one implementation,” “one implementation,” and any other variations appearing in various places throughout this application do not necessarily all refer to the same example.
[0115] Furthermore, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory. Acquiring may include receiving, retrieving, constructing, generating, and / or determining.
[0116] Furthermore, this application may refer to “accessing” various pieces of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.
[0117] Furthermore, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, as is "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory). Moreover, "receiving" generally involves, in some way, actions such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0118] For example, please understand that in the cases of "A / B", "A and / or B", and "at least one of A and B", any use of " / ", "and / or", and "at least one of" below includes the selection of only the first enumerated option (A), or only the second enumerated option (B), or both options (A and B). As further examples, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording includes the selection of only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or only the second and third enumerated options (B and C), or all three options (A, B, and C). This can be extended to the same number of items as listed, as will be obvious to those skilled in the art.
[0119] Furthermore, as used herein, the term "signal" specifically means indicating something to the corresponding decoder. Encoder signals may include, for example, a low-complexity configuration. In this way, in one example, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can send specific parameters to the decoder (explicit signaling), and the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters and others, signaling can be used without transmission (implicit signaling), simply allowing the decoder to know and select the specific parameters. Bit savings are achieved in various examples by avoiding the transmission of any actual function. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc., are used in various examples to signal information to the corresponding decoder. The above relates to the verb form of the word "signal," but the word "signal" may also be used as a noun in this specification.
[0120] As will be obvious to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be stored or transmitted. This information may include, for example, instructions for carrying out the method or data generated by one of the described implementations. For example, a signal may be formatted to carry the bitstream of the described example. Such a signal may be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a wide variety of wired or wireless links, as is known. The signal may be stored on, accessed from, or received from a processor-readable medium.
[0121] This specification describes many examples. The features of the examples may be given individually or in any combination across various claim categories and types. Furthermore, the examples may include one or more of the features, devices, or embodiments described herein individually or in any combination across various claim categories and types. For example, the features described herein may be implemented in a bitstream or signal containing information generated as described herein. The information may enable a decoder to decode the bitstream, and encoders, bitstreams, and / or decoders according to any embodiment are described. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented in a method, process, apparatus, medium for storing instructions, medium for storing data, or signal. For example, the features described herein may be implemented by a TV, set-top box, cellphone, tablet, or other electronic device that performs decoding. The TV, set-top box, cellphone, tablet, or other electronic device can display the resulting image (e.g., an image from residual reconstruction of a video bitstream) (e.g., using a monitor, screen, or other type of display). TVs, set-top boxes, cell phones, tablets, or other electronic devices can receive and decode signals containing encoded images.
[0122] Systems, methods, and means for defining and using low-complexity multiplane image (MPI) profiles (for example, for the MPEG immersive video (MIV) standard) are disclosed.
[0123] In the example, the device can generate MIV-encoded volumetric video. MIV-encoded volumetric video can be encoded using the MPI format. The device can determine a set of MPI parameters and / or a set of MIV parameters associated with the MIV-encoded volumetric video.
[0124] A device can determine that a set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration and can be transmitted to a device with low computing power. The device can transmit a representation of MIV-encoded volumetric video, an indication of the set of MPI parameters, and / or a set of MIV parameters (e.g., in a bitstream). MPI parameters may include the number of MPI layers, an indication of the MPI layer arrangement, and / or the number of reference cameras. MIV parameters may include the total number of layers, an indication of frame packing, and / or multiple level limits.
[0125] In the example, the device may send an indication (e.g., a tag) that a set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration.
[0126] An indication of a set of MPI parameters and / or a set of MIV parameters can indicate the level of the MIV profile.
[0127] In the example, the device can use SEI (supplemental enhancement information) messages to send indications of a set of MPI parameters and / or a set of MIV parameters. In the example, the set of MPI parameters and / or MIV parameters can indicate that the set of parameters is associated with a low-complexity configuration.
[0128] In the example, the decoder may receive (for example, in a bitstream) a representation of a MIV-encoded volumetric video encoded using the MPI format, an indication of a set of MPI parameters, and / or a set of MIV parameters associated with the MIV-encoded volumetric video. In the example, the decoder may have limited computing power to decode the MIV-encoded MPI content.
[0129] In the example, the decoder can determine that MPI parameters and / or MIV parameters are associated with a low-complexity configuration. In the example, the decoder can receive an indication that a set of MPI parameters and / or a set of MIV parameters are associated with a low-complexity configuration.
[0130] The decoder can decode and / or render the representation. The representation can be decoded and / or rendered using a set of MPI parameters and / or a set of MIV parameters associated with a low-complexity configuration.
[0131] Volumetric video can be considered for virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) applications. Volumetric video can include sequences of 3D frames (e.g., in 3D formats such as point clouds, meshes, and / or multiview plus depth video).
[0132] 3D volumetric information may be converted to a set of 2D images and / or associated data (e.g., to encode volumetric frames). The converted 2D images can be encoded using a 2D video encoder. Associated data can be encoded into a metadata stream (e.g., an additional metadata stream). The encoded images and / or associated metadata can be decoded. The encoded images and / or associated metadata can be used to reconstruct the 3D volumetric information. Video standards incorporating a 2D video compatibility approach (e.g., as described herein) can incorporate multiview plus depth inputs.
[0133] Figure 5 shows an example of a volumetric scene containing a set of MPI slices (e.g., a perspective projection). MPI can be a hierarchical representation of a volumetric scene. A volumetric scene (e.g., in a hierarchical representation) can be represented as a set of slices sampled at different depths (e.g., from a given reference viewpoint). As shown in Figure 5, for example, if a camera projection model is used as a perspective projection, the slices may be front-parallel planes.
[0134] Figure 6 shows an example of a multi-sphere image (MSI) slice (e.g., equirectangular projection). As shown in Figure 6, for example, if the camera projection model is an equirectangular projection model, the slice may be concentric spheres.
[0135] MPI can be defined by a set of slices / layers. A layer (e.g., each layer) may be a frame obtained by projecting a portion of the 3D scene contained within the layer onto a reference camera. The reference camera may be positioned at a given reference viewpoint. Each layer (e.g., each layer) may contain an indication of transparency information (e.g., a scalar value per pixel ranging from 0 to 1). The transparency information can represent the level of transparency of each pixel in the layer frame.
[0136] Immersive video standards (e.g., MIV) may support MPI input formats through extended profiles (e.g., the MIV Extended-Restricted Geometry profile). MPI representation of 3D scenes can enable efficient rendering based on alpha blending (e.g., known as the reverse painter algorithm). The reverse painter algorithm may be suitable for end-user devices with low computing power (e.g., WTRU).
[0137] Immersive video formats can be flexible with respect to the parameters of MPI input content (e.g., about those parameters). There may be no limit to the number of depth layers. For example, the number of depth layers could be between 100 and 300 to accurately sample the depth of a complex 3D scene and enable high-quality rendering of the virtual view. The number of depth layers may be fewer for simple 3D scenes with small depth amplitudes. Layers may be sampled at equal distances from each other (e.g., regularly). Layers may be placed at arbitrary positions (e.g., regularly). One or more MPIs (e.g., multiple) may be transmitted in an immersive video bitstream (e.g., the same MIV bitstream). Each MPI (e.g., each MPI) may have its own reference camera position (e.g., collectively covering most of the 3D scene).
[0138] Immersive video standards (e.g., MIV) may transmit MPI input content as multiple (e.g., two) video bitstreams. MPI input content may include a patch atlas of textures and / or a patch atlas of transparency information. Patch representations may allow for a high pixel rate reduction (e.g., by pruning sparse or empty portions of the MPI layer). Multiple (e.g., two) video bitstreams may be packed into a single video bitstream that allows for a single decoder instantiation (e.g., for user devices with limited or reduced processing power).
[0139] Table 1 shows an example of an MPI profile (e.g., MIV) for an immersive video standard. Exemplary features of an MPI profile are shown in Table 1. In the example, the features of an MPI profile may include one or more of the following: allowed attribute types (e.g., texture and / or transparency), patches signaled with a constant depth value, or the absence of a geometric (e.g., depth map) video component.
[0140] [Table 1-1]
[0141] [Table 1-2]
[0142] One or more (e.g., several) end-user devices (e.g., WTRUs with too low computational complexity) may have difficulty or be unable to decode and / or render in real time a detailed MPI scene representation encoded in an immersive video standard (e.g., MIV). A set of constraints may be specified to limit the complexity of decoding and / or rendering an MPI representation encoded in an immersive video standard (e.g., MIV). The decoder may be notified (e.g., via the bitstream) whether the bitstream conforms to a set of constraints (e.g., low complexity constraints).
[0143] Volumetric video encoded in the MPI format immersive video standard (e.g., MIV) can be signaled to MIV-compliant end-user devices (e.g., MIV-compliant WTRUs). Volumetric MPI format encoded in MIV may satisfy constraints that limit the complexity of decoding and / or rendering.
[0144] Constraints (e.g., limits) can be provided on MPI parameters (e.g., number of layers, layer positions, reference camera, etc.) and / or parameters encoded in immersive video standards (e.g., patch pruning, frame packing, level limits, etc.). Constraints may be specified (e.g., as a low-complexity MPI profile) and / or transmitted via SEI messages.
[0145] An end-user device compliant with an immersive video standard (e.g., an MIV-compliant WTRU) may signal that volumetric video in the MPI format encoded with the immersive video standard satisfies constraints (e.g., one or more constraints that may limit the complexity of decoding and / or rendering). These constraints may be referred to as low-complexity MPI configurations.
[0146] Features related to constraints on MPI representation are provided herein.
[0147] A set of constraints can be determined for 3D scene modeling using MPI. Constraints for a low-complexity MPI configuration may include constraints on one or more parameters such as the number of MPI layers, the arrangement of layers, and / or the number of reference cameras.
[0148] Constraints on the number of MPI layers may apply to the MPI representation of immersive video. Patches (e.g., patches obtained from a patch atlas representation) can have a certain depth value. Patches can be signaled with patch metadata. During the decoding stage, the reconstructed MPI may have the same number of layers as the different patch depth values in the patch atlas (e.g., one layer for each patch depth value).
[0149] High-quality depth sampling of a 3D scene with complex depth content may involve 100-300 depth layer values (e.g., to ensure high-quality rendering of virtual views for a user-immersive experience). A limited immersive experience in a simple 3D scene can be provided with fewer depth layers (e.g., 16 or 32 depth layers). If the renderer receives the maximum number of layers in advance (e.g., signaled) without having to wait for the complete decoding of the bitstream, the renderer can allocate sufficient memory for MPI reconstruction before virtual view compositing. Low-complexity configurations can specify the maximum number of depth layers (e.g., 16 or 32).
[0150] Constraints on layer placement may apply to MPI representations of immersive video. Layer placement may be restricted to fixed intervals, so that the position of each layer can be inferred from the number of layers, as well as the minimum and / or maximum depth values of the scene (e.g., as signaled by the depth quantization parameters of an immersive video standard (e.g., MIV)). In an example, a low-complexity configuration may be specified so that layers are placed at regular or even intervals (e.g., between the minimum and maximum scene depth values).
[0151] Constraints on the number of reference cameras may apply to the MPI representation of immersive video. A basic MPI may be associated with a single reference camera (e.g., a camera located at the center of the viewing space in which the user virtually resides). A basic MPI may provide a virtual viewpoint offset from the center position. Patches of an immersive video standard (e.g., MIV) representation may be associated with a (e.g., unique) camera. In the example, the patch may be associated with the camera's intrinsic and / or extrinsic parameters. For example, if there is no restriction on the number of reference cameras, one or more (e.g., multiple) MPIs may be transmitted in an immersive video standard bitstream (e.g., the same MIV bitstream). The complexity of rendering a virtual viewpoint from one or more (e.g., multiple) MPI samples of a 3D scene can increase (e.g., significantly). In the example, a low-complexity configuration may specify that there is a (e.g., unique) reference camera (e.g., a unique immersive video standard view).
[0152] Features related to constraints on immersive video standard (e.g., MIV) coding are provided herein.
[0153] The representation of immersive video standards and simplification of video encoding can be achieved through patch atlases. Constraints on immersive video standard (e.g., MIV) encoding may include constraints on one or more of the complete layer, frame packing, and level limits.
[0154] The constraint on complete layers may apply to the encoding of immersive video standards. Immersive video standards can provide an efficient patch-based representation where empty (e.g., completely transparent) patches from the input MPI are pruned (e.g., filtered, extracted, omitted, and / or similar). The remaining patches of the efficient patch-based representation (e.g., patches with information important to the renderer) can be organized into a patch atlas. Organizing the patch atlas may allow for a reduction in the pixel rate. At the decoder stage, the pruned MPI can be reconstructed by removing patches from the atlas and / or by placing the patches in their original positions (e.g., their depth layers) in a reference view. The original positions (e.g., for each patch) may be signaled in the patch metadata. If pruning is not performed in the encoder and / or different layers are transmitted and the video is encoded as complete layers, MPI reconstruction can be simplified and / or accelerated (e.g., at the expense of the pixel rate). Simplification may be feasible when the number of layers is limited. A patch atlas can have as many patches (e.g., patches with a certain depth) as there are depth layers. In the example, there may be (e.g., one) patch for each depth layer. In the example, a low-complexity configuration can (e.g., provide) patches with a certain depth corresponding to complete (e.g., unpruned) layers.
[0155] Constraints on frame packing may apply to immersive video standard (e.g., MIV) encoding. The frame packing features of immersive video standards can allow various video components to be packed into a single video bitstream (e.g., a single decoder instantiation). A single video bitstream may be suitable for user devices (e.g., WTRU) with limited or reduced processing power (e.g., within that scope). For example, a low-complexity configuration might specify that texture and / or transparency components be packed into a single video bitstream.
[0156] Level constraints can be applied to the encoding of immersive video standards. For example, a low-complexity configuration can specify level constraints on one or more of the following: a single atlas, pixel rate, or bitrate.
[0157] Low-complexity configurations can limit the use of atlases (e.g., a single atlas). Immersive video standards (e.g., no restrictions on the number of patch atlases) may allow for splitting patch-based representations into multiple atlases (e.g., resulting in many video bitstreams). Low-complexity configurations can still be used with atlases (e.g., a single atlas).
[0158] Low-complexity configurations can limit the pixel rate. The maximum number of samples per frame (e.g., atlas size) and / or the maximum number of samples per second may be specified (e.g., in low-complexity configurations).
[0159] Low-complexity configurations allow you to specify the bitrate (for example, the maximum video bitrate may be fixed).
[0160] Features related to profiling immersive video standards (e.g., MIV) for low-complexity MPI are provided herein.
[0161] Low complexity constraints (e.g., constraints as described herein) can be implemented using dedicated immersive video standard (e.g., MIV) profiles. These dedicated immersive video standards (e.g., MIV) may also be limitations of existing profiles. The features of a dedicated immersive video standard (e.g., MIV) profile may include one or more of the following (e.g., using exemplary MIV syntax elements): the number of MPI depth layers does not exceed the maximum number of layers (e.g., MaxNumLayers) (e.g., equal to 16 or 32); depth layers are equally spaced from one another; there is a (e.g., single) reference camera (e.g., mvp_num_views_minus1=0); patches correspond to complete layers (e.g., ave_view_complete_in_atlas_flag[atlasID][0]=1); or texture and / or transparency video components are packed into (e.g., a single) video frame (e.g., VpsPackingInformationPresentFlag=1 and / or vps_packed_video_present_flag[atlasID]=1).
[0162] A dedicated immersive video standard profile may include one or more of the following constraints associated with at least one level: maximum atlas frame size, pixel rate (e.g., samples per second), or bitrate value. In the example, the levels of the immersive video standard profile (e.g., each level) may be associated with a specified maximum atlas frame size, pixel rate, and / or bitrate.
[0163] In the example, a low-complexity configuration may be signaled using SEI messages. In the example, one or more constraints signaled by SEI messages (e.g., MPI representation constraints and / or MVI coding constraints) may indicate a profile (e.g., a low-complexity MPI profile).
[0164] User devices with low computing power can decode and / or render immersive content encoded in the MPI Immersive Video Standard, which is tagged as low complexity and can provide an immersive experience.
[0165] While features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. Furthermore, the methods described herein may be executed by computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, magnetic media such as ROM, RAM, registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multipurpose disks (DVDs). Processors working with software may be used to implement radio frequency transceivers for use in WTRUs, UEs, terminals, base stations, RNCs, or any host computer.
Claims
1. A video decoding device equipped with a processor, The aforementioned processor, The means of receiving a representation of volumetric video encoded in MPEG Immersive Video (MIV), which is encoded using the Multiplane Image (MPI) format, wherein the MIV-encoded volumetric video representation is received via a single video bitstream. Receiving indications of a set of MPI parameters and a set of MIV parameters associated with a volumetric video encoded in MIV, The determination that the set of MPI parameters and the set of MIV parameters are associated with a low-complexity configuration, wherein the low-complexity configuration limits the complexity of rendering the volumetric video encoded in the MIV. Decoding the volumetric video representation encoded in MIV based on the set of MPI parameters and the set of MIV parameters, A video decoding device configured to perform the following.
2. The aforementioned processor, Receiving at least one constraint on the set of MPI parameters and the set of MIV parameters, Determining that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration based on the at least one constraint, The video decoding device according to claim 1, configured to further perform the following:
3. The aforementioned processor, Receiving an indication of an MIV profile, wherein the MIV profile is associated with the low-complexity configuration, Based on the indication of the MIV profile, determine at least one parameter from the set of MPI parameters and the set of MIV parameters. A video decoding device according to claim 1 or 2, configured to further perform the following:
4. The video decoding apparatus according to any one of claims 1 to 3, wherein the indication of the set of MPI parameters and the set of MIV parameters is received via a SEI (supplemental enhancement information) message.
5. The video decoding apparatus according to any one of claims 1 to 4, wherein the set of MPI parameters associated with the volumetric video encoded in MIV includes at least one of the number of MPI layers, the arrangement of the layers, and the number of reference cameras.
6. The set of MPI parameters includes the number of MPI layers, the number of MPI layers is associated with the maximum number of layers based on the low-complexity configuration, and the processor is The video decoding apparatus according to any one of claims 1 to 4, further configured to perform determining, based on the number of MPI layers, that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration.
7. The video decoding apparatus according to claim 5 or 6, wherein the number of MPI layers is 16 or 32.
8. The video decoder according to any one of claims 1 to 7, wherein the set of MIV parameters associated with the volumetric video encoded in MIV includes at least one of the number of complete layers, frame packing indication, and level limit.
9. The video decoder according to any one of claims 1 to 7, wherein the set of MIV parameters associated with the volumetric video encoded in MIV includes a level limit, the level limit is associated with a single atlas, pixel rate, or bitrate.
10. The aforementioned processor, Receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches is associated with texture information or transparency information of at least one layer of the volumetric video encoded in MIV. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is decoded, A video decoding device according to any one of claims 1 to 9, configured to further perform the following:
11. The aforementioned processor, The process involves receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches represents a complete layer. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is decoded, A video decoding device according to any one of claims 1 to 9, configured to further perform the following:
12. A video encoding device equipped with a processor, The aforementioned processor, The means of receiving a representation of volumetric video encoded in MPEG Immersive Video (MIV), which is encoded using the Multiplane Image (MPI) format, wherein the MIV-encoded volumetric video representation is received via a single video bitstream. Receiving indications of a set of MPI parameters and a set of MIV parameters associated with a volumetric video encoded in MIV, The determination that the set of MPI parameters and the set of MIV parameters are associated with a low-complexity configuration, wherein the low-complexity configuration limits the complexity of rendering the volumetric video encoded in the MIV. Encoding a representation of a volumetric video encoded in MIV based on the set of MPI parameters and the set of MIV parameters, A video encoding device configured to perform the following.
13. The aforementioned processor, Receiving at least one constraint on the set of MPI parameters and the set of MIV parameters, Determining that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration based on the at least one constraint, The video encoding apparatus according to claim 12, configured to further perform the following:
14. The aforementioned processor, Receiving an indication of an MIV profile, wherein the MIV profile is associated with the low-complexity configuration, Based on the indication of the MIV profile, determine at least one parameter from the set of MPI parameters and the set of MIV parameters. A video encoding apparatus according to claim 12 or 13, configured to further perform the following:
15. The video encoding apparatus according to any one of claims 12 to 14, wherein the indication of the set of MPI parameters and the set of MIV parameters is received via a SEI (supplemental enhancement information) message.
16. The video encoding apparatus according to any one of claims 12 to 15, wherein the set of MPI parameters associated with the volumetric video encoded in MIV is associated with at least one of the number of MPI layers, the arrangement of the layers, and the number of reference cameras.
17. The set of MPI parameters includes the number of MPI layers, the number of MPI layers is associated with the maximum number of layers based on the low-complexity configuration, and the processor is The video coding apparatus according to any one of claims 12 to 15, further configured to perform determining, based on the number of MPI layers, that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration.
18. The video encoding apparatus according to claim 16 or claim 17, wherein the number of MPI layers is 16 or 32.
19. The video encoding apparatus according to any one of claims 12 to 18, wherein the set of MIV parameters associated with the volumetric video encoded in MIV includes at least one of the number of complete layers, frame packing indication, and level limit.
20. The video encoding apparatus according to any one of claims 12 to 18, wherein the set of MIV parameters includes a level limit, the level limit is associated with a single atlas, pixel rate, or bitrate.
21. The aforementioned processor, Receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches is associated with texture information or transparency information of at least one layer of the volumetric video encoded in MIV. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is encoded, A video encoding apparatus according to any one of claims 12 to 20, configured to further perform the following:
22. The aforementioned processor, The process involves receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches represents a complete layer. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is encoded, A video encoding apparatus according to any one of claims 12 to 20, configured to further perform the following:
23. A video decoding method, The means of receiving a representation of volumetric video encoded in MPEG Immersive Video (MIV), which is encoded using the Multiplane Image (MPI) format, wherein the MIV-encoded volumetric video representation is received via a single video bitstream. Receiving indications of a set of MPI parameters and a set of MIV parameters associated with a volumetric video encoded in MIV, The determination that the set of MPI parameters and the set of MIV parameters are associated with a low-complexity configuration, wherein the low-complexity configuration limits the complexity of rendering the volumetric video encoded in the MIV. Decoding the volumetric video representation encoded in MIV based on the set of MPI parameters and the set of MIV parameters, Video decoding methods, including...
24. Receiving at least one constraint on the set of MPI parameters and the set of MIV parameters, Determining that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration based on the at least one constraint, The video decoding method according to claim 23, further comprising:
25. Receiving an indication of an MIV profile, wherein the MIV profile is associated with the low-complexity configuration, Based on the indication of the MIV profile, determine at least one parameter from the set of MPI parameters and the set of MIV parameters. The video decoding method according to claim 23 or claim 24, further comprising:
26. The video decoding method according to any one of claims 23 to 25, wherein the indication of the set of MPI parameters and the set of MIV parameters is received via a SEI (supplemental enhancement information) message.
27. The video decoding method according to any one of claims 23 to 26, wherein the set of MPI parameters associated with the volumetric video encoded in MIV includes at least one of the number of MPI layers, the arrangement of the layers, and the number of reference cameras.
28. The set of MPI parameters includes the number of MPI layers, the number of MPI layers is associated with the maximum number of layers based on the low complexity configuration, and the method is The video decoding method according to any one of claims 23 to 26, further comprising determining, based on the number of MPI layers, that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration.
29. The video decoding method according to claim 27 or claim 28, wherein the number of MPI layers is 16 or 32.
30. The video decoding method according to any one of claims 23 to 29, wherein the set of MIV parameters associated with the volumetric video encoded in MIV includes at least one of the number of complete layers, frame packing indication, and level limit.
31. The video decoding method according to any one of claims 23 to 29, wherein the set of MIV parameters associated with the volumetric video encoded in the MIV includes a level limit, the level limit is associated with a single atlas, pixel rate, or bitrate.
32. Receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches is associated with texture information or transparency information of at least one layer of the volumetric video encoded in MIV. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is decoded, A video decoding method according to any one of claims 23 to 31, further comprising:
33. The process involves receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches represents a complete layer. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is decoded, A video decoding method according to any one of claims 23 to 31, further comprising:
34. A video encoding method, The means of receiving a representation of volumetric video encoded in MPEG Immersive Video (MIV), which is encoded using the Multiplane Image (MPI) format, wherein the MIV-encoded volumetric video representation is received via a single video bitstream. Receiving indications of a set of MPI parameters and a set of MIV parameters associated with a volumetric video encoded in MIV, The determination that the set of MPI parameters and the set of MIV parameters are associated with a low-complexity configuration, wherein the low-complexity configuration limits the complexity of rendering the volumetric video encoded in the MIV. Encoding a representation of a volumetric video encoded in MIV based on the set of MPI parameters and the set of MIV parameters, Video encoding methods, including...
35. Receiving at least one constraint on the set of MPI parameters and the set of MIV parameters, Determining that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration based on the at least one constraint, The video encoding method according to claim 34, further comprising:
36. Receiving an indication of an MIV profile, wherein the MIV profile is associated with the low-complexity configuration, Based on the indication of the MIV profile, determine at least one parameter from the set of MPI parameters and the set of MIV parameters. The video encoding method according to claim 34 or claim 35, further comprising:
37. The video encoding method according to any one of claims 34 to 36, wherein the indication of the set of MPI parameters and the set of MIV parameters is received via a SEI (supplemental enhancement information) message.
38. The video encoding method according to any one of claims 34 to 37, wherein the set of MPI parameters associated with the volumetric video encoded in MIV is associated with at least one of the number of MPI layers, the arrangement of the layers, and the number of reference cameras.
39. The set of MPI parameters includes the number of MPI layers, the number of MPI layers is associated with the maximum number of layers based on the low complexity configuration, and the method is The video coding method according to any one of claims 34 to 37, further comprising determining, based on the number of MPI layers, that the set of MPI parameters and the set of MIV parameters are associated with the low complexity configuration.
40. The video encoding method according to claim 38 or 39, wherein the number of MPI layers is 16 or 32.
41. The video encoding method according to any one of claims 34 to 40, wherein the set of MIV parameters associated with the volumetric video encoded in MIV includes at least one of the number of complete layers, frame packing indication, and level limit.
42. The video encoding method according to any one of claims 34 to 40, wherein the set of MIV parameters includes a level limit, the level limit is associated with a single atlas, pixel rate, or bitrate.
43. Receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches is associated with texture information or transparency information of at least one layer of the volumetric video encoded in MIV. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is encoded, A video encoding method according to any one of claims 34 to 42, further comprising:
44. The process involves receiving a patch atlas using the volumetric video encoded in MIV, wherein the patch atlas includes one or more patches, and each of the one or more patches represents a complete layer. Furthermore, based on the aforementioned patch atlas, the representation of the volumetric video encoded in MIV is encoded, A video encoding method according to any one of claims 34 to 42, further comprising:
45. A computer program product comprising program code instructions stored in a non-temporary computer-readable medium and, when executed by at least one processor, causing to perform a method according to at least one of claims 23 to 33 or claims 34 to 44.
46. A computer program comprising program code instructions that, when executed by a processor, cause to perform a method according to at least one of claims 23 to 33 or claims 34 to 44.
47. Video data including information representing an encoded output generated according to one of the methods described in any one of claims 34 to 44.