Normalized streaming infrastructure for encoding point cloud attributes
By employing normalized flow subprocesses and feature enhancement techniques, the problem of low encoding efficiency in point cloud data is solved, enabling efficient point cloud data compression and transmission, and adapting to the application requirements of various data formats.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTERDIGITAL CE PATENT HOLDINGS SAS
- Filing Date
- 2024-07-23
- Publication Date
- 2026-04-24
Smart Images

Figure CN121925680A_ABST
Abstract
Description
[0001] Cross-reference to related applications This application claims the benefit of European Patent Application No. EP23306306 entitled “NORMALIZING FLOW SMALL ARCHITECTURES TO CODE POINTCLOUD ATTRIBUTES”, filed on 28 July 2023, which is incorporated herein by reference in its entirety.
[0002] Cross-references to other applications The following applications are incorporated herein by reference in their entirety: European patent application entitled “METHODS AND APPARATUSES FOR ENCODING AND DECODING A POINT CLOUD”, filed on September 6, 2022, with serial number EP22306317 (“'317 application”); and European patent application entitled “METHODS AND APPARATUSES FOR ENCODING AND DECODING A POINT CLOUD”, filed on March 13, 2023, with serial number EP23305352 (“'352 application”). Background Technology
[0003] The use of 3D applications is becoming increasingly popular every day. To leverage these applications, different data formats are being used. One such data format is a point cloud. A point cloud is an unordered collection of points with coordinates (x, y, z) corresponding to their locations in space and their attributes, such as color and normal vectors. Summary of the Invention
[0004] The embodiments described herein include methods used in video encoding and decoding (collectively, “encoding”).
[0005] Example methods according to some embodiments may include: obtaining information corresponding to a point cloud; performing feature enhancement on point cloud data corresponding to the information; generating data corresponding to a latent space based on the information corresponding to the point cloud, wherein generating data corresponding to the latent space may include: using the feature-enhanced point cloud data as input data for a first loop through a subprocess; executing the subprocess twice, wherein the subprocess may include: passing the input data through a voxel shuffling layer to generate a voxel shuffling layer output; performing convolution on the voxel shuffling layer output to generate convolutional output data; passing the convolutional output data through one or more coupling layers to generate a subprocess output; and using the subprocess output as input data for a second loop through the subprocess; and passing the second loop subprocess output through an attention layer to generate data corresponding to the latent space; and encoding the data corresponding to the latent space as a bitstream.
[0006] In some embodiments of the example method, the sub-procedure may be executed three times.
[0007] In some embodiments of the example method, the sub-process further includes performing an average back projection process on the output of the one or more coupling layers.
[0008] For some embodiments of the example method, performing the average back projection process may include performing at least one of the following: performing an averaging process, a copying process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0009] In some embodiments of the example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0010] In some embodiments of the example method, the subprocess is a normalized flow (NF) subprocess.
[0011] In some embodiments of the example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0012] Example apparatuses according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0013] A second example method according to some embodiments may include: obtaining information corresponding to a point cloud; generating data corresponding to a latent space based on the information corresponding to the point cloud, wherein generating the data corresponding to the latent space may include: using feature-enhanced point cloud data as input data for a first loop through a subprocess; executing the subprocess twice, wherein the subprocess may include: passing the input data through a voxel shuffling layer to generate a voxel shuffling layer output; performing convolution on the voxel shuffling layer output to generate convolutional output data; passing the convolutional output data through one or more coupling layers to generate a subprocess output; and using the subprocess output as input data for a second loop through the subprocess; and generating data corresponding to the latent space based on the second loop subprocess output; and encoding the data corresponding to the latent space into a bitstream.
[0014] In some embodiments of the second example method, the sub-process is executed three times.
[0015] In some embodiments of the second example method, the sub-process further includes performing an average back projection process on the output of the one or more coupling layers.
[0016] For some embodiments of the second example method, performing the average back projection process may include performing at least one of the following: performing an averaging process, a copying process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0017] In some embodiments of the second example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0018] In some embodiments of the second example method, the subprocess is a normalized flow (NF) subprocess.
[0019] In some embodiments of the second example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0020] A second example apparatus according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0021] A third example method according to some embodiments may include: obtaining an encoded bitstream; decoding the encoded bitstream to generate a latent space; and reconstructing a point cloud from the generated latent space, wherein reconstructing the point cloud from the generated latent space may include: passing the generated latent space through an attention layer to generate input data for a sub-process; executing the sub-process twice, wherein the sub-process may include: passing the input data through one or more coupling layers to generate coupling layer outputs; performing convolution on the coupling layer outputs to generate convolution output data; passing the convolution output data through a voxel shuffling layer to generate a sub-process output; and using the sub-process output as input data for a second loop through the sub-process; and performing feature enhancement on the output of the second loop sub-process to generate the reconstructed point cloud.
[0022] In some embodiments of the third example method, the sub-process may be executed three times.
[0023] In some embodiments of the third example method, the sub-process may further include performing a copy back projection process on the output of the voxel shuffling layer.
[0024] For some embodiments of the third example method, performing the copy back projection process may include at least one of: performing a copy process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0025] In some embodiments of the third example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0026] In some embodiments of the third example method, the sub-process is a normalized flow (NF) sub-process.
[0027] In some embodiments of the third example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0028] A third example apparatus according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0029] A fourth example method / apparatus according to some embodiments may include: obtaining an encoded bitstream; decoding the encoded bitstream to generate a latent space; and reconstructing a point cloud from the generated latent space, wherein reconstructing the point cloud from the generated latent space may include: obtaining input data for a sub-process based on the generated latent space; executing the sub-process twice, wherein the sub-process may include: passing the input data through one or more coupling layers to generate coupling layer outputs; performing convolution on the coupling layer outputs to generate convolution output data; passing the convolution output data through a voxel shuffling layer to generate a sub-process output; and using the sub-process output as input data for a second loop through the sub-process; and generating the reconstructed point cloud based on the second loop sub-process output.
[0030] In some embodiments of the fourth example method, the sub-process may be executed three times.
[0031] In some embodiments of the fourth example method, the sub-process may further include performing a copy back projection process on the output of the voxel shuffling layer.
[0032] For some embodiments of the fourth example method, performing the copy back projection process includes at least one of performing a copy process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0033] In some embodiments of the fourth example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0034] In some embodiments of the fourth example method, the subprocess is a normalized flow (NF) subprocess.
[0035] In some embodiments of the fourth example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0036] A fourth example apparatus according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0037] Another example apparatus according to some embodiments may include at least one processor configured to perform one of the methods listed above.
[0038] Another example apparatus according to some embodiments may include a computer-readable medium storing instructions for causing one or more processors to perform one of the methods listed above.
[0039] Another example apparatus according to some embodiments may include at least one processor and at least one non-transient computer-readable medium storing instructions for causing the at least one processor to perform any of the methods listed above.
[0040] Example computer-readable media according to some embodiments may include: storing a scene description file generated according to one of the methods listed above.
[0041] Example signals according to some embodiments may include a scene description file generated according to any of the methods listed above.
[0042] In additional embodiments, encoder and decoder devices are provided to perform the methods described herein. The encoder or decoder device may include a processor configured to perform the methods described herein. The device may include a computer-readable medium (e.g., a non-transient medium) storing instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., a non-transient medium) stores video encoded using any of the methods described herein.
[0043] One or more of these embodiments also provide a computer-readable storage medium having instructions stored thereon for performing bidirectional optical flow, encoding or decoding video data according to any of the methods described above. This embodiment also provides a computer-readable storage medium having a bitstream generated according to the methods described above stored thereon. This embodiment also provides a method and apparatus for transmitting a bitstream generated according to the methods described above. This embodiment also provides a computer program product including instructions for performing any of the described methods. Attached Figure Description
[0044] Figure 1A This is a system diagram illustrating an example communication system according to some embodiments.
[0045] Figure 1B The illustration shows a configuration according to some embodiments. Figure 1A The diagram shows a system diagram of an example wireless transmit / receive unit (WTRU) used in a communication system.
[0046] Figure 1C This is a system diagram illustrating a set of example interfaces for a system according to some embodiments.
[0047] Figure 2A This is a block diagram illustrating an example system according to some embodiments.
[0048] Figure 2B This is a system diagram illustrating an example set of interfaces for networking according to some embodiments.
[0049] Figure 2C This is a diagram of message bit fields in an example group according to some embodiments.
[0050] Figure 3 This is a process diagram illustrating an example normalized stream architecture for compressing and decompressing point cloud data according to some embodiments.
[0051] Figure 4A This is a flowchart illustrating an example feature enhancement layer according to some embodiments.
[0052] Figure 4B This is a diagram illustrating an example feature enhancement process according to some embodiments.
[0053] Figure 4C This is a process diagram illustrating an example coupling layer according to some embodiments.
[0054] Figure 4D This is a flowchart illustrating an example transformation block for a coupling layer according to some embodiments.
[0055] Figure 5 This is a flowchart illustrating an example process for encoding point clouds according to some embodiments.
[0056] Figure 6 This is a flowchart illustrating an example process for reconstructing a point cloud according to some embodiments.
[0057] Figure 7 This is a process diagram illustrating an example normalized flow non-average process according to some embodiments.
[0058] Figure 8A This is a process diagram illustrating an example average back projection process according to some embodiments.
[0059] Figure 8B This is a process diagram illustrating an example copy back projection process according to some embodiments.
[0060] Figure 9 This is a process diagram illustrating an example normalized flow back projection process according to some embodiments.
[0061] Figure 10 The diagram illustrates an example peak signal-to-noise ratio (PSNR) versus bits per input point for a first test scenario, according to some embodiments.
[0062] Figure 11 The figure illustrates an example peak signal-to-noise ratio (PSNR) versus bits per input point for a second test scenario according to some embodiments.
[0063] Figure 12 The figure illustrates an example peak signal-to-noise ratio (PSNR) versus bits per input point for a third test scenario according to some embodiments.
[0064] Figure 13 This is a flowchart illustrating an example process for encoding point clouds according to some embodiments.
[0065] Figure 14 This is a flowchart illustrating an example process for decoding a point cloud according to some embodiments.
[0066] As examples, not limitations, are presented the entities, connections, arrangements, etc., depicted in—and described in conjunction with—the various figures. Therefore, any and all statements or other indications concerning what is “depicted” in a particular figure, what a particular element or entity “is” or “has” in a particular figure, and any and all similar statements—which may be interpreted in isolation and out of context as absolute and therefore restrictive—may be properly interpreted only as being preceded by a clause such as “In at least one embodiment, …”. For the sake of brevity and clarity, this implies the unconventional repetition of the preceding clause in the specific embodiments. Detailed Implementation
[0067] Figure 1A This diagram illustrates an example communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content (such as voice, data, video, messaging, broadcasting, etc.) to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources (including wireless bandwidth). For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.
[0068] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106, Public Switched Telephone Network (PSTN) 108, Internet 110, and other networks 112. Although it will be appreciated, the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0069] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN 106, the Internet 110, and / or other networks 112. As an example, base stations 114a and 114b may be any of a base transceiver station (BTS), Node-B, eNode B, home node B, home eNode B, gNB, NR NodeB, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are depicted as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0070] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a specific geographic area for a radio service, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In one embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology, and multiple transceivers may be used for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.
[0071] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).
[0072] More specifically, as noted above, communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113, and WTRUs 102a, 102b, and 102c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish the air interface 116. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0073] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 116.
[0074] In one embodiment, base station 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as NR radio access, which can use a new radio (NR) to establish an air interface 116.
[0075] In one embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement multiple radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access, for example, using the dual connectivity (DC) principle. Therefore, the air interface utilized by WTRUs 102a, 102b, and 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0076] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c can implement the following radio technologies, such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.
[0077] Figure 1ABase station 114b can be, for example, a wireless router, a home node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a commercial area, home, vehicle, campus, industrial facility, air corridor (e.g., for drone use), road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. Figure 1A As shown, base station 114b may have a direct connection to Internet 110. Therefore, base station 114b may not be required to access Internet 110 via CN 106.
[0078] RAN 104 / 113 can communicate with CN 106, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, and / or perform advanced security functions, such as user authentication. Although... Figure 1A As not shown, but will be understood, RAN 104 / 113 and / or CN106 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113, which can utilize NR radio technology, CN 106 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0079] CN 106 may also act as a gateway for WTRUs 102a, 102b, 102c, and 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.
[0080] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can use cellular-based radio technology and a base station 114b that can use IEEE 802 radio technology.
[0081] Figure 1B This is a system diagram illustrating example WTRU 102. (Example:) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138, etc. It will be appreciated that WTRU 102 may include any sub-combination of the above-described elements while remaining consistent with the embodiments.
[0082] Processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 may be coupled to transceiver 120, which may be coupled to transmitting / receiving element 122. Although... Figure 1B The processor 118 and transceiver 120 are depicted as separate components, but it will be understood that the processor 118 and transceiver 120 can be integrated together in an electronic package or chip.
[0083] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF and optical signals. It will be appreciated that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0084] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.
[0085] Transceiver 120 can be configured to modulate signals to be transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As noted above, WTRU 102 can have multi-mode capability. Thus, for example, transceiver 120 may include multiple transceivers for enabling WTRU 102 to communicate via multiple RATs (such as NR and IEEE 802.11).
[0086] The processor 118 of WTRU 102 can be coupled to the speaker / microphone 124, keypad 126, and / or display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit), and can receive user input data from them. The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Additionally, the processor 118 can access information from any type of suitable memory (such as non-removable memory 130 and / or removable memory 132), and store data in that memory. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, processor 118 may access information from memory that is not physically located on WTRU 102 (such as on a server or home computer (not shown)) and store data in that memory.
[0087] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control the power going to other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0088] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to, or instead of, information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information using any suitable location determination method, while remaining consistent with the embodiments.
[0089] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripherals 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripherals 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0090] WTRU 102 may include a full-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or via signal processing (e.g., a separate processor (not shown) or via processor 118). In one embodiment, WTRU 102 may include a half-duplex radio, for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.
[0091] Despite WTRU in Figures 1A to 1B While described as a wireless terminal, it is envisioned that, in some representative embodiments, such a terminal may use (e.g., temporarily or permanently) a wired communication interface with a communication network.
[0092] In a representative embodiment, the other network 112 may be a WLAN.
[0093] Given Figures 1A to 1B The description, along with its corresponding information, indicates that one or more of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0094] Simulation devices can be designed to perform tests on one or more other devices in a laboratory environment and / or a carrier network environment. For example, the one or more simulation devices can perform one or more or all of their functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices can perform one or more or all of their functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices can be directly coupled to another device for testing purposes and / or can use over-the-air wireless communication to perform tests.
[0095] The one or more simulation devices can perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation devices can be used in test scenarios in a test laboratory and / or in non-deployed (e.g., testing) wired and / or wireless communication networks to perform testing on one or more components. The one or more simulation devices can be test rigs. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) can be used by the simulation devices to transmit and / or receive data.
[0096] Figure 1C This is a system diagram illustrating an example set of interfaces for a system according to some embodiments. Extended reality display devices, together with their control electronics, can be implemented using a system such as the system of Figure 1D. System 150 can be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 150 can be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 150 are distributed across multiple ICs and / or discrete components. In various embodiments, system 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 150 is configured to implement one or more of the aspects described in this document.
[0097] System 150 includes: at least one processor 152 configured to execute instructions loaded therein for implementing various aspects, such as those described in this document. Processor 152 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 150 includes at least one memory 154 (e.g., a volatile memory device and / or a non-volatile memory device). System 150 may include: a storage device 158 which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. Storage device 158 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices, as non-limiting examples.
[0098] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 156 may include its own processor and memory. The encoder / decoder module 156 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 156 may be implemented as a separate element of system 150, or may be incorporated into processor 152 as a combination of hardware and software as known to those skilled in the art.
[0099] Program code to be loaded onto processor 152 or encoder / decoder 156 to execute the various aspects described herein may be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. According to various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder / decoder module 156 may store one or more items of various kinds during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0100] In some embodiments, the memory within processor 152 and / or encoder / decoder module 156 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 152 or encoder / decoder module 152) is used for one or more of these functions. External memory may be memory 154 and / or storage device 158, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Universal Video Coding: a new standard developed by the Joint Video Experts Team JVET)).
[0101] Inputs to the components of system 172 can be provided through various input devices as indicated in box 150. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives RF signals transmitted over the air, for example, by a broadcaster; (ii) component (COMP) input terminals (or a collection of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1C Other examples not shown include composite video.
[0102] In various embodiments, the input device of block 172 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal or limiting a signal to a frequency band); (ii) down-converting the selected signal; (iii) further limiting the frequency band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select the desired stream of data packets. The RF section in various embodiments includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various embodiments rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0103] Additionally, the USB and / or HDMI endpoints may include corresponding interface processors for connecting system 150 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or, if necessary, within processor 152. Similarly, aspects of USB or HDMI interface processing may be implemented, either within a separate interface IC or, if necessary, within processor 152. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 152 and encoder / decoder 156, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0104] Various components of system 150 can be provided within an integrated housing, where the components can be interconnected and data can be transmitted therebetween using a suitable connection arrangement 174, such as an internal bus as known in the art, including inter-IC (I2C) bus, wiring, and printed circuit board.
[0105] System 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 162. The communication interface 160 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 162 may be implemented over, for example, wired and / or wireless media.
[0106] In various embodiments, data is streamed or otherwise provided to system 150 using a wireless network such as WiFi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). In these embodiments, the Wi-Fi signal is received on a communication channel 162 and a communication interface 160 adapted for Wi-Fi communication. The communication channel 162 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data over an HDMI connection in input box 172 to provide streaming data to system 150. Still other embodiments use an RF connection in input box 172 to provide streaming data to system 150. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0107] System 150 can provide output signals to various output devices, including display 176, speaker 178, and other peripheral devices 180. Display 176 in various embodiments includes one or more of the following: for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 176 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. Display 176 can also be integrated with other components (e.g., as in smartphones) or separate (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 180 include one or more of a standalone digital video disc (or digital multifunction disc) (DVR, for both terms), disc player, stereo system, and / or lighting system. Various embodiments use one or more peripheral devices 180 that provide functionality based on the output of system 150. For example, a disc player performs the function of playing the output of system 150.
[0108] In various embodiments, signaling such as AV is used to transmit control signals between system 150 and display 176, speaker 178, or other peripheral devices 180. Device-to-device control links, consumer electronics control (CEC), or other communication protocols are implemented with or without user intervention. Output devices can be communicatively coupled to system 150 via dedicated connections through corresponding interfaces 164, 166, and 168. Alternatively, output devices can be connected to system 150 via communication interface 160 using communication channel 162. Display 176 and speaker 178 can be integrated into a single unit with other components of system 150 in electronic devices such as, for example, televisions. In various embodiments, display interface 164 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0109] Display 176 and speaker 178 can alternatively be separated from one or more other components, for example, if the RF section of input 172 is part of a separate set-top box. In various embodiments where display 176 and speaker 178 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0110] System 150 may include one or more sensor devices 168. Examples of usable sensor devices include one or more GPS sensors, gyroscope sensors, accelerometers, light sensors, cameras, depth sensors, microphones, and / or magnetometers. Such sensors can be used to obtain information such as the user's position and orientation. Where system 150 is used as a control module (such as control modules 124, 132) for an extended reality display, the user's position and orientation can be used to determine how image data is presented, so that the user perceives the correct portion of a virtual object or scene from the correct viewpoint. In the case of a head-mounted display device, the device's own position and orientation can be used to determine the user's position and orientation for the purpose of presenting virtual content. In the case of other display devices such as telephones, tablets, computer monitors, or televisions, other inputs can be used to determine the user's position and orientation for the purpose of presenting content. For example, a user can select and / or adjust the desired viewpoint and / or viewing direction using a touchscreen, keypad or keyboard, trackball, joystick, or other inputs. When the display device has sensors such as accelerometers and / or gyroscopes, the viewpoint and orientation used for the purpose of presenting content can be selected and / or adjusted based on the movement of the display device.
[0111] The embodiments may be implemented by computer software, hardware, or a combination of hardware and software, as implemented by processor 152. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, memory 154 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 152 may be of any type suitable for the technical environment and may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0112] The use of 3D applications is becoming increasingly popular every day. To leverage these applications, different data formats are being used. One such data format is a point cloud. A point cloud is an unordered collection of points with coordinates (x, y, z) corresponding to their locations in space and their attributes, such as color and normal vectors.
[0113] The use of these new types of data will likely require new compression methods to efficiently store and transmit the data, especially since point clouds can potentially contain millions of points. Among the various proposed methods, learning-based architectures are gaining momentum. These architectures are often extensions of learning-based methods explored in the 2D image domain.
[0114] Conference paper by L. Dihn, J. Sohl-Dickstein, and S. Bengio, Density Estimation Using Real NVP ARXIV: 1605.08803v3 (2016) is understood to explore the use of normalized flow as a compression architecture in the 2D image compression domain. However, this method uses a squeezing layer that is not adapted for point cloud data structures. Normalized flow is a type of architecture that generates a latent space from the input. The latent space is a representation of the input with different coefficients. The goal is to generate a latent space that is easier to compress than the original input.
[0115] This document presents two architectures for performing compression, modified from example embodiments of the '317 application. For some embodiments, these new architectures aim to enhance performance and reduce network complexity, such as memory complexity. Some example implementations of the example embodiments of the '317 application have 278 million parameters.
[0116] Figure 2A This is a block diagram illustrating an example system according to some embodiments. Figure 2A A block diagram of a system in which aspects of this embodiment can be implemented according to another embodiment is illustrated. Figure 2AAn embodiment of an apparatus 200 for encoding or decoding point clouds or point cloud attributes, as described in any of the embodiments described herein, is shown. The apparatus may include a processor 210 and may be interconnected to a memory 220 via at least one port. Both the processor 210 and the memory 220 may also have one or more additional interconnects to external connections.
[0117] Using any of the embodiments described herein, processor 220 is also configured to use a reversible neural network to encode one or more attributes of the point cloud. For example, processor 210 is configured using a computer program product comprising code instructions that implement any of the embodiments described herein.
[0118] Figure 2B This is a system diagram illustrating an example set of interfaces for networking according to some embodiments. Figure 2B In one embodiment illustrated in the diagram, in a transmission scenario between two remote devices A and B 230, 234 on a communication network NET 232, device A 230 includes a processor associated with RAM and ROM, configured to implement methods for encoding point clouds, as described above. Figure 3-13 As described, and device B includes a processor associated with memory RAM and ROM, configured to implement methods for decoding point clouds, as per [reference to...]. Figure 3-13 As described. According to one example, the network is: a broadcast network adapted to broadcast / transmit encoded point clouds from device A to decoding devices including device B.
[0119] Figure 2C This is a diagram of message bit fields in an example group according to some embodiments. Figure 2C An example of the syntax for a signal transmitted over a packet-based transport protocol is shown. Each transmitted packet P includes a header H 260 and a payload PAYLOAD 262. In some embodiments, the payload PAYLOAD 262 may include encoded point cloud data according to any of the embodiments described above. In a variation, the signal includes a tag indicating a deep learning method for decoding the point cloud or for decoding one or more attributes of the point cloud.
[0120] Figure 3This is a process diagram illustrating an example normalized flow (NF) architecture for compressing and decompressing point cloud data according to some embodiments. The reversibility of the normalized flow architecture allows it to differ from many other architecture types. After the latent space is created, the original input can be reconstructed by applying the architecture in a reversed or inverse manner. This property is particularly interesting considering that data compression can be naturally viewed as an inverse problem. To enable the use of this architecture, a squash operation is performed, thereby efficiently compromising the spatial size of the input relative to the channel so that the operation can be performed. (See '317 application and journal article Pinheiro, R. Borba, et al.) NF-PCAC: Normalizing Flow Based Point Cloud Attribute Compression , 2023 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS,SPEECH, AND SIGNAL PROCESSING(ICASSP) (2023) (“ Pinheiro The paper “” uses a reversible neural network (flow) block. According to the paper by Xie, Y., Cheng, KL, and Chen, Q., Enhanced Invertible Encoding for Learned Image Compression, ARXIV:2108.03690v1 (2021) (“ Xie The architecture has been adapted. Figure 3 The architecture is shown in the diagram.
[0121] Figure 3 The diagram illustrates the encoding and decoding processes that generate bitstreams between each step. Figure 3 A normalized streaming architecture adapted for compression of color attributes of 3D point clouds is illustrated. In some embodiments, the point cloud is input into an example architecture 300, which includes a feature enhancement block 302, a reversible neural network (INN) (or streaming) block 304, a channel averaging block 306, and an attention layer 308.
[0122] On the encoding side, the normalized streaming architecture generates a latent space encoded by an entropy encoder 312 to produce a bitstream. Figure 3 In the example, the entropy encoder 312 is a neural network-based encoder coupled to the hyperprior encoder 310. For Figure 3The example configuration shown receives a previous version of the output of attention layer 308 and the output of super-prior decoder 314 as input to entropy encoder 312. Super-prior encoder 310 may include a neural network with sparse convolutions for transforming the output of attention layer 308 into a bitstream with side information. This side information is used by entropy encoder 312 to output the main bitstream. In some embodiments, the purpose of super-prior encoder 310 is to generate side information for use by the (main) entropy encoder. In some embodiments, the output of super-prior encoder 310 is passed through super-prior decoder 314, whose output is used as input to entropy encoder 312. The output of attention layer 308 corresponds to the latent space of the original signal. The output of super-prior decoder 314, after passing through super-prior encoder 310, provides contextual information for encoding the original latent space. In terms of data shape, the super-prior decoder output can provide a mean and scale for each coefficient in the latent space. Therefore, the super-prior may have doubled the size of the latent space output by attention layer 308.
[0123] On the decoder side, the bitstream is entropy-decoded using, for example, a neural network-based entropy decoder 316 coupled with a super-prior decoder 314 to provide a decoded latent space. The decoded latent space is passed to an example normalized stream architecture to produce a reconstructed point cloud, which includes an attention layer 318, a channel copy layer 320, a reversible neural network 304, and a feature enhancement block 322.
[0124] In some embodiments, the feature enhancement layer aims to extract more nonlinear features from the original point cloud. Figure 3 In the architecture shown, the feature enhancement layer provides features to INN block 304.
[0125] In some embodiments, the reversible neural network 304 includes three sets: a voxel shuffling layer, a 1×1 convolution, and a set of coupling layers. Because each layer has its own weights and biases, the three sets are not represented as a loop repeated three times. The number of channels changes as the data passes through the voxel shuffling layer in each set. As explained below, the size of the filters in subsequent convolutions also changes. Xie The architecture contains a sequence of four repetitions of reversible blocks, which includes a pixel shuffling layer, a 1×1 convolution, and three coupling layers.
[0126] exist Figure 3In this architecture, the number of reversible blocks is reduced compared to other architectures (such as the example architecture shown in '317 application). This reduction is driven by the evolution from 2D architectures to 3D domains. When new dimensions are added without reducing the number of repetitions of reversible blocks, the number of coefficients explodes, and the usability of the network may be negatively affected. The number of coupling layers in reversible blocks may also be reduced for the same reason. For example, a squash operation may include a 1×1 sparse convolution followed by two coupling layers.
[0127] To achieve the desired performance and accelerate convergence of INNs, voxel shuffling layers can be specifically designed for sparse 3D data. In some embodiments, voxel shuffling layers aim to efficiently compromise the spatial dimension of the channel without losing any information. In some embodiments, 1×1 convolutions of reversible blocks aim to enhance the feature representation for the coupling layer. Figure 3 In the INN diagram above, the 1×1 convolution is a sparse 1×1 convolution, not... Xie The 1×1 convolution shown is an example. The coupling layer can be a series of invertible transformations applied to the input tensor. Figure 3 In some embodiments, all convolutions used in the coupling layer can be sparse 3D convolutions.
[0128] In some embodiments, the channel averaging layer is a layer in which the number of channels for the latent space is reduced by averaging all channels at each spatial location. The voxel shuffling layer can be specifically designed for 3D sparse data. Without such a design for 3D sparse data, channel averaging can consider several zeros, thereby passing the distortion coefficients to the attention layer.
[0129] In some embodiments, the attention layer aims to focus on "more important" regions in the point cloud. In some embodiments, the attention layer uses a sigmoid function as part of a weighting function for the encoder to allocate more bits to certain regions of the point cloud data. Figure 3 The attention blocks shown in the diagram are different. Xie This is because Xie All regular 2D convolutions are replaced by sparse 3D convolutions, which enables the process to handle the sparsity and extra dimensions of point clouds.
[0130] The sparse nature and high number of points in point cloud representations are typically not well-suited for use with conventional 3D convolutions. Therefore, Figure 3 The architecture illustrated in the diagram uses sparse convolutions. In some embodiments, a specially designed 3D voxel shuffling layer is provided, which allows... Figure 3 The system shown in the diagram converges faster.
[0131] In some embodiments, layers of a network that can become very large with more than 270 million parameters may not necessarily be used in compression. In some embodiments, the objective may be to enhance the performance of the architecture in the high bit rate domain while also reducing the number of coefficients.
[0132] In this regard, two new architectures (e.g., such as) are adapted from the example of point cloud attribute compression based on normalized flow. Pinheiro As shown in the examples, and for example, the embodiment according to '317 application, reduces the network size. Such architectures can be standardized for purposes without sacrificing performance. These architectures can improve performance compared to other models and provide better control over the number of channels in the model. This control was considered impossible with previous normalized flow architectures.
[0133] The first new architecture is in relation to, for example Figure 3 The architecture shown does not increase the size of the super-prior model while adding no new blocks, and excludes the averaging layer and the last block in the core of the normalized flow (NF) architecture.
[0134] The second architecture uses the data from Haris, M., Shakhnarovich, G., and Ukita, N. Deep Back-Projection Networks for Super-Resolution , 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2018) (“ Haris The strategy used in super-resolution fields, known as back projection, is described. This second architecture achieves better control over the number of channels while attempting to avoid information loss. Both the first and second architectures encode video pairwise point cloud attributes.
[0135] Figure 4A This diagram illustrates the process of an example feature enhancement layer according to some embodiments. The feature enhancement layer 400 aims to help extract more non-linear features from the original point cloud. Figure 4A The example of the feature enhancement layer shown in the figure consists of a dense block architecture 402 followed by three 3D sparse convolutions (404, 406, 408) with a kernel size of 3×3×3, followed by another dense block architecture 410.
[0136] Figure 4B This is a diagram illustrating an example feature enhancement process according to some embodiments. Figure 4BThe diagram illustrates an example of a sparse-dense block 400 that can be used in a feature enhancement layer. The dense block 400 aims to preserve initial features along the convolution by concatenating (cat) the outputs of previous convolutions into the output of the current convolution. This block has been shown to enhance the performance of learning-based architectures. Figure 7 In the architecture illustrated, the feature enhancement layer provides better features for the INNs at the core of the illustrated network. The feature enhancement layer consists of... Xie One of the architectures inspired by this, but which uses sparse 3D convolutions instead of... Xie The 2D regular convolution is used. In some embodiments, the convolution can be performed by 3D sparse convolution 3×3×3 blocks 422, 426, 430, 434, 438. In some embodiments, LeakyReLU blocks 424, 428, 432, 436 can be used between the 3D sparse convolution 3×3×3 blocks 422, 426, 430, 434, 438 and the cat blocks.
[0137] Figure 4C This diagram illustrates the process of an example coupling layer according to some embodiments. The coupling layer 450 is a series of fully reversible transformations applied to the input tensor. According to... Figure 4C In the scheme illustrated above, the input tensor 452 is split into two parts, and each part undergoes its own transformation. Figure 4C Above, the input tensor 452 is split into two parts: x1(456) and x2(454, 458), which are then transformed in y1(460, 464) and y2(462) before being concatenated again to provide the output y470. Each of the transformations G1(476), G2(472), H1(478), and H2(474) consists of three sparse 3D convolutions.
[0138] Figure 4D This is a flowchart illustrating an example transformation block for a coupling layer according to some embodiments. Figure 4D An example of a transform block 480 for the coupled layer, which can be used to transform G1, G2, H1, and H2, is illustrated. Within the coupled layer, each transform consists of three levels (482, 486, 490) of 3D sparse convolutions with a kernel of 3×3×3 and LeakyReLU 484, 488. LeakyReLU is a deep learning block. LeakyReLU blocks are similar to ReLU blocks, except that they replace only positive values; LeakyReLU also "leaks" negative values. In some embodiments, the amount of "leakage" can be indicated using parameters.
[0139] The arrangement of the transformations ensures reversibility. In the embodiments illustrated in Figures 2 and 4C, all convolutions used in the coupling layer are sparse 3D convolutions.
[0140] The channel averaging layer is a layer that reduces the number of channels for the latent space by averaging all channels at each spatial location. The voxel shuffling layer, specifically designed for 3D sparse data (described further below), is particularly important here because without it, channel averaging would take into account several zeros, thus passing the distortion coefficients to the attention layer.
[0141] The attention layer serves to help the architecture focus on more important regions in the point cloud. It uses a sigmoid function to provide weights, informing the encoder which regions of the point cloud will require more bits to be encoded. This is further enhanced by replacing [the previous method] with sparse 3D convolutions. Xie All regular 2D convolutions are used to modify the attention layer block illustrated in Figure 2 to handle the sparsity and extra dimensions of the point cloud.
[0142] The sparse nature and high number of points in point cloud representations are not well adapted to the use of conventional 3D convolutions. Therefore, the architecture illustrated in Figure 2 uses sparse convolutions.
[0143] Figure 5 This is a flowchart illustrating an example process for encoding point clouds according to some embodiments. Figure 5 The above diagram illustrates an example of using a reversible neural network to encode one or more attributes of a point cloud. Figure 5 The diagram illustrates an example block diagram of a method for encoding one or more attributes of a point cloud. At 502, the point cloud is provided as input to the encoding system. The encoding system includes at least a reversible neural network configured to encode at least one attribute of the point cloud. In some variations, the encoding system includes a geometric encoding module configured to encode the geometry of the point cloud. At 504, a latent representation of at least one attribute of the point cloud is obtained using at least the reversible neural network. At 506, the latent representation is encoded to produce a bitstream, for example, using a neural network-based entropy encoder.
[0144] Figure 6 This is a flowchart illustrating an example process for reconstructing a point cloud according to some embodiments. Figure 6 The above illustration shows another embodiment for encoding one or more attributes of a point cloud using a reversible neural network. Figure 6The diagram illustrates an example block diagram of a method for decoding one or more attributes of a point cloud. At 602, a bitstream is provided to a decoding system. The bitstream includes at least encoded data representing at least one attribute of the point cloud. The decoding system includes at least a reversible neural network configured to decode at least one attribute of the point cloud. In some variations, the bitstream also includes encoded data representing the geometry of the point cloud, and the decoding system includes a geometry decoding module configured to decode and reconstruct the geometry of the point cloud. At 604, a latent representation of the at least one attribute is obtained by decoding the portion of the bitstream representing the at least one attribute of the point cloud, for example, using an entropy decoder based on a neural network. At 606, the at least one attribute is reconstructed using at least a reversible neural network.
[0145] Normalized Flow Non-Average (NF-No-Avg) Architecture Figure 7 This is a process diagram illustrating an example normalized flow non-average process according to some embodiments. In some embodiments, for ease of description, this example process may be referred to as "NF-No Avg," although various designs and architectures are conceived. Architecture 700 is... Figure 3 A smaller version of the architecture shown. For some embodiments, this leads to irreversibility. Figure 3 The main component in the initial architecture is the channel averaging layer. Figure 3 In the example, the channel averaging block is connected to the output of the INN block, which, in some embodiments, can be designated as the NF core. This strong irreversibility of the network can lead to saturation of the results for high bit rates. This saturation refers to the rate distortion curve. Negatively, even with a very large bit rate allocated for a given point cloud, the reconstruction quality does not change much and may have degraded rate distortion performance. Combining the channel averaging layer with the substitution of empty voxels using the average of their neighbors results in some smoothing of the properties of the original point cloud. The substitution of empty voxels may occur in the sparse voxel shuffling block. The sparse voxel shuffling block can be used in various architectures and can contribute to performance. Figure 3 In the example, sparse voxel shuffling is performed in the voxel shuffling layer.
[0146] Figure 3 The voxel shuffling layer shown is maintained as in Figure 7 The new architecture remains intact. However, the last group of INN 704 blocks, including the voxel shuffling layer, 1×1 sparse convolutions, and coupling layer set, is removed to limit the increase in the number of channels. This allows the removal of the channel averaging layer while maintaining a reasonable amount of computational complexity and coefficient number. Depending on some implementations... Figure 7 The example architecture has 25 million parameters, which is smaller than, depending on some implementations... Figure 310% of the 278 million parameters.
[0147] In some embodiments, point clouds are input into an example architecture 700, which includes a feature augmentation block 702, a reversible neural network (INN) (or stream) block 704, and an attention layer 706. Figure 7 The example configuration shown receives a previous version of the output of the attention layer 706 and the output of the super-prior decoder 712 as input to the entropy encoder 708. In some embodiments, the output of the super-prior encoder 710 is passed through the super-prior decoder 712, and its output is used as input to the entropy encoder 708.
[0148] On the decoder side, the bitstream is entropy-decoded using, for example, a neural network-based entropy decoder 714 coupled with a super-prior decoder 712 to provide a decoded latent space. The decoded latent space is passed to an example normalized stream architecture to produce a reconstructed point cloud, which includes an attention layer 716, a reversible neural network 704, and a feature enhancement block 718.
[0149] Normalized Flow Back Projection (NF-BP) Architecture Figure 8A This is a process diagram illustrating an example average back projection process according to some embodiments. Figure 8B This is a process diagram illustrating an example copy backprojection process according to some embodiments. The second architecture uses average backprojection blocks and copy backprojection blocks, such as... Figure 8A and 8B As shown in the diagram. In some embodiments, these blocks serve the purpose of reducing the number of channels in the inner layers of the normalized flow block (or core) while attempting to retain all desired information in the channels. Average backprojection blocks and copy backprojection blocks are used in the super-resolution architecture to enhance the reconstruction by utilizing details lost in the averaging process.
[0150] Figure 8A The “Av↓” (down arrow) block 804 performs averaging on the channel of the tensor, thereby reducing the dimensionality. Figure 8A and 8B The “Cp↑” (up arrow) blocks 806 and 854 perform the inverse operation and copy the channel to restore the same dimension as before.
[0151] exist Figure 8AIn architecture 800, the output of “Cp↑” block 806 is subtracted from the original input 802 to form the input to “Cs↓” block 808. Taking into account the residuals generated by the difference between the original input tensor and the tensor to be reconstructed in a copy after averaging, “Cs↓” block 808 performs downsampling. The output of “Cs↓” block 808 is added to the averaging tensor to produce the output 810 of the ABP block. In some embodiments, “Cs↓” block 808 may include a convolutional layer that produces a residual (a portion of the “Cs↓” block output) to be added to the averaging vector (“Av↓” block output) to produce the output of the average backprojection block, which, in some embodiments, aims to obtain a downsampled tensor (“Cs↓” block output) that is more representative of the original data than the naive average.
[0152] Figure 8B The copy back projection block architecture 850 is complete. Figure 8A The inverse operation of the average backprojection block is performed. This generates an upsampled version of the data, as output 860. The output of the “Cs↓” block 856 is subtracted from the original input 852 to form the input to the “Cs↑” block 858. Taking into account the residual generated by the difference between the original input tensor and the tensor to be reconstructed in downsampling after copying, the “Cs↑” block 858 performs upsampling. The output of the “Cs↑” block 858 is added to the copied tensor to produce the output 860 of the CBP block. In some embodiments, the “Cs↑” block 858 may include a convolutional layer that produces a residual (a portion of the “Cs↑” block output) to be added to the copied tensor (the “Cp↑” block output) to produce an output (the “Cs↑” block output) of an upsampled tensor that is more representative of the original data than a naive copy.
[0153] Figure 9 This is a process diagram illustrating an example normalized flow back projection process according to some embodiments. In some embodiments, for ease of description, this example process 900 may be referred to as "NF-BP," although various designs and architectures are conceived. Figure 9 On the encoding side of the bitstream, an average back projection (ABP) block is inserted after each sequence of voxel shuffling, 1×1 sparse convolution, and coupling layer set. Figure 9 On the decoding side of the bitstream, a copy back projection (CBP) block is inserted before each sequence of voxel shuffling, 1×1 sparse convolution, and coupling layer set.
[0154] These ABP and CBP blocks help reduce the size of the original model by progressively reducing the number of channels along the core of the network, which can be done for each sequence of voxel shuffling, 1×1 sparse convolutions, and coupling layer sets. Depending on some implementations... Figure 9 The example architecture uses 18 million coefficients, which is also smaller than according to Figure 3Some implementations use 10% of the original 278 million parameters. For some embodiments, only two sequences of voxel shuffling, 1×1 sparse convolution, coupling layer sets, and ABP or CBP blocks are performed (instead of...). Figure 9 (The three sequences shown). In some embodiments, four or more sequences of voxel shuffling, 1×1 sparse convolution, coupling layer set, and ABP or CBP block are performed (instead of...). Figure 9 The three sequences shown.
[0155] In some embodiments, point clouds are input into an example architecture 900, which includes a feature augmentation block 902, a reversible neural network (INN) (or stream) block 904, and an attention layer 906. Figure 9 The example configuration shown receives a previous version of the output of the attention layer 906 and the output of the super-prior decoder 912 as input to the entropy encoder 908. In some embodiments, the output of the super-prior encoder 910 is passed through the super-prior decoder 912, and its output is used as input to the entropy encoder 908.
[0156] On the decoder side, the bitstream is entropy-decoded using, for example, a neural network-based entropy decoder 914 coupled with a super-prior decoder 912 to provide a decoded latent space. The decoded latent space is passed to an example normalized stream architecture to produce a reconstructed point cloud, which includes an attention layer 916, a reversible neural network 904, and a feature enhancement block 918.
[0157] Comparison with the original architecture Figure 10 The figure illustrates a graph 1000 showing an example peak signal-to-noise ratio (PSNR) versus bits per input point for a first test scenario according to some embodiments. Figure 11 The figure 1100 illustrates an example peak signal-to-noise ratio (PSNR) versus bits per input point for a second test scenario according to some embodiments. Figure 12 This is a graph 1200 illustrating an example peak signal-to-noise ratio (PSNR) versus bits per input point for a third test scenario according to some embodiments. Regarding... Figure 10 and 12 See d'Eon, E., et al., 8i Voxelized Full Bodies - A Voxelized Point Cloud Dataset ISO / IEC JTC1 / SC29 JOINT WG11 / WG1 (MPEG / JPEG) INPUTDOCUMENTWG11M40059 / WG1M74006 (2017). About Figure 11 See Xu, Y., et al. Owlii Dynamic Human Mesh Sequence Dataset , ISO / IEC JTC1 / SC29 / WG11 M41658, 120TH MPEGMEETING (2017).
[0158] Figure 7 The NF-No Avg architecture shown outperforms other learning-based processes applied in the high bit rate domain. The removal of averaging allows the network to perform well over the high bit rate range. However, for low bit rates, other methods may outperform the NF-No Avg architecture. This result may be due to the additional unnecessary coefficients created by voxel shuffling that are not filtered in the channel averaging. For some embodiments, Figure 10 The GPCC curve can be used as a basis, where, for example, the first two or three data points in the curve represent a low bit rate and the last two or three data points represent a high bit rate. For example, in... Figure 10 In the GPCC curve, three data points less than 0.1 bits per input point can be considered as low bit rate, and three data points greater than 0.35 bits per input point can be considered as high bit rate.
[0159] On the other hand, the NF-BP architecture performs better in the low bit rate domain but is outperformed in the high bit rate domain. The progressive averaging of the channel by the network may lead to the loss of texture details in the point cloud.
[0160] Signaling / Syntax For rate-distortion optimization (RDO) schemes, different configurations are tested for a specific bit rate. For codecs implementing RDO schemes, they can be tested at rates higher than... Pinheiro Smaller networks (e.g., fewer parameters) achieve good performance over a wide range of bit rates. In some embodiments, two network configurations are tested, and the network configuration with better (or best) reconstruction quality is selected. Additional bits can be added to indicate to the decoder side which architecture was used on the encoder side. In some embodiments, signaling flags (which can be one or more bits) can be used to indicate the architecture configuration. In some embodiments, [the following can be utilized] Figures 7 to 9 The configuration shown replaces the normalized flow block in application '352'.
[0161] Current activity within MPEG AI-PCC is focused on defining AI models for compressing and decompressing point cloud geometry and photometry. The MPEG group currently handles photometry separately from geometry. The process described in this paper increases the number of available architectures for compression to perform photometry encoding / decoding. For example, the process described in this paper can be applied to deep learning-based extensions to the G-PCC coding standard.
[0162] The process described in this paper can be transformed into an RDO-based option for the codec by selecting an encoder that yields better rate distortion. Signaling tags can be added to the bitstream to indicate which depth decoder architecture to use. Table 1 shows an example mapping of architectures to be used for bitstreams based on point cloud (PC) learning.
[0163] The NF-No-Avg architecture is a smaller version of the normalization architecture used for point cloud compression with sparse convolution. The NF-BP architecture is used for point cloud compression with average back projection (ABP) and copy back projection (CBP) to improve reconstruction quality. Photometric codec architecture 2-bit tag describe 00 Using a variational autoencoder (VAE) 01 Use a normalized flow (NF) architecture (e.g., see the example architecture disclosed in '317 application). 10 Using the NF-No-Avg architecture (for example, its example is in) Figure 7 (as shown in the image) 11 Using the NF-BP architecture (for example, its example is in) Figure 9 (as shown in the image)
[0164] Figure 13 This is a flowchart illustrating an example process for encoding a point cloud according to some embodiments. In some embodiments, example process 1300 may include: obtaining 1302 information corresponding to the point cloud. In some embodiments, example process 1300 may further include: performing 1304 feature enhancement on the point cloud data corresponding to the information. In some embodiments, example process 1300 may further include: generating 1306 data corresponding to the latent space based on the information corresponding to the point cloud. In some embodiments of example process 1300, generating 1308 data corresponding to the latent space may include: using the feature-enhanced point cloud data as input data for a first loop of a subprocess; executing the subprocess twice, wherein the subprocess may include: passing the input data through a voxel shuffling layer to generate a voxel shuffling layer output; performing convolution on the voxel shuffling layer output to generate convolutional output data; passing the convolutional output data through one or more coupling layers to generate a subprocess output; and using the subprocess output as input data for a second loop of the subprocess; and passing the second loop subprocess output through an attention layer to generate data corresponding to the latent space. In some embodiments, example process 1300 may further include encoding data corresponding to the latent space 1310 into a bit stream.
[0165] Figure 14This is a flowchart illustrating an example process for decoding a point cloud according to some embodiments. In some embodiments, example process 1400 may include: obtaining an encoded bitstream at 1402. In some embodiments, example process 1400 may further include: decoding the encoded bitstream at 1404 to generate a latent space. In some embodiments, example process 1400 may further include: reconstructing the point cloud from the generated latent space at 1406. In some embodiments of example process 1400, reconstructing the point cloud from the generated latent space at 1408 may include: passing the generated latent space through an attention layer to generate input data for a sub-process; executing the sub-process twice, wherein the sub-process may include: passing the input data through one or more coupling layers to generate coupling layer outputs; and performing convolution on the coupling layer outputs to generate convolutional output data; passing the convolutional output data through a voxel shuffling layer to generate a sub-process output; and using the sub-process output as input data for a second loop following the sub-process. In some embodiments, example process 1400 may further include performing feature enhancement 1410 on the output of the second loop subprocess to generate a reconstructed point cloud.
[0166] While methods and systems according to some embodiments have been generally discussed in the context of extended reality (XR), some embodiments can be applied to any XR context such as, for example, virtual reality (VR) / mixed reality (MR) / augmented reality (AR). Furthermore, although the term "head-mounted display (HMD)" is used herein according to some embodiments, some embodiments can be applied to wearable devices (which may or may not be attached to the head) that have, for example, XR, VR, AR, and / or MR capabilities for some embodiments.
[0167] Example methods according to some embodiments may include: obtaining information corresponding to a point cloud; performing feature enhancement on point cloud data corresponding to the information; generating data corresponding to a latent space based on the information corresponding to the point cloud, wherein generating data corresponding to the latent space may include: using the feature-enhanced point cloud data as input data for a first loop through a subprocess; executing the subprocess twice, wherein the subprocess may include: passing the input data through a voxel shuffling layer to generate a voxel shuffling layer output; performing convolution on the voxel shuffling layer output to generate convolutional output data; passing the convolutional output data through one or more coupling layers to generate a subprocess output; and using the subprocess output as input data for a second loop through the subprocess; and passing the second loop subprocess output through an attention layer to generate data corresponding to the latent space; and encoding the data corresponding to the latent space as a bitstream.
[0168] In some embodiments of the example method, the sub-procedure may be executed three times.
[0169] In some embodiments of the example method, the sub-process further includes performing an average back projection process on the output of the one or more coupling layers.
[0170] For some embodiments of the example method, performing the average back projection process may include performing at least one of the following: performing an averaging process, a copying process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0171] In some embodiments of the example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0172] In some embodiments of the example method, the subprocess is a normalized flow (NF) subprocess.
[0173] In some embodiments of the example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0174] Example apparatuses according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0175] A second example method according to some embodiments may include: obtaining information corresponding to a point cloud; generating data corresponding to a latent space based on the information corresponding to the point cloud, wherein generating the data corresponding to the latent space may include: using feature-enhanced point cloud data as input data for a first loop through a subprocess; executing the subprocess twice, wherein the subprocess may include: passing the input data through a voxel shuffling layer to generate a voxel shuffling layer output; performing convolution on the voxel shuffling layer output to generate convolutional output data; passing the convolutional output data through one or more coupling layers to generate a subprocess output; and using the subprocess output as input data for a second loop through the subprocess; and generating data corresponding to the latent space based on the second loop subprocess output; and encoding the data corresponding to the latent space into a bitstream.
[0176] In some embodiments of the second example method, the sub-process is executed three times.
[0177] In some embodiments of the second example method, the sub-process further includes performing an average back projection process on the output of the one or more coupling layers.
[0178] For some embodiments of the second example method, performing the average back projection process may include performing at least one of the following: performing an averaging process, a copying process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0179] In some embodiments of the second example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0180] In some embodiments of the second example method, the subprocess is a normalized flow (NF) subprocess.
[0181] In some embodiments of the second example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0182] A second example apparatus according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0183] A third example method according to some embodiments may include: obtaining an encoded bitstream; decoding the encoded bitstream to generate a latent space; and reconstructing a point cloud from the generated latent space, wherein reconstructing the point cloud from the generated latent space may include: passing the generated latent space through an attention layer to generate input data for a sub-process; executing the sub-process twice, wherein the sub-process may include: passing the input data through one or more coupling layers to generate coupling layer outputs; performing convolution on the coupling layer outputs to generate convolution output data; passing the convolution output data through a voxel shuffling layer to generate a sub-process output; and using the sub-process output as input data for a second loop through the sub-process; and performing feature enhancement on the output of the second loop sub-process to generate the reconstructed point cloud.
[0184] In some embodiments of the third example method, the sub-process may be executed three times.
[0185] In some embodiments of the third example method, the sub-process may further include performing a copy back projection process on the output of the voxel shuffling layer.
[0186] For some embodiments of the third example method, performing the copy back projection process may include at least one of: performing a copy process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0187] In some embodiments of the third example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0188] In some embodiments of the third example method, the subprocess is a normalized flow (NF) subprocess.
[0189] In some embodiments of the third example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0190] A third example apparatus according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0191] A fourth example method / apparatus according to some embodiments may include: obtaining an encoded bitstream; decoding the encoded bitstream to generate a latent space; and reconstructing a point cloud from the generated latent space, wherein reconstructing the point cloud from the generated latent space may include: obtaining input data for a sub-process based on the generated latent space; executing the sub-process twice, wherein the sub-process may include: passing the input data through one or more coupling layers to generate coupling layer outputs; performing convolution on the coupling layer outputs to generate convolution output data; passing the convolution output data through a voxel shuffling layer to generate a sub-process output; and using the sub-process output as input data for a second loop through the sub-process; and generating the reconstructed point cloud based on the second loop sub-process output.
[0192] In some embodiments of the fourth example method, the sub-process may be executed three times.
[0193] In some embodiments of the fourth example method, the sub-process may further include performing a copy back projection process on the output of the voxel shuffling layer.
[0194] For some embodiments of the fourth example method, performing the copy back projection process includes at least one of performing a copy process, a downsampling process, and replacing at least one empty voxel with the average of two or more neighboring voxels.
[0195] In some embodiments of the fourth example method, the information corresponding to the point cloud further corresponds to one or more voxels.
[0196] In some embodiments of the fourth example method, the subprocess is a normalized flow (NF) subprocess.
[0197] In some embodiments of the fourth example method, the encoded bit stream may include signaling tags for indicating architectural configuration.
[0198] A fourth example apparatus according to some embodiments may include: a processor; and a non-transient computer-readable medium storing instructions that, when executed by the processor, operate to cause the apparatus to perform any of the methods listed above.
[0199] Another example apparatus according to some embodiments may include at least one processor configured to perform one of the methods listed above.
[0200] Another example apparatus according to some embodiments may include a computer-readable medium storing instructions for causing one or more processors to perform one of the methods listed above.
[0201] Another example apparatus according to some embodiments may include at least one processor and at least one non-transient computer-readable medium storing instructions for causing the at least one processor to perform any of the methods listed above.
[0202] Example computer-readable media according to some embodiments may include: storing a scene description file generated according to one of the methods listed above.
[0203] Example signals according to some embodiments may include a scene description file generated according to any of the methods listed above.
[0204] This disclosure describes a variety of aspects, including tools, features, embodiments, models, schemes, etc. Many of these aspects are described in detail and are often described in a manner that may sound limiting, at least to illustrate individual characteristics. However, this is for clarity of purpose and does not limit the disclosure or scope of those aspects. Indeed, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, this aspect can also be combined and interchanged with aspects described in earlier filings.
[0205] The aspects described and contemplated in this disclosure can be implemented in many different forms. Although some embodiments are specifically illustrated, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having bitstreams generated according to any of the described methods stored thereon.
[0206] In this disclosure, the terms "reconstructed" and "decoded" may be used interchangeably, as may the terms "pixel" and "sample," and the terms "image," "picture," and "frame." Typically, but not necessarily, the term "reconstructed" is used on the encoder side and "decoded" is used on the decoder side.
[0207] The terms HDR (High Dynamic Range) and SDR (Standard Dynamic Range) often convey specific values of dynamic range to those skilled in the art. However, additional embodiments are contemplated, wherein a reference to HDR is understood to mean "higher dynamic range" and a reference to SDR is understood to mean "lower dynamic range." Such additional embodiments are not constrained by any specific values of dynamic range that may often be associated with the terms "high dynamic range" and "standard dynamic range."
[0208] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." The use of such terms does not imply a sequence of modified operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but may occur, for example, before, during, or in the time period overlapping with the second decoding.
[0209] For example, various numerical values may be used in this disclosure. Specific values are for illustrative purposes only, and the aspects described are not limited to these specific values.
[0210] The embodiments described herein may be implemented by computer software or other hardware implemented by a processor, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The processor may be of any type suitable for the technical environment and may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures, as a non-limiting example.
[0211] Various implementations involve decoding. As used in this disclosure, “decoding” can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, dequantization, inverse transform, and differential decoding. In various embodiments, such a process also or alternatively includes processes performed by a decoder of the various implementations described in this disclosure, such as: extracting images from chunks (packed) of images, determining an upsampling filter to use and then upsampling the images, and flipping the images back to their intended orientation.
[0212] As a further example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description.
[0213] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” the term “encoding,” as used herein, can encompass all or part of a process performed, for example, on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such a process also, or alternatively, includes processes performed by an encoder of the various implementations described herein.
[0214] As a further example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in yet another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the context of the specific description.
[0215] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0216] Various embodiments involve rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often due to computational complexity constraints. Rate distortion optimization is generally formulated to minimize a rate distortion function, which is a weighted sum of rate and distortion. Different schemes exist for solving the rate distortion optimization problem. For example, a scheme can be based on an extensive test of all encoding options, including all considered modes or encoding parameter values, which provides a complete evaluation of their encoding costs and associated distortions in the reconstructed signal after encoding and decoding. Faster schemes can also be used to save encoding complexity, particularly by utilizing the computation of approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed signal. A hybrid of these two schemes can also be used, such as using approximate distortion for only some of the possible encoding options and full distortion for the others. Other schemes evaluate only a subset of the possible encoding options. More generally, many schemes employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortions.
[0217] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features in question can be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in appropriate hardware, software, and firmware. Methods can be implemented, for example, in a processor, where processor generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end users.
[0218] References to "an embodiment" or "an embodiment" or "an implementation" or "an implementation," and other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with an embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in one embodiment" or "in one implementation," and any variations appearing throughout this disclosure, do not necessarily all refer to the same embodiment.
[0219] Additionally, this disclosure may refer to "determining" each piece of information. Determining information may include one or more of, for example, estimation information, calculation information, prediction information, or information retrieved from memory.
[0220] Furthermore, this disclosure may refer to "accessing" each piece of information. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0221] Additionally, this disclosure may refer to "receiving" individual pieces of information. As with "access," receiving is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., from memory). Further, "receiving" typically refers to actions performed during operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0222] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of” (e.g., in the cases of “A / B”, “A and / or B”, and “at least one of A and B”) is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this phrase is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to as many items as are listed.
[0223] Moreover, as used herein, the term “signaling” refers, among other things, to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a specific one of several parameters for selecting region-based filter parameters used for artifact removal filtering. In this way, in one embodiment, the same parameter is used at both the encoder and decoder sides. Thus, for example, the encoder can transmit (explicitly signal) the specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, then signaling can be used without transmission (implicitly signal) to simply allow the decoder to know and select the specific parameter. Bit saving is achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in a variety of ways. For example, in various embodiments, one or more syntax elements, tags, etc., are used to signal information to the corresponding decoder. Although the signature refers to the verb form of the term “signaling,” the term “signaling” may also be used as a noun herein.
[0224] Implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored on a processor-readable medium.
[0225] We have described several embodiments. These embodiments may be provided individually or in any combination across various claim classes and types. Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim classes and types: • A bitstream or signal that includes one or more of the described syntax elements or their variants. • Bitstream or signal, which includes grammatical communication information generated according to any of the described embodiments. • Create and / or transmit and / or receive and / or decode bitstreams or signals, including one or more of the described syntax elements or their variants. • Create and / or transmit and / or receive and / or decode according to any of the described embodiments. • The method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the described embodiments. • TV, set-top box, cellular phone, tablet, or other electronic device that adapts the filter parameters according to any of the described embodiments. • TV, set-top box, cellular phone, tablet, or other electronic device that performs filter parameter adaptation and displays (e.g., using a monitor, screen, or other type of display) the resulting image according to any of the described embodiments. • TV, set-top box, cellular phone, tablet, or other electronic device that selects (e.g., using a tuner) a channel to receive signals including encoded images and performs filter parameter adaptations according to any of the described embodiments. TVs, set-top boxes, cellular phones, tablets, or other electronic devices that receive signals that encode images and perform filter parameter adaptations according to any of the described embodiments.
[0226] Note that the various hardware elements of one or more in the described embodiments are referred to as “modules” that implement (i.e., perform, execute, etc.) the functions described herein in conjunction with the corresponding Module 2. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that are considered suitable by those skilled in the art for a given implementation. Each described module may also include instructions executable to implement one or more functions described as being implemented by the corresponding module, and it should be noted that such instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and may be stored in any one or more suitable non-transient computer-readable media (such as collectively referred to as RAM, ROM, etc.).
[0227] Although features and elements have been described above in specific combinations, those skilled in the art will appreciate that each feature or element may be used individually or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magnetic-optical media, and optical media (such as CD-ROMs and digital multifunction discs (DVDs)). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method comprising: Obtain information corresponding to the point cloud; Data corresponding to the latent space is generated based on the information corresponding to the point cloud. The data generated corresponding to the latent space includes: The point cloud data corresponding to the information in the point cloud is used as the input data for the first loop of the subprocess; Execute the sub-procedure twice. The sub-processes mentioned above include: The input data is passed through a voxel shuffle layer to generate a voxel shuffle layer output. Perform convolution on the output of the voxel shuffle layer to generate convolutional output data; The convolutional output data is passed through one or more coupling layers to generate a subprocess output; and The output of the subprocess is used as input data for the second loop that passes through the subprocess; and Data corresponding to the implicit space is generated based on the output of the second loop subprocess; and The data corresponding to the hidden space is encoded into a bit stream.
2. The method of claim 1, wherein the sub-process is executed three times.
3. The method of any one of claims 1-2, wherein the sub-process further comprises: An average back projection process is performed on the output of the one or more coupling layers.
4. The method of any one of claims 1-3, wherein performing the average back projection process comprises: Perform at least one of the following: averaging, copying, downsampling, and replacing at least one empty voxel with the average of two or more neighboring voxels.
5. The method of any one of claims 1-4, wherein the information corresponding to the point cloud further corresponds to one or more voxels.
6. The method of any one of claims 1-5, wherein the subprocess is a normalized flow (NF) subprocess.
7. The method of any one of claims 1-6, wherein the encoded bit stream includes signaling tags for indicating architectural configuration.
8. An apparatus comprising: processor; as well as A non-transient computer-readable medium storing instructions that, when executed by the processor, cause the apparatus to perform the method as described in any one of claims 1 to 7.
9. A method comprising: Obtain the encoded bitstream; Decode the encoded bitstream to generate the hidden space; as well as Reconstructing the point cloud from the generated latent space, The point cloud reconstructed from the generated latent space includes: The input data for the subprocess is obtained based on the generated latent space; Execute the sub-procedure twice. The sub-processes mentioned above include: The input data is passed through one or more coupling layers to generate a coupling layer output; Perform convolution on the output of the coupling layer to generate convolution output data; The convolutional output data is passed through a voxel shuffling layer to generate the subprocess output; and The output of the subprocess is used as input data for the second loop that passes through the subprocess; and The reconstructed point cloud is generated based on the output of the second loop subprocess.
10. The method of claim 9, wherein the sub-process is executed three times.
11. The method of any one of claims 9-10, wherein the sub-process further comprises: A copy back projection process is performed on the output of the voxel shuffle layer.
12. The method of claim 11, wherein performing the copy back projection process comprises: Perform at least one of the following: a copying process, a downsampling process, and a substitution of at least one empty voxel using the average of two or more neighboring voxels.
13. The method of any one of claims 9-12, wherein the information corresponding to the point cloud further corresponds to one or more voxels.
14. The method of any one of claims 9-13, wherein the subprocess is a normalized flow (NF) subprocess.
15. The method of any one of claims 9-14, wherein the encoded bit stream includes signaling tags for indicating architectural configuration.
16. An apparatus comprising: processor; as well as A non-transient computer-readable medium storing instructions that, when executed by the processor, cause the apparatus to perform the method as described in any one of claims 9 to 15.
17. A method comprising: Obtain information corresponding to the point cloud; Perform feature enhancement on the point cloud data corresponding to the information; Data corresponding to the latent space is generated based on the information corresponding to the point cloud. The data generated corresponding to the latent space includes: The feature-enhanced point cloud data is used as input data for the first loop of the subprocess; Execute the sub-procedure twice. The sub-processes mentioned above include: The input data is passed through a voxel shuffle layer to generate a voxel shuffle layer output. Perform convolution on the output of the voxel shuffle layer to generate convolutional output data; The convolutional output data is passed through one or more coupling layers to generate a subprocess output; and The output of the subprocess is used as input data for the second loop that passes through the subprocess; and The output of the second loop subprocess is passed through the attention layer to generate data corresponding to the latent space; and The data corresponding to the hidden space is encoded into a bit stream.
18. The method of claim 17, wherein the sub-process is executed three times.
19. The method of any one of claims 17-18, wherein the sub-process further comprises: An average back projection process is performed on the output of the one or more coupling layers.
20. The method of any one of claims 17-19, wherein performing the average back projection process comprises: Perform at least one of the following: averaging, copying, downsampling, and replacing at least one empty voxel with the average of two or more neighboring voxels.
21. The method of any one of claims 17-20, wherein the information corresponding to the point cloud further corresponds to one or more voxels.
22. The method of any one of claims 17-21, wherein the subprocess is a normalized flow (NF) subprocess.
23. The method of any one of claims 17-22, wherein the encoded bit stream includes signaling tags for indicating architectural configuration.
24. An apparatus comprising: processor; as well as A non-transient computer-readable medium storing instructions that, when executed by the processor, cause the apparatus to perform the method as described in any one of claims 17 to 23.
25. A method comprising: Obtain the encoded bitstream; Decode the encoded bitstream to generate the hidden space; as well as Reconstructing the point cloud from the generated latent space, The point cloud reconstructed from the generated latent space includes: The generated latent space is passed through an attention layer to generate input data for the subprocesses; Execute the sub-procedure twice. The sub-processes mentioned above include: The input data is passed through one or more coupling layers to generate a coupling layer output; Perform convolution on the output of the coupling layer to generate convolution output data; The convolutional output data is passed through a voxel shuffling layer to generate the subprocess output; and The output of the subprocess is used as input data for the second loop that passes through the subprocess; and Feature enhancement is performed on the output of the second loop subprocess to generate a reconstructed point cloud.
26. The method of claim 25, wherein the sub-process is executed three times.
27. The method of any one of claims 25-26, wherein the sub-process further comprises: A copy back projection process is performed on the output of the voxel shuffle layer.
28. The method of claim 27, wherein performing the copy back projection process comprises: Perform at least one of the following: a copying process, a downsampling process, and a substitution of at least one empty voxel using the average of two or more neighboring voxels.
29. The method of any one of claims 25-28, wherein the information corresponding to the point cloud further corresponds to one or more voxels.
30. The method of any one of claims 25-29, wherein the subprocess is a normalized flow (NF) subprocess.
31. The method of any one of claims 25-30, wherein the encoded bit stream includes signaling tags for indicating architectural configuration.
32. An apparatus comprising: processor; as well as A non-transient computer-readable medium storing instructions that, when executed by the processor, cause the apparatus to perform the method as described in any one of claims 25 to 31.
33. An apparatus comprising at least one processor configured to perform the method as described in any one of claims 1-7, 9-15, 17-23 and 25-31.
34. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method as described in any one of claims 1-7, 9-15, 17-23, and 25-31.
35. An apparatus comprising at least one processor and at least one non-transient computer-readable medium storing instructions for causing the at least one processor to perform the method as claimed in any one of claims 1-7, 9-15, 17-23 and 25-31.
36. A computer-readable medium storing a scene description file generated according to any one of claims 1-7, 9-15, 17-23 and 25-31.
37. A signal comprising a scene description file generated according to any one of claims 1-7, 9-15, 17-23 and 25-31.