Bit-wise deep octree coding based on sparse tensors
The proposed method for bitwise lossless compression of point clouds using sparse tensor processing and deep learning addresses inefficiencies in existing methods, achieving reduced storage and computational costs for point cloud data processing.
Patent Information
- Application Number
- JP2025517948
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-03
- Publication Date
- 2025-11-12
AI Technical Summary
Existing methods for compressing and processing point cloud data are inefficient, particularly for large and dynamic point clouds, leading to high computational costs and resource consumption, especially in consumer devices with limited computing power.
A method for bitwise lossless compression of voxelized point cloud data using sparse tensor processing and deep learning, which involves encoding and decoding point clouds hierarchically based on context modeling and sparse tensor operations to estimate occupancy probabilities.
Achieves efficient compression and processing of point clouds with reduced storage requirements and computational costs, enabling real-time handling of dynamic point clouds and improved performance on consumer devices.
Smart Images

Figure 2025536880000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical Field [1] The present embodiments generally relate to methods and apparatus for compressing and processing point clouds. [Background technology]
[0002] background [2] The point cloud (PC) data format is a versatile data format across multiple business domains, such as autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, computer graphics, and even the animation / film industry. 3D LiDAR (Light Detection and Ranging) sensors are being deployed in autonomous vehicles, and affordable LiDAR sensors have been released by Velodyne Velabit, Apple iPad Pro 2020, and Intel RealSense LiDAR Camera L515. Advances in sensing technology make 3D point cloud data more practical than ever before and are expected to be the ultimate enabler for the applications described herein. Summary of the Invention
[0003] overview [3] According to one embodiment, there is provided a method for encoding or decoding point cloud data, the method comprising: obtaining features associated with point cloud data corresponding to a point cloud, the point cloud data being represented in a sparse tensor format at a Level of Detail (LoD); processing the features associated with the LoD to match a resolution of another LoD subsequent to the LoD; for each occupied voxel of the LoD, encoding or decoding a plurality of child voxels in the other LoD based on the processed features to obtain occupancy information of the other LoD; and updating the processed features based on the occupancy information of the other LoD to generate updated features associated with the other LoD.
[0004] [4] According to another embodiment, there is provided an apparatus for encoding or decoding point cloud data, the apparatus including one or more processors and at least one memory coupled to the one or more processors, configured to obtain features associated with point cloud data corresponding to a point cloud, the point cloud data being represented in a sparse tensor format at a level of detail (LoD); process the features associated with the LoD to match a resolution of another LoD subsequent to the LoD; for each occupied voxel of the LoD, encode or decode a plurality of child voxels in the another LoD based on the processed features to obtain occupancy information of the another LoD; and update the processed features based on the occupancy information of the another LoD to generate updated features associated with the another LoD.
[0005] [5] One or more embodiments also provide a computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any of the embodiments described herein. One or more of the presented embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding point cloud data according to the methods described herein.
[0006] [6] One or more embodiments also provide a computer-readable storage medium having stored thereon point cloud data generated according to the methods described above. One or more embodiments also provide methods and apparatus for transmitting or receiving point cloud data generated according to the methods described herein. [Brief explanation of the drawings]
[0007] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] [7] Figure 7 shows a block diagram of a system in which aspects of the present embodiments may be implemented. [Figure 2A] [8] presents a point-based point cloud representation. [Figure 2B] [8] present a voxel-based point cloud representation. [Figure 2C][8] present a point cloud representation based on sparse voxels. [Figure 3] [9] shows the LoD structure of the point cloud. [Figure 4]
[10] Shows the voxel coding order / steps. [Figure 5]
[11] presents bitwise coding using context modeling. [Figure 6]
[12] presents bitwise decoding using context modeling. [Figure 7]
[13] An example of a point cloud used for encoding / decoding is shown below. [Figure 8]
[14] Upsampling reveals features inherited from previous levels. [Figure 9]
[15] illustrates the encoding of the first bit / voxel group according to one embodiment. [Figure 10]
[16] illustrates the encoding of the second bit / voxel group according to one embodiment. [Figure 11]
[17] illustrates the encoding of the third bit / voxel group according to one embodiment. [Figure 12]
[18] shows the following LoD feature aggregation according to one embodiment: [Figure 13]
[19] illustrates the decoding of the first bit / voxel group according to one embodiment. [Figure 14]
[20] illustrates the decoding of the second bit / voxel group according to one embodiment. [Figure 15]
[21] illustrates the decoding of the third bit / voxel group according to one embodiment. [Figure 16A]
[22] illustrates a representation of location as context information according to one embodiment. [Figure 16B]
[22] shows child voxel positions as context information, according to one embodiment. [Figure 17]
[23] present cascaded sparse convolutional layers for feature aggregation. [Figure 18]
[24] present a ResNet block for feature aggregation. [Figure 19]
[25] present an Inception ResNet block for feature aggregation. [Figure 20]
[26] present a transformer block for feature aggregation. [Figure 21]
[27] present the architecture of the self-attention block. [Figure 22]
[28] demonstrate the cascading of several feature aggregation modules. [Figure 23]
[29] shows the probability estimation module. [Figure 24]
[30] illustrates a stochastic training strategy, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Detailed Description
[31] Figure 1 illustrates a block diagram of an example system in which various aspects and embodiments can be implemented. System 100 may be implemented as a device including various elements described below and configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 100, singly or in combination, may be implemented on a single integrated circuit, multiple ICs, and / or separate elements. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or separate elements. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to perform one or more aspects described herein.
[0009]
[32] System 100 includes at least one processor 110 configured to execute loaded instructions, for example, to implement various aspects described herein. Processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., volatile and / or non-volatile memory devices). System 100 includes storage device 140, which may include non-volatile and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage device 140 may include, by way of non-limiting example, an internal storage device, an external storage device, and / or a network-accessible storage device.
[0010]
[33] System 100 includes encoder / decoder module 130, which may include its own processor and memory, configured to process data to provide, for example, encoded or decoded video. Encoder / decoder module 130 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include either or both encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software, as known to those skilled in the art.
[0011]
[34] Program code to be loaded into the processor 110 or the encoder / decoder 130 to perform various aspects described herein may be stored in the storage device 140 and then loaded into the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more various objects during the execution of the processes described herein. Such stored objects may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results resulting from the processing of equations, formulas, operations, and arithmetic logic.
[0012]
[35] In some embodiments, memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either processor 110 or encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as MPEG-2, MPEG-I, JPEG Pleno, HEVC, or VVC.
[0013]
[36] Inputs to the elements of system 100 may be provided via various input devices, as shown in block 105. Such input devices may include, but are not limited to, (i) an RF section for receiving RF signals, e.g., transmitted wirelessly by a broadcast station, (ii) a composite input, (iii) a USB input, and / or (iv) an HDMI input.
[0014]
[37] In various embodiments, the input devices of block 105 are associated with respective input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as signal selection or bandlimiting the signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which may be referred to (for example) as a channel in certain embodiments, (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) selecting a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a bandlimiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs these various functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0015]
[38] The USB and / or HDMI terminals may also include respective interface processors that connect system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within processor 110, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, in a separate interface IC or within processor 110, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data stream as desired for presentation to an output device.
[0016]
[39] The various elements of system 100 may be provided within a unified housing in which the various elements are interconnected and data may be transmitted between them using suitable connection structures 115, such as internal buses known in the art, including I2C buses, wiring, and printed circuit boards.
[0017]
[40] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented in a wired and / or wireless medium, for example.
[0018]
[41] In various embodiments, data is streamed to system 100 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received via communication channel 190 and communication interface 150, which are adapted for Wi-Fi communication. Communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 100 using a set-top box that delivers data via an HDMI connection in input block 105. Still other embodiments provide streamed data to system 100 using an RF connection in input block 105.
[0019]
[42] System 100 can provide output signals to various output devices, including display 165, speakers 175, and other peripherals 185. Other peripherals 185, in various exemplary embodiments, include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are communicated between system 100 and display 165, speakers 175, or other peripherals 185 using signaling such as AV, Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speakers 175 may be integrated with other elements of system 100 within an electronic device, such as a television, as a single unit. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0020]
[43] Display 165 and speakers 175 may alternatively be separate from one or more other elements, for example if the RF portion of input 105 is part of a separate set-top box. In various embodiments where display 165 and speakers 175 are external elements, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0021]
[44] Point cloud data may consume a large portion of network traffic, for example between cars connected via 5G networks, and in immersive communications (VR / AR). Understanding and communicating point clouds requires efficient representation formats. In particular, raw point cloud data needs to be properly organized and processed for world modeling and sensing purposes. Compression of raw point clouds is essential if the data needs to be stored and transmitted in relevant scenarios.
[0022]
[45] Furthermore, point clouds can represent sequential scans of the same scene containing multiple moving objects. These are called dynamic point clouds, as compared to static point clouds captured from a static scene or object. Dynamic point clouds are typically organized into multiple frames, with different frames captured at different times. Dynamic point clouds may require real-time or low-latency processing and compression.
[0023]
[46] The automotive industry and autonomous vehicles are areas where point clouds are used. Autonomous vehicles must be able to "explore" their environment and make good driving decisions based on the reality that immediately surrounds them. Typical sensors such as LiDAR generate (dynamic) point clouds that are used by perception engines. These point clouds are not intended to be seen by the naked eye, are typically sparse, not necessarily colored, are captured frequently, and are dynamic. They may have other attributes such as reflectivity provided by LiDAR, as these may indicate the material of the detected object and be useful for making decisions.
[0024]
[47] Virtual reality (VR) and immersive worlds are predicted by many to be the future of 2D flat video. Unlike standard TV, where the viewer can only see the virtual world in front of them, VR and immersive worlds immerse the viewer in their entire environment. There are several gradations of immersion depending on the viewer's degree of freedom within the environment. Point clouds are an excellent candidate format for delivering VR worlds. Point clouds used in VR can be static or dynamic and typically do not exceed an average size, e.g., a few million points at a time.
[0025]
[48] Point clouds may also be used for various purposes, such as cultural heritage / architecture, where 3D scanning of objects such as statues or buildings is performed to share the spatial configuration of the object without the need to transmit or visit the object. Point clouds may also be used to ensure the preservation of knowledge of an object, such as a temple, when it is at risk of being destroyed by an earthquake. Such point clouds are typically static, colored, and large.
[0026]
[49] Another use case is terrain analysis and mapping using three-dimensional representations, where maps are not limited to flat surfaces and may include reliefs. Google Maps is a good example of a three-dimensional map, but it uses meshes rather than point clouds. Nevertheless, because point clouds can be a suitable data format for three-dimensional maps, such point clouds are typically static, colored, and large.
[0027]
[50] World modeling and sensing via point clouds can be a useful technique to enable machines to gain knowledge about the three-dimensional world around them for the applications described herein.
[0028]
[51] 3D point cloud data are essentially discrete samples of the surface of an object or scene. To fully represent the real world with point samples requires a vast number of points. For example, a typical VR immersive scene contains millions of points, whereas a point cloud typically contains hundreds of millions of points. Therefore, processing such large point clouds is computationally expensive, especially for consumer devices with limited computing power, such as smartphones, tablets, and car navigation systems.
[0029]
[52] Efficient storage methods are required to perform processing or inference on point clouds. One solution for storing and processing an input point cloud at a realistic computational cost is to first downsample the point cloud so that the downsampled point cloud summarizes the geometry of the input point cloud while having far fewer points. The downsampled point cloud is then fed to a subsequent machine task for further utilization. However, further reductions in storage space can be achieved by converting the raw point cloud data (original or downsampled) into a bitstream via entropy coding techniques for lossless compression. Improved entropy models result in smaller bitstreams, thus improving compression efficiency. It is also possible to combine the entropy model with downstream tasks, allowing the entropy encoder to preserve task-specific information during compression.
[0030]
[53] In addition to lossless coding, many scenarios call for lossy coding, which significantly improves compression ratios while maintaining a certain level of distortion.
[0031]
[54] We propose a method for bitwise lossless compression of voxelized point cloud data based on sparse tensor processing and deep learning. In the following, we first consider voxel-based representations of point cloud data, since octree coding relies on voxel-based representations of point clouds. We then review several octree coding methods, focusing on bitwise octree coding methods.
[0032]
[55] Voxel-based representation
[56] In a voxel-based representation, 3D point coordinates, e.g., as shown in Figure 2A, are uniformly quantized by a quantization step size. Each point corresponds to an occupied voxel of size equal to the quantization step, e.g., as shown in Figure 2B. Typically, occupied voxels are assigned a "1" and empty voxels a "0," and these voxels are arranged in a 3D array for random access.
[0033]
[57] However, because the majority of voxels are empty, a naive voxel representation can be memory inefficient. To solve this problem, we introduce a sparse voxel representation, in which occupied voxels are arranged in a sparse tensor format. A sparse tensor only tracks the locations and features within its filled / occupied entries, allowing for efficient storage and processing if the majority of entries are empty. An example of a sparse voxel representation is shown in Figure 2C, where empty voxels (shown as dotted lines) consume no memory, and only occupied voxels (shown as solid lines) need to be stored. Note that Figure 2 and the remaining figures are shown in two dimensions primarily for simplicity.
[0034]
[58] Point clouds are represented as 3D voxels, so they can be processed and interpreted using 3D convolutional neural networks, inspired by the successful application of 2D convolutional neural networks to 2D images. In conventional 3D convolution, a 3D kernel is superimposed at every location specified in the stride step, regardless of whether the voxel is occupied or empty. To avoid the computation and memory consumption caused by empty voxels, sparse 3D convolutional layers can be applied if the point cloud voxels are represented by sparse tensors.
[0035]
[59] Point cloud compression via octree coding
[60] A voxelized point cloud can be represented via an octree decomposition tree. Initially, the root node covers the entire 3D space within a bounding box. The space is then divided equally in all directions, i.e., x, y, and z directions, resulting in 2 x 2 x 2 = 8 voxels. For each voxel, if there is at least one point, the voxel is marked as occupied and represented by "1", otherwise it is marked as empty and represented by "0". This step results in the first Level of Detail (LoD) of the input point cloud. The voxel division then continues, resulting in a second size of the input point cloud, 2 3 ×2 3 ×2 3 The LoD is obtained as follows: The voxel can be further divided until a pre-specified condition is met.
[0036]
[61] Byte-wise Octree Coding: A popular approach to octree coding is to encode each occupied voxel with an 8-bit value, i.e., one byte, which indicates the occupancy of each individual octet. Thus, the root voxel node is first encoded with an 8-bit value. Then, for each occupied voxel in the next level, an 8-bit occupancy symbol is encoded before moving on to the next level. This type of octree coding algorithm that encodes 8-bit occupancy symbols is called the byte-wise octree coding method.
[0037]
[62] Bit-wise Octree Coding: Another perspective on encoding the octree is to directly encode the binary occupancy bits of all voxels. At each LoD, we encode a set of occupancy bits that represent the voxels at the current LoD, and then encode the next LoD. This kind of approach is called the bit-wise Octree coding approach. Our method is proposed for this kind of approach.
[0038]
[63] Comparing these two encoding methods, we can see that in the byte-wise octree encoding method, the encoding of the current voxel is essentially the encoding of the occupancy symbols of its child voxels. In other words, in the bit-wise octree encoding method, the encoding of the current voxel is actually the encoding of its own binary occupancy bits. Below, we will consider the bit-wise octree encoding approach in more detail.
[0039]
[64] Bitwise Octree Coding
[65] In the paper "Neural Network Modeling of Probabilities for Coding the Octree Representation of Point Clouds" by Kaya, Emre Can et al., MMSP 2021 (hereafter referred to as "Kaya"), the authors predict the occupancy probability of the current voxel using the occupancy bits of neighboring voxels at the same LoD as context information. The prediction is performed through a neural network containing a simple multilayer perceptron (MLP) layer. After achieving the probability prediction, an adaptive arithmetic coder is applied to encode the occupancy bits.
[0040]
[66] In the SparsePCGC paper "Sparse Tensor-Based Multiscale Representation for Point Cloud Geometry Compression," arXiv preprint arXiv:2111, 10633, 2021 (hereafter "SparsePCGC"), the authors also use neural networks to predict occupancy probabilities. In contrast to Kaya, SparsePCGC uses sparse 3D convolutional layers to construct a neural network for probability estimation. However, SparsePCGC employs a complex multi-stage design, where each stage contains a dedicated convolutional layer designed to estimate the probability of a specific child voxel. Furthermore, in SparsePCGC, the probability estimation for a specific LoD does not take into account features from the previous LoD, which prevents it from fully utilizing the benefits of neural networks.
[0041]
[67] A jointly owned patent application (Attorney Docket No. 2021PF00298) also introduces bitwise octree coding. The application proposes to summarize the available contextual information of a voxel into a more concise / condensed representation that is more amenable to probability estimation. This summarization process can be either non-learning-based or learning-based. The method proposed in the application differs in two aspects. First, the proposed method uses a fully learning-based approach. Second, the proposed method considers the point cloud to be coded as a sparse tensor and uses sparse tensor operators to effectively and efficiently estimate occupancy probabilities.
[0042]
[68] Another jointly owned patent application (Attorney Docket No. 2022PF00245) proposes a learning-based bitwise octree coding. However, it proposes to estimate the occupancy probability of the current LoD on the voxel grid of the previous LoD. In other words, it estimates the occupancy probability of a higher resolution based on the lower resolution features. The main motivation for choosing such a design is the reduction of computational cost, which is different from the method proposed here and SparsePCGC.
[0043]
[69] As mentioned above, the proposed method targets a bitwise octree coding scheme. In the following, we first outline the bitwise octree coding system and then describe our proposal in detail.
[0044]
[70] Hierarchical coding structure
[71] aims to compress the octree hierarchically by directly encoding the binary occupancy of voxels. Given an input point cloud of bit depth N, we hierarchically encode and decode the input point cloud as shown in Figure 3.
[0045]
[72] The encoder first constructs the coarsest voxel representation at the first LoD, PC1, which is first encoded and transmitted as the first bitstream, BS1. Then, it constructs the next LoD, PC2. By comparing PC2 with PC1, it is determined that to encode PC2, only the hatched voxels need to be coded, since checking PC1 ensures that the white voxels are empty. Therefore, the hatched voxels of PC2 are coded and transmitted as the second bitstream, BS2. Next, it constructs an even finer LoD, PC3. Again, by comparing PC3 with PC2, it is determined that only the hatched voxels of PC3 need to be coded, resulting in the third bitstream, BS3. This procedure is repeated until the finest bitdepth, N, of the point cloud is reached.
[0046]
[73] Similarly, the decoder first reconstructs the coarsest LoD of the point cloud, i.e., PC1, by decoding the first bitstream BS1. By referencing the already decoded PC1, it is found that the hatched voxels of PC2 are included in the second bitstream BS2. Therefore, BS2 is decoded and the decoded bits are placed in the hatched voxels of PC2 for reconstruction. Similarly, the third bitstream BS3 is decoded and the decoded bits are assigned to the hatched voxels of PC3 for reconstruction. This procedure is repeated until the finest bit depth of the point cloud is reached.
[0047]
[74] Context-based bitwise octree coding
[75] To achieve high compression ratios when encoding a particular LoD, the bitwise octree coding algorithm relies on an arithmetic coder and an efficient mechanism to estimate the occupancy probability of every voxel to be encoded / decoded. As an example of how to perform LoD encoding / decoding using occupancy probability estimation and arithmetic coding, we compress the second LoD, i.e., PC2, as shown in Figure 3.
[0048]
[76] First, we divide PC2 compression into eight steps (in the three-dimensional case), each of which encodes a group of bits / voxels at a specific location relative to the parent voxel. This design is shown as a two-dimensional example in Figure 4. In this two-dimensional case, the first step encodes the group of bits / voxels labeled "1" because they are all in the upper-left corner of the parent voxel, then the second step encodes the group of bits / voxels labeled "2" because they are all in the upper-right corner of the parent voxel, and so on. Note that in this two-dimensional case, four steps are sufficient to encode all voxels, as shown in Figure 5, where each step covers a group of voxels. However, eight steps are required in three dimensions because each parent voxel is divided into eight child voxels in three dimensions, i.e., there are eight groups of voxels in three dimensions.
[0049]
[77] A block diagram of the actual encoding process of PC2 is shown in Figure 5. As mentioned above, the whole encoding process is divided into several steps. In the ith step (dotted block), the context modeling module (510) estimates the occupancy probability of the ith bit / voxel group. The estimation of the occupancy probability is based on the bits already coded and other known information about the voxel group, e.g., their coordinates. According to the estimated probabilities ([p1p2] in Figure 5), the arithmetic encoder (520) encodes the input bits associated with the current step to generate the sub-bitstream BS i The final coded bitstream BS is obtained by combining all the sub-bitstreams.
[0050]
[78] The decoding process reverses the encoding as shown in the block diagram of Figure 6. Again, the decoding is divided into several steps. In the ith step (dotted block), the context modeling module (610) estimates the occupancy probability of the ith voxel group. The estimation of the occupancy probability is based on the already decoded bits and other known information about the voxel group, such as their coordinates. According to the estimated probability ([p1p2] in Figure 6), the arithmetic decoder (620) outputs the occupancy bit associated with the current step. The final decoded bit is obtained by concatenating all the decoded bits obtained in each step.
[0051]
[79] Note that the probabilities output by the context modeling module during decoding are the same as the probabilities output during encoding, i.e., how the occupied bits are losslessly compressed. Furthermore, since the context model essentially models the entropy of the bitstream, the more accurate the occupancy probabilities, the smaller the output bitstream BS should be. Therefore, it is essential to have a good context modeling method in octree coding.
[0052]
[80] Given that voxels have already been coded / decoded, the proposed method aims to model occupancy probabilities based on voxel-wise features inherited from the previous LoD by applying operations on sparse tensors.
[0053]
[81] Encoder
[82] We will demonstrate our encoder step with a concrete example. As shown in Figure 7, the point cloud PC 1FThat is, suppose that the first LoD of the input point cloud has already been encoded, containing geometric features f1 and f2 for each occupied voxel. We aim to encode the second LoD of the input point cloud, i.e., PC2. These geometric features are abstract, high-level features generated by a deep neural network. These features describe the occupancy of voxels in the current LoD and its neighborhood. The geometric features f1 and f2 can be represented as vectors, e.g., vectors of length 64. As mentioned above, PC2 is 1F By checking the value of PC2, it is guaranteed that the remaining voxels are empty, so only the hatched voxels of PC2 need to be coded. 1F Both PC1 and PC2 are represented in sparse tensor form. Therefore, the voxels enclosed by the dotted line in Figure 7 do not consume memory / storage.
[0054]
[83] At the start of encoding the first LoD of the point cloud (PC1), the previous LoD, i.e., PC 0F Note that the points from are simply a single voxel representing the entire space, in which case the geometric features are set to constants, e.g., features that are all 1's.
[0055]
[84] The preparation step for encoding is to install a voxel upsampling module (810) on the PC as shown in Figure 8. 1F and the upsampled point cloud PC UP The upsampled point cloud PC UP has the same resolution / LoD as PC2, but its own features are 1F This upsampling step is a common preparation for the subsequent step of encoding each group of voxels in PC2.
[0056]
[85] Next, we begin the actual encoding process. We begin with the encoding of the third group of bits of PC2 (bits labeled 3 in Figure 4), which involves a more general encoding pipeline than the first or second group of bits of PC2. The complete encoding pipeline for encoding the third group of bits of PC2 is shown in Figure 11. In this case, the first and second groups of bits of PC2 have already been encoded. See PC2 in Figure 11, where four voxels, shown in bold, have already been encoded. We intend to use these already-encoded bits for context modeling.
[0057]
[86] First, the coordinate reading module (1110) identifies all empty voxels from PC2 that have already been coded. In this example, voxels at locations (0,3) and (2,1) are identified by the coordinate reading module. These voxels are then pruned from PC2 using the voxel pruning module (1120). UP Removed / pruned from the pruned point cloud PC' UP This step is done by using the PC UP The geometry of the
[0058]
[87] Second, a context construction module 1130 is used to construct the context point cloud PC CTX Build a PC CTX contains all the bit-wise / voxel-wise discriminative information for predicting occupancy probability. CTX is also expressed in sparse tensor form, and PC' UP In one embodiment, we construct binary context information for occupancy probability prediction. CTX For each occupied voxel in PC1, if the co-located voxel in PC2 has already been coded and occupied, then assign a "1" to that voxel, and if the co-located voxel in PC2 has not yet been coded, then assign a "0" to that voxel. CTXRecall that for a voxel that is occupied in PC1, the co-located voxel in PC2 cannot be both coded and unoccupied, since it has already been removed in a previous step by the voxel pruning module.
[0059]
[88] Third, the context point cloud PC CTX and the pruned point cloud PC' UP are connected (1140), the voxel-wise features and local context information are combined into a new point cloud PC UP Specifically, in this step, the connectivity module computes the PC CTX and PC' UP Connect the corresponding features in the PC UP Generate a PC UP The output of the feature aggregation module (1150) is then fed into a neural network module to further refine / refine the voxel-wise features. The output of the feature aggregation module is then also fed into a PC UP The probability estimation module (1160) is a neural network module that estimates the occupancy probability of voxels in the probability point cloud PC p The feature aggregation module mainly contains sparse convolutional layers, while the probability estimation module mainly contains multi-layer perceptron (MLP) layers.
[0060]
[89] Finally, the occupancy probability of the third group of voxels is serialized by a serialization module (1170). The serialization module serializes the estimated occupancy probability of the third group of voxels, i.e., p 13 and p 23 The probability point cloud PC p In the example of Figure 11, the array [p 13 p 23]. This serialization step is performed in preparation for the next arithmetic coding. Meanwhile, the ground truth occupancy status of the third voxel group is also serialized. In the example of Figure 11, array
[11] is obtained. The serialized occupancy probabilities and serialized ground truth occupancies are then fed to an arithmetic encoder (1180) to generate a sub-bitstream for the third voxel group BS3.
[0061]
[90] Note that the encoding process of the second voxel group, which has the same procedure as in Figure 11, is also shown in Figure 10. The encoding process of the first voxel group is a special (simplified) case of the encoding of the other voxel groups, as shown in Figure 9. Since there are no voxels / bits already encoded at the start of encoding, the PC UP There are no voxels that need to be pruned from PC UP is fed directly to the concatenation module instead of its pruned version. Furthermore, since there are no encoded voxels, the context construction module uses the PC CTX , and also directly outputs the . Note that in a simplified implementation, voxel groups other than the first voxel group may also follow the process shown in Figure 9, i.e., voxel pruning (1120) is not performed and / or all contexts are set to zero. In this case, information from previously coded bits is partially or not used. However, the geometric information of the previous LoD is still used in the PC UP Since the information from previously coded bits is propagated to the current LoD via the geometric features of the previous coded bits, the occupancy probability estimation module can still perform reasonably well even if the information from previously coded bits is not fully utilized.
[0062]
[91] Once the encoding of the current LoD is complete, the final step is to compute per-voxel features to prepare for encoding the next LoD, as shown in Figure 12. This step is similar to the other steps for encoding voxel groups, except that the process stops immediately after the feature aggregation module PC. 2F The output of the ,contains voxel-wise features of the next LoD.,This is the PC in Figure 8. 1F is upsampled in the same way.
[0063]
[92] Decoder
[93] The decoding process reverses the encoding process, with a number of decoding steps and operations being the same as the encoder.
[0064]
[94] As shown in Figure 7, the point cloud PC 1F Suppose we have already decoded the first LoD of a point cloud, i.e., the occupied voxels contain geometric features f1 and f2. We aim to decode the second LoD of the input point cloud PC2 based on the received bitstream BS. Note that BS contains several sub-bitstreams, BS1, BS2, ..., BS8 (or BS4 in this 2D example), each corresponding to one group of voxels / bits, as shown in Figure 4.
[0065]
[95] At the start of decoding the first LoD of point cloud (PC1), the points from the previous LoD, i.e., PC 0F Note that is just a single voxel representing the entire space. In this case, the shape feature of the point cloud is set to a constant. Note that this must be the same as the constant feature used on the encoder side.
[0066]
[96] As with the encoding, as shown in Figure 8, 1F Decoding also requires a preparatory step of applying a voxel upsampling module to the PC2 voxels. This upsampling step serves to prepare the PC2 voxels for the subsequent decoding steps.
[0067]
[97] Next, the actual decoding process begins. Similar to encoding, we begin with the decoding of the third group of bits in PC2 (the bits labeled "3" in Figure 4), but this involves a more general decoding pipeline. The complete decoding pipeline for decoding the third group of bits in PC2 is shown in Figure 15. In this case, the first and second groups of bits in PC2 have already been decoded. These decoded bits are then added to the intermediate decoding point cloud PC2 as shown in Figure 15. DEC2 It forms a PC for context modeling and decoding. DEC2 It is intended to be used.
[0068]
[98] To decode the third group of bits, their occupancy probabilities must be estimated. The steps to estimate these occupancy probabilities are the same as those in encoding (Fig. 11). First, the coordinate reading module (1510) reads all empty voxels already decoded from the PC DEC2 These voxels are then pruned into the PC by the voxel pruning module (1520). UP Removed / pruned from the pruned point cloud PC' UP is obtained.
[0069]
[99] Second, a context construction module 1530 is used to construct the context point cloud PC CTX In one embodiment, binary context information for occupancy probability prediction is constructed. CTX For each occupied voxel in DEC2 If the voxel located at the same location in the DEC2 If a co-located voxel in has not yet been decoded, it is assigned a "0".
[0070]
[0100] Third, the context point cloud PC CTX and the pruned point cloud PC' UPare concatenated (1540), which brings together the voxel-wise features and local context information into the point cloud PC UP PC” UP is then fed to a feature aggregation module (1550) and subsequently to a probability estimation module (1560) to generate a probability point cloud PC p is obtained.
[0071]
[0101] Finally, the occupancy probabilities for the third group of voxels are serialized (1570) and fed to an arithmetic decoder (1580). The arithmetic decoder receives the occupancy probabilities and the sub-bitstream BS3 as input and decodes the array, i.e., occupancy bits. These occupancy bits are then deserialized by a deserialization module (1590), which returns the bits to their associated voxels. From this step, a decoded point cloud, i.e., the PCs in FIG. 15, is obtained. DEC3 An updated version of
[0072]
[0102] The decoding process of the second voxel group, which has the same procedure as in Figure 15, is also shown in Figure 14. Similar to encoding, the decoding of the first voxel group is also a simplified case of the decoding of other voxel groups, as shown in Figure 13. Since there are no voxels / bits already decoded at the start of decoding, the PC UP There are no voxels that need to be pruned from PC UP is fed directly to the concatenation module instead of its pruned version. Furthermore, since there are no decoded voxels, the context construction module uses the PC CTX Similarly to the encoder side, in a simplified implementation, voxel groups other than the first voxel group may also follow the process shown in Figure 13, i.e., voxel pruning (1520) is not performed and / or all contexts are set to zero. In this case, information from previously decoded bits is partially or not used. However, the geometric information of the previous LoD is used in PCUP Since the information from previously decoded bits is propagated to the current LoD via the geometric features of the previous LoD, the occupancy probability estimation module can still perform reasonably well even if the information from previously decoded bits is not fully utilized.
[0073]
[0103] Similar to encoding, once the decoding of the current LoD is completed, the final step is to compute the voxel-wise features to decode the next LoD, as shown in FIG.
[0074]
[0104] Building Context
[0105] In addition to the binary context examples shown above, the context point cloud PC CTX may also contain other contextual information that is useful for probability estimation.
[0075]
[0106] In one embodiment, the context information is further extended with x, y, and z coordinates. Specifically, for an occupied voxel A located at (x, y, z), its feature is the vector f=[xyz].
[0076]
[0107] In another embodiment, normalized coordinates are used as context information. n If , the dimension is 2 n ×2 n ×2 n and therefore PC CTX The feature vector associated with voxel (x, y, z) in n y / 2 n z / 2 n ].
[0077]
[0108] Furthermore, rather than using Euclidean coordinates, one may use spherical coordinates, which are useful when processing LiDAR sweeps. To do this, apply the following equation:
number
[0078]
[0109] In another embodiment, the context information is also n In this case, PC' CTX The feature vector of is simply a scalar f=n.
[0079]
[0110] In one embodiment, the context information may be the position of a child voxel relative to its parent voxel. For example, "front" and "back," "left" and "right," and "top" and "bottom" may be represented as "0" and "1," respectively, as shown in FIG. 16A. Thus, child voxels located in front, to the right, and above a parent voxel may be represented by a PC CTX 16B, the child voxels located after, to the right, and above the parent voxel will have f=
[0010] as shown in Figure 16B. In another embodiment, the three additional bits representing position are further converted to decimal, e.g., f=
[0010] becomes the scalar f=2, and f=
[0110] becomes the scalar f=6.
[0080]
[0111] In one embodiment, the extended context information may be some or any combination and permutation of all of the above examples.
[0081]
[0112] Feature aggregation
[0113] The purpose of the feature aggregation module is to refine the input features to better serve occupancy probability estimation.
[0082]
[0114] In one embodiment, the feature aggregation module is simply a series of sparse 3D convolution layers, with each sparse 3D convolution followed by a ReLU activation function, as shown in Figure 17. Note that "CONVD" represents a sparse 3D convolution layer with D output channels.
[0083]
[0115] In another embodiment, the feature aggregation module employs a ResNet architecture as shown in Figure 18. In this example, the figure shows the architecture of a ResNet block that aggregates features with D channels. Compared to Figure 17, Figure 18 shows the residual connections from the input being additively connected to the output of the convolutional layer.
[0084]
[0116] In another embodiment, the feature aggregation module employs an Inception ResNet (IRN) architecture, as shown in Figure 19. In this example, the figure shows the architecture of an IRN block that aggregates features with D channels.
[0085]
[0117] In another embodiment, we employ a transformer architecture similar to the voxel transformer proposed in the paper "Voxel transformer for 3D object detection" by Mao, Jiageng et al., Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021. A diagram of the transformer block is shown in Figure 20, which includes a self-attention block with residual connections and an MLP block (including an MLP layer) with residual connections. A block diagram of the self-attention block is shown in Figure 21, and is described in more detail below.
[0086]
[0118] The current feature vector f associated with voxel location A A , and voxel position A i k neighboring features f associated with AiGiven, the self-attention block calculates all the neighboring features f Ai Based on the feature f A where Ai (0≦i≦k-1) are the k nearest neighbors of A in the input sparse tensor. First, we try to update point A i is obtained by k-nearest neighbor (kNN) search based on the coordinates of A. Then, the query embedding Q for A A is calculated by the following formula: Q A =MLP Q (f A ) Then, all the nearest neighbor key embeddings of A, K Ai and value embedding V Ai is calculated using the following formula: K Ai =MLPK(f Ai )+E Ai , V Ai =MLPV(f Ai )+E Ai , 0≦i≦k-1, Here is MLP Q (·), MLP K (·), and MLP V (·) are MLP layers for retrieving queries, keys, and values, respectively, and E Ai is voxel A and A i and is calculated by the following formula: E Ai =MLP P (P A -P Ai ) Here is MLP P (·) is the MLP layer for obtaining positional coding, and P A and P Ai are the three-dimensional coordinates of voxels A and A i The output feature of the self-attention block at position A is as follows:
number
[0087]
[0119] The transformer block similarly updates the features of all occupied positions in the sparse tensor and then outputs the updated sparse tensor. Q (·), MLP K (·), MLP V (·), and MLP P Note that (·) can contain only one fully connected layer, which corresponds to a linear projection.
[0088]
[0120] In one embodiment, as shown in Figure 22, several feature aggregation blocks (Figures 17, 18, 19, and 20) are cascaded together to further improve performance. The feature aggregation blocks may be of the same type, e.g., all Transformer blocks, in which case the parameters of their neural network layers may or may not be shared. The feature aggregation blocks may also be a mix of different types of feature aggregation blocks, e.g., a mix of IRN blocks and Transformer blocks.
[0089]
[0121] Probability Estimation
[0122] In one embodiment, the probability estimation module includes a series of multi-layer perceptron (MLP) layers. Assuming that the input point cloud contains vector features of length D1 that reside in its voxels, the MLP uses the channel dimensions (D1, D2, ..., D) for classification. k-1 , 1) has k layers. The result of the MLP is then fed into a softmax function which converts the output of the MLP into a probability value in the range 0 to 1.
[0090]
[0123] Enhanced Voxel Upsampling
[0124] In one embodiment, during the preparatory step in encoding / decoding (FIG. 8), an additional feature aggregation module is added immediately after the voxel upsampling module to aggregate and refine the initial features, which can further improve compression performance.
[0091]
[0125] Training Strategies
[0126] To obtain neural network parameters suitable for compression, an initial training step is required. To efficiently train the neural network module, we also propose a training strategy called the stochastic training strategy.
[0092]
[0127] A block diagram of the proposed probabilistic training strategy according to one embodiment is shown in Figure 24. In each training step, we first randomly select one LoD from the octree for training (2410). Specifically, given an input point cloud with a total bit depth of N, this step randomly generates LoDs from 1 to N. The selected LoDs may be drawn according to a predetermined probability distribution. In one embodiment, they are drawn according to a uniform distribution. In another embodiment, they are drawn according to a predetermined multinomial distribution.
[0093]
[0128] Then, we randomly select some voxel groups from the selected LoD (2420). Specifically, we first randomly select a number m in the range of 0 to 7 (note that there are 8 voxel groups in the LoD), and then randomly select m voxel groups from all 8 groups. We assume that these selected m voxel groups are known for context modeling, i.e., have already been encoded / decoded.
[0094]
[0129] Next, the occupancy probabilities of all remaining (8-m) voxel groups are calculated according to the probability estimation method shown in FIG. 14 (or FIG. 15) if m≠0, or in FIG. 13 if m=0 (2430).
[0095]
[0130] Finally, a loss function, e.g., binary cross-entropy loss between the estimated probability and the ground truth occupancy at the selected LoD of the octree, is calculated (2440). Note that binary cross-entropy loss is a typical loss function for binary classification that characterizes the discrepancy between the occupancy probability and the ground truth occupancy. Backpropagation is performed using the calculated binary cross-entropy loss to update the neural network parameters (2450). This training procedure is repeated until a predetermined condition is met, e.g., until a predetermined number of training steps is reached.
[0096]
[0131] Experimental results
[0132] We apply the proposed method to losslessly encode the Ford point cloud sequence. The Ford dataset is a test dataset recommended by the MPEG-PCC Common Test Criteria (CTC). The dataset contains 4500 LiDAR frames collected from a moving vehicle for autonomous driving applications. In this experiment, 1500 LiDAR frames were used to train the neural network, while the remaining 3000 LiDAR frames were reserved for testing.
[0097]
[0133] An embodiment including voxel pruning was experimented with during encoding and decoding (FIGS. 11 and 15), where the context information additionally includes normalized coordinates and the current bit depth / LoD.
[0098]
[0134] First, the MPEG-PCC octree (standardized by MPEG, non-learning method) achieves a result of 22.35 bpp (bits per point, the smaller the better), while the sparse PCGC achieves a result of 5.6 × 10 6 Using a neural network model with parameters, 20.36 bpp was obtained, whereas the proposed model in this application achieved 4.9 × 10 6A neural network model with parameters was used to obtain 20.21 bpp, thus achieving the best compression performance in the proposed method. Compared with sparse PCGC, a smaller neural network model can achieve better compression performance.
[0099]
[0135] Various numerical values are used in this application, and the specific values are for illustrative purposes only and the above-described aspects are not limited to these specific values.
[0100]
[0136] Various methods are described herein, each of which includes one or more steps or actions that achieve the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," and the like may be used in various embodiments to refer to elements, components, steps, operations, etc., e.g., "first decoding" and "second decoding." The use of such terms does not imply a modified order of operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but may occur before, during, or overlap with the second decoding, for example.
[0101]
[0137] The implementations and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if described in the context of only one type of implementation (e.g., described only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. A method may be implemented in, for example, an apparatus, such as a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that enable communication of information between end users.
[0102]
[0138] References to "one embodiment" or "one embodiment" or "one implementation mode" or "one implementation mode" and "embodiment" and other variations mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in one embodiment" or "in one implementation mode" or "in one implementation mode" appearing in various places throughout this application, and any other variations, are not necessarily all referring to the same embodiment.
[0103]
[0139] Additionally, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0104]
[0140] Additionally, the application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, replicating information, calculating information, determining information, predicting information, or inferring information.
[0105]
[0141] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended as a broad term. Receiving information may include, for example, one or more of accessing the information or retrieving the information (e.g., from a memory). Furthermore, "receiving" typically includes, in some manner, during operation, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0106]
[0142] It should be understood that the use of any of the following terms, " / ," "and / or," and "at least one of," is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B), for example, in the case of "A / B," "A and / or B," and "at least one of A and B." As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third alternatives (A and C), or the selection of only the second and third alternatives (B and C), or the selection of all three alternatives (A, B, and C). This may be extended to all items listed, as would be apparent to one of skill in the art.
[0107]
[0143] As will be apparent to those skilled in the art, multiple implementations may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for performing a method or data generated by one of the implementations described above. For example, a signal may be formatted to carry a bitstream of the above-described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
Claims
1. 1. A method for encoding or decoding point cloud data, comprising: obtaining point cloud data corresponding to a point cloud, the point cloud data being represented in a sparse tensor format at a level of detail (LoD); processing the features associated with the LoD to match the resolution of another LoD subsequent to the LoD; For each occupied voxel of the LoD, encoding or decoding a plurality of child voxels in the other LoD based on the processed features to obtain occupancy information of the other LoD; updating the processed features based on occupancy information of the other LoD to generate updated features associated with the other LoD.
2. 2. The method of claim 1, wherein the encoding or decoding of the plurality of child voxels at the different LoD comprises, for a current child voxel belonging to the plurality of child voxels at the different LoD, obtaining occupancy information of a child voxel that is encoded or decoded earlier among the plurality of child voxels in the different LoD; obtaining context information for encoding or decoding the current child voxel; generating an augmented feature by associating the context information with features of the current child voxel; aggregating another feature of the current child voxel based on the extended feature; generating an occupancy probability for the current child voxel based on the further feature; and encoding or decoding occupancy information of the current child voxel based on the occupancy probability of the current child voxel.
3. 3. The method of claim 1 or 2, The method further comprising pruning the processed features based on the occupancy information of the previously encoded or decoded child voxels, wherein the extended features are based on the pruned features.
4. The method according to any one of claims 1 to 3, wherein the features associated with the LoD are upsampled to match the resolution of the other LoD.
5. The method of claim 4 , wherein feature aggregation is performed on the upsampled features.
6. The method according to any one of claims 1 to 5, wherein the features associated with the LoD and the further LoD are geometric features.
7. The method according to any one of claims 1 to 6, The method further includes obtaining occupancy information and characteristics at a LoD subsequent to the another LoD based on the occupancy information and characteristics at the another LoD.
8. 8. The method of claim 2, wherein the context information indicates that a child voxel is either (1) coded / decoded and occupied, or (2) not coded / decoded.
9. The method of claim 8 , wherein the context information further indicates alignment information, a current LoD bit depth, and a position of a child voxel relative to a corresponding parent voxel.
10. The method of any one of claims 2 to 9, wherein the aggregating is based on at least one of a Transformer block with cascaded sparse convolutional layers, a ResNet, an Inception ResNet, and a self-attention block.
11. The method of claim 10 , wherein the aggregating is performed by cascading multiple feature aggregation modules.
12. The method according to any one of claims 2 to 11, wherein the occupancy probabilities are generated based on multiple multi-layer perceptron (MLP) layers.
13. 1. A method for training neural network parameters, comprising: Randomly selecting one LoD from the octree for training; randomly selecting a plurality of voxel groups from the selected LoD and assuming that the selected plurality of voxel groups are known for context modeling; calculating the occupancy probabilities of the remaining voxel groups; calculating a loss function between the estimated probability and the ground truth occupancy of the selected LoD of the octree; and performing backpropagation to update the neural network parameters.
14. 14. An apparatus comprising one or more processors and at least one memory coupled to said one or more processors, said one or more processors configured to perform a method according to any one of claims 1 to 13.
15. A signal containing video data formed by carrying out a method according to any one of claims 1 to 12.
16. A computer readable storage medium storing instructions for encoding or decoding a point cloud according to the method of any one of claims 1 to 12.