Bit-by-bit depth octree coding based on sparse tensor
Patent Information
- Application Number
- CN202380072245.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-03
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology is difficult to efficiently compress and process large-scale point cloud data, especially in the world of autonomous driving, virtual reality and immersive, resulting in increased computing costs and storage requirements.
The bit-by-bit encoding method based on sparse tensor processing and deep learning is used to perform lossless compression of point cloud data. This method constructs a multi-level point cloud representation, uses sparse 3D convolutional layers and neural networks to estimate the occupancy probability, and combines an arithmetic encoder to achieve efficient voxelized point cloud data compression.
It realizes efficient compression of point cloud data, significantly reduces storage space and transmission traffic, while maintaining data quality and accuracy, and is suitable for consumer devices with limited computing power.
Smart Images

Figure CN120035842A_ABST
Abstract
Description
Technical Field
[0001] The present embodiments generally relate to methods and apparatus for point cloud compression and processing. Background Art
[0002] Point cloud (PC) data format is a common data format across multiple business fields, such as autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, computer graphics, and animation / film industries. 3D LiDAR (Light Detection and Ranging) sensors have been deployed on autonomous vehicles, and affordable LiDAR sensors have been released by Velodyne Velabit, Apple iPadPro 2020, and Intel RealSense LiDAR Camera L515. With the advancement of sensing technology, 3D point cloud data has become more practical than ever and is expected to be the ultimate enabler of the applications discussed in this article. Summary of the invention
[0003] According to one embodiment, a method for encoding or decoding point cloud data is provided, comprising: obtaining features associated with point cloud data of a point cloud, the point cloud data being represented in a sparse tensor format at a level of detail (LoD); processing the features associated with the LoD to match the resolution of another LoD, wherein the other LoD is subsequent to the LoD; for each occupied voxel in the LoD, encoding or decoding a plurality of sub-voxels at the other LoD based on the processed features to obtain occupancy information at the other LoD; and updating the processed features based on the occupancy information at the other LoD to generate updated features associated with the other LoD.
[0004] According to another embodiment, a device for encoding or decoding point cloud data is provided, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: obtain features associated with point cloud data of a point cloud, the point cloud data being represented in a sparse tensor format at a level of detail (LoD); process the features associated with the LoD to match the resolution of another LoD, wherein the other LoD is after the LoD; for each occupied voxel in the LoD, encode or decode a plurality of sub-voxels at the other LoD based on the processed features to obtain occupancy information at the other LoD; and update the processed features based on the occupancy information at the other LoD to generate updated features associated with the other LoD.
[0005] One or more embodiments also provide a computer program including instructions, which, when executed by one or more processors, cause the one or more processors to perform an encoding method or a decoding method according to any embodiment described herein. One or more of the embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding point cloud data according to the method described herein.
[0006] One or more embodiments further provide a computer-readable storage medium on which point cloud data generated according to the above method is stored. One or more embodiments further provide a method and apparatus for transmitting or receiving point cloud data generated according to the method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 A block diagram of a system is shown in which aspects of the present embodiments may be implemented.
[0008] Figure 2A , Figure 2B and Figure 2C Point-based, voxel-based, and sparse voxel-based point cloud representations are shown, respectively.
[0009] Figure 3 The LoD construction of the point cloud is shown.
[0010] Figure 4 The encoding order / steps of the voxels are shown.
[0011] Figure 5 Bit-by-bit encoding with context modeling is shown.
[0012] Figure 6 Bit-by-bit decoding with context modeling is shown.
[0013] Figure 7 An example of a point cloud for encoding / decoding is shown.
[0014] Figure 8 Shows the features inherited from previous layers through upsampling.
[0015] Fig. 9 An encoding of a first set of bits / voxel is shown according to an embodiment.
[0016] Fig.10 An encoding of a second set of bits / voxel is shown according to an embodiment.
[0017] Fig.11 An encoding of a third set of bits / voxel is shown according to an embodiment.
[0018] Fig.12 Feature aggregation for the next LoD according to an embodiment is shown.
[0019] Fig.13 Decoding of a first set of bits / voxels is shown according to an embodiment.
[0020] Fig.14 Decoding of a second set of bits / voxels is shown according to an embodiment.
[0021] Fig.15 Decoding of a third set of bits / voxels is shown according to an embodiment.
[0022] Fig.16A and Fig. 16B The position representation and the subvoxel position as context information according to an embodiment are respectively shown.
[0023] Fig.17 Cascaded sparse convolutional layers for feature aggregation are shown.
[0024] Fig.18 A ResNet block for feature aggregation is shown.
[0025] Fig.19 An Inception-ResNet block for feature aggregation is shown.
[0026] Fig. 20 A transformer block for feature aggregation is shown.
[0027] Fig.21 The architecture of the self-attention block is shown.
[0028] Fig. 22 A cascade of several feature aggregation modules is shown.
[0029] Fig.23 A probability estimation module is shown.
[0030] Fig.24 A probabilistic training strategy according to an embodiment is shown. DETAILED DESCRIPTION
[0031] Figure 1A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. System 100 can be embodied as a device including various components described below, and is configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances and servers. The elements of system 100 can be embodied in a single integrated circuit, multiple ICs and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed over multiple ICs and / or discrete components. In various embodiments, system 100 is coupled to other systems or other electronic devices via, for example, a communication bus or by a dedicated input and / or output port communication. In various embodiments, system 100 is configured to implement one or more aspects described in this application.
[0032] The system 100 includes at least one processor 110, which is configured to execute instructions loaded therein, for implementing various aspects described in the present application, for example. The processor 110 may include embedded memory, input and output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include a non-volatile memory and / or a volatile memory, including but not limited to an EEPROM, a ROM, a PROM, a RAM, a DRAM, an SRAM, a flash memory, a magnetic disk drive, and / or an optical disk drive. As a non-limiting example, the storage device 140 may include an internal storage device, an additional storage device, and / or a network accessible storage device.
[0033] The system 100 includes an encoder / decoder module 130, which is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 130 may be implemented as a separate element of the system 100, or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.
[0034] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0035] In several embodiments, memory internal to the processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be a memory 120 and / or a storage device 140, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as for MPEG-2, MPEG-1, JPEG Pleno, HEVC, or VVC.
[0036] As indicated at block 105, input may be provided to the elements of system 100 via various input devices. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, such as by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0037] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the desired frequency band again by filtering, down-conversion and perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements that perform similar or different functions. Adding element can include inserting element between existing elements, for example, inserting amplifier and analog-to-digital converter. In various embodiments, the RF part includes antenna.
[0038] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within the processor 110 as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 110 as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and the encoder / decoder 130, which operate in combination with memory and storage elements to process the data streams as desired for presentation on an output device.
[0039] The various elements of the system 100 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and data may be transferred between them using suitable connection means 115, such as an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.
[0040] The system 100 includes a communication interface 150 that enables communication with other devices through a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented in, for example, a wired and / or wireless medium.
[0041] In various embodiments, a Wi-Fi network such as IEEE 802.11 is used to stream data to the system 100. The Wi-Fi signals of these embodiments are received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to the system 100, which passes data through the HDMI connection of the input block 105. Still other embodiments use the RF connection of the input block 105 to provide streaming data to the system 100.
[0042] The system 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripherals 185. In various examples of embodiments, the other peripherals 185 include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system 100. In various embodiments, control signals are transmitted between the system 100 and the display 165, the speaker 175, or other peripherals 185 using signaling such as AV Link, CEC, or other communication protocols, which can achieve device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to the system 100 using a communication channel 190 through the communication interface 150. The display 165 and the speaker 175 can be integrated into a single unit with other components of the system 100 in an electronic device (e.g., a television). In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0043] For example, if the RF portion of input 105 is part of a separate set-top box, the display 165 and speaker 175 may alternatively be separate from one or more other components. In various embodiments where the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.
[0044] It is conceivable that point cloud data may consume a large portion of network traffic, for example, in connected cars on 5G networks and in immersive communications (VR / AR). Efficient representation formats are necessary for point cloud understanding and communication. In particular, raw point cloud data needs to be properly organized and processed for the purpose of world modeling and sensing. Compression of raw point clouds is essential when data needs to be stored and transmitted in relevant scenarios.
[0045] In addition, point clouds can represent continuous scans of the same scene containing multiple moving objects. Compared with static point clouds captured from static scenes or static objects, they are called dynamic point clouds. Dynamic point clouds are usually organized into frames, with different frames captured at different times. Dynamic point clouds may require real-time or low-latency processing and compression.
[0046] The automotive industry and autonomous vehicles are areas where point clouds can be used. Autonomous vehicles should be able to "probe" their environment to make good driving decisions based on the reality around them in real time. Typical sensors like LiDAR produce (dynamic) point clouds that are used by perception engines. These point clouds are not intended to be viewed by the human eye, and they are usually sparse, not necessarily in color, and dynamic with a high capture frequency. They can have other properties like reflectivity provided by LiDAR, as this property indicates the material of the object being sensed and can help in making decisions.
[0047] Virtual Reality (VR) and immersive worlds are foreseen by many as the future of 2D flat-panel video. With VR and immersive worlds, the viewer is immersed in the environment around the viewer, as opposed to standard TV where the viewer can only see the virtual world in front of the viewer. There are several levels of immersion, depending on the degree of freedom the viewer has in the environment. Point clouds are a good candidate format for distributing VR worlds. Point clouds used in VR can be static or dynamic and are typically of average size, e.g., no more than millions of points at a time.
[0048] Point clouds can also be used for various purposes such as cultural heritage / architecture, where objects like statues or buildings are 3D scanned in order to share the spatial configuration of the object without sending or accessing the object. In addition, point clouds can also be used to ensure that knowledge of the object is preserved in cases where the object may be destroyed, for example, earthquakes causing damage to temples. Such point clouds are usually static, colorful, and huge.
[0049] Another use case is in topography and cartography where 3D representations are used, and maps are not limited to flat surfaces and can include undulations. Google Maps is a good example of a 3D map, but it uses a mesh instead of a point cloud. However, a point cloud may be a suitable data format for a 3D map, and such point clouds are often static, colorful, and huge.
[0050] World modeling and sensing via point clouds can be a useful technique that allows machines to gain knowledge about the 3D world around them for the applications discussed here.
[0051] 3D point cloud data is essentially discrete samples on the surface of an object or scene. In order to fully represent the real world with point samples, a large number of points are actually required. For example, a typical VR immersive scene contains millions of points, and point clouds usually contain hundreds of millions of points. Therefore, the processing of such large-scale point clouds is computationally expensive, especially for consumer devices with limited computing power, such as smartphones, tablets, and car navigation systems.
[0052] In order to perform processing or reasoning on point clouds, efficient storage methods are required. In order to store and process the input point cloud at an affordable computational cost, one solution is to first downsample the point cloud, where the downsampled point cloud generalizes the geometric structure of the input point cloud while having far fewer points. The downsampled point cloud is then fed to subsequent machine tasks for further use. However, further reductions in storage space can be achieved by converting the raw point cloud data (original or downsampled) into a bitstream through entropy coding techniques for lossless compression. Better entropy models produce smaller bitstreams and are therefore more efficient in compression. In addition, entropy models can also be paired with downstream tasks, which allows the entropy encoder to maintain task-specific information while compressing.
[0053] In addition to lossless encoding, many scenarios seek lossy encoding to significantly improve the compression ratio while keeping the introduced distortion at a certain quality level.
[0054] We propose a lossless compression method for voxelized point cloud data in a bit-by-bit manner based on sparse tensor processing and deep learning. In the following, we first review the voxel-based representation of point cloud data, since octree encoding relies on the voxel-based representation of point clouds. We then review some octree encoding methods, focusing on bit-by-bit octree encoding methods.
[0055] Voxel-based representation
[0056] In voxel-based representations, for example, Figure 2A As shown in , the 3D point coordinates are uniformly quantized by the quantization step size. Figure 2BAs shown, each point corresponds to an occupied voxel of size equal to the quantization step size. Typically, "1" will be assigned to an occupied voxel, while "0" will be assigned to an empty voxel, and the voxels are arranged into a 3D array for random access.
[0057] However, since most voxels are empty, a simple voxel representation may not be efficient in memory usage. To address this issue, a sparse voxel representation is introduced, where the occupied voxels are arranged in a sparse tensor format. A sparse tensor only keeps track of the locations and features in its filled / occupied entries, enabling efficient storage and processing when most entries are empty. Figure 2C An example of a sparse voxel representation is depicted in , where empty voxels (indicated by dashed lines) consume no memory, and only occupied voxels (indicated by solid lines) need to be stored. Note that FIG. 2 and the remaining figures are shown in 2D only for simplicity.
[0058] Point clouds are already represented as 3D voxels, which can be processed / digested with 3D convolutional neural networks - inspired by the successful application of 2D convolutional neural networks to 2D images. For regular 3D convolutions, the 3D kernel is overlaid on every location specified by the stride stride, regardless of whether the voxel is occupied or empty. To avoid the computation and memory consumption caused by empty voxels, sparse 3D convolution layers can be applied if the point cloud voxels are represented by sparse tensors.
[0059] Point cloud compression via octree encoding
[0060] The voxelized point cloud can be represented via an octree decomposition tree. First, the root node covers the entire 3D space in the bounding box. Then, the space is divided equally along each direction, namely the x, y and z directions, resulting in 2×2×2=8 voxels. For each voxel, if there is at least one point, the voxel is marked as occupied and represented by "1"; otherwise, it is marked as empty and represented by "0". This step results in the first level of detail (LoD) of the input point cloud. Voxel segmentation is then continued to obtain the second LoD of the input point cloud, which is of size 2 3 ×2 3 ×2 3 . Voxels can be further segmented until a pre-specified condition is met.
[0061] Byte-by-byte octree encoding: A popular method to encode an octree is to encode each occupied voxel with an 8-bit value, i.e. 1 byte. It indicates the occupancy of its respective octants. In this way, we first encode the root voxel node with an 8-bit value. Then, for each occupied voxel in the next level, we encode its 8-bit occupancy sign and move to the next level. We call this octree encoding algorithm that encodes 8-bit occupancy sign as byte-by-byte octree encoding method.
[0062] Bit-by-bit octree encoding: An alternative view of encoding the octree is by directly encoding the binary occupancy bits of each voxel. At each LoD, we encode a sequence of occupancy bits representing the voxel at the current LoD, and then we encode the next LoD. We call this approach the bit-by-bit octree encoding method. Our method is proposed for this type of method.
[0063] Comparing these two types of encoding methods, we see that in the byte-by-byte octree encoding method, the encoding of the current voxel is essentially encoding the occupancy sign of its child voxels. The difference is that in the bit-by-bit octree encoding method, the encoding of the current voxel is actually encoding its own binary occupancy bits. In the following, we review the bit-by-bit octree encoding method in detail.
[0064] Bit-by-bit octree encoding
[0065] In the MMSP 2021 paper titled “Neural Network Modeling of Probabilitiesfor Coding the Octree Representation of Point Clouds” by Kaya, Emre Can et al. (hereinafter referred to as “Kaya”), the authors use the occupancy bits of neighboring voxels in the same LoD as context information to predict the occupancy probability of the current voxel. The prediction is performed by a neural network containing a simple multi-layer perceptron (MLP) layer. After the probability prediction is completed, an adaptive arithmetic encoder is applied to encode the occupancy bits.
[0066] In the article SparsePCGC (arXiv preprint arXiv:2111.10633, 2021) (hereinafter referred to as "SparsePCGC") by Wang, Jianqiang et al., titled "Sparse Tensor-based Multiscale Representation for Point Cloud Geometry Compression", the authors also used neural networks to predict occupancy probabilities. In contrast to Kaya, SparsePCGC uses sparse 3D convolutional layers to build neural networks for probability estimation. However, the design of SparsePCGC adopts a complex multi-level design, in which each level involves dedicated convolutional layers that are designed to estimate the probability of a specific sub-voxel. In addition, in SparsePCGC, the probability estimation of a specific LoD does not take into account the features from the previous LoD, which cannot fully utilize the advantages of neural networks.
[0067] A jointly owned patent application (attorney docket number 2021PF00298) also introduces a method for bit-by-bit octree encoding. It proposes to summarize the available contextual information of a voxel into a more concise / condensed representation that is more friendly to probability estimation. This summarization process can be non-learning-based or learning-based. The method proposed here differs in two aspects. First, the proposed method uses a completely learning-based approach. Second, the proposed method treats the point cloud to be encoded as a sparse tensor and utilizes sparse tensor operators to effectively and efficiently estimate occupancy probabilities.
[0068] Another commonly owned patent application (Attorney Docket No. 2022PF00245) proposes a learning-based bit-by-bit octree encoding. However, it proposes to estimate the occupancy probability of the current LoD in the voxel grid of the previous LoD. In other words, it estimates the occupancy probability of a higher resolution based on the features of the lower resolution. The main motivation for this design choice is to reduce computational cost, and it is different from the method proposed here, and also different from SparsePCGC.
[0069] As mentioned above, the proposed method targets a bit-by-bit octree coding scheme. In the following, we first provide a systematic overview on bit-by-bit octree coding and then elaborate on our proposal.
[0070] Hierarchical coding structure
[0071] We propose to compress the octree hierarchically by directly encoding the binary occupancy state of voxels. Given an input point cloud of bit depth N, we hierarchically encode and decode the input point cloud, as Figure 3 shown.
[0072] On the encoder side, we first construct its coarsest voxel representation PC at the first LoD 1 , and PC 1 It is first encoded and used as the first bit stream BS 1 Send. Then build the next LoD, PC 2 By comparing PC 2 and PC 1 We know that we need to 2 To encode, only its shadow voxels need to be encoded, because by checking PC 1 It is guaranteed that white voxels are empty. Therefore, PC 2 The shadow voxels are encoded and sent as the second bitstream BS 2 Next, we build even finer LoDs, PC 3 Again, by comparing PC 3 and PC 2 , we only encode PC3 The shadow voxels in the , thus generating the third bit stream BS 3 The process is repeated until the finest bit depth N of the point cloud is reached.
[0073] Similarly, on the decoder side, we first decode the first bitstream BS by 1 To rebuild the point cloud PC 1 The coarsest LoD. By referring to the decoded PC 1 , we know that PC 2 The shadow voxels in the second bitstream BS are included 2 Therefore, we have BS 2 decodes the bits and puts them into the PC 2 Similarly, the third bitstream BS 3 is also decoded, and the decoded bits are assigned to the PC 3 The shadow voxels of are used for reconstruction. This process is repeated until the finest bit depth of the point cloud is reached.
[0074] Context-based bit-by-bit octree coding
[0075] In order to achieve high compression ratios when encoding a certain LoD, the bit-by-bit octree coding algorithm relies on an arithmetic encoder and an efficient mechanism to estimate the occupancy probability of each voxel to be encoded / decoded. Figure 3 Second LoD shown, PC 2 The compression of is taken as an example to illustrate how to use occupancy probability estimation and arithmetic coding to perform LoD encoding / decoding.
[0076] First, PC 2 The compression of is divided into 8 steps (for 3D), where each step encodes a set of bits / voxel at a specific position relative to its parent voxel. We Figure 4 This design is illustrated in 2D as an example. In this 2D case, we encode the bit / voxel group marked as ‘1’ in the first step because they are all located at the top left corner of their parent voxel, and then encode the bit / voxel group marked as ‘2’ in the second step because they are all located at the top right corner of their parent voxel, and so on. Note that in this 2D case, only four steps are sufficient to encode all voxels, as Figure 5 As shown, each step covers a group of voxels. However, in 3D, eight steps are required because in 3D each parent voxel will be divided into eight child voxels, i.e., there are eight groups of voxels in 3D.
[0077] Figure 5 Shows PC 2As mentioned above, the whole encoding process is divided into several steps. In the i-th step (dashed box), the context modeling module (510) estimates the occupancy probability of the i-th group of bits / voxels. The estimation of the occupancy probability is based on the bits that have been encoded, as well as other known information about the voxel group, such as their coordinates. Using the estimated probability ( Figure 5 [p 1 p 2 ]), the arithmetic encoder (520) encodes the input bits associated with the current step and outputs a sub-bitstream BS i The final coded bitstream BS is obtained by combining all sub-bitstreams.
[0078] The decoding process reverses the encoding, such as Figure 6 As shown in the block diagram of . Again, decoding is divided into several steps. In the i-th step (dashed box), the context modeling module (610) estimates the occupancy probability of the i-th group of voxels. The estimation of the occupancy probability is based on the decoded bits and other known information about the voxel group, such as their coordinates. Using the estimated probability ( Figure 6 [p 1 p 2 ]), the arithmetic decoder (620) outputs the occupied bits associated with the current step. The final decoded bits are obtained by concatenating all the decoded bits obtained in each step.
[0079] We note that the probabilities output by the context modeling module during decoding are the same as those output during encoding, which is how occupied bits can be losslessly compressed. Moreover, the context model essentially models the entropy of the bitstream, and the more accurate the occupancy probabilities are, the smaller the output bitstream BS will be. Therefore, having a good context modeling method is crucial in octree coding.
[0080] Given already encoded / decoded voxels, our proposed method aims to model occupancy probabilities based on the voxel-by-voxel features inherited from the previous LoD by applying operations on the sparse tensor.
[0081] Encoder
[0082] We use a concrete example to illustrate our encoder steps. Figure 7 As shown, assuming that we have encoded the point cloud PC 1F The first LoD of the input point cloud is equipped with a geometric feature f in each occupied voxel 1 and f 2 We aim to encode the second LoD of the input point cloud, PC 2 These geometric features are abstract high-level features generated by deep neural networks. They describe the occupancy status of nearby voxels at the current LoD.1 and f 2 can be represented as vectors, for example, they can be vectors of length 64. As discussed above, only PC 2 The shadow voxels in need to be encoded because by checking the PC 1F It ensures that the rest of the voxels are empty. In addition, we represent PC in sparse tensor format 1F and PC 2 Both. Therefore, Figure 7 Voxels enclosed by dashed lines in FIG. 5 consume no memory / storage.
[0083] Note that when you start encoding the first LoD (PC 1 ), the point cloud PC from its previous LoD 0F is just a voxel representing the entire space. In this case, its geometric features are set to constants, e.g., features that are all 1.
[0084] The preparatory step for encoding is to apply the voxel upsampling module (810) to the PC 1F , generating the upsampled point cloud PC UP ,like Figure 8 As shown. Upsampled point cloud PC UP With PC 2 Same resolution / LoD, but features directly from PC 1F Obtained / inherited. This upsampling step is used as a PC 2 Each group of voxels in the encoding is a common preparation for the subsequent steps.
[0085] Next, we begin the actual encoding process. We start from the PC 2 The third group of bits in Figure 4 We begin our description of the encoding of the bit marked 3 in the PC 2 The first or second group of bits in a more general encoding pipeline. Fig.11 Shows the PC 2 The third group of bits in the entire encoding pipeline are encoded. In this case, PC 2 The first and second groups of bits in have been encoded. Fig.11 PC in 2 , where the four voxels in bold have been encoded. And we intend to use these encoded bits for context modeling.
[0086] First, the coordinate reader module (1110) locates the PC 2 All the encoded voxels in are empty. In this example, the coordinate reader module locates the voxels at positions (0, 3) and (2, 1). Thereafter, the voxel pruning module (1120) is used to extract the voxels from the PCUP Remove / trim these voxels to obtain the pruned point cloud PC' UP , which is used to refine the PC based on the bits that have been encoded UP The geometric structure of .
[0087] Second, we use the context building module (1130) to build the context point cloud PC CTX . PC CTX Includes all bit-by-bit / voxel-by-voxel discriminative information used to predict occupancy probability. Contextual Point Cloud PC CTX is also represented in sparse tensor format, and it is consistent with PC' UP share the same voxel geometry. In one embodiment, we construct binary context information for occupancy probability prediction. CTX For each occupied voxel in PC 2 The co-located voxel in PC is already coded and occupied, we place a "1" in this voxel; if it is in 2 The co-located voxel in has not been encoded yet, so we place a “0” in this voxel. Note that for PC CTX The occupied voxels in PC 2 Co-located voxels in cannot be both encoded and unoccupied, because such voxels have been removed by the voxel pruning module in the previous step.
[0088] Third, the context point cloud PC CTX And the trimmed point cloud PC' UP Cascade (1140), which puts together the voxel-by-voxel features and local context information and produces a new point cloud PC" UP Specifically, in this step, the cascade module cascades PCs for each occupied voxel CTX and PC' UP The corresponding features in to generate PC" UP Then, connect the PC UP The output of the feature aggregation module is then fed to the probability estimation module (1160), which is also a neural network module, to estimate the PC” UP The occupancy probability of the voxels in . It outputs the probability point cloud PC p , where each voxel contains its own estimated occupancy probability. The feature aggregation module mainly consists of sparse convolutional layers, while the probability estimation module mainly consists of multi-layer perceptron (MLP) layers.
[0089] Finally, the serialization module (1170) serializes the occupancy probability of the third voxel group. The serialization module is used to obtain the occupancy probability of the third voxel group from the probability point cloud PC.p The estimated occupancy probability of the third voxel group (i.e., p 13 and p 23 ) and put it into a one-dimensional array. Fig.11 In the example, it outputs the array [p 13 p 23 ]. This serialization step is to prepare for the arithmetic coding that will be started next. On the other hand, the reference truth occupancy state of the third voxel group is also serialized. Fig.11 In the example of , it results in the array [1 1]. The serialized occupancy probabilities and the serialized ground truth occupancy are then fed to an arithmetic encoder (1180) to generate a sub-bitstream BS for the third voxel group 3 .
[0090] Note that the encoding process of the second voxel group is also Fig.10 It is shown in Fig.11 The same process. Fig. 9 As shown, the encoding process of the first voxel group is a special (and simplified) case of the encoding of the other voxel groups. Because at the beginning of the encoding, there are no voxels / bits that have been encoded, so there is no need to UP Therefore, PC UP is directly fed to the cascade module instead of its pruned version. In addition, since there is no encoded voxel, the context building module also directly outputs the PC with all zeros on its voxel CTX Note that in a simplified implementation, the voxel groups other than the first voxel group can also follow Fig. 9 The process shown, i.e., no voxel pruning (1120) is performed and / or all contexts are set to zero. In this case, information from previously encoded bits is partially used or not used. However, since the geometry information of the previous LoD is passed through the PC UP The geometric features in are propagated to the current LoD, so even if the information from the previously encoded bits is not fully utilized, the occupancy probability estimation module can still perform reasonably well.
[0091] After completing the encoding of the current LoD, the final step is to calculate the voxel-by-voxel features to prepare for the encoding of the next LoD, such as Fig.12 This step is similar to the other steps of encoding voxel groups, except that processing stops immediately after the feature aggregation module. The output PC of the feature aggregation module is 2F Contains the voxel-by-voxel features of the next LoD. It will be Figure 8 PC in 1F The same effect is achieved with upsampling.
[0092] Decoder
[0093] The decoding process reverses the encoding process, where quite a few decoding steps and operations are the same as the encoder.
[0094] like Figure 7 As shown, assuming that we have decoded the point cloud PC 1F ——The first LoD of the point cloud, equipped with geometric features f1 and f in its occupied voxels 2 We aim to decode the second LoD of the input point cloud based on the received bitstream BS, PC 2 Note that BS consists of several sub-bitstreams BS 1 ,BS 2 ,…,BS 8 (or BS in this 2-D example 4 ), where each sub-bitstream corresponds to a set of voxels / bits, such as Figure 4 shown.
[0095] Note that when you start decoding the first LoD (PC 1 ), the point cloud PC from its previous LoD 0F is just one voxel representing the entire space. In this case, its geometric feature is set to a constant. Note that it must be the same constant feature used on the encoder side.
[0096] Similar to encoding, the voxel upsampling module also needs to be applied to the PC in decoding 1F The preparation steps are as follows: Figure 8 This upsampling step is used as the PC 2 Each group of voxels in the decoder is a common preparation for the subsequent steps.
[0097] Next, we start the actual decoding process. Similar to encoding, we start from the PC 2 The third group of bits in Figure 4 We begin our description of the decoding of the byte marked '3' in ), which involves a more general decoding pipeline. Fig.15 shows the decoded PC 2 The entire decoding pipeline for the third group of bits in . In this case, PC 2 The first and second groups of bits in have been decoded. These decoded bits form the intermediate decoded point cloud PC DEC2 ,like Fig.15 As shown, and we intend to use PC DEC2 Perform context modeling and decoding.
[0098] To decode the third group of bits, we need to estimate their occupancy probabilities. The steps to estimate these occupancy probabilities are similar to the steps during encoding ( Fig.11 ) is the same. First, the coordinate reader module (1510) locates the coordinates from the PCDEC2 All decoded voxels are empty. Thereafter, the voxel pruning module (1520) is used to extract the decoded voxels from the PC UP Remove / trim these voxels to obtain the pruned point cloud PC' UP .
[0099] Second, we use the context building module (1530) to build the context point cloud PC CTX In one embodiment, we construct binary context information for occupancy probability prediction. CTX For each occupied voxel in PC DEC2 The co-located voxel in PC has been decoded and occupied, we place a "1" in this voxel; if it is in DEC2 The co-located voxel in has not been decoded yet, and we place a “0” in this voxel.
[0100] Third, contextual point cloud PC CTX and the trimmed point cloud PC' UP are concatenated (1540), which puts together the voxel-by-voxel features and local context information and produces a point cloud PC" UP Then, connect the PC UP Feed to the feature aggregation module (1550), followed by the probability estimation module (1560), to obtain the probability point cloud PC containing the estimated occupancy probability p .
[0101] Finally, the occupancy probabilities of the third voxel group are serialized (1570) and fed to the arithmetic decoder (1580). The arithmetic decoder converts the occupancy probabilities and the sub-bitstream BS 3 as input and decodes the array - the occupied bits. These occupied bits are then deserialized (1590) by the deserialization module, which puts the bits back into their associated voxels. This step results in an updated version of the decoded point cloud, Fig.15 PC in DEC3 .
[0102] The decoding process of the second voxel group is also Fig.14 It is shown in Fig.15 The same process. Similar to encoding, the decoding of the first voxel group is also a simplified case of the decoding of other voxel groups, such as Fig.13 Because there are no decoded voxels / bits at the start of decoding, there is no need to UP Therefore, PC UP is directly fed to the cascade module instead of its pruned version. In addition, since there are no decoded voxels, the context building module also directly outputs the PC with all zeros on its voxels CTXSimilar to the encoder side, in a simplified implementation, the voxel groups other than the first voxel group can also follow the following Fig.13 The process shown, i.e., no voxel pruning (1520) is performed and / or all contexts are set to zero. In this case, information from previously decoded bits is partially used or not used. However, since the geometry information of the previous LoD is passed through the PC UP The geometric features in are propagated to the current LoD, so even if the information from previously decoded bits is not fully utilized, the occupancy probability estimation module can still perform reasonably well.
[0103] Similar to encoding, after decoding the current LoD, the final step is to calculate the voxel-by-voxel features for decoding the next LoD, such as Fig.12 shown.
[0104] Context Building
[0105] In addition to the binary context examples shown above, the context point cloud PC CTX Other contextual information that aids in the probability estimation may also be included.
[0106] In one embodiment, the context information is further enhanced by x, y and z coordinates. Specifically, for an occupied voxel A located at (x, y, z), it is characterized by a vector f = [xy z].
[0107] In another embodiment, the normalized coordinates are used as context information. n , whose dimension is 2 n ×2 n ×2 n , then with PC CTX The eigenvector associated with the voxel (x, y, z) in is f = [x / 2 n y / 2 n z / 2 n ].
[0108] Furthermore, we can work with spherical coordinates instead of Euclidean coordinates, which is useful in the case of working with LiDAR scans. To do this, we apply the following formula:
[0109]
[0110] φ=arccos((y-2 n-1 ) / r)
[0111] θ=arctan((y-2 n-1 ) / (x-2 n-1 ))
[0112] where r is the radial distance, is the elevation angle, and θ is the azimuth angle. Then the vector f becomes or If the distance is normalized.
[0113] In another embodiment, the context information may also be PC n The current bit depth / LoD of PC is n. In this case, PC' CTX The eigenvectors of are just scalars f=n.
[0114] In one embodiment, the context information may also be the position of the child voxel relative to its parent voxel. For example, "front" and "back", "left" and "right", "up" and "down" may be represented as "0" and "1", respectively, as shown in FIG. Fig.16A Then the child voxels located in front, right and above their parent voxels should have PC CTX The context feature f = [0 1 0] in , while the child voxels located behind, to the right and above their parent voxels should have f = [1 1 0], as Fig. 16B In another embodiment, the three additional bits representing the position are further converted into a decimal number, for example, f=[0 1 0] becomes a scalar f=2, and f=[1 1 0] becomes a scalar f=6.
[0115] In one embodiment, the enhanced contextual information may be a portion of all of the foregoing examples or any combination and permutation.
[0116] Feature Aggregation
[0117] The purpose of the feature aggregation module is to refine the input features so that they can better serve the occupancy probability estimation.
[0118] In one embodiment, it is just a series of sparse 3D convolutional layers with a ReLU activation function after each sparse 3D convolution, such as Fig.17 Note that “CONV D” represents a sparse 3D convolutional layer with D output channels.
[0119] In another embodiment, the feature aggregation module adopts a ResNet architecture, such as Fig.18 In this case, it shows the architecture of a ResNet block for aggregating features with D channels. Fig.17 compared to, Fig.18 Residual connections from the input are introduced and the outputs of the convolutional layers are added.
[0120] In another embodiment, the feature aggregation module adopts the Inception-ResNet (IRN) architecture, such as Fig.19 In this example, it shows the architecture of an IRN block for aggregating features with D channels.
[0121] In another embodiment, a transformer architecture similar to the voxel transformer proposed by Mao, Jiageng et al. in the article entitled “Voxel transformer for 3D object detection” in the 2021 IEEE / CVF International Conference on Computer Vision Proceedings is adopted. Fig. 20 A diagram of a Transformer block consisting of a self-attention block with residual connections and an MLP block (consisting of MLP layers) with residual connections is shown. Fig.21 A block diagram of the self-attention block is shown. The details are described below.
[0122] Given the current feature vector f associated with voxel location A A and its relationship with the voxel position A i The associated k adjacent features f Ai , the self-attention block strives to be based on all neighboring features f Ai To update the feature f A , where A i (0≤i≤k-1) are the k nearest neighbors of A in the input sparse tensor. First, we obtain point A by performing a k nearest neighbor (kNN) search based on the coordinates of A. i Then, the query embedding Q of A is calculated by the following formula: A :
[0123] Q A =MLP Q (f A ),
[0124] Afterwards, the key embeddings K of all the nearest neighbors of A are calculated Ai The sum value is embedded in V Ai :
[0125]
[0126] Among them, MLP Q (·),MLP K (·) and MLP V (·) are MLP layers for obtaining query, key, and value, respectively, and E Ai are voxels A and A i The position code between is calculated by the following formula:
[0127]
[0128] Among them, MLP P (·) is the MLP layer used to obtain the positional encoding, PA and P Ai are the three-dimensional coordinates, which are voxels A and A i The output feature of the position A of the self-attention block is:
[0129]
[0130] Where σ(@) is the Softmax normalization function, d is the feature vector f A The length of , and c is a predefined constant.
[0131] The transformer block updates the features of all occupied positions in the sparse tensor in the same way, and then outputs the updated sparse tensor. Note that in the simplified embodiment, the MLP Q (·),MLP K (·),MLP V (·) and MLP P (·) can contain only one fully connected layer, which corresponds to a linear projection.
[0132] In one embodiment, several feature aggregation blocks ( Fig.17 , Fig.18 , Fig.19 and Fig. 20 ) can be cascaded together to further enhance performance, such as Fig. 22 As shown. The feature aggregation blocks can be of the same type, for example, they are all transformer blocks. In this case, the parameters of their neural network layers may or may not be shared. The feature aggregation blocks can also be a mixture of different types of feature aggregation blocks, such as a mixture of IRN blocks and transformer blocks.
[0133] Probability Estimation
[0134] In one embodiment, the probability estimation module is composed of a series of multi-layer perceptron (MLP) layers. Assume that the input point cloud contains a length D of pixels residing on its voxels. 1 The vector features of the MLP have k channels (D 1 ,D 2 ,…,D k-1 ,1) layer is used for classification. The MLP result is then fed into the Softmax function, which converts the MLP output into a range from 0 to 1, representing a probability value.
[0135] Enhanced voxel upsampling
[0136] In one embodiment, during the preparation step of encoding / decoding ( Figure 8), an additional feature aggregation module is attached just after the voxel upsampling module for initial feature aggregation and refinement. With this additional feature aggregation module, the compression performance can be further improved.
[0137] Training strategy
[0138] In order to obtain suitable neural network parameters for compression, training needs to be performed first. In order to efficiently train the neural network module, we also propose a training strategy, which we call the probabilistic training strategy.
[0139] According to one embodiment, Fig.24 A block diagram of the proposed probabilistic training strategy is shown in . In each training step, we first randomly select (2410) a LoD from the octree for training. Specifically, given an input point cloud with a total bit depth of N, this step randomly generates LoDs from 1 to N. The selected LoD can be drawn according to a predefined probability distribution. In one embodiment, it is drawn according to a uniform distribution. In another embodiment, it is drawn according to a predefined polynomial distribution.
[0140] After that, we randomly select (2420) groups of voxels from the selected LoD. Specifically, we first randomly select a number m from the range of 0 to 7 (note that there are 8 groups of voxels in the LoD), and then randomly select m groups of voxels from all 8 groups of voxels. We assume that these selected m groups of voxels are known for context modeling, i.e., they have been encoded / decoded.
[0141] Next, if m≠0, according to Fig.14 (or Fig.15 ) and if m = 0, according to Fig.13 As shown in the probability estimation process, we calculate (2430) the occupancy probabilities of all remaining (8-m) groups of voxels.
[0142] Finally, a loss function, such as a binary cross entropy loss, is calculated (2440) between the estimated probability and the ground truth occupancy state at the selected LoD of the octree. Note that the binary cross entropy loss is a typical loss function for binary classification, which characterizes the difference between the occupancy probability and the ground truth occupancy state. The calculated binary cross entropy loss is used to perform back propagation to update (2450) the neural network parameters. The training process is repeated until a predefined condition is met, such as a predefined number of training steps.
[0143] Experimental Results
[0144] The proposed method is applied to lossless encoding of Ford point cloud sequences. The Ford dataset is a test dataset recommended by MPEG G-PCC Common Test Conditions (CTC). It contains 4500 LiDAR frames collected from a driving car for autonomous driving applications. In this experiment, we use 1500 LiDAR frames to train the neural network, while the remaining 3000 LiDAR frames are reserved for testing.
[0145] We experimented with the following implementation, where voxel pruning was included during encoding and decoding ( Fig.11 and Fig.15 ), and the context information additionally includes normalized coordinates and current bit depth / LoD.
[0146] First, the result provided by MPEG G-PCC octree (MPEG's standardized method, which is a non-learning method) is 22.35bpp (bits per point, the smaller the better). On the other hand, SparsePCGC uses a 5.6×10 6 The neural network model with 10 parameters provides 20.36 bpp, while our proposal uses a neural network with 4.9×10 6 The neural network model with 2 parameters provides 20.21 bpp. Therefore, our proposal achieves the best compression performance. Compared with SparsePCGC, we provide better compression performance with a smaller neural network model.
[0147] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0148] Various methods are described herein, and each method includes one or more steps or actions for realizing the described method.Unless the correct operation of the method requires a specific order of steps or actions, otherwise the order and / or use of specific steps and / or actions can be modified or combined.In addition, terms such as "first", "second" etc. can be used to modify elements, components, steps, operations, etc. in various embodiments, for example, such as "first decoding" and "second decoding".Unless specifically required, the use of these terms does not mean the sequencing of the operation modified.Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can occur in, for example, before, during or in a time period overlapping with the second decoding.
[0149] The implementations and aspects described herein can be implemented in, for example, methods or processes, devices, software programs, data streams, or signals. Even if only discussed in the context of a single implementation form (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., devices or programs). Devices can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.
[0150] Reference to "one embodiment" or "an embodiment" or "an implementation" or "implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation", as well as any other variations appearing in various places throughout the application, are not necessarily all referring to the same embodiment.
[0151] Additionally, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0152] Additionally, this application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0153] Furthermore, this application may refer to "receiving" various pieces of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is generally involved in one way or another.
[0154] It should be understood that any of the following uses of " / ", "and / or", and "at least one of" are intended to encompass selecting only the first listed option (A), or only the second listed option (B), or both options (A and B), for example, in the case of "A / B", "A and / or B", and "at least one of A and B". As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to many of the listed items, as will be apparent to one of ordinary skill in this and related arts.
[0155] It will be apparent to one of ordinary skill in the art that embodiments may generate a variety of signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.
Claims
1. A method for encoding or decoding point cloud data, include: Obtaining features associated with point cloud data of a point cloud, the point cloud data represented in a sparse tensor format at a level of detail (LoD); processing the features associated with the LoD to match a resolution of another LoD, wherein the other LoD is subsequent to the LoD; For each occupied voxel in the LoD, encoding or decoding a plurality of sub-voxels at the other LoD based on the processed features to obtain occupancy information at the other LoD; as well as The processed features are updated based on the occupancy information at the other LoD to generate updated features associated with the other LoD.
2. The method according to claim 1, wherein the encoding or decoding of the plurality of sub-voxels at the other LoD comprises, for a current sub-voxel belonging to the plurality of sub-voxels at the other LoD, performing the following operations: obtaining occupancy information of previously encoded or decoded sub-voxels of the plurality of sub-voxels at the another LoD; Obtaining context information for encoding or decoding the current subvoxel; generating an enhanced feature by associating the context information with a feature of the current subvoxel; aggregating another feature of the current subvoxel based on the enhanced feature; generating an occupancy probability of the current subvoxel based on the another feature; as well as Based on the occupancy probability of the current sub-voxel, occupancy information of the current sub-voxel is encoded or decoded.
3. The method according to claim 1 or 2, further comprising: include: Processed features are pruned based on the occupancy information of the previously encoded or decoded sub-voxels, wherein the enhanced features are based on the pruned features.
4. The method of any one of claims 1-3, wherein the features associated with the LoD are upsampled to match the resolution of the another LoD. The method of claim 4 , wherein feature aggregation is performed on the upsampled features.
6. The method according to any one of claims 1-5, wherein the features associated with the LoD and the another LoD are geometric features.
7. The method according to any one of claims 1 to 6, further comprising: include: Based on the occupancy information and the feature at the another LoD, occupancy information and the feature at a LoD subsequent to the another LoD are obtained.
8. The method of any one of claims 2-7, wherein the context information indicates that the subvoxel is one of: (1) coded / decoded and occupied and (2) not coded / decoded. 9 . The method according to claim 8 , wherein the context information further indicates coordinate information, a bit depth of a current LoD, and a position of a child voxel relative to a corresponding parent voxel.
10. The method according to any one of claims 2-9, wherein the aggregation is based on at least one of cascaded sparse convolutional layers, ResNet, Inception-ResNet, and a transformer block with a self-attention block. The method according to claim 10 , wherein the aggregation is performed by cascading a plurality of feature aggregation modules.
12. The method according to any one of claims 2-11, wherein the occupancy probability is generated based on a plurality of multi-layer perceptron (MLP) layers.
13. A method for training neural network parameters, include: Randomly select a LoD from the octree for training; Randomly select multiple groups of voxels from the selected LoD, and assume that the selected groups of voxels are known for context modeling; Calculate the occupancy probability of the remaining set of voxels; Calculating a loss function between the estimated probability and the ground truth occupancy at the selected LoD of the octree; as well as Back-propagation is performed to update the neural network parameters.
14. An apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 1-13.
15. A signal comprising video data, formed by performing a method according to any one of claims 1-12.
16. A computer-readable storage medium having stored thereon instructions for encoding or decoding a point cloud according to the method of any one of claims 1-12.
Citation Information
Cited By
Dynamic and static combined multi-scale point cloud adaptive characterization method and system
CN120852764A
Key part point cloud segmentation method and system for oil sample collection robot navigation
CN120976552A