Learning-Based Predictive Coding for Dynamic Point Clouds
A two-tiered learning-based approach for point cloud compression classifies blocks in dynamic point clouds into inter-mode and intra-mode, using point-based neural networks for efficient motion estimation and residual feature extraction, improving compression efficiency and accuracy.
Patent Information
- Application Number
- JP2025541613
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-16
- Filing Date
- 2024-01-03
- Publication Date
- 2026-01-21
Smart Images

Figure 2026502304000001_ABST
Abstract
Description
[Technical Field]
[0001] The present embodiments generally relate to methods and apparatus for point cloud compression and processing. [Background technology]
[0002] The point cloud (PC) data format is a universal data format across various business fields, from autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, and computer graphics to the animation / film industry. 3D Light Detection and Ranging (LiDAR) sensors are being installed in autonomous vehicles, and affordable LiDAR sensors have been released by the Velodyne Velabit, Apple iPad Pro 2020, and Intel RealSense LiDAR Camera L515. Advances in sensing technology make 3D point cloud data more practical than ever and are expected to become the ultimate enabler for the applications discussed herein. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Tingyu Fan, et al., “D-DPCC: Deep Dynamic Point Cloud Compression via 3D Motion Prediction,” pp. 898-904, IJCAI 2022 [Non-patent document 2] Anique Akhtar, et al., “Inter-Frame Compression for Dynamic Point Cloud Geometry Coding,” arXiv preprint arXiv:2207.12554 (2022) [Non-patent document 3] Xingyu Liu, et al. “FlowNet3D: Learning Scene Flow in 3D Point Clouds,” In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 529-537, 2019 [Non-patent document 4] Haiyan Wang et al., “FESTA: Flow Estimation via Spatial-Temporal Attention for Scene Point Clouds,” In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 14173-14182, 2021 [Non-patent document 5] Paul J. Besl and Neil D. McKay, "A Method for Registration of 3D Shapes", PAMI, 1992 [Non-patent document 6] Y. Chen and GG Medioni, "Object modeling by registration of multiple range images", Image and Vision Computing, 10(3), 1992 Summary of the Invention
[0004] According to one embodiment, a method for decoding point cloud frames of a dynamic point cloud sequence is presented, the method comprising: decoding motion information of a current point cloud block, wherein the current point cloud block is inter-coded; and Obtaining a reference point cloud block of a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining features representative of the reference point cloud block using at least one point-based neural network; Decoding residual feature information of the current point cloud block, which represents a difference between the current point cloud block and the reference point cloud block; obtaining features of the current point cloud based on the decoded residual feature information and features representing the reference point cloud block; reconstructing 3D point positions of the current point cloud block based on features of the current point cloud block using at least another point-based neural network; and Includes:
[0005] According to another embodiment, an apparatus for decoding a point cloud frame of a dynamic point cloud sequence is presented, the apparatus comprising: one or more processors; and at least one memory coupled to the one or more processors, the one or more processors: decoding motion information of a current point cloud block, wherein the current point cloud block is inter-coded; and Obtaining a reference point cloud block of a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining features representative of the reference point cloud block using at least one point-based neural network; Decoding residual feature information of the current point cloud block, which represents a difference between the current point cloud block and the reference point cloud block; obtaining features of the current point cloud based on the decoded residual feature information and features representing the reference point cloud block; reconstructing 3D point positions of the current point cloud block based on features of the current point cloud block using at least another point-based neural network; and is configured to run
[0006] According to another embodiment, a method for encoding point cloud frames of a dynamic point cloud sequence is presented, the method comprising: estimating motion information of the current point cloud block based on the current point cloud frame and a previously reconstructed point cloud frame; obtaining a reference point cloud block from a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining, using at least a point-based neural network, a first feature representing the current point cloud block and a second feature representing the reference point cloud block; obtaining a difference between a first feature representing the current point cloud block and a second feature representing the reference point cloud block to form a residual feature; encoding the residual features; Includes:
[0007] According to another embodiment, an apparatus for encoding point cloud frames of a dynamic point cloud sequence is presented, the apparatus comprising: one or more processors; and at least one memory coupled to the one or more processors, the one or more processors: estimating motion information of the current point cloud block based on the current point cloud frame and a previously reconstructed point cloud frame; obtaining a reference point cloud block from a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining, using at least a point-based neural network, a first feature representing the current point cloud block and a second feature representing the reference point cloud block; obtaining a difference between a first feature representing the current point cloud block and a second feature representing the reference point cloud block to form a residual feature; encoding the residual features; is configured to run
[0008] One or more embodiments also provide a computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any of the embodiments described herein. One or more embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding point cloud data according to the methods described herein.
[0009] One or more embodiments also provide a computer-readable storage medium having stored thereon point cloud data generated according to the methods described above. One or more embodiments also provide methods and apparatus for transmitting or receiving point cloud data generated according to the methods described herein. [Brief explanation of the drawings]
[0010] [Figure 1] 1 shows a block diagram of a system in which aspects of the present embodiments may be implemented; [Figure 2] 1 shows learning-based PCC. [Figure 3] We present a learning-based dynamic PCC. [Figure 4] 1 illustrates an encoder for the proposed learning-based dynamic PCC according to one embodiment. [Figure 5] 1 illustrates a decoder for the proposed learning-based dynamic PCC according to one embodiment. [Figure 6] 1 illustrates block motion estimation for an ME / D module according to one embodiment. [Figure 7] 1 illustrates reference block estimation for an ME / D module according to one embodiment. [Figure 8] 10 illustrates block-wise residual subtraction of the BAinter module according to one embodiment. [Figure 9] 1 illustrates block-level compositing with inter and intra modes in a frame, according to one embodiment. [Figure 10] 1 illustrates a BSinter module with motion compensation according to one embodiment. [Figure 11]1 illustrates a voxel analysis module (VA), according to one embodiment. [Figure 12] 1 illustrates a voxel synthesis module (VS), according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] FIG. 1 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device including the various components described below and configured to perform one or more aspects described herein. Examples of such devices include various electronic devices, such as, but not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 100, alone or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described herein.
[0012] System 100 includes at least one processor 110 configured to execute loaded instructions to implement various aspects described herein, for example. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes storage device 140, which may include non-volatile memory and / or volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage device 140 may include, by way of non-limiting example, internal storage, attached storage, and / or network-accessible storage.
[0013] System 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 130 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software, as is known to those skilled in the art.
[0014] Program code to be loaded into processor 110 or encoder / decoder 130 to implement various aspects described herein may be stored in storage device 140 and then loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results of the processing of equations, formulas, arithmetic, and operational logic.
[0015] In some embodiments, memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, JPEG Pleno, MPEG-I, HEVC, or VVC.
[0016] Input to the elements of system 100 may be provided through various input devices shown in block 105. Such input devices include, but are not limited to, (i) an RF section for receiving RF signals transmitted over the air by, for example, a broadcast station, (ii) a composite input, (iii) a USB input, and / or (iv) an HDMI® input.
[0017] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select a signal frequency band, which in some embodiments may be referred to as a channel (for example), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner for performing various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0018] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within processor 110, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110 and encoder / decoder 130 operating in combination with memory and storage elements, to process the data stream as desired for presentation on an output device.
[0019] The various elements of system 100 may be provided within an integrated housing in which the various elements may be interconnected and data transmitted between them using suitable connection arrangements 115, such as internal buses known in the art, including an I2C bus, wiring, and printed circuit boards.
[0020] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented in a wired and / or wireless medium, for example.
[0021] In various embodiments, data is streamed to system 100 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received via communication channel 190 and communication interface 150, which are adapted for Wi-Fi communication. Communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, enabling streaming applications and other over-the-top communications. Other embodiments provide streaming data to system 100 using a set-top box that sends data through an HDMI connection in input block 105. Still other embodiments provide streaming data to system 100 using an RF connection in input block 105.
[0022] System 100 can provide output signals to various output devices, including display 165, speakers 175, and other peripheral devices 185. In various example embodiments, other peripheral devices 185 include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are communicated between system 100 and display 165, speakers 175, or other peripheral devices 185 using AV-like signaling. Links, CEC, or other communication protocols enable control between devices with or without user intervention. Output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 using communication channel 190 via communication interface 150. Display 165 and speakers 175 may be integrated in a single unit with other components of system 100 within an electronic device, such as a television. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0023] Alternatively, display 165 and speakers 175 may be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where display 165 and speakers 175 are external components, the output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.
[0024] It is contemplated that point cloud data may consume a large portion of network traffic, for example, between connected cars over 5G networks and during immersive communications (VR / AR). Understanding and communicating point clouds requires efficient representation formats. In particular, raw point cloud data needs to be properly organized and processed for world modeling and sensing purposes. Compression of raw point clouds is essential if the relevant scenarios require data storage and transmission.
[0025] Furthermore, point clouds may represent successive scans of the same scene containing multiple moving objects. These are called dynamic point clouds, as compared to static point clouds captured from a static scene or object. Dynamic point clouds are typically organized into frames, with different frames captured at different times. Dynamic point clouds may require real-time or low-latency processing and compression.
[0026] The automotive industry and autonomous vehicles are areas where point clouds can be used. Autonomous vehicles need to be able to "probe" their environment to make appropriate driving decisions based on their surroundings. Typical sensors like LiDAR generate (dynamic) point clouds that are used by perception engines. These point clouds are not intended for human viewing; they are typically sparse, not necessarily colored, and are captured frequently and dynamically changing. They may also have other attributes, such as reflectivity provided by LiDAR, which indicates the material of the sensed object and can be useful in determining its shape.
[0027] Virtual reality (VR) and immersive worlds are predicted by many to be the future of 2D flat video. In VR and immersive worlds, the viewer is immersed in the environment around them, unlike traditional television, where the viewer can only see the virtual world in front of them. There are various degrees of immersion depending on the viewer's degree of freedom within the environment. Point clouds are a good candidate format for delivering VR worlds. Point clouds used in VR can be static or dynamic and are typically of average size, e.g., not exceeding a few million points at any one time.
[0028] Point clouds can also be used for various purposes, such as cultural heritage / architecture, where 3D scanning of objects such as statues or buildings allows sharing of the spatial configuration of the objects without transmitting or visiting them. Point clouds can also be used to ensure preservation of knowledge of objects, for example, in case an earthquake destroys an object such as a temple. Such point clouds are typically static, colored, and large in size.
[0029] Another use case is in topographical mapping and cartography, where using a 3D representation the map is not limited to a flat surface but can include relief. Google® Maps is a good example of a 3D map, but it uses a mesh instead of a point cloud. Nevertheless, a point cloud can be an appropriate data format for a 3D map, and such point clouds are typically static, colored, and large in size.
[0030] Modeling and sensing the world via point clouds can be a useful technique for enabling machines to gain knowledge about the 3D world around them for the applications described herein.
[0031] 3D point cloud data is essentially a set of discrete samples of the surface of an object or scene. A complete representation of the real world using point samples requires a vast number of points. For example, a typical VR immersive scene contains millions of points, while a point cloud typically contains hundreds of millions of points. Therefore, processing such large point clouds is computationally prohibitively expensive, especially for consumer devices such as smartphones, tablets, and in-car navigation systems, which have limited computing power.
[0032] To perform processing or inference on point cloud data, an efficient storage method is required. To store and process input point cloud data at an affordable computational cost, one solution is to first downsample the point cloud data. The downsampled point cloud data summarizes the geometry of the input point cloud while significantly reducing the number of points. The downsampled point cloud data is then sent to a subsequent machine task for further processing. However, storage space can be further reduced by converting the raw point cloud data (original or downsampled) into a bitstream using entropy coding techniques for lossless compression.
[0033] In addition to lossless coding, many scenarios require lossy coding, which significantly improves the compression ratio while maintaining the induced distortion below a certain quality level. To improve coding efficiency, an efficient point feature extractor can be useful to improve the accuracy of reconstruction within a given resource budget.
[0034] With the growing interest in the application of sparse convolution, learning-based point cloud compression (PCC) frameworks are shifting their focus to dynamic point clouds. Dynamic coding is a well-known approach in 2D video compression, but it is a new approach in the field of point cloud compression, especially in learning-based PCC. With the evolution of sparse convolution, a new research area is forming around DPCC.
[0035] In traditional dynamic coding approaches, using a previously reconstructed frame to decode the current frame is useful for reducing the size of the bitstream when compressing point cloud data. However, points in the current frame may be newly introduced or may disappear from the previous (reference) frame. Therefore, to deal with this complex situation, a more sophisticated approach is needed, such as classifying points and blocks of points into inter-mode (dynamic interpretation with the previous frame) and intra-mode.
[0036] Learning-based dynamic point cloud compression (DPCC) is a recent active research topic with several prior works. In the following, we describe related work on DPCC and the two-tier PCC architecture.
[0037] Sparsity and non-uniformity in point clouds (PCs) are problems encountered in many learning-based PCC approaches. One solution is to design a two-tiered architecture by separating point feature learning into voxel-based and point-based methods. Voxel-based approaches are often realized by applying CNN-based designs to extract feature vectors on each downsampled 3D coordinate. However, this approach is inefficient for highly sparse PCs because features on voxels are less likely to propagate to their neighbors. For example, if a downsampled voxel merges only one point, there is less neighbor information to learn. On the other hand, point-based methods can be useful for bridging this gap by analyzing larger blocks at a time in floating space. A block segmentation module first assigns highly sparse and non-uniform points to different blocks, and then point features for each point are extracted using a shared MLP. Points within each block are grouped together for block-level feature extraction (e.g., PointNet). The collection of these extracted block features forms a feature map and is fed to a voxel-based network for further feature aggregation. The final aggregated feature maps are then coded into the bitstream. At the decoder, the process is reversed: the coded feature maps are first sent to voxel-based upsampling and then to block-level point synthesis via a shared MLP.
[0038] Before describing the state-of-the-art in DPCC, a general learning-based PCC framework is illustrated in Figure 2. First, the current PC frame is sent to a voxel-based downsampling module (210 “voxel down convolution”) to generate a feature map f c This f c is encoded into a bitstream by the entropy encoder (220), and then decoded into feature maps by the entropy decoder (230).
[0039]
number
[0040] This feature map is then upsampled via a voxel-based upsampling module (240, "voxel upconvolution") in the decoder, finally reconstructing the current PC frame. Note that in this framework, the decoding of the current PC frame does not depend on previously decoded PC frames. We call such a learning-based coding method intra-coding.
[0041] In a typical approach in DPCC, a previously decoded PC frame can be used as a reference frame to eliminate redundancy. As shown in Figure 3, this reference PC frame is first processed by a voxel-based CNN (310, "voxel-down convolution") together with the current PC frame to transform the point cloud into the feature domain. These feature maps f r (Reference PC) and f c (for the current PC) is used to estimate the motion information m from the reference frame to the current frame. Then, the motion compensation module (330) calculates the motion information m from the reference frame f r Based on the feature map of and the estimated motion information m, c predictor(f p This predictor f p is the feature f embedded in the current frame c is subtracted from (350), and the feature residual f R This f R and the motion information m are entropy coded (340) and converted into a bitstream.
[0042] On the decoder side, the previously decoded point cloud, i.e., the reference frame, is fed to a “voxel down convolution” module (370) to generate a feature map f r is obtained. The "voxel down convolution" modules of the encoder and decoder are the same. Then, f ris calculated by the motion compensation module (380, same as the one in the encoder) together with the decoded (360) motion information m to obtain the predictor f p Note that the addition (385) and subtraction operations in Figure 3 are simply vector addition and subtraction. Meanwhile, the decoder-side "voxel upconvolution" module (390) is a voxel-based CNN that upsamples the input feature maps with convolutional layers.
[0043] In the end-to-end DPCC framework (see Non-Patent Document 1), the architecture is designed according to Figure 3. In this method, the features of each frame are directly embedded through sparse convolutional layers and downsampling. Then, the feature map of the entire frame is predicted through motion estimation and the feature map of the reference PC. In Non-Patent Document 1, the motion estimation module consists of a series of sparse convolutional layers with downsampling to calculate motion information, and the motion information is encoded. In the decoder, the decoded motion information and the feature map of the reference PC are sent to the motion compensation module.
[0044] This approach can have potential problems. First, downsampling completely embeds the current frame and the reference frame, which can lead to over-generalized extracted features. Some regions within a frame may not be suitable for inter-coding. Second, motion estimation further downsamples the motion field via a sparse CNN. This can oversimplify the motion information and make it difficult to properly decode.
[0045] Another study on DPCC (see Non-Patent Document 2) directly uses sparse CNN to estimate the current PC frame f using only the reference PC. c This method predicts the feature map of f without explicit motion estimation. r f by mapping cThe decoder reconstructs the current frame hierarchically by progressively rescaling the feature embeddings.
[0046] Similar to Non-Patent Document 1, the proposal in Non-Patent Document 2 also relies entirely on voxel-based operations and implicitly injects motion information into the embedding features. This may cause similar problems as those described in Non-Patent Document 1. In the above method, feature embedding and motion information are generated by a voxel-based method. This may lead to inefficiencies when encoding dynamic point clouds with various sparsities and non-uniformities. These drawbacks can be mitigated by the proposed architecture described below.
[0047] Overall architecture of the proposed DPCC framework
[0048] An encoder diagram of our proposal, according to one embodiment, is shown in Figure 4. Considering the current point cloud PCc, it is first sent to a "block partitioning module" (410), which uniformly partitions the current point cloud in 3D space to generate a set of occupied blocks. An occupied block means that there is at least one 3D point in the block; otherwise, it is an empty block that has no 3D points in it and can be ignored in encoding / decoding. The locations of all occupied blocks are sent to a first entropy encoder EE1 (430) for encoding, which generates a first bitstream BS1.
[0049] Compared to previous works, our proposal provides the flexibility to assign different coding modes (intra or inter) to different occupied blocks in the current point cloud frame. An occupied block that is encoded / decoded without depending on other blocks in different frames is called an intra-mode block, while a block that depends on other blocks in different frames for encoding / decoding is called an inter-mode block.
[0050] Next, we introduce the motion estimation and mode decision module ME / D (420), whose purpose is to decide whether an occupied block should be encoded in intra or inter mode; for occupied blocks classified as inter blocks, additional information is output for the subsequent inter coding step.
[0051] In particular, the 3D position B from the current frame c For convenience, we will refer to this occupied block as b c 3D position B c , current frame PC c , and the reference frame PC r is sent to the ME / D module, which then generates three quantities: i) b c The mode (inter or intra) associated with c (denoted as ii)b c The estimated motion vector (MV c ), and iii) the position of the reference block in the reference frame PCr (B r For convenience, the position B of the reference frame is r The reference block of b r Essentially, b r is the reference frame PC r From b c is a predictor of
[0052] Next, all occupied blocks b c Encoding mode MODE c are assembled as a mode map and passed to the second entropy encoder EE2 (431) to generate the second bitstream BS2. Separately, the motion vectors MV for all inter mode blocks are c is also passed to a third entropy encoder EE3 (432) to produce a third bitstream BS3.
[0053] Next, we encode the actual content in the occupied block. There are two branches that operate differently depending on the classification of the occupied block. Blocks classified as intra-mode blocks are coded as follows: c Regarding the coordinate B c is the “Intra-Block Analysis Module” (460,BA intra B c Apart from that, B.A. intra is the current point cloud PC c It also receives as input the c According to PC c From Block B c and launch a point-based neural network to find b c The purpose of this is to extract the feature vector of f. c It is written as follows.
[0054] On the other hand, block b c If is classified as an intermodal block, its coordinate B c is the "Interblock Analysis Module" (470, BA inter B c Apart from that, B.A. inter Also, the reference block position B r , current point cloud PC c , and reference point group PC r It takes as input BA inter The purpose of is to find its predictor b in the feature space for encoding in order to improve compression performance. r b against c Specifically, the residual of BA inter is input position B c and B r According to b c and b r Then, launch a shared point-based neural network (e.g., PointNet) and c and b r The features of each are extracted and two feature vectors f c and f p(The subscript "p" stands for "predictor"), and then we generate a residual feature vector f r is f c From f p This residual feature vector f r BA inter is output by
[0055] Compared to previous studies, our study introduces a point-based neural network to process raw point clouds, allowing for effective analysis of the input point cloud even if it is very sparse. Furthermore, the "motion estimation and mode decision module" has access to the raw point cloud. While previous studies only used block-level features for motion estimation, our study achieves more accurate motion estimation, contributing to improved compression rates.
[0056] Whether a block is classified as an inter-block or an intra-block, the feature vector f c or f r Note that we obtain c The feature encodings of F can be combined. Specifically, the feature vectors of the occupied blocks are combined / assembled as a feature map (450). The combined feature map is c In previous studies that did not combine inter-block and intra-block features, the feature map was a homogeneous feature map. In this study, the feature map F c is a heterogeneous feature map.
[0057] F c is sent to a "voxel analysis module" (440, denoted VA), which consists of several convolutional layers operating in the voxel domain. The VA module (440) generates a mode map (indicating which blocks are inter-mode and which are intra-mode blocks) and a feature map F c as input and the feature map F cWe downsample F to further exploit the correlation between neighboring blocks. c Because BS1 is a heterogeneous feature map composed of features from both inter-mode and intra-mode blocks (indicated by the mode map sent to the VA), the VA module processes Fc by taking this mode assignment into account, which was not the case in previous studies. In one embodiment, the VA module expanded the feature map Fc by one dimension indicating whether the associated block is an inter-mode or intra-mode block, and then processed it in a convolutional layer. The downsampled feature map is passed to a fourth entropy encoder (433), which generates a fourth bitstream, BS4. In one embodiment, BS1, BS2, BS3, and BS4 can be multiplexed into one bitstream.
[0058] Figure 5 is a diagram of the proposed decoder according to one embodiment. First, BS1, BS2, BS3, and BS4 are input to the first, second, third, and fourth entropy decoders (510, 511, 512, 513, ED1, ED2, ED3, and ED4), which generate occupied block positions, mode maps, motion vectors, and downsampled feature maps, respectively.
[0059] The downsampled feature maps are then sent to a "voxel synthesis module" (520, denoted VS). The VS module consists of convolutional layers operating in the voxel domain, whose purpose is to upsample the input feature maps to obtain all individual feature vectors associated with the occupied blocks. The upsampled feature maps are then c Note that the VS module receives as input not only the downsampled feature maps but also the mode maps of the occupied blocks. This is because the occupied blocks have different coding modes (either intra or inter), so the decoded feature map F c ´ is the encoder-side feature map Fc The feature map F is useful for the quality of the reconstructed point cloud. c To obtain F′, the VS module also takes into account the mode map, unlike previous work. In one embodiment, the VS module converts the input feature map into F c Then, we upsample it to match the size of,F,, and then extend it by one dimension that indicates whether the associated block is an inter-mode or intra-mode block.,The obtained feature map is passed to a convolutional layer, which produces the output feature map,F,. c ´ is obtained.
[0060] Next, position B c Considering the occupied blocks located in the c There are two branches for decoding depending on whether MODE is inter or intra. c If is intra, meaning that the associated block is an intra-mode block, then the associated feature (F c ) is used for decoding by the "intra-block synthesis module" (BS intra BS intra is block b c In one embodiment, the BS intra applies a series of MLP layers to decode the 3D point locations. intra is position B c It takes as input, transforms the decoded points, and stores them in block position B c This is because the coordinates of the decoded point are c This is achieved by simply adding
[0061] However, MODE c If is inter, meaning that the associated block is an inter mode block, then the motion estimation module is used to estimate the PC r Reference block B rThe motion estimation module (530) calculates the position of B r =B c +MV c where MV c is block b c Next, B r , PC r , and block b c The (residual) features F associated with c Based on the "Interblock Synthesis Module" (550, BS inter ) is block b c is applied to decode BS. inter works as follows: First, the reference point group PC r From B r Reference block b located at r Cut out. b r is block b c Note that BS is a predictor of inter is b r is passed to a point-based neural network for feature extraction, and the resulting feature vector f p Note that the point-based neural network here is the same as the point-based neural network on the encoder side. Next, we calculate the residual F c f p In addition, block b c The decoded features (f c Finally, we derive the BS intra Similarly, we apply a series of MLP layers to decode the 3D point locations, transform the decoded points, and map them to block locations B c This is because the coordinates of the decoded point are c This is achieved by simply adding
[0062] The current point cloud PC is reconstructed by assembling all of the decoded occupied blocks (both inter-mode and intra-mode blocks). c ´ is finally obtained.
[0063] Block Motion Estimation and Mode Decision (ME / D) Module
[0064] As shown in Figure 6, the "ME / D" module consists of two sub-modules, "Scene Flow Estimation" (620) and "Reference Block Estimation" (630), and an optional sub-module, "Global Motion Estimation" (610).
[0065] The "Scene Flow Estimation" module (620) calculates the current point cloud (PC) for the current frame. c ), reference (previous) point cloud PC r , and block position B c as input, and then calculates the block-wise motion vectors (MV c ) and each MV c The confidence or error generated during the calculation of those MODE c The learning-based scene flow estimation method (see Non-Patent Documents 3 and 4) is an example of this module.
[0066] The first method (Non-Patent Document 3) estimates scene flow from a pair of consecutive PC frames. This method introduces a "flow embedding" layer that learns the correlation between two consecutive point clouds and a "set upconv" layer that learns to propagate features from one PC to another. The second method (Non-Patent Document 4) spatially expands the selection of points in the group query to improve the flow correspondence between two frames. Furthermore, it temporarily refines the point search region of the flow embedding step.
[0067] The "reference block estimation" module (630) c , MODE c , and B c as input and B is the block position in the reference frame corresponding to the current frame. r Outputs B r MODE c If and only if it is determined to be inter, the corresponding B c and MV c It is calculated by adding
[0068] B r Further details of how to calculate PC are shown in Figure 7. c In the case of a frame, the mode of each block is MODE c An example block position B is defined by c is PC c The arrows indicate the B c The motion vector MV estimated above c It represents the current frame (PC c ) and the reference frame (PC r ) are the same size, so B c and MV c can be transferred to the reference PC frame. The corresponding reference block location B r is B c and MV c It can be calculated by adding both B r is the output of the ME / D module.
[0069] In another embodiment, a "global motion estimation" module can be added before the scene flow estimation.
[0070] The "global motion estimation" solves the following equation:
[0071]
number
[0072] are the transposed 4D current point and 4D reference point, respectively, with their homogeneous coordinates filled with ones. g is the unknown global transformation matrix that needs to be approximated.
[0073] Point-to-point or point-to-plane ICP (see Non-Patent Document 5, Non-Patent Document 6) is an example of this global transformation. In general, ICP algorithms iterate through two steps: 1) a reference point cloud;
[0074]
number
[0075] the correspondence set K (the selection of points to be considered as correspondences between the current frame and the reference frame) and the current transformation matrix M g The current point cloud transformed by
[0076]
number
[0077] and 2) find the objective function E(M g ) to minimize the transformation matrix M g The point-to-point ICP algorithm (Non-Patent Document 5) updates the objective function
[0078]
number
[0079] and point-to-plane ICP (Non-Patent Document 6) uses
[0080]
number
[0081] where:
[0082]
number
[0083] is the normal at the reference point of K.
[0084] The "Scene Flow Estimation" submodule actually transforms the current PC frame
[0085]
number
[0086]
number
[0087] as input without homogeneous coordinates for both PC frames.
[0088] Intermode block-level analysis module (BA inter )
[0089] As shown in Figure 7, the Inter mode (for temporal analysis) and the Intra mode (for spatial analysis) are selected via the "ME / D" module in Figure 4. The Intra mode block features the current PC frame (PC c ) and aggregated directly. Meanwhile, the residual features of inter-mode blocks are extracted from the current frame (PC c ) and the reference frame (PC r ) are extracted and aggregated. The block-wise motion vectors estimated by the "Scene Flow Estimation" module (see Figure 6) identify the location of the predicted block. The aggregated features (f r ) and the corresponding aggregated feature of the current point in the current block (f c ) is extracted.
[0090] Residual features of intermodal blocks (f R ) is converted to f via the subtraction module. c From f r For more information, see BA interThe module is shown in Figure 8. In one embodiment, the subtraction module
[0091]
number
[0092] is a simple vector subtraction, i.e., f R =f c -f r Note that in another embodiment, the subtraction module is implemented by a neural network. First, two features f r and f c are concatenated and then passed through a series of convolutional layers, the output of which is the residual feature f R is.
[0093] In particular, block-wise features f c and f r are extracted by two feature extractors similar to PointNet. As shown in Figure 8, each block (B c or B r ), pointwise features are extracted via shared MLP (810, 815) and then aggregated by max pooling (820, 825) to obtain aggregated features (fa c and fa r ) occurs. fa c and fa r are embedded through block-wise shared MLP (830). These features generated for the current block and the reference block are respectively c and f r By subtracting these features (840), we obtain block-wise residual features f R is obtained.
[0094] Block-Level Synthesis Module
[0095] Block-level compositing is shown in Figure 9. inter or B.S. intra) depends on the block mode (such as, but not limited to, inter or intra) analyzed during the encoding process. c is PC c Determine the mode of each block in block position B c defines the position of each block. c For , the corresponding upsampled F c These block attributes are given by BS inter Modules and BS intra The current reconstructed point cloud is passed to both modules. c will be output.
[0096] For intra-mode, block residual synthesis is simply implemented as a series of MLP layers, as previously described in co-owned application PCT / US2022 / 052861, entitled "A Scalable Framework for Point Cloud Compression." Specifically, given the block-by-block features of an intra-mode block, they are sent to the MLP layer for decoding. The MLP layer first generates 3D coordinates for points within the decoded block. These generated 3D coordinates are local coordinates, i.e., they are relative to the center position of the intra-mode block. Therefore, the decoder finally reconstructs the intra-mode block by adding the center position of the intra-mode block to the 3D points generated by the MLP layer.
[0097] BS with block motion compensation inter Module
[0098] In one embodiment, BS inter The block diagram of the module is shown in Figure 10. The motion estimation module in Figure 5 uses the reconstructed reference point cloud PC r Predict the block position above. This block position B r and reference point cloud PC r is the feature of the reconstructed reference frame
[0099]
number
[0100] The points present in the predicted block are sent to the "Block Feature Extractor" (1010), which was previously shown in Figure 8. This feature extractor extracts block-wise features on the reference frame.
[0101]
number
[0102] Extract.
[0103] At the same time, the decoded and upsampled residual features F c is input to the block-wise MLP (1020),
[0104]
number
[0105] These two features are the summation module
[0106]
number
[0107] are added together via (1030), and the current block-wise features
[0108]
number
[0109] In one embodiment, this addition module is a simple vector addition:
[0110]
number
[0111] Note that in another embodiment, this summing module is performed by a neural network. The neural network first calculates two features:
[0112]
number
[0113] Then, the concatenated features are passed to a series of convolutional layers, and the output of the convolutional layers is the feature
[0114]
number
[0115] This current feature is unpooled (1040)
[0116]
number
[0117] This is then input to the final point-wise MLP (1050) to obtain the point-wise local coordinates
[0118]
number
[0119] These local coordinates corresponding to the reconstruction points within the block can be reconstructed at each corresponding block position B c Added to the current PC c The final position of is reconstructed (1060).
[0120] Voxel analysis and voxel synthesis
[0121] A block diagram of one embodiment of the "voxel analysis module" (i.e., VA) is shown in Figure 11. The VA module generates a mode map (which indicates which blocks are inter-mode blocks and which are intra-mode blocks) and a feature map F c As input, the mode map receives the feature map F. In one embodiment, the mode map is a binary map using 0 to indicate an inter-mode block and 1 to indicate an intra-mode block. In another embodiment, 1 to indicate an inter-mode block and 0 to indicate an intra-mode block. In FIG. 11, first, the feature map F c The and mode maps are concatenated (1110) to obtain an augmented feature map, which is passed to a series of convolutional layers (1120) for aggregation and downsampling to obtain a downsampled feature map, denoted as Fc2.
[0122] A block diagram of a "voxel synthesis module" (i.e., VS) according to one embodiment is shown in Figure 12. The VS module generates a mode map and a downsampled feature map F'. c2 The VS module first takes the feature map F' as input. c2 The size of the encoder-side feature map F c To match the feature map F' c2 , which performs voxel upsampling (1210) on F. The output feature map is concatenated (1220) with the mode map to obtain an augmented feature map. The augmented feature map is passed to a series of convolutional layers (1230) for aggregation and upsampling to obtain the output feature map F'. c is obtained.
[0123] Various numerical values are used in this application, and the specific values are for illustrative purposes only and the described aspects are not limited to these particular values.
[0124] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a particular order of steps or actions is required for proper operation of the method, the order and / or use of particular steps and / or actions may be varied or combined. Furthermore, terms such as “first,” “second,” etc. may be used in various embodiments to vary elements, components, steps, actions, etc., e.g., “first decode” and “second decode.” The use of these terms does not imply a varied order of actions unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but may be performed before, during, or during an overlapping period with the second decode.
[0125] The embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed in the context of only a single form of embodiment (e.g., only a method), the discussed functional implementation may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, or firmware. A method may be implemented in an apparatus, e.g., a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the transfer of information between end users.
[0126] Reference to "one embodiment" or "one embodiment" or "one implementation" or "one embodiment," as well as other variations thereof, means that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in one embodiment" or "in one embodiment" or "in one embodiment" in various places throughout this application, as well as any other variations, are not necessarily all referring to the same embodiment.
[0127] Additionally, this application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0128] Additionally, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, transferring information, copying information, computing information, determining information, predicting information, or inferring information.
[0129] Additionally, this application may refer to "receiving" various information. Receiving, like "accessing," is a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is typically involved in some way in an operation such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information.
[0130] For example, in the cases of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" should be understood to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second alternatives (A and B), or the selection of only the first and third alternatives (A and C), or the selection of only the second and third alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded by any number of items listed, as will be apparent to those skilled in this and related arts.
[0131] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a high frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
1. 1. A method for decoding point cloud frames of a dynamic point cloud sequence, comprising: decoding motion information of a current point cloud block, wherein the current point cloud block is inter-coded; and Obtaining a reference point cloud block of a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining features representative of the reference point cloud block using at least one point-based neural network; Decoding residual feature information of the current point cloud block, which represents a difference between the current point cloud block and the reference point cloud block; obtaining features of the current point cloud based on the decoded residual feature information and features representing the reference point cloud block; reconstructing 3D point positions of the current point cloud block based on features of the current point cloud block using at least another point-based neural network; and A method comprising:
2. 1. An apparatus for decoding a point cloud frame of a dynamic point cloud sequence, the apparatus comprising: one or more processors; and at least one memory coupled to the one or more processors, the one or more processors comprising: decoding motion information of a current point cloud block, wherein the current point cloud block is inter-coded; and Obtaining a reference point cloud block of a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining features representative of the reference point cloud block using at least one point-based neural network; Decoding residual feature information of the current point cloud block, which represents a difference between the current point cloud block and the reference point cloud block; obtaining features of the current point cloud based on the decoded residual feature information and features representing the reference point cloud block; reconstructing 3D point positions of the current point cloud block based on features of the current point cloud block using at least another point-based neural network; and 20. An apparatus configured to:
3. 3. The method of claim 1, further comprising decoding information indicating a location of an occupation block in the point cloud frame, or the apparatus of claim 2, wherein one or more processors are further configured to perform decoding information indicating a location of an occupation block in the point cloud frame.
4. 4. The method of claim 3, comprising decoding information indicating an encoding mode of each occupied block of the point cloud frame, or the apparatus of claim 3, wherein the one or more processors are further configured to perform decoding information indicating an encoding mode of each occupied block of the point cloud frame.
5. Obtaining features representative of the current point cloud block includes: Concatenating the residual feature information and features representative of the reference point cloud block; applying a series of convolutional layers to the concatenated features to form features representing the current point cloud block; 5. The method of any one of claims 1, 3 and 4 or the apparatus of any one of claims 2 to 4, comprising:
6. The method of claim 1 or any one of claims 3 to 5 or the apparatus of claim 2 to 5, wherein the at least another point-based neural network comprises a series of MLP layers.
7. Decoding the residual feature information of the current point cloud block includes: decoding a downsampled feature map, the downsampled feature map being augmented with data indicating whether intra mode or inter mode is used for a corresponding block of the current point cloud frame; and upsampling and aggregating the downsampled feature maps to obtain the residual feature information; 7. The method of any one of claims 1 and 3 to 6, comprising: One or more processors decoding a downsampled feature map, the downsampled feature map being augmented with data indicating whether intra mode or inter mode is used for a corresponding block of the current point cloud frame; and upsampling and aggregating the downsampled feature maps to obtain the residual feature information; 7. An apparatus according to claim 2, configured to perform the following:
8. The method of claim 1 or any one of claims 3 to 7 or the apparatus of claim 2 to 7, wherein another block of the current point cloud frame is decoded in intra mode without a reference point cloud frame.
9. 1. A method for encoding point cloud frames of a dynamic point cloud sequence, comprising: estimating motion information of the current point cloud block based on the current point cloud frame and a previously reconstructed point cloud frame; obtaining a reference point cloud block from a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining, using at least a point-based neural network, first features representing the current point cloud block and second features representing the reference point cloud block; obtaining a difference between a first feature representing the current point cloud block and a second feature representing the reference point cloud block to form a residual feature; encoding the residual features; A method comprising:
10. 1. An apparatus for encoding point cloud frames of a dynamic point cloud sequence, the apparatus comprising: one or more processors; and at least one memory coupled to the one or more processors, the one or more processors: estimating motion information of the current point cloud block based on the current point cloud frame and a previously reconstructed point cloud frame; obtaining a reference point cloud block from a previously reconstructed point cloud frame for predicting the current point cloud block based on the motion information of the current point cloud block; obtaining, using at least a point-based neural network, first features representing the current point cloud block and second features representing the reference point cloud block; obtaining a difference between a first feature representing the current point cloud block and a second feature representing the reference point cloud block to form a residual feature; encoding the residual features; 20. An apparatus configured to:
11. 11. The method of claim 9, further comprising encoding information indicative of a location of an occupation block in the point cloud frame, or the apparatus of claim 10, wherein one or more processors are further configured to perform encoding information indicative of a location of an occupation block in the point cloud frame.
12. 12. The method of claim 11, further comprising encoding information indicating an encoding mode for each occupied block of the point cloud frame, or the apparatus of claim 11, wherein the one or more processors are further configured to perform encoding information indicating an encoding mode for each occupied block of the point cloud frame.
13. Obtaining the difference Concatenating a first feature representing the current point cloud block and a second feature representing the reference point cloud block; applying a series of convolutional layers to the concatenated features to form the difference; 13. The method of any one of claims 9, 11 and 12 or the apparatus of any one of claims 10 to 12, comprising:
14. 14. The method of claim 9 or any one of claims 11 to 13, or the apparatus of claim 10 to 13, wherein the first feature representing the current point cloud block and the second feature representing the reference point cloud block are obtained by a shared point-based neural network.
15. downsampling and aggregating feature maps to form downsampled feature maps, the feature maps being augmented with data indicating whether intra mode or inter mode is used for corresponding blocks of the current point cloud frame; and encoding the downsampled feature map; 15. The method of any one of claims 9 and 11 to 14, further comprising: One or more processors downsampling and aggregating feature maps to form downsampled feature maps, the feature maps being augmented with data indicating whether intra mode or inter mode is used for corresponding blocks of the current point cloud frame; and encoding the downsampled feature map; 15. The apparatus of claim 10, further configured to:
16. 16. The method of claim 9 or any one of claims 11 to 15 or the apparatus of claim 10 to 15, wherein another block of the current point cloud frame is encoded in intra mode without a reference point cloud frame.
17. 17. A signal comprising point cloud data formed by carrying out the method of any one of claims 9 and 11 to 16.
18. A computer-readable storage medium having stored thereon instructions for encoding or decoding a point cloud according to the method of any one of claims 1, 3 to 9 and 11 to 16.