Learning-based predictive coding for dynamic point clouds
By combining point-based neural networks and voxel networks, block partitioning and pattern decision-making are performed on dynamic point clouds, solving the problem of low compression efficiency of sparse and non-uniform point clouds in existing technologies, and achieving more efficient and accurate point cloud encoding and decoding.
Patent Information
- Application Number
- CN202480018125.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-16
- Filing Date
- 2024-01-03
- Publication Date
- 2025-11-14
AI Technical Summary
Existing dynamic point cloud compression methods are inefficient when dealing with sparse and non-uniform point clouds, have inaccurate motion estimation, and are difficult to effectively utilize previous frames for efficient encoding and decoding.
A method combining point-based neural networks and voxel-based networks is adopted. Through block partitioning and pattern decision modules, the point cloud frame-intra-frame patterns are distinguished and processed. The point-based neural network is used to extract features, and combined with motion estimation and pattern decision modules, heterogeneous feature maps are generated for encoding and decoding.
It improves the efficiency and accuracy of point cloud compression, can better handle sparse and non-uniform point clouds, reduces redundancy in encoding and decoding, and improves the quality of point cloud reconstruction.
Smart Images

Figure CN120958486A_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally relates to a method and apparatus for point cloud compression and processing. Background Technology
[0002] Point cloud (PC) data format is a common data format across several business sectors, such as autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, computer graphics, and the animation / film industry. 3D LiDAR (light detection and ranging) sensors have been deployed in autonomous vehicles, and affordable LiDAR sensors have been released by Velodyne Velabit, Apple iPad Pro 2020, and Intel RealSense LiDAR camera L515. With advancements in sensing technology, 3D point cloud data is becoming more practical than ever before and is expected to be the ultimate enabler in the applications discussed in this article. Summary of the Invention
[0003] According to one embodiment, a method for decoding point cloud frames in a dynamic point cloud sequence is proposed. The method includes: decoding motion information of a current point cloud block, wherein the current point cloud block is inter-frame encoded; obtaining a reference point cloud block in a previously reconstructed point cloud frame based on the motion information of the current point cloud block to predict the current point cloud block; obtaining features representing the reference point cloud block using at least a point-based neural network; decoding residual feature information of the current point cloud block, the residual feature information representing the difference between the current point cloud block and the reference point cloud block; obtaining features of the current point cloud block based on the decoded residual feature information and the features representing the reference point cloud block; and reconstructing the positions of 3D points in the current point cloud block based on the features of the current point cloud block using at least another point-based neural network.
[0004] According to another embodiment, an apparatus for decoding point cloud frames in a dynamic point cloud sequence is provided. The apparatus includes one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: decode motion information of a current point cloud block, wherein the current point cloud block is inter-frame encoded; obtain a reference point cloud block in a previously reconstructed point cloud frame based on the motion information of the current point cloud block to predict the current point cloud block; obtain features representing the reference point cloud block using at least a point-based neural network; decode residual feature information of the current point cloud block, the residual feature information representing the difference between the current point cloud block and the reference point cloud block; obtain features of the current point cloud block based on the decoded residual feature information and the features representing the reference point cloud block; and reconstruct the positions of 3D points in the current point cloud block based on the features of the current point cloud block using at least another point-based neural network.
[0005] According to another embodiment, a method for encoding point cloud frames in a dynamic point cloud sequence is proposed. The method includes: estimating motion information of a current point cloud block based on a current point cloud frame and a previously reconstructed point cloud frame; obtaining a reference point cloud block from the previously reconstructed point cloud frame based on the motion information of the current point cloud block to predict the current point cloud block; obtaining a first feature representing the current point cloud block and a second feature representing the reference point cloud block using at least a point-based neural network; obtaining the difference between the first feature representing the current point cloud block and the second feature representing the reference point cloud block to form a residual feature; and encoding the residual feature.
[0006] According to another embodiment, an apparatus for encoding point cloud frames in a dynamic point cloud sequence is provided. The apparatus includes one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: estimate motion information of a current point cloud block based on a current point cloud frame and a previously reconstructed point cloud frame; obtain a reference point cloud block from the previously reconstructed point cloud frame based on the motion information of the current point cloud block to predict the current point cloud block; obtain a first feature representing the current point cloud block and a second feature representing the reference point cloud block using at least a point-based neural network; obtain the difference between the first feature representing the current point cloud block and the second feature representing the reference point cloud block to form a residual feature; and encode the residual feature.
[0007] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any embodiment described herein. One or more of these embodiments also provide a computer-readable storage medium having instructions stored thereon for encoding or decoding point cloud data according to the methods described herein.
[0008] One or more embodiments also provide a computer-readable storage medium storing point cloud data generated according to the method described above. One or more embodiments also provide methods and apparatus for transmitting or receiving point cloud data generated according to the methods described herein. Attached Figure Description
[0009] Figure 1 A block diagram of a system in which aspects of this embodiment can be implemented is shown.
[0010] Figure 2 The diagram illustrates learning-based PCC.
[0011] Figure 3 The diagram illustrates a learning-based dynamic PCC.
[0012] Figure 4 The figure illustrates a learning-based dynamic PCC encoder according to one embodiment.
[0013] Figure 5 The figure illustrates a proposed learning-based dynamic PCC decoder according to one embodiment.
[0014] Figure 6 The illustration shows block motion estimation of an ME / D module according to one embodiment.
[0015] Figure 7 The illustration shows a reference block estimation of an ME / D module according to one embodiment.
[0016] Figure 8 The illustration shows a BA according to one embodiment. inter Block-by-block residual subtraction of the module.
[0017] Figure 9 The illustration shows block-level composition in a frame using inter-frame and intra-frame modes according to one embodiment.
[0018] Figure 10 The figure illustrates a motion-compensated BA according to one embodiment. inter Module.
[0019] Figure 11 The illustration shows a voxel analysis module (VA) according to one embodiment.
[0020] Figure 12 The illustration shows a voxel synthesis module (VS) according to one embodiment. Detailed Implementation
[0021] Figure 1 The diagram illustrates an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.
[0022] System 100 includes at least one processor 110 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.
[0023] System 100 includes an encoder / decoder module 130, which is configured to process data, for example, to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. Encoder / decoder module 130 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100, or it may be incorporated within processor 110 as a combination of hardware and software known to those skilled in the art.
[0024] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0025] In several embodiments, memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, JPEG Pleno, MPEG-1, HEVC, or VVC.
[0026] Inputs can be provided to the components of system 100 through various input devices as indicated in box 105. Such input devices include, but are not limited to: (i) an RF section that receives, for example, RF signals transmitted over the air by a broadcasting company, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0027] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or limiting a signal band to a certain band); (ii) down-converting the selected signal; (iii) further limiting the band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the aforementioned (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0028] Additionally, USB and / or HDMI terminals may include corresponding interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, in a separate input processing IC or in processor 110. Similarly, as needed, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.
[0029] Various components of system 100 can be provided within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data between them using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit board.
[0030] System 100 includes a communication interface 150, which is capable of communicating with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card, and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.
[0031] In various embodiments, a Wi-Fi network such as IEEE 802.11 is used to stream data to system 100. The Wi-Fi signals in these embodiments are received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100, delivering data via an HDMI connection to input box 105. Still other embodiments use an RF connection to input box 105 to provide streaming data to system 100.
[0032] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various examples of embodiments, the other peripheral devices 185 include one or more of the following: a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are transmitted between system 100 and the display 165, speaker 175, or other peripheral devices 185 using signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention). Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. The display 165 and speaker 175 can be integrated into a single unit within an electronic device (e.g., a television set) along with other components of system 100. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0033] For example, if the RF portion of input 105 is part of a separate set-top box, then display 165 and speaker 175 may alternatively be separated from one or more other components. In various embodiments where display 165 and speaker 175 are external components, output signals may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0034] It is foreseeable that point cloud data can consume a significant portion of network traffic, for example, between cars connected via 5G networks, and in immersive communications (VR / AR). Efficient representation formats are essential for point cloud understanding and communication. In particular, raw point cloud data needs to be properly organized and processed for world modeling and sensing. Compression of the raw point cloud is crucial when data needs to be stored and transmitted in relevant scenarios.
[0035] Furthermore, point clouds can represent sequential scans of the same scene containing multiple moving objects. These point clouds are called dynamic point clouds compared to static point clouds captured from static scenes or objects. Dynamic point clouds are typically organized into frames, where different frames are captured at different times. Dynamic point clouds may require real-time or low-latency processing and compression.
[0036] The automotive industry and autonomous vehicles are among the sectors where point clouds can be used. Autonomous vehicles should be able to "detect" their environment, enabling them to make good driving decisions based on the actual conditions of their immediate surroundings. Typical sensors, such as LiDAR, generate (dynamic) point clouds that are used by perception engines. These point clouds are not intended to be seen by the human eye, and they are typically sparse, not necessarily colored, and dynamic at high capture frequencies. They can possess other properties, such as the reflectivity provided by LiDAR, as this property indicates the material of the sensed object and can aid in decision-making.
[0037] Virtual reality (VR) and immersive worlds are envisioned by many as the future of 2D flat video. In VR and immersive worlds, the viewer is immersed in the environment surrounding them, unlike standard TV where the viewer only sees the virtual world in front of them. Depending on the viewer's degree of freedom within the environment, there are several gradations in terms of immersion. Point clouds are an excellent candidate format for distributing VR worlds. Point clouds used for VR can be static or dynamic and typically have an average size, for example, no more than a few million points at a time.
[0038] Point clouds can also be used for various purposes, such as cultural heritage / buildings, where objects (such as statues or buildings) are 3D scanned to share their spatial arrangement without sending or visiting the objects. Furthermore, point clouds can be used to ensure the preservation of knowledge about an object in cases where it may be damaged (e.g., a temple destroyed by an earthquake). Such point clouds are typically static, colored, and large in size.
[0039] Another use case is in the fields of topography and cartography, where maps are not limited to a flat plane and can include relief by using 3D representation. Google Maps is an excellent example of a 3D map, but it uses a grid instead of a point cloud. Nevertheless, point clouds can be a suitable data format for 3D maps, and such point clouds are typically static, colored, and large.
[0040] World modeling and sensing via point clouds can be a useful technique to allow machines to acquire knowledge about the 3D world around them for the applications discussed in this paper.
[0041] 3D point cloud data is essentially discrete samples of the surface of an object or scene. To fully represent the real world using these point samples, a massive number of points are actually required. For example, a typical VR immersive scene contains millions of points, while a point cloud typically contains hundreds of millions. Therefore, processing such large-scale point clouds is computationally expensive, especially for consumer devices with limited computing power, such as smartphones, tablets, and car navigation systems.
[0042] To perform processing or inference on point clouds, efficient storage methods are required. One solution to store and process input point clouds at affordable computational cost is to first downsample the point cloud, where the downsampled point cloud generalizes the geometry of the input point cloud while having far fewer points. The downsampled point cloud is then fed to subsequent machine tasks for further processing. However, further reductions in storage space can be achieved by converting the raw point cloud data (raw or downsampled) into a bitstream for lossless compression using entropy coding techniques.
[0043] Besides lossless coding, many scenarios seek lossy coding to significantly improve compression ratios while maintaining the resulting distortion at a specific quality level. To improve coding efficiency, efficient point feature extractors can help improve reconstruction accuracy within a given resource budget.
[0044] With the growing interest in applying sparse convolutions, learning-based point cloud compression (PCC) frameworks have shifted their focus to dynamic point clouds. Dynamic coding is a well-known method in 2D video compression, but it is relatively new in the field of point cloud compression, especially for learning-based PCC. Consistent with the evolution of sparse convolutions, a new research area has emerged around DPCC.
[0045] Similar to traditional dynamic coding methods, using previously reconstructed frames to decode the current frame helps reduce the bitstream size when compressing point cloud data. However, points in the current frame may be newly appearing or may have disappeared in previous (reference) frames. Therefore, more sophisticated methods are needed, such as classifying points and blocks of points into inter-frame modes (utilizing the dynamic interpretation of previous frames) and intra-frame modes, to handle this complexity.
[0046] Learning-based dynamic point cloud compression (DPCC) is a recent hot research topic, in which some previous work has been incorporated. The following section describes the relevant work on DPCC along with a two-level PCC architecture.
[0047] Sparsity and non-uniformity in point clouds (PCs) are problems encountered in many learning-based PCC methods. One solution is to design a two-tier architecture by splitting point feature learning into a voxel-based approach and a point-based approach. Voxel-based methods are typically implemented by applying a CNN-based design to extract feature vectors at each downsampled 3D coordinate, but this approach is inefficient for very sparse PCs because features on voxels become difficult to propagate to their neighbors. For example, if downsampled voxels only merge a single point, there is little neighbor information available for learning. On the other hand, point-based methods can help compensate for this deficiency by analyzing larger blocks at once in a floating space. The block partitioning module first assigns very sparse and non-uniform points to different blocks, and then point features are extracted for each point using a shared MLP. Points within each block are grouped together for block-level feature extraction (e.g., PointNet). The set of these extracted block features forms a feature map and is fed into a voxel-based network for further feature aggregation. The final aggregated feature map is encoded as a bitstream. On the decoder, this process is reversed. First, the encoded feature map is sent to a voxel-based upsampling system, and then block-level point synthesis is performed via a shared MLP.
[0048] Before describing the current state of the art regarding DPCC, Figure 2 The paper describes a general learning-based PCC framework. First, the current PC frame is fed into a voxel-based downsampling module (210, "voxel subconvolution"), and a feature map f is generated. c The f cThe bitstream is encoded by the entropy encoder (220) and then decoded by the entropy decoder (230) into a feature map. . The current PC frame is upsampled back via a voxel-based upsampling module (240, "voxel upconvolution") in the decoder to ultimately reconstruct it. It's important to note that in this framework, decoding the current PC frame does not depend on any previously decoded PC frames. This learning-based coding method is called intra-frame coding.
[0049] A typical approach to DPCC is to use previously decoded PC frames as a reference to eliminate redundancy. For example... Figure 3 As shown, the reference PC frame and the current PC frame are first processed by a voxel-based CNN (310, "voxel-subconvolution") to transform the point cloud to the feature domain. These feature maps f r (For reference PC) and f c (For the current PC) is used to estimate motion information m from the reference frame to the current frame. Then, the motion compensation module (330) uses the feature map f of the reference frame. r And the estimated motion information m generates f c The predicted value (denoted as f) p Embedding features f from the current frame c Subtract (350) from the predicted value f p To generate characteristic residual f R The f R The motion information m is entropy encoded (340) into a bit stream.
[0050] On the decoder side, the previously decoded point cloud (i.e., the reference frame) is fed into the "voxel subconvolution" module (370) to obtain the feature map f. r It's important to note that the "voxel-based subconvolution" module is the same in both the encoder and decoder. Then, f r Together with the decoded (360) motion information m, a predicted value f is generated via a motion compensation module (370, the same as the motion compensation module on the encoder). p It is important to note that, Figure 3 The addition (380) and subtraction operations in the code are simply vector addition and subtraction. On the other hand, the "voxel-on-convolution" module (390) on the decoder side is a voxel-based CNN that upsamples the input feature map through convolutional layers.
[0051] In the end-to-end DPCC framework (see Tingyu Fan et al., “D-DPCC: Deep Dynamic PointCloud Compression via 3D Motion Prediction,” pp. 898-904, IJCAI 2022, hereinafter referred to as “Fan”), the architecture is based on... Figure 3 The design approach embeds features from each frame directly via sparse convolutional layers and downsampling. The feature map for the entire frame is then predicted using the feature map of the reference frame PC and motion estimation. In Fan, the motion estimation module consists of a series of sparse convolutional layers with downsampling to compute motion information, which is then encoded. In the decoder, the feature map of the reference frame PC and the decoded motion information are fed into the motion compensation module.
[0052] This approach may have problems. First, using fully embedded downsampled current and reference frames may overgeneralize the extracted features. Some regions within a frame may not be suitable for inter-frame encoding. Second, motion estimation further downsamples the motion field using a sparse CNN. This may oversimplify the motion information, making it difficult to decode accurately.
[0053] Another DPCC work (see Anique Akhtar et al., “Inter-Frame Compression for Dynamic Point Cloud Geometry Coding,” arXiv preprint arXiv:2207.12554 (2022), hereinafter referred to as “Akhtar”) proposes to predict the feature map f of the current PC frame by directly employing a sparse CNN and using only the reference PC. c This method involves f r f is generated by mapping to the downsampled voxel position of the current frame. c The decoder generates predicted values without explicit motion estimation. It reconstructs the current frame hierarchically by progressively re-expanding the feature embeddings.
[0054] Similar to Fan, Akhtar's proposal also relies entirely on voxel-based operations and implicitly incorporates motion information into the embedded features. This can lead to problems similar to those described in Fan. In the aforementioned methods, feature embeddings and motion information are generated using voxel-based methods. Therefore, this can be inefficient when encoding dynamic point clouds with various sparsities and non-uniformities. These drawbacks can be mitigated by the proposed architecture, which will be described below.
[0055] Overall architecture of the proposed DPCC framework According to one embodiment, Figure 4 The diagram illustrates our proposed encoder. Given the current point cloud PC... c In this case, it is first fed to the "block partitioning module" (410), which uniformly divides the current point cloud in 3D space to obtain a set of occupied blocks. An occupied block means that there is at least one 3D point in the block; otherwise, the block is an empty block, which can be ignored during encoding / decoding because there are no 3D points in it. The positions of all occupied blocks are sent to the first entropy encoder EE1 (430) for encoding to obtain the first bitstream BS1.
[0056] Compared to previous work, our proposal offers the flexibility to assign different encoding / decoding modes (intra-frame or inter-frame) to different occupancy blocks within the current point cloud frame. Occupancy blocks that are encoded / decoded without relying on another block from a different frame are called intra-frame mode blocks; while blocks that rely on another block from a different frame for encoding / decoding are called inter-frame mode blocks.
[0057] Next, we introduce the motion estimation and mode decision module ME / D (420). Its purpose is to determine whether an occupied block should be encoded in intra-frame mode or inter-frame mode, and for occupied blocks classified as inter-frame blocks, it will output additional information for subsequent inter-frame coding steps.
[0058] Specifically, assume that there is a source located at position B in 3D. c The current frame's occupied block. For convenience, b will also be used below. c This represents the occupied block. Then, the 3D position B... c Current frame PC c and reference frame PC r The feed is sent to the ME / D module, which outputs three quantities: i) and b c The associated mode (inter-frame or intra-frame) is represented as MODE. c ;ii) with b c The associated estimated motion vector is denoted as MV. c ; and iii) the reference block in the reference frame PC r The position above is represented as B. r For convenience, position B on the reference frame is used. r The reference block at that location is also represented as b. r Essentially, b r It comes from the reference frame PC r b c The predicted value.
[0059] All occupying blocks b cEncoding / decoding mode MODE c They are then assembled into a pattern graph and passed to the second entropy encoder EE2 (431) to generate the second bitstream BS2. In addition, the motion vectors MV of all inter-frame pattern blocks are... c It is also passed to the third entropy encoder EE 3 (432) to generate the third bitstream BS3.
[0060] Next, the actual content within the occupied block is encoded. Depending on the classification of the occupied block, there are two branches that operate in different ways. For block b, which is classified as an intra-frame mode block... c Its coordinates B c It will be passed to the "Intra-Block Analysis Module" (460, denoted as BA) intra ). Except for B c In addition, BA intra Also, the current point cloud PC c As input. Its purpose is based on the provided location B. c block b c From PC c The image is cropped from the middle, and then a point-based neural network is activated to extract b. c The eigenvectors obtained are represented as f. c .
[0061] On the other hand, if block b c If it is classified as an inter-frame mode block, then its coordinates B c It will be passed to the "Inter-Block Analysis Module" (470, denoted as BA) inter ). Except for B c In addition, BA inter Also refer to block position B r Current point cloud PC c and reference point cloud PC r As input. BA inter The goal is to compute b in the feature space. c Relative to its predicted value b r The residuals are used for encoding, thereby enhancing compression performance. Specifically, BA... inter Based on input position B c and B r Visit b c and b r Then, it launches a shared point-based neural network (e.g., PointNet) to extract b. c and b r The characteristics are used to obtain two feature vectors f. c and f p (The subscript "p" refers to the "predicted value"). Then, by using fc Subtract f from the middle p To calculate the residual eigenvector f r Then, by BA inter Output residual eigenvector f r .
[0062] It is important to note that, compared to previous work, our work introduces a point-based neural network to digest the raw point cloud, thus enabling efficient analysis even when the input point cloud is very sparse. Furthermore, the "Motion Estimation and Pattern Decision Module" has access to the raw point cloud. Compared to previous work that only used block-by-block features for motion estimation, our work provides more accurate estimated motion, which further benefits compression.
[0063] It is important to note that regardless of whether a block is classified as an inter-frame block or an intra-frame block, it always obtains a feature vector f. c or f r Therefore, for PC c The encoding of features can be unified. Specifically, the feature vectors of the occupying blocks are combined / assembled (450) into a feature map. The combined feature map is represented as F. c In previous works without inter-frame / intra-frame combination, the feature maps were homogeneous. In this work, feature map F... c It is a feature map of heterogeneous.
[0064] F c The data is then fed into the "voxel analysis module" (440, denoted as VA), which consists of several convolutional layers operating in the voxel domain. The VA module (440) combines the pattern map (indicating which blocks are inter-frame pattern blocks and which are intra-frame pattern blocks) and the feature map F. c As input, and for feature map F c Downsampling is performed to further utilize the correlation between adjacent blocks. Due to the feature map F c It is a heterogeneous feature map—this feature map is composed of features from both inter-frame and intra-frame mode blocks (represented by a mode map fed to the VA), so the VA module processes F by taking this mode allocation into account. c However, this was not the case in previous work. In one embodiment, the VA module will feature map F c An additional dimension is added—indicating whether the associated block is an inter-frame mode block or an intra-frame mode block—and then processed using convolutional layers. The downsampled feature map is then passed to a fourth entropy encoder (433) to obtain a fourth bitstream BS4. In one embodiment, BS1, BS2, BS3, and BS4 can be multiplexed into a single bitstream.
[0065] According to one embodiment, Figure 5 The proposed decoder diagram is provided. First, BS1, BS2, BS3 and BS4 are fed into the first, second, third and fourth entropy decoders (510, 511, 512, 513, ED1, ED2, ED3 and ED4) to generate the occupied block location, pattern map, motion vector and downsampled feature map, respectively.
[0066] Subsequently, the downsampled feature map is fed into the "voxel analysis module" (520, denoted as VS). The VS module consists of convolutional layers operating in the voxel domain, and its purpose is to upsample the input feature map to obtain the individual feature vector associated with each occupied block. The upsampled feature map is denoted as F. c It's important to note that the VS module takes not only the downsampled feature map as input, but also the pattern map of the occupant block. This is because the occupant block has different encoding / decoding modes (intra-frame or inter-frame), therefore the decoded feature map F... c 'Is the feature map F on the encoder side c Similar heterogeneous feature maps. To obtain feature maps F that are beneficial to the quality of the reconstructed point cloud. c The VS module also considers pattern maps, which differs from previous work. In one embodiment, the VS module upsamples the input feature map to make its size equal to F. c The matched blocks are then augmented with a dimension indicating whether the associated block is an inter-frame mode block or an intra-frame mode block. The resulting feature map is then passed to a convolutional layer to obtain the output feature map F. c '.
[0067] Next, given the location at position B c In the case of a occupied block, according to the associated MODE c Whether it's inter-frame mode or intra-frame mode, there are two branches for decoding. If MODE c If it is intra-mode (meaning the associated block is an intra-mode block), then the associated feature (represented as F) c It is fed to the "intra-block composition module" (represented as BS) intra Decode the code. BS intra It is used for block b c A point-based neural network for decoding associated 3D points. In one embodiment, it applies a series of MLP layers to decode the positions of the 3D points. BS intra Position B will also be c As input, the decoded points are translated so that they are located at block position B. c Inside. This is by using B cThis is achieved by simply adding the coordinates of the decoded point.
[0068] However, if MODE c If it is an inter-frame mode (meaning the associated block is an inter-frame mode block), then the motion prediction module is used to calculate the PC. r Reference block B on r The position. The motion prediction module (530) simply calculates B. r = B c +MV c MV c Is with block b c The associated motion vectors. Then, based on B... r PC r and with block b c Associated (residual) characteristics F c The "Inter-Frame Block Composition Module" (550, denoted as BS) is applied. inter ) to block b c Decode. BS inter The following steps are performed. First, it retrieves the reference point cloud from the PC. r Cut out the part located at B r Reference block b r It should be noted that b r It's a block of b c The predicted value. Then, BS inter b r The data is passed to a point-based neural network for feature extraction, thereby obtaining its feature vector f. p Note that the point-based neural network here is the same as the point-based neural network on the encoder side. Next, the residual F... c Add to f p Obtain block b c The decoded features are represented as f c Finally, with BS intra Similarly, it uses a series of MLP layers to decode the positions of 3D points, then translates the decoded points so that they are located at block position B. c Inside. This is by using B c This is achieved by simply adding the coordinates of the decoded point.
[0069] By assembling all the decoded occupancy blocks (both inter-frame mode blocks and intra-frame mode blocks), the reconstructed current point cloud PC is finally obtained. c '.
[0070] Block Motion Estimation and Pattern Decision (ME / D) Module like Figure 6As shown, the “ME / D” module consists of two sub-modules, “Scene Flow Estimation” (620) and “Reference Block Estimation” (630), and there is also an optional sub-module, “Global Motion Estimation” (610).
[0071] The "Scene Flow Estimation" module (620) will estimate the current point cloud (PC). c ), referencing (previous) point cloud PC r and the block position B of the current frame c As input, then output the block-by-block motion vector (MV). c ) and by each MV c The confidence level or error generated during the calculation determines its MODE. c The learning-based scene flow estimation method (see “FlowNet3D: Learning Scene Flow in 3D Point Clouds,” Xingyu Liu et al., “In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition,” pp. 529-537, 2019, hereinafter referred to as “Liu,” and “FESTA: Flow Estimation via Spatial-TemporalAttention for Scene Point Clouds,” Haiyan Wang et al., “In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition,” pp. 14173-14182, 2021, hereinafter referred to as “Wang”) is an example of this module.
[0072] The first method (Liu) estimates scene flow from a pair of consecutive PC frames. It introduces a "flow embedding" layer that learns to associate two consecutive point clouds, and a "Set Upconv" layer that learns to propagate features from one PC to another. The second method (Wang) spatially enhances the selection of group query points for better flow correspondence between two frames. Additionally, it temporally improves the point search region used for the flow embedding step.
[0073] The “Reference Block Estimation” module (630) will estimate the MV. c MODE c and B c As input, then output B. r B rIt is the block position in the reference frame corresponding to the current frame. If MODE c If determined to be in inter-frame mode, then B r It is by passing the corresponding B c and MV c It is calculated by addition.
[0074] Figure 7 The diagram illustrates how to calculate B. r Further details for PC. c Frames, the mode of each block is determined by MODE. c Definition. Example block location B c Described as PC c Points within an inter-frame mode block in a frame. The arrow indicates the point in B. c The estimated motion vector MV c Because the current frame (PC) c ) and reference frame (PC) r If the sizes of B and B are the same, then B can be used. c and MV c Move to the reference PC frame. The corresponding reference block position is B. r By B c and MV c The two are added together to calculate. B r This is the output of the ME / D module.
[0075] In another embodiment, the "global motion estimation" module can be added before scene flow estimation.
[0076] The "Global Motion Estimation" solves the following equation: ,in and These are the transposed current and reference 4D points, respectively, with homogeneous coordinates padded with 1s. Matrix It is an unknown global transformation matrix that needs to be approximated.
[0077] Point-to-point or point-to-plane ICP (see Paul J. Besl and Neil D. McKay, “A Method for Registration of 3D Shapes”, PAMI, 1992, hereinafter referred to as “Besl”; Y. Chen and GG Medioni, “Object modelling by registration of multiple range images”, Image and Vision Computing, 10(3), 1992, hereinafter referred to as “Chen”) is an example of such a global transformation. Generally, the ICP algorithm iterates through two steps: 1) from the reference point cloud and using the current change matrix Transformed current point cloud Find the correspondence set K (selecting points that are considered as correspondences between the current frame and the reference frame); and 2) minimize the objective function defined on the correspondence set K. To update the transformation The point-to-point ICP algorithm (Besl) uses an objective function. Furthermore, the point-to-plane ICP algorithm (Chen) uses the objective function. ,in It is the normal vector of the reference point in K.
[0078] The "Scene Flow Estimation" submodule will actually transform the current PC frame. and reference PC frame As input, there are no homogeneous coordinates for these two PC frames.
[0079] Block-level analysis module for inter-frame mode (BA) inter ) like Figure 7 As shown, inter-frame mode (for temporal analysis) and intra-frame mode (for spatial analysis) are transmitted via... Figure 4 The "ME / D" module in the frame is used for selection. The characteristic of an intra-frame mode block is that it is directly selected from the current PC frame (PC). c Extracted and aggregated. On the other hand, the residual features of inter-frame mode blocks are extracted and aggregated from the current frame (PC). c ) and reference frame (PC) r The differences between the scenes are extracted and aggregated. This is done by the "Scene Flow Estimation" module (see [link]). Figure 6 The estimated block-by-block motion vectors are used to locate the prediction block. Aggregate features (f) of reference points within the prediction block are extracted. r ) and the corresponding aggregated features of the current point within the current block (f c ).
[0080] Next, from f via the subtraction module c Subtract f from the middle r To calculate the residual features of inter-frame mode blocks (using f) R (This indicates that...) For more details, Figure 8 The diagram in the middle shows BA. inter Module. Note that in one embodiment, the subtraction module ⊖ is implemented through simple vector subtraction (i.e., f R =f c -f r This is achieved through [the following]. In another embodiment, the subtraction module is implemented using a neural network. It first combines two features f [f]. r and f c The data is cascaded and then passed to a series of convolutional layers. The output of each convolutional layer is the residual feature f. R .
[0081] In particular, block-by-block features f c and f r It is extracted by two PointNet-like feature extractors. For example... Figure 8 As shown, for each block (B) c Or B r Pointwise features are extracted using a shared MLP (810, 815), and then aggregated using max pooling (820, 825). c andfa r ) are aggregated. Then, fa c andfa r The MLP (830) is embedded through block-by-block sharing. These features generated in the current block and the reference block are defined as f, respectively. c and f r By subtracting these features (840), the block-by-block residual feature f is obtained. R Used for frame-level analysis ( Figure 4 (VA module in the middle).
[0082] Block-level composition module Figure 9 The diagram illustrates block-level composition. Point composition method (BS) inter or BS intra The MODE varies depending on the block mode (intra-frame or inter-frame, but not limited to) analyzed during the encoding process. c Decision PC c The pattern of each block, and block position B. c Define the position of each block. For each B c The corresponding upsampling F is given. c These block attributes are assigned to the BS. inter Modules and BS intraBoth modules output the reconstructed current point cloud PC. c .
[0083] For intra-frame mode, block residual synthesis is simply implemented as a series of MLP layers, similar to the implementation in the co-owned application PCT / US2022 / 052861 entitled "Scalable Framework for Point Cloud Compression". Specifically, given the block-by-block features of an intra-frame block, it is fed into the MLP layers for decoding. The MLP layers first generate the 3D coordinates of the center point of the decoded block. These generated 3D coordinates are local coordinates, i.e., they are relative coordinates with respect to the center position of the intra-frame mode block. Therefore, by adding the center position of the intra-frame mode block back to the 3D points generated by the MLP layers, the intra-frame mode block is finally reconstructed on the decoder side.
[0084] BS with block motion compensation inter Module According to one embodiment, Figure 10 The diagram in the middle shows BS. inter Block diagram of the module. Figure 5 The motion prediction module in the middle predicts the reconstructed reference point cloud PC. r The block location is B. r and PC r It can be used as input to extract features from the reconstructed reference frame. Points residing in the predicted block are fed into the "block feature extractor" (1010), which was previously used in... Figure 8 The image is depicted in the image. The extractor extracts block-by-block features from the reference frame. .
[0085] Meanwhile, the encoded and upsampled residual features F c It is fed into the block-by-block MLP (1020) for output. These two features are added together via the summation module ⊕ (1030) to form the current block-wise feature. Note that in one embodiment, the summation module ⊕ is achieved through simple vector addition (i.e., This is achieved through [the following]. In another embodiment, the summation module is implemented using a neural network. It first combines two features... and The features are then cascaded and passed to a series of convolutional layers, with the output of each layer being the feature vector. The current feature can be pooled (1040) to become... This is then fed into the final pointwise MLP (1050) to reconstruct the pointwise local coordinates. These local coordinates corresponding to the reconstructed points in the block are added to each corresponding block location B. c And the current PC c The final position was reconstructed (1060).
[0086] Voxel analysis and voxel synthesis According to one embodiment, Figure 11 The diagram shows a block diagram of the "Voxel Analysis Module" (i.e., VA). The VA module combines the pattern map (indicating which blocks are inter-frame mode blocks and which are intra-frame mode blocks) and the feature map F. c As input. In one embodiment, the pattern graph is a binary graph that uses 0 to indicate inter-frame mode blocks and 1 to indicate intra-frame mode blocks. In another embodiment, it uses 1 to indicate inter-frame mode blocks and 0 to indicate intra-frame mode blocks. Finally, in Figure 11 In the middle, feature map F c The pattern map is concatenated (1110) to obtain an expanded feature map. The expanded feature map is then passed to a series of convolutional layers (1120) for aggregation and downsampling to obtain a downsampled feature map, denoted as F. c2 .
[0087] According to one embodiment, Figure 12 The diagram shows a block diagram of the "voxel synthesis module" (i.e., VS). The VS module combines the pattern map and the downsampled feature map F' c2 As input, it first processes the feature map F' c2 Perform voxel upsampling (1210) to size it relative to the feature map F on the encoder side. c Matching is performed. Then, the output feature map is concatenated with the pattern map (1220) to obtain the augmented feature map. The augmented feature map is then passed to a series of convolutional layers (1230) for aggregation and upsampling to obtain the output feature map F'. c .
[0088] Various numerical values are used in this application. Specific values are for illustrative purposes only, and the aspects described are not limited to these specific values.
[0089] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first," "second," etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of these terms does not imply a modified ordering of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.
[0090] The implementations and aspects described herein can be implemented, for example, as methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features under discussion can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in apparatuses, such as processors, which generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0091] References to "an embodiment" or "an implementation" or "an implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in one embodiment" or "in one implementation" or "in one implementation" appearing in various places throughout this application, and any other variations, do not necessarily all refer to the same embodiment.
[0092] Additionally, this application may relate to "determining" various information segments. Determining information may include one or more of the following: for example, estimation information, calculation information, prediction information, or information retrieved from memory.
[0093] Furthermore, this application may relate to "accessing" various information segments. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0094] Additionally, this application may relate to "receiving" various segments of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., from memory). Further, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0095] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of…”—for example, in the cases of “A / B”, “A and / or B”, and “at least one of A and B”—is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As further examples, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). Those skilled in the art and related fields will appreciate that this can be extended to as many items as possible listed.
[0096] It will be apparent to those skilled in the art that implementations can generate various signals formatted to carry information that can, for example, be stored or transmitted. For example, the information may include instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. For example, formatting may include encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.
Claims
1. A method for decoding point cloud frames in a dynamic point cloud sequence, the method comprising: Decode the motion information of the current point cloud block, wherein the current point cloud block is inter-frame encoded; Based on the motion information of the current point cloud block, a reference point cloud block is obtained in the previously reconstructed point cloud frame to predict the current point cloud block. At least a point-based neural network is used to obtain features representing the reference point cloud block; The residual feature information of the current point cloud block is decoded, and the residual feature information represents the difference between the current point cloud block and the reference point cloud block; Based on the decoded residual feature information and the features representing the reference point cloud block, the features of the current point cloud block are obtained; and At least one other point-based neural network is used to reconstruct the positions of 3D points in the current point cloud block based on the features of the current point cloud block.
2. An apparatus for decoding point cloud frames in a dynamic point cloud sequence, the apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: Decode the motion information of the current point cloud block, wherein the current point cloud block is inter-frame encoded; Based on the motion information of the current point cloud block, a reference point cloud block is obtained in the previously reconstructed point cloud frame to predict the current point cloud block. At least a point-based neural network is used to obtain features representing the reference point cloud block; The residual feature information of the current point cloud block is decoded, and the residual feature information represents the difference between the current point cloud block and the reference point cloud block; Based on the decoded residual feature information and the features representing the reference point cloud block, the features of the current point cloud block are obtained; and At least one other point-based neural network is used to reconstruct the positions of 3D points in the current point cloud block based on the features of the current point cloud block.
3. The method of claim 1, further comprising the steps of, or the apparatus of claim 2, wherein the one or more processors are further configured to perform the following steps: The information is decoded, indicating the location of the occupied block in the point cloud frame.
4. The method of claim 3, further comprising the steps of: or the apparatus of claim 3, wherein the one or more processors are further configured to perform the following steps: The information is decoded, and the information indicates the encoding / decoding mode of the corresponding occupied block in the point cloud frame.
5. The method according to any one of claims 1, 3, and 4, or the apparatus according to any one of claims 2 to 4, wherein obtaining the feature representing the current point cloud block further comprises: The residual feature information and the feature representing the reference point cloud block are concatenated; as well as A series of convolutional layers are applied to the cascaded features to form the features representing the current point cloud block.
6. The method according to any one of claims 1 and 3 to 5, or the apparatus according to any one of claims 2 to 5, wherein the at least one other point-based neural network comprises a series of MLP layers.
7. The method according to any one of claims 1 to 3 to 6, wherein decoding the residual feature information of the current point cloud block comprises the following steps, or the apparatus according to any one of claims 2 to 6, wherein the one or more processors are configured to decode the residual feature information by performing the following steps: Decoding the downsampled feature map, wherein the downsampled feature map is augmented with data indicating whether intra-frame or inter-frame mode was used for the corresponding block in the current point cloud frame; and The downsampled feature map is upsampled and aggregated to obtain the residual feature information.
8. The method according to any one of claims 1 and 3 to 7, or the apparatus according to any one of claims 2 to 7, wherein another block in the current point cloud frame is decoded in intra-frame mode without a reference point cloud frame.
9. A method for encoding point cloud frames in a dynamic point cloud sequence, the method comprising: Based on the current point cloud frame and the previously reconstructed point cloud frame, estimate the motion information of the current point cloud block; Based on the motion information of the current point cloud block, a reference point cloud block is obtained from the previously reconstructed point cloud frame to predict the current point cloud block; At least a point-based neural network is used to obtain a first feature representing the current point cloud block and a second feature representing the reference point cloud block; The difference between the first feature representing the current point cloud block and the second feature representing the reference point cloud block is obtained to form a residual feature; as well as The residual features are encoded.
10. An apparatus for encoding point cloud frames in a dynamic point cloud sequence, the apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: Based on the current point cloud frame and the previously reconstructed point cloud frame, estimate the motion information of the current point cloud block; Based on the motion information of the current point cloud block, a reference point cloud block is obtained from the previously reconstructed point cloud frame to predict the current point cloud block; At least a point-based neural network is used to obtain a first feature representing the current point cloud block and a second feature representing the reference point cloud block; The difference between the first feature representing the current point cloud block and the second feature representing the reference point cloud block is obtained to form a residual feature; as well as The residual features are encoded.
11. The method of claim 9, further comprising the step of, or the apparatus of claim 10, wherein the one or more processors are further configured to perform the following steps: The information is encoded to indicate the location of an occupied block in the point cloud frame.
12. The method of claim 11, further comprising the step of, or the apparatus of claim 11, wherein the one or more processors are further configured to perform the following steps: The information is encoded, and the information indicates the encoding / decoding mode of the corresponding occupied block in the point cloud frame.
13. The method according to any one of claims 9, 11, and 12, or the apparatus according to any one of claims 10 to 12, wherein obtaining the difference comprises: The first feature representing the current point cloud block is concatenated with the second feature representing the reference point cloud block; as well as A series of convolutional layers are applied to the cascaded features to form the differences.
14. The method of any one of claims 9 and 11 to 13, or the apparatus of any one of claims 10 to 13, wherein the first feature representing the current point cloud block and the second feature representing the reference point cloud block are obtained by a shared point-based neural network.
15. The method according to any one of claims 9 and 11 to 14, the method further comprising the step of, or the apparatus according to any one of claims 10 to 14, wherein the one or more processors are further configured to perform the following steps: The feature maps are downsampled and aggregated to form downsampled feature maps, wherein the feature maps are augmented with data indicating whether intra-frame or inter-frame mode is used for the corresponding blocks in the current point cloud frame; and The downsampled feature map is encoded.
16. The method according to any one of claims 9 and 11 to 15, or the apparatus according to any one of claims 10 to 15, wherein another block in the current point cloud frame is encoded in intra-frame mode without a reference point cloud frame.
17. A signal comprising point cloud data, the point cloud data being formed by performing the method according to any one of claims 9 to 16.
18. A computer-readable storage medium having instructions stored thereon for encoding or decoding a point cloud using the method according to any one of claims 1, 3 to 9 and 11 to 16.