Method and device for point cloud processing and medium
By performing sparse downsampling and sparse convolutional network prediction on dynamic point clouds, the problems of low intra-frame prediction accuracy and information redundancy in dynamic point cloud compression are solved, achieving more efficient encoding and decoding results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-10-06
- Publication Date
- 2026-05-05
AI Technical Summary
Existing learning-based dynamic point cloud compression methods have low intra-frame prediction accuracy, resulting in significant redundancy between the current frame and the reference frame information. They fail to fully utilize the information in the encoded frame and ignore the constraints of prediction information, thus affecting the encoding/decoding rate and decoding quality.
By sparsely downsampling the point clouds of the current frame and the reference frame, sparse point clouds and features are obtained. Sparse convolutional networks are used for prediction. By combining multi-stage downsampling and neural network methods, intra-frame and inter-frame prediction is optimized, information redundancy is reduced, and encoding and decoding efficiency is improved.
It improves intra-frame prediction accuracy, reduces information redundancy, enhances encoding and decoding efficiency and decoding quality, and optimizes inter-frame compression performance.
Smart Images

Figure CN121986359A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to point cloud processing techniques, and more specifically, to dynamic point cloud encoding and decoding based on reference frames. Background Technology
[0002] A point cloud is a collection of data points in a three-dimensional (3D) plane, where each point has defined coordinates on the X, Y, and Z axes. Therefore, point clouds can be used to represent the physical content of three-dimensional space. For a wide range of immersive applications, from augmented reality to autonomous vehicles, point clouds have proven to be a promising way to represent 3D visual data.
[0003] Point cloud encoding and decoding standards have largely evolved from the well-known MPEG organization. MPEG stands for Moving Picture Experts Group, one of the main standardization groups for multimedia processing. In 2017, the MPEG 3D Graphics Codec Group (3DG) released a Call for Proposals (CFP) document to begin developing point cloud encoding and decoding standards. The final standard will encompass two categories of solutions. Video-based point cloud compression (V-PCC or VPCC) is suitable for point sets with relatively uniform point distribution. Geometry-based point cloud compression (G-PCC or GPCC) is suitable for sparser distributions. However, the overall expectation is to further improve the encoding and decoding efficiency of conventional point cloud encoding and decoding techniques. Summary of the Invention
[0004] The embodiments of this disclosure provide a solution for point cloud processing.
[0005] In a first aspect, a method for point cloud processing is proposed. The method includes: a conversion between a current point cloud (PC) sample of a point cloud sequence and a bitstream of the point cloud sequence; determining a prediction of current features for the current PC sample based on a current sparse PC sample of the current PC sample, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; and performing the conversion based on the prediction of the current features.
[0006] Based on the method according to the first aspect of this disclosure, the prediction of the current features of the current PC sample is determined based on the current sparse PC samples of the current PC sample, the sparse PC sample set of at least one reference PC sample of the current PC sample, and the feature set. Further, the current PC sample is encoded and decoded based on the prediction of the current features. Compared with conventional solutions, the proposed method can advantageously make full use of the information of the encoded and decoded PC samples to reduce redundancy between the information of the current PC sample and the information of the encoded and decoded PC samples. In this way, encoding and decoding efficiency can be improved.
[0007] In a second aspect, an apparatus for point cloud processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0008] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.
[0009] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence generated by a method performed by an apparatus for point cloud processing. The method includes: determining a prediction of current features for the current PC sample based on current sparse PC samples of the current PC sample in the point cloud sequence, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; and generating a bitstream based on the prediction of the current features.
[0010] In the fifth aspect, a method for storing bitstreams of point cloud sequences is proposed. The method includes: determining a prediction of current features for the current PC sample based on a current sparse PC sample of the current PC sample of the point cloud sequence, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; generating a bitstream based on the prediction of the current features; and storing the bitstream in a non-transitory computer-readable recording medium.
[0011] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0012] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0013] Figure 1 This is a block diagram illustrating an example point cloud encoding / decoding system that can utilize the techniques disclosed herein; Figure 2 A block diagram of an example point cloud encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example point cloud decoder according to some embodiments of the present disclosure is shown; Figure 4 The process of dynamic point cloud compression based on multiple reference frames is shown; Figure 5 A flowchart of a method for point cloud processing according to embodiments of the present disclosure is shown; and Figure 6 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0014] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation
[0015] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0016] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0017] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0018] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0019] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0020] Example Environment Figure 1 This is a block diagram illustrating an example point cloud encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the point cloud encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a point cloud encoding device, and the destination device 120 may also be referred to as a point cloud decoding device. In operation, the source device 110 may be configured to generate encoded point cloud data, and the destination device 120 may be configured to decode the encoded point cloud data generated by the source device 110. The techniques of this disclosure are generally directed to encoding and / or decoding point cloud data, i.e., supporting point cloud compression. Encoding and decoding can be effective in compressing and / or decompressing point cloud data.
[0021] Source device 100 and destination device 120 may include any of a variety of devices, including desktop computers, laptops, tablets, set-top boxes, handsets (such as smartphones and mobile phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land or sea vehicles, spacecraft, aircraft, etc.), robots, LiDAR devices, satellites, extended reality devices, etc. In some cases, source device 100 and destination device 120 may be equipped for wireless communication.
[0022] Source device 100 may include a data source 112, a memory 114, a GPCC encoder 116, and an input / output (I / O) interface 118. Destination device 120 may include an input / output (I / O) interface 128, a GPCC decoder 126, a memory 124, and a data consumer 122. According to this disclosure, the GPCC encoder 116 of source device 100 and the GPCC decoder 126 of destination device 120 may be configured to apply the point cloud encoding / decoding techniques of this disclosure. Therefore, source device 100 represents an example of an encoding device, and destination device 120 represents an example of a decoding device. In other examples, source device 100 and destination device 120 may include other components or arrangements. For example, source device 100 may receive data (e.g., point cloud data) from an internal or external source. Similarly, destination device 120 may interface with an external data consumer rather than including the data consumer in the same device.
[0023] Generally, data source 112 represents a source of point cloud data (i.e., raw, unencoded point cloud data) and can provide a continuous series of "frames" of point cloud data to GPCC encoder 116, which encodes the point cloud data for each frame. In some examples, data source 112 generates point cloud data. The data source 112 of source device 100 may include point cloud acquisition devices, such as any of various cameras or sensors, such as one or more cameras, an archive containing previously acquired point cloud data, a 3D scanner or light detection and ranging (LIDAR) device, and / or a data feed interface that receives point cloud data from a data content provider. Thus, in some examples, data source 112 may generate point cloud data based on signals from a LIDAR device. Alternatively or additionally, point cloud data may be generated by a computer from scanners, cameras, sensors, or other data. For example, data source 112 may generate point cloud data, or produce a combination of real-time point cloud data, archived point cloud data, and computer-generated point cloud data. In each case, the GPCC encoder 116 encodes the acquired, pre-acquired, or computer-generated point cloud data. The GPCC encoder 116 can rearrange the frames of the point cloud data from the receiving order (sometimes referred to as the "display order") to an encoding / decoding order for encoding and decoding. The GPCC encoder 116 can generate one or more bitstreams comprising the encoded point cloud data. The source device 100 can then output the encoded point cloud data via I / O interface 118 for reception and / or retrieval by, for example, the I / O interface 128 of the destination device 120. The encoded point cloud data can be directly transmitted to the destination device 120 via I / O interface 118 through network 130A. The encoded point cloud data can also be stored on storage medium / server 130B for access by the destination device 120.
[0024] The memory 114 of the source device 100 and the memory 124 of the destination device 120 may represent general-purpose memory. In some examples, memory 114 and memory 124 may store raw point cloud data, such as raw point cloud data from data source 112 and raw, decoded point cloud data from GPCC decoder 126. Additionally or alternatively, memory 114 and memory 124 may store software instructions executable by, for example, GPCC encoder 116 and GPCC decoder 126. Although memory 114 and memory 124 are shown separately from GPCC encoder 116 and GPCC decoder 126 in this example, it should be understood that GPCC encoder 116 and GPCC decoder 126 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memory 114 and memory 124 may store encoded point cloud data, such as encoded point cloud data output from GPCC encoder 116 and input to GPCC decoder 126. In some examples, portions of memory 114 and memory 124 may be allocated as one or more caches, for example, to store raw point cloud data, decoded and / or encoded point cloud data. For example, memory 114 and memory 124 may store point cloud data.
[0025] I / O interfaces 118 and 128 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where I / O interfaces 118 and 128 include wireless components, they may be configured to transmit data, such as encoded point cloud data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc. In some examples where I / O interface 118 includes a wireless transmitter, they may be configured to transmit data, such as encoded point cloud data, according to other wireless standards such as the IEEE 802.11 specification. In some examples, source device 100 and / or destination device 120 may include corresponding system-on-chip (SoC) devices. For example, source device 100 may include a SoC device for performing functions belonging to GPCC encoder 116 and / or I / O interface 118, and destination device 120 may include a SoC device for performing functions belonging to GPCC decoder 126 and / or I / O interface 128.
[0026] The techniques disclosed herein can be applied to encoding and decoding to support any of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors and processing devices (e.g., local or remote servers), geographic mapping, or other applications.
[0027] The I / O interface 128 of the destination device 120 receives an encoded bitstream from the source device 110. The encoded bitstream may include signaling information defined by the GPCC encoder 116, which is also used by the GPCC decoder 126, such as syntax elements having values representing the point cloud. The data consumer 122 uses the decoded data. For example, the data consumer 122 may use the decoded point cloud data to determine the location of physical objects. In some examples, the data consumer 122 may include a display for presenting images based on the point cloud data.
[0028] The GPCC encoder 116 and GPCC decoder 126 can each be implemented as any of a variety of suitable encoder circuitry and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the GPCC encoder 116 and GPCC decoder 126 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the GPCC encoder 116 and / or GPCC decoder 126 may include one or more integrated circuits, microprocessors, and / or other types of devices.
[0029] The GPCC encoder 116 and GPCC decoder 126 can operate according to encoding / decoding standards such as the Video Point Cloud Compression (VPCC) standard or the Geometric Point Cloud Compression (GPCC) standard. Generally, this disclosure may refer to the encoding and decoding of frames (e.g., encoding and decoding) to include the process of encoding or decoding data. Encoded bitstreams typically include a series of values for syntax elements representing encoding / decoding decisions (e.g., encoding / decoding modes).
[0030] A point cloud can contain a set of points in 3D space and can have attributes associated with those points. Attributes can be color information, such as R, G, B or Y, Cb, Cr, or reflectivity information, or other attributes. Point clouds can be acquired by various cameras or sensors, such as LiDAR sensors and 3D scanners, and can also be computer-generated. Point cloud data is used in a variety of applications, including but not limited to architecture (modeling), graphics (3D models for visualization and animation), and the automotive industry (LiDAR sensors for navigation aids).
[0031] Figure 2 This is a block diagram illustrating an example of a GPCC encoder 200 according to some embodiments of the present disclosure. The GPCC encoder 200 may be... Figure 1 An example of a GPCC encoder 116 in system 100 is shown. Figure 3 This is a block diagram illustrating an example of a GPCC decoder 300 according to some embodiments of the present disclosure. The GPCC decoder 300 may be... Figure 1 An example of the GPCC decoder 126 in the system 100 shown.
[0032] In both the GPCC encoder 200 and GPCC decoder 300, point cloud locations are encoded and decoded first. Attribute encoding and decoding depend on the decoded geometry. Figure 2 and Figure 3 In this configuration, Region Adaptive Hierarchical Transformation (RAHT) unit 218, Surface Approximation Analysis unit 212, RAHT unit 314, and Surface Approximation Synthesis unit 310 are options typically used for Category 1 data. Level of Detail (LOD) Generation unit 220, Lifting unit 222, LOD Generation unit 316, and Inverse Lifting unit 318 are options typically used for Category 3 data. All other units are common between Category 1 and Category 3.
[0033] For Category 3 data, the compressed geometry is typically represented as an octree from the root down to the leaf level of each voxel. For Category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than voxels) plus a model for approximating the surface within each leaf node of the pruned octree. In this way, both Category 1 and Category 3 data share the octree encoding / decoding mechanism, while Category 1 data can additionally utilize the surface model to approximate the voxels within each leaf node. The surface model used is a triangulation of each block comprising 1 to 10 triangles, producing a triangle soup. Therefore, the Category 1 geometry codec is called a triangle soup geometry codec, while the Category 3 geometry codec is called an octree geometry codec.
[0034] exist Figure 2 In the example, the GPCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometric reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.
[0035] like Figure 2 As shown in the example, the GPCC encoder 200 can receive a set of locations and a set of attributes. Locations can include the coordinates of points in the point cloud. Attributes can include information about the points in the point cloud, such as the colors associated with those points.
[0036] The coordinate transformation unit 202 can apply transformations to the coordinates of a point to transform the coordinates from the initial domain to the transformation domain. The transformed coordinates can be referred to as transformed coordinates. The color transformation unit 204 can apply transformations to convert the color information of an attribute to different domains. For example, the color transformation unit 204 can convert color information from the RGB color space to the YCbCr color space.
[0037] In addition, Figure 2 In the example, voxelization unit 206 can voxelize the transformed coordinates. Voxelization of the transformed coordinates can include quantization and removal of some points in the point cloud. In other words, multiple points in the point cloud can be grouped into a single "voxel," which can then be considered a point in some respects. Furthermore, octree analysis unit 210 can generate an octree based on the voxelized transformed coordinates. Additionally, in Figure 2 In the example, the surface approximation analysis unit 212 can analyze points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 can perform arithmetic coding on syntax elements representing information about an octree and / or information about the surface determined by the surface approximation analysis unit 212. The GPCC encoder 200 can output these syntax elements in a geometric bitstream.
[0038] The geometric reconstruction unit 216 can reconstruct the transformed coordinates of points in the point cloud based on an octree, data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometric reconstruction unit 216 may differ from the original number of points in the point cloud. The resulting points may be referred to as reconstructed points. The attribute transfer unit 208 can transfer attributes of the original points in the point cloud to the reconstructed points in the point cloud data.
[0039] Furthermore, RAHT unit 218 can apply RAHT encoding to the attributes of the reconstructed points. Alternatively or additionally, LOD generation unit 220 and lifting unit 222 can apply LOD processing and lifting to the attributes of the reconstructed points, respectively. RAHT unit 218 and lifting unit 222 can generate coefficients based on the attributes. Coefficient quantization unit 224 can quantize the coefficients generated by RAHT unit 218 or lifting unit 222. Arithmetic encoding unit 226 can apply arithmetic encoding to the syntax elements representing the quantized coefficients. GPCC encoder 200 can output these syntax elements in the attribute bitstream.
[0040] exist Figure 3 In the example, the GPCC decoder 300 may include a geometric arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometric reconstruction unit 312, a RAHT unit 314, an LOD generation unit 316, an inverse lifting unit 318, a coordinate inverse transformation unit 320, and a color inverse transformation unit 322.
[0041] The GPCC decoder 300 can acquire geometric bitstreams and attribute bitstreams. The geometric arithmetic decoding unit 302 of the decoder 300 can apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to the syntax elements in the geometric bitstream. Similarly, the attribute arithmetic decoding unit 304 can apply arithmetic decoding to the syntax elements in the attribute bitstream.
[0042] Octree synthesis unit 306 can synthesize octrees based on syntax elements parsed from the geometric bitstream. In the case of using surface approximation in the geometric bitstream, surface approximation synthesis unit 310 can determine the surface model based on syntax elements parsed from the geometric bitstream and based on the octree.
[0043] Furthermore, the geometric reconstruction unit 312 can perform reconstruction to determine the coordinates of points in the point cloud. The inverse coordinate transformation unit 320 can apply an inverse transformation to the reconstructed coordinates to transform the reconstructed coordinates (positions) of points in the point cloud from the transformation domain back to the initial domain.
[0044] Additionally, in Figure 3 In the example, the dequantization unit 308 can dequantize the attribute value. The attribute value can be based on syntax elements obtained from the attribute bitstream (e.g., including syntax elements decoded by the attribute arithmetic decoding unit 304).
[0045] Depending on how the attribute values are encoded, RAHT unit 314 can perform RAHT decoding to determine the color value for a point in the point cloud based on the dequantized attribute values. Alternatively, LOD generation unit 316 and inverse boosting unit 318 can use level-of-detail (LOD) based techniques to determine the color value for a point in the point cloud.
[0046] In addition, Figure 3 In the example, the color inverse transformation unit 322 can apply an inverse color transformation to color values. The inverse color transformation can be the inverse of the color transformation applied by the color transformation unit 204 of the encoder 200. For example, the color transformation unit 204 can transform color information from the RGB color space to the YCbCr color space. Correspondingly, the color inverse transformation unit 322 can transform color information from the YCbCr color space to the RGB color space.
[0047] Figure 2 and Figure 3 Various units are shown to aid in understanding the operations performed by encoder 200 and decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit is a circuit that provides a specific function and is preset with respect to the operations that can be performed. A programmable circuit is a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by instructions in the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.
[0048] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, while some embodiments are described with reference to GPCC or other specific point cloud codecs, the disclosed techniques are also applicable to other point cloud codec techniques. Additionally, although some embodiments describe point cloud codec steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. 1. Summary of the Invention This disclosure relates to dynamic point cloud encoding and decoding techniques. Specifically, it relates to learning-based dynamic point cloud geometry compression. This idea can be combined with point cloud encoding and decoding standards, such as geometry-based point cloud compression (G-PCC), which is under development.
[0050] 2. Abbreviation G-PCC is a geometry-based point cloud compression technology. MPEG Moving Picture Experts Group 3. Introduction In point cloud compression, traditional octree encoding / decoding, mesh encoding / decoding, mapping encoding / decoding, and attribute encoding / decoding provide the basic ideas and framework for compression. Following their encoding principles and modular structures, various signal processing methods are used to design new modules or optimize and enhance existing ones. The same applies to learning-based point cloud compression, which can use neural network models to replace traditional modules and optimize model parameters based on data-driven optimization.
[0051] Currently, learning-based static point cloud compression significantly outperforms traditional compression methods, thus the goal is to migrate learning-based methods to dynamic point cloud compression. Similar to static compression, dynamic compression also needs to retain as much information about the point cloud as possible in each frame while maintaining the required bit rate. However, unlike static point cloud compression, dynamic point cloud compression needs to consider not only intra-frame compression within a single frame but also inter-frame compression, which is the core of dynamic compression. Intra-frame prediction uses information from a reference frame to predict the information in the current frame and is a crucial component of inter-frame compression. The more accurate the intra-frame prediction, the more efficient the intra-frame compression will be, and the lower the bit rate consumed per frame will be. Considering the similar correlation between neighboring frames, most dynamic point cloud compression methods use the previous frame as a reference frame to predict the current frame, and all of these methods have achieved good prediction performance.
[0052] 3.1 Sparse Convolution To leverage the sparsity of point clouds, researchers have explored various approaches, such as octree-based CNNs and sparse CNNs. In sparse CNNs, the data tensor is represented by a set of coordinates C and associated features F. Convolution only aggregates the features occupying the coordinates. It is defined as:
[0053] in These are the input coordinates and the output coordinates. Coordinates The input feature vector and the output feature vector at the location. Define 3D convolution kernels, covering... The set of locations centered on, where It has an offset. Indicates offset The kernel value at the point cloud. This sparse convolution leverages the sparsity of the point cloud to reduce complexity and performs computation only on the voxels that are currently occupied.
[0054] 4. Question Existing learning-based dynamic point cloud compression methods have the following problems: 1. Current dynamic point cloud compression methods have low intra-frame prediction accuracy, resulting in significant redundancy between the current frame information to be transmitted and the reference frame information.
[0055] 2. Most dynamic point cloud compression methods do not make full use of the information in the encoded frame, which reduces the available information for reference and results in redundancy between the information of the current frame to be transmitted and the information of the encoded frame.
[0056] 3. Current learning-based dynamic point cloud compression methods only consider the balance between the residual information of the current frame and the bit rate of the prediction information and the quality of the decoded current frame, while ignoring the constraints of the prediction information. This may affect the encoding / decoding rate and reduce the decoding quality of the current frame.
[0057] 5. Detailed Solution To address the issues mentioned above and other unmentioned problems, the methods outlined below are disclosed. The solutions should be considered as examples for explaining general concepts, not as narrow interpretations. Furthermore, these solutions can be applied individually or in combination in any way.
[0058] In the following discussion, the term "encoder" refers to a model that encodes information to be transmitted via a signal. The term "decoder" refers to a model that decodes compressed bits to obtain the information transmitted via a signal.
[0059] 1) A method is proposed to downsample the point cloud of the current frame to obtain the real sparse point cloud and corresponding real features of the current frame.
[0060] a. In one example, traditional non-learning downsampling methods can be used to obtain sparse point clouds and corresponding features.
[0061] i. In one example, the farthest point sampling can be used to obtain a sparse point cloud and its corresponding features.
[0062] ii. In one example, uniform sampling can be used to obtain sparse point clouds and corresponding features.
[0063] b. In one example, a learning-based downsampling method can be used to obtain sparse point clouds and their corresponding features.
[0064] i. In one example, a neural network-based downsampling method can be used to obtain sparse point clouds and their corresponding features.
[0065] 1. In one example, sparse point clouds and their corresponding features can be obtained through a multi-stage downsampling network.
[0066] a. In one example, the number of stages can be N. For example, N equals 3.
[0067] 2. In one example, sparse point clouds and corresponding features from any stage of a multi-stage downsampling network can be output.
[0068] a. In one example, the sparse point cloud and corresponding features of the last stage in a multi-stage downsampling network can be used as the true sparse point cloud and corresponding true features of the current frame.
[0069] 3. In one example, sparse point clouds and corresponding features from all stages of a multi-stage downsampling network can be output.
[0070] 4. In one example, sparse convolution can be used as a basic operation in convolutional networks.
[0071] a. In one example, the stride of a sparse convolution can be N. For example, N equals 2.
[0072] b. In one example, N can be predefined.
[0073] c. In one example, N can be transmitted via a signal.
[0074] 2) A method is proposed to downsample the point cloud of the reference frame to obtain the reference sparse point cloud and reference features.
[0075] a. In one example, the reference frame point cloud can be a frame point cloud preceding the current frame point cloud or multiple frame point clouds.
[0076] i. In one example, the reference frame point cloud could be the point clouds of the two frames preceding the current frame.
[0077] b. In one example, the reference sparse point cloud and reference features can be obtained using a learning-based downsampling method.
[0078] i. In one example, the reference sparse point cloud and reference features can be obtained using a neural network-based downsampling method.
[0079] 1. In one example, the reference sparse point cloud and reference features can be obtained through a multi-stage downsampling network.
[0080] a. In one example, the number of stages can be N. For example, N equals 3.
[0081] 2. In one example, the reference sparse point cloud and reference features of any stage in the multi-stage downsampling network can be output.
[0082] 3. In one example, reference sparse point clouds and reference features from any stage in a multi-stage downsampling network can be output.
[0083] a. In one example, the reference sparse point cloud and reference features of the last two stages in a multi-stage downsampling network can be used as the reference sparse point cloud and reference features of the reference frame.
[0084] 4. In one example, reference sparse point clouds and reference features for all stages in a multi-stage downsampling network can be output.
[0085] 5. In one example, sparse convolution can be used as a basic operation in convolutional networks.
[0086] a. In one example, the stride of a sparse convolution can be N. For example, N equals 2.
[0087] b. In one example, N can be predefined.
[0088] c. In one example, N can be transmitted via a signal.
[0089] 3) It proposes using reference sparse point clouds and reference features to predict the features of the current frame.
[0090] a. In one example, the sparse point cloud and features of a reference frame are used to predict the features of the current frame.
[0091] i. In one example, the features of the current frame can be predicted using the sparse point cloud and features from the last stage of the reference frame.
[0092] 1. In one example, the features of the current frame can be directly predicted using the sparse point cloud and features of the last stage of the reference frame.
[0093] 2. In one example, the sparse point cloud and features of the last stage of the reference frame can be used to perform a refined secondary prediction on the features of the current frame.
[0094] a. In one example, the sparse point cloud and features from the last stage of the reference frame can be used to make initial predictions about the features of the current frame.
[0095] b. In one example, the first predicted feature of the current frame and the feature of the reference frame can be used to perform a secondary prediction on the feature of the current frame.
[0096] c. In one example, a neural network-based approach can be used to predict the features of the current frame.
[0097] i. In one example, sparse convolution and sparse convolution on the target coordinates can be used as basic operations in convolutional networks.
[0098] 3. In one example, a neural network-based prediction method can be used to predict the features of the current frame.
[0099] ii. In one example, the multi-stage sparse point cloud and features of the reference frame can be used to predict the features of the current frame.
[0100] 1. In one example, the multi-stage sparse point cloud and features of the reference frame can be downsampled to align with the next stage point cloud and fused with the aligned features until downsampling to the last stage.
[0101] 2. In one example, the prediction of the current frame features can be accomplished by using the fused last-stage features of the reference frame.
[0102] 3. In one example, a neural network-based method can be used to perform feature prediction for the current frame.
[0103] a. In one example, sparse convolution and sparse convolution on the target coordinates can be used as basic operations in convolutional networks.
[0104] b. In one example, sparse point clouds and features from multiple reference frames can be used to predict features of the current frame.
[0105] i. In one example, the sparse point cloud and features of two reference frames can be used to predict the features of the current frame.
[0106] 1. In one example, the sparse point cloud and features of the last stage of the reference frame can be used to directly predict the features of the current frame.
[0107] a. In one example, the sparse point cloud and features of the first reference frame can be fused with the sparse point cloud and features of the second reference frame.
[0108] b. In one example, the fused features can be used to predict the features of the current frame.
[0109] 2. In one example, the sparse point cloud and features of the last stage of the reference frame can be used to perform a refined secondary prediction on the features of the current frame.
[0110] a. In one example, the features of the current frame can be predicted directly using the first reference frame.
[0111] b. In one example, a second reference frame can be used to directly predict the features of the current frame.
[0112] c. In one example, features predicted using two reference frames can be fused.
[0113] d. In one example, the fused features can be used to predict the current frame again.
[0114] 3. In one example, the multi-stage sparse point cloud and features of the reference frame can be used to predict the features of the current frame.
[0115] a. In one example, the multi-stage sparse point cloud and features of the reference frame can be downsampled to align with the next stage point cloud and fused with the aligned features until downsampling to the last stage.
[0116] b. In one example, the prediction of the current frame features can be accomplished by using the fused last-stage features of the reference frame.
[0117] c. In one example, a neural network-based approach can be used to predict the features of the current frame.
[0118] 4) The method proposes to encode and decode the sparse point cloud of the current frame and transmit it to the decoder via signal.
[0119] a. In one example, the sparse point cloud of the current frame can be encoded and decoded using a point cloud codec.
[0120] i. In one example, the point cloud codec could be G-PCC, V-PCC, Draco, etc.
[0121] 5) The method proposes to encode and decode the residual between the predicted features and the true features of the current frame and transmit it to the decoder via a signal.
[0122] a. In one example, features can be encoded or decoded using fixed-length encoding / decoding, unary encoding / decoding, rounding unary encoding / decoding, etc.
[0123] b. In one example, features can be encoded and decoded in a predictive manner.
[0124] 6) A method is proposed to reconstruct the point cloud of the current frame based on the features and coordinates of the sparse point cloud obtained in the current frame.
[0125] a. In one example, the accurate features of the current frame can be obtained by adding the residuals of the features predicted based on the reference frame and the features decoded by the decoder.
[0126] b. In one example, the accurate features acquired in the current frame can be used for upsampling reconstruction.
[0127] c. In one example, the reconstructed point cloud of the current frame can be obtained directly through a single upsampling.
[0128] d. In one example, the reconstructed point cloud of the current frame can be directly obtained through multiple progressive upsamplings.
[0129] i. In one example, N upsampling operations can be used to reconstruct the point cloud of the current frame, such as N = 3.
[0130] 1. In one example, N can be predefined.
[0131] 2. In one example, N can be transmitted to the decoder through a signal.
[0132] a. In one example, fixed-length coding and decoding, unary coding and decoding, truncated unary coding and decoding, etc. can be used to encode and decode N.
[0133] b. In one example, N can be encoded and decoded in a predictive manner.
[0134] [[ID=!]]e. In one example, generative convolution based on sparse convolution can be used to implement point cloud upsampling.
[0135] f. In one example, during the training of upsampling, multi-stage loss functions with different granularities can be used to constrain the neural network.
[0136] i. In one example, binary cross-entropy value can be used as the loss function in the first stage.
[0137] ii. In one example, in different stages, the number of points used in the loss function can be different.
[0138] iii. In one example, the number of points used in the loss function of each stage can be indicated by at least one indicator.
[0139] 1. In one example, there are 3 stages and 3 indicators, M, N, and K.
[0140] a. In one example, M% of the points of the point cloud obtained by voxel sampling from the real point cloud can be used to constrain the reconstructed point cloud in the first stage.
[0141] b. In one example, N% of the points of the point cloud obtained by voxel sampling from the real point cloud can be used to constrain the reconstructed point cloud in the first stage.
[0142] c. In one example, K% of the points of the point cloud obtained by voxel sampling from the real point cloud can be used to constrain the reconstructed point cloud in the last stage.
[0143] d. In one example, M < N < K, such as M = 12.5, N = 50, K = 100.
[0144] iv. In one example, the instruction can be predefined.
[0145] v. In one example, an instruction can be transmitted via a signal.
[0146] 7) Whether and / or how the methods disclosed above can be applied to transmit signals from the encoder to the decoder in a bitstream / frame / slice / strip / octree / etc.
[0147] 8) Whether and / or how to apply the methods disclosed above may depend on the encoded / decoded information, such as dimensions, color format, color components, and strip / image type.
[0148] 6. Example The following is an example of the workflow of a dynamic point cloud compression method based on multiple reference frames. First, the high-resolution original point clouds of the current frame and the reference frames are processed through a learnable adaptive downsampling network to obtain sparse point clouds and corresponding features. Second, the features of the current frame are predicted using the acquired features of the reference frames. Then, an octree lossless encoder-decoder and an entropy encoder-decoder compress the sparse point cloud coordinates of the current frame and the residual between the true and predicted features of the current frame, respectively. Finally, the current frame is upsampled and reconstructed using the decoded sparse point cloud coordinates and the corrected features.
[0149] Figure 4 The image shows an example of an encoding / decoding process for dynamic point cloud compression based on multiple reference frames. Figure 4 The process of dynamic point cloud compression based on multiple reference frames is shown.
[0150] Further details of embodiments of this disclosure, relating to dynamic point cloud encoding and decoding based on one or more reference frames, will now be described. The embodiments of this disclosure should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments may be applied individually or in any combination.
[0151] As used herein, the term "point cloud sequence" can refer to a sequence of one or more point clouds. The term "point cloud frame" or "frame" can refer to a point cloud within a point cloud sequence. The term "point cloud (PC) sample" can refer to a frame, a sub-region within a frame, an image, a slice, a subframe, a sub-image, a piece, a segment, or any other suitable processing unit.
[0152] Figure 5A flowchart of a method 500 for point cloud processing according to some embodiments of the present disclosure is shown. Method 500 can be implemented during the conversion between a current PC sample of a point cloud sequence and a bitstream of the point cloud sequence. At 502, a prediction for the current features of the current PC sample is determined based on a current sparse PC sample of the current PC sample, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set. For example, the at least one reference PC sample may have been encoded / decoded prior to the current PC sample.
[0153] In some embodiments, the features of the PC sample correspond to sparse PC samples and can be generated by performing a downsampling process on the PC sample. The number of points in the sparse PC sample is less than the number of points in the PC sample. This will be described in detail below.
[0154] In some embodiments, a first model based on a neural network (NN) can be used to determine a prediction for the current feature. As used herein, the term "model" refers to the relationship between inputs and outputs learned from training data, thus enabling the generation of a corresponding output for a given input after training. The generation of the model can be based on machine learning techniques, such as neural network techniques. Typically, an NN-based model can be built that receives input information and makes predictions based on that input information.
[0155] For example, the first model based on NN may include at least one sparse convolution, at least one sparse convolution on the target coordinates, etc.
[0156] At position 504, a transformation is performed based on the prediction for the current feature. In some embodiments, the transformation may include encoding the current PC sample into a bitstream. In this case, the residual between the current feature and the prediction for the current feature can be determined as the difference between the current feature and the prediction. The residual and the current sparse PC sample can be encoded into the bitstream.
[0157] Alternatively or additionally, the transformation may include decoding the current PC sample from the bitstream. In this case, the residual between the current feature and the prediction for the current feature can be decoded from the bitstream, and the current sparse PC sample can be decoded from the bitstream. Furthermore, the current PC sample can be reconstructed based on the prediction and the current sparse PC sample. This will be described in detail below.
[0158] Based on the above, a prediction of the current features for the current PC sample is determined based on the current sparse PC sample, the sparse PC sample set of at least one reference PC sample of the current PC sample, and the feature set. Furthermore, the current PC sample is encoded and decoded based on the prediction of the current features. Compared with traditional solutions, the proposed method can effectively utilize the information of the encoded and decoded PC samples to reduce redundancy between the information of the current PC sample and the information of the encoded and decoded PC samples. In this way, encoding and decoding efficiency can be improved.
[0159] In some embodiments, at least one reference PC sample may include a single reference PC sample. In one example embodiment, a downsampling process with at least one stage may be applied to a single reference PC sample. The sparse PC sample set may include sparse PC samples generated at the last stage of the single reference PC sample in at least one stage, and the feature set may include features generated at the last stage of the single reference PC sample.
[0160] In one example, a prediction for the current feature can be determined by directly using the current sparse PC sample of the current PC sample, the sparse PC sample of a single reference PC sample, and the feature. For example, the sparse PC sample of a single reference PC sample and the feature can be input into a neural network-based model, and the neural network-based model outputs a prediction for the current feature.
[0161] Alternatively, a refined secondary prediction process can be performed based on the current sparse PC samples of the current PC samples, the sparse PC samples of a single reference PC sample, and the features to determine the prediction for the current feature. As an example, and not a limitation, an initial prediction for the current feature can be generated based on the current sparse PC samples of the current PC samples, the sparse PC samples of a single reference PC sample, and the features. Furthermore, a secondary prediction for the current feature can be generated based on the initial prediction and the features of the single reference PC sample to obtain the final prediction for the current feature. The secondary prediction can be viewed as a refinement of the initial prediction. In this way, the quality of the prediction can be improved.
[0162] In another example embodiment, a downsampling process with multiple stages can be applied to a single reference PC sample. The sparse PC sample set may include more than one sparse PC sample generated at more than one stage in the multiple stages of the single reference PC sample, and the feature set may include more than one feature generated at more than one stage of the single reference PC sample.
[0163] In one example, sparse PC samples and features generated in a first stage (more than one stage) from a single reference PC sample can be downsampled. The downsampled sparse PC samples and features can then be aligned and fused with sparse PC samples and features generated in a second stage (more than one stage) from the same reference PC sample. The second stage follows the first stage. Through alignment, a mapping is established between points in the downsampled sparse PC samples and points in the sparse PC samples for the second stage, and a mapping is established between elements in the downsampled features and elements in the features for the second stage. Subsequently, points in the downsampled sparse PC samples can be fused with points in the sparse PC samples for the second stage to obtain fused sparse PC samples. Similarly, elements in the downsampled features can be fused with elements in the features for the second stage to obtain fused features. The fusion operation can be implemented using accumulation, weighted sum, weighted average, etc.
[0164] Furthermore, predictions for the current features are determined based on the current sparse PC samples of the current PC samples, the fused sparse PC samples of a single reference PC sample at the last stage in more than one stage, and the fused features.
[0165] In some other embodiments, at least one reference PC sample may include multiple reference PC samples. For example, the number of multiple reference PC samples may be two, three, etc. In one example embodiment, a downsampling process with at least one stage may be applied to multiple reference PC samples. The sparse PC sample set may include sparse PC samples generated at the last stage of at least one stage of multiple reference PC samples, and the feature set may include multiple features generated at the last stage of multiple reference PC samples.
[0166] In one example, a prediction for the current feature can be determined by directly using the current sparse PC sample of the current PC sample, the sparse PC samples of multiple reference PC samples, and the features. As an example, and not a limitation, the sparse PC samples of multiple reference PC samples can be fused, and the features of multiple reference PC samples can also be fused. A prediction for the current feature can be generated based on the current sparse PC sample of the current PC sample, the fused sparse PC samples, and the fused features.
[0167] In another example, a refined secondary prediction process can be performed based on the current sparse PC sample of the current PC sample, sparse PC samples of multiple reference PC samples, and features to determine a prediction for the current feature. Assume the multiple reference PC samples include a first reference PC sample and a second reference PC sample. A first prediction for the current feature can be generated based on the current sparse PC sample of the current PC sample, the sparse PC samples of the first reference PC sample, and features. A second prediction for the current feature can be generated based on the current sparse PC sample of the current PC sample, the sparse PC samples of the second reference PC sample, and features. Furthermore, a prediction for the current feature can be generated based on the result of fusing the first and second predictions. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0168] In some other embodiments, a downsampling process with multiple stages can be applied to multiple reference PC samples. The sparse PC sample set may include more than one sparse PC sample generated at more than one stage of the multiple reference PC samples, and the feature set may include more than one feature generated at more than one stage of the multiple reference PC samples.
[0169] For example, sparse PC samples and features generated in the first stage of more than one stage from multiple reference PC samples can be downsampled. The downsampled sparse PC samples and downsampled features can then be aligned and fused with sparse PC samples and features generated in the second stage of more than one stage from multiple reference PC samples. The second stage follows the first stage. The above operation can be repeated for more than one stage, and a prediction for the current feature can be determined based on the current sparse PC sample of the current PC sample, the fused sparse PC samples at the last stage of more than one stage, and the fused features.
[0170] In some embodiments, method 500 may further include: obtaining the current sparse point cloud and current features by performing a first downsampling process on the current PC samples. In one example embodiment, the first downsampling process may not be based on machine learning. For example, the first downsampling process may include farthest-distance point sampling, uniform sampling, etc.
[0171] In another example embodiment, the first downsampling process may be based on machine learning. For example, the first downsampling process may utilize a second neural network-based model. In some embodiments, the first downsampling process may include multiple stages. For example, the number of stages may be 2, 3, 4, etc.
[0172] In some embodiments, sparse point clouds and features generated at any stage of multiple stages for the current PC sample can be output. Additionally or alternatively, sparse point clouds and features generated at all stages of multiple stages for the current PC sample can be output. For example, the current sparse point cloud can be determined as the sparse point cloud generated at the last stage of multiple stages for the current PC sample, and the current feature can be determined as the feature generated at the last stage for the current PC sample.
[0173] In some embodiments, the NN-based second model may include at least one sparse convolution. For example, the stride of the at least one sparse convolution may be 2, 3, etc. The stride of the at least one sparse convolution may be predetermined or indicated in the bitstream.
[0174] In some embodiments, method 500 may further include: obtaining a sparse set of PC samples and a feature set of at least one reference PC sample by performing a second downsampling process on at least one reference PC sample. As an example and not a limitation, the second downsampling process may be based on machine learning. For example, the second downsampling process may utilize a third NN-based model.
[0175] In some embodiments, the second downsampling process may include multiple stages. For example, the number of stages may be 2, 3, 4, etc. In one example embodiment, sparse point clouds and features generated at one or more of the multiple stages for at least one reference PC sample may be output. Alternatively, sparse point clouds and features generated at all of the multiple stages for at least one reference PC sample may be output.
[0176] In some embodiments, the sparse PC sample set of at least one reference PC sample may include sparse PC samples generated at the last two stages of a plurality of stages from at least one reference PC sample. Additionally, the feature set of at least one reference PC sample may include features generated at the last two stages from at least one reference PC sample.
[0177] In some embodiments, the NN-based third model may include at least one sparse convolution. For example, the stride of the at least one sparse convolution may be 2. The stride of the at least one sparse convolution may be predetermined or indicated in the bitstream.
[0178] In some embodiments, the current sparse PC sample can be encoded into a bitstream. Additionally or alternatively, the current sparse PC sample can be decoded from the bitstream. For example, the current sparse PC sample can be encoded and decoded using a point cloud codec. The point cloud codec can be based on geometry-based point cloud compression (G-PCC), video-based point cloud compression (V-PCC), Draco, or similar technologies.
[0179] In some embodiments, the residual between the current feature and the prediction for the current feature can be encoded into a bitstream. Additionally or alternatively, the residual can be decoded from the bitstream. For example, the residual can be encoded or decoded using fixed-length encoding / decoding, unary encoding / decoding, or rounded unary encoding / decoding. The residual can be encoded or decoded in a predictive manner.
[0180] In some embodiments, at 504, the current PC sample can be reconstructed based on the prediction for the current feature and the coordinates of the current sparse PC sample. For example, the current feature can be obtained by adding the prediction for the current feature and the residual obtained from the bitstream. Alternatively, an upsampling process can be performed on the current sparse PC sample and the current feature to reconstruct the current PC sample.
[0181] In some embodiments, the upsampling process may include a single upsampling operation. Alternatively, the upsampling process may include multiple upsampling operations. For example, the number of multiple upsampling operations may be 2, 3, 5, etc. In one example, the number of multiple upsampling operations may be predetermined. In another example, the number of multiple upsampling operations may be indicated in the bitstream. For example, the number of multiple upsampling operations may be encoded or decoded using fixed-length encoding / decoding, unary encoding / decoding, or rounding unary encoding / decoding. The number of multiple upsampling operations may be encoded or decoded predictively.
[0182] In some embodiments, the upsampling process can be implemented using a neural network-based fourth model. By way of example and not limitation, the neural network-based fourth model may include at least one generative convolution based on sparse convolution. For example, the neural network-based fourth model may be trained using multi-stage loss functions with different granularities. In one example, the binary cross-entropy value may be used as the loss function in the first stage.
[0183] In some embodiments, the number of points used in the loss function may differ for different stages. For example, at least one indicator may be used to indicate the number of points used in the loss function for each stage. The at least one indicator may be predetermined or indicated in the bitstream.
[0184] By way of example and not limitation, a third-stage loss function can be used to train a fourth NN-based model. In this case, at least one indicator may include M, N, and K. M% of the points of the PC samples obtained from the original PC samples through voxel sampling can be used in the first-stage loss function, N% of the points of the PC samples obtained from the original PC samples through voxel sampling can be used in the second-stage loss function, and K% of the points of the PC samples obtained from the original PC samples through voxel sampling can be used in the final-stage loss function. Each of M, N, and K can be a non-negative number. In one example, M can be less than N, and N can be less than K. For example, M can be equal to 12.5, N can be equal to 50, and K can be equal to 100. It should be understood that the above description is for illustrative purposes only, and the specific values listed herein are intended to be exemplary and not to limit the scope of this disclosure. The scope of this disclosure is not limited in this respect.
[0185] In some embodiments, whether and / or how the method is applied may be indicated in the bitstream at one of the following levels: frame level, slice level, stripe level, or octree level. Additionally or alternatively, whether and / or how the method is applied may depend on the encoded / decoded information of the current PC sample.
[0186] In some embodiments, the proposed method can be used to encode and decode geometric information (e.g., coordinates of points in the point cloud sequence). In this case, only the geometric information of the point cloud sequence is encoded and decoded and transmitted as a signal in the bitstream. Alternatively, the proposed method can also be used to encode and decode both geometric and attribute information of the point cloud sequence. The scope of this disclosure is not limited in this respect.
[0187] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve encoding and decoding efficiency and encoding and decoding quality.
[0188] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence generated by a method performed by an apparatus for point cloud processing. The method includes: determining a prediction of current features for the current PC sample based on a current sparse PC sample of the point cloud sequence, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; and generating a bitstream based on the prediction of the current features.
[0189] According to further embodiments of this disclosure, a method for storing a bitstream of a point cloud sequence is provided. The method includes: determining a prediction of current features for the current PC sample based on a current sparse PC sample of the point cloud sequence, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; generating a bitstream based on the prediction of the current features; and storing the bitstream in a non-transitory computer-readable recording medium.
[0190] The implementation of this disclosure can be described according to the following entries, and the features of these entries can be combined in any reasonable manner.
[0191] Item 1. A method for point cloud processing, comprising: a conversion between a current point cloud (PC) sample of a point cloud sequence and a bitstream of the point cloud sequence; determining a prediction of a current feature for the current PC sample based on a current sparse PC sample of the current PC sample, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; and performing the conversion based on the prediction of the current feature.
[0192] Item 2. According to the method of Item 1, at least one reference PC sample includes a single reference PC sample.
[0193] Item 3. According to the method of Item 2, wherein a downsampling process having at least one stage is applied to a single reference PC sample, the sparse PC sample set includes sparse PC samples of the single reference PC sample generated at the last stage of at least one stage, and the feature set includes features of the single reference PC sample generated at the last stage.
[0194] Item 4. According to the method of Item 3, the prediction for the current feature is determined by directly using the current sparse PC sample of the current PC sample, the sparse PC sample of a single reference PC sample, and the feature.
[0195] Item 5. According to the method of Item 3, the prediction for the current feature is determined by performing a refined secondary prediction process based on the current sparse PC sample of the current PC sample, the sparse PC sample of a single reference PC sample, and the feature.
[0196] Item 6. According to the method of Item 3, wherein determining the prediction for the current feature includes: generating an initial prediction for the current feature based on the current sparse PC sample of the current PC sample, the sparse PC sample of a single reference PC sample, and the feature; and generating a secondary prediction for the current feature based on the initial prediction for the current feature and the feature of the single reference PC sample to obtain the prediction for the current feature.
[0197] Item 7. The method of any one of items 3-6, wherein the prediction for the current feature is determined using a first model based on a neural network (NN).
[0198] Item 8. According to the method of Item 7, the first NN-based model includes at least one sparse convolution or at least one sparse convolution on the target coordinates.
[0199] Item 9. According to the method of Item 2, wherein a downsampling process with multiple stages is applied to a single reference PC sample, and the sparse PC sample set includes more than one sparse PC sample generated at more than one stage in the multiple stages of the single reference PC sample, and the feature set includes more than one feature generated at more than one stage of the single reference PC sample.
[0200] Item 10. According to the method of Item 9, wherein sparse PC samples and features generated in a first stage of more than one stage of a single reference PC sample are downsampled, and the downsampled sparse PC samples and downsampled features are aligned and fused with sparse PC samples and features generated in a second stage of more than one stage of a single reference PC sample, and the second stage follows the first stage.
[0201] Item 11. According to the method of Item 10, wherein the prediction for the current feature is determined based on the current sparse PC sample of the current PC sample, the fused sparse PC sample of a single reference PC sample at the last stage in more than one stage, and the fused features.
[0202] Item 12. The method of any of Items 9-11, wherein the prediction for the current feature is determined using a first model based on a neural network (NN).
[0203] Item 13. According to the method of Item 12, wherein the first NN-based model includes at least one sparse convolution or at least one sparse convolution on the target coordinates.
[0204] Item 14. According to the method of Item 1, at least one reference PC sample includes a plurality of reference PC samples.
[0205] Item 15. According to the method of Item 14, the number of multiple reference PC samples is 2.
[0206] Item 16. The method according to any one of Items 14-15, wherein a downsampling process having at least one stage is applied to a plurality of reference PC samples, the sparse PC sample set includes sparse PC samples generated at the last stage of the plurality of reference PC samples in at least one stage, and the feature set includes features generated at the last stage of the plurality of reference PC samples.
[0207] Item 17. According to the method of Item 16, the prediction for the current feature is determined by directly using the current sparse PC sample of the current PC sample, the sparse PC sample of multiple reference PC samples, and the feature.
[0208] Item 18. According to the method of Item 17, wherein determining the prediction for the current feature includes: merging sparse PC samples from multiple reference PC samples; merging features from multiple reference PC samples; and generating a prediction for the current feature based on the current sparse PC sample, the merged sparse PC sample, and the merged features.
[0209] Item 19. According to the method of Item 16, the prediction for the current feature is determined by performing a refined secondary prediction process based on the current sparse PC sample of the current PC sample, the sparse PC sample of multiple reference PC samples, and the feature.
[0210] Item 20. According to the method of Item 16, wherein the plurality of reference PC samples include a first reference PC sample and a second reference PC sample, and determining a prediction for the current feature includes: generating a first prediction for the current feature based on the current sparse PC sample of the current PC sample, the sparse PC sample of the first reference PC sample, and the feature; generating a second prediction for the current feature based on the current sparse PC sample of the current PC sample, the sparse PC sample of the second reference PC sample, and the feature; and generating a prediction for the current feature based on the result of fusing the first prediction and the second prediction.
[0211] Item 21. The method according to any one of Items 14-15, wherein a downsampling process with multiple stages is applied to multiple reference PC samples, and the sparse PC sample set includes more than one sparse PC sample generated at more than one stage of the multiple reference PC samples, and the feature set includes more than one feature generated at more than one stage of the multiple reference PC samples.
[0212] Item 22. According to the method of Item 21, sparse PC samples and features generated in a first stage of more than one stage of multiple reference PC samples are downsampled, and the downsampled sparse PC samples and downsampled features are aligned and fused with sparse PC samples and features generated in a second stage of more than one stage of multiple reference PC samples, respectively, and the second stage is after the first stage.
[0213] Item 23. According to the method of Item 22, wherein the prediction for the current feature is determined based on the current sparse PC sample of the current PC sample, the fused sparse PC sample at the last stage in more than one stage, and the fused feature.
[0214] Item 24. The method of any of Items 15-23, wherein the prediction for the current feature is determined using a neural network (NN) based model.
[0215] Item 25. The method according to any one of items 1-24 further includes: obtaining the current sparse point cloud and the current features by performing a first downsampling process on the current PC sample.
[0216] Item 26. The method of Item 25, wherein the first downsampling process is not based on machine learning.
[0217] Item 27. According to the method of Item 26, wherein the first downsampling process includes sampling at the furthest point or a uniform sampling process.
[0218] Item 28. The method of Item 25, wherein the first downsampling process is based on machine learning.
[0219] Item 29. According to the method of Item 28, the first downsampling process utilizes a second NN-based model.
[0220] Item 30. The method according to any one of items 28-29, wherein the first downsampling process comprises multiple stages.
[0221] Item 31. According to the method of Item 30, the number of multiple stages is 3.
[0222] Item 32. The sparse point cloud and features generated at any stage in multiple stages of the current PC sample are output according to the method of any one of items 30-31.
[0223] Item 33. The method according to any one of items 30-32, wherein the current sparse point cloud is determined as the sparse point cloud generated at the last of multiple stages of the current PC sample, and the current feature is determined as the feature generated at the last stage of the current PC sample.
[0224] Item 34. The method of any one of items 30-33, wherein the sparse point cloud and features generated at all stages in multiple phases of the current PC sample are output.
[0225] Item 35. According to the method of Item 29, the second NN-based model includes at least one sparse convolution.
[0226] Item 36. According to the method of Item 35, the stride of at least one sparse convolution is 2.
[0227] Item 37. The method according to any one of items 35-36, wherein the stride of at least one sparse convolution is predetermined or indicated in the bitstream.
[0228] Item 38. According to the method of any one of items 1-37, at least one reference PC sample is encoded or decoded before the current PC sample.
[0229] Item 39. According to the method of Item 38, at least one reference PC sample includes two PC samples that were encoded and decoded before the current PC sample.
[0230] Item 40. The method according to any one of items 1-39 further includes: obtaining a sparse PC sample set and a feature set of at least one reference PC sample by performing a second downsampling process on at least one reference PC sample.
[0231] Item 41. The method of Item 40, wherein the second downsampling process is based on machine learning.
[0232] Item 42. According to the method of Item 41, the second downsampling process utilizes a third NN-based model.
[0233] Item 43. The method according to any one of items 41-42, wherein the second downsampling process comprises multiple stages.
[0234] Item 44. According to the method of Item 43, the number of multiple stages is 3.
[0235] Item 45. According to the method of any one of Items 43-44, a sparse point cloud and features generated at one or more of multiple stages from at least one reference PC sample are output.
[0236] Item 46. The method according to any one of Items 43-45, wherein the sparse PC sample set of at least one reference PC sample includes sparse PC samples generated at the last two stages of a plurality of stages of at least one reference PC sample, and the feature set of at least one reference PC sample includes features generated at the last two stages of at least one reference PC sample.
[0237] Item 47. According to the method of any one of items 43-46, a sparse point cloud and features generated at all stages in multiple stages from at least one reference PC sample are output.
[0238] Item 48. According to the method of Item 42, the NN-based third model includes at least one sparse convolution.
[0239] Item 49. According to the method of Item 48, the stride of at least one sparse convolution is 2.
[0240] Item 50. The method according to any one of items 48-49, wherein the stride of at least one sparse convolution is predetermined or indicated in the bitstream.
[0241] Item 51. The method according to any one of items 1-50, wherein the current sparse PC sample is encoded into or decoded from the bitstream.
[0242] Item 52. According to the method of Item 51, wherein the current sparse PC sample is encoded and decoded using a point cloud codec.
[0243] Item 53. The method according to Item 52, wherein the point cloud codec is based on: geometry-based point cloud compression (G-PCC), video-based point cloud compression (V-PCC), or Draco.
[0244] Item 54. The method of any one of items 1-53, wherein the residual between the current feature and the prediction for the current feature is encoded into or decoded from the bitstream.
[0245] Item 55. According to the method of Item 54, wherein the residual is encoded or decoded using fixed-length encoding / decoding, unary encoding / decoding, or rounding unary encoding / decoding.
[0246] Item 56. According to the method of Item 55, the residuals are encoded and decoded in a predictive manner.
[0247] Item 57. The method of any one of items 1-56, wherein performing the transformation includes: reconstructing the current PC sample based on the prediction for the current feature and the coordinates of the current sparse PC sample.
[0248] Item 58. According to the method of Item 57, reconstructing the current PC sample includes obtaining the current feature by adding the prediction for the current feature to the residual obtained from the bitstream.
[0249] Item 59. According to the method of Item 58, the reconstruction of the current PC sample further includes: performing an upsampling process on the current sparse PC sample and the current features to reconstruct the current PC sample.
[0250] Item 60. The method according to Item 59, wherein the upsampling process comprises a single upsampling operation.
[0251] Item 61. The method according to Item 59, wherein the upsampling process includes multiple upsampling operations.
[0252] Item 62. According to the method of Item 61, the number of multiple upsampling operations is 3.
[0253] Item 63. The method according to any one of items 61-62, wherein the number of multiple upsampling operations is predetermined.
[0254] Item 64. The method of any of Items 61-62, wherein the number of multiple upsampling operations is indicated in the bitstream.
[0255] Item 65. According to the method of Item 64, the number of multiple upsampling operations is encoded using fixed-length encoding / decoding, unary encoding / decoding, or rounding unary encoding / decoding.
[0256] Item 66. According to the method of Item 64, the number of multiple upsampling operations is encoded and decoded in a predictive manner.
[0257] Item 67. The method of any of Items 59-66, wherein the upsampling process is implemented using a fourth model based on a neural network.
[0258] Item 68. According to the method of Item 67, the fourth NN-based model includes at least one generative convolution based on sparse convolution.
[0259] Item 69. The method of any of Items 67-68, wherein the fourth NN-based model is trained using a multi-level loss function with different granularities.
[0260] Item 70. According to the method of Item 69, the binary cross-entropy value is used as the loss function in the first stage.
[0261] Item 71. The method of any of Items 69-70, wherein the number of points used in the loss function varies for different stages.
[0262] Item 72. The method according to any of Items 69-70, wherein the number of points used in the loss function for each stage is indicated by at least one indicator.
[0263] Item 73. According to the method of Item 72, wherein the NN-based fourth model is trained using a 3-stage loss function, at least one indicator including M, N, and K, where M% of the points of the PC samples obtained from the original PC samples through voxel sampling are used in the first-stage loss function, N% of the points of the PC samples obtained from the original PC samples through voxel sampling are used in the second-stage loss function, and K% of the points of the PC samples obtained from the original PC samples through voxel sampling are used in the last-stage loss function, and each of M, N, and K is a non-negative number.
[0264] Item 74. According to the method of Item 73, where M is less than N and N is less than K.
[0265] Item 75. The method according to any one of items 72-74, wherein at least one indication is predetermined or indicated in the bitstream.
[0266] Item 76. The method of any one of items 1-75, wherein whether and / or how the method is applied is indicated in the bitstream at one of the following levels: frame level, slice level, stripe level, or octree level.
[0267] Item 77. The method according to any one of items 1-76, wherein whether and / or how the method is applied depends on the encoded / decoded information of the current PC sample.
[0268] Item 78. The method according to any one of items 1-77, wherein the PC sample is one of the following: frame, picture, slice, subframe, subpicture, piece, or segment.
[0269] Item 79. The method according to any one of items 1-78, wherein the transformation includes encoding the current PC sample into a bitstream.
[0270] Item 80. The method according to any one of items 1-78, wherein the conversion includes decoding the current PC sample from the bitstream.
[0271] Item 81. An apparatus for point cloud processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1-80.
[0272] Item 82. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method according to any one of items 1-80.
[0273] Item 83. A non-transitory computer-readable recording medium storing a bitstream of a point cloud sequence generated by a method performed by means of a point cloud processing apparatus, wherein the method includes: determining a prediction of a current feature for the current PC sample based on a current sparse PC sample of a current PC sample of the point cloud sequence, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; and generating a bitstream based on the prediction of the current feature.
[0274] Item 84. A method for storing a bitstream of a point cloud sequence, comprising: determining a prediction of a current feature for the current PC sample based on a current sparse PC sample of a current PC sample of the point cloud sequence, a set of sparse PC samples of at least one reference PC sample of the current PC sample, and a feature set; generating a bitstream based on the prediction of the current feature; and storing the bitstream in a non-transitory computer-readable recording medium.
[0275] Example device Figure 6 A block diagram of a computing device 600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 600 may be implemented as a source device 110 (or GPCC encoder 114 or 200) or a destination device 120 (or GPCC decoder 124 or 300), or may be included in a source device 110 (or GPCC encoder 114 or 200) or a destination device 120 (or GPCC decoder 124 or 300).
[0276] It should be understood that, Figure 6 The computing device 600 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0277] like Figure 6 As shown, computing device 600 includes general-purpose computing device 600. Computing device 600 may include at least one or more processors or processing units 610, memory 620, storage unit 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.
[0278] In some embodiments, computing device 600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 600 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0279] Processing unit 610 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 600. Processing unit 610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0280] Computing device 600 typically includes various computer storage media. Such media can be any media accessible by computing device 600, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 630 can be any removable or non-removable media and can include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 600.
[0281] The computing device 600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 6 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0282] Communication unit 640 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 600 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0283] Input device 650 can be one or more of a variety of input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 660 can be one or more of a variety of output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 640, computing device 600 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 600 can also communicate with one or more devices that enable a user to interact with computing device 600, or, if needed, with any device that enables computing device 600 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).
[0284] In some embodiments, some or all of the components of computing device 600 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although to users they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0285] In embodiments of this disclosure, computing device 600 may be used to implement point cloud encoding / decoding. Memory 620 may include one or more point cloud processing modules 625 having one or more program instructions. These modules are accessible and executable by processing unit 610 to perform the functions of the various embodiments described herein.
[0286] In an example embodiment of point cloud encoding, input device 650 may receive point cloud data as input 670 to be encoded. The point cloud data may be processed, for example, by point cloud processing module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via output device 660.
[0287] In an example embodiment of point cloud decoding, input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by point cloud processing module 625 to generate decoded point cloud data. The decoded point cloud data may be provided as output 680 via output device 660.
[0288] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for point cloud processing, comprising: For the conversion between the current point cloud (PC) sample of the point cloud sequence and the bit stream of the point cloud sequence, a prediction of the current feature for the current PC sample is determined based on the current sparse PC sample of the current PC sample, the sparse PC sample set of at least one reference PC sample of the current PC sample, and the feature set. as well as The transformation is performed based on the prediction for the current feature.
2. The method of claim 1, wherein the at least one reference PC sample comprises a single reference PC sample.
3. The method of claim 2, wherein a downsampling process having at least one stage is applied to the single reference PC sample, the sparse PC sample set comprising sparse PC samples of the single reference PC sample generated at the last stage of the at least one stage, and the feature set comprising features of the single reference PC sample generated at the last stage.
4. The method of claim 3, wherein the prediction for the current feature is determined by directly using the current sparse PC sample of the current PC sample, the sparse PC sample of the single reference PC sample, and the feature.
5. The method of claim 3, wherein the prediction for the current feature is determined by performing a refined secondary prediction process based on the current sparse PC sample of the current PC sample, the sparse PC sample of the single reference PC sample, and the feature.
6. The method of claim 3, wherein determining the prediction for the current feature comprises: Based on the current sparse PC sample of the current PC sample, the sparse PC sample of the single reference PC sample, and the feature, an initial prediction is generated for the current feature; as well as Based on the initial prediction for the current feature and the feature of the single reference PC sample, a secondary prediction for the current feature is generated to obtain the prediction for the current feature.
7. The method according to any one of claims 3-6, wherein the prediction for the current feature is determined using a first model based on a neural network (NN).
8. The method of claim 7, wherein the NN-based first model comprises at least one sparse convolution or at least one sparse convolution on the target coordinates.
9. The method of claim 2, wherein a downsampling process having multiple stages is applied to the single reference PC sample, and the sparse PC sample set includes more than one sparse PC sample generated at more than one stage of the multiple stages of the single reference PC sample, and the feature set includes more than one feature generated at more than one stage of the single reference PC sample.
10. The method of claim 9, wherein the sparse PC sample and features of the single reference PC sample generated at a first stage of the more than one stage are downsampled, the downsampled sparse PC sample and the downsampled features are respectively aligned and fused with the sparse PC sample and features of the single reference PC sample generated at a second stage of the more than one stage, and the second stage follows the first stage.
11. The method of claim 10, wherein the prediction for the current feature is determined based on the current sparse PC sample of the current PC sample, the fused sparse PC sample of the single reference PC sample at the last stage of the more than one stage, and the fused features.
12. The method according to any one of claims 9-11, wherein the prediction for the current feature is determined using a first model based on a neural network (NN).
13. The method of claim 12, wherein the NN-based first model comprises at least one sparse convolution or at least one sparse convolution on the target coordinates.
14. The method of claim 1, wherein the at least one reference PC sample comprises a plurality of reference PC samples.
15. The method of claim 14, wherein the number of the plurality of reference PC samples is 2.
16. The method according to any one of claims 14-15, wherein a downsampling process having at least one stage is applied to the plurality of reference PC samples, the sparse PC sample set comprising sparse PC samples of the plurality of reference PC samples generated at the last stage of the at least one stage, and the feature set comprising features of the plurality of reference PC samples generated at the last stage.
17. The method of claim 16, wherein the prediction for the current feature is determined by directly using the current sparse PC sample of the current PC sample, the sparse PC sample of the plurality of reference PC samples, and the feature.
18. The method of claim 17, wherein determining the prediction for the current feature comprises: The sparse PC sample that fuses the multiple reference PC samples; The features of the multiple reference PC samples are fused together; as well as Based on the current sparse PC sample, the fused sparse PC sample, and the fused features, the prediction for the current features is generated.
19. The method of claim 16, wherein the prediction for the current feature is determined by performing a refined secondary prediction process based on the current sparse PC sample of the current PC sample, the sparse PC sample of the plurality of reference PC samples, and the feature.
20. The method of claim 16, wherein the plurality of reference PC samples includes a first reference PC sample and a second reference PC sample, and determining the prediction for the current feature comprises: Based on the current sparse PC sample of the current PC sample, the sparse PC sample of the first reference PC sample, and the features, a first prediction is generated for the current features; Based on the current sparse PC sample of the current PC sample, the sparse PC sample of the second reference PC sample, and the features, a second prediction is generated for the current features; as well as Based on the result of fusing the first prediction and the second prediction, the prediction for the current feature is generated.
21. The method according to any one of claims 14-15, wherein a downsampling process having multiple stages is applied to the plurality of reference PC samples, and the sparse PC sample set includes more than one sparse PC sample generated at more than one stage of the plurality of reference PC samples, and the feature set includes more than one feature generated at more than one stage of the plurality of reference PC samples.
22. The method of claim 21, wherein the sparse PC samples and features of the plurality of reference PC samples generated at a first stage of the more than one stage are downsampled, the downsampled sparse PC samples and the downsampled features are respectively aligned and fused with the sparse PC samples and features of the plurality of reference PC samples generated at a second stage of the more than one stage, and the second stage follows the first stage.
23. The method of claim 22, wherein the prediction for the current feature is determined based on the current sparse PC sample of the current PC sample, the fused sparse PC sample at the last stage of the more than one stage, and the fused feature.
24. The method according to any one of claims 15-23, wherein the prediction for the current feature is determined using a neural network (NN) based model.
25. The method according to any one of claims 1-24, further comprising: The current sparse point cloud and the current features are obtained by performing a first downsampling process on the current PC sample.
26. The method of claim 25, wherein the first downsampling process is not based on machine learning.
27. The method of claim 26, wherein the first downsampling process comprises sampling at the furthest point or a uniform sampling process.
28. The method of claim 25, wherein the first downsampling process is based on machine learning.
29. The method of claim 28, wherein the first downsampling process is applied using a second NN-based model.
30. The method according to any one of claims 28-29, wherein the first downsampling process comprises multiple stages.
31. The method of claim 30, wherein the number of the plurality of stages is 3.
32. The method according to any one of claims 30-31, wherein the sparse point cloud and features of the current PC sample generated at any of the plurality of stages are output.
33. The method according to any one of claims 30-32, wherein the current sparse point cloud is determined to be the sparse point cloud of the current PC sample generated at the last of the plurality of stages, and the current feature is determined to be the feature of the current PC sample generated at the last stage.
34. The method according to any one of claims 30-33, wherein the sparse point cloud and features of the current PC sample generated at all of the plurality of stages are output.
35. The method of claim 29, wherein the second NN-based model comprises at least one sparse convolution.
36. The method of claim 35, wherein the stride of the at least one sparse convolution is 2.
37. The method according to any one of claims 35-36, wherein the stride of the at least one sparse convolution is predetermined or indicated in the bitstream.
38. The method according to any one of claims 1-37, wherein the at least one reference PC sample is encoded or decoded prior to the current PC sample.
39. The method of claim 38, wherein the at least one reference PC sample comprises two PC samples that were encoded or decoded prior to the current PC sample.
40. The method according to any one of claims 1-39, further comprising: By performing a second downsampling process on the at least one reference PC sample, the sparse PC sample set and the feature set of the at least one reference PC sample are obtained.
41. The method of claim 40, wherein the second downsampling process is based on machine learning.
42. The method of claim 41, wherein the second downsampling process is applied using a third NN-based model.
43. The method according to any one of claims 41-42, wherein the second downsampling process comprises multiple stages.
44. The method of claim 43, wherein the number of the plurality of stages is 3.
45. The method according to any one of claims 43-44, wherein the sparse point cloud and features generated at one or more of the plurality of stages of the at least one reference PC sample are output.
46. The method according to any one of claims 43-45, wherein the sparse PC sample set of the at least one reference PC sample includes sparse PC samples of the at least one reference PC sample generated at the last two stages of the plurality of stages, and the feature set of the at least one reference PC sample includes features of the at least one reference PC sample generated at the last two stages.
47. The method according to any one of claims 43-46, wherein the sparse point cloud and features generated at all stages of the plurality of stages of the at least one reference PC sample are output.
48. The method of claim 42, wherein the NN-based third model comprises at least one sparse convolution.
49. The method of claim 48, wherein the stride of the at least one sparse convolution is 2.
50. The method according to any one of claims 48-49, wherein the stride of the at least one sparse convolution is predetermined or indicated in the bitstream.
51. The method according to any one of claims 1-50, wherein the current sparse PC sample is encoded into or decoded from the bitstream.
52. The method of claim 51, wherein the current sparse PC sample is encoded and decoded using a point cloud codec.
53. The method of claim 52, wherein the point cloud codec is based on geometry-based point cloud compression (G-PCC), video-based point cloud compression (V-PCC), or Draco.
54. The method according to any one of claims 1-53, wherein the residual between the current feature and the prediction for the current feature is encoded into or decoded from the bitstream.
55. The method of claim 54, wherein the residual is encoded using fixed-length encoding / decoding, unary encoding / decoding, or rounding unary encoding / decoding.
56. The method of claim 55, wherein the residual is encoded and decoded in a predictive manner.
57. The method according to any one of claims 1-56, wherein performing the conversion comprises: The current PC sample is reconstructed based on the prediction for the current feature and the coordinates of the current sparse PC sample.
58. The method of claim 57, wherein reconstructing the current PC sample comprises: The current feature is obtained by adding the prediction for the current feature to the residual obtained from the bitstream.
59. The method of claim 58, wherein reconstructing the current PC sample further comprises: An upsampling process is performed on the current sparse PC sample and the current feature to reconstruct the current PC sample.
60. The method of claim 59, wherein the upsampling process comprises a single upsampling operation.
61. The method of claim 59, wherein the upsampling process comprises a plurality of upsampling operations.
62. The method of claim 61, wherein the number of the plurality of upsampling operations is 3.
63. The method according to any one of claims 61-62, wherein the number of the plurality of upsampling operations is predetermined.
64. The method according to any one of claims 61-62, wherein the number of the plurality of upsampling operations is indicated in the bitstream.
65. The method of claim 64, wherein the number of the plurality of upsampling operations is encoded using a fixed-length codec, a unary codec, or a rounded unary codec.
66. The method of claim 64, wherein the number of said plurality of upsampling operations is encoded and decoded in a predictive manner.
67. The method according to any one of claims 59-66, wherein the upsampling process is implemented using a fourth model based on a neural network.
68. The method of claim 67, wherein the NN-based fourth model comprises at least one generative convolution based on sparse convolution.
69. The method according to any one of claims 67-68, wherein the NN-based fourth model is trained using a multi-level loss function with different granularities.
70. The method of claim 69, wherein the binary cross-entropy value is used as the loss function in the first stage.
71. The method according to any one of claims 69-70, wherein the number of points used in the loss function varies for different stages.
72. The method according to any one of claims 69-70, wherein the number of points used in the loss function for each stage is indicated using at least one indicator.
73. The method of claim 72, wherein the NN-based fourth model is trained using a 3-level loss function, the at least one indicator comprising M, N, and K, wherein M% of the points of the PC samples obtained from the original PC samples through voxel sampling are used in the first-stage loss function, N% of the points of the PC samples obtained from the original PC samples through voxel sampling are used in the second-stage loss function, and K% of the points of the PC samples obtained from the original PC samples through voxel sampling are used in the last-stage loss function, and each of M, N, and K is a non-negative number.
74. The method of claim 73, wherein M is less than N and N is less than K.
75. The method according to any one of claims 72-74, wherein the at least one indication is predetermined or indicated in the bit stream.
76. The method according to any one of claims 1-75, wherein whether and / or how the method is applied is indicated in the bitstream in one of the following: Frame level, Film level, strip level, or Octree level.
77. The method according to any one of claims 1-76, wherein whether and / or how the method is applied depends on the encoded / decoded information of the current PC sample.
78. The method according to any one of claims 1-77, wherein the PC sample is one of the following: frame, picture, slice, Subframe, Sub-images, film, or part.
79. The method according to any one of claims 1-78, wherein the conversion comprises encoding the current PC sample into the bitstream.
80. The method according to any one of claims 1-78, wherein the conversion comprises decoding the current PC sample from the bitstream.
81. An apparatus for point cloud processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-80.
82. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-80.
83. A non-transitory computer-readable recording medium storing a bitstream of a point cloud sequence generated by a method performed by means of a point cloud processing apparatus, wherein the method comprises: Based on the current sparse PC sample of the current PC sample in the point cloud sequence, the sparse PC sample set of at least one reference PC sample of the current PC sample, and the feature set, a prediction for the current feature of the current PC sample is determined. as well as The bitstream is generated based on the prediction for the current feature.
84. A method for storing a bitstream of a point cloud sequence, comprising: Based on the current sparse PC sample of the current PC sample in the point cloud sequence, the sparse PC sample set of at least one reference PC sample of the current PC sample, and the feature set, a prediction for the current feature of the current PC sample is determined. The bitstream is generated based on the prediction for the current feature; as well as The bitstream is stored in a non-transitory computer-readable recording medium.