Encoding method and device, encoder, code stream, equipment and storage medium

CN120226355APending Publication Date: 2025-06-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080467.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The amount of point cloud data is large, and direct storage and transmission consume a lot of resources. Existing compression technology is difficult to effectively improve the compression rate of point clouds, leading to the problem of insufficient bandwidth.

Method used

By migrating the first point cloud features of the first scale of the first frame to the second point cloud of the first scale of the second frame, the inter-frame information is fused, and the neural network is used for encoding to generate a point cloud code stream, which improves Point cloud compression rate.

Benefits of technology

Significantly improves the compression rate of point clouds, saves coding streams, reduces storage and transmission resource consumption, while maintaining a low distortion rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226355A_ABST
    Figure CN120226355A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an encoding method and device, an encoder, a code stream, equipment and a storage medium, and the encoding method comprises the steps: migrating a first point cloud feature of a first scale of a first frame to a second point cloud of a first scale of a second frame, and obtaining a third point cloud; and according to the third point cloud, encoding a fourth point cloud of a second scale of the second frame to obtain a point cloud code stream.
Need to check novelty before this filing date? Find Prior Art

Description

Coding method and device, encoder, code stream, device, storage medium Technical Field

[0001] The embodiments of the present application relate to point cloud technology, including but not limited to encoding methods and devices, encoders, code streams, devices, and storage media. Background Art

[0002] Point clouds are a type of three-dimensional data, referring to a collection of vectors in a three-dimensional coordinate system. These vectors are usually expressed in the form of (x, y, z) three-dimensional coordinates and also include information such as color, material, and / or reflection intensity. Point clouds are generally used as a representation of three-dimensional objects or scenes. With the rapid development of emerging technologies such as augmented reality, virtual reality, autonomous driving, and robotics, point clouds have become one of the main data forms of these emerging technologies due to their concise representation of three-dimensional space. However, the amount of point cloud data is huge, and directly storing point cloud data consumes a lot of memory and is not conducive to data transmission. Currently, there is insufficient bandwidth to transmit point cloud data directly at the network layer, so it is necessary to improve the compression rate of point clouds.

[0003] Summary of the Invention

[0004] The encoding method and apparatus, encoder, code stream, device, and storage medium provided in the embodiments of the present application can improve the compression rate of point clouds, thereby saving the encoding code stream. The encoding method and apparatus, encoder, code stream, device, and storage medium provided in the embodiments of the present application are implemented as follows:

[0005] According to one aspect of an embodiment of the present application, an encoding method is provided, the method comprising: migrating a first point cloud feature of a first scale of a first frame to a second point cloud of a first scale of a second frame to obtain a third point cloud; and encoding a fourth point cloud of a second scale of the second frame based on the third point cloud to obtain a point cloud code stream.

[0006] According to one aspect of an embodiment of the present application, an encoding device is provided, comprising: a feature migration module configured to migrate features of a first point cloud of a first scale of a first frame to a second point cloud of a first scale of a second frame to obtain a third point cloud; and an encoding module configured to encode a fourth point cloud of a second scale of the second frame based on the third point cloud to obtain a point cloud code stream.

[0007] According to one aspect of an embodiment of the present application, an encoder is provided, comprising: a memory and a processor; wherein the memory is used to store a computer program that can be run on the processor; and the processor is used to execute the encoding method described in the embodiment of the present application when running the computer program.

[0008] According to one aspect of an embodiment of the present application, a point cloud code stream is provided, which is generated by the encoding method described in the embodiment of the present application.

[0009] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor adapted to execute a computer program; and a computer-readable storage medium storing a computer program, wherein when the computer program is executed by the processor, the encoding method described in the embodiment of the present application is implemented.

[0010] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the encoding method described in the embodiment of the present application is implemented.

[0011] In an embodiment of the present application, the first point cloud features of the first scale of the first frame are migrated to the second point cloud of the first scale of the second frame to obtain a third point cloud; based on the third point cloud, the fourth point cloud of the second scale of the second frame is encoded to obtain a point cloud code stream; in this way, when encoding the fourth point cloud of the second scale of the second frame, it is based on the fusion of the point cloud data of the first frame and the second frame, that is, the inter-frame information is taken into account, and there may be a large amount of redundant information between frames, which is beneficial to improving the compression rate when performing point cloud encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings herein are incorporated into and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, serve to illustrate the technical solutions of the present application. Obviously, the drawings described below are merely some embodiments of the present application. Those skilled in the art can, without inventive effort, derive other drawings from these drawings.

[0013] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0014] FIG1 is a flow chart of G-PCC coding;

[0015] FIG2 is a flow chart of G-PCC decoding;

[0016] FIG3 is a schematic diagram of an implementation flow of the encoding method provided in an embodiment of the present application;

[0017] FIG4 is a schematic diagram of the structure of a feature extraction network provided in an embodiment of the present application;

[0018] FIG5 is a schematic diagram of the structure of a feature extraction network provided in an embodiment of the present application;

[0019] FIG6 is a schematic diagram of the structure of the residual layer provided in an embodiment of the present application;

[0020] FIG7 is a schematic diagram of an implementation flow of the encoding method provided in an embodiment of the present application;

[0021] FIG8 is a schematic diagram of an implementation flow of a method for multi-scale geometric lossless compression of dynamic lidar point clouds based on a neural network according to an embodiment of the present application;

[0022] FIG9 is a schematic diagram of the structure of a feature migration network provided in an embodiment of the present application;

[0023] FIG10 is a schematic diagram showing a comparison of experimental results of various encoding methods provided in an embodiment of the present application;

[0024] FIG11 is a schematic structural diagram of an encoding device provided in an embodiment of the present application;

[0025] FIG12 is a schematic diagram of the structure of the encoder provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0028] In the following description, references to “some embodiments,” “this embodiment,” “embodiments of the present application,” and examples, etc., describe a subset of all possible embodiments. However, it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.

[0029] The descriptions such as "first, second, third" appearing in the embodiments of this application are only for illustration and distinction of the described objects. There is no order, nor does it indicate any special limitation on the number of devices in the embodiments of this application, and cannot constitute any limitation on the embodiments of this application.

[0030] The term "voxel" used in the embodiments of this application is short for volume element, the smallest unit of digital data used in three-dimensional space. Voxels allow for the gridding of 3D space and the assignment of characteristics to each grid. For example, a voxel can be a fixed-size cube in three-dimensional space. Voxels are widely used in fields such as 3D imaging, scientific data, and medical imaging.

[0031] Point cloud compression algorithms include two schemes developed by the Moving Picture Experts Group (MPEG): Video-based Point Cloud Compression (V-PCC) and Geometry-based Point Cloud Compression (G-PCC). G-PCC primarily implements geometry compression using an octree model and / or a triangular surface model. V-PCC primarily uses 3D-to-2D projection and video compression.

[0032] With the development of artificial intelligence (AI) technology, neural networks have been applied to geometry-based point cloud compression techniques. Neural network-based point cloud geometry compression techniques can be broadly categorized into lossy and lossless compression. Lossless compression algorithms primarily focus on designing prediction models for voxel occupancy probabilities. Voxel data representations typically utilize octree models, volumetric models, or sparse tensor representations. For lossless geometry compression, the encoder often uses surrounding context, such as parent nodes and / or neighbor nodes, as input. This input is processed through neural network (e.g., convolutional and / or fully connected) layers to output the occupancy probability (also understood as occupancy probability) of each voxel in the point cloud's geometric data. An entropy encoder is then used to convert the voxel occupancy symbol corresponding to each voxel's occupancy probability into a bitstream. Correspondingly, on the decoder side, the occupancy probability of each voxel is predicted using the same process. Based on the predicted occupancy probability, an entropy decoder is used to decode the voxel occupancy symbol from the bitstream, thereby reconstructing the point cloud's geometric data.

[0033] In order to facilitate the understanding of the technical solutions provided by the embodiments of the present application, a flowchart of G-PCC encoding and a flowchart of G-PCC decoding are first provided. It should be noted that the flowchart of G-PCC encoding and the flowchart of G-PCC decoding described in the embodiments of the present application are only for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of point cloud compression technology and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to point cloud encoding and decoding architectures similar to G-PCC. The point cloud compressed by the embodiments of the present application can be a point cloud in a video, or it can be obtained based on a point cloud collected by a lidar sensor, but is not limited to this.

[0034] In the point cloud G-PCC encoder framework, the input point cloud data is sliced ​​and each slice is encoded independently.

[0035] As shown in the G-PCC encoding flow diagram in Figure 1, the encoder first divides the point cloud data to be encoded into multiple slices using striping. Within each slice, the point cloud's geometric and attribute information are encoded separately. During the geometric encoding process, the geometric information undergoes coordinate transformation so that the entire point cloud is contained within a bounding box. Quantization then proceeds. Quantization primarily serves a scaling purpose. Due to quantization rounding, the geometric information of some point clouds remains the same, allowing for parameter-based decision-making regarding whether to remove duplicate points. This process of quantization and removing duplicate points is also known as voxelization. The bounding box is then partitioned into an octree. In the octree-based geometric information encoding process, the bounding box is divided into eight equal parts. The sub-cubes that are not empty (contain points in the point cloud) are further divided into eight equal parts until the resulting leaf nodes are 1x1x1 unit cubes. The points in the leaf nodes are then arithmetic-coded to generate a binary geometric bitstream, or geometry codestream.

[0036] In the process of geometric information encoding based on triangle soup (trisoup), octree partitioning must also be performed first. However, unlike the geometric information encoding based on octree, the trisoup does not need to divide the point cloud into unit cubes with a side length of 1x1x1 step by step. Instead, the division stops when the sub-block (block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, the surface and the twelve edges of the block are obtained. The vertices are arithmetically encoded (surface fitting is performed based on the intersections) to generate a binary geometric bit stream, that is, a geometric code stream. Vertex is also used to implement the process of geometric reconstruction, and the reconstructed geometric information is used when encoding the attributes of the point cloud.

[0037] During the attribute encoding process, color conversion is performed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. The reconstructed geometric information is then used to recolor the point cloud so that the unencoded attribute information corresponds to the reconstructed geometric information. There are two main transformation methods for color information encoding: a distance-based lifting transform that relies on level of detail (LOD) partitioning, and a direct region-adaptive hierarchical transform (RAHT) transformation. Both methods convert color information from the spatial domain to the frequency domain, obtaining high-frequency and low-frequency coefficients. These coefficients are then quantized (i.e., quantized coefficients). The geometrically encoded data, which has undergone octree partitioning and surface fitting, is then sliced ​​and synthesized with the attribute-encoded data processed with the quantized coefficients. The vertex coordinates of each block are then encoded sequentially (i.e., arithmetic encoding), generating a binary attribute bitstream, i.e., the attribute codestream.

[0038] The flowchart of G-PCC decoding shown in Figure 2 is applied to the decoder. The decoder obtains the binary code stream and independently decodes the geometric bit stream (i.e., geometric code stream) and attribute bit stream in the binary code stream. When decoding the geometric bit stream, the geometric information of the point cloud is obtained through arithmetic decoding-octree synthesis-surface fitting-reconstruction geometry-inverse coordinate transformation; when decoding the attribute bit stream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD-based inverse lifting or RAHT-based inverse transformation-inverse color conversion, and the three-dimensional image model of the point cloud data to be encoded is restored based on the geometric information and attribute information.

[0039] The encoding method of the embodiment of the present application can be applied to the arithmetic coding of the geometric information encoding process of the G-PCC as shown in Figure 1, and the first point cloud features of the first scale of the first frame are migrated to the second point cloud of the first scale of the second frame to obtain a third point cloud; according to the third point cloud, the fourth point cloud of the second scale of the second frame is encoded (such as entropy coding) to obtain a point cloud code stream / geometric code stream.

[0040] An embodiment of the present application provides an encoding method, which is applied to an encoder. FIG3 is a schematic diagram of an implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG3 , the method includes the following steps 301 to 302:

[0041] Step 301 , migrating features of a first point cloud at a first scale of a first frame to a second point cloud at a first scale of a second frame to obtain a third point cloud;

[0042] Step 302: Encode the fourth point cloud of the second frame at the second scale according to the third point cloud to obtain a point cloud code stream.

[0043] In an embodiment of the present application, when encoding the fourth point cloud of the second scale of the second frame, the basis is to fuse the point cloud data of the first frame and the second frame, that is, the inter-frame information is taken into account, and there may be a large amount of redundant information between frames, which is beneficial to improving the compression rate when performing point cloud encoding.

[0044] The following describes further optional implementations and related terms of each of the above steps.

[0045] In step 301 , features of a first point cloud at a first scale of a first frame are transferred to a second point cloud at a first scale of a second frame to obtain a third point cloud.

[0046] In some embodiments, the encoder may perform feature migration on the first point cloud of the first scale of the first frame using a pre-trained feature migration network to obtain a seventh point cloud; wherein the feature migration network includes at least a convolutional layer; and merge the seventh point cloud with the second point cloud of the first scale of the second frame to obtain the third point cloud.

[0047] In the embodiment of the present application, the relationship between the first frame and the second frame is not limited and can be adjacent frames or non-adjacent frames. In the case of non-adjacent frames, the number of frames between the first frame and the second frame is not limited and can be one frame, two frames, or three frames.

[0048] In some embodiments, the first frame is captured before the second frame. For example, the first frame is the t-1th frame, and the second frame is the tth frame; where t is greater than or equal to 1.

[0049] In the embodiment of the present application, there is no limitation on how the first point cloud of the first scale is obtained. The first point cloud of the first scale can be obtained based on the point cloud collected by a point cloud collector / point cloud sensor (such as a lidar sensor), or can be obtained by downsampling the point cloud of the first frame that is higher than the first scale, or by upsampling the point cloud of the first frame that is lower than the first scale.

[0050] In some embodiments, feature extraction may be performed on the fifth point cloud of the second scale of the first frame to obtain the first point cloud of the first scale of the first frame.

[0051] It can be understood that the purpose of feature extraction is to help the encoder predict the occupancy probability more accurately, which is beneficial to improving the compression performance, that is, the distortion is small at a high compression rate.

[0052] In the embodiments of the present application, the relationship between the second scale and the first scale is not limited; the second scale can be higher or lower than the first scale. It is understood that when the second scale is higher than the first scale, feature extraction performs a downsampling function; when the second scale is lower than the first scale, feature extraction performs an upsampling function.

[0053] For example, in some embodiments, the encoder may extract features from the fifth point cloud at the second scale of the first frame using a pre-trained feature extraction network to obtain the first point cloud at the first scale of the first frame; wherein the feature extraction network includes at least a convolutional layer, an activation function, a residual layer, and a downsampling layer. Furthermore, in some embodiments, the feature extraction network is a sparse convolution-based network, and the convolutional layer is a sparse convolutional layer.

[0054] In the embodiments of the present application, there is no restriction on the size of the convolution kernel used in the convolution layer. For example, the size of the convolution kernel is 2×2×2. In some embodiments, the downsampling layer can implement voxel downsampling by pooling. There is no restriction on the step size used for pooling, which of course depends on the relationship between the first scale and the second scale.

[0055] In the embodiments of the present application, there is no limitation on the number of convolutional layers, activation functions, residual layers, and downsampling layers. In other words, there is no limitation on the network structure of the feature extraction network. For example, in some embodiments, as shown in FIG4 , the feature extraction network includes a downsampling module 40, which includes four sparse convolutional layers, two activation functions, two residual layers, and one downsampling layer.

[0056] Furthermore, in some embodiments, the feature extraction network further includes an upsampling layer. For example, in other embodiments, as shown in FIG5 , the feature extraction network 5 includes two downsampling modules 40 and one upsampling module 50; wherein the upsampling module 50 includes four sparse convolutional layers, two activation functions, two residual layers, and one upsampling layer.

[0057] In the embodiments of the present application, the residual layer plays the role of feature extraction, which affects the final compression rate. Specifically, in some embodiments, the structure of the residual layer is shown in Figure 6, including two sparse convolutional layers and two activation functions. The output of the residual layer can be the sum of the two input point clouds. Of course, the structure of the residual layer is not limited to this, that is, there is no limit on the number of sparse convolutional layers and the number of activation functions.

[0058] In some embodiments, the network parameters of the feature extraction network can be pre-trained based on a first threshold, where the first threshold is a performance parameter that characterizes the encoding. That is, the encoding performance parameter (e.g., compression ratio) of the encoding method for the sample point cloud can be determined, and a corresponding loss function value can be determined based on the first threshold and the encoding performance parameter. The network parameters of the feature extraction network are then trained / adjusted based on the loss function value, and this process is repeated until the loss function value of the input sample point cloud satisfies the first threshold.

[0059] It can be understood that for scenes where the first scale is higher than / larger than the second scale, the number of upsampling layers is greater than the number of downsampling layers, and the number of upsampling modules is greater than the number of downsampling modules. For scenes where the first scale is lower than / smaller than the second scale, the number of upsampling layers is less than the number of downsampling layers, and the number of upsampling modules is less than the number of downsampling modules.

[0060] In the embodiment of the present application, the fifth point cloud of the second scale of the first frame may be a point cloud after motion compensation / self-motion compensation, or may be a point cloud without self-motion compensation.

[0061] In some embodiments, the encoder may perform motion compensation on the sixth point cloud of the second scale of the first frame based on the posture change information of the second frame relative to the first frame to obtain the fifth point cloud; wherein the sixth point cloud may be obtained based on the point cloud collected by the point cloud sensor, for example, the sixth point cloud may be obtained after preprocessing the point cloud collected by the point cloud sensor.

[0062] In an embodiment of the present application, the purpose of motion compensation is to convert the sixth point cloud of the second scale of the first frame to the coordinate system of the second frame, so that the first point cloud and the second point cloud are in the same coordinate system after feature migration, that is, the first point cloud and the second point cloud included in the third point cloud are in the same coordinate system; in this way, the error between the geometric coordinates of the same physical space point in the first point cloud and the geometric coordinates in the second point cloud can be reduced, which is beneficial for the neural network to extract more redundant information, and further beneficial for improving the compression rate of the point cloud.

[0063] It can be understood that the device where the sensor for collecting point clouds is located usually also carries sensors such as inertial navigation. For example, an autonomous driving car used to collect point cloud data is not only equipped with a lidar sensor for collecting point cloud data, but also an inertial navigation sensor. While the lidar sensor collects point cloud data, the inertial navigation sensor also synchronously collects posture information / position coordinates. In an embodiment of the present application, the posture change information refers to the change from the posture coordinates of the first frame to the posture coordinates of the second frame, that is, the inertial navigation sensor establishes a coordinate system at the center point in the first frame, and re-establishes a coordinate system at the center point in the second frame. The posture change information is the rotation and translation of the coordinate system from the first frame to the second frame, that is, the rotation and translation of the inertial navigation sensor from the first frame to the second frame, which is equivalent to / approximate to the rotation and translation of the lidar sensor from the first frame to the second frame.

[0064] In step 302, a fourth point cloud of the second frame at the second scale is encoded according to the third point cloud to obtain a point cloud code stream.

[0065] In some embodiments, the encoder may upsample the third point cloud to obtain an eighth point cloud of the second scale; predict the occupancy probability of the voxels of the fourth point cloud based on the eighth point cloud to obtain the voxel occupancy probability of the fourth point cloud; and encode the fourth point cloud based on the voxel occupancy probability of the fourth point cloud to obtain the point cloud code stream.

[0066] Exemplarily, in some embodiments, encoding the fourth point cloud according to the voxel occupancy probability of the fourth point cloud to obtain the point cloud code stream includes: determining the occupancy symbol of the corresponding voxel according to the voxel occupancy probability of the fourth point cloud; encoding the occupancy symbol to obtain the point cloud code stream.

[0067] In an embodiment of the present application, for the voxelization process, a point in the point cloud can correspond to an occupied voxel (i.e., a non-empty voxel), and an unoccupied voxel (i.e., an empty voxel) indicates that there is no point in the point cloud at the voxel position. In some embodiments, the occupied voxels can be marked with a first occupancy symbol (e.g., the symbol is 1), and the unoccupied voxels can be marked with a second occupancy symbol (e.g., the symbol is 0). In this way, the voxelized point cloud can represent the geometric data of the point cloud by the occupancy symbols of the voxels at each position in the voxel grid. If the occupancy probability of a voxel is greater than or equal to a preset probability threshold, indicating that the probability of the voxel containing a point in the point cloud is relatively high, then the occupancy probability is represented by 1, and the preset occupancy symbol 1 corresponds to the voxel corresponding to the occupancy probability. Conversely, if the occupancy probability of a voxel is less than the preset probability threshold, indicating that the probability of the voxel containing a point in the point cloud is relatively low, then the occupancy probability is represented by 0, and the preset occupancy symbol 0 corresponds to the voxel corresponding to the occupancy probability.

[0068] The encoder encodes the occupancy symbols of the voxels to obtain corresponding first encoded data, that is, occupancy indication information corresponding to the voxels, thereby achieving lossless compression of the geometric data of the point cloud.

[0069] In some embodiments, entropy coding may employ, but is not limited to, a context-based adaptive binary arithmetic coding (CABAC) algorithm. According to the principle of entropy coding, the more accurate the prediction of occupancy probability, the lower the information entropy, and the greater the actual bit rate and bandwidth savings.

[0070] The embodiment of the present application further provides an encoding method. FIG7 is a schematic diagram of an implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG7 , the method includes the following steps 701 to 705:

[0071] Step 701: Perform motion compensation on the sixth point cloud of the second scale of the first frame based on the pose change information of the second frame relative to the first frame to obtain the fifth point cloud of the second scale of the first frame.

[0072] Step 702: Perform feature extraction on the fifth point cloud of the second scale of the first frame using a pre-trained feature extraction network to obtain a first point cloud of the first scale of the first frame; wherein the feature extraction network includes at least a convolutional layer, an activation function, a residual layer, a downsampling layer, and an upsampling layer;

[0073] Step 703 , migrating the first point cloud features of the first frame at the first scale to the second point cloud of the second frame at the first scale to obtain a third point cloud;

[0074] Step 704: Input the third point cloud into a probability prediction network to perform occupancy probability prediction on voxels of the fourth point cloud at the second scale of the second frame to obtain voxel occupancy probability of the fourth point cloud;

[0075] Step 705 : Input the voxel occupancy probability of the fourth point cloud and the fourth point cloud into an entropy codec to encode the fourth point cloud to obtain the point cloud code stream.

[0076] Point cloud compression algorithms, such as video-based point cloud compression (V-PCC) and geometry-based point cloud compression (G-PCC). Geometric compression in G-PCC is primarily implemented using an octree model and / or a triangular surface model. V-PCC is primarily implemented through 3D to 2D projection and video compression.

[0077] In recent years, neural networks and deep learning technologies have also been widely used in point cloud geometry compression. These can be categorized into volumetric model compression techniques based on 3D convolutional neural networks (3D CNNs), compression techniques that directly apply point coordinate sets to neural networks based on multi-layer perceptrons (MLPs), and compression techniques that use MLPs or 3D CNNs to perform probability estimation and entropy coding on octree node symbols. These methods have demonstrated superior compression performance compared to traditional methods, but they currently target single-frame point clouds. Point cloud acquisition typically results in a time-domain sequence, resulting in a significant amount of redundant information between frames in dynamic point cloud sequences.

[0078] Based on this, the following describes an exemplary application of the embodiment of the present application in a practical application scenario.

[0079] For the dynamic radar point cloud sequence collected by the lidar sensor, in an embodiment of the present application, the relative motion posture information of the vehicle collected by the sensor (such as IMU) on the autonomous driving vehicle is used to perform self-motion compensation on the radar point cloud, so that the neural network can more easily discover redundant information in the time domain of the dynamic radar point cloud, thereby improving the compression performance of the radar point cloud.

[0080] In an embodiment of the present application, a method for geometric lossless compression of dynamic radar point clouds is provided using a neural network, aiming to improve the lossless compression performance of dynamic radar point clouds. The method comprises the following steps (1) to (5):

[0081] Step (1) takes the n-th scale point cloud of the t-1th frame of the dynamic lidar point cloud sequence, that is, (i.e. an example of the sixth point cloud), using the known pose information of the t-1th frame, Perform self-motion compensation and obtain (i.e., an example of the fifth point cloud); wherein the t-1th frame is an example of the first frame, and the nth scale is an example of the second scale;

[0082] Step (2), Input the sparse convolution-based feature extraction network to achieve feature extraction and downsampling, and obtain the n-1 scale point cloud corresponding to the t-1 frame (i.e., an example of the first point cloud); wherein the n-1th scale is an example of the above-mentioned first scale;

[0083] Step (3), And the point cloud of the n-1th scale corresponding to the tth frame (i.e., an example of the second point cloud) is input into the feature transfer network based on the target convolution to realize feature transfer and obtain (i.e., an example of the third point cloud); wherein, the tth frame is an example of the second frame;

[0084] Step (4), The sparse convolution-based probability prediction module is input to predict the voxel occupancy / occupancy probability of the point cloud at the nth scale corresponding to the tth frame. The actual occupancy symbol of the corresponding voxel is losslessly compressed using entropy coding based on the voxel occupancy probability, thereby outputting the final encoded bitstream. This can improve the encoding and decoding efficiency of the entropy encoder and save bitstream.

[0085] Step (5): Repeat the above steps (1) to (4) from the lowest scale to the highest scale, thereby achieving lossless compression of the dynamic lidar point cloud sequence.

[0086] Furthermore, in step (1), the steps of self-motion compensation are as follows: first read the corresponding pose T of the t-1 frame t-1 and the corresponding pose T of the t-th frame t ,use Calculate the relative pose transformation (i.e. pose change information) from the t-1th frame to the tth frame, and then Convert to homogeneous coordinate matrix The point cloud P t-1 Convert to homogeneous coordinates The two calculate the point cloud after self-motion compensation

[0087] Furthermore, the sparse convolution-based feature extraction network described in step (2) includes two downsampling modules and one upsampling module, and the downsampling module includes two residual layers and one downsampling layer. The upsampling module includes two residual layers and one upsampling layer. The residual layer is a residual connection based on sparse convolution, the downsampling layer is implemented by a sparse convolution layer with a convolution kernel size and a convolution step size of 2×2×2, and the upsampling layer is implemented by a transposed convolution layer with a convolution kernel size and a convolution step size of 2×2×2.

[0088] Furthermore, the target convolution-based feature transfer network described in step (3) includes a target convolution layer.

[0089] As can be understood, in the embodiments of this application, the time domain information of the lidar point cloud is utilized, and a neural network is used to extract features from this time domain information and intra-frame information, thereby constructing a multi-scale compression model and achieving lossless compression of dynamic radar point cloud sequences. Secondly, the position information of the autonomous vehicle is used to perform self-motion compensation on the radar point cloud, which facilitates feature extraction from the neural network and greatly improves the efficiency of point cloud compression.

[0090] In an embodiment of the present application, a multi-scale geometric lossless compression method (i.e., encoding method) of a dynamic lidar point cloud based on a neural network is implemented as shown in FIG8. The n-scale point cloud of the t-1th frame of the dynamic lidar point cloud sequence is taken, and the self-motion compensation is performed on the n-scale point cloud of the t-1th frame using the known pose information of the t-1th frame to obtain the n-scale point cloud of the t-1th frame after motion compensation; the point cloud is then input into a feature extraction network based on sparse convolution to extract features and downsample to obtain the n-1th scale point cloud corresponding to the t-1th frame; then, the n-1th scale point cloud of the t-1th frame and the n-1th scale point cloud corresponding to the t-1th frame are input into a feature transfer network based on target convolution to achieve feature transfer; finally, the n-1th scale point cloud corresponding to the t-1th frame is input into a probability prediction network based on sparse convolution to predict the occupancy probability of the n-1th scale corresponding to the t-1th frame, and the actual occupancy symbol of the voxel is losslessly compressed using entropy coding according to the occupancy probability, thereby outputting the final encoded bitstream. This process is repeated from the lowest scale to the highest scale to achieve lossless compression of the dynamic lidar point cloud sequence. This encoding method can be used between multiple scales, and the compression of each scale is independent of each other, which can achieve scale-scalable encoding with strong flexibility.

[0091] In some embodiments, the structure of the downsampling module in the feature extraction network is shown in FIG5 , where the encoder includes two residual layers and one downsampling layer, as well as four sparse convolution layers and two activation functions.

[0092] In some embodiments, as shown in FIG5 , the upsampling module in the feature extraction network includes two residual layers and one upsampling layer, and also includes four sparse convolution layers and two activation functions.

[0093] Among them, the residual layer is a residual connection based on sparse convolution. As shown in Figure 6, the residual layer consists of two sparse convolution layers and residual connections. Among them, downsampling is achieved through a sparse convolution layer with a convolution kernel size and a convolution step size of 2×2×2, and upsampling is achieved through a transposed convolution layer with a convolution kernel size and a convolution step size of 2×2×2.

[0094] In some embodiments, the structure of the feature transfer network is shown in FIG9 , and the feature transfer network includes a target convolution layer.

[0095] The embodiments of this application were compared with a sparse convolution-based multi-scale intra-frame point cloud compression scheme (SparsePCGC) and a traditional encoding and decoding method (G-PCC). The experimental results are shown in Figure 10. The comparison indicator for point cloud lossless compression is bit rate (bits per input point, bpp), which can also be understood as compression rate. The rate-distortion curve of the lossy compression of the point cloud is also given. The test data is the lidar point cloud data in the SemanticKITTI dataset, and the acquisition sensor is a Velodyne 64-line lidar.

[0096] The experimental results are shown in Table 1:

[0097] Table 1

[0098] Point cloud dataset G-PCCSparsePCGCOursKITTI_07_vox1mm19.38918.203(-6.12%)14.439(-25.53%)

[0099] In Table 1, the values ​​in parentheses represent the rate-distortion compared to G-PCC. It can be understood that the rate-distortion performance gain is negatively correlated with codec performance. Smaller gains, such as those expressed as negative values, indicate better codec performance. As can be seen, the embodiments of the present application achieve higher BD-rate gains compared to the traditional G-PCC codec method, with a maximum improvement of 25.53%. This data demonstrates improved codec performance.

[0100] It can be seen from the experimental results that the embodiments of the present application are superior to the current mainstream methods in terms of quantitative performance indicators.

[0101] Based on the foregoing embodiments, the encoding device provided in the embodiments of the present application, including the modules included and the units included in each module, can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0102] FIG11 is a schematic diagram of the structure of an encoding device provided in an embodiment of the present application. As shown in FIG11 , the encoding device 11 includes:

[0103] A feature migration module 111 is configured to migrate features of a first point cloud of a first scale in a first frame to a second point cloud of a first scale in a second frame to obtain a third point cloud;

[0104] The encoding module 112 is configured to encode the fourth point cloud of the second frame at the second scale according to the third point cloud to obtain a point cloud code stream.

[0105] In some embodiments, the feature migration module 111 is configured to: perform feature migration on the first point cloud through a pre-trained feature migration network to obtain a seventh point cloud; wherein the feature migration network includes at least a convolutional layer; and merge the seventh point cloud with the second point cloud to obtain the third point cloud.

[0106] In some embodiments, the encoding device 11 further includes a feature extraction module, which is configured to: perform feature extraction on the fifth point cloud of the second scale of the first frame to obtain the first point cloud.

[0107] Furthermore, in some embodiments, the feature extraction module is configured to: perform feature extraction on the fifth point cloud through a pre-trained feature extraction network to obtain the first point cloud; wherein the feature extraction network includes at least a convolutional layer, an activation function, a residual layer and a downsampling layer.

[0108] In some embodiments, the feature extraction network further includes an upsampling layer.

[0109] Exemplarily, in some embodiments, the feature extraction network is a sparse convolution-based network, and the convolution layer is a sparse convolution layer.

[0110] In some embodiments, the network parameters of the feature extraction network are obtained by training according to a first threshold; the first threshold is a performance parameter that characterizes the encoding.

[0111] In some embodiments, the encoding device 11 further includes a motion compensation module, which is configured to perform motion compensation on the sixth point cloud of the second scale of the first frame according to the posture change information of the second frame relative to the first frame to obtain the fifth point cloud.

[0112] In some embodiments, the first scale is smaller than the second scale, and the number of upsampling layers is smaller than the number of downsampling layers.

[0113] In some embodiments, the encoding module 112 is configured to: upsample the third point cloud to obtain an eighth point cloud of the second scale; predict the occupancy probability of the voxels of the fourth point cloud based on the eighth point cloud to obtain the voxel occupancy probability of the fourth point cloud; and encode the fourth point cloud based on the voxel occupancy probability of the fourth point cloud to obtain the point cloud code stream.

[0114] Furthermore, in some embodiments, the encoding module 112 is configured to: determine an occupancy symbol of a corresponding voxel according to the voxel occupancy probability of the fourth point cloud; and encode the occupancy symbol to obtain the point cloud code stream.

[0115] In some embodiments, the point cloud includes geometric information; wherein the point cloud refers to a first point cloud, a second point cloud, a third point cloud, a fourth point cloud, a fifth point cloud, a sixth point cloud, a seventh point cloud, and an eighth point cloud.

[0116] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.

[0117] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0118] It should be noted that the division of modules in the device described in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or they can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. It can also be implemented in the form of a combination of software and hardware.

[0119] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0120] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, an encoding method such as that on an encoder side is implemented.

[0121] An embodiment of the present application provides a point cloud code stream, which is generated by the encoding method described in the embodiment of the present application.

[0122] The present application provides an encoder. FIG12 is a schematic diagram of the structure of the encoder provided by the embodiment of the present application. As shown in FIG12 , the encoder 120 includes: a communication interface 121, a memory 122, and a processor 123; each component is coupled together through a bus system 124. It can be understood that the bus system 124 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 124 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus systems 124 in FIG12. Among them,

[0123] Communication interface 121, used for sending and receiving signals when sending and receiving information with other external devices;

[0124] Memory 122, for storing computer programs that can be run on processor 123;

[0125] The processor 123 is configured to, when running the computer program, execute:

[0126] Migrating the first point cloud features of the first frame at the first scale to the second point cloud of the second frame at the first scale to obtain a third point cloud;

[0127] According to the third point cloud, a fourth point cloud of the second frame at the second scale is encoded to obtain a point cloud code stream.

[0128] It is understood that the memory 122 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 122 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0129] The processor 123 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 123 or software instructions. The above-mentioned processor 123 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 122 , and the processor 123 reads the information in the memory 122 and completes the steps of the above encoding method in combination with its hardware.

[0130] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0131] Optionally, as another embodiment, the processor 123 is further configured to execute the aforementioned encoding method embodiment when running the computer program.

[0132] An embodiment of the present application provides an electronic device, comprising: a processor adapted to execute a computer program; and a computer-readable storage medium storing the computer program. When the computer program is executed by the processor, the computer program implements the encoding and / or decoding methods described in the embodiments of the present application. The electronic device can be any type of device capable of point cloud encoding and / or decoding, such as a mobile phone, tablet computer, laptop computer, personal computer, television, projection device, or monitoring device.

[0133] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0134] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.

[0135] The term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, object A and / or object B can mean: object A exists alone, object A and object B exist at the same time, and object B exists alone.

[0136] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0137] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0138] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed across multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of this embodiment.

[0139] In addition, all functional modules in the embodiments of the present application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0140] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0141] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks or optical disks.

[0142] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0143] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0144] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0145] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A coding method, the method comprising: Migrating a first point cloud feature of a first scale of a first frame to a second point cloud of a first scale of a second frame to obtain a third point cloud; According to the third point cloud, a fourth point cloud of the second frame at a second scale is encoded to obtain a point cloud code stream.

2. The method according to claim 1, wherein: The method further comprises: Perform feature extraction on a fifth point cloud of the second scale of the first frame to obtain the first point cloud.

3. The method according to claim 2, wherein: The method further comprises: According to the posture change information of the second frame relative to the first frame, motion compensation is performed on the sixth point cloud of the second scale of the first frame to obtain the fifth point cloud.

4. The method according to claim 2, wherein: The step of extracting features from the fifth point cloud of the second scale of the first frame to obtain the first point cloud includes: The first point cloud is obtained by performing feature extraction on the fifth point cloud through a pre-trained feature extraction network; wherein the feature extraction network at least includes a convolutional layer, an activation function, a residual layer and a downsampling layer.

5. The method according to claim 4, wherein: The feature extraction network also includes an upsampling layer.

6. The method according to claim 5, wherein: The first scale is smaller than the second scale, and the number of layers of the upsampling layer is smaller than the number of layers of the downsampling layer.

7. The method according to claim 4, wherein: The feature extraction network is a sparse convolution-based network, and the convolution layer is a sparse convolution layer.

8. The method according to any one of claims 4 to 7, wherein: The network parameters of the feature extraction network are obtained by training according to a first threshold; the first threshold is a performance parameter that characterizes the encoding.

9. The method according to claim 1, wherein: The step of migrating the first point cloud features of the first frame at the first scale to the second point cloud of the second frame at the first scale to obtain the third point cloud includes: Performing feature migration on the first point cloud through a pre-trained feature migration network to obtain a seventh point cloud; wherein the feature migration network includes at least a convolutional layer; The seventh point cloud is merged with the second point cloud to obtain the third point cloud.

10. The method according to claim 1, wherein: The step of encoding the fourth point cloud of the second frame at the second scale according to the third point cloud to obtain a point cloud code stream includes: Upsampling the third point cloud to obtain an eighth point cloud of the second scale; According to the eighth point cloud, predicting the occupancy probability of the voxels of the fourth point cloud to obtain the voxel occupancy probability of the fourth point cloud; The fourth point cloud is encoded according to the voxel occupancy probability of the fourth point cloud to obtain the point cloud code stream.

11. The method according to claim 10, wherein: The step of encoding the fourth point cloud according to the voxel occupancy probability of the fourth point cloud to obtain the point cloud code stream includes: Determining an occupancy symbol of a corresponding voxel according to the voxel occupancy probability of the fourth point cloud; The occupancy symbol is encoded to obtain the point cloud code stream.

12. The method according to any one of claims 1 to 11, wherein: The point cloud includes geometric information.

13. A coding device, comprising: a feature migration module configured to migrate features of a first point cloud of a first scale of a first frame to a second point cloud of a first scale of a second frame to obtain a third point cloud; The encoding module is configured to encode the fourth point cloud of the second frame at the second scale according to the third point cloud to obtain a point cloud code stream.

14. An encoder, comprising: A memory and a processor; wherein, The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the method according to any one of claims 1 to 12 when running the computer program.

15. A point cloud code stream, wherein the point cloud code stream is generated by the encoding method according to any one of claims 1 to 12.

16. An electronic device, comprising: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the encoding method according to any one of claims 1 to 12 is implemented.

17. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed, the encoding method according to any one of claims 1 to 12 is implemented.