Three-dimensional point cloud prediction geometric coding method and system based on deep learning

By introducing deep learning-based predictive geometric coding methods and differential evolution quantization step size selection methods in LiDAR point cloud compression technology, the problem of ignoring long-distance geometric correlation and quantitative parameter optimization in the existing technology is solved, and more efficient point cloud coding and better rate distortion performance are achieved.

CN120111256AActive Publication Date: 2025-06-06SHANDONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510254666.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The existing LiDAR point cloud compression technology models a simple linear relationship between adjacent points in the prediction tree, ignoring the long-distance geometric correlation between different coordinates, resulting in an increase in distortion at low bit rate, which is not as good as the octree encoding method, and the rate distortion optimization problem of the quantization parameters of the spherical coordinate system is not discussed.

Method used

A three-dimensional point cloud prediction geometric coding method based on deep learning is proposed. By constructing two prediction trees (based on radar parameters and threshold segmentation), and using the pitch angle prediction model of the LSTM network in a high-code rate mode, combining differential coding and entropy coding, it adapts to the encoding mode at different code rates. In addition, a quantized step size selection method based on differential evolution is adopted to optimize the encoding parameters to improve rate distortion performance.

Benefits of technology

The rate distortion performance of LiDAR point cloud encoding is improved. At the same bit rate, the decoded reconstruction point cloud has higher fidelity, and better quantization results are achieved under the spherical coordinate system, reducing the loss of encoding speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111256A_ABST
    Figure CN120111256A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional point cloud prediction geometric coding method and system based on deep learning. The method comprises the following steps: 1) constructing a prediction tree; constructing a prediction tree based on the radar parameters; or constructing a prediction tree by adopting a threshold segmentation method; 2) predictive coding; a high-bit-rate mode and a low-bit-rate mode are designed, and the coding requirements under different bit rates are met; and 3) selecting a coding parameter based on differential evolution: converting a coding parameter selection problem into an optimization problem under a constraint condition through a quantization step length selection method based on differential evolution, and obtaining an approximately optimal coding parameter through iterative optimization on a small data set only containing ten point clouds, and finally, other point clouds are coded through the coding parameters. According to the method, the rate distortion performance of Lidar point cloud coding is improved, and compared with a previous method, the reconstructed point cloud after decoding has higher fidelity under the same bit rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a three-dimensional point cloud prediction geometry coding method and system based on deep learning, and belongs to the technical field of image processing. Background Art

[0002] LiDAR can accurately collect 3D scene information and is widely used in fields such as autonomous driving, robot navigation, and geographic information systems. The point cloud data collected by LiDAR generally contains geometric information (such as Cartesian coordinates x, y, z) and attribute information (such as reflectivity). However, due to the huge amount of LiDAR point cloud data, efficient compression technology is urgently needed to reduce the storage and transmission costs of LiDAR point clouds.

[0003] In recent years, the Moving Picture Experts Group (MPEG) has released the Geometry-Based Point Cloud Compression (G-PCC) standard and is still researching more efficient point cloud compression technologies. G-PCC includes two LiDAR point cloud geometry information encoding methods: octree encoding method and predictive geometry encoding method. Among them, the predictive geometry encoding method makes full use of the LiDAR acquisition principle and is more efficient than the octree encoding method.

[0004] Specifically, LiDAR collects point cloud data through multiple built-in lasers with fixed elevation angles and fixed azimuth resolution. This acquisition principle causes the point cloud to have stronger correlation in the spherical coordinate system. In order to take advantage of this feature, the angle mode in the predictive geometry coding method converts the geometric information of the point cloud into a spherical coordinate system and groups the points according to the laser number. Within each group, the points are connected in order according to the size of their azimuth angles to construct a prediction tree. Subsequently, starting from the root node, the geometric information is sequentially predicted through the encoded neighboring points, and the residual is entropy encoded.

[0005] However, the predictive geometry coding method in G-PCC still has the following shortcomings. First, this method models a simple linear relationship between adjacent points in the prediction tree to predict the geometric information of the current point, but ignores the long-distance geometric correlation between different coordinates. Second, at low bit rates, the distortion increases and the geometric correlation decreases accordingly, resulting in the performance of this method being inferior to that of the octree coding method. Finally, in the spherical coordinate system, the distortion caused by the quantization of the residual on each axis has different effects on the reconstruction distortion, but this method does not explore the rate-distortion (RD) optimization problem of the quantization parameter (QP) in the spherical coordinate system, so the optimal RD performance cannot be achieved. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention provides a three-dimensional point cloud prediction geometry coding method based on deep learning;

[0007] like Figure 1As shown, first, according to the different known information of the point cloud to be encoded, the present invention proposes two prediction tree construction methods, and selects the most suitable one to construct the prediction tree in the spherical coordinate system. Since the lidar point cloud has different annotation contents about the acquisition device related information in addition to geometry and attributes, and the arrangement order of the geometric coordinates of the points in the point cloud file is also different, this method has two built-in prediction tree construction methods. For point cloud files containing the pitch angles of each laser scanner of the lidar and the height of each laser radar scanner coordinate system relative to the radar coordinate system, such as the Ford point cloud dataset, the prediction tree construction method based on radar parameters is preferably used. For point cloud files that do not contain the above-mentioned related parameters, but the coordinates of the points are arranged in the point cloud file according to the height order of the laser scanner and the acquisition order of the laser scanner, such as the SemanticKITTI dataset, the threshold segmentation method is preferably used to construct the prediction tree.

[0008] Then, as the bit rate decreases, the correlation of the points will be weakened. Therefore, the present invention proposes a high bit rate coding mode and a low bit rate coding mode. In the high bit rate coding mode, the method trains a LSTM-based pitch angle prediction model by exploring the relationship between the radius and the pitch angle of the Lidar point cloud, and encodes the prediction residual of the pitch angle; the azimuth angle and the radial distance are directly differentially encoded. In the low bit rate mode, the radial distance is encoded by a differential autoencoder; the azimuth angle and the pitch angle are directly differentially encoded.

[0009] Since the best rate-distortion performance cannot be obtained by processing the residuals of the pitch angle, azimuth angle and radius with the same quantization step size in a spherical coordinate system, the present invention proposes a quantization step size selection method based on differential evolution, thereby further improving the coding efficiency.

[0010] The present invention also provides a three-dimensional point cloud prediction geometry coding system based on deep learning.

[0011] Terminology explanation:

[0012] 1. High bit rate mode: The high bit rate mode compresses the original point cloud into a larger bit stream, so that the fidelity of the reconstructed point cloud after decoding is higher.

[0013] 2. Entropy coding, that is, coding without losing any information during the encoding process according to the entropy principle.

[0014] The technical solution of the present invention is:

[0015] A 3D point cloud prediction geometry coding method based on deep learning, including:

[0016] 1) Construct a prediction tree; for a point cloud file containing the pitch angles of each laser scanner of the laser radar and the height of each laser scanner coordinate system relative to the radar coordinate system, a prediction tree is constructed based on radar parameters; otherwise, for a point cloud file in which the coordinates of the points are arranged in the order of the height of the laser scanner and the acquisition order of the laser scanner, a prediction tree is constructed using the threshold segmentation method;

[0017] 2) Predictive coding;

[0018] In high bitrate mode, in each prediction tree, starting from the root node, the coordinates of the next point are predicted one by one, and the quantized residuals are encoded. When the coordinates of the coded point in the spherical coordinate system are obtained, the radial distance and azimuth are differentially encoded first. At the same time, the relationship between the pitch angle and the radius is fitted through the pitch angle prediction model based on the LSTM network, and the reconstructed coordinates of the coded point and the pitch angle and radial distance of the current point are used to predict the pitch angle, and the quantized prediction residuals are encoded.

[0019] In low bitrate mode, the azimuth angle is encoded in the same way as in high bitrate mode, and the elevation angle is represented by differential encoding of the reconstructed value; for the radius r, it is compressed by an autoencoder based on an entropy model;

[0020] 3) Selecting encoding parameters based on differential evolution;

[0021] Through the quantization step selection method based on differential evolution, the coding parameter selection problem is transformed into an optimization problem under constraints. The approximately optimal coding parameters are obtained through iterative optimization on a small data set containing only ten point clouds, and finally other point clouds are encoded using these coding parameters.

[0022] Preferably, according to the present invention, the problem of selecting the encoding parameters is transformed into an optimization problem under constraints, including:

[0023] The problem of selecting encoding parameters is transformed into a single-objective optimization problem under bit rate constraints:

[0024]

[0025] Where Q is a set of quantization parameters, Q = (q θ ,q r ,q δ ,φ unit ), q θ represents the quantization step size for the pitch angle θ, q r represents the quantization step size for radius R, q δ ,v unit is a parameter used to represent the azimuth angle, D g (Q,P i) indicates that the reconstructed point cloud is in the point cloud P i The mean square error measured on θ Indicates the bit rate of the coded elevation angle, R r represents the bit rate of the coding radius, R φ Indicates the bit rate of the coded azimuth, R(Q,P i ) represents the point cloud P i The total bit rate is also tested on a small dataset.

[0026] Preferably, according to the present invention, a prediction tree is constructed based on radar parameters; comprising:

[0027] The process of converting the coordinates (x, y, z) of point n in the Cartesian coordinate system to the spherical coordinate system is expressed as:

[0028]

[0029] φ=atan(y,x);

[0030]

[0031] Where r represents the radial distance, φ represents the azimuth angle, i and j represent the numbers of the laser scanners, represents the height of laser scanner j under the lidar scanner, θ(j) represents the preset pitch angle of laser scanner j, and N represents the number of laser scanners;

[0032] Through the above calculation, the radial distance, azimuth and laser scanner number i of point n in the spherical coordinate system are obtained;

[0033] According to the laser scanner number i, the height θ(i) of the laser scanner under the laser radar scanner is obtained, and the pitch angle θ of point n is calculated:

[0034]

[0035] At this point, the coordinates of point n in the spherical coordinate system are obtained (φ, θ, r);

[0036] According to the laser scanner number i of each point, all points are divided into N groups. Within each group, they are sorted according to the value of φ. Each group of points constitutes a prediction tree.

[0037] Preferably, according to the present invention, a threshold segmentation method is used to construct a prediction tree; comprising:

[0038] First, transform the coordinates (x, y, z) of point n in the Cartesian coordinate system to the spherical coordinate system:

[0039]

[0040] φ=atan(y,x);

[0041] θ1=atan(z,r);

[0042] At this point, the coordinates of point n in the spherical coordinate system are obtained (φ, θ1, r);

[0043] Next, group the points, which means: starting from the azimuth φ of the first point in the point cloud file 0 Start traversal and calculate the threshold difference between the current point and its adjacent points in the order of point cloud file storage point by point; when point n in the point cloud file i The azimuth of point n i=0 When the difference in the azimuth angle is greater than the preset threshold, they are grouped here, and each group constitutes a prediction tree:

[0044]

[0045] Among them, G . represents the jth group, G .A0 represents the j+1th group, t represents the preset threshold, Represents the azimuth of the i-th point in the point cloud file, Indicates the azimuth of the i-1th point in the point cloud file.

[0046] After the prediction tree is constructed by the above method, the encoding process begins. Preferably, the radial distance is differentially encoded; including:

[0047] First, when differentially encoding the radial distance, the radial distance r of the root node in each prediction tree 0 Directly perform entropy coding;

[0048] Then, click n i The predicted value of the radial distance Through differential encoding, we get:

[0049]

[0050] in, Indicates point n i=0 The reconstructed value of the radial distance is then point n i The prediction residual and quantization residual of the radial distance are:

[0051]

[0052] Among them, res r,i is the prediction residual of the radial distance of the point cloud, q r is the quantization step size for radial distance, is the quantized residual; then point n iThe reconstructed value of the radial distance It is expressed as:

[0053]

[0054] Finally, for r 0 and the set of quantized residuals of radial distances Perform entropy coding.

[0055] Preferably, according to the present invention, differential encoding is performed on the azimuth angle; comprising:

[0056] When encoding the azimuth of the point cloud, the azimuth of each point is φ i It is expressed as:

[0057] φ i =φ unit ×s i +δ i ;

[0058] Among them, δ i Represents an angle, s i is an integer, φ unit Indicates the set unit azimuth, φ unit Set to the LiDAR azimuth resolution x is a preset integer;

[0059] At this time, for φ i The encoding is converted into s i and δ i The encoding of

[0060] For s i , differential coding is followed by entropy coding; for δ i , then quantize first and then perform entropy coding;

[0061]

[0062] in, Denotes δ i The quantitative value of is the reconstruction value, q δ is the quantization step size; φ i The reconstruction value It is expressed as:

[0063]

[0064] The data that needs entropy encoding is: s i The differential encoding and δ i The set of quantized residuals

[0065] Preferably, according to the present invention, the relationship between the pitch angle and the radius is fitted by a pitch angle prediction model based on an LSTM network, the pitch angle is predicted by using the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point, and the quantized prediction residual is encoded; including:

[0066] The pitch angle prediction model based on LSTM network is used to fit the relationship between pitch angle and radius under the prediction tree sequence. The pitch angle prediction model based on LSTM network includes 3 layers of LSTM network and 5 layers of fully connected network. At the same time, parallel calculation of multiple prediction trees is realized.

[0067] First, when encoding a point n in a prediction tree i When the pitch angle is i=0 to n i=[\ The reconstruction information of each point is input into the LSTM network. j∈[i-50,i-1], where, l . Indicates point n . The prediction tree number to which it belongs; fill in the missing points by adding padding;

[0068] After feature extraction from the LSTM network and feature aggregation from the fully connected layer, the partial reconstruction information of the current point Stitched together, are the reconstruction values ​​of the current point, For point n i=0 Reconstructed value of the pitch angle;

[0069] Subsequently, the concatenated features are aggregated again through two layers of fully connected networks, and the final output point n i Predicted value of pitch angle Then point n i The predicted residual and quantized residual of the pitch angle are:

[0070]

[0071] Among them, res θ,i is the predicted residual of the pitch angle, q θ is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the pitch angle is expressed as:

[0072]

[0073] Finally, for θ 0 And the set of quantized residuals of the pitch angle Perform entropy coding.

[0074] According to the preferred embodiment of the present invention, the pitch angle prediction model based on the LSTM network is trained by calculating the prediction value and the true value θ i The MSE loss between them constructs the loss function l dML :

[0075]

[0076] Where N represents the number of points in the point cloud.

[0077] Preferably, according to the present invention, in the low bit rate mode, the pitch angle is represented by differential coding of the reconstructed value; including:

[0078] Point n i The predicted value of the pitch angle Through differential encoding, we get:

[0079]

[0080] in, Indicates point n i=0 The reconstructed value of the radial distance is then point n i The prediction residual and quantization residual of the radial distance are:

[0081]

[0082] Among them, res θ,i is the prediction residual of the radial distance of the point cloud, q θ is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the radial distance is expressed as:

[0083]

[0084] Finally, for θ 0 and the set of quantized residuals of radial distances Perform entropy coding; the pitch angle is expressed as

[0085] Preferably, according to the present invention, the specific implementation process of the quantization step size selection method based on differential evolution includes:

[0086] Step 1: Population initialization

[0087] First, generate i individuals X in the 4-dimensional feasible solution space to form the initial population;

[0088] X(0)=(X 0 ,X 2 ,X T ,…,Xi )=S(x 0,0 ,x 0,2 ,x 0,T ,x 0,k ),…,(x i,0 ,x i,2 ,x i,T ,x i,k )U;

[0089] In the formula, X(0) represents the initial population, X i represents the i-th individual in the population, x i,. is the jth element of the i-th individual in the population; where X 0 One-to-one correspondence with the elements in Q, that is, x i,0 Corresponding to q θ , that is, x i,2 Corresponding to q r , x i,T Corresponding to q δ , x i,k Corresponding to φ unit , unified with X i Represents the individuals of the population; based on empirical values, set:

[0090]

[0091] After the population is randomly initialized according to the above-set range, the fitness value of each individual X is first calculated, that is, ∑ i D g (Q,P i ) and the corresponding bit rate ∑ i R(Q,P i );For∑ i R(Q,P i ) is greater than R # The fitness value of the individual is reassigned to positive infinity;

[0092] Step 2: Mutation Operation

[0093] First, randomly select two different individuals to make a difference, as shown in the following formula:

[0094] B i,. =X i (t)-X . (t),

[0095] X i (t) represents the individual in the population after t iterations, i and j represent the individual numbers in the population after t iterations, and the difference vector B i,. After weighted processing and another iteration t times, the individual X p (t) sum, and we get the offspring variant individual V i(t+1); the specific mutation operations are as follows:

[0096] V i (t+1)=X p (t)+μB i,. ,

[0097] Where i is the current target individual number, i, j and k are the numbers of individuals in the t-th generation population that are different from the current target number in X(t), that is, i≠j≠k; X p (t) is the mutation vector V i (t+1) basis vector; μ∈[0,2] is the mutation scale factor parameter of the differential evolution algorithm;

[0098] Step 3: Crossover Operation

[0099] Test individual U i Every element u in (t+1) i,. The calculation of (t+1) is as follows:

[0100]

[0101] In the formula, v i,. (t+1) is the variant individual V i The j-th dimension component of (t+1), x i,. (t) is the j-th dimension component of the target individual in the parent population, i = 1, 2, ..., NP, j = 1, 2, ..., d; r i,. is the random number corresponding to the j-dimensional component and satisfies the normal distribution between (0,1); CR is the crossover probability factor of the algorithm, with a value range of [0,1], m∈{1,2,..,d}, ensuring that the experimental individual U i At least one dimension in (t+1) comes from the variant individual V i (t+1);

[0102] When the random number r corresponding to the jth component of the population individual i,. When the crossover rate CR is less than or j = m, the test individual U i The j-th dimension component of (t+1) is composed of the variant individual V i (t+1) is provided, otherwise the parent target individual X i (t) provide;

[0103] Step 4: Select an action

[0104] According to the experimental individual U i (t+1) and the parent target individual X iThe fitness value of (t) and whether it meets the constraint condition are determined, wherein the constraint condition is the bit rate measured on the small data set, and the fitness value is the total distortion measured on the small data set; individuals that meet the bit rate constraint and have a better fitness value are selected to enter the next generation population;

[0105]

[0106] Through multiple rounds of iterations of the differential algorithm, the approximately optimal quantization parameter Q under the bit rate is obtained;

[0107] The quantization parameter Q is then used directly to compress the entire point cloud at the bit rate.

[0108] A 3D point cloud prediction geometry coding system based on deep learning, including:

[0109] The prediction tree construction module is configured to: for a point cloud file containing the pitch angles of each laser scanner of the laser radar and the height of each laser scanner coordinate system relative to the radar coordinate system, construct a prediction tree based on radar parameters; otherwise, for a point cloud file in which the coordinates of the points are arranged in the point cloud file according to the height order of the laser scanner and the acquisition order of the laser scanner, construct a prediction tree using a threshold segmentation method;

[0110] The prediction coding module is configured as follows: in high bit rate mode, in each prediction tree, starting from the root node, the coordinates of the next point are predicted one by one, and the quantized residual is encoded; when the coordinates of the encoded point in the spherical coordinate system are obtained, the radial distance and azimuth are differentially encoded first; at the same time, the relationship between the pitch angle and the radius is fitted by the pitch angle prediction model based on the LSTM network, and the pitch angle is predicted by the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point, and the quantized prediction residual is encoded; in low bit rate mode, the azimuth is encoded by the same method as in the high bit rate mode, and the pitch angle is represented by the differential encoding of the reconstructed value; for the radius r, it is compressed by an autoencoder based on the entropy model;

[0111] The coding parameter selection module is configured as follows: through a quantization step selection method based on differential evolution, the coding parameter selection problem is transformed into an optimization problem under constraints, and the approximately optimal coding parameters are obtained through iterative optimization on a small data set containing only ten point clouds, and finally other point clouds are encoded using the coding parameters.

[0112] The beneficial effects of the present invention are:

[0113] 1. The present invention improves the rate-distortion performance of Lidar point cloud encoding, that is, compared with previous methods, at the same bit rate, the decoded reconstructed point cloud has higher fidelity.

[0114] 2. The present invention proposes a method for calculating the quantization parameters of the residuals of each coordinate axis of the point cloud in a spherical coordinate system, which can further improve the rate-distortion performance of point cloud coding and can also be applied in similar spherical coordinate coding methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] Figure 1 It is a block diagram of the implementation of the three-dimensional point cloud prediction geometry coding method based on deep learning of the present invention;

[0116] Figure 2 is a schematic diagram of the relationship between the radius and the pitch angle of a point originating from the same laser;

[0117] Figure 3 is a second schematic diagram of the relationship between the radius and the pitch angle of a point originating from the same laser;

[0118] Figure 4 is a third schematic diagram of the relationship between the radius and the pitch angle of a point originating from the same laser;

[0119] Figure 5 It is a fourth schematic diagram of the relationship between the radius and the pitch angle of a point originating from the same laser;

[0120] Figure 6 It is a structural schematic diagram of the pitch angle prediction model based on LSTM network proposed in the present invention;

[0121] Figure 7 It is a schematic diagram of the structure of the radius compression model in low bit rate mode;

[0122] Figure 8 This is a schematic diagram of the rate-distortion curve based on D1-PSNR measured on the Semantic dataset;

[0123] Fig. 9 This is a second schematic diagram of the rate-distortion curve based on D1-PSNR measured on the Semantic dataset;

[0124] Fig.10 It is a schematic diagram of the rate-distortion curve based on D2-PSNR measured on the FORD dataset;

[0125] Fig.11 This is a second schematic diagram of the rate-distortion curve based on D2-PSNR measured on the FORD dataset. DETAILED DESCRIPTION

[0126] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.

[0127] Example 1

[0128] 3D point cloud prediction geometry encoding method based on deep learning, such as Figure 1 As shown, including:

[0129] 1) Construct a prediction tree; because the lidar point cloud has different annotations about the acquisition device in addition to geometry and attributes, and the order of the geometric coordinates of the points in the point cloud file is also different, this method has two built-in prediction tree construction methods. For point cloud files containing the pitch angles of each lidar laser scanner and the height of each lidar scanner coordinate system relative to the radar coordinate system, such as the Ford point cloud dataset, a prediction tree is constructed based on radar parameters; otherwise, for point cloud files that do not contain the above-mentioned related parameters, and for point cloud files in which the coordinates of the points are arranged in the order of the height of the laser scanner and the acquisition order of the laser scanner, such as the SemanticKITTI dataset, a threshold segmentation method is used to construct a prediction tree;

[0130] 2) Predictive coding;

[0131] In high bit rate mode, after the prediction tree is constructed, in each prediction tree, starting from the root node, the coordinates of the next point are predicted one by one in order, and the quantized residuals are encoded; when the coordinates of the encoded point in the spherical coordinate system are obtained, the radial distance and azimuth are differentially encoded first; at the same time, the relationship between the pitch angle and the radius is fitted by the pitch angle prediction model based on the LSTM network, and the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point are used to predict the pitch angle, and the quantized prediction residuals are encoded; since the prediction coding only depends on the information in the prediction tree, multiple prediction trees can be encoded in parallel.

[0132] In low bitrate mode, the correlation between points is destroyed, and it is difficult for predictive coding to accurately predict the coordinates of adjacent points. Therefore, in low bitrate mode, the azimuth angle is encoded in the same way as in high bitrate mode, and the elevation angle is represented by differential coding of the reconstructed value; for the radius r, it is compressed by an autoencoder based on an entropy model;

[0133] 3) Selecting encoding parameters based on differential evolution;

[0134] In the predictive coding process, there are θ ,q r ,q δ ,φ unitFour parameters that affect the rate-distortion performance of point cloud coding. In the quantization of coordinate residuals in the Cartesian coordinate system, the residuals from different coordinate axes can be quantized using the same quantization step size, because the residuals of different coordinate axes in the Cartesian coordinate system have the same effect on the distortion of the reconstructed point cloud. However, in spherical coordinates, the pitch angle, azimuth angle, and radial distance need to be converted to the Cartesian coordinate system first. This conversion is nonlinear, which makes the residuals of the above spherical coordinates have different effects on the distortion of the reconstructed point cloud. The same quantization step size cannot obtain the optimal quantization result. However, the quantization step size has a large value range, forming a lot of combinations. It will undoubtedly take a very high time and computational cost to obtain the optimal quantization step size by exhaustive method. Therefore, through the quantization step size selection method based on differential evolution, the selection problem of coding parameters is transformed into an optimization problem under constraints, and the approximately optimal coding parameters are obtained through iterative optimization on a small data set containing only ten point clouds, and finally the coding parameters are used to encode other point clouds. Compared with the precoding method, this method has less rate-distortion performance loss, but can significantly improve the encoding speed.

[0135] Example 2

[0136] The three-dimensional point cloud prediction geometry coding method based on deep learning is different in that:

[0137] The problem of selecting encoding parameters is transformed into a single-objective optimization problem under bit rate constraints:

[0138]

[0139] Where Q is a set of quantization parameters, Q = (q θ ,q r ,q δ ,φ unit ), q θ represents the quantization step size for the pitch angle θ, q r represents the quantization step size for radius R, q δ ,φ unit is a parameter used to represent the azimuth angle, D g (Q,P i ) indicates that the reconstructed point cloud is in the point cloud P i The mean square error measured on θ Indicates the bit rate of the coded elevation angle, R r represents the bit rate of the coding radius, R φ Indicates the bit rate of the coded azimuth, R(Q,P i ) represents the point cloud P i The total bit rate is also tested on a small dataset.

[0140] Build a prediction tree based on radar parameters; including:

[0141] This method is the same as the prediction tree construction method in the G-PCC point cloud coding standard. The process of converting the coordinates (x, y, z) of point n in the Cartesian coordinate system to the spherical coordinate system is expressed as:

[0142]

[0143] φ=atan(y,x);

[0144]

[0145] Where r represents the radial distance, φ represents the azimuth angle, i and j represent the numbers of the laser scanners, represents the height of laser scanner j under the lidar scanner, θ(j) represents the preset pitch angle of laser scanner j, and N represents the number of laser scanners;

[0146] Through the above calculation, the radial distance, azimuth and laser scanner number i of point n in the spherical coordinate system are obtained;

[0147] According to the laser scanner number i, the height θ(i) of the laser scanner under the laser radar scanner is obtained, and the pitch angle θ of point n is calculated:

[0148]

[0149] At this point, the coordinates of point n in the spherical coordinate system are obtained (φ, θ, r);

[0150] According to the laser scanner number i of each point, all points are divided into N groups. Within each group, they are sorted according to the value of φ, and each group of points constitutes a prediction tree. This method makes full use of the parameter information of the laser radar, and the constructed prediction tree is the most accurate. For point cloud files containing the pitch angles of each laser scanner of the laser radar and the height of each laser scanner coordinate system relative to the radar coordinate system, the prediction tree construction method based on radar parameters is preferred.

[0151] The prediction tree is constructed using the threshold segmentation method; including:

[0152] This method is applicable to point cloud files where the coordinates of the points are arranged in the order of the laser scanner and the acquisition order inside the laser scanner. First, convert the coordinates (x, y, z) of point n in the Cartesian coordinate system to the spherical coordinate system:

[0153]

[0154] φ=atan(y,x);

[0155] θ1=atan(z,r);

[0156] At this point, the coordinates of point n in the spherical coordinate system are obtained (φ, θ1, r);

[0157] Next, we group the points. Similar to the above method, we need to group the points from the same laser scanner into the same group. In the point cloud file, points from the same laser scanner are arranged adjacent to each other, and points from the same laser scanner are arranged in the order of acquisition (i.e., azimuth from small to large). Therefore, we only need to find the boundary between the points from the two laser scanners to achieve grouping. i With point n i=0 They are the first and last points scanned by two different laser scanners in a scanning cycle, so the azimuth difference between the two points is often large. i The azimuth of point n i=0 When the difference between the azimuth angles is greater than the set threshold, it can be divided into two different laser scanners. It means: the azimuth angle φ of the first point in the point cloud file 0 Start traversal and calculate the threshold difference between the current point and its adjacent points in the order of point cloud file storage point by point; when point n in the point cloud file i The azimuth of point n i=0 When the difference in the azimuth angle is greater than the preset threshold, they are grouped here, and each group constitutes a prediction tree:

[0158]

[0159] Among them, G . represents the jth group, G .A0 represents the j+1th group, t represents the preset threshold, which is generally set to 60. Represents the azimuth of the i-th point in the point cloud file, Indicates the azimuth of the i-1th point in the point cloud file. This method utilizes the potential laser scanner grouping information in the point cloud file. Although the final division result is not accurate enough and there may be an erroneous division in which the number of predicted trees is not equal to the number of actual laser scanners, a large number of points from the same laser scanner will still be divided into the same prediction tree.

[0160] In high bit rate mode, radial distance is differentially encoded; including:

[0161] First, when differentially encoding the radial distance, the radial distance r of the root node in each prediction tree 0 Directly perform higher precision entropy coding;

[0162] Then, click n i The predicted value of the radial distance Through differential encoding, we get:

[0163]

[0164] in, Indicates point n i=0 The reconstructed value of the radial distance is then point n i The prediction residual and quantization residual of the radial distance are:

[0165]

[0166] Among them, res r,i is the prediction residual of the radial distance of the point cloud, q r is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the radial distance It is expressed as:

[0167]

[0168] Finally, for r 0 and the set of quantized residuals of radial distances Perform entropy coding.

[0169] Differential encoding of azimuth; including:

[0170] When encoding the azimuth of the point cloud, the azimuth of each point is φ i It is expressed as:

[0171] φ i =φ unit ×s i +δ i ;

[0172] Among them, δ i Represents an angle, s i is an integer, φ unit Indicates the set unit azimuth, φ unit Set to the LiDAR azimuth resolution x is a preset integer; the advantage of this representation is that there is no empty part s in the prediction tree i With s i=0 The difference between them is often an integer multiple of x or x, and the information entropy is small. At the same time, the difference between the azimuths is processed into an integer that is easy to encode, reducing δ i The number of bits consumed by the encoding.

[0173] At this time, for φ i The encoding is converted into s i and δ i The encoding of

[0174] For s i , differential coding is followed by entropy coding; for δ i , then quantize first and then perform entropy coding;

[0175]

[0176] in, Denotes δ i The quantitative value of is the reconstruction value, q δ is the quantization step size; φ i The reconstruction value It is expressed as:

[0177]

[0178] The data that needs entropy encoding is: s i The differential encoding and δ i The set of quantized residuals It is also the result of differential encoding of the azimuth angle.

[0179] The relationship between the pitch angle and radius is fitted through the pitch angle prediction model based on the LSTM network, the pitch angle is predicted using the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point, and the quantized prediction residual is encoded; including:

[0180] like Figure 2 , Figure 3 , Figure 4 and Figure 5 As shown in the figure, the abscissa represents the radius and the ordinate represents the pitch angle. The radius and pitch angle of the point cloud show a significant negative correlation. This relationship is mainly caused by the calibration process inside the lidar. However, due to the internal errors of the lidar and the influence of the external environment, it is difficult to effectively model the relationship between the radius and pitch angle of the point cloud only by function fitting.

[0181] Through the pitch angle prediction model based on LSTM network, the relationship between pitch angle and radius under the prediction tree sequence is fitted through the long- and short-range modeling capabilities of LSTM network. The pitch angle prediction model based on LSTM network is based on lightweight design, such as Figure 6 As shown in FIG. 1 , the pitch angle prediction model based on the LSTM network includes a 3-layer LSTM network and a 5-layer fully connected network; the LSTM network and the fully connected network are common and widely used neural network module structures in this field, and simultaneously realize the parallel calculation of multiple prediction trees;

[0182] First, when encoding a point n in a prediction tree i When the pitch angle isi=0 to n i=[\ The reconstruction information of each point is input into the LSTM network. j∈[i-50,i-1], where, l . Indicates point n . The prediction tree number to which it belongs; since the first 50 points under the prediction tree do not have enough encoded points as the input of the neural network, the missing points are supplemented by adding padding; specifically, the first point under the prediction tree can be decoded without relying on adjacent points, so 49 points are generated according to the coordinates of the first point and placed in front of the first point of the prediction tree to supplement the input of 50 points.

[0183] After feature extraction from the LSTM network and feature aggregation from the fully connected layer, the partial reconstruction information of the current point Stitched together, are the reconstruction values ​​of the current point, For point n i=0 Reconstructed value of the pitch angle;

[0184] Subsequently, the concatenated features are aggregated again through two layers of fully connected networks, and the final output point n i Predicted value of pitch angle Then point n i The predicted residual and quantized residual of the pitch angle are:

[0185]

[0186] Among them, res θ,i is the predicted residual of the pitch angle, q θ is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the pitch angle is expressed as:

[0187]

[0188] Finally, for θ 0 And the set of quantized residuals of the pitch angle Perform entropy coding.

[0189] When training the pitch angle prediction model based on the LSTM network, the prediction value is calculated and the true value θ i The MSE loss between them constructs the loss function l dML :

[0190]

[0191] Where N represents the number of points in the point cloud.

[0192] In low bitrate mode, the pitch angle is represented by differential encoding of the reconstructed value; including:

[0193] Point n i The predicted value of the pitch angle Through differential encoding, we get:

[0194]

[0195] in, Indicates point n i=0 The reconstructed value of the radial distance is then point n i The prediction residual and quantization residual of the radial distance are:

[0196]

[0197] Among them, res θ,i is the prediction residual of the radial distance of the point cloud, q θ is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the radial distance is expressed as:

[0198]

[0199] Finally, for θ 0 and the set of quantized residuals of radial distances Perform entropy coding; the pitch angle is expressed as

[0200] For radius r, compression is performed through an autoencoder that introduces an entropy model, such as Figure 7 As shown, the autoencoder consists of a convolutional neural network (CNN), a generalized normalization activation function (GDN), and a linear rectification function (ReLU) stack. It includes:

[0201] First, the radius values ​​of each point are arranged in the order of the octree and form a radius matrix R;

[0202] Then, the radius matrix R is input into the autoencoder and the latent representation Y is extracted;

[0203] Subsequently, the entropy model models the distribution of Y and entropy encodes Y;

[0204] At the decoding end, the decoder obtains the reconstructed matrix through the reconstructed Y Thus the radius of the point cloud is reconstructed.

[0205] The specific implementation process of the quantization step size selection method based on differential evolution includes:

[0206] Step 1: Population initialization

[0207] In the present invention, the differential evolution algorithm is mainly used in the case where Q = (q θ, q r ,q δ ,φ unit ) in the 4-dimensional real parameter space. First, i individuals X are generated in the 4-dimensional feasible solution space to form an initial population. Here, the generated population size is 20.

[0208] X(0)=(X 0 ,X 2 ,X T ,…,X i )=S(x 0,0 ,x 0,2 ,x 0,T ,x 0,k ),…,(x i,0 ,x i,2 ,x i,T ,x i,k )U;

[0209] In the formula, X(0) represents the initial population, X i represents the i-th individual in the population, x i,. is the jth element of the i-th individual in the population; where X 0 One-to-one correspondence with the elements in Q, that is, x i,0 Corresponding to q θ , that is, x i,2 Corresponding to q r , x i,T Corresponding to q δ , x i,k Corresponding to φ unit , unified with X i Represents the individuals of the population; based on empirical values, set:

[0210]

[0211] After the population is randomly initialized according to the above-set range, the fitness value of each individual X is first calculated, that is, ∑ i D g (Q,P i ) and the corresponding bit rate ∑ i R(Q,P i );For∑ i R(Q,P i ) is greater than R # The fitness value of the individual is reassigned to positive infinity;

[0212] Step 2: Mutation Operation

[0213] In order to maintain the diversity of the population, each individual needs to be mutated. The mutation operation of the differential evolution algorithm uses the differences between individuals to add perturbations to other individuals. First, two different individuals are randomly selected for difference, as shown in the following formula:

[0214] B i,. =X i (t)-X . (t),

[0215] X i (t) represents the individual in the population after t iterations, i and j represent the individual numbers in the population after t iterations, and the difference vector B i,. After weighted processing and another iteration t times, the individual X p (t) sum, and we get the offspring variant individual V i (t+1); the specific mutation operations are as follows:

[0216] V i (t+1)=X p (t)+μB i,. ,

[0217] Where i is the current target individual number, i, j and k are the numbers of individuals in the t-th generation population that are different from the current target number in X(t), that is, i≠j≠k; X p (t) is the mutation vector V i (t+1) basis vector; μ∈[0,2] is the mutation scale factor parameter of the differential evolution algorithm; when V i When the value of the element in (t+1) is greater than the preset value range, the maximum value of the value range is taken; when the value is less than the preset value range, the minimum value of the value range is taken. According to the number of individuals in the current population, the minimum population number NP of differential evolution is 4, otherwise the mutation operation cannot be performed.

[0218] Step 3: Crossover Operation

[0219] After the compilation operation, the population needs to be cross-operated. The mutant individuals and the target individuals exchange some of their components in a certain proportion to form the experimental individuals U i (t+1). In order to ensure that the target individual X i (t) To achieve evolution, it is necessary to ensure that the experimental individual U i At least one element in (t+1) comes from the variant individual V i (t+1). Test individual U i Every element u in (t+1) i,. The calculation of (t+1) is as follows:

[0220]

[0221] In the formula, v i,. (t+1) is the variant individual V i The j-th dimension component of (t+1), x i,. (t) is the j-th dimension component of the target individual in the parent population, i = 1, 2, ..., NP, j = 1, 2, ..., d; r i,. is the random number corresponding to the j-dimensional component and satisfies the normal distribution between (0,1); CR is the crossover probability factor of the algorithm, with a value range of [0,1], which directly affects the variant individual V i (t+1) In the experimental individual U i The proportion of (t+1); m∈{1,2,..,d}, to ensure that the experimental individual U i At least one dimension in (t+1) comes from the variant individual V i (t+1);

[0222] When the random number r corresponding to the jth component of the population individual i,. When the crossover rate CR is less than or j = m, the test individual U i The j-th dimension component of (t+1) is composed of the variant individual V i (t+1) is provided, otherwise the parent target individual X i (t) provide;

[0223] Step 4: Select an action

[0224] In order to ensure that the algorithm approaches the global optimum, the fitness value selects the next generation of individuals. Specifically, according to the experimental individual U i (t+1) and the parent target individual X i The fitness value of (t) and whether it meets the constraint condition are determined, wherein the constraint condition is the bit rate measured on the small data set, and the fitness value is the total distortion measured on the small data set; individuals that meet the bit rate constraint and have a better fitness value are selected to enter the next generation population;

[0225]

[0226] Through multiple rounds of iterations of the differential algorithm, the approximately optimal quantization parameter Q under the bit rate is obtained;

[0227] The quantization parameter Q is then used directly to compress the entire point cloud at the bit rate.

[0228] Table 1 is a comparison table of the rate-distortion performance of the present invention and other methods on the FORD dataset.

[0229] Table 1

[0230]

[0231] Table 2 is a comparison of the rate-distortion performance of the present invention and other methods on the SemanticKITTI dataset.

[0232] Table 2

[0233]

[0234] like Figure 8 , Fig. 9 , Fig.10 , Fig.11 As shown in Table 1 and Table 2, the horizontal axis represents the bit rate and the vertical axis represents the peak signal-to-noise ratio. Combined with Table 1 and Table 2, compared with other methods, the method proposed in the present invention achieves the best rate-distortion performance. On the Ford dataset, compared with the currently best performing SCP-EHEM, D1BD-Rate is reduced by 6.8%, and D2BD-Rate is reduced by 8.5%. Compared with the point cloud coding standard G-PCC, D1BD-Rate is reduced by 24.9%, and D2BD-Rate is reduced by 21.7%. On the SemanticKITTI dataset, compared with SCP-EHEM, D1BD-Rate is reduced by 2.5%, and D2BD-Rate is reduced by 1.9%. Compared with G-PCC, the BD-Rates based on D1 and D2 are reduced by -23.1% and -23.9%, respectively.

[0235] Example 3

[0236] A 3D point cloud prediction geometry coding system based on deep learning, including:

[0237] The prediction tree construction module is configured as follows: Since the lidar point cloud has different annotation contents about the acquisition device related information in addition to geometry and attributes, and the arrangement order of the geometric coordinates of the points in the point cloud file is also different, this method has two built-in prediction tree construction methods. For point cloud files containing the pitch angles of each lidar laser scanner and the height of each lidar scanner coordinate system relative to the radar coordinate system, such as the Ford point cloud dataset, a prediction tree is constructed based on the radar parameters; otherwise, for point cloud files that do not contain the above-mentioned related parameters and for point cloud files in which the coordinates of the points are arranged in the order of the height of the laser scanner and the acquisition order of the laser scanner, such as the SemanticKITTI dataset, a threshold segmentation method is used to construct a prediction tree;

[0238] The prediction coding module is configured as follows: in high bit rate mode, after the prediction tree is constructed, in each prediction tree, starting from the root node, the coordinates of the next point are predicted one by one, and the quantized residual is encoded; when the coordinates of the encoded point in the spherical coordinate system are obtained, the radial distance and azimuth are differentially encoded first; at the same time, the relationship between the pitch angle and the radius is fitted by the pitch angle prediction model based on the LSTM network, and the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point are used to predict the pitch angle, and the quantized prediction residual is encoded; since the prediction coding only depends on the information in the prediction tree, multiple prediction trees can be encoded in parallel. In low bit rate mode, since the correlation between points is destroyed, it is difficult for the prediction coding to accurately predict the coordinates of adjacent points. Therefore, at low bit rate, the azimuth is encoded in the same way as at high bit rate, and the pitch angle is represented by differential encoding of the reconstructed value; for the radius r, it is compressed by an autoencoder based on an entropy model;

[0239] The encoding parameter selection module is configured to: in the predictive encoding process, there are q θ ,q r ,q δ ,φ unit Four parameters that affect the rate-distortion performance of point cloud coding. In the quantization of coordinate residuals in the Cartesian coordinate system, the residuals from different coordinate axes can be quantized using the same quantization step size, because the residuals of different coordinate axes in the Cartesian coordinate system have the same effect on the distortion of the reconstructed point cloud. However, in spherical coordinates, the pitch angle, azimuth angle, and radial distance need to be converted to the Cartesian coordinate system first. This conversion is nonlinear, which makes the residuals of the above spherical coordinates have different effects on the distortion of the reconstructed point cloud. The same quantization step size cannot obtain the optimal quantization result. However, the quantization step size has a large value range, forming a lot of combinations. It will undoubtedly take a very high time and computational cost to obtain the optimal quantization step size by exhaustive method. Therefore, through the quantization step size selection method based on differential evolution, the selection problem of coding parameters is transformed into an optimization problem under constraints, and the approximately optimal coding parameters are obtained through iterative optimization on a small data set containing only ten point clouds, and finally the coding parameters are used to encode other point clouds. Compared with the precoding method, this method has less rate-distortion performance loss, but can significantly improve the encoding speed.

Claims

1. A three-dimensional point cloud prediction geometric coding method based on deep learning, characterized in that: include: 1) Construct a prediction tree; For point cloud files containing the pitch angles of each laser scanner of the laser radar and the height of each laser scanner coordinate system relative to the radar coordinate system, a prediction tree is constructed based on the radar parameters; otherwise, for point cloud files in which the coordinates of the points are arranged in the order of the height of the laser scanner and the acquisition order of the laser scanner in the point cloud file, a prediction tree is constructed using the threshold segmentation method; 2) Predictive coding; In high bitrate mode, in each prediction tree, starting from the root node, the coordinates of the next point are predicted one by one, and the quantized residuals are encoded. When the coordinates of the coded point in the spherical coordinate system are obtained, the radial distance and azimuth are differentially encoded first. At the same time, the relationship between the pitch angle and the radius is fitted through the pitch angle prediction model based on the LSTM network, and the reconstructed coordinates of the coded point and the pitch angle and radial distance of the current point are used to predict the pitch angle, and the quantized prediction residuals are encoded. In low bit rate mode, the azimuth angle is encoded, and the elevation angle is represented by differential encoding of the reconstructed value; for the radius r, it is compressed by an autoencoder based on the entropy model; 3) Selecting encoding parameters based on differential evolution; Through the quantization step selection method based on differential evolution, the coding parameter selection problem is transformed into an optimization problem under constraints. The approximately optimal coding parameters are obtained through iterative optimization on a small data set containing only ten point clouds, and finally other point clouds are encoded using these coding parameters.

2. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: The problem of selecting encoding parameters is transformed into an optimization problem under constraints, including: The problem of selecting encoding parameters is transformed into a single-objective optimization problem under bit rate constraints: Where Q is a set of quantization parameters, Q = (q θ ,q r ,q δ ,φ unit ), q θ represents the quantization step size for the pitch angle θ, q r represents the quantization step size for radius R, q δ ,v unit is a parameter used to represent the azimuth angle, D g (Q,P i ) indicates that the reconstructed point cloud is in the point cloud P i The mean square error measured on θ Indicates the bit rate of the coded elevation angle, R r represents the bit rate of the coding radius, R φ Indicates the bit rate of the coded azimuth, R(Q,P i ) represents the point cloud P i The total bit rate is also tested on a small dataset.

3. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: Build a prediction tree based on radar parameters; including: The process of converting the coordinates (x, y, z) of point n in the Cartesian coordinate system to the spherical coordinate system is expressed as: φ=atan(y,x); Where r represents the radial distance, φ represents the azimuth angle, i and j represent the numbers of the laser scanners, represents the height of laser scanner j under the lidar scanner, θ(j) represents the preset pitch angle of laser scanner j, and N represents the number of laser scanners; Through the above calculation, the radial distance, azimuth and laser scanner number i of point n in the spherical coordinate system are obtained; According to the laser scanner number i, the height θ(i) of the laser scanner under the laser radar scanner is obtained, and the pitch angle θ of point n is calculated: At this point, the coordinates of point n in the spherical coordinate system are obtained (φ, θ, r); According to the laser scanner number i of each point, all points are divided into N groups. Within each group, they are sorted according to the value of φ. Each group of points constitutes a prediction tree. Further preferably, a threshold segmentation method is used to construct a prediction tree; comprising: First, transform the coordinates (x, y, z) of point n in the Cartesian coordinate system to the spherical coordinate system: φ=atan(y,x); θ1=atan(z,r); At this point, the coordinates of point n in the spherical coordinate system are obtained (φ, θ1, r); Next, grouping the points means: starting from the azimuth φ0 of the first point in the point cloud file, calculating the threshold difference between the current point and its adjacent points in the order of point cloud file storage point by point; when point n in the point cloud file i The azimuth of point n i=0 When the difference in the azimuth angle is greater than the preset threshold, they are grouped here, and each group constitutes a prediction tree: Among them, G j represents the jth group, G j+1 represents the j+1th group, t represents the preset threshold, Represents the azimuth of the i-th point in the point cloud file, Indicates the azimuth of the i-1th point in the point cloud file.

4. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: Differential encoding of radial distance; including: First, when differentially encoding the radial distance, the radial distance r0 of the root node in each prediction tree is directly entropy encoded; Then, click n i The predicted radial distance Through differential encoding, we get: in, Indicates point n i=0 The reconstructed value of the radial distance is then point n i The prediction residual and quantization residual of the radial distance are: Among them, res r,i is the prediction residual of the radial distance of the point cloud, q r is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the radial distance It is expressed as: Finally, the set of quantized residuals of r0 and radial distance is Perform entropy coding.

5. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: Differential encoding of azimuth; including: When encoding the azimuth of the point cloud, the azimuth of each point is φ i It is expressed as: f i =φ unit ×s i +d i ; Among them, δ i Represents an angle, s i is an integer, φ unit Indicates the set unit azimuth, φ unit Set to the LiDAR azimuth resolution x is a preset integer; At this time, for φ i The encoding is converted into s i and δ i The encoding of For s i , differential coding is followed by entropy coding; for δ i , then quantize first and then perform entropy coding; in, Denotes δ i The quantitative value of is the reconstruction value, q δ is the quantization step size; φ i The reconstruction value It is expressed as: The data that needs entropy encoding is: s i The differential encoding and δ i The set of quantized residuals 6. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: The relationship between the pitch angle and radius is fitted through the pitch angle prediction model based on the LSTM network, the pitch angle is predicted using the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point, and the quantized prediction residual is encoded; including: The pitch angle prediction model based on LSTM network is used to fit the relationship between pitch angle and radius under the prediction tree sequence. The pitch angle prediction model based on LSTM network includes 3 layers of LSTM network and 5 layers of fully connected network. At the same time, parallel calculation of multiple prediction trees is realized. First, when encoding a point n in a prediction tree i When the pitch angle is i=0 to n i=[\ The reconstruction information of each point is input into the LSTM network. Among them, l j Indicates point n j The prediction tree number to which it belongs; fill in the missing points by adding padding; After feature extraction from the LSTM network and feature aggregation from the fully connected layer, the partial reconstruction information of the current point Stitched together, l i are the reconstruction values ​​of the current point, For point n i=0 Reconstructed value of the pitch angle; Subsequently, the concatenated features are aggregated again through two layers of fully connected networks, and the final output point n i Predicted value of pitch angle Then point n i The predicted residual and quantized residual of the pitch angle are: Among them, res θ,i is the predicted residual of the pitch angle, q θ is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the pitch angle is expressed as: Finally, the set of quantized residuals of θ0 and pitch angle is Perform entropy coding.

7. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: When training the pitch angle prediction model based on the LSTM network, the prediction value is calculated and the true value θ i The MSE loss between them constructs the loss function l dML : Where N represents the number of points in the point cloud.

8. The three-dimensional point cloud prediction geometry coding method based on deep learning according to claim 1, characterized in that: In low bit rate mode, the pitch angle is represented by differential encoding of the reconstructed value; include: Point n i The predicted value of the pitch angle Through differential encoding, we get: in, Indicates point n i=0 The reconstructed value of the radial distance is then point n i The prediction residual and quantization residual of the radial distance are: Among them, res θ,i is the prediction residual of the radial distance of the point cloud, q θ is the quantization step size for radial distance, is the quantized residual; then point n i The reconstructed value of the radial distance is expressed as: Finally, the set of quantized residuals of θ0 and radial distance is Perform entropy coding; the pitch angle is expressed as 9. The three-dimensional point cloud prediction geometric coding method based on deep learning according to any one of claims 1 to 8, characterized in that: The specific implementation process of the quantization step size selection method based on differential evolution includes: Step 1: Population initialization First, generate i individuals X in the 4-dimensional feasible solution space to form the initial population; X(0)=(X0,X2,X T ,…,X i )=S(x 0,0 ,x 0,2 ,x 0,T ,x 0,j ),…,(x i,0 ,x i,2 ,x i,T ,x i,j )U; In the formula, X(0) represents the initial population, X i represents the i-th individual in the population, x i,j is the jth element of the i-th individual in the population; where X0 corresponds one-to-one to the elements in Q, that is, x i,0 Corresponding to q θ , that is, x i,2 Corresponding to q r , x i,T Corresponding to q δ , x i,j Corresponding to φ unit , unified with X i Represents the individuals of the population; based on empirical values, set: After the population is randomly initialized according to the above-set range, the fitness value of each individual X is first calculated, that is, ∑ i D g (Q,P i ) and the corresponding bit rate ∑ i R(Q,P i );For∑ i R(Q,P i ) is greater than R # The fitness value of the individual is reassigned to positive infinity; Step 2: Mutation Operation First, randomly select two different individuals to make a difference, as shown in the following formula: B i,j =X i (t)-X j (t), X i (t) represents the individual in the population after t iterations, i and j represent the individual numbers in the population after t iterations, and the difference vector B i,j After weighted processing and another iteration t times, the individual X p (t) sum, and we get the offspring variant individual V i (t+1); the specific mutation operations are as follows: V i (t+1)=X p (t)+μB i,j , Where i is the current target individual number, i, j and k are the numbers of individuals in the t-th generation population that are different from the current target number in X(t), that is, i≠j≠k; X p (t) is the mutation vector V i (t+1) basis vector; μ∈[0,2] is the mutation scale factor parameter of the differential evolution algorithm; Step 3: Crossover Operation Test individual U i Every element u in (t+1) i,j The calculation of (t+1) is as follows: In the formula, v i,j (t+1) is the variant individual V i The j-th dimension component of (t+1), x i,j (t) is the j-th dimension component of the target individual in the parent population, i = 1, 2, ..., NP, j = 1, 2, ..., d; r i,j is the random number corresponding to the j-dimensional component and satisfies the normal distribution between (0,1); CR is the crossover probability factor of the algorithm, with a value range of [0,1], m∈{1,2,..,d}, ensuring that the experimental individual U i At least one dimension in (t+1) comes from the variant individual V i (t+1); When the random number r corresponding to the jth component of the population individual i,j When the crossover rate CR is less than or j = m, the test individual U i The j-th dimension component of (t+1) is composed of the variant individual V i (t+1) is provided, otherwise the parent target individual X i (t) provide; Step 4: Select an Action According to the experimental individual U i (t+1) and the parent target individual X i The fitness value of (t) and whether it meets the constraint condition are determined, wherein the constraint condition is the bit rate measured on the small data set, and the fitness value is the total distortion measured on the small data set; individuals that meet the bit rate constraint and have a better fitness value are selected to enter the next generation population; Through multiple rounds of iterations of the differential algorithm, the approximately optimal quantization parameter Q under the bit rate is obtained; The quantization parameter Q is then used directly to compress the entire point cloud at the bit rate.

10. A three-dimensional point cloud prediction geometry coding system based on deep learning, characterized in that: include: The prediction tree construction module is configured to: for a point cloud file containing the pitch angles of each laser scanner of the laser radar and the height of each laser scanner coordinate system relative to the radar coordinate system, construct a prediction tree based on radar parameters; otherwise, for a point cloud file in which the coordinates of the points are arranged in the point cloud file according to the height order of the laser scanner and the acquisition order of the laser scanner, construct a prediction tree using a threshold segmentation method; The prediction coding module is configured as follows: in high bit rate mode, in each prediction tree, starting from the root node, the coordinates of the next point are predicted one by one, and the quantized residual is encoded; when the coordinates of the encoded point in the spherical coordinate system are obtained, the radial distance and azimuth are differentially encoded first; at the same time, the relationship between the pitch angle and the radius is fitted by the pitch angle prediction model based on the LSTM network, and the pitch angle is predicted by the reconstructed coordinates of the encoded point and the pitch angle and radial distance of the current point, and the quantized prediction residual is encoded; in low bit rate mode, the azimuth is encoded, and the pitch angle is represented by the differential encoding of the reconstructed value; for the radius r, it is compressed by the autoencoder based on the entropy model; The coding parameter selection module is configured as follows: through a quantization step selection method based on differential evolution, the coding parameter selection problem is transformed into an optimization problem under constraints, and the approximately optimal coding parameters are obtained through iterative optimization on a small data set containing only ten point clouds, and finally other point clouds are encoded using the coding parameters.

Citation Information

Patent Citations

  • Self-adaptive point cloud geometric coding and decoding method and device

    CN114913253A

  • Predictive Encoding / Decoding Method and Apparatus for Azimuth Information of Point Cloud

    US20240070922A1

  • Large-scale point cloud-oriented two-dimensional regularized planar projection and encoding and decoding method

    US20240298039A1

  • Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning

    WO2024130776A1