Three-dimensional surface direction estimation method based on deep learning
Through self-supervised learning and spherical convolutional neural network, a twin self-supervised network model is built to process three-dimensional point cloud data, solving the problem of low recognition accuracy in the existing technology when rotation has not been seen, and achieving higher three-dimensional shape recognition and adaptability.
Patent Information
- Application Number
- CN202411950073.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
AI Technical Summary
Existing 3D surface orientation estimation methods based on deep learning are poor in handling no rotations, making it difficult to adapt to the diversity of shape changes and complexity, resulting in a decrease in recognition and classification accuracy.
Self-supervised learning combined with spherical convolutional neural network is used to construct a twin self-supervised network model, and the S2 spherical convolutional residual module and the SO(3) rotary group convolutional residual module are used to process three-dimensional point cloud data and output three-dimensional surface direction features.
It improves the ability to adapt to various unknown deformations and rotations, improves the recognition rate and accuracy of three-dimensional shapes, reduces the dependence on labeled data, and improves the universality and robustness of the model.
Smart Images

Figure BDA0005214651890000032 
Figure BDA0005214651890000062 
Figure BDA0005214651890000097
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and deep learning, and particularly to a three-dimensional surface direction estimation method based on deep learning. Background Art
[0002] In computer vision and robotics, reliably defining and finding the standard direction of a three-dimensional surface is crucial for many applications, such as object recognition, manipulation, and scene understanding. Such tasks are usually solved by manually designed algorithms that rely on geometric cues considered unique and robust by the designers. However, these traditional algorithms have design limitations and are often difficult to adapt to the diversity of shape changes and complexities. The assumptions and intuitions of the designers may not cover all actual situations, resulting in poor performance of these algorithms in specific cases and a lack of general applicability.
[0003] Humans can learn the inherent direction of three-dimensional objects through experience, and machines also have the potential to acquire this ability through similar learning methods. Existing deep learning-based point cloud processing methods, such as PointNet and AtlasNet, although performing well in processing static images, usually rely on data augmentation techniques to sample all possible rotation ranges during training. Although data augmentation can enhance the generalization ability of the model in some aspects, in practical applications, these methods still perform poorly when facing unseen rotations. Specifically, the model is difficult to handle object postures significantly different from the training stage, resulting in a decrease in recognition and classification accuracy. This insufficient adaptability to unseen rotations limits the effective application of deep learning methods in practical scenarios. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a three-dimensional surface direction estimation method based on deep learning that can efficiently learn the inherent direction of three-dimensional objects, improve the adaptability to various unknown deformations and rotations through self-supervised learning combined with a spherical convolutional neural network, and thus enhance the recognition rate and accuracy of three-dimensional shapes.
[0005] The technical solution adopted by the present invention to solve the above technical problem is: A three-dimensional surface direction estimation method based on deep learning, comprising the following steps:
[0006] Step 1: Obtain the original 3D point cloud data from the original 3D point cloud database, map the original 3D point cloud data to the spherical coordinate system to obtain the first set of 3D point clouds to be trained. After obtaining the azimuth angle, inclination angle, and radial distance to the center of each point in the first set of 3D point clouds to be trained, convert the first set of 3D point clouds to the first set of spherical signals; after performing a rotation operation on the first set of 3D point clouds to be trained at an arbitrary angle, obtain the second set of 3D point clouds to be trained. After obtaining the azimuth angle, inclination angle, and radial distance to the center of each point in the second set of 3D point clouds to be trained, convert the second set of 3D point clouds to the second set of spherical signals;
[0007] Step 2: Construct a twin self-supervised network model to be trained. The twin self-supervised network model to be trained includes two branches with the same structure. Each branch includes S 2 spherical convolutional residual module and SO(3) rotation group convolutional residual module, where S 2 The spherical convolutional residual module includes a first spherical convolutional path and a second spherical convolutional path connected in mutual residuals. The first spherical convolutional path consists of a first S 2 spherical convolutional layer, a first batch normalization layer, and a first ReLU activation function, a first SO(3) rotation group convolutional layer, a second batch normalization layer, and a second ReLU activation function. The second spherical convolutional path consists of a second S 2 spherical convolutional layer, a third batch normalization layer, and a third ReLU activation function; The SO(3) rotation group convolutional residual module includes a first rotation group path and a second rotation group path connected in mutual residuals. The first rotation group path consists of two layers of second SO(3) rotation group convolutional layers, a fourth batch normalization layer, a fifth batch normalization layer, a fourth ReLU activation function, and a fifth ReLU activation function. The second rotation group path consists of a single layer of third SO(3) rotation group convolutional layer, a sixth batch normalization layer, and a sixth ReLU activation function;
[0008] Step 3: Input the first set of spherical signals into the first branch of the twin self-supervised network model to be trained, and process them sequentially through the S 2 spherical convolutional residual module and SO(3) rotation group convolutional residual module to output the first branch feature; Input the second set of spherical signals into the second branch of the twin self-supervised network model to be trained, and process them sequentially through the S 2 spherical convolutional residual module and SO(3) rotation group convolutional residual module to output the second branch feature;
[0009] Step 4: Process the first-branch feature using the soft_argmax function to obtain the most significant feature in the first-branch feature. The most significant feature in the first-branch feature represents the three-dimensional surface direction corresponding to the current input first set of spherical signals in the SO(3) three-dimensional rotation space; process the second-branch feature using the soft_argmax function to obtain the most significant feature in the second-branch feature. The most significant feature in the second-branch feature represents the three-dimensional surface direction corresponding to the current input second set of spherical signals in the SO(3) three-dimensional rotation space;
[0010] Step 5: Set the loss function for the twin self-supervised network model to be trained. Rotate the three-dimensional surface direction corresponding to the current input first set of spherical signals in the SO(3) three-dimensional rotation space by the same rotation angle as the current input second set of three-dimensional point clouds to be trained, and then calculate the distance in the SO(3) three-dimensional rotation space with the three-dimensional surface direction corresponding to the current input second set of spherical signals in the SO(3) three-dimensional rotation space. Take the calculated distance as the network error, backpropagate the network error to the twin self-supervised network model to be trained through the backpropagation algorithm, and use the gradient descent method to adjust the network parameters to iteratively update the twin self-supervised network model to be trained until the set maximum number of iterations is reached and stop the iterative update process, and obtain the trained twin self-supervised network model;
[0011] Step 6: Obtain the three-dimensional point cloud data to be measured from the target three-dimensional point cloud database, map the three-dimensional point cloud data to the spherical coordinate system to obtain the mapped three-dimensional point cloud to be estimated, obtain the azimuth angle, inclination angle, and radial distance to the center of each point of the mapped three-dimensional point cloud to be estimated, and then convert the mapped three-dimensional point cloud to be estimated into a spherical signal to be estimated and input it into the trained twin self-supervised network model to obtain the three-dimensional surface direction of the spherical signal to be estimated in the SO(3) three-dimensional rotation space, and complete the estimation process of the three-dimensional surface direction of the three-dimensional point cloud data to be measured.
[0012] Compared with the prior art, the advantages of the present invention are that the whole process utilizes the rotational equivariance property of spherical convolution, which can maintain the corresponding rotation of the feature map during three-dimensional rotation operations, making the present invention have advantages in dealing with rotation-equivariant tasks; in addition, S 2The design of the spherical convolution residual module and the SO(3) rotation group convolution residual module not only enhances the accuracy of signal processing but also improves the performance of the network; the twin self-supervised network model to be trained uses a self-supervised method during training, automatically adapting to different rotations and deformations, and can learn to predict the best direction without explicit supervision, which reduces the dependence on labeled data and improves the generality of the model. In this way, the robustness of 3D shape recognition, classification, and processing can be significantly improved, enabling computer vision and robotics to work more ideally in more complex and variable environments.
[0013] Specifically, the specific process of converting the first group of 3D point clouds to be trained into the first group of spherical signals in step 1 is as follows:
[0014] Step 1-1: Convert each point in the first group of 3D point clouds to be trained from the Euclidean coordinate system to the spherical coordinate system. Define the spherical coordinates of the point in the spherical coordinate system as (α, β, d), where α represents the azimuth angle, indicating the angle of the point on the horizontal plane; β represents the elevation angle, indicating the angle of the point in the vertical direction; and d represents the radial distance from the point to the center.
[0015] Step 1-2: In the spherical coordinate system, construct a discretized 3D grid for quantifying the point cloud as follows: First, divide the unit sphere in the spherical coordinate system into multiple grid cells according to the azimuth angle and elevation angle, and then divide the unit sphere along the radial direction into multiple concentric spherical shells, each spherical shell representing a depth range. Set the discretization resolution of the azimuth angle to I, the discretization resolution of the elevation angle to I, and the discretization resolution of the radial distance from the point to the center to K. Then each grid cell is numbered by the triple (α[i], β[j], d[k]) as the corresponding number in the discretized 3D grid index, and the discretized 3D grid is formed by arranging them in order, where,
[0016] Step 1-3: Assign the points in the first group of 3D point clouds to be trained to the discretized 3D grid to generate the first group of spherical signals. The specific steps are as follows: According to the spherical coordinates of each point, assign it to the corresponding grid cell to generate a single-channel spherical signal. The signal value of the spherical signal is defined as the point density and geometric features. Multiple single-channel spherical signals along the same radial direction are used as a multi-channel spherical signal in the first group of spherical signals.
[0017] Specifically, in step 3, the S 2 The specific process of the spherical convolution residual module processing the first group of spherical signals is as follows:
[0018] Step 3-1: The first S in the first spherical convolution path 2The spherical convolution layer convolves the first set of spherical signals and maps them into the SO(3) three-dimensional rotation space to obtain the first mapped signal. The number of channels of the input first set of spherical signals is 4, and the bandwidth is 32. After convolution, the number of channels of the obtained signal increases to 40, and the bandwidth is 32. The first mapped signal is processed by the first batch normalization layer and the first ReLU activation function to obtain the first activation signal, and the first activation signal is input into the first SO(3) rotation group convolution layer;
[0019] Step 3-2: The first SO(3) rotation group convolution layer convolves the first activation signal and outputs the first convolution signal. The number of channels of the first convolution signal is 40, and the bandwidth is 32. The first convolution signal is processed by the second batch normalization layer and the second ReLU activation function to obtain the output result of the first spherical convolution path;
[0020] Step 3-3: The second S 2 The spherical convolution layer convolves the first set of spherical signals to obtain the second convolution signal. The number of channels of the second convolution signal is 40, and the bandwidth is 32. The second convolution signal is processed by the third batch normalization layer and the third ReLU activation function to obtain the output result of the second spherical convolution path. The output result of the second spherical convolution path is added to the output result of the first spherical convolution path, and then the added result is processed by the seventh batch normalization layer and the seventh ReLU activation function to obtain S 2 The SO(3) three-dimensional rotation space signal output by the spherical convolution residual module, where the number of channels of the SO(3) three-dimensional rotation space signal is 40 and the bandwidth is 32.
[0021] Specifically, in step 3, the specific processing process of the SO(3) rotation group convolution residual module is as follows:
[0022] Step 3-4: The first layer of the second SO(3) rotation group convolution layer in the first rotation group path convolves the SO(3) three-dimensional rotation space signal to obtain the third convolution signal. The number of channels of the third convolution signal is 20, and the bandwidth is 32. The third convolution signal is processed by the fourth batch normalization layer and the fourth ReLU activation function to obtain the second activation signal;
[0023] Step 3-5: The second SO(3) rotation group convolution layer of the second layer in the first rotation group path convolves the second activation signal to obtain a fourth convolution signal, the fourth convolution signal has a channel number of 1 and a bandwidth of 32; the fourth convolution signal is processed by the fifth batch normalization layer and the fifth ReLU activation function to obtain the output result of the first rotation group path; Step 3-6: The third SO(3) rotation group convolution layer in the second rotation group path convolves the SO(3) three-dimensional rotation space signal to obtain a fifth convolution signal, the channel number of the fifth convolution signal is 1 and the bandwidth is 32; the fifth convolution signal is processed by the sixth batch normalization layer and the sixth ReLU activation function to obtain the output result of the second rotation group path, the output result of the second rotation group path is added to the output result of the first rotation group path, and the result of the addition is processed by the eighth batch normalization layer and the eighth ReLU activation function to obtain the output of the SO(3) rotation group convolution residual module, that is, the first branch feature.
[0024] Specifically, the specific process of step 4 is as follows:
[0025] Step 4-1: Let the first branch feature be θ, and then find the optimal parameter point in the θ feature map through the soft_argmax layer function, that is, the position of the maximum value in the θ feature map, and record the optimal parameter point as C R ;
[0026] Step 4-2: Find C in θ R The corresponding rotation matrix R P , R P The three-dimensional surface direction corresponding to the first set of spherical signals currently input in the SO(3) three-dimensional rotation space.
[0027] Specifically, in step 5, the loss function of the twin self-supervised network model to be trained is recorded as The three-dimensional surface direction corresponding to the first set of spherical signals currently input in the three-dimensional rotation space of SO(3) is recorded as R P , the rotation operation in the second set of 3D point clouds to be trained is obtained by rotating the first set of 3D point clouds to be trained at any angle and recorded as the rotation matrix R, then Represents R P The three-dimensional surface direction obtained after R rotation, R W Indicates the three-dimensional surface direction corresponding to the second set of spherical signals currently input in the SO(3) three-dimensional rotation space, Represents R W The transpose of express and The trace of the matrix after the multiplication.
[0028] Specifically, the learning rate of the twin self-supervised network model to be trained is set to 0.0005, the batch size is set to 16, and the number of training epochs is set to 100 epochs. Detailed implementation manners
[0029] The present invention will be further described in detail below in conjunction with specific embodiments.
[0030] A three-dimensional surface direction estimation method based on deep learning includes the following steps:
[0031] Step 1: Obtain the original three-dimensional point cloud data from the original three-dimensional point cloud database, map the original three-dimensional point cloud data to the spherical coordinate system to obtain the first group of three-dimensional point clouds to be trained, and after obtaining the azimuth angle, inclination angle, and radial distance from the center of each point in the first group of three-dimensional point clouds to be trained, convert the first group of three-dimensional point clouds to the first group of spherical signals; after performing a rotation operation on the first group of three-dimensional point clouds to be trained at an arbitrary angle, obtain the second group of three-dimensional point clouds to be trained, and after obtaining the azimuth angle, inclination angle, and radial distance from the center of each point in the second group of three-dimensional point clouds to be trained, convert the second group of three-dimensional point clouds to the second group of spherical signals; specifically, the specific process of converting the first group of three-dimensional point clouds to the first group of spherical signals in Step 1 is as follows:
[0032] Step 1-1: Convert each point in the first group of three-dimensional point clouds to be trained from the Euclidean coordinate system to the spherical coordinate system. Define the spherical coordinates of the point in the spherical coordinate system as (α, β, d), where α represents the azimuth angle, indicating the angle of the point on the horizontal plane; β represents the elevation angle, indicating the angle of the point in the vertical direction; d represents the radial distance from the point to the center; through this conversion, each point in the first group of three-dimensional point clouds to be trained can be represented by spherical coordinates, thereby mapping the point cloud data to the sphere.
[0033] Step 1-2: In the spherical coordinate system, construct a discretized three-dimensional grid for quantifying the point cloud, specifically as follows: First, divide the unit sphere in the spherical coordinate system into multiple grid cells according to the azimuth angle and elevation angle, and then divide the unit sphere along the radial direction into multiple concentric spherical shells, each spherical shell representing a depth range; set the discretization resolution of the azimuth angle to I, the discretization resolution of the elevation angle to I, and the discretization resolution of the radial distance from the point to the center to K. Then each grid cell is numbered by the triple (α[i], β[j], d[k]) corresponding to the discretized three-dimensional grid index, and is arranged in order according to the number to form the discretized three-dimensional grid, where,
[0034] Step 1-3: Assign the points in the first group of 3D point clouds to be trained to the discretized 3D grids to generate the first group of spherical signals. The specific steps are as follows: According to the spherical coordinates of each point, assign it to the corresponding grid cell to generate a single-channel spherical signal. The signal value of the spherical signal is defined as the point density and geometric features. Take multiple single-channel spherical signals along the same radial direction as a multi-channel spherical signal in the first group of spherical signals. In this way, map the point cloud signal into a spherical signal, which is convenient for subsequent spherical convolution processing.
[0035] Step 2: Construct the twin self-supervised network model to be trained. The twin self-supervised network model to be trained includes two branches with the same structure. Each branch includes S 2 spherical convolutional residual module and SO(3) rotation group convolutional residual module. SO(3) represents the special orthogonal group in 3D space (the full English name is Special Orthogonal Group in 3D), which is the set of all 3D rotation matrices. Among them, S 2 The spherical convolutional residual module includes a first spherical convolutional path and a second spherical convolutional path with mutual residual connection. The first spherical convolutional path consists of a first S 2 spherical convolutional layer, the first batch normalization layer and the first ReLU activation function, the first SO(3) rotation group convolutional layer, the second batch normalization layer and the second ReLU activation function. The second spherical convolutional path consists of a second S 2 spherical convolutional layer, the third batch normalization layer and the third ReLU activation function; The SO(3) rotation group convolutional residual module includes a first rotation group path and a second rotation group path with mutual residual connection. The first rotation group path consists of two layers of the second SO(3) rotation group convolutional layer, the fourth batch normalization layer, the fifth batch normalization layer, the fourth ReLU activation function and the fifth ReLU activation function. The second rotation group path consists of a single layer of the third SO(3) rotation group convolutional layer, the sixth batch normalization layer and the sixth ReLU activation function. Build the twin self-supervised network model to be trained in this way. S 2 The spherical convolutional residual module and the SO(3) rotation group convolutional residual module can build a deeper twin self-supervised network model to be trained and improve the generalization ability of the twin self-supervised network model to be trained.
[0036] Step 3: Input the first group of spherical signals into the first branch of the twin self-supervised network model to be trained, and successively pass through S 2 spherical convolutional residual module and SO(3) rotation group convolutional residual module for processing, and output the first branch feature; Input the second group of spherical signals into the second branch of the twin self-supervised network model to be trained, and successively pass through S 2Processed by the spherical convolution residual module and the SO(3) rotation group convolution residual module to obtain the second branch feature; S 2 The specific process of the spherical convolution residual module for processing the first set of spherical signals is as follows:
[0037] Step 3-1: The first S in the first spherical convolution path 2 The spherical convolution layer convolves the first set of spherical signals and maps them into the SO(3) three-dimensional rotation space to obtain the first mapped signal. The number of channels of the input first set of spherical signals is 4, the bandwidth is 32, and the number of channels of the signal obtained after convolution increases to 40, and the bandwidth is 32. After processing the first mapped signal through the first batch normalization layer and the first ReLU activation function, the first activation signal is obtained, and the first activation signal is input into the first SO(3) rotation group convolution layer.
[0038] Step 3-2: The first SO(3) rotation group convolution layer outputs the first convolution signal after convolving the first activation signal. The number of channels of the first convolution signal is 40, and the bandwidth is 32. After processing the first convolution signal through the second batch normalization layer and the second ReLU activation function, the output result of the first spherical convolution path is obtained.
[0039] Step 3-3: The second S 2 The spherical convolution layer convolves the first set of spherical signals to obtain the second convolution signal. The number of channels of the second convolution signal is 40, and the bandwidth is 32. After processing the second convolution signal through the third batch normalization layer and the third ReLU activation function, the output result of the second spherical convolution path is obtained. The output result of the second spherical convolution path is added to the output result of the first spherical convolution path, and then the result of the addition is processed through the seventh batch normalization layer and the seventh ReLU activation function to obtain the S 2 The SO(3) three-dimensional rotation space signal output by the spherical convolution residual module, and the number of channels of the SO(3) three-dimensional rotation space signal is 40, and the bandwidth is 32.
[0040] The specific processing process of the SO(3) rotation group convolution residual module is as follows:
[0041] Step 3-4: The first layer of the second SO(3) rotation group convolution layer in the first rotation group path convolves the SO(3) three-dimensional rotation space signal to obtain the third convolution signal. The number of channels of the third convolution signal is 20, and the bandwidth is 32; after processing the third convolution signal through the fourth batch normalization layer and the fourth ReLU activation function, the second activation signal is obtained.
[0042] Step 3-5: The second SO(3) rotation group convolutional layer in the first rotation group path convolves the second activation signal to obtain a fourth convolutional signal. The number of channels of the fourth convolutional signal is 1, and the bandwidth is 32. The fourth convolutional signal is processed through the fifth batch normalization layer and the fifth ReLU activation function to obtain the output result of the first rotation group path.
[0043] Step 3-6: The third SO(3) rotation group convolutional layer in the second rotation group path convolves the SO(3) three-dimensional rotation space signal to obtain a fifth convolutional signal. The number of channels of the fifth convolutional signal is 1, and the bandwidth is 32. The fifth convolutional signal is processed through the sixth batch normalization layer and the sixth ReLU activation function to obtain the output result of the second rotation group path. The output result of the second rotation group path is added to the output result of the first rotation group path, and then the added result is processed through the eighth batch normalization layer and the eighth ReLU activation function to obtain the output of the SO(3) rotation group convolutional residual module, that is, the first branch feature.
[0044] Step 4: The first branch feature is processed using the soft_argmax function to obtain the most significant feature in the first branch feature. The most significant feature in the first branch feature represents the three-dimensional surface direction corresponding to the current input first group of spherical signals in the SO(3) three-dimensional rotation space. The second branch feature is processed using the soft_argmax function to obtain the most significant feature in the second branch feature. The most significant feature in the second branch feature represents the three-dimensional surface direction corresponding to the current input second group of spherical signals in the SO(3) three-dimensional rotation space. The specific process is as follows:
[0045] Step 4-1: Denote the first branch feature as θ, and then find the optimal parameter point in the θ feature map through the soft_argmax layer function, that is, the position of the maximum value in the θ feature map, and denote the optimal parameter point as C R ;
[0046] Step 4-2: Find the rotation matrix R corresponding to C in θ R and use R P as the three-dimensional surface direction corresponding to the current input first group of spherical signals in the SO(3) three-dimensional rotation space. P
[0047] Step 5: Set the loss function for the twin self-supervised network model to be trained. Rotate the three-dimensional surface direction corresponding to the current input first set of spherical signals in the SO(3) three-dimensional rotation space by the same rotation angle as the current input second set of three-dimensional point clouds to be trained, and then calculate the distance in the SO(3) three-dimensional rotation space between the rotated three-dimensional surface direction and the three-dimensional surface direction corresponding to the current input second set of spherical signals in the SO(3) three-dimensional rotation space. Take the calculated distance as the network error, backpropagate the network error to the twin self-supervised network model to be trained through the backpropagation algorithm, and use the gradient descent method to adjust the network parameters. Iteratively update the twin self-supervised network model to be trained until the set maximum number of iterations is reached and stop the iterative update process, and obtain the trained twin self-supervised network model; specifically, denote the loss function for the twin self-supervised network model to be trained as wherein, denote the three-dimensional surface direction corresponding to the current input first set of spherical signals in the SO(3) three-dimensional rotation space as R P , denote the rotation operation for obtaining the second set of three-dimensional point clouds to be trained after rotating the first set of three-dimensional point clouds to be trained by an arbitrary angle as the rotation matrix R, then represents the three-dimensional surface direction obtained after R P is rotated by R, and R W represents the three-dimensional surface direction corresponding to the current input second set of spherical signals in the SO(3) three-dimensional rotation space, represents the transpose of R W , represents and the trace of the matrix after multiplication, and from the above, is actually defined as the distance between W and R. Among them, the learning rate of the twin self-supervised network model to be trained is set to 0.0005, the batch size is set to 16, and the number of training epochs is set to 100 epochs.
[0048] Step 6: Obtain the three-dimensional point cloud data to be measured from the target three-dimensional point cloud database, map the three-dimensional point cloud data to be measured to the spherical coordinate system to obtain the mapped three-dimensional point cloud to be estimated, obtain the azimuth angle, inclination angle and radial distance from the center of each point of the mapped three-dimensional point cloud to be estimated, and then convert the mapped three-dimensional point cloud to be estimated into the spherical signal to be estimated and input it into the trained twin self-supervised network model to obtain the three-dimensional surface direction of the spherical signal to be estimated in the SO(3) three-dimensional rotation space, and complete the estimation process of the three-dimensional surface direction of the three-dimensional point cloud data to be measured.
[0049] The effectiveness of the three-dimensional surface direction estimation method of this embodiment is compared through specific comparison examples below, where the three-dimensional surface direction estimation method of this embodiment is abbreviated as this method.
[0050] Table 1 shows the comparison of the classification accuracy of this method and other traditional point cloud processing networks in the classification task on the ModelNet40 dataset when dealing with the rotation task.
[0051] Table 1
[0052]
[0053] As can be seen from Table 1, after using this method, the performance on the test rotation dataset has been significantly improved; when the training dataset is not rotated and the test dataset is randomly rotated, the classification accuracy is improved to 90.38%.
[0054] Table 2 shows the comparison results of the intersection over union (IoU) of this method and other traditional point cloud processing networks in the partial segmentation task of the ShapeNet dataset when dealing with the rotation task and the training dataset is not rotated while the test dataset is randomly rotated.
[0055] Table 2
[0056]
[0057] In Table 2, "this method + PointNet" means that the three-dimensional surface direction of the point cloud to be measured is estimated by this method first, and after rotating the point cloud to be measured according to this direction, it is then sent to the PointNet network for training and testing. The performance of the PointNet network on the test rotation dataset has been significantly improved. As can be seen from Table 2, the average instance IoU and average class IoU of this method have achieved excellent results. Among them, avg.inst. represents the average instance IoU, and avg.cls. represents the average class IoU. When this method and the PointNet method are used to test airplanes, bags, and hats, the IoU has increased by more than 63%, 36%, and 40% respectively; when compared with the PCNN method for testing airplanes, bags, and hats, the IoU has increased by more than 59%, 36%, and 48% respectively;
[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional surface direction estimation method based on deep learning, characterized in that The following steps are involved: Step 1: Obtain original 3D point cloud data from an original 3D point cloud database, map the original 3D point cloud data to a spherical coordinate system to obtain a first set of 3D point clouds to be trained, obtain the azimuth, inclination and radial distance to the center of each point in the first set of 3D point clouds to be trained, and then convert the first set of 3D point clouds to be trained into a first set of spherical signals; After performing an arbitrary angle rotation operation on the first group of three-dimensional point clouds to be trained, a second group of three-dimensional point clouds to be trained is obtained, and after obtaining the azimuth, inclination and radial distance to the center of each point of the second group of three-dimensional point clouds to be trained, the second group of three-dimensional point clouds to be trained is converted into a second group of spherical signals; Step 2: Construct the twin self-supervised network model to be trained. The twin self-supervised network model to be trained includes two branches with the same structure, each branch includes S 2 Spherical convolution residual module and SO(3) rotation group convolution residual module, where S 2 The spherical convolution residual module includes a first spherical convolution path and a second spherical convolution path which are mutually residually connected. The first spherical convolution path is composed of a first S 2 The first spherical convolution layer, the first normalization layer and the first ReLU activation function, the first SO(3) rotation group convolution layer, the second normalization layer and the second ReLU activation function, and the second spherical convolution path consists of the second S 2 The spherical convolution layer, the third batch normalization layer and the third ReLU activation function are composed; the SO(3) rotation group convolution residual module includes a first rotation group path and a second rotation group path which are residually connected to each other, the first rotation group path is composed of two layers of the second SO(3) rotation group convolution layer, the fourth batch normalization layer, the fifth batch normalization layer, the fourth ReLU activation function and the fifth ReLU activation function, and the second rotation group path is composed of a single layer of the third SO(3) rotation group convolution layer, the sixth batch normalization layer and the sixth ReLU activation function; Step 3: Input the first set of spherical signals into the first branch of the twin self-supervised network model to be trained, and pass through S 2 The spherical convolution residual module and the SO(3) rotation group convolution residual module are processed to output the first branch feature; the second set of spherical signals are input into the second branch of the twin self-supervised network model to be trained, and then pass through S 2 The spherical convolution residual module and the SO(3) rotation group convolution residual module are processed and the output is the second branch feature; Step 4: Process the first branch feature using the soft_argmax function to obtain the most significant feature in the first branch feature, the most significant feature in the first branch feature represents the three-dimensional surface direction corresponding to the first group of spherical signals currently input in the SO(3) three-dimensional rotation space; process the second branch feature using the soft_argmax function to obtain the most significant feature in the second branch feature, the most significant feature in the second branch feature represents the three-dimensional surface direction corresponding to the second group of spherical signals currently input in the SO(3) three-dimensional rotation space; Step 5: Set the loss function for the twin self-supervised network model to be trained, rotate the three-dimensional surface direction corresponding to the current input first set of spherical signals in the SO(3) three-dimensional rotation space at the same rotation angle as the current input second set of three-dimensional point clouds to be trained, and then calculate the distance in the SO(3) three-dimensional rotation space with the three-dimensional surface direction corresponding to the current input second set of spherical signals in the SO(3) three-dimensional rotation space. Use the calculated distance as the network error, and propagate the network error back to the twin self-supervised network model to be trained through the back propagation algorithm, and use the gradient descent method to adjust the network parameters. Iteratively update the twin self-supervised network model to be trained until the iterative update process is stopped when the set maximum number of iterations is reached, and the trained twin self-supervised network model is obtained. Step 6: Obtain the 3D point cloud data to be measured from the target 3D point cloud database, map the 3D point cloud data to be measured to the spherical coordinate system to obtain the mapped 3D point cloud to be estimated, obtain the azimuth, inclination and radial distance to the center of each point of the mapped 3D point cloud to be estimated, and then convert the mapped 3D point cloud to be estimated into the spherical signal to be estimated and input it into the trained twin self-supervised network model to obtain the 3D surface direction of the spherical signal to be estimated in the SO(3) 3D rotation space, thus completing the estimation process of the 3D surface direction of the 3D point cloud data to be measured.
2. A three-dimensional surface direction estimation method based on deep learning according to claim 1, characterized in that The specific process of converting the first set of three-dimensional point clouds to be trained into the first set of spherical signals in step 1 is as follows: Step 1-1: Convert each point in the first set of 3D point clouds to be trained from the Euclidean coordinate system to the spherical coordinate system, and define the spherical coordinates of the point in the spherical coordinate system as (α, β, d), where α represents the azimuth, which indicates the angle of the point on the horizontal plane; β represents the elevation, which indicates the angle of the point in the vertical direction; d represents the radial distance from the point to the center; Step 1-2: In the spherical coordinate system, a discretized three-dimensional grid for quantizing the point cloud is constructed as follows: first, the unit sphere in the spherical coordinate system is divided into multiple grid cells according to the azimuth and elevation angles, and then the unit sphere is divided radially into multiple concentric spherical shells, each of which represents a depth range; the discretization resolution of the azimuth angle is set to I, the discretization resolution of the elevation angle is set to I, and the discretization resolution of the radial distance from the point to the center is set to K. Then, each grid cell is composed of a triplet (α[i], β[j], d[k]) as the corresponding number in the discretized three-dimensional grid index, and is arranged in order according to the number to form a discretized three-dimensional grid, where, Step 1-3: Assign the points in the first group of three-dimensional point clouds to be trained to the discretized three-dimensional grid to generate the first group of spherical signals. The specific steps are as follows: According to the spherical coordinates of each point, assign it to the corresponding grid unit to generate a single-channel spherical signal. The signal value of the spherical signal is defined as the point density and geometric characteristics. Multiple single-channel spherical signals along the same radial direction are used as a multi-channel spherical signal in the first group of spherical signals.
3. A three-dimensional surface direction estimation method based on deep learning according to claim 2, characterized in that In the step 3, the S 2 The specific process of the spherical convolution residual module processing the first group of spherical signals is as follows: Step 3-1: The first S in the first spherical convolution path 2 The spherical convolution layer convolves the first group of spherical signals and maps them into the SO(3) three-dimensional rotation space to obtain a first mapping signal. The number of channels of the input first group of spherical signals is 4 and the bandwidth is 32. The number of channels of the signal obtained after convolution increases to 40 and the bandwidth is 32. The first mapping signal is processed by the first batch of normalization layers and the first ReLU activation function to obtain a first activation signal, and the first activation signal is input into the first SO(3) rotation group convolution layer; Step 3-2: The first SO(3) rotation group convolution layer convolves the first activation signal to output the first convolution signal. The number of channels of the first convolution signal is 40 and the bandwidth is 32. The first convolution signal is processed by the second batch normalization layer and the second ReLU activation function to obtain the output result of the first spherical convolution path; Step 3-3: Second S 2 The spherical convolution layer convolves the first group of spherical signals to obtain a second convolution signal. The number of channels of the second convolution signal is 40 and the bandwidth is 32. The second convolution signal is processed by the third batch normalization layer and the third ReLU activation function to obtain the output result of the second spherical convolution path. The output result of the second spherical convolution path is added to the output result of the first spherical convolution path, and then the result of the addition is processed by the seventh batch normalization layer and the seventh ReLU activation function to obtain S 2 The spherical convolution residual module outputs the SO(3) three-dimensional rotational spatial signal, which has 40 channels and a bandwidth of 32.
4. A three-dimensional surface direction estimation method based on deep learning according to claim 3, characterized in that In step 3, the specific processing process of the SO(3) rotation group convolution residual module is as follows: Step 3-4: The first second SO(3) rotation group convolution layer in the first rotation group path convolves the SO(3) three-dimensional rotation space signal to obtain a third convolution signal, where the number of channels of the third convolution signal is 20 and the bandwidth is 32; the third convolution signal is processed by the fourth batch normalization layer and the fourth ReLU activation function to obtain a second activation signal; Step 3-5: The second SO(3) rotation group convolution layer in the first rotation group path convolves the second activation signal to obtain a fourth convolution signal, where the fourth convolution signal has 1 channel and a bandwidth of 32; the fourth convolution signal is processed by a fifth batch normalization layer and a fifth ReLU activation function to obtain the output result of the first rotation group path; Step 3-6: The third SO(3) rotation group convolution layer in the second rotation group path convolves the SO(3) three-dimensional rotation space signal to obtain a fifth convolution signal, the number of channels of the fifth convolution signal is 1, and the bandwidth is 32; the fifth convolution signal is processed by the sixth batch normalization layer and the sixth ReLU activation function to obtain the output result of the second rotation group path, the output result of the second rotation group path is added to the output result of the first rotation group path, and the result of the addition is processed by the eighth batch normalization layer and the eighth ReLU activation function to obtain the output of the SO(3) rotation group convolution residual module, that is, the first branch feature.
5. A three-dimensional surface direction estimation method based on deep learning according to claim 4, characterized in that The specific process of step 4 is as follows: Step 4-1: Let the first branch feature be θ, and then find the optimal parameter point in the θ feature map through the soft_argmax layer function, that is, the position of the maximum value in the θ feature map, and record the optimal parameter point as C R ; Step 4-2: Find C in θ R The corresponding rotation matrix R P , R P The three-dimensional surface direction corresponding to the first set of spherical signals currently input in the SO(3) three-dimensional rotation space.
6. A three-dimensional surface direction estimation method based on deep learning according to claim 5, characterized in that In step 5, the loss function of the twin self-supervised network model to be trained is recorded as The three-dimensional surface direction corresponding to the first set of spherical signals currently input in the three-dimensional rotation space of SO(3) is recorded as R P , the rotation operation in the second set of 3D point clouds to be trained is obtained by rotating the first set of 3D point clouds to be trained at any angle and recorded as the rotation matrix R, then Represents R P The three-dimensional surface direction obtained after R rotation, R W Indicates the three-dimensional surface direction corresponding to the second set of spherical signals currently input in the SO(3) three-dimensional rotation space, Represents R W The transpose of express and The trace of the matrix after the multiplication.
7. The method for estimating a three-dimensional surface direction based on deep learning according to claim 6, characterized in that: The learning rate of the twin self-supervised network model to be trained is set to 0.0005, the batch size is set to 16, and the training rounds are set to 100 rounds.