Display feature embedding-based point cloud analysis method
By introducing technical means such as display feature embedding, local spatial feature coding and residual connection in the point cloud analysis method, the problem of low point cloud classification and segmentation efficiency and difficulty in local detail identification in the existing technology is solved, and high-precision three-dimensional point cloud data processing and model generalization capabilities are improved.
Patent Information
- Application Number
- CN202510289794.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing deep learning methods are inefficient in point cloud classification and segmentation, and local details recognition is easy to be confused, making it difficult to identify three-dimensional point cloud models with high accuracy, and it is difficult to extract point cloud features.
A point cloud analysis method based on display feature embedding (Point MLP-FE) is proposed, combining furthest point sampling (FPS), K nearest neighbor (KNN), and Point Net++ algorithms, and introducing the relative position information of the domain points through the local spatial feature encoding module, encoding the position and structure information in the neighborhood, and improving the model's local structural information representation and generalization ability through residual connection and soft voting mechanisms.
By enhancing the fusion of local feature differences and global feature, the classification effect and generalization capabilities of the model are improved, and high-precision three-dimensional point cloud data processing is achieved.
Smart Images

Figure CN120219320A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to fields such as autonomous driving, industrial measurement, and surveying and mapping. Based on the Farthest Point Sampling (FPS) point cloud sampling algorithm and the K-Nearest Neighbor (KNN) point cloud grouping algorithm, Point Net++ proposes a point cloud analysis method based on explicit feature embedding (Point MLP-FE). Through a novel local feature encoding mode and global feature extraction mechanism, high-quality point cloud feature extraction, fusion, and analysis are completed. Finally, a voting network is applied to calculate the optimal classification result. These technologies show superiority in processing 3D reconstruction, object detection, and semantic segmentation of complex scenes, and can have advantages such as high efficiency, strong robustness, high precision, and wide adaptability. Through the application of these innovative technologies, the present invention has achieved remarkable results in realizing high-precision 3D point cloud data processing. Background Art
[0002] Applying deep learning in 3D point cloud data analysis is a current research hotspot. Existing deep learning methods have disadvantages such as low efficiency in point cloud classification and segmentation and easy confusion in local detail recognition. How to effectively and highly accurately identify 3D point cloud models and extract point cloud features has become an urgent problem to be solved.
[0003] The point cloud analysis method based on explicit feature embedding (Point MLP-FE) combines the Farthest Point Sampling (FPS), K-Nearest Neighbor (KNN), and Point Net++ algorithms. To improve the accuracy of point cloud feature extraction, this paper further improves on the basis of Point Net++. By using explicit feature embedding to distinguish the original input features, the position information and structural information within the neighborhood are encoded, a set of learnable parameters is introduced, and a point-by-point linear transformation is performed to increase the local feature difference and obtain highly distinguishable feature information.
[0004] In terms of the fusion of local features and global features, a residual connection is introduced. By connecting the input layer and the output layer and skipping the same feature extraction main body, the model convergence speed is accelerated, and the problem of model gradient disappearance is effectively alleviated.
[0005] The MLP layer classification and voting network fuse the prediction results of multiple model trainings through a soft voting mechanism. By using different initial sampling points, different network models are obtained, which alleviates the situation that a single model is easily trapped in a local optimum, thereby improving the generalization ability of the model. Summary of the Invention
[0006] When the geometric structures of point clouds are similar, the similarity between the original input features is too high, resulting in the network model being unable to capture feature information with high discrimination, and further leading to the model being unable to achieve a satisfactory classification effect. To address this problem, a point cloud analysis method based on explicit feature embedding (Point MLP-FE) is proposed. This method combines farthest point sampling (FPS), K-nearest neighbor (KNN), and Point Net++ algorithms. Through a local spatial feature encoding module, the relative position information of neighboring points is introduced, and the position information and structural information within the neighborhood are encoded. A set of learnable parameters is introduced to perform a linear transformation point by point, effectively enhancing the representational ability of the local structural information of the model while increasing the local feature differences. In terms of fusing local features and global features, a residual connection is introduced. By connecting the input layer and the output layer and skipping the same feature extraction main body, the model convergence speed is accelerated, and the problem of model gradient disappearance is effectively alleviated. The MLP layer classification and voting network fuse the prediction results of multiple model trainings through a soft voting mechanism. By using different initial sampling points, different network models are obtained, alleviating the situation where a single model is easily trapped in a local optimum, thereby improving the generalization ability of the model.
[0007] Specific invention content:
[0008] A point cloud analysis method based on explicit feature embedding (Point MLP-FE), which is characterized by including the following steps:
[0009] Step 1) In point cloud analysis and processing, it is first necessary to sample the original point cloud into a specified number of target points to reduce the hardware calculation overhead, which includes the following steps:
[0010] Step 1.1) Given that the input point cloud has M points, select a point p0 from the set to obtain the initial sampling point set S = {p0}; use an array L to record the Euclidean distances d from all M points to p0 in the sampling point set S; select the point p1 corresponding to the maximum value of the distance d and add it to the sampling point set S = {p0, p1};
[0011]
[0012] Step 1.2) Calculate the distances d from all points to point p1. For each point p i , if this distance d is less than L i , then assign it to L i , and the array L i always records the minimum distance from each point P to the sampling point set S; select the point p2 corresponding to the maximum value and add it to the sampling point set
[0013] S = {p0, p1, p2}
[0014] Step 1.3) Repeat Step 1.2) until N target sampling points are sampled;
[0015] Step 2) Group feature extraction based on KNN
[0016] Step 2.1) Perform Step 1) on the initial point cloud to obtain 1024 points, then use KNN to divide K nearest neighbor points and extract the original features In this method, K is set to 24;
[0017] Step 2.2) Perform relative position encoding:
[0018] For each point p i of the K nearest neighbor points perform encoding
[0019]
[0020] where p i and are the spatial coordinate positions of the sampling point and its Kth nearest neighbor point, is the Contact operation to splice the matrix along the last dimension, calculate the Euclidean distance between the center point and the nearest neighbor points;
[0021] Step 2.3) Use the geometric affine module to perform a linear transformation point by point instead of the shared MLP layer, and introduce a set of learnable parameters alpha and beta:
[0022]
[0023] Step 2.4) Feature enhancement: Connect the original local features of with the embedded relative position offset code to obtain an enhanced feature vector f i k ;
[0024] Step 3) Extract local feature weights and aggregate depth features to obtain the final feature g i
[0025] g i = Φpos(Α(Φpre(f i k ),|j = 1,…,K))
[0026] Α(·) feature aggregation uses the max pooling operation, and f i k represents the feature of the kth neighboring point of the ith sampling point; Φpre is used to learn the shared weights from the local area, and Φpos is used for depth aggregation of features;
[0027] Step 3.1) Max pooling operation, performing pooling processing on the obtained feature matrix:
[0028]
[0029] M(i, j) is the (i, j) element of the output feature map, max represents the maximum value operation, u and v are indices that vary within the range [0, f - 1] and are used to traverse all elements within the pooling window, s is the stride, which defines the distance by which the pooling window moves on the input feature map, and Α(i×s + u, j×s + v) is the element of the local window in the input feature map Α corresponding to the output feature map M(i, j);
[0030] Step 3.2) Extracting local region weights:
[0031] Φpre(x) = MLP(x)
[0032] where the MLP layer consists of a linear transformation, normalization, and activation function:
[0033] Linear transformation:
[0034] y = W·x + b
[0035] x is the input vector, W is the weight matrix, and b is the bias term, all of which are learnable parameters, and a linear transformation is performed to reconstruct the feature distribution that the original network is supposed to learn;
[0036] Normalization:
[0037]
[0038] is the average value for normalization, and by learning the γ and β parameters through the neural network, standardization can be achieved, which can accelerate the network training efficiency;
[0039] Activation function:
[0040] f(x) = max(0, x)
[0041] Step 3.3) Deep aggregating features, continuously reducing the number of input points, extracting features, and completing feature aggregation
[0042] Φpos = MLP(x) + x
[0043] Step 3.4) Repeat steps 3.1), 3.2), and 3.3) four times to obtain a 1024 - dimensional feature matrix g i ;
[0044] Step 4) Point cloud classification, through the extracted feature matrix g iGradually classified into target categories, the OverallAccuracy (OA) and Average accuracy (AA) metrics are calculated. Finally, the model is run three times and the voting algorithm is used to achieve the optimal output of the algorithm;
[0045] Step 4.1) Gradually classify into 40 target categories (the number of target categories can be set according to the actual situation) through the MLP layer;
[0046] Step 4.2) Calculate the OA and AA metrics. OA represents the proportion of all correctly predicted samples to the total number of all predicted samples; AA represents the calculation of the average precision, which is the ratio between the number of correctly predicted samples in each category and the total number of that category. Finally, the average of the precision of each category is taken; The formula is as follows:
[0047]
[0048] Where TP is True Positive, True represents that the actual and the prediction are the same, and Positive represents that the prediction is a positive sample; False Positive (FP) represents that the actual category and the predicted class label are different, and the predicted category is a positive sample while the actual category is a negative sample; False Negative (FN) represents that the actual category and the predicted class label are different, and the predicted category is a negative sample while the actual category is a positive sample; True Negative (TP) represents that the actual category and the predicted class label are the same, and both the predicted category and the actual category are negative samples;
[0049] Step 4.3) Use the voting algorithm to achieve the optimal output of the model;
[0050] Calculate that the probability that the sample x belongs to each category output by the model is p i (j) (j = 1, 2, 3), and the final classification result C is calculated by the following formula:
[0051]
[0052] w i Is the model weight, and in this method, it can be taken as 1 when the algorithms are the same.
[0053] The present invention has the following advantages and beneficial effects:
[0054] The original input features are distinguished by using display feature embedding, the positional information and structural information within the neighborhood are encoded, a set of learnable parameters are introduced, a point-by-point linear transformation is performed, the local feature difference is increased, and high-discrimination feature information is obtained. In terms of the fusion of local features and global features, a residual connection is introduced. By connecting the input layer and the output layer, the same feature extraction main body is skipped, the model convergence speed is accelerated, and the problem of model gradient disappearance is effectively alleviated.
[0055] The MLP layer classification and voting network fuses the prediction results of multiple model trainings through a soft voting mechanism. By using different initial sampling points, a differentiated network model is obtained, the situation where a single model is prone to falling into a local optimum is alleviated, and the generalization ability of the model is improved. Brief Description of the Drawings
[0056] Figure 1 It is the overall flowchart of the technical field of the point cloud analysis method based on display feature embedding of the present invention. Detailed Embodiments
[0057] Next, in combination with the drawings and the implementation method, the advantages and purposes of the present invention will be further described. It should be understood that the description here only explains the present invention and is not used to limit the present invention.
[0058] A point cloud analysis method (Point MLP-FE) based on display feature embedding combines farthest point sampling (FPS), K-nearest neighbor (KNN), and Point Net++ algorithms. Through a local spatial feature encoding module, the relative position information of neighboring points is introduced, and the positional information and structural information within the neighborhood are encoded. A set of learnable parameters are introduced, and a point-by-point linear transformation is performed. While increasing the local feature difference, the representation ability of the local structure information of the model is effectively enhanced. In terms of the fusion of local features and global features, a residual connection is introduced. By connecting the input layer and the output layer, the same feature extraction main body is skipped, the model convergence speed is accelerated, and the problem of model gradient disappearance is effectively alleviated. The MLP layer classification and voting network fuses the prediction results of multiple model trainings through a soft voting mechanism. By using different initial sampling points, a differentiated network model is obtained, the situation where a single model is prone to falling into a local optimum is alleviated, and thus the generalization ability of the model is improved.
[0059] Specific Invention Content:
[0060] The point cloud analysis method (Point MLP-FE) based on display feature embedding is characterized by including the following steps:
[0061] Step 1) In point cloud analysis and processing, it is necessary to sample the original point cloud into a specified number of target points to reduce the hardware calculation overhead, which includes the following steps:
[0062] Step 1) In point cloud analysis and processing, it is first necessary to sample the original point cloud into a specified target number of points to reduce the hardware calculation overhead, which includes the following steps:
[0063] Step 1.1) The input point cloud has M points. Select a point p0 from the set to obtain the initial sampling point set S = {p0}; use the array L to record the Euclidean distance d from all M points to p0 in the sampling point set S; select the point p1 corresponding to the maximum value of the distance d and add it to the sampling point set S = {p0, p1};
[0064]
[0065] Step 1.2) Calculate the distance d from all points to point p1. For each point p i , if the distance d is less than L i , then assign it to L i , the array L i always records the minimum distance from each point P to the sampling point set S; select the point p2 corresponding to the maximum value and add it to the sampling point set
[0066] S = {p0, p1, p2}
[0067] Step 1.3) Repeat Step 1.2) until N target sampling points are sampled;
[0068] Step 2) Group feature extraction based on KNN
[0069] Step 2.1) Perform Step 1) on the initial point cloud to obtain 1024 points, and then use KNN to divide K nearest neighbor points and extract the original features In this method, K is set to 24;
[0070] Step 2.2) Perform relative position encoding:
[0071] For each K nearest neighbor points of p i perform encoding
[0072]
[0073] where p i and are the spatial coordinate positions of the sampling point and its Kth nearest neighbor point, is the Contact operation to splice the matrix along the last dimension, calculate the Euclidean distance between the center point and the nearest neighbor points;
[0074] Step 2.3) Use the geometric affine module to perform a point-by-point linear transformation to replace the shared MLP layer, and introduce a set of learnable parameters alpha and beta:
[0075]
[0076] Step 2.4) Feature enhancement: Connect the original local features of with the embedded relative position offset code to obtain an enhanced feature vector f i k ;
[0077] Step 3) Extract local feature weights and aggregate depth features to obtain the final feature g i
[0078] g i = Φpos(Α(Φpre(f i k ), |j = 1, …, K))
[0079] The Α(·) feature aggregation uses a max pooling operation, and f i k represents the feature of the k-th neighboring point of the i-th sampling point; Φpre is used to learn shared weights from the local region, and Φpos is used for depth aggregation of features;
[0080] Step 3.1) Max pooling operation, perform pooling on the obtained feature matrix:
[0081]
[0082] M(i, j) is the (i, j) element of the output feature map, max represents the maximum value operation, u and v are indices that vary within the range [0, f - 1] and are used to traverse all elements within the pooling window, s is the stride, which defines the distance by which the pooling window moves on the input feature map, and Α(i × s + u, j × s + v) is the element of the local window in the input feature map Α corresponding to the output feature map M(i, j);
[0083] Step 3.2) Extract local region weights:
[0084] Φpre(x) = MLP(x)
[0085] where the MLP layer consists of a linear transformation, normalization, and activation function:
[0086] Linear transformation:
[0087] y = W · x + b
[0088] x is the input vector, W is the weight matrix, and b is the bias term, all of which are learnable parameters. A linear transformation is performed to reconstruct the feature distribution that the original network is supposed to learn;
[0089] Normalization:
[0090]
[0091] is the average value after normalization. By learning the γ and β parameters through the neural network, standardization can be achieved, which can accelerate the network training efficiency;
[0092] Activation function:
[0093] f(x) = max(0, x)
[0094] Step 3.3) Deeply aggregate features, continuously reduce the number of input points, extract features, and complete feature aggregation
[0095] Φpos = MLP(x) + x
[0096] Step 3.4) Repeat steps 3.1), 3.2), and 3.3) four times to obtain a 1024-dimensional feature matrix g i ;
[0097] Step 4) Point cloud classification. Through the extracted feature matrix g i Gradually classify into target categories, calculate the OverallAccuracy (OA) and Average accuracy (AA) metrics, and finally repeat running the model three times and use the voting algorithm to achieve the optimal output of the algorithm;
[0098] Step 4.1) Gradually classify into 40 target categories (the number of target categories can be set according to the actual situation) through the MLP layer;
[0099] Step 4.2) Calculate the OA and AA metrics. OA represents the proportion of all correctly predicted samples in the total number of all predicted samples; AA represents the calculation of the average precision, which is the ratio between the number of correctly predicted samples in each category and the total number of that category. Finally, take the average of the precision of each category; The formula is as follows:
[0100]
[0101] Among them, TP is True Positive, where True means the actual and the prediction are the same, and Positive means the prediction is a positive sample; False Positive (FP) means that the actual category and the predicted label are different, and the predicted category is a positive sample while the actual category is a negative sample; False Negative (FN) means that the actual category and the predicted label are different, and the predicted category is a negative sample while the actual category is a positive sample; True Negative (TN) means that the actual category and the predicted label are the same, and both the predicted category and the actual category are negative samples;
[0102] Step 4.3) Use the voting algorithm to achieve the optimal output of the model;
[0103] Calculate that the probabilities of the model outputting that the sample x belongs to each category are p i (j) (j = 1, 2, 3), and the final classification result C is calculated by the following formula:
[0104]
[0105] w i is the model weight, and in this method, it can be taken as 1 when the algorithms are the same.
Claims
1. Point cloud analysis method based on display feature embedding (Point MLP-FE), characterized by The steps include: Step 1) In point cloud analysis and processing, the original point cloud needs to be sampled into a specified target number of points to reduce hardware computing overhead, which includes the following steps: Step 1.1) Input point cloud has M points, select a point P0 from the set, and get the initial sampling point Set S = {P0}; use array L to record the Euclidean distance d of all M points to P0 in the sampling point set S; select the point P1 corresponding to the maximum value of the distance d and add it to the sampling point set S = {P0, P1}; Step 1.2) Calculate the distance d from all points to point P1. For each point P i , if the distance d is less than L i , then assign it to L i , array L i Always record the minimum distance from each point P to the sampling point set S; select the point P2 corresponding to the maximum value and add it to the sampling point set S = {P0, P1, P2} Step 1.3) Repeat step 1.2) until N target sampling points are sampled; Step 2) Group feature extraction based on KNN Step 2.1) Perform step 1) on the initial point cloud to obtain 1024 points, then use KNN to divide K neighboring points and extract the original features In this method, K is set to 24; Step 2.2) Perform relative position encoding: For each p i The K nearest neighbors of Encoding where p i and is the spatial coordinate position of the sampling point and its Kth neighbor point, Concatenate the matrices along the last dimension for the Contact operation, Calculate the Euclidean distance between the center point and its neighboring points; Step 2.3) Use the geometric affine module to perform linear transformation point by point instead of the shared MLP layer, and introduce a set of learnable parameters alpha and beta: Step 2.4) Feature enhancement: The original local features The relative position of the embedded code Connect them together to get an enhanced feature vector f i k ; Step 3) Extract local feature weights and aggregate deep features to obtain the final feature g i g i =Φpos(Α(Φpre(f i k ),|j=1,…,K)) Α(·) feature aggregation uses the maximum pooling operation, f i k Represents the features of the kth neighboring point of the i-th sampling point; Φpre learns shared weights from local regions, and Φpos is used for deep feature aggregation; Step 3.1) Maximum pooling operation, pooling the obtained feature matrix: M(i,j) is the (i,j)th element of the output feature map, max represents the maximum value operation, u and v are indices that vary in the range [0,f-1], which are used to traverse each element in the pooling window, s is the step size, which defines the distance the pooling window moves on the input feature map, Α(i×s+u,j×s+v) is the element of the local window in the input feature map Α that corresponds to the output feature map M(i,j); Step 3.2) Local area weight extraction: Φpre(x)=MLP(x) The MLP layer consists of linear transformation, normalization, and activation function: Linear transformation: y=W·x+b x is the input vector, W is the weight matrix, and b is the bias term, all of which are learnable parameters. Linear transformation is performed to reconstruct the feature distribution that the original network wants to learn. Normalization: The normalized average value is used to learn the γ and β parameters through the neural network to achieve standardization, which can speed up the network training efficiency. Activation function: f(x)=max(0,x) Step 3.3) Deeply aggregate features, continuously reduce the number of input points, extract features, and complete feature aggregation Φpos = MLP (x) + x Step 3.4) Repeat steps 3.1) 3.2) 3.3) four times to obtain a 1024-dimensional feature matrix g i ; Step 4) Point cloud classification: gradually classify the point cloud into target categories through the extracted feature matrix, calculate the OverallAccuracy (OA) and AverageAccuracy (AA) indicators, and finally repeat the model three times and use the voting algorithm to achieve the optimal output of the algorithm; Step 4.1) Classify into 40 target categories step by step through the MLP layer (the number of target categories can be set according to actual conditions); Step 4.2) Calculate the OA and AA indicators. OA represents the ratio of all correctly predicted samples to the total number of predicted samples. AA represents the calculation of average accuracy, which is the ratio between the number of correctly predicted samples in each category and the total number of samples in that category. Finally, the average accuracy of each category is taken. The following table shows: Among them, TP stands for True Positive, True means that the actual and predicted are the same, and Positive means that the prediction is a positive sample; False Positive (FP) means that the actual category and the predicted category label are different, and the predicted category is a positive sample, while the actual category is a negative sample; False Negative (FN) means that the actual category and the predicted category label are different, and the predicted category is a negative sample, while the actual category is a positive sample; True Negative (TP) means that the actual category and the predicted category label are the same, and both the predicted category and the actual category are negative samples; Step 4.3) Use voting algorithm to achieve optimal model output; The calculated model output is the probability p that the sample x belongs to each category i (j)(j=1,2,3), the final classification result C is calculated by the following formula: w i is the model weight, and in this method, the algorithm is the same and can be set to 1.