Feature structure adjustable reduction-based few-sample 3D point cloud identification method
Through the method of adjustable feature structure reduction, the feature structure is dynamically adjusted using the Dirichlet process model and DGCNN network, which solves the problem of high model generalization capability and high computing storage requirements in 3D point cloud few sample learning, and achieves efficient recognition performance and speed improvement.
Patent Information
- Application Number
- CN202510502245.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to effectively use a very small number of labeled samples for feature representation in 3D point cloud learning, resulting in low generalization capabilities and recognition accuracy, and high computing and storage requirements, limiting the scalability of the model.
Through the method of adjustable feature structure reduction, the Dilicre process model and DGCNN network are used to dynamically adjust the feature structure, reduce the complexity of the model and improve the feature diversity, and enhance the generalization ability of the neural network.
It significantly improves the accuracy and speed of 3D point cloud recognition for few samples, reduces the computing and storage requirements of the model, adapts to different data sets, and solves the poor performance problem in small samples learning.
Smart Images

Figure CN120356006A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a few-shot 3D point cloud recognition method based on feature structure adjustable reduction. Background Art
[0002] 3D point cloud data has been widely used in many fields, such as 3D object detection and recognition, scene understanding, autonomous driving, and remote sensing target recognition. However, obtaining and annotating large-scale 3D point cloud datasets is both time-consuming and expensive and requires professional skills. In practical applications, usually only a small number of annotated samples can be obtained. Therefore, how to solve the few-shot learning problem of 3D point clouds has important practical significance. In addition, the complexity of 3D point cloud data further increases the difficulty of few-shot learning because this data is usually sparse, noisy, and there are changes in perspective and scale. These factors make it more challenging to extract robust feature representations from a small number of annotated samples.
[0003] Traditional deep learning models, such as convolutional neural networks (CNNs) and graph convolutional networks (GCNs), are widely used in processing and classifying 3D point cloud data. However, in the few-shot learning scenario, the number of samples available for training is usually small, making it difficult to meet the training requirements of deep networks, resulting in insufficient training effects of the models, weak generalization ability and interpretability, and it is difficult to play a role in practical applications with high performance requirements. Meta-learning, as a technique for learning how to learn, has achieved success in few-shot learning. In the few-shot learning of 3D point clouds, meta-learning is introduced to improve the generalization ability of the model for few-shot samples. However, meta-learning usually requires designing a relatively complex model structure and additional learning strategies, and this complexity will bring an increase in the number of parameters and computational burden. Especially when dealing with large-scale 3D point cloud data, the excessive computational overhead may limit its application and scalability. Reinforcement learning, as a learning method based on a reward mechanism, has also attracted attention in the few-shot learning of 3D point clouds. By designing appropriate reward mechanisms and strategies, reinforcement learning can help the model accumulate experience in a small number of samples and improve the model performance through interaction with the environment. However, in the few-shot learning of 3D point clouds, designing a suitable reward function still faces challenges. Due to the complexity of 3D point cloud data in terms of feature expression and semantic understanding, it is not easy to define a reasonable reward signal to guide the model learning. In addition, the sparsity of the data may also cause the sparse reward problem, affecting the learning and optimization process of the model. Although large-scale deep learning models have great potential for few-shot learning in theory, especially in terms of model depth, their large number of parameters and complex structures will bring greater computational and storage requirements. When dealing with large-scale 3D point cloud data, more computational resources and storage space are required for training and inference, which will limit the scalability and practical application of the model.
[0004] Currently, the deficiencies of traditional deep learning methods mainly include: having a strong dependence on large-scale data and storage resources, being difficult to learn rich feature representations from extremely small amounts of labeled samples, and having poor generalization ability for other unseen targets of the same category, resulting in low point cloud recognition accuracy for few-shot 3D targets. Summary of the Invention
[0005] The present invention proposes a few-shot 3D point cloud recognition method based on feature structure adjustable reduction. This method reduces the feature structure of the 3D target point cloud input into the network model, which not only increases the diversity of training samples but also reduces the limitation of the feature structure on the network model, thereby reducing the complexity of model learning, enhancing the learning ability of the model, and ultimately improving the recognition effect of the network model on 3D target point clouds.
[0006] A few-shot 3D point cloud recognition method based on feature structure adjustable reduction, which includes the following steps:
[0007] Step 1: Define the spatial coordinates of the target point cloud input into the network during the 3D point cloud recognition process as its feature structure;
[0008] Step 2: Map the reduction intensity to an infinite-dimensional space through Dirichlet processing so that it can be learned by the network, and obtain an adjustable reduction intensity through the learning of the network model;
[0009] Step 3: Reduce the training of the model and the feature structure through a sampler with adjustable reduction intensity to obtain a dynamically changing new feature structure of the point cloud;
[0010] Step 4: Input the point cloud information after reducing the target point cloud, select DGCNN as the feature encoder, and obtain the target point cloud features;
[0011] Step 5: Use the reduced features as input and output the final target recognition accuracy through a classifier.
[0012] Preferably, in Step 1, let the 3D target point cloud be the set Point = {p1, p2,..., p N}, N represents the point cloud volume, then it can be represented in three-dimensional space as (X, Y, Z) = {(x1, y1, z1), (x2, y2, z2),..., (x N , y N , z N ). (X, Y, Z) represents the feature structure information of the point cloud, and the target feature structure constrained by this point cloud is R = (X, Y, Z),
[0013] Preferably, in step 2, a Gaussian mixture model in the Dirichlet process model (DPM) is selected, and its likelihood expression is:
[0014]
[0015] In the above formula, μ k represents the mean vector of the k-th parameter, and λ k represents the precision sequence. A Gaussian - gamma distribution with parameters μ0, r0, ω0, and β0 is used as its conjugate prior.
[0016] Preferably, step 3 includes the following steps:
[0017] Step 3.1: Based on the feature structure R, establish a DPM model, and the expression is as follows:
[0018]
[0019] In the above formula, indicates the feature information of each element in the point cloud set before reduction. This information is constrained by the distribution of the category vector π, and this distribution satisfies the infinite mixture ratio feature. At the same time, the Gaussian distribution models the describable target point cloud feature structure R, and then fits the likelihood function, where is the mean vector of the c i -th sequence, is the precision matrix of the c i -th sequence. The re - planned point cloud data is fed back to the input end of the network model to assist in generating a new feature structure of the point cloud target, and the DGCNN model redesigned the re - generated point cloud data.
[0020] The DGCNN guidance model reconstructs the feature structure R from the original latent variable Z. The DGCNN in the neural network structure designed in the present invention generates a corresponding set according to the extracted point cloud features. Specifically, the 3D point cloud exists as unordered elements in a set, and then is mapped to a one - to - one sampling matrix according to the DPM sampling model of dynamic programming. The assigned parameter c i is constrained by the sampler, so that each sampling set can be reduced and a new feature structure can be established, and a DGCNN can be formed, and then an MLP is connected;
[0021] Step 3.2: Under the framework of the present invention, let any measured value r iAll follow the Gaussian distribution constraint and are recursively obtained from the probabilistic latent space. Through the hidden descriptions dispersed in the dissimilar space, the general acquisition module can obtain a feature matrix subject to dissimilar dispersion. The present invention draws on the matrix modeling idea in NLP, disperses the feature matrix into multiple branches, and regenerates part of the matrix content at each time sequence. Intuitively explained, the feature structure is fitted into N vectors, that is, each vector is composed of adjacent feature sets and satisfies S (1) +S (2) +…+S (N) =S. The hidden elements corresponding to the N vectors Therefore, the present invention sequentially obtains N vectors from a probabilistic perspective, which can be expressed as:
[0022]
[0023] In the above formula, n = 1, 2,..., N, and the measurement value is recursively obtained using the hidden elements in the previous few paragraphs. The measurement value of the nth segment is modeled using a Gaussian distribution μ and where the mean vector and the diagonal covariance matrix are determined by the functions f μ and is simplified to a constant to reduce the complexity of the model parameters;
[0024] Step 3.3: In the differential Bayesian framework, the cost function of the model can be effectively optimized using gradient update algorithms such as the SGD optimizer. It is worth mentioning that the prior of z i is a learnable parameter related to sampling and depends on the sampling assignment c i ;
[0025] Preferably, the DGCNN encoder in step 4 can effectively model the local features and global relationships of point cloud data by introducing a dynamic graph and is applicable to the feature extraction and encoding of point cloud targets. The network consists of an input layer, an EdgeConv layer, a pooling layer, a fully connected layer, and an output layer.
[0026] Preferably, in step 5, according to the task requirements, the softmax function is used for classification or other methods are used for point cloud segmentation.
[0027] The present invention has the following advantages:
[0028] The adjustable reduction strategy improves the diversity of feature structures, thereby enhancing the generalization ability of the neural network. At the same time, the reduction greatly eliminates the differences in feature structures within the target class, essentially solving the problem of poor performance in few-shot 3D point cloud recognition. In addition, using the stick-breaking model in the Dirichlet process, the reduction is mapped to a decreasing trend in the infinite-dimensional spatial domain, making it learnable by the network model, so that this feature structure reduction method can be flexibly adjusted to adapt to different few-shot data sets. Omnidirectional experiments on the general 3D point cloud data set verify the efficiency of the network model designed in the present invention in terms of recognition performance compared with some SoTA method models of the same data type.
[0029] More importantly, based on flexibly adjusting the reduction intensity, the present invention greatly improves the recognition accuracy of few-shot 3D point clouds, and the recognition speed is more than 100 times that of some classical methods, establishing an effective theoretical method and model framework for the efficient recognition and classification of few-shot 3D point cloud data sets with different sample sizes, and essentially solving the problem of poor performance in few-shot 3D point cloud recognition. Brief Description of the Drawings
[0030] Figure 1 is the flow chart of a few-shot 3D point cloud recognition method based on adjustable reduction of feature structure according to the present invention;
[0031] Figure 2 is the model architecture of DPM in the present invention;
[0032] Figure 3 is the target point cloud after reducing the point cloud feature structure of two pairs of targets in the present invention;
[0033] Figure 4 is the detailed network architecture of the encoder in the present invention;
[0034] Figure 5 is the generative model composed of DGCNN and classification MLP in the present invention. Detailed Embodiments
[0035] The present invention will be further described below in conjunction with the drawings and specific embodiments:
[0036] A few-shot 3D point cloud recognition method based on adjustable reduction of feature structure, which includes the following steps:
[0037] As Figures 1 to 5 shown,
[0038] S1: Let the 3D target point cloud be the set Point = {p1, p2,..., p N}, Let \(N\) denote the point cloud volume, which can be represented in three-dimensional space as \((X, Y, Z)=\{(x_1, y_1, z_1), (x_2, y_2, z_2), \cdots, (x N , y N , z N )\}. \((X, Y, Z)\) represents the characteristic structure information of the point cloud, and the target characteristic structure constrained by the point cloud is \(R=(X, Y, Z)\).
[0039] S2: Select the Gaussian mixture model in the Dirichlet process model (DPM), and its likelihood expression is:
[0040]
[0041] In the above formula, \(\mu k \) represents the mean vector of the \(k\)-th parameter, and \(\lambda k \) represents the precision sequence. The Gaussian-gamma distribution with parameters \(\mu_0\), \(r_0\), \(\omega_0\) and \(\beta_0\) is used as its conjugate prior.
[0042] S3: The sampler with adjustable reduction intensity is used to reduce the training of the model and the characteristic structure to obtain a dynamically changing new characteristic structure of the point cloud;
[0043] S3.1: Establish a DPM model based on the characteristic structure \(T\), and the expression is as follows:
[0044]
[0045] In the above formula, indicates the characteristic information of each element in the point cloud set before reduction. This information is constrained by the distribution of the category vector \(\pi\), and this distribution satisfies the infinite mixture ratio characteristic. At the same time, the Gaussian distribution models the target point cloud characteristic structure \(R\) that can be described, and then fits the likelihood function, where is the mean vector of the \(c i -th sequence, is the precision matrix of the \(c i -th sequence. The re-planned point cloud data is fed back to the input end of the network model to assist in the generation of the new characteristic structure of the point cloud target again, and the DGCNN model redesigned the re-generated point cloud data.
[0046] The DGCNN guiding model reconstructs the characteristic structure \(R\) from the original latent variable \(Z\). As Figure 5 shown, the DGCNN in the neural network structure designed by the present invention generates a corresponding set according to the extracted point cloud characteristics. Specifically, the 3D point cloud exists as unordered elements in a set, and then is mapped to a one-to-one sampling matrix according to the DPM sampling model of dynamic programming. Assign the parameter \(c iConstrained by the sampler, each sampling set can be reduced and a new feature structure can be established, and a DGCNN can be formed, followed by connecting an MLP;
[0047] S3.2: Under the framework of the present invention, let any measurement value r i all follow the Gaussian distribution constraint and be recursively obtained from the probability latent space. Through the hidden descriptions dispersed in the dissimilar space, the general acquisition module can obtain a feature matrix that obeys dissimilar dispersion. The present invention draws on the matrix modeling idea in NLP, disperses the feature matrix into multiple branches, and regenerates part of the content of the matrix at each time series. Intuitively explained, the feature structure is fitted to N vectors, that is, each vector is composed of adjacent feature sets and satisfies S (1) +S (2) +…+S (N) =S. The hidden elements corresponding to the N vectors correspond, and its spatial domain range is Therefore, the present invention sequentially obtains N vectors from a probabilistic perspective, which can be expressed as:
[0048]
[0049] In the above formula, n = 1, 2,..., N, and the measurement value is recursively obtained using the hidden elements in the previous few paragraphs . The measurement value of the nth segment is modeled using the Gaussian distribution , where the mean vector and the diagonal covariance matrix are determined by the functions f μ and . The above two functions are both non-linear functions and can be implemented using a neural network. Specifically, f μ is provided by the DGCNN feature encoder, simplified to a constant to reduce the complexity of the model parameters;
[0050] S3.3: In the differential Bayesian framework, the cost function of the model can be effectively optimized using gradient update algorithms such as the SGD optimizer. It is worth mentioning that the prior of z i is a learnable parameter related to sampling and depends on the sampling assignment c i ;
[0051] S4: By introducing a dynamic graph, the DGCNN encoder can effectively model the local features and global relationships of point cloud data and is suitable for feature extraction and encoding of point cloud targets. The network consists of an input layer, an EdgeConv layer, a pooling layer, a fully connected layer, and an output layer, and the network architecture is as Figure 4as shown
[0052] S5: According to the task requirements, use the softmax function for classification or use other methods for point cloud segmentation.
[0053] Experimental Results and Analysis:
[0054] 1. Experimental Environment
[0055] The experimental development environment is: Intel(R) Xeon(R) Platinum 8270 CPU + NVIDIA Ampere A40, 48GB memory GPU; the development tool is: Pytorch framework.
[0056] 2. Experimental Datasets
[0057] ModelNet40 is a commonly used 3D object recognition dataset for training and evaluating the performance of deep learning models in 3D object classification tasks. This dataset was created by the ShapeNet project team at Stanford University and released in 2015. It contains 3D object models of 40 different categories, and these categories cover common object categories such as chairs, tables, beds, monitors, cabinets, etc. Each category has approximately 1,024 different 3D model samples, with a total of approximately 40,000 samples. These model samples are extracted from the real 3D model library in the ShapeNet project, including objects with various viewpoints and poses.
[0058] The Sydney Urban Objects dataset contains various common urban road objects scanned using a Velodyne HDL-64E lidar in the central business district of Sydney, Australia. The dataset includes 631 objects from different scans, covering categories such as vehicles, pedestrians, signs, and trees. The collection of this dataset aims to test matching and classification algorithms. It aims to provide the non-ideal perception conditions of an actual urban perception system, including a large number of variations in viewpoints and occlusions.
[0059] To verify the effectiveness of the adjustable feature structure reduction proposed in the present invention, we conduct tests on the above two mainstream international general 3D point cloud datasets.
[0060] 3. Experimental Results and Analysis
[0061] Experiment 1 was first tested on the above two mainstream international 3D point cloud datasets. The comparison methods were some classic neural network-based 3D point cloud recognition methods in recent years. Only the present invention has a feature structure adjustable reduction. The results show that the present invention performs best on the ModelNet40 dataset with an accuracy of 94.8%, and achieves excellent results of 93.9% on the SydneyUrban Object dataset. The RS-CNN method also performs prominently on the ModelNet40 dataset with an accuracy of 93.1% and 89.0% on the Sydney Urban Object dataset. The accuracies of methods such as PVNet, PointNet++, and DGCNN all exceed 92%. The performances of VoxNet and PointNet are relatively low, especially on the SydneyUrban Object dataset, with performances of 79.8% and 88.4% respectively. Generally speaking, the present invention performs well on these two datasets, with the accuracy leading other methods. The proposed feature structure adjustable reduction method has a significant improvement in the recognition accuracy of the entire dataset. Since the present invention uses DGCNN as the feature encoder of the proposed method, the effectiveness of the proposed method of the present invention is more significant compared with the DGCNN method without feature structure adjustable reduction.
[0062] Experiment 2 was conducted for few-shot learning tasks, and the above two international mainstream datasets were also selected. In these two datasets, few-shot datasets were randomly selected respectively. We evaluated our method under the few-shot learning setting according to previous studies. The typical setting is "K-way N-shot". First, K categories are randomly selected, and then 20 new objects are sampled for each category. The model is trained on K×N samples (support set) and evaluated on the remaining 20K new samples (query set). In our experiment, we tested the performance of "5-way 10-shot", "5-way 20-shot", "10-way 10-shot" and "10-way 20-shot". We conducted 5 independent experiments under each setting and reported the average performance of 5 runs. In addition, in the experiment, we used 2048 points in each 3D point cloud sample, and the initial feature constraint of each point only considered the three-dimensional coordinates. The results show that the present invention performs outstandingly in all settings. Especially, it achieved accuracies of 92.5% and 94.0% in the 5-way 10-shot and 20-shot tasks of the ModelNet40 dataset respectively, and achieved scores of 92.0% and 94.5% in the 10-way 10-shot and 20-shot tasks of the Sydney UrbanObjects dataset respectively. The method of Nie et al. performed excellently in the 10-way 10-shot task of the ModelNet40 dataset, reaching the highest accuracy of 96.5%, while the PointBert method reached 92.5% in the 5-way 20-shot task of Sydney Urban Objects. Generally speaking, the present invention performs well in multiple few-shot tasks and has significant advantages in the few-shot learning scenario. In addition, it can be seen from the experimental results that the method proposed in the present invention has better recognition accuracy than other SoTA methods in the few-shot 3D point cloud learning task, indicating the effectiveness of the feature structure adjustable reduction strategy proposed in the present invention.
[0063] The present invention uses the features after adjustable reduction of the feature structure as the network input. Compared with the data constrained by the original feature structure, the feature data to be processed in the network is less, greatly reducing the complexity and the number of parameters of the model. In Experiment 3, the training time and the number of parameters of the network model for the 5-way, 10-shot few-shot dataset were tested on the same experimental platform. The average experimental results of five times showed that the present invention exhibited significant advantages in terms of training time and model complexity. The training time was only 3.0 seconds, and the number of parameters was only 0.18 million. In contrast, the training time and the number of parameters of other methods were relatively large. For example, the training time of the PointBert method was 305.3 seconds, and the number of parameters was 2.89 million. The training times of the SS-FLS, Nie etal., and Ye et al. methods all exceeded 150 seconds, and the number of parameters was between 1.77 and 2.06 million. It is obvious from the experimental results that the network model of the few-shot shape recognition method based on adjustable reduction of the feature structure proposed by the present invention has fewer parameters and faster training time, and thus it is easy to know that it has an extremely fast recognition speed, providing strong real-time support for the practical application of 3D point cloud recognition.
[0064] Based on the above experimental results, the experimental results in multiple aspects on the mainstream synthetic and real 3D point cloud datasets have shown the effectiveness of the method proposed by the present invention compared with some SoTA methods in terms of recognition performance. In addition, on the basis of effectively improving the recognition accuracy, the recognition speed of the present invention is up to 100 times higher than that of some classical methods, providing strong theoretical support for the practical application of few-shot 3D point clouds, especially in terms of real-time performance.
Claims
1. A few-shot 3D point cloud recognition method based on feature structure adjustable reduction, characterized in that: It includes the following steps: Step 1: Define the spatial coordinates of the target point cloud input to the network during 3D point cloud recognition as its feature structure; Step 2: Map the reduced intensity to an infinite-dimensional space through Dirichlet processing so that it can be learned by the network, and obtain an adjustable reduced intensity through the learning of the network model; Step 3: Reduce the training of the model and the feature structure through a sampler with adjustable reduced intensity to obtain a dynamically changing new feature structure of the point cloud; Step 4: Input the point cloud information after reducing the target point cloud, select DGCNN as the feature encoder, and obtain the target point cloud features; Step 5: Use the reduced features as input and output the final target recognition accuracy through a classifier; In step 1, let the 3D target point cloud be the set Point = {p1, p2,..., p N}, If N represents the point cloud volume, then it can be represented in three-dimensional space as (X, Y, Z) = {(x1, y1, z1), (x2, y2, z2),..., (x N , y N , z N )}; (X, Y, Z) represents the characteristic structure information of the point cloud, and the target characteristic structure constrained by the point cloud is R = (X, Y, Z). In step 2, the Gaussian mixture model in the Dirichlet process model (DPM) is selected, and its likelihood expression is: In the above formula, μ k represents the mean vector of the k-th parameter, and λ k represents the precision sequence; the Gaussian-Gamma distribution with parameters μ0, r0, ω0, and β0 is used as its conjugate prior; Step 3 includes the following steps: Step 3.1: Establish a DPM model based on the constraint R, and the expression is as follows: In the above formula, indicates the feature information of each element in the point cloud set before reduction The information is constrained by the distribution of the category vector π, and this distribution satisfies the infinite mixture ratio feature; at the same time, the Gaussian distribution models the describable target point cloud feature structure R, and then fits the likelihood function, where is the mean vector of the c i -th sequence, is the precision matrix of the c i -th sequence; the feedback of the re-planned point cloud data to the input end of the network model helps to generate a new feature structure of the point cloud target again, and the DGCNN model redesigned the re-generated point cloud data; The DGCNN guiding model reconstructs the feature structure R from the original latent variable Z. The DGCNN in the neural network structure designed by the present invention generates corresponding sets according to the extracted point cloud features. Specifically, the 3D point cloud exists as unordered elements in a set, and then is mapped to a one-to-one corresponding sampling matrix according to the DPM sampling model of dynamic programming; the parameter c is assigned. i Constrained by the sampler, each sampling set can be reduced and a new feature structure can be established, and a DGCNN can be formed, and then an MLP is connected. Step 3.
2. Under the framework of the present invention, let any measured value r i all follow the Gaussian distribution constraint and be recursively obtained from the probabilistic latent space; through the hidden descriptions dispersed in the dissimilar space, the general acquisition module can obtain a feature matrix that obeys dissimilar dispersion; the present invention draws on the matrix modeling idea in NLP, disperses the feature matrix into multiple branches, and regenerates part of the content of the matrix at each time series; intuitively explained, the feature structure is fitted into N vectors, that is each vector is composed of adjacent feature sets and satisfies S (1) +S (2) +…+S (N) =S; corresponding to the hidden elements corresponding to the N vectors, the spatial domain range thereof is Therefore, the present invention sequentially obtains N vectors from the perspective of probability, which can be expressed as: In the above formula, n = 1, 2, ..., N, the measured value is recursively obtained using the hidden elements of the previous few segments The measured value of the nth segment is modeled using a Gaussian distribution where the mean vector and the diagonal covariance matrix are determined by functions f μ and Both of the above two functions are non-linear functions and can be implemented using a neural network; specifically, f μ is provided by the DGCNN feature encoder is simplified to a constant to reduce the complexity of the model parameters; Step 3.
3. In the differential Bayesian framework, the cost function of the model can be effectively optimized using gradient update algorithms such as the SGD optimizer; it is worth mentioning that the prior of z i is a learnable parameter related to sampling, which depends on the sampling assignment c i ; The DGCNN encoder in step 4 can effectively model the local features and global relationships of point cloud data by introducing a dynamic graph, and is suitable for feature extraction and encoding of point cloud targets; this network consists of an input layer, an EdgeConv layer, a pooling layer, a fully connected layer, and an output layer; In step 5, according to the task requirements, use the softmax function for classification or use other methods for point cloud segmentation.