Point cloud segmentation method for cross-domain active learning based on prototype guidance
Through the cross-domain active learning method based on prototype guidance, pseudo-labels are generated and the difference scores and uncertainty scores are calculated, and only some target domain point cloud data are annotated. Through the construction of source prototypes and fine-tuning model methods, the problems of insufficient generalization ability and high labeling cost of three-dimensional point cloud semantic segmentation technology on cross-domain data sets are solved, achieving more efficient domain adaptation and active learning effects.
Patent Information
- Application Number
- CN202510175684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The existing three-dimensional point cloud semantic segmentation technology has insufficient generalization capabilities on cross-domain data sets and is highly dependent on labeled data, resulting in high labeling costs and unstable model performance.
A cross-domain active learning point cloud segmentation method based on prototype guidance is proposed. By obtaining target point cloud data and inputting the trained point cloud segmentation model, a pseudo-label is generated and the difference score and uncertainty score are calculated. Only some target domain point cloud data are annotated, and the model adaptability and generalization ability are improved through the method of building source prototypes and fine-tuning the model.
It effectively reduces the amount of labeled data, reduces the cost of labeling, and improves the performance of the model in the target domain and the generalization ability of the cross-domain data sets by reasonably selecting and labeling target samples.
Smart Images

Figure CN120107586A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of point cloud data processing, and mainly refers to a point cloud segmentation method based on prototype-guided cross-domain active learning. Background Art
[0002] Point cloud semantic segmentation plays a vital role in 3D environment perception and is widely used in fields such as autonomous driving, navigation, and robotics. In recent years, point cloud segmentation technology based on fully supervised learning has made significant breakthroughs with the support of dense annotation. However, for 3D point cloud data, large-scale point-level annotation is extremely time-consuming and costly. In practical applications, it is often unrealistic to annotate semantic information for all points to train a high-precision semantic segmentation model. In addition, the generalization performance of pre-trained models on different datasets varies significantly, which limits their practical application effects.
[0003] In order to solve these problems, a variety of unsupervised domain adaptation (UDA) methods have been proposed in recent years to promote the effective transfer of knowledge from the source domain to the target domain. For example, Yuan et al. proposed a density-guided converter to adjust the density distribution of point clouds between domains and incorporated it into a two-stage self-training framework. Li et al. proposed Context Collaboration Learning (CCL), which bridges the context differences across domains through a local context layer and a global prototype-based attention mechanism, thereby enhancing the learning ability of domain-invariant features and improving the adaptability of the model. However, although these UDA methods reduce the dependence on labeled data, their performance still has a significant gap compared to fully supervised methods.
[0004] In order to further improve the performance of the target domain, active learning (AL) has gradually become an important means. The recent active domain adaptation (ADA) method combines the idea of active learning and greatly improves the adaptability performance by selecting and labeling a small number of information-rich target samples. However, the existing ADA method mainly relies on the prediction results of the segmentation model to select key target points, without fully considering the differences between the source domain and the target domain. Since the segmentation model is trained in the source domain, the sample selection based on the prediction results can easily lead to the distribution of the selected point cloud samples being too localized, thereby reducing the diversity and representativeness of the data and having an adverse effect on the model performance.
[0005] In summary, although deep learning-based 3D point cloud semantic segmentation technology has made great progress, the generalization ability of the model on cross-domain datasets still needs to be further improved. This requires continuous exploration of more efficient domain adaptation and active learning methods while reducing annotation costs to provide more robust solutions for practical applications. Summary of the invention
[0006] In order to solve the problems existing in the background technology, one aspect of the present invention provides a point cloud segmentation method based on prototype-guided cross-domain active learning, comprising: acquiring target point cloud data, inputting the target point cloud data into a trained point cloud segmentation model, and obtaining a segmentation result of the target point cloud data; wherein the training process of the point cloud segmentation model includes:
[0007] S1: Receive a source domain point cloud dataset with labels and an unlabeled target domain point cloud dataset;
[0008] S2: Use the source domain point cloud dataset to train the point cloud segmentation model to obtain a source domain segmentation model, and use the source domain segmentation model to extract the semantic features of the source domain point cloud data;
[0009] S3: Construct the source prototype of each category according to the mean value of the semantic features of the source domain point cloud data under each category;
[0010] S4: Input the target domain point cloud data into the source domain segmentation model to predict the pseudo labels of the target domain point cloud data and extract the semantic features of the target domain point cloud data;
[0011] S5: Generate a difference score of the target domain point cloud data based on the distance between the semantic features of the target domain point cloud data and the source prototypes of each category; and calculate the uncertainty score of the target domain point cloud data based on the pseudo-labels of the target domain point cloud data;
[0012] S7: combining the difference score and uncertainty score of the target domain point cloud data to label some points in the target domain point cloud data to obtain partial target domain point cloud data with labels;
[0013] S8: Extract part of the source domain point cloud data and part of the target domain point cloud data with labels from the source domain point cloud data to form new training samples, and fine-tune the source domain segmentation model to obtain a trained point cloud segmentation model.
[0014] Another aspect of the present invention provides a point cloud segmentation device based on prototype-guided cross-domain active learning, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory, and is used to execute the computer program stored in the memory, so that the point cloud segmentation device based on prototype-guided cross-domain active learning performs the point cloud segmentation method based on prototype-guided cross-domain active learning.
[0015] Yet another aspect of the present invention is a computer-readable storage medium storing a program, which, when executed by a processor, implements the point cloud segmentation method based on prototype-guided cross-domain active learning.
[0016] The present invention has at least the following beneficial effects
[0017] The present invention utilizes an unlabeled target domain point cloud dataset, generates pseudo labels and calculates difference scores and uncertainty scores, and only labels part of the target domain point cloud data. Compared with the fully supervised learning method, the amount of labeled data is greatly reduced, thereby reducing the labeling cost. The present invention combines the idea of active learning and proposes a cross-domain active learning method based on prototype guidance. By reasonably selecting and labeling a small number of information-rich target samples, the adaptive performance is greatly improved. Compared with the traditional unsupervised domain adaptation method, it can effectively improve the performance of the model in the target domain. The present invention constructs a source prototype and generates a difference score based on the distance between the semantic features of the target domain point cloud data and the source prototype. It comprehensively considers the differences between the source domain and the target domain, improves the diversity and representativeness of the selected point cloud samples, and thus enhances the generalization ability of the model on cross-domain datasets, providing a more robust solution for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0019] Figure 2 It is a schematic diagram of the process framework of the present invention;
[0020] Figure 3 This is a visual comparison diagram of the segmentation results of the method of the present invention. DETAILED DESCRIPTION
[0021] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0022] See also Figure 1 and Figure 2One aspect of the present invention provides a point cloud segmentation method based on prototype-guided cross-domain active learning, comprising: acquiring target point cloud data, inputting the target point cloud data into a trained point cloud segmentation model, and obtaining a segmentation result of the target point cloud data; wherein the training process of the point cloud segmentation model includes:
[0023] S1: Receive a source domain point cloud dataset with labels and an unlabeled target domain point cloud dataset;
[0024] In an embodiment of the present invention, the target point cloud data is a set of point cloud data to be segmented, and the category of each point in the set needs to be marked, and the target point cloud data is segmented according to the category. The set of point cloud data to be segmented can be collected by a collection device on an autonomous driving vehicle, for example.
[0025] S2: Use the source domain point cloud dataset to train the point cloud segmentation model to obtain a source domain segmentation model, and use the source domain segmentation model to extract the semantic features of the source domain point cloud data;
[0026] Preferably, the point cloud segmentation model comprises: a feature extraction module, used to extract the semantic features of each point in the input point cloud data; and a segmentation head, used to predict the category of each point according to the semantic features of each point in the point cloud data.
[0027] S3: Construct the source prototype of each category according to the mean value of the semantic features of the source domain point cloud data under each category;
[0028] Preferably, constructing the source prototype of each category according to the semantic feature mean of the source domain point cloud data under each category comprises: inputting the source domain point cloud data into the source domain segmentation model in batches, and gradually calculating the semantic feature mean of the source domain point cloud data using an incremental average algorithm according to the category labels of the points in the source domain point cloud data to construct the source prototype of the category:
[0029]
[0030] in, represents the mean value of the semantic features of c categories in the source domain point cloud data of the bth batch, represents the number of points in the c categories in the source domain point cloud data of the bth batch; f i c The semantic features of the i-th point under the c-th category in the b-th batch of source domain point cloud data; C represents the number of categories; when b = B, represents the source prototype of the cth category, and B represents the number of batches.
[0031] S4: Input the target domain point cloud data into the source domain segmentation model to predict the pseudo labels of the target domain point cloud data and extract the semantic features of the target domain point cloud data;
[0032] S5: Generate a difference score of the target domain point cloud data based on the distance between the semantic features of the target domain point cloud data and the source prototypes of each category; and calculate the uncertainty score of the target domain point cloud data based on the pseudo-labels of the target domain point cloud data;
[0033] Preferably, the difference score of the target domain point cloud data includes:
[0034] S51: Calculate the distance between the semantic features of each point in the target domain point cloud data and the source prototype of each category:
[0035]
[0036] Among them, f i t represents the semantic features of the i-th point in the target domain point cloud data, p c represents the source prototype of the cth category; represents the distance between the semantic feature of the i-th point in the target domain point cloud data and the source prototype of the c-th category; D i Represents the distance set between the semantic features of the i-th point in the target domain point cloud data and the source prototypes of each category;
[0037] S52: taking the derivative of the distance between the semantic feature of each point in the target domain point cloud data and the source prototype of each category, and processing it through the softmax function to obtain the similarity between the semantic feature of each point in the target domain point cloud data and the source prototype of each category;
[0038] S53: Calculate the difference score of the target domain point cloud data according to the maximum and second largest similarity between the semantic features of each point in the target domain point cloud data and the source prototypes of each category:
[0039]
[0040] Among them, f(.) represents the softmax function, and max(.) represents the maximum value; represents the difference score of the i-th point in the target domain point cloud data.
[0041] Preferably, calculating the uncertainty score of the target domain point cloud data according to the pseudo-labels of the target domain point cloud data comprises:
[0042]
[0043] in, Represents the uncertainty score of the i-th point in the target domain point cloud data; represents the predicted pseudo label of the i-th point in the target domain point cloud data.
[0044] S7: combining the difference score and uncertainty score of the target domain point cloud data to label some points in the target domain point cloud data to obtain partial target domain point cloud data with labels;
[0045] Preferably, the step S7 comprises:
[0046] S71: Calculate the comprehensive score of the target domain point cloud by combining the difference score and uncertainty score of the target domain point cloud data:
[0047]
[0048] in, represents the comprehensive score of the i-th point in the target domain point cloud data, α is the weight parameter, and its value range is 0~1. represents the uncertainty score of the i-th point in the target domain point cloud data, Represents the uncertainty score of the i-th point in the target domain point cloud data;
[0049] S72: Sort the points in the target domain point cloud data according to the comprehensive scores of the points in the target domain point cloud data, select k points with the largest comprehensive scores as candidate points, and annotate the selected k candidate points to obtain partial target domain point cloud data with labels.
[0050] Preferably, the number of points in the partial source domain point cloud data extracted from the source domain point cloud data and the partial target domain point cloud data with labels is the same.
[0051] S8: Extract part of the source domain point cloud data and part of the target domain point cloud data with labels from the source domain point cloud data to form new training samples, and fine-tune the source domain segmentation model to obtain a trained point cloud segmentation model.
[0052] In an embodiment of the present invention, in view of the high cost and annotation redundancy caused by the existing semantic segmentation neural network relying on a large amount of point-level annotation data, and considering the differences between cross-domain data sets, a point cloud semantic segmentation neural network method based on active learning is proposed, aiming to reduce the annotation cost and improve the generalization performance of the model. The method first pre-trains the point cloud semantic segmentation neural network through deep learning technology to obtain a pre-trained model. The model is used to perform semantic prediction on the source domain and target domain data sets respectively, and the difference scores between the two are calculated. The candidate points for the target domain point cloud segmentation are screened out by combining the uncertainty score of the model on the target domain. Subsequently, a balanced source-target hybrid strategy is adopted to randomly select points from the source domain point cloud that are equal to the target domain annotation data for matching, and the two are mixed to generate an intermediate data set, which is used for fine-tuning the segmentation model. The fine-tuned segmentation model can generate a predicted category and a predicted probability for each point cloud data, where the predicted probability reflects the confidence of the corresponding category. By combining the predicted probability and category information, the segmentation result of the target point cloud data can be accurately obtained.
[0053] In order to better illustrate the cross-domain point cloud semantic segmentation neural network method based on active learning of the present invention, the present invention describes in detail the active learning point cloud semantic segmentation neural network and its corresponding training process.
[0054] In the point cloud domain adaptive segmentation task, there is a labeled source domain dataset and an unlabeled target domain dataset. The goal of this method is to achieve high-precision segmentation on the target domain dataset based on an active learning strategy through the pre-trained model trained on the source domain dataset. Figure 2 is a schematic diagram of point cloud semantic segmentation based on active learning implemented in the present invention, such as Figure 2 shown.
[0055] In the example of the present invention, a segmentation network (MinkNet) constructed in Minkowski space is used to extract the semantic features of point clouds. MinkNet uses sparse convolution operations to effectively reduce computational redundancy, which is very suitable for sparse data such as point clouds and supports multi-scale feature extraction.
[0056] This method focuses on the cross-domain adaptive point cloud segmentation task. Based on the pre-trained segmentation model trained on the source domain data set, the present invention fine-tunes the segmentation model by adopting an active learning strategy to achieve high-precision segmentation of the target domain data set. The source domain data set and the target domain data set are two different data sets, the source domain data set is annotated, and the target domain data set is not annotated, which is the data set to be segmented.
[0057] In this method, in order to ensure the reliability and scalability of data selection, the difference between the target domain data point cloud and the source domain dataset point cloud needs to be considered. In order to describe the structural characteristics of the source domain dataset point cloud, the source prototype is first constructed based on the feature tree of each category. For a class of points in the source, the segmentation model is used to extract the feature vector of the point cloud, and its average feature vector is used as the prototype of the class.
[0058] like Figure 3 As shown in Figure 1, this method shows the visual segmentation results from SysLidAR to SemanticKITTI. Compared with the true label, target domain only, Margin, Entropy and random methods, the segmentation effect of this method is significantly better than other methods.
[0059] Another aspect of the present invention provides a point cloud segmentation device based on prototype-guided cross-domain active learning, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory, and is used to execute the computer program stored in the memory, so that the point cloud segmentation device based on prototype-guided cross-domain active learning performs the point cloud segmentation method based on prototype-guided cross-domain active learning.
[0060] Yet another aspect of the present invention is a computer-readable storage medium storing a program, which, when executed by a processor, implements the point cloud segmentation method based on prototype-guided cross-domain active learning.
[0061] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0062] In summary, the present invention utilizes an unlabeled target domain point cloud dataset, generates pseudo labels and calculates difference scores and uncertainty scores, and only labels part of the target domain point cloud data. Compared with the fully supervised learning method, the amount of labeled data is greatly reduced, thereby reducing the labeling cost. The present invention combines the idea of active learning and proposes a prototype-guided cross-domain active learning method, which greatly improves the adaptive performance by reasonably selecting and labeling a small number of information-rich target samples. Compared with the traditional unsupervised domain adaptation method, it can effectively improve the performance of the model in the target domain. The present invention constructs a source prototype and generates a difference score based on the distance between the semantic features of the target domain point cloud data and the source prototype. It comprehensively considers the differences between the source domain and the target domain, improves the diversity and representativeness of the selected point cloud samples, and thus enhances the generalization ability of the model on cross-domain datasets, providing a more robust solution for practical applications.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A point cloud segmentation method based on prototype-guided cross-domain active learning, characterized in that: include: Obtain target point cloud data, input the target point cloud data into a trained point cloud segmentation model, and obtain a segmentation result of the target point cloud data; wherein the training process of the point cloud segmentation model includes: S1: Receive a source domain point cloud dataset with labels and an unlabeled target domain point cloud dataset; S2: Use the source domain point cloud dataset to train the point cloud segmentation model to obtain a source domain segmentation model, and use the source domain segmentation model to extract the semantic features of the source domain point cloud data; S3: Construct the source prototype of each category according to the mean value of the semantic features of the source domain point cloud data under each category; S4: Input the target domain point cloud data into the source domain segmentation model to predict the pseudo labels of the target domain point cloud data and extract the semantic features of the target domain point cloud data; S5: Generate a difference score of the target domain point cloud data based on the distance between the semantic features of the target domain point cloud data and the source prototypes of each category; and calculate the uncertainty score of the target domain point cloud data based on the pseudo-labels of the target domain point cloud data; S7: combining the difference score and uncertainty score of the target domain point cloud data to label some points in the target domain point cloud data to obtain partial target domain point cloud data with labels; S8: Extract part of the source domain point cloud data and part of the target domain point cloud data with labels from the source domain point cloud data to form new training samples, and fine-tune the source domain segmentation model to obtain a trained point cloud segmentation model.
2. According to claim 1, a point cloud segmentation method based on prototype-guided cross-domain active learning is characterized in that: The method of constructing a source prototype of each category according to the semantic feature mean of the source domain point cloud data includes: inputting the source domain point cloud data into the source domain segmentation model in batches, and using an incremental average algorithm to gradually calculate the semantic feature mean of the source domain point cloud data according to the category labels of the points in the source domain point cloud data to construct the source prototype of the category: Among them, P c b represents the mean value of the semantic features of c categories in the source domain point cloud data of the bth batch, represents the number of points in the c categories in the source domain point cloud data of the bth batch; f i c The semantic features of the i-th point under the c-th category in the b-th batch of source domain point cloud data; C represents the number of categories; when b = B, represents the source prototype of the cth category, and B represents the number of batches.
3. The point cloud segmentation method based on prototype-guided cross-domain active learning according to claim 1, characterized in that: The difference scores of the target domain point cloud data include: S51: Calculate the distance between the semantic features of each point in the target domain point cloud data and the source prototype of each category: Among them, f i t represents the semantic features of the i-th point in the target domain point cloud data, p c represents the source prototype of the cth category; represents the distance between the semantic feature of the i-th point in the target domain point cloud data and the source prototype of the c-th category; D i Represents the distance set between the semantic features of the i-th point in the target domain point cloud data and the source prototypes of each category; S52: taking the derivative of the distance between the semantic feature of each point in the target domain point cloud data and the source prototype of each category, and processing it through the softmax function to obtain the similarity between the semantic feature of each point in the target domain point cloud data and the source prototype of each category; S53: Calculate the difference score of the target domain point cloud data according to the maximum and second largest similarity between the semantic features of each point in the target domain point cloud data and the source prototypes of each category: Among them, f(.) represents the softmax function, and max(.) represents the maximum value; represents the difference score of the i-th point in the target domain point cloud data.
4. The point cloud segmentation method based on prototype-guided cross-domain active learning according to claim 1, characterized in that: Calculating the uncertainty score of the target domain point cloud data according to the pseudo-label of the target domain point cloud data includes: in, represents the uncertainty score of the i-th point in the target domain point cloud data; P i t represents the predicted pseudo label of the i-th point in the target domain point cloud data.
5. The point cloud segmentation method based on prototype-guided cross-domain active learning according to claim 1, characterized in that: The step S7 comprises: S71: Calculate the comprehensive score of the target domain point cloud by combining the difference score and uncertainty score of the target domain point cloud data: in, represents the comprehensive score of the i-th point in the target domain point cloud data, α is the weight parameter, and its value range is 0~1. represents the uncertainty score of the i-th point in the target domain point cloud data, Represents the uncertainty score of the i-th point in the target domain point cloud data; S72: Sort the points in the target domain point cloud data according to the comprehensive scores of the points in the target domain point cloud data, select k points with the largest comprehensive scores as candidate points, and annotate the selected k candidate points to obtain partial target domain point cloud data with labels.
6. The point cloud segmentation method based on prototype-guided cross-domain active learning according to claim 1, characterized in that: The number of points in the part of the source domain point cloud data extracted from the source domain point cloud data and the part of the target domain point cloud data with labels is the same.
7. The point cloud segmentation method based on prototype-guided cross-domain active learning according to claim 1, characterized in that: The point cloud segmentation model includes: a feature extraction module, which is used to extract the semantic features of each point in the input point cloud data; and a segmentation head, which is used to predict the category of each point according to the semantic features of each point in the point cloud data.
8. A point cloud segmentation device based on prototype-guided cross-domain active learning, characterized in that: It includes a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the point cloud segmentation device based on prototype-guided cross-domain active learning performs the point cloud segmentation method based on prototype-guided cross-domain active learning described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, a point cloud segmentation method based on prototype-guided cross-domain active learning as described in any one of claims 1 to 7 is implemented.