Method for training point cloud sequence semantic segmentation model and application thereof
By constructing intra-frame semantic affinity graphs and category affinity graphs, and combining cross-graph convolution and self-distillation mechanisms, pseudo-labels are propagated across frames. This solves the problems of low pseudo-label accuracy and error accumulation in sequence point cloud semantic segmentation, improves semantic segmentation accuracy, and approaches the performance of fully supervised methods.
Patent Information
- Application Number
- CN202511280572.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing semantic segmentation methods for sequential point clouds fail to make sufficient use of spatiotemporal information during pseudo-label propagation, resulting in low pseudo-label accuracy and severe error accumulation, making them difficult to apply effectively in large-scale real-world application scenarios.
We construct intra-frame semantic affinity graphs and category affinity graphs, combine cross-graph convolution and self-distillation mechanisms, propagate pseudo-labels across frames by predicting the flow field, use cross-graph convolution and self-distillation mechanisms to propagate and update feature information, and design a weighted confidence fusion mechanism to optimize pseudo-label quality.
It significantly improves semantic segmentation accuracy with extremely low annotation cost, and its performance is close to that of fully supervised methods. It effectively solves the problems of insufficient utilization of spatiotemporal information and error accumulation in the process of pseudo-label propagation.
Smart Images

Figure CN120765946B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer vision and artificial intelligence, more specifically, relates to a point cloud sequence semantic segmentation model training method and application. BACKGROUND
[0002] Sequence point cloud semantic segmentation is an important task to realize spatio-temporal understanding of dynamic three-dimensional scenes, and has wide application value in the fields of autonomous driving, intelligent robots, space monitoring, etc. With the rapid development of dynamic sensors such as multi-line laser radar, it has become more efficient and routine to obtain continuous three-dimensional point cloud sequence data, providing a rich data basis for spatio-temporal modeling and dynamic semantic analysis.
[0003] In recent years, thanks to the rapid development of deep learning, the method of semantic segmentation of sequence point cloud has made continuous progress. Typical methods mostly rely on the fully supervised training paradigm, training the model by fine labeling each point in the continuous frames. However, sequence point cloud data has the characteristics of large volume, high frame density, complex scene, etc. Point-by-point labeling not only has high cost and long time, but also is difficult to extend to large-scale real application scenarios.
[0004] In order to alleviate the labeling burden, weakly supervised learning methods have gradually become an important research direction in sequence point cloud semantic segmentation. Weakly supervised methods attempt to assist the training process of a large number of unlabeled frames with a small amount of labeled data, and improve the model's ability to utilize unlabeled data by designing pseudo-label generation mechanisms, spatio-temporal information propagation structures, and self-supervised feature optimization strategies. In sequence point cloud, due to the strong correlation and motion consistency between frames, how to reasonably utilize the spatio-temporal continuity information to guide the pseudo-label propagation and avoid label error accumulation becomes a core problem faced by weakly supervised methods.
[0005] Currently, existing research attempts to model the intra-frame spatial relationship through graph structure, estimate the inter-frame point correspondence relationship using scene flow, and combine confidence filtering and feature distillation mechanism to improve the quality of pseudo-labels. However, existing methods still have great performance bottlenecks when facing problems such as label drift between frames, prediction uncertainty, and class imbalance, making it difficult to fully exploit the potential of sequence structure in semantic enhancement. Therefore, how to construct a robust intra-frame-inter-frame structure representation, improve the stability of label propagation, and improve the learning efficiency under weak supervision, is an important technical problem to be solved in the field of sequence point cloud semantic segmentation. SUMMARY
[0006] In view of the deficiencies of the prior art, the purpose of the present application is to provide a point cloud sequence semantic segmentation model training method and application, to solve the problems of insufficient spatio-temporal information utilization, low pseudo-label precision and error accumulation in the pseudo-label propagation process in the prior art.
[0007] Based on the above objectives, a first aspect of the present invention provides a training method for a point cloud sequence semantic segmentation model, characterized by comprising: acquiring a point cloud sequence, the point cloud sequence containing multiple point cloud frames, some of which are unlabeled; performing supervoxel segmentation on the point cloud frames, and constructing an intra-frame semantic affinity map and an intra-frame class affinity map based on the segmented supervoxels; processing the intra-frame semantic affinity map and the intra-frame class affinity map respectively according to the cross-graph convolution mechanism and the self-distillation mechanism, so as to propagate and update the feature information of the supervoxels in the intra-frame semantic affinity map and the intra-frame class affinity map; matching supervoxels between adjacent point cloud frames to obtain a predicted flow field, and performing pseudo-label cross-frame propagation based on the predicted flow field to generate pseudo-labels for unlabeled point cloud frames; and training a point cloud sequence semantic segmentation model using labeled point cloud frames and pseudo-labeled point cloud frames in the point cloud sequence.
[0008] Preferably, the constructed intra-frame semantic affinity graph is as follows:
[0009] ;
[0010] The constructed intra-frame class affinity graph is as follows:
[0011] ;
[0012] in, Let be the element in the i-th row and j-th column of the affinity matrix of the intra-frame semantic affinity graph. Let be the element in the i-th row and j-th column of the affinity matrix of the category semantic affinity graph. As the first hyperparameter, This is the second hyperparameter. As the first fixed parameter, As the second fixed parameter, Let be the Euclidean distance between supervoxel i and supervoxel j. The feature difference between supervoxel i and supervoxel j. Let k be the set of the neighbors of supervox i in frame t. , Let be the class prediction probability vectors of supervoxel i and supervoxel j, respectively. , Let i and j be the predicted probabilities of supervoxels i and j corresponding to class b, respectively, with superscripts indicating their respective probabilities. Indicates the first Next iteration, superscript Indicates the first Frame-by-frame cloud.
[0013] Preferably, the intra-frame semantic affinity graph and the intra-frame category affinity graph are processed according to a cross-graph convolution mechanism, specifically including: alternately performing the following two graph convolution processing; performing graph convolution processing on the intra-frame semantic affinity graph guided by the affinity matrix of the intra-frame category affinity graph; performing graph convolution processing on the intra-frame category affinity graph guided by the affinity matrix of the intra-frame semantic affinity graph.
[0014] Preferably, the intra-frame semantic affinity graph and the intra-frame category affinity graph are processed according to a self-distillation mechanism, specifically including: performing exponential weighted moving average processing on the confidence matrix and the feature matrix of the intra-frame semantic affinity graph and the confidence matrix and the feature matrix of the intra-frame category affinity graph after cross-graph convolution processing, to update each matrix.
[0015] Preferably, pseudo label cross-frame propagation is performed according to the predicted flow field to generate pseudo labels for the unlabeled point cloud frame, specifically including: calculating a pseudo-labeled predicted frame corresponding to the current point cloud frame according to the predicted flow field and the previous point cloud frame; for each super voxel in the current point cloud frame: searching for a plurality of neighbor points corresponding to the super voxel in the predicted frame; calculating the feature similarity between the super voxel and each neighbor point, and weighting and fusing the pseudo labels of each neighbor point with the corresponding feature similarity as the weight; the class with the highest confidence in the weighted and fused pseudo label is taken as the pseudo label of the super voxel.
[0016] Preferably, the point cloud sequence semantic segmentation model is trained, specifically including: training the point cloud sequence semantic segmentation model with the goal of minimizing the total loss; the total loss includes: supervised loss and unsupervised loss; the unsupervised loss includes: scene flow loss, semantic affinity graph loss and category affinity graph loss.
[0017] Preferably, the semantic affinity graph loss is:
[0018] ;
[0019] wherein, is the semantic affinity graph loss, is a weight factor; is a first parameter, if the labels of points i and j are the same, , if the labels of points i and j are different, ; 、 are the feature vectors of points i and j in the t-1 frame in the mth iteration, respectively, is the prototype feature of the corresponding category of point i in the t frame in the mth iteration, is a temperature coefficient, is the feature prototype corresponding to the category c in the t frame in the mth iteration, is the prototype memory bank.
[0020] The second aspect of the present application provides a point cloud sequence semantic segmentation method, comprising: inputting a to-be-processed point cloud sequence into the trained point cloud sequence semantic segmentation model as described above to obtain a semantic class label of each point in the to-be-processed point cloud sequence.
[0021] The third aspect of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method as described above when executing the program.
[0022] The fourth aspect of the present application provides a non-transitory computer readable storage medium storing computer instructions for causing a computer to execute the method as described above.
[0023] Compared with the prior art, the advantages of the present application include: a point cloud sequence semantic segmentation model training method is provided, the propagation and correction of pseudo labels in the point cloud sequence are realized by constructing an intra-frame semantic affinity graph and an intra-frame class affinity graph, combining a prediction flow field guiding mechanism between adjacent point cloud frames; at the same time, the cross-graph convolution mechanism and the self-distillation mechanism are combined, the confidence weighted fusion and class prototype comparison mechanism in the mechanism are used to effectively enhance the feature expression ability and class separability, and the problems of insufficient use of space-time information, low precision of pseudo labels, and error accumulation in the pseudo label propagation process in the prior art are solved, thereby significantly improving the semantic segmentation precision at a very low annotation cost, and the performance of the point cloud sequence semantic segmentation model trained by the semi-supervised method is close to that of the model trained by the full supervision method. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 The flowchart of the point cloud sequence semantic segmentation model training method provided by the embodiment of the present application.
[0025] Figure 2 The system framework diagram of the method provided by the embodiment of the present application. Figure 1
[0026] Figure 3 The schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0027] In view of the deficiencies in the prior art, the present application has been proposed after long-term research and a large number of practices. The technical solution, its implementation process and principles will be further explained as follows.
[0028] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description.
[0029] In addition, in the description of the present application, it needs to be understood that the terms "upper", "lower", "inner", "outer", "horizontal", "vertical" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0030] In the description of the present application, the description of the terms "one embodiment", "an embodiment", "the embodiment" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present description, the illustrative description of the above terms is not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.
[0031] The embodiment of the present application provides a training method of a point cloud sequence semantic segmentation model. Figure 1 The training method of the point cloud sequence semantic segmentation model comprises steps S100-S500.
[0032] Step S100, acquiring a point cloud sequence, the point cloud sequence comprising a plurality of point cloud frames, and some point cloud frames being unlabeled.
[0033] The point cloud sequence comprises a small amount of labeled point cloud frames and a large amount of unlabeled point cloud frames, and the label of the labeled point cloud frame is generated by artificial labeling, for example.
[0034] Step S200, performing super voxel segmentation on the point cloud frame, and constructing an intra-frame semantic affinity graph and an intra-frame class affinity graph according to the segmented super voxels.
[0035] For each point cloud frame in the point cloud sequence: performing super voxel segmentation on the point cloud frame, and the parameters related to each super voxel after segmentation are as follows:
[0036] ;
[0037] wherein, , , are the coordinates, features and class probability matrix of the super voxel j, is the number of points in the super voxel j, , , are respectively the spatial coordinates, eigenvectors, predicted posterior probability vectors belonging to each category of the midpoint i in the super voxel j; superscript represents the first iteration, i.e. the initial value; superscript represents the first frame point cloud frame.
[0038] For each point cloud frame in the point cloud sequence: construct an intra-frame semantic affinity graph and an intra-frame category affinity graph according to the segmented super voxels, so as to capture the intra-frame structural consistency and category consistency.
[0039] The semantic affinity graph is constructed based on the spatial distance and the feature distance between super voxels, the category affinity graph is constructed based on the consistency of the category probability vectors, and the similarity weight of the affinity graph is weighted using a Gaussian function.
[0040] Specifically, the spatial distance is obtained by calculating the Euclidean distance between super voxels, the feature distance is obtained by calculating the difference measure of the corresponding features between super voxels, and the category affinity graph is constructed by calculating the similarity of the category probability vectors of different super voxels in the same category. When calculating the similarity weight of the affinity graph, a Gaussian function is used to weight the distance similarity and the feature similarity, so as to ensure that the similarity of the graph has a higher weight between super voxels with a closer distance, thereby effectively promoting the information transmission between adjacent super voxels.
[0041] In a more preferred embodiment, the constructed intra-frame semantic affinity graph is:
[0042] ;
[0043] The constructed intra-frame category affinity graph is:
[0044] ;
[0045] wherein, is the element of the i-th row and the j-th column in the affinity matrix of the intra-frame semantic affinity graph, is the element of the i-th row and the j-th column in the affinity matrix of the category semantic affinity graph, is a first hyperparameter, is a second hyperparameter, is a first fixed parameter, is a second fixed parameter, is the Euclidean distance between super voxel i and super voxel j, is the feature difference between super voxel i and super voxel j, is the k-neighbor point set of super voxel i in the t frame, , respectively the class prediction probability vector of the super voxel i, j, , respectively the prediction probability of the super voxel i, j corresponding to the class b, the superscript denotes the th iteration, the superscript denotes the frame point cloud frame.
[0046] affinity matrix of the intra-frame semantic affinity graph the result after normalization is:
[0047] ;
[0048] affinity matrix of the intra-frame class affinity graph the result after normalization is:
[0049] ;
[0050] wherein, , are the corresponding diagonal matrices. each diagonal element in , each diagonal element in . is the value between the super voxels i and j corresponding to the affinity matrix of the semantic affinity graph in the t frame in the mth iteration.
[0051] Step S300, according to the cross-graph convolution mechanism and the self-distillation mechanism, respectively processing the intra-frame semantic affinity graph and the intra-frame class affinity graph to propagate and update the feature information of the super voxels in the intra-frame semantic affinity graph and the intra-frame class affinity graph.
[0052] The cross-graph convolution operation is respectively performed on the intra-frame semantic affinity graph and the intra-frame class affinity graph, the pseudo-label and the feature information are propagated based on the adjacency structure, and the pseudo-label confidence and the node feature representation are dynamically adjusted through the self-distillation strategy. This step helps to enhance the class consistency and aggregate the semantic space in the frame, thereby improving the pseudo-label quality and the feature distinguishing ability.
[0053] In a more preferred embodiment, the intra-frame semantic affinity graph and the intra-frame class affinity graph are respectively processed according to the cross-graph convolution mechanism, specifically including: alternately performing the following two kinds of graph convolution processing; performing graph convolution processing on the intra-frame semantic affinity graph guided by the affinity matrix of the intra-frame class affinity graph; performing graph convolution processing on the intra-frame class affinity graph guided by the affinity matrix of the intra-frame semantic affinity graph.
[0054] The cross-graph convolution operation can gradually update the feature information of the super voxel according to the similarity information in the graph structure, so that the semantic information and the category information of the super voxel can be propagated and enhanced in the graph. The cross-graph convolution of the two graphs in the mth iteration is as follows:
[0055] ;
[0056] ;
[0057] wherein, , are the category affinity graph node confidence matrix and the semantic affinity graph node feature matrix in the mth iteration, respectively, is an activation function, is a first parameter matrix, is a second parameter matrix. The confidence matrix and the feature matrix exchange information based on the normalized semantic affinity matrix and the normalized category affinity matrix .
[0058] In a more preferred embodiment, the intra-frame semantic affinity graph and the intra-frame category affinity graph are processed according to the self-distillation mechanism, which specifically includes: performing exponential weighted moving average processing on the confidence matrix and the feature matrix of the intra-frame semantic affinity graph after cross-graph convolution processing, and on the confidence matrix and the feature matrix of the intra-frame category affinity graph, to update each matrix.
[0059] The self-distillation update mechanism selectively transmits information between the semantic affinity graph and the category affinity graph through a dynamic gating mechanism, and adjusts the intensity of the information flow according to the confidence value, so as to ensure that the pseudo label with high confidence can preferentially affect the update of the label with low confidence.
[0060] Taking the confidence matrix of the intra-frame category affinity graph as an example, the update rule under the self-distillation mechanism is as follows:
[0061] ;
[0062] wherein, is the probability vector of the super voxel i after EMA update in the mth iteration t frame, is the probability vector of the super voxel i before update in the iteration t frame; is a function that returns 1 when the result in the parentheses is true, and returns 0 otherwise; is a third fixed parameter, is a dynamic threshold.The purpose of is to adapt to the category imbalance problem, which is expressed as follows:
[0063] ;
[0064] wherein, is a fourth fixed parameter, , respectively represent the learning effect of the current class c, b, is calculated as follows:
[0065] ;
[0066] wherein, is the class prediction vector of the initial super voxel i in the t frame. The dynamic threshold is to make the model pay more attention to difficult-to-classify samples and reduce overfitting to easy-to-classify samples (usually with too high confidence).
[0067] The updating rules of the confidence matrix of the intra-frame semantic affinity graph, the feature matrix, and the feature matrix of the intra-frame class affinity graph under the self-distillation mechanism are the same, and will not be repeated here.
[0068] In step S400, the super voxels between adjacent point cloud frames are matched to obtain a prediction flow field, and pseudo labels are cross-frame propagated according to the prediction flow field to generate pseudo labels for the unlabeled point cloud frame.
[0069] The scene flow network is used to estimate the motion of the super voxels between adjacent point cloud frames to obtain a prediction flow field of the inter-frame point cloud; the prediction flow field is used to guide the propagation of the pseudo labels of the previous frame to the current frame, and a feature similarity weighting mechanism is introduced into the propagation result to construct a weighted confidence fusion strategy. This mechanism can improve the accuracy and stability of the inter-frame pseudo labels and effectively alleviate the label drift and semantic mismatch problems.
[0070] In a more preferred embodiment, the pseudo labels are cross-frame propagated according to the prediction flow field to generate pseudo labels for the unlabeled point cloud frame, specifically including: calculating a pseudo-labeled prediction frame corresponding to the current point cloud frame according to the prediction flow field and the previous point cloud frame; for each super voxel in the current point cloud frame: searching for a plurality of neighbor points corresponding to the super voxel in the prediction frame; calculating the feature similarity between the super voxel and each neighbor point, and weighting and fusing the pseudo labels of each neighbor point with the corresponding feature similarity as the weight; and taking the class with the highest confidence in the weighted and fused pseudo labels as the pseudo label of the super voxel.
[0071] Specifically, the scene flow network is used to estimate the motion of the super voxels between adjacent point cloud frames to obtain a prediction flow field of the inter-frame point cloud. A prediction frame can be roughly calculated through the prediction flow field, which is represented as follows:
[0072] ;
[0073] wherein, is calculated by the previous point cloud frame corresponding to the current point cloud frame predicted frame, is a predicted flow field.
[0074] Then, the k nearest voxels in the predicted frame are searched using the K Nearest Neighbor (KNN) method, denoted as follows:
[0075] ;
[0076] wherein, is the t-th voxel in the current point cloud frame, and the k neighbors of the t-th voxel in the current point cloud frame are searched in the predicted frame . is the set of k nearest voxels of the i-th voxel in the t-th frame in the predicted frame .
[0077] Then, the feature similarity between the i-th voxel and its k neighbors is calculated:
[0078] ;
[0079] ;
[0080] Then, the pseudo-labels of the i-th voxel are fused by weighting according to the similarity weights between the i-th voxel and its k neighbors, so as to achieve more refined pseudo-label refinement:
[0081] ;
[0082] wherein, is a hyperparameter, is the feature similarity between the i-th voxel and its neighbor j.
[0083] Step S500, training the point cloud sequence semantic segmentation model using the labeled point cloud frames and the pseudo-labeled point cloud frames in the point cloud sequence.
[0084] In a more preferred embodiment, the point cloud sequence semantic segmentation model is trained, specifically comprising: training the point cloud sequence semantic segmentation model with the objective of minimizing the total loss; the total loss comprises: supervised loss and unsupervised loss; the unsupervised loss comprises: scene flow loss, semantic affinity graph loss and class affinity graph loss.
[0085] A joint loss function is constructed by a supervised loss and an unsupervised loss, wherein the supervised loss is used to guide the learning of labeled data, and the unsupervised part includes a scene flow loss, a semantic affinity graph loss and a class affinity graph loss, which are used to improve the inter-frame matching accuracy and the pseudo label quality, and optimize the weakly supervised learning effect of the model as a whole. The point cloud sequence semantic segmentation model is, for example, a model constructed based on a sparse convolutional network and having intra-frame graph modeling and cross-frame guiding capabilities.
[0086] The labeled data (i.e. labeled point cloud frames) are input into the point cloud sequence semantic segmentation model to obtain a labeled segmentation result, and a supervised loss of the model is calculated based on the labeled segmentation result and a corresponding true value label. The calculation formula of the supervised loss is as follows:
[0087] ;
[0088] Wherein, is the number of labeled points, is the true label of the i-th point, is the prediction vector of the model.
[0089] For the semantic affinity graph loss, intra-frame contrast learning is performed on the previous frame and inter-frame contrast learning is performed on the current frame. The inter-frame loss adopts a prototype-based contrast loss. For this purpose, a prototype memory bank is constructed, and a feature prototype belonging to a class c (generically referred to as a class) is constructed based on feature averaging, and is expressed as follows:
[0090] ;
[0091] The prototype bank can be dynamically updated by exponential moving average of historical features, and is expressed as follows:
[0092] ;
[0093] Wherein, the left side is the updated parameter, the right side is the parameter before updating, is a fourth fixed parameter, is the pseudo label of the i-th point in the t-1 frame in the m-th iteration.
[0094] The semantic affinity graph loss introduces an inter-frame and intra-frame contrast-based contrast learning method to enhance the consistency between features of the same class and the separability between features of different classes. Intra-frame learning can be achieved by class prototypes.
[0095] In a more preferred embodiment, the semantic affinity graph loss is as follows:
[0096] ;
[0097] wherein, is the semantic affinity graph loss, is a weight factor; E is the number of point pair sets; is a first parameter, if the labels of super voxel i and super voxel j are the same, , if the labels of point i and point j are different, ; , are the feature vectors of point i and point j in the t-1 frame in the mth iteration, respectively, is the prototype feature of the corresponding category of point i in the t frame in the mth iteration, is a temperature coefficient, is the feature prototype corresponding to category c in the t frame in the mth iteration, is the prototype memory bank.
[0098] Category affinity graph loss Then, based on the super voxels with high confidence pseudo labels in the current frame, the influence of the noise labels is further reduced through a dynamic weighted cross-entropy loss, as follows:
[0099] ;
[0100] wherein, is the pseudo label obtained through the aforementioned confidence fusion mechanism, is the prediction probability of super voxel i in the t frame in the mth iteration in category .
[0101] Total loss The calculation formula of the total loss is as follows:
[0102] ;
[0103] wherein, is the scene flow loss, , , respectively represent the corresponding weight coefficients.
[0104] The point cloud sequence semantic segmentation model training method provided by the application performs super voxel segmentation on continuous point cloud frames, constructs a semantic affinity graph and a class affinity graph, optimizes semantic features by using cross-graph convolution and a self-distillation mechanism, enhances intra-frame class consistency and structure representation capability, designs a pseudo-label propagation strategy based on scene flow, realizes inter-frame super voxel matching in combination with a predicted flow field, further proposes a weighted confidence fusion mechanism to adaptively balance pseudo-label information of a current frame and a previous frame, and improves label confidence and stability. In addition, prototype contrast learning and graph contrast loss are used to constrain feature clustering effect, and model parameters are iteratively updated in combination with a supervised loss and a pseudo-supervised loss. Without relying on large-scale manual labeling, the method effectively alleviates error accumulation in the label propagation process, enhances the perception ability of the model to spatiotemporal consistency, and outperforms existing weakly supervised methods in segmentation accuracy on multiple point cloud sequence datasets, close to the full supervision level.
[0105] The technical solutions of the application will be further described in detail below in combination with several preferred embodiments and the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the application. The test method is not specified in the following embodiments, and is usually performed according to conventional conditions.
[0106] In some embodiments, the iterations in the point cloud sequence semantic segmentation model training method can be performed in batches, that is, the training set is divided into multiple batches, and each batch corresponds to one iteration. The number of data in a batch can be several greater than 1, for example, 2-128, etc., or can be only one, which is equivalent to not being divided into batches.
[0107] The specific execution steps of the point cloud sequence semantic segmentation model training method are as follows, and the principle block diagram of the implementation is shown in Figure 2
[0108] Step 1: Create and independently initialize a basic model. The training set is composed of a small amount of labeled data and a large amount of unlabeled data, the labeled data and the unlabeled data are divided into batches of a certain size, and are input into the model for training in turn.
[0109] Step 2: Group the point cloud sequence according to every m, and only label the first frame of each group. The point cloud data in each group is continuous, and the divided data is input into the model, which can be expressed as:
[0110] ;
[0111] Wherein, Xi,t represents the i-th point of the t-th frame in the point cloud data training set, denote the corresponding predicted probability distribution, denote the corresponding model predicted output features, denote the network parameters of the model.
[0112] Since the initial annotation is too little to train a better model, the embodiment uses super voxel segmentation for preliminary segmentation. In the same super voxel: if there is only one labeled point, the remaining points in the same super voxel are labeled with the same label; if it contains multiple labeled points, a label is assigned through majority voting, expanding the initial annotation, and the super voxel without any annotation remains unlabeled.
[0113] Step 3: Construct semantic affinity graph and class affinity graph based on the semantic relationship and class relationship between super voxels, and perform cross-graph convolution operation on the semantic affinity graph and the class affinity graph respectively.
[0114] After that, the self-distillation update mechanism selectively transmits information between the semantic affinity graph and the class affinity graph through a dynamic gating mechanism, and adjusts the intensity of information flow according to the confidence value, ensuring that high-confidence pseudo labels can preferentially affect the update of low-confidence labels. The updated feature matrix and confidence matrix are used to update the edge weights of the semantic affinity graph and the class affinity graph respectively.
[0115] Step 4: The semantic affinity graph is used for scene flow estimation.
[0116] The pre-trained or jointly trained scene flow network is used to match the super voxels in the two consecutive frames of point cloud frames, and the predicted position of each super voxel in the next frame is obtained, that is, a flow field of the point cloud scene is constructed. The flow field describes the temporal correspondence at the super voxel level and provides a basis for cross-frame information propagation. The flow field is used to guide the propagation of pseudo labels from the previous frame to the current frame, and a feature similarity weighting mechanism is introduced in the propagation result to construct a weighted confidence fusion strategy.
[0117] After that, the pseudo label of the i-th super voxel is refined according to the similarity of its neighbor points. After updating the pseudo label, the super voxels with high confidence in the predicted probability vector are selected, and the class with the highest prediction probability in the confidence is taken as the pseudo label of the super voxel.
[0118] Step 5: A joint loss function composed of supervised loss and unsupervised loss is constructed. The supervised loss is used to guide the effective learning of a small amount of labeled data; the unsupervised loss is used to enhance the consistency of inter-frame point cloud matching, improve the aggregation of semantic features in unlabeled data, and optimize the consistency and discriminability of class prediction. The scene flow loss includes chamfer loss, smooth loss, and consistency loss.
[0119] Step 6: Combine the supervised loss and the unsupervised loss, and update the parameters of the base model by gradient descent optimization, repeat steps 2-5 until the base model converges, thereby obtaining a segmentation model that can be used for point cloud semantic prediction.
[0120] Further, the reliability of the point cloud sequence semantic segmentation model trained by the training method is verified by using known network data, and the training method is compared with existing weakly supervised and fully supervised point cloud segmentation training methods, as follows.
[0121] In a first aspect, the training method provided by the present application is compared with existing weakly supervised and fully supervised point cloud segmentation training methods, and is verified on a first data set.
[0122] The first point cloud sequence data set is designed for semantic segmentation tasks in urban driving environments, and covers multiple urban road scenes with rich scene changes, including different weather conditions, time of day, and urban structures, making it suitable for evaluating point cloud data processing methods in autonomous driving technology. The training set contains 9 sequences, a total of 19130 frames; the validation set contains 1 sequence, a total of 4071 frames; and the test set contains all 9 sequences, a total of 20351 frames.
[0123] In this embodiment, the backbone network uses the Minkowski network. First, pre-train the backbone model and the scene flow model with sparse initial annotations. Training is performed using the adaptive moment estimation optimizer, with a learning rate of 0.01, a batch size of 4, and loss weight coefficients set to =0.1, =0.2, =0.7. The comparison results are shown in Table 1.
[0124] Table 1
[0125] ;
[0126] The values in the third column of Table 1 are the average intersection over union (%) for all classes, representing the average segmentation accuracy of the model. As can be seen from Table 1, the training method provided in this embodiment effectively improves the accuracy of the weakly supervised point cloud segmentation model, making it close to or even reaching the performance of the fully supervised training method.
[0127] In a second aspect, the training method provided by the present application is compared with existing weakly supervised and fully supervised point cloud segmentation training methods, and is verified on a second data set.
[0128] The second dataset is a point cloud dataset aimed at capturing the pedestrian viewpoint of a city, consisting of 6 different scenes, with a total of 2988 frames. The dataset is divided into a training set containing 2488 frames and a test set containing 500 frames. Unlike the first dataset, which focuses on high-speed environments, the second dataset focuses on low-speed urban environments, such as campuses and pedestrian streets. It includes various urban features, including buildings, green spaces, streetlights, and pedestrians, as well as slowly moving objects such as pedestrians, electric scooters, and cyclists.
[0129] In this embodiment, the backbone network adopts the Minkowski network. First, pre-train the backbone model and the scene flow model with sparse initial annotations. Training is performed using the adaptive moment estimation optimizer, with a learning rate of 0.01, a batch size of 4, and loss weight coefficients set to =0.1, =0.2, =0.7. The comparison results are shown in Table 2.
[0130] Table 2
[0131] ;
[0132] The values in the third column of Table 2 represent the average segmentation accuracy of the 6 regions. As can be seen from Table 2, the training method provided in this embodiment is also effective in improving the performance of weakly supervised point cloud segmentation models in large-scale outdoor scenes.
[0133] Based on the same inventive concept, corresponding to any of the above embodiment methods, the present application also provides a point cloud sequence semantic segmentation method, the method comprising: inputting a to-be-processed point cloud sequence into the point cloud sequence semantic segmentation model trained by the training method of the point cloud sequence semantic segmentation model in the above embodiment, to obtain the semantic class label of each point in the to-be-processed point cloud sequence.
[0134] Based on the same inventive concept, corresponding to any of the above embodiment methods, the present application also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the search model establishment method based on the heterogeneous graph neural network and / or the search method based on the heterogeneous graph neural network of any one of the embodiments.
[0135] Figure 3 A more specific hardware structure diagram of an electronic device provided by the present embodiment is shown, which can include a processor 310, a memory 320, an input / output interface 330, a communication interface 340, and a bus 350. The processor 310, the memory 320, the input / output interface 330, and the communication interface 340 are connected to each other through the bus 350 for communication within the device.
[0136] The processor 310 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided by the embodiments of the present specification.
[0137] The memory 320 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 320 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, relevant program codes are stored in the memory 320 and called and executed by the processor 310.
[0138] The input / output interface 330 is configured to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0139] The communication interface 340 is configured to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0140] The bus 350 includes a channel for transmitting information between various components (such as the processor 310, the memory 320, the input / output interface 330, and the communication interface 340) of the device.
[0141] It should be noted that although the above device only shows the processor 310, the memory 320, the input / output interface 330, the communication interface 340, and the bus 350, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include components necessary for implementing the solutions of the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0142] The electronic device of the above-mentioned embodiments is used to implement the training method of the point cloud sequence semantic segmentation model and / or the point cloud sequence semantic segmentation method of the corresponding any one of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here again.
[0143] Based on the same inventive concept, the present application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the training method of the point cloud sequence semantic segmentation model and / or the point cloud sequence semantic segmentation method of any one of the above-mentioned embodiments.
[0144] The computer-readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0145] The computer instructions stored in the storage medium of the above-mentioned embodiments are used to cause the computer to perform the training method of the point cloud sequence semantic segmentation model and / or the point cloud sequence semantic segmentation method of any one of the above-mentioned embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here again.
[0146] Those skilled in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope (including claims) of the present application is limited to these examples; under the idea of the present application, the above embodiments or technical features between different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the present application as described above. In order to be brief, they are not provided in detail.
[0147] Additionally, to simplify the description and discussion, and so as not to obscure the embodiments of the application with details that are well known to those skilled in the art, some conventional attributes of integrated circuit (IC) chips and other components can or can not be shown or described in detail. Moreover, devices can be shown in block diagram form in order to avoid obscuring the embodiments of the application, and this also acknowledges the fact that details of implementation of such block devices are highly dependent on the platform within which the embodiments of the application are to be implemented (i.e., these details should be apparent to those skilled in the art). Where specific details of such implementation are set forth in order to describe an illustrative embodiment of the application, it will be apparent to one skilled in the art that the embodiment of the application can be practiced without, or with variation of, these specific details. Thus, the description is to be considered as illustrative only and not restrictive of the application.
[0148] Although the application has been described in conjunction with specific embodiments thereof, numerous alternatives, modifications, and variations will be readily apparent to those of ordinary skill in the art. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0149] It is therefore intended that the embodiments of the application be covered by all such alternatives, modifications and variations that fall within the broad scope of the appended claims. Accordingly, any and all such alternatives, modifications and variations should be included within the scope of the present application.
Claims
1. A method for training a point cloud sequence semantic segmentation model, characterized in that, The method comprises the following steps: acquiring a point cloud sequence comprising a plurality of point cloud frames, some of which are unlabeled; performing super voxel segmentation on the point cloud frames, and constructing an intra-frame semantic affinity graph and an intra-frame category affinity graph according to the segmented super voxels; processing the intra-frame semantic affinity graph and the intra-frame category affinity graph according to a cross-graph convolution mechanism, and then processing the intra-frame semantic affinity graph and the intra-frame category affinity graph according to a self-distillation mechanism to propagate and update the feature information of the super voxels in the intra-frame semantic affinity graph and the intra-frame category affinity graph; matching the super voxels between adjacent point cloud frames to obtain a predicted flow field, and performing pseudo-label cross-frame propagation according to the predicted flow field to generate pseudo-labels for the unlabeled point cloud frames; training a point cloud sequence semantic segmentation model using the labeled point cloud frames and the pseudo-labeled point cloud frames in the point cloud sequence; wherein the processing of the intra-frame semantic affinity graph and the intra-frame category affinity graph according to the self-distillation mechanism comprises: performing exponential weighted moving average processing on the confidence matrix and the feature matrix of the intra-frame semantic affinity graph and the confidence matrix and the feature matrix of the intra-frame category affinity graph after cross-graph convolution processing to update the matrices.
2. The method of claim 1, wherein the method further comprises: The constructed intra-frame semantic affinity graph is: ; The constructed intra-frame category affinity graph is: ; in, Let be the element in the i-th row and j-th column of the affinity matrix of the intra-frame semantic affinity graph. Let be the element in the i-th row and j-th column of the affinity matrix of the category semantic affinity graph. As the first hyperparameter, This is the second hyperparameter. As the first fixed parameter, As the second fixed parameter, Let be the Euclidean distance between supervoxel i and supervoxel j. The feature difference between supervoxel i and supervoxel j. Let k be the set of the neighbors of supervox i in frame t. , Let be the class prediction probability vectors of supervoxel i and supervoxel j, respectively. , Let i and j be the predicted probabilities of supervoxels i and j corresponding to class b, respectively, with superscripts indicating their respective probabilities. Indicates the first Next iteration, superscript Indicates the first Frame-by-frame cloud.
3. The method of claim 1, wherein the method further comprises: The processing of the intra-frame semantic affinity graph and the intra-frame category affinity graph according to the cross-graph convolution mechanism comprises: alternately performing the following two kinds of graph convolution processing; performing graph convolution processing on the intra-frame semantic affinity graph guided by the affinity matrix of the intra-frame category affinity graph; performing graph convolution processing on the intra-frame category affinity graph guided by the affinity matrix of the intra-frame semantic affinity graph.
4. The method of claim 1, wherein, The pseudo-label cross-frame propagation according to the predicted flow field to generate pseudo-labels for the unlabeled point cloud frames comprises: calculating a pseudo-labeled predicted frame corresponding to the current point cloud frame according to the predicted flow field and the previous point cloud frame; for each super voxel in the current point cloud frame: searching for a plurality of neighbor points corresponding to the super voxel in the predicted frame; calculating the feature similarity between the super voxel and each neighbor point; and performing weighted fusion on the pseudo-labels of each neighbor point with the corresponding feature similarity as the weight; and taking the class with the highest confidence in the weighted fused pseudo-labels as the pseudo-label of the super voxel.
5. The method of claim 1, wherein, The training of the point cloud sequence semantic segmentation model comprises: training the point cloud sequence semantic segmentation model to minimize the total loss. The total loss comprises: a supervised loss and an unsupervised loss. The unsupervised loss comprises: a scene flow loss, a semantic affinity graph loss, and a category affinity graph loss.
6. The method of claim 5, wherein the method further comprises: The semantic affinity graph loss is: ; wherein, is the semantic affinity graph loss, is a weight factor; is a first parameter, if the labels of the super-voxel i and the super-voxel j are the same, , if the labels of the point i and the point j are different, ; , is the feature vector of the point i in the t-1 frame of the mth iteration, is the prototype feature of the corresponding class of the point i in the t frame of the mth iteration, is a temperature coefficient, is the feature prototype corresponding to the class c in the t frame of the mth iteration, is the prototype memory bank.
7. A method for semantic segmentation of a sequence of point clouds, the method comprising: The method comprises the following steps: inputting a to-be-processed point cloud sequence into the point cloud sequence semantic segmentation model trained by the method of any one of claims 1-6 to obtain the semantic category labels of each point in the to-be-processed point cloud sequence.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the method of any one of claims 1-7.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to make the computer execute the method of any one of claims 1-7.
Citation Information
Patent Citations
Semantic segmentation model training method, semantic segmentation method and related device
CN117974996A
Semi-supervised semantic segmentation model training method based on multistage label correction and application
CN118379593A