A Small Sample Point Cloud Semantic Segmentation Method and System Based on Multi-View Text Supervision

By employing a few-sample point cloud semantic segmentation method guided by multi-view text supervision, and combining word embeddings and a large language model, the method enhances the support set feature representation capability, solves the problems of high data annotation cost and insufficient generalization ability in few-sample point cloud segmentation, and achieves efficient semantic segmentation results.

CN120726330BActive Publication Date: 2025-12-02RENMIN ZHONGKE (JINAN) INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511149365.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-02
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing point cloud semantic segmentation methods rely on large-scale manual annotation, which is costly and lacks the ability to generalize to unseen categories in open set scenarios. Traditional text supervision is also unable to fully characterize the complexity of category semantics.

Method used

A few-sample point cloud semantic segmentation method supervised by multi-view text is proposed. This method constructs a point cloud semantic segmentation model based on deep learning, combines word embedding model and large language model to generate multi-view text descriptions, uses point cloud self-attention and cross-attention mechanisms to enhance support set features, and utilizes category-specific cross-attention modules to aggregate query set features for few-sample point cloud semantic segmentation.

Benefits of technology

It significantly improves the model's ability to understand rare categories and complex scenes and its segmentation accuracy, reduces data acquisition and annotation costs, enables rapid knowledge transfer to process new categories, and provides a more universal and efficient small-sample point cloud semantic segmentation solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726330B_ABST
    Figure CN120726330B_ABST
Patent Text Reader

Abstract

This invention relates to the field of point cloud data processing technology, specifically providing a few-sample point cloud semantic segmentation method and system based on multi-view text supervision. The method includes: acquiring and preprocessing a point cloud semantic segmentation dataset, dividing it into disjoint training and testing category subsets, and constructing corresponding support and query sets; constructing a deep learning-based point cloud semantic segmentation model as the backbone network and pre-training it; constructing a few-sample point cloud semantic segmentation model based on multi-view text supervision and training it using the support and query sets of the training categories; and outputting segmentation results using a test dataset and evaluating its performance. This invention fully leverages the capabilities of large language models in generating diverse semantics. By generating multi-view text descriptions for the same category, it effectively improves the model's semantic understanding and segmentation capabilities for unknown categories and complex scenes, thereby significantly improving point cloud semantic segmentation performance under few-sample conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of point cloud data processing technology, specifically relating to a small sample point cloud semantic segmentation method and system based on multi-view text supervision guidance. Background Technology

[0002] The goal of point cloud semantic segmentation is to determine a semantic category for each point in a 3D point cloud. Point cloud semantic segmentation is a challenging task in the field of computer vision, providing crucial support for 3D scene understanding in intelligent and automated applications such as environmental perception for autonomous vehicles, navigation and obstacle avoidance for intelligent robots, and enhancement of realism in virtual reality environments.

[0003] With the rapid development of deep learning technology and the accumulation of large-scale 3D scene datasets, fully supervised point cloud semantic segmentation methods based on deep neural networks have achieved significant performance breakthroughs. However, these methods have two inherent drawbacks: first, they heavily rely on large-scale manual annotation, resulting in high data collection and annotation costs and cumbersome processes; second, they lack the ability to generalize segmentation to unseen categories in open-set scenarios. In recent years, few-shot point cloud semantic segmentation technology has emerged. This paradigm effectively solves the limitations of traditional methods by introducing a small number of labeled samples (support set) and samples to be segmented (query set), significantly reducing annotation requirements while enabling the model to generalize semantic segmentation to new categories. Most existing few-shot segmentation methods focus on feature mining of the support set point cloud, and complete the category assignment and semantic segmentation of the query set by measuring the similarity between the features of the query point cloud and the support set.

[0004] In the field of semantic segmentation of small point clouds, due to the scarcity of support set samples, the semantic representation is not only incomplete—only covering local areas of the target category—but also biased—it is difficult to balance the differences in the shape and scale of objects within the class, which makes the model prone to information loss and cognitive bias during feature learning.

[0005] While existing technologies have made preliminary explorations into text-supervised point cloud segmentation, they are typically based on simple textual information (such as Word2vec word embedding techniques), and single-dimensional textual descriptions are insufficient to fully characterize the complexity of category semantics. Therefore, this invention provides a few-sample point cloud semantic segmentation method and system based on multi-view text supervision guidance. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing methods and provide a few-sample point cloud semantic segmentation method and system based on multi-view text supervision guidance.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0008] The first objective of this invention is to provide a few-sample point cloud semantic segmentation method based on multi-view text supervision guidance, comprising:

[0009] (1) Obtain the point cloud semantic segmentation dataset and preprocess it. Divide the point cloud semantic segmentation dataset into a training set and a test set. Establish a small sample support set and a query set for the training set and the test set respectively.

[0010] (2) Construct a point cloud semantic segmentation model based on deep learning, use the training set to iteratively train the point cloud semantic segmentation model, and use the cross-entropy loss function to optimize it, so as to obtain a pre-trained backbone segmentation model with basic semantic segmentation capabilities.

[0011] (3) Based on the pre-trained backbone segmentation model, a small-sample point cloud semantic segmentation model with multi-view text supervision is constructed, and fine-tuned using the support set and query set in the training set. Specifically, this includes:

[0012] Point cloud features of the support set and query set are extracted using a pre-trained backbone segmentation model. The support set features are then enhanced using a point cloud self-attention mechanism. Finally, the support set point cloud features for each category are obtained by multiplying the support set features and the support set binary mask.

[0013] By combining a word embedding model (Word2Vec) and a large language model (LLM) with 3D command prompts, multi-view text description information is generated for each semantic category, and the corresponding multi-view text features are extracted.

[0014] The enhanced support set representation is obtained by fusing text features and support set point cloud features through a text-guided support set enhancement mechanism; the enhancement mechanism includes a text self-attention mechanism and a cross-attention mechanism to improve semantic alignment capability.

[0015] By using a category-specific cross-attention module, the query set point cloud features are aggregated with the enhanced support set features to generate a category-guided query set feature representation.

[0016] The query set features are input into a classifier for segmentation prediction. The classifier includes a convolutional layer and a softmax layer, which are used to output the point-by-point prediction probability of the query set.

[0017] The prediction error is calculated based on the cross-entropy loss function, and the Adam optimizer is used to optimize and update the overall small sample point cloud semantic segmentation model.

[0018] (4) The trained model is tested using the support set and query set in the test set, and the model performance is evaluated by calculating the mIoU index based on the prediction results.

[0019] Furthermore, in step (1), the point cloud semantic segmentation dataset is divided into known categories and unknown categories according to semantic categories. The training set includes known categories, and the test set includes unknown categories.

[0020] Furthermore, the deep learning-based point cloud semantic segmentation model includes a feature extractor composed of a dynamic graph CNN architecture, an attention learner composed of a self-attention network, and a metric learner composed of multiple perceptron layers, which are respectively used to extract local geometric features, aggregate global context features, and map features into the manifold space.

[0021] Furthermore, in step (4), the testing of the trained model using the support set and query set in the test set specifically includes:

[0022] Extract point cloud features from the support set and query set of the test set to obtain category-specific support set point cloud features;

[0023] For each new category, a large language model combining word embedding and 3D command prompts is used to generate multi-view text features;

[0024] The semantic expressive power of the support set is enhanced by using a text-guided support set enhancement mechanism;

[0025] Use a category-specific cross-attention mechanism to obtain feature representations of a specific class in the query set;

[0026] The query set segmentation result is obtained using convolutional layers and softmax layers.

[0027] Another objective of this invention is to provide a few-sample point cloud semantic segmentation system based on multi-view text supervision guidance, comprising:

[0028] The data acquisition and preprocessing module is used to acquire and preprocess the point cloud semantic segmentation dataset, divide the point cloud semantic segmentation dataset into training set and test set, and generate small sample support set and query set for the training set and test set respectively.

[0029] The pre-training module is used to train the deep learning-based point cloud semantic segmentation model and optimize it using the cross-entropy loss function.

[0030] The model building and training module is used to construct a few-sample point cloud semantic segmentation model based on multi-view text supervision. The model is fine-tuned using the support set and query set of the training set to obtain the trained few-sample point cloud semantic segmentation model. Specifically, a pre-trained backbone segmentation model is used to extract point cloud features from the support set and query set, and the support set features are enhanced through a point cloud self-attention mechanism. Finally, the support set point cloud features for each category are obtained by multiplying the support set features and the support set binary mask. Combining a word embedding model and a large language model using 3D command prompts, multi-view text description information is generated for each semantic category, and the corresponding multi-view text features are extracted. The model is then guided by text... A support set enhancement mechanism fuses text features with support set point cloud features to obtain an enhanced support set representation. This enhancement mechanism includes a text self-attention mechanism and a cross-attention mechanism to improve semantic alignment. A category-specific cross-attention module is used to aggregate query set point cloud features with the enhanced support set features, generating a category-guided query set feature representation. The query set features are then input into a classifier for segmentation prediction. This classifier includes convolutional layers and a softmax layer to output the point-by-point prediction probability of the query set. The prediction error is calculated based on the cross-entropy loss function, and the Adam optimizer is used to optimize and update the overall small-sample point cloud semantic segmentation model.

[0031] The testing and evaluation module is used to test the small sample point cloud semantic segmentation model using the support set and query set of the test set, and evaluate the model performance using the mIoU metric.

[0032] Furthermore, the model building and training module includes:

[0033] A multi-view text supervision generation unit is used to combine word embedding models and large language models to generate multi-view text features for semantic categories.

[0034] The point cloud feature enhancement unit is used to optimize the support set and query set features through the point cloud self-attention mechanism and generate category-specific support set features;

[0035] Text-guided support set enhancement units are used to fuse text features and support set features using text self-attention and cross-attention enhancement mechanisms.

[0036] Category-specific attention units are used to aggregate query set features and enhanced support set features through a cross-attention mechanism;

[0037] The segmented output subunit is used to output the query set segmentation results through convolutional layers and softmax layers.

[0038] Another object of the present invention is to provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the computer program to implement the small sample point cloud semantic segmentation method based on multi-view text supervision guidance provided by the first object of the present invention.

[0039] Another object of the present invention is to provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the few-sample point cloud semantic segmentation method based on multi-view text supervision guidance provided by the first object of the present invention.

[0040] Another object of the present invention is to provide a server comprising at least one processor and a memory communicatively connected to the processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the processor to cause the at least one processor to perform the small sample point cloud semantic segmentation method based on multi-view text supervision guidance provided in the first object of the present invention.

[0041] In combination with the above technical solutions, the beneficial effects of the present invention compared with the prior art are as follows:

[0042] This invention introduces a large language model to construct a multi-view text supervision system. Leveraging the multi-dimensional semantic generation capabilities of LLM, it can output diverse text descriptions of the same category from perspectives such as appearance description, word choice and sentence construction reflecting characteristics, and comprehensive description. This mechanism not only breaks through the semantic expression bottleneck of traditional point cloud data but also provides structured semantic knowledge information for semantic segmentation in small-sample scenarios through the complementarity of multi-view texts. This significantly improves the model's understanding of rare categories and complex scenes, as well as its segmentation accuracy, achieving a paradigm shift from geometric feature-driven to semantic-geometric joint-driven approaches.

[0043] This invention trains a few-shot point cloud segmentation model guided by multi-view text supervision using only a small number of labeled point clouds. This model generates more comprehensive and richer multi-view text supervision information for semantic categories to enhance the support set point cloud, thereby assisting in the segmentation of the query set. Furthermore, for new categories not encountered during the training phase, this invention can significantly improve the performance of few-shot point cloud segmentation with only a small number of labeled point clouds.

[0044] Building upon traditional few-sample point cloud segmentation methods that rely on a small number of labeled support clouds to reduce data acquisition and labeling costs, this invention introduces a large language model and the Word2vec word embedding technique to construct a multi-view text supervision system. This system generates text descriptions covering appearance, characteristics, and other dimensions for semantic categories, forming a structured semantic set. By deeply complementing textual information with point cloud geometric features, the semantic representation capability of the support set is greatly enhanced, addressing the semantic gaps and biases inherent in traditional point cloud data and improving query set segmentation accuracy. When faced with new categories not encountered during training, this invention, leveraging the generalized semantic framework constructed through multi-view text supervision, requires only a small amount of labeled point cloud data to quickly transfer knowledge, significantly improving the model's segmentation performance for new categories and providing a more universal and efficient solution for few-sample point cloud semantic segmentation. Attached Figure Description

[0045] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0046] Figure 1 This is a flowchart of a few-sample point cloud semantic segmentation method based on multi-view text supervision provided in an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the small-sample point cloud semantic segmentation method based on multi-view text supervision provided in this embodiment of the invention;

[0048] Figure 3 This is a schematic diagram of the architecture principle of the model training stage provided in the embodiment of the present invention. Detailed Implementation

[0049] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0050] Example 1:

[0051] like Figure 1 The image shows an embodiment of the few-sample point cloud semantic segmentation method based on multi-view text supervision provided by this invention, which specifically includes the following steps:

[0052] S1: Obtain the point cloud semantic segmentation dataset and preprocess it. Divide the point cloud semantic segmentation dataset into a training set and a test set. Establish a small sample support set and a query set for the training set and the test set, respectively.

[0053] S2: Construct a point cloud semantic segmentation model based on deep learning, use the training set to iteratively train the point cloud semantic segmentation model, and use the cross-entropy loss function to optimize it, so as to obtain a pre-trained backbone segmentation model with basic semantic segmentation capabilities.

[0054] S3: Based on the pre-trained backbone segmentation model, construct a few-sample point cloud semantic segmentation model guided by multi-view text supervision, and fine-tune it using the support set and query set in the training set.

[0055] S4: Test the trained model using the support set and query set in the test set, and evaluate the model performance based on the mIoU index calculated based on the prediction results.

[0056] Specifically, such as Figure 2 As shown, this invention specifically includes a model training phase and a model testing phase. First, point cloud data is acquired and preprocessed, and training and testing sets are created. In the model training phase, a point cloud semantic segmentation network is pre-trained using the training set, and a few-sample point cloud semantic segmentation model is fine-tuned. Specifically, after extracting point cloud features from the support set and training set using the point cloud semantic segmentation network, this invention jointly uses a large language model (LLM) and word embedding (Word2vec) to generate multi-view text supervision. A text-guided support set enhancement module fuses support set point cloud features and text features, and a category-specific cross-attention module measures the relationship between the query set and the support set. Finally, a classifier determines the segmentation result. During the training phase, the entire model is optimized using a cross-entropy loss function. In the model testing phase, the proposed few-sample point cloud semantic segmentation model is used to segment new categories in the test set. The steps include: generating multi-view text supervision for the support set; fusing text and support set point cloud features using a text-guided support set enhancement module; and finally, using a category-specific cross-attention module and a classifier to obtain the final segmentation result for performance evaluation.

[0057] Specifically, step S1 in this embodiment of the invention includes the following steps:

[0058] S1.1, Obtain the point cloud semantic segmentation dataset: including two indoor point cloud segmentation datasets, S3DIS and ScanNet. S3DIS covers 272 rooms in 6 indoor environments and is labeled with 13 semantic categories, while ScanNet contains 1513 point clouds from 707 indoor scenes and is labeled with 20 semantic categories.

[0059] S1.2, Data Preprocessing: Considering the large number of points in the original room, this method uses a 1m×1m non-overlapping sliding window to divide the room into blocks on the xy plane, resulting in 7,547 and 36,350 blocks for S3DIS and ScanNet, respectively. During training and testing, 2048 points are randomly selected from each block, and each point is represented by a 9D vector, including XYZ coordinates, RGB values, and normalized 3D spatial coordinates.

[0060] S1.3, Dividing the Training and Test Sets: In addition to a fixed background category, this method randomly divides the semantic class into two non-overlapping subsets, S0 and S1. This method employs cross-validation, using S0 as the known category to construct the training set and S1 as the new category to construct the test set, and vice versa. Although the same point cloud may appear in both S0 and S1, the categories of interest labeled for that point cloud will differ.

[0061] S1.4, Constructing a small sample support set and query set: First, construct the support set and query set for the training set, that is, randomly select combinations of C categories and randomly select labeled point clouds and point clouds to be segmented for each selected category to form the support set S and query set Q. Then, construct the support set and query set for the test set using a similar method. The only difference is that all combinations of C target classes in the new category need to be exhaustively enumerated, instead of randomly selecting C classes.

[0062] Specifically, step S2 in this embodiment of the invention includes the following steps:

[0063] S2.1, Constructing a Point Cloud Semantic Segmentation Network: The deep learning-based point cloud segmentation network constructed in this method comprises three main modules: a feature extractor composed of a Dynamic Graph CNN architecture (DGCNN), an attention learner composed of a Self-Attention Network (SAN), and a metric learner composed of multilayer perceptron (MLP) layers, which are respectively used to extract local geometric features, aggregate global context features, and map features to the manifold space. The output features of the above modules are concatenated together as the output features of the entire segmentation network;

[0064] S2.2, Pre-training the point cloud semantic segmentation network: Using the training set, train the point cloud semantic segmentation network described in S2.1, wherein the batch size is 32 and the training is iterated for 100 rounds;

[0065] S2.3, Optimize the objective function: Use the cross-entropy loss function and the Adam optimizer with a learning rate of 0.001 to optimize the network.

[0066] Specifically, step S3 in this embodiment of the invention includes the following steps:

[0067] S3.1, Extracting point cloud features: Given a support set of point clouds and its GT mask Query Jidian Cloud and its GT mask ,in , and These represent the number of categories, the number of points contained in the support point cloud, and the number of points contained in the query point cloud, respectively. This invention uses the point cloud semantic segmentation network pre-trained in step S2 as a shared point cloud encoder to extract the support set and query set. The features of the 3D point cloud are denoted as follows: and Optimize using point cloud self-attention module and Furthermore, it will support multiplying point cloud features with ground truth masks to calculate the specific class support set features corresponding to each class. ;

[0068] S3.2, Multi-view Text Supervision Generation: To expand the semantic expressive power of the support set S, this invention considers utilizing the semantic knowledge expressed by the text as supplementary supervision information. Specifically, this invention generates a different set of text descriptions for each semantic category, referred to as "multi-view text supervision," which can be generated in two ways:

[0069] (1) Word2vec: Word2vec is a word embedding technique used to map words to a high-dimensional vector space, aiming to capture the semantic relationships between words. Specifically, this invention uses the name of the category as text information, i.e. ,For example The Word2vec model, pre-trained on a large-scale text corpus, is used to extract reliable text features for each category name; and fully connected layers (FC) are further used to reduce the feature dimension to [value missing]. The resulting text features can be represented as: ;

[0070] (2) 3D Command Prompt Large Language Model (LLM): Large Language Model (LLM) refers to a language model with a huge parameter scale, such as GPT-3, which understands and generates natural language text through pre-training on a large amount of text data. This invention first uses multiple sets of different heuristic 3D command prompts (GPT) to generate text with rich 3D semantics as supplementary category information. The 3D prompt commands used include:

[0071]

[0072]

[0073]

[0074] Inputting the obtained multi-view text representation into a text encoder yields a multi-view text embedding, denoted as . ,in Indicates the number of 3D prompt commands used.

[0075] Step S3.2 combines word embedding (Word2vec) and the 3D command prompt large language model (LLM) to obtain the final multi-view text embedding, which can be represented as follows: ,in This represents the total number of views for each class.

[0076] S3.3, Text-Guided Support Set Enhancement Module: Designed for each category Integrating its point cloud support set features and text features To improve semantic expressiveness, it specifically includes a text self-attention module and a cross-attention enhancement module:

[0077] (1) Text self-attention module: Implemented by the self-attention mechanism of Transformer, it enables the model to focus on the feature distribution of multiple views within the class, that is, the text features of multiple views are optimized by passing through a self-attention layer. ;

[0078] (2) Cross-attention enhancement module: Implemented by the Transformer's cross-attention mechanism, this module allows the model to enhance its attention to each point cloud feature by considering information from the text. Specifically, it focuses on point cloud features of specific classes. Extract and from text features Extract Enhanced specific point cloud features are obtained through a cross-attention layer. ;

[0079] S3.4, Category-Specific Cross-Attention Module: Implemented by the Transformer's cross-attention mechanism, it integrates the query set's point cloud features. Support features for specific classes Aggregation, from and Extract separately and This generates attention for a specific class, ultimately yielding a representation for that specific class. ;

[0080] S3.5, obtain the segmentation result: a representation of a specific class of the query set. Use two Convolutional layer and one The layer yields the model's prediction results for a single category c. For all Each category can be used to obtain prediction results for each category separately. , … , Finally, it can be spliced ​​together. Prediction results for each category ,… , }, and use a The layer yields prediction results for all categories. ;

[0081] S3.6, Optimize the objective function: During training, use the cross-entropy loss function. The network is optimized using the Adam optimizer with a learning rate of 0.001, as shown in the formula:

[0082] .

[0083] Specifically, step S4 in this embodiment of the invention includes the following steps:

[0084] S4.1 Extract point cloud features: Similar to step S3.1, extract point cloud features for the support set and query set of the test set; and extract category-specific support set features;

[0085] S4.2, Multi-view text supervision generation: Similar to step S3.2, multi-view text supervision is generated for the new category using the large language model LLM and word embedding Word2vec as auxiliary guidance information for the support set;

[0086] S4.3, Text-guided support set enhancement module: Similar to step S3.3, the multi-view text supervision information is optimized through the text self-attention module; the multi-view text supervision information is optimized for support set features through the cross-attention enhancement module;

[0087] S4.4, Category-specific cross-attention module: Similar to step S3.4, obtain category-specific point cloud representations for the query set;

[0088] S4.5, obtain the segmentation result: similar to step S3.5, using a fully connected layer and Layers are used to obtain the segmentation results;

[0089] S4.6, Evaluate Segmentation Performance: The mIoU metric is used to evaluate the segmentation performance on the test set. Evaluation results on the S3DIS and ScanNet datasets are shown in Tables 1 and 2, respectively. Experimental results demonstrate that the proposed few-sample point cloud semantic segmentation method based on multi-view text supervision outperforms other existing algorithms.

[0090] Table 1. Comparison with existing few-sample point cloud semantic algorithms on the S3DIS dataset.

[0091]

[0092] Table 2 Comparison with existing few-sample point cloud semantic algorithms on the ScanNet dataset.

[0093]

[0094] In particular, step S4 of the present invention is similar to step S3, the main difference being that the input of S3 is a training support set and a training query set containing known categories, while the input of S4 is a test support set and a test query set containing only new categories.

[0095] Specifically, step S4 freezes all parameters of the model and does not calculate any loss function;

[0096] In summary, the embodiments of the present invention can train a few-shot point cloud segmentation model based on multi-view text supervision guidance with only a small amount of labeled point cloud data. Using this model, more comprehensive and richer multi-view text supervision information can be generated for semantic categories to enhance the support set point cloud, thereby assisting in the segmentation of the query set. Furthermore, for new categories not appearing during the training phase, the present invention can significantly improve the performance of few-shot point cloud segmentation with only a small amount of labeled point cloud data.

[0097] Example 2:

[0098] This invention provides a few-sample point cloud semantic segmentation system based on multi-view text supervision guidance, comprising:

[0099] The data acquisition and preprocessing module is used to acquire and preprocess the point cloud semantic segmentation dataset, divide the point cloud semantic segmentation dataset into training set and test set, and generate small sample support set and query set for the training set and test set respectively.

[0100] The pre-training module is used to train the deep learning-based point cloud semantic segmentation model and optimize it using the cross-entropy loss function.

[0101] The model building and training module is used to construct a few-sample point cloud semantic segmentation model based on multi-view text supervision. The model is fine-tuned using the support set and query set of the training set to obtain the trained few-sample point cloud semantic segmentation model. Specifically, a pre-trained backbone segmentation model is used to extract point cloud features from the support set and query set, and the support set features are enhanced through a point cloud self-attention mechanism. Finally, the support set point cloud features for each category are obtained by multiplying the support set features and the support set binary mask. Combining a word embedding model and a large language model using 3D command prompts, multi-view text description information is generated for each semantic category, and the corresponding multi-view text features are extracted. The model is then guided by text... A support set enhancement mechanism fuses text features with support set point cloud features to obtain an enhanced support set representation. This enhancement mechanism includes a text self-attention mechanism and a cross-attention mechanism to improve semantic alignment. A category-specific cross-attention module is used to aggregate query set point cloud features with the enhanced support set features, generating a category-guided query set feature representation. The query set features are then input into a classifier for segmentation prediction. This classifier includes convolutional layers and a softmax layer to output the point-by-point prediction probability of the query set. The prediction error is calculated based on the cross-entropy loss function, and the Adam optimizer is used to optimize and update the overall small-sample point cloud semantic segmentation model.

[0102] The testing and evaluation module is used to test the small sample point cloud semantic segmentation model using the support set and query set of the test set, and evaluate the model performance using the mIoU metric.

[0103] Preferably, the model building and training module in this embodiment of the invention includes:

[0104] A multi-view text supervision generation unit is used to combine word embedding models and large language models to generate multi-view text features for semantic categories.

[0105] The point cloud feature enhancement unit is used to optimize the support set and query set features through the point cloud self-attention mechanism and generate category-specific support set features;

[0106] Text-guided support set enhancement units are used to fuse text features and support set features using text self-attention and cross-attention enhancement mechanisms.

[0107] Category-specific attention units are used to aggregate query set features and enhanced support set features through a cross-attention mechanism;

[0108] The segmented output subunit is used to output the query set segmentation results through convolutional layers and softmax layers.

[0109] Example 3: This embodiment of the invention provides an electronic device, including a processor and a memory storing a computer program. When the processor executes the computer program, it implements the small sample point cloud semantic segmentation method based on multi-view text supervision guidance provided in Example 1 of the invention.

[0110] Example 4: This embodiment of the invention provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the small sample point cloud semantic segmentation method based on multi-view text supervision guidance provided in Example 1 of the invention.

[0111] Example 5: This embodiment of the invention provides a server, including at least one processor and a memory communicatively connected to the processor. The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the processor to cause the at least one processor to execute the small sample point cloud semantic segmentation method based on multi-view text supervision guidance provided in Example 1 of the invention.

[0112] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in the present invention, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0113] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0114] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A few-sample point cloud semantic segmentation method based on multi-view text supervision, characterized in that, The method includes: (1) Obtain the point cloud semantic segmentation dataset and preprocess it. Divide the point cloud semantic segmentation dataset into a training set and a test set. Establish a small sample support set and a query set for the training set and the test set respectively. (2) Construct a point cloud semantic segmentation model based on deep learning, use the training set to iteratively train the point cloud semantic segmentation model, and use the cross-entropy loss function to optimize it, so as to obtain a pre-trained backbone segmentation model with basic semantic segmentation capabilities. (3) Based on the pre-trained backbone segmentation model, a small-sample point cloud semantic segmentation model with multi-view text supervision is constructed, and fine-tuned using the support set and query set in the training set. Specifically, this includes: Point cloud features of the support set and query set are extracted using a pre-trained backbone segmentation model. The support set features are then enhanced using a point cloud self-attention mechanism. Finally, the support set point cloud features for each category are obtained by multiplying the support set features and the support set binary mask. By combining word embedding models and large language models using 3D command prompts, multi-view text description information is generated for each semantic category, and corresponding multi-view text features are extracted. The enhanced support set representation is obtained by fusing text features and support set point cloud features through a text-guided support set enhancement mechanism; the enhancement mechanism includes a text self-attention mechanism and a cross-attention mechanism to improve semantic alignment capability. By using a category-specific cross-attention module, the query set point cloud features are aggregated with the enhanced support set features to generate a category-guided query set feature representation. The query set features are input into a classifier for segmentation prediction. The classifier includes a convolutional layer and a softmax layer, which are used to output the point-by-point prediction probability of the query set. The prediction error is calculated based on the cross-entropy loss function, and the Adam optimizer is used to optimize and update the overall small sample point cloud semantic segmentation model. (4) The trained model is tested using the support set and query set in the test set, and the model performance is evaluated by calculating the mIoU index based on the prediction results.

2. The few-sample point cloud semantic segmentation method based on multi-view text supervision guidance according to claim 1, characterized in that, In step (1), the point cloud semantic segmentation dataset is divided into known categories and unknown categories according to semantic categories. The training set includes known categories, and the test set includes unknown categories.

3. The few-sample point cloud semantic segmentation method based on multi-view text supervision guidance according to claim 1, characterized in that, The deep learning-based point cloud semantic segmentation model includes a feature extractor composed of a dynamic graph CNN architecture, an attention learner composed of a self-attention network, and a metric learner composed of multiple perceptron layers, which are used to extract local geometric features, aggregate global context features, and map features into the manifold space, respectively.

4. The few-sample point cloud semantic segmentation method based on multi-view text supervision guidance according to claim 1, characterized in that, In step (4), the testing of the trained model using the support set and query set in the test set specifically includes: Extract point cloud features from the support set and query set of the test set to obtain category-specific support set point cloud features; For each new category, a large language model combining word embedding and 3D command prompts is used to generate multi-view text features; The semantic expressive power of the support set is enhanced by using a text-guided support set enhancement mechanism; Use a category-specific cross-attention mechanism to obtain feature representations of a specific class in the query set; The query set segmentation result is obtained using a classifier containing convolutional layers and softmax layers.

5. A few-sample point cloud semantic segmentation system based on multi-view text supervision, characterized in that, The system includes: The data acquisition and preprocessing module is used to acquire and preprocess the point cloud semantic segmentation dataset, divide the point cloud semantic segmentation dataset into training set and test set, and generate small sample support set and query set for the training set and test set respectively. The pre-training module is used to train the deep learning-based point cloud semantic segmentation model and optimize it using the cross-entropy loss function. The model building and training module is used to construct a few-sample point cloud semantic segmentation model based on multi-view text supervision. The model is fine-tuned using the support set and query set of the training set to obtain the trained few-sample point cloud semantic segmentation model. Specifically, a pre-trained backbone segmentation model is used to extract point cloud features from the support set and query set, and the support set features are enhanced through a point cloud self-attention mechanism. Finally, the support set point cloud features for each category are obtained by multiplying the support set features and the support set binary mask. Combining a word embedding model and a large language model using 3D command prompts, multi-view text description information is generated for each semantic category, and the corresponding multi-view text features are extracted. The model is then guided by text... A support set enhancement mechanism fuses text features with support set point cloud features to obtain an enhanced support set representation. This enhancement mechanism includes a text self-attention mechanism and a cross-attention mechanism to improve semantic alignment. A category-specific cross-attention module is used to aggregate query set point cloud features with the enhanced support set features, generating a category-guided query set feature representation. The query set features are then input into a classifier for segmentation prediction. This classifier includes convolutional layers and a softmax layer to output the point-by-point prediction probability of the query set. The prediction error is calculated based on the cross-entropy loss function, and the Adam optimizer is used to optimize and update the overall small-sample point cloud semantic segmentation model. The testing and evaluation module is used to test the small sample point cloud semantic segmentation model using the support set and query set of the test set, and evaluate the model performance using the mIoU metric.

6. The few-sample point cloud semantic segmentation system based on multi-view text supervision and guidance according to claim 5, characterized in that, The model building and training module includes: The multi-view text supervision generation unit is used to jointly generate multi-view text features for semantic categories by combining word embedding models and large language models. The multi-view text supervision refers to generating diverse text descriptions of the same semantic category through word embedding models and large language models to enhance the model's understanding and generalization ability of semantic features. The point cloud feature enhancement unit is used to optimize the support set and query set features through the point cloud self-attention mechanism and generate category-specific support set features; Text-guided support set enhancement units are used to fuse text features and support set features using text self-attention and cross-attention enhancement mechanisms. Category-specific attention units are used to aggregate query set features and enhanced support set features through a cross-attention mechanism; The segmented output subunit is used to output the query set segmentation results through convolutional layers and softmax layers.

7. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the few-sample point cloud semantic segmentation method based on multi-view text supervision guidance as described in any one of claims 1 to 4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the few-sample point cloud semantic segmentation method based on multi-view text supervision guidance as described in any one of claims 1 to 4.

9. A server, characterized in that: The method includes at least one processor and a memory communicatively connected to the processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the processor to cause the at least one processor to perform the few-sample point cloud semantic segmentation method based on multi-view text supervision guidance as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Cross-modal single-sample three-dimensional point cloud segmentation method

    CN114529757A