Few-sample increment point cloud classification method based on multi-modal feature matching
By adopting a multimodal feature matching method in point cloud classification, combining the CLIP model and PointNet network, cross-modal feature fusion is solved, and the generalization ability and learning ability in the existing technology is insufficient, and efficient point cloud classification and generalization ability is achieved.
Patent Information
- Application Number
- CN202510055433.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The existing technology is difficult to effectively generalize to different fields, and lacks the ability to incrementally learn, making it difficult to adapt to scenarios where label data and multi-sample are lacking in point cloud classification tasks.
Using a multimodal feature matching method, the CLIP model is used to generate text description features of point clouds, and the image modal features are extracted through Vision Transformer, and three-dimensional modal features are obtained in combination with the PointNet network to perform cross-modal fusion to enhance classification accuracy and generalization.
It realizes efficient point cloud classification with few samples, improves the generalization ability and continuous learning ability of the method, and can adapt to scenarios of multiple samples and new categories.
Smart Images

Figure CN120107649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and three-dimensional computer vision, and in particular to a few-sample incremental point cloud classification method based on multimodal feature matching. Background Art
[0002] With the rapid development of 3D sensors and 3D reconstruction and generation technologies, the scale and types of 3D assets are growing rapidly. How to efficiently process 3D data has become one of the main bottlenecks in its application. Take point cloud as an example. As one of the important forms of 3D data, it has the characteristics of disorder, large data volume, and loose structure. Therefore, reliably extracting point cloud features is the core challenge of downstream tasks such as classification and segmentation. At the same time, point cloud classification often relies on a large amount of labeled data for fully supervised training, which brings great difficulties to emerging fields that lack relevant data sets. In addition, with the rapid increase in the number of point cloud categories, for example, in industrial defect detection, common defect categories such as breakage, cracking, and dirt may not cover all types of defects that appear in subsequent processes, which poses a severe challenge to the generalization ability and continuous learning ability of point cloud classification methods. Therefore, there is an urgent need for a point cloud classification method that can effectively generalize to different fields and has the ability of incremental learning with few samples.
[0003] Traditional point cloud classification methods usually extract point cloud features layer by layer by constructing neighborhood graphs, and then integrate these features into global features for downstream tasks. However, these methods rely on labeled fully supervised training, and each category requires a large number of sample support, which makes it difficult to adapt and generalize to few-sample classification tasks. In addition, some works have adopted three-dimensional distillation methods based on multimodal pre-training. Although they have made some progress in zero-sample and few-sample tasks, their classification stability on new category samples still needs to be improved. In addition, the categories and number of samples of point clouds in actual production applications are often unpredictable, which puts higher requirements on the robustness and continuous learning ability of existing methods. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings and defects of the prior art, and proposes a few-sample incremental point cloud classification method based on multimodal feature matching, which uses the widely used contrastive visual-language pre-training model (CLIP model) to generate point cloud text description features as prototype features for point cloud classification, and uses CLIP's Vision Transformer to extract image modal features from the two-dimensional depth map obtained from the point cloud projection. At the same time, the image modal features and the three-dimensional modal features obtained through the PointNet network are cross-modally fused to enhance the accuracy and generalization of point cloud classification. Finally, a binary classifier is used to distinguish between base class and new class samples, and effective classification is achieved by maximizing the feature distance between the two.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: a few-sample incremental point cloud classification method based on multimodal feature matching, which uses three methods: cross-modal feature fusion, image modality enhancement and parallel dual branches to achieve few-sample incremental point cloud classification, wherein the cross-modal feature fusion uses a cross-attention to fuse the three-dimensional modal features and image modality features of the point cloud, and then uses a self-attention to further enhance the fused cross-features; the image modality enhancement is to learn an adaptive background color based on the three-dimensional modality features of the point cloud, and enhance the point cloud foreground in the two-dimensional depth map through additional background information to highlight the point cloud body contour; the parallel dual branches are to dynamically train a binary classifier to distinguish between base class samples and new class samples to the greatest extent, and use two sets of different parameters to process the samples respectively;
[0006] The specific implementation of the method includes the following steps:
[0007] 1) For each input point cloud, the point cloud is first rendered into a 2D depth map, and the image modality features corresponding to the point cloud are extracted from the 2D depth map using a pre-trained depth map feature extractor. Then, a PointNet network is used to process the point cloud data to obtain the 3D modality features of the point cloud;
[0008] 2) A cross-attention is used to perform cross-modal feature fusion of the 3D modal features of the point cloud and the image modal features to obtain the cross-features of the point cloud. A self-attention is used to process the cross-features. At the same time, a mask ratio is introduced to perform a mask operation on the attention weights generated in the self-attention. By minimizing the cosine similarity difference between the two features generated by the masked and unmasked operations, an enhanced self-attention feature is obtained. Subsequently, a background color is generated according to the 3D modal features of the point cloud, and it is overlaid on the 2D depth map for image modality enhancement to obtain a background enhancement map. The obtained background enhancement map and the text description features of the point cloud are used as the input of the CLIP model to obtain a set of enhanced probability distribution values.
[0009] 3) On each incremental task, a binary classifier is dynamically trained based on only one sample from each base class and a few samples from the new class to maximize the margin between the base class and the new class, and two sets of different parameters are used to perform parallel dual-branch processing on the base class and new class samples;
[0010] 4) Calculate the dot product between the self-attention feature of the point cloud and the text description feature of the point cloud to obtain a set of probability distributions, fuse the probability distributions with the above-mentioned enhanced probability distribution values to obtain the final probability distribution, and finally select the maximum value of each point cloud in the above probability distribution as the prediction result.
[0011] Further, the step 1) comprises the following steps:
[0012] 1.1) Use a renderer that takes point cloud data as input, projects the point cloud onto a two-dimensional plane, and renders the corresponding two-dimensional depth map of the point cloud:
[0013] I p =Renderer(P)
[0014] Where P is the 3D point cloud data of shape [32, 1024, 3], 32 is the batch size, representing processing 32 point cloud data at a time, 1024 is the number of points contained in the point cloud, and 3 is the 3D coordinate of each point; Renderer is the renderer used to render the 3D point cloud into a 2D depth map, and I p is the two-dimensional depth map corresponding to the point cloud P;
[0015] 1.2) Use the pre-trained depth map feature extractor to extract depth map features from the 2D depth map of the point cloud:
[0016] F I =Encoder(I p )
[0017] Where Encoder is a depth map feature extractor pre-trained based on CLIP Vision Transformer, and F I is the obtained point cloud depth map feature, that is, the image modality feature of the point cloud;
[0018] 1.3) Acquisition of 3D modal features: A PointNet network is used to process the point cloud data to obtain the 3D modal features of the point cloud:
[0019] F 3D =PointNet(P)
[0020] In the formula, F 3D Represents the 3D modal features of the point cloud.
[0021] Further, the step 2) comprises the following steps:
[0022] 2.1) Cross-modal feature fusion: Project the 3D modal features and image modal features to the same feature dimension, and use a cross attention to fuse the features of the two modalities together to obtain the cross-features of the point cloud:
[0023] F cross =CrossAttention(F 3D ,F I )
[0024] In the formula, CrossAttention represents the cross attention module, which uses the three-dimensional modal features F of the point cloud 3D As the query, take the image modality feature F I As the key and value, the cross feature F is calculated cross ;
[0025] 2.2) Self-attention feature extraction: Use a self-attention pair cross feature F cross Processing is performed and a mask ratio R is introduced m The attention weights generated in the self-attention are masked, and more representative point cloud features are extracted by comparing the generated self-attention features with the masked self-attention features:
[0026]
[0027] In the formula, MaskedSelfAttention represents a masked self-attention module, which is based on the cross feature F cross And the mask ratio R m As input, R m is a floating point number ranging from 0 to 1. The mask operation represents when R m When it is greater than 0, the self-attention weights are sorted by R m The value of R is set to zero, and the non-mask operation represents m The value of is 0, that is, the mask is not used to modify the attention weight, and a set of paired masked self-attention features are calculated through the masked and unmasked operations of the masked self-attention and the self-attention feature F self , the masked self-attention module minimizes F self and The cosine similarity difference between them is used to learn more representative features, where Loss s Defined as cosine similarity difference, cos represents the cosine similarity between two features, with a value of 0 to 1. Only used to calculate the self-attention loss, F self is the final self-attention feature;
[0028] 2.3) Generation of background enhancement map: Generate a set of adaptive RGB values based on the 3D modal features of the point cloud as the background color of the 2D depth map of the point cloud:
[0029] R,G,B=ColorNet(F 3D )
[0030] In the formula, ColorNet is an MLP network that takes the three-dimensional modal features of the point cloud as input and generates a set of RGB values ranging from 0 to 1, that is, the values of the three channels of R, G, and B are all 0 to 1, corresponding to a background color;
[0031] The two-dimensional depth map of the point cloud I p Fusion with the background color to get the background enhancement image
[0032]
[0033] 2.4) Enhance the background image And the text description feature T corresponding to the point cloud p Input into the CLIP model to obtain the enhanced probability distribution value logits of the point cloud for each category aux , and have The text description feature T corresponding to the point cloud p Generate using the following template: "An image of a[class]", where class represents the category to which the point cloud belongs.
[0034] Further, the step 3) comprises the following steps:
[0035] 3.1) On each incremental task, first pre-train a binary classifier based on the sample with only one sample from each class in the base class and the few sample data of the incremental task, and then use this binary classifier to distinguish whether each sample of the current task belongs to the base class or the new class in the test phase;
[0036] 3.2) Generate corresponding position masks for base class samples, pass the base class data to the base class branch network with frozen parameters for processing, and send the new class data to the unfrozen current network for processing;
[0037] mask=MG(F 3D )
[0038]
[0039] In the formula, MG represents a binary classifier, mask is a Boolean value, which is converted to 0 or 1 when used, where 0 represents that the sample belongs to the new class and 1 represents that the sample belongs to the base class. mask The samples after the mask operation are then sent to the base class network. B Processing, while sending new class samples to the new class network Network N As another branch for parallel processing, the results calculated by the two branches are finally summarized to obtain a set of logits values, which represent the probability distribution of samples in each category.
[0040] Further, the step 4) comprises the following steps:
[0041] 4.1) By calculating the self-attention feature F of the point cloud self and its corresponding text description feature T p The dot product between them is used to obtain a set of probability distributions of the point cloud on the current category, and then the result is combined with logits aux The sum is a vector of shape [32, class'], where 32 is the batch size and class' is the total number of categories of the current task;
[0042] logits = F self ·T p +logits aux
[0043] 4.2) Take the maximum value in the class' dimension to obtain a vector of shape
[32] , which represents the classification labels corresponding to the 32 samples, that is, the final classification result.
[0044] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0045] 1. Cross-modal feature fusion includes a cross attention and a self-attention, which can effectively integrate information from multiple modalities to facilitate further feature matching. It also has good reusability and scalability and can be deployed to other fields involving cross-modal feature fusion.
[0046] 2. Image modality enhancement uses a simple, effective and lightweight method to enhance the features of the image modality, which can make more full use of the prior knowledge of CLIP and can also be deployed as a plug-and-play general module to other tasks.
[0047] 3. Parallel dual branch is based on the idea of parallel processing and uses two sets of parameters to process data from the base class and the new class, thereby effectively improving the performance of the method.
[0048] 4. The present invention surpasses the existing best methods in multiple few-sample classification tasks, and can be deployed in industrial scenarios such as defect detection and autonomous driving at low cost, realizing a deep integration of machine learning and industrial production practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 The data processing flow chart of the method of the present invention.
[0050] Figure 2 Schematic diagram of the cross-modal feature fusion.
[0051] Figure 3Schematic diagram of image modality enhancement.
[0052] Figure 4 This is the schematic diagram of the parallel double branch. DETAILED DESCRIPTION
[0053] The present invention is further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0054] like Figures 1 to 4 As shown, this embodiment discloses a method for classification of small-sample incremental point clouds based on multimodal feature matching, which is developed on the Ubuntu distribution of the Linux system using the Python programming language with version number 3.7, and can be run in an environment equipped with the Pytorch deep learning framework. This embodiment can be run on a single RTX 4090D graphics card. The method uses three methods, namely, cross-modal feature fusion, image modality enhancement, and parallel dual branches, to achieve classification of small-sample incremental point clouds, wherein the cross-modal feature fusion uses a cross attention to fuse the three-dimensional modality features and image modality features of the point cloud, and then uses a self-attention to further enhance the fused cross-features; the image modality enhancement is to learn an adaptive background color based on the three-dimensional features of the point cloud, and enhance the point cloud foreground in the two-dimensional depth map through additional background information to highlight the contour of the point cloud body; the parallel dual branches are to dynamically train a binary classifier to distinguish the base class samples from the new class samples to the greatest extent, and use two sets of different parameters to process the samples respectively; the specific implementation of this method includes the following steps:
[0055] 1) For each input point cloud, first render the point cloud into a two-dimensional depth map, and use a pre-trained depth map feature extractor to extract the image modality features corresponding to the point cloud from the two-dimensional depth map, and then use a PointNet network to process the point cloud data to obtain the three-dimensional modality features of the point cloud; including the following steps:
[0056] 1.1) Use a renderer that takes point cloud data as input, projects the point cloud onto a two-dimensional plane, and renders the corresponding two-dimensional depth map of the point cloud:
[0057] I p =Renderer(P)
[0058] Where P is the 3D point cloud data of shape [32, 1024, 3], 32 is the batch size, representing processing 32 point cloud data at a time, 1024 is the number of points contained in the point cloud, and 3 is the 3D coordinate of each point; Renderer is the renderer used to render the 3D point cloud into a 2D depth map, and I p is the two-dimensional depth map corresponding to the point cloud P;
[0059] 1.2) Use the pre-trained depth map feature extractor to extract image modality features from the 2D depth map of the point cloud:
[0060] F I =Encoder(I p )
[0061] Where Encoder is a depth map feature extractor pre-trained based on CLIP Vision Transformer. The CLIP version used in this embodiment is ViT-B / 32. I is the obtained point cloud depth map feature, that is, the image modality feature of the point cloud;
[0062] 1.3) Acquisition of 3D modal features: A PointNet network is used to process the point cloud data to obtain the 3D modal features of the point cloud:
[0063] F 3D =PointNet(P)
[0064] In the formula, F 3D The PointNet used in this embodiment includes four one-dimensional convolutional layers, where every two convolutional layers form a convolutional block, and each convolutional block uses batch normalization and ReLU activation function to process data.
[0065] 2) Using a cross attention to cross-modally fuse the three-dimensional modal features of the point cloud with the image modal features to obtain the cross features of the point cloud, using a self-attention to process the cross features, and introducing a mask ratio to mask the attention weights generated in the self-attention, and obtaining an enhanced self-attention feature by minimizing the cosine similarity difference between the two features generated by the mask and non-mask operations, and then generating a background color according to the three-dimensional modal features of the point cloud, and overlaying it on the two-dimensional depth map for image modality enhancement to obtain a background enhancement map, and using the obtained background enhancement map and the text description features of the point cloud as the input of the CLIP model to obtain a set of enhanced probability distribution values; comprising the following steps:
[0066] 2.1) Cross-modal feature fusion: Project the 3D modal features and image modal features to the same feature dimension, and use a cross attention to fuse the features of the two modalities together to obtain the cross-features of the point cloud:
[0067] F cross =CrossAttention(F 3D ,F I )
[0068] In the formula, CrossAttention represents the cross attention module, which uses the three-dimensional modal features F of the point cloud 3D As the query, take the image modality feature F I As the key and value, the cross feature F is calculated cross In this embodiment, the cross attention module uses the three-dimensional features F of the point cloud 3D As the query, the depth map feature F I As the key and value, the attention weight and value are then calculated to obtain the cross feature F cross , the cross attention is regularized using dropout with a value of 0.1.
[0069] 2.2) Self-attention feature extraction: Use a self-attention pair cross feature F cross Processing is performed and a mask ratio R is introduced m The attention weights generated in the self-attention are masked, and more representative point cloud features are extracted by comparing the generated self-attention features with the masked self-attention features:
[0070]
[0071]
[0072] In the formula, MaskedSelfAttention represents a masked self-attention module, which is based on the cross feature F cross And the mask ratio R m As input, R m is a floating point number ranging from 0 to 1. The mask operation represents when R m When it is greater than 0, the self-attention weights are sorted by R m The value of R is set to zero, and the non-mask operation represents m The value of is 0, that is, the mask is not used to modify the attention weight; a set of paired masked self-attention features are calculated through the masked and unmasked operations of the masked self-attention and the self-attention feature F self , the masked self-attention module minimizes F self and The cosine similarity difference between them is used to learn more representative features, where Loss s Defined as cosine similarity difference, cos represents the cosine similarity between two features, with a value of 0 to 1. Only used to calculate the self-attention loss, F self is the final self-attention feature;
[0073] 2.3) Generation of background enhancement map: Generate a set of adaptive RGB values based on the 3D modal features of the point cloud as the background color of the 2D depth map of the point cloud:
[0074] R,G,B=ColorNet(F 3D )
[0075] In the formula, ColorNet is an MLP network that takes the three-dimensional modal features of the point cloud as input and generates a set of RGB values ranging from 0 to 1, that is, the values of the three channels of R, G, and B are all 0 to 1, corresponding to a background color;
[0076] The two-dimensional depth map of the point cloud I p Fusion with the background color to get the background enhancement image
[0077]
[0078] 2.4) Enhance the background image And the text description feature T corresponding to the point cloud p Input into the CLIP model to obtain the enhanced probability distribution value logits of the point cloud for each category aux , and have The text description feature T corresponding to the point cloud p Use the following template to generate: "An image of a[class]", where class represents the category to which the point cloud belongs. For example, for a point cloud belonging to the car category, its corresponding text description is: "An image of a car". The CLIP version used in this example is ViT-B / 32, and the parameters of the Vision Transformer and Text Encoder contained in it are consistent with this version.
[0079] 3) On each incremental task, dynamically train a binary classifier based on only one sample from each class of the base class and a few samples from the new class to maximize the margin between the base class and the new class, and use two different sets of parameters to perform parallel dual-branch processing on the base class and new class samples; including the following steps:
[0080] 3.1) On each incremental task, a binary classifier MG is first pre-trained based on the sample with only one sample from each class of the base class and the few sample data of the incremental task. Then, the binary classifier is used to distinguish whether each sample of the current task belongs to the base class or the new class during the test phase. The binary classifier used in this example takes the 3D modal features of the point cloud as input and is optimized by calculating the binary cross entropy loss between the generated mask and the true value.
[0081] 3.2) When MG training is completed, it generates the corresponding mask for the input sample, passes the base class data to the base class branch network with frozen parameters for processing, and sends the new class data to the current network that is not frozen for processing;
[0082] mask=MG(F 3D )
[0083]
[0084] In the formula, MG represents a binary classifier, mask is a Boolean value, which is converted to 0 or 1 when used, where 0 represents that the sample belongs to the new class and 1 represents that the sample belongs to the base class. mask The samples after the mask operation are then sent to the base class network. B Processing, while sending new class samples to the new class network Network N As another branch for parallel processing, the results of the two branches are finally aggregated to obtain a set of logits values, representing the probability distribution of samples in each category. B After training is completed, its parameters are frozen, and the new class network Network N Keep all its parameter training configuration unchanged.
[0085] 4) Calculate the dot product between the self-attention feature of the point cloud and the text description feature of the point cloud to obtain a set of probability distributions, fuse the probability distributions with the above-mentioned enhanced probability distribution values to obtain the final probability distribution, and finally select the maximum value of each point cloud in the above probability distribution as the prediction result. The following steps are included:
[0086] 4.1) By calculating the self-attention feature F of the point cloud self and its corresponding text description feature T p The dot product between them is used to obtain a set of probability distributions of the point cloud on the current category, and then the result is combined with logits aux The sum is a vector of shape [32, class'], where 32 is the batch size and class' is the total number of categories of the current task;
[0087] logits = F self ·T p +logits aux
[0088] 4.2) Take the maximum value in the class' dimension to obtain a vector of shape
[32] , which represents the classification labels corresponding to the 32 samples, that is, the final classification result.
[0089] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A few-sample incremental point cloud classification method based on multimodal feature matching, characterized in that: The method uses three methods: cross-modal feature fusion, image modality enhancement and parallel dual branches to achieve small-sample incremental point cloud classification. The cross-modal feature fusion uses a cross-attention to fuse the three-dimensional modality features and image modality features of the point cloud, and then uses a self-attention to further enhance the fused cross-features; the image modality enhancement is to learn an adaptive background color based on the three-dimensional modality features of the point cloud, and enhance the point cloud foreground in the two-dimensional depth map through additional background information to highlight the point cloud body contour; the parallel dual branches are to dynamically train a binary classifier to distinguish the base class samples from the new class samples to the greatest extent, and use two different sets of parameters to process the samples respectively; The specific implementation of the method includes the following steps: 1) For each input point cloud, the point cloud is first rendered into a 2D depth map, and the image modality features corresponding to the point cloud are extracted from the 2D depth map using a pre-trained depth map feature extractor. Then, a PointNet network is used to process the point cloud data to obtain the 3D modality features of the point cloud; 2) A cross-attention is used to perform cross-modal feature fusion of the 3D modal features of the point cloud and the image modal features to obtain the cross-features of the point cloud. A self-attention is used to process the cross-features. At the same time, a mask ratio is introduced to perform a mask operation on the attention weights generated in the self-attention. By minimizing the cosine similarity difference between the two features generated by the masked and unmasked operations, an enhanced self-attention feature is obtained. Subsequently, a background color is generated according to the 3D modal features of the point cloud, and it is overlaid on the 2D depth map for image modality enhancement to obtain a background enhancement map. The obtained background enhancement map and the text description features of the point cloud are used as the input of the CLIP model to obtain a set of enhanced probability distribution values. 3) On each incremental task, a binary classifier is dynamically trained based on only one sample from each base class and a few samples from the new class to maximize the margin between the base class and the new class, and two sets of different parameters are used to perform parallel dual-branch processing on the base class and new class samples; 4) Calculate the dot product between the self-attention feature of the point cloud and the text description feature of the point cloud to obtain a set of probability distributions, fuse the probability distributions with the above-mentioned enhanced probability distribution values to obtain the final probability distribution, and finally select the maximum value of each point cloud in the above probability distribution as the prediction result.
2. The method for small sample incremental point cloud classification based on multimodal feature matching according to claim 1, characterized in that: Step 1) The following steps are involved: 1.1) Use a renderer that takes point cloud data as input, projects the point cloud onto a two-dimensional plane, and renders the corresponding two-dimensional depth map of the point cloud: IN p =Renderer(P) Where P is the 3D point cloud data of shape [32, 1024, 3], 32 is the batch size, representing processing 32 point cloud data at a time, 1024 is the number of points contained in the point cloud, and 3 is the 3D coordinate of each point; Renderer is the renderer used to render the 3D point cloud into a 2D depth map, and I p is the two-dimensional depth map corresponding to the point cloud P; 1.2) Use the pre-trained depth map feature extractor to extract depth map features from the 2D depth map of the point cloud: F I =Encoder(I p ) Where Encoder is a depth map feature extractor pre-trained based on CLIP Vision Transformer, and F I is the obtained point cloud depth map feature, that is, the image modality feature of the point cloud; 1.3) Acquisition of 3D modal features: A PointNet network is used to process the point cloud data to obtain the 3D modal features of the point cloud: F 3D =PointNet(P) In the formula, F 3D Represents the 3D modal features of the point cloud.
3. The method for small sample incremental point cloud classification based on multimodal feature matching according to claim 2, characterized in that: The step 2) comprises the following steps: 2.1) Cross-modal feature fusion: Project the 3D modal features and image modal features to the same feature dimension, and use a cross attention to fuse the features of the two modalities together to obtain the cross-features of the point cloud: F cross =CrossAttention(F 3D ,F I ) In the formula, CrossAttention represents the cross attention module, which uses the three-dimensional modal features F of the point cloud 3D As the query, take the image modality feature F I As the key and value, the cross feature F is calculated cross ; 2.2) Self-attention feature extraction: Use a self-attention pair cross feature F cross Processing is performed and a mask ratio R is introduced m The attention weights generated in the self-attention are masked, and more representative point cloud features are extracted by comparing the generated self-attention features with the masked self-attention features: In the formula, MaskedSelfAttention represents a masked self-attention module, which is based on the cross feature F cross And the mask ratio R m As input, R m is a floating point number ranging from 0 to 1. The mask operation represents when R m When it is greater than 0, the self-attention weights are sorted by R m The value of R is set to zero, and the non-mask operation represents m The value of is 0, that is, the mask is not used to modify the attention weight; a set of paired masked self-attention features are calculated through the masked and unmasked operations of the masked self-attention and the self-attention feature F self , the masked self-attention module minimizes and F self The cosine similarity difference between them is used to learn more representative features, where Loss s Defined as cosine similarity difference, cos represents the cosine similarity between two features, with a value of 0 to 1. Only used to calculate the self-attention loss, F self is the final self-attention feature; 2.3) Generation of background enhancement map: Generate a set of adaptive RGB values based on the 3D modal features of the point cloud as the background color of the 2D depth map of the point cloud: R,G,B=ColorNet(F 3D ) In the formula, ColorNet is an MLP network that takes the three-dimensional modal features of the point cloud as input and generates a set of RGB values ranging from 0 to 1, that is, the values of the three channels of R, G, and B are all 0 to 1, corresponding to a background color; The 2D depth map of the point cloud I p Fusion with the background color to get the background enhancement image 2.4) Enhance the background image And the text description feature T corresponding to the point cloud p Input into the CLIP model to obtain the enhanced probability distribution value logits of the point cloud for each category aux , and have The text description feature T corresponding to the point cloud p Generate using the following template: "An image of a[class]", where class represents the category to which the point cloud belongs.
4. The method for small sample incremental point cloud classification based on multimodal feature matching according to claim 3, characterized in that: The step 3) comprises the following steps: 3.1) On each incremental task, first pre-train a binary classifier based on the sample with only one sample from each class in the base class and the few sample data of the incremental task, and then use this binary classifier to distinguish whether each sample of the current task belongs to the base class or the new class in the test phase; 3.2) Generate corresponding position masks for base class samples, pass the base class data to the base class branch network with frozen parameters for processing, and send the new class data to the unfrozen current network for processing; mask=MG(F 3D ) In the formula, MG represents a binary classifier, mask is a Boolean value, which is converted to 0 or 1 when used, where 0 represents that the sample belongs to the new class and 1 represents that the sample belongs to the base class. mask The samples after the mask operation are then sent to the base class network. B Processing, while sending new class samples to the new class network Network N As another branch, the results calculated by the two branches are finally aggregated to obtain a set of logits values, which represent the probability distribution of samples in each category.
5. The method for small sample incremental point cloud classification based on multimodal feature matching according to claim 4, characterized in that: The step 4) comprises the following steps: 4.1) By calculating the self-attention feature F of the point cloud self and its corresponding text description feature T p The dot product between them is used to obtain a set of probability distributions of the point cloud on the current category, and then the result is combined with logits aux The sum is a vector of shape [32, class'], where 32 is the batch size and class' is the total number of categories of the current task; logits=F self ·T p +logits aux 4.2) Take the maximum value in the class' dimension to obtain a vector of shape [32], which represents the classification labels corresponding to the 32 samples, that is, the final classification result.
Citation Information
Patent Citations
Three-dimensional model identification method based on point cloud multi-view fusion
CN112347932A
Cross-modal single-sample three-dimensional point cloud segmentation method
CN114529757A
Industrial scene point cloud multi-modal registration method and system based on deep learning
CN118071805A
Traffic rail deformation detection method and device based on multi-modal three-dimensional point cloud fusion
CN118485898A
Multiscale point cloud classification method and system
US20230222768A1