Multi-modal Feature Fusion-based Point Cloud Data Classification Method and Device

By adopting the multimodal feature fusion method in 3D point cloud data classification, the features of the multi-view convolutional neural network model and the point cloud Transformer model are fused, which solves the shortcomings of traditional models in information loss and global feature extraction, and achieves a more efficient point cloud data classification effect.

CN114494708BActive Publication Date: 2025-06-27SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210085153.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-06-27
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The existing technology has problems such as information loss and difficulty in extracting global features in 3D point cloud data classification. Traditional CNN models have strong capabilities in modeling underlying information, but weaker processing capabilities for global associations. The Transformer model pays more attention to high-level semantic information and is difficult to take into account both.

Method used

The multimodal feature fusion method is adopted to fuse the image features extracted from the multi-view convolutional neural network model and the point cloud features extracted from the point cloud Transformer model. The feature fusion module is used to superimpose intermediate results and reduce dimensionality, combining the advantages of both to achieve better point cloud classification effect.

Benefits of technology

Through multimodal feature fusion, features at different levels can be extracted and fused more effectively, the accuracy and effectiveness of 3D point cloud data classification can be improved, and the shortcomings of a single model in information loss and global feature extraction are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494708B_ABST
    Figure CN114494708B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for classifying point cloud data based on multi-modal feature fusion. The method includes the following steps: extracting image features using a pre-established multi-view convolutional neural network model; extracting point cloud features from point cloud data using a pre-established point cloud Transformer model; performing multi-modal feature fusion on the image features and point cloud features using a feature fusion module, and obtaining a classification result of the point cloud data according to the fused features; the feature fusion module includes a first path and a second path, the first path superimposes the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module at the middle using corresponding sizes; the second path is the feature obtained from the point cloud Transformer model, and also superimposes the output of the original point cloud Transformer model and the intermediate result at the corresponding scale position in the middle. The present invention complements the advantages and disadvantages of the two models through the designed multi-modal feature fusion module, thereby improving the classification effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision 3D point cloud data classification, and particularly relates to a method and device for classifying point cloud data based on multi-modal feature fusion. Background Art

[0002] The 3D point cloud data classification task is one of the important tasks in point cloud data processing. With the increasing number of 3D point cloud data acquisition channels, the processing of 3D point cloud data has gradually become popular. Due to the disorder, noise interference, and occlusion relationship of 3D point cloud data, it is very challenging to process such data. In the previously proposed 3D point cloud data classification models, there are mainly three methods: multi-view based, voxel based, and point based. The multi-view based method mainly extracts features through convolution of images from multiple perspectives for classification. The voxel based method mainly uses volume representation. However, the volume data may grow rapidly, be very large in scale, and take a long time to process. The point based method can be further divided into point-by-point fully connected networks, convolution based, graph based, high-level data structure based, and other typical methods. Recently, the proposed Transformer method has achieved good results in natural language processing. Since it pays attention to global information and is insensitive to the input order, it has advantages in processing point cloud data.

[0003] In the prior art, the ICCV2015 paper "Multi-view Convolutional Neural Networks for 3D Shape Recognition" proposed a multi-view convolutional neural network for classifying 3D point cloud data. This model projects the point cloud data through multiple angles to obtain multi-view 2D images, and then the convolutional neural network extracts features. Then, the maximum value of each element of multiple views is retained as the output of view-pooling. Finally, another convolutional neural network is used for classification. Another paper from (Computational Visual Media), "PCT: Point Cloud Transformer", proposed a method for directly processing 3D point cloud data. The paper uses the Transformer model to directly process 3D point cloud data. Since the core part of the Transformer is the self-attention mechanism, it is insensitive to the input sequence order, that is, no matter in what order the input is, it can extract effective information. Therefore, it is very helpful for processing unordered data such as 3D point clouds. After encoding through the attention mechanism, the hidden space representation of the point cloud data is obtained. Then, decoding this hidden space representation can perform different tasks, such as point cloud classification, point cloud segmentation, etc.

[0004] However, the method using multi-view convolution will cause a certain information loss problem because it retains the maximum value of the features of a certain view, making it difficult to take into account the information of some other un-retained views. At the same time, the convolution operation depends on the receptive field provided by the convolution kernel and it is difficult to extract global features. The method using point cloud transformers can better focus on global correlations, but the transformer's ability to model low-level information is not as good as that of traditional CNN networks, such as translational and rotational invariance, etc. Transformer pays more attention to high-level semantic information, that is, how to reasonably combine various elements to form an object. Summary of the Invention

[0005] The main object of the present invention is to overcome the disadvantages and deficiencies of the prior art, and to provide a multi-modal feature fusion-based point cloud data classification method and device. The present invention extracts multi-modal features from point cloud data and grayscale image data generated by point cloud rendering, and fuses the features of different modalities to achieve complementarity to achieve better results in point cloud classification tasks.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] On the one hand, the present invention provides a multi-modal feature fusion-based point cloud data classification method, including the following steps:

[0008] Use a pre-established multi-view convolutional neural network model to extract image features;

[0009] Use a pre-established point cloud Transformer model to extract point cloud features from point cloud data;

[0010] Perform multi-modal feature fusion on the image features and point cloud features using a feature fusion module, and obtain a point cloud data classification result according to the fused features; the feature fusion module includes a first path and a second path. The first path inputs the feature map obtained from the multi-view convolutional neural network model, and superimposes the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module at the corresponding size in the middle; the second path is the feature obtained from the point cloud Transformer model, and also superimposes the output of the original point cloud Transformer model and the intermediate result at the corresponding scale position in the middle.

[0011] As a preferred technical solution, when performing image feature extraction, given the projections at K different positions of the input to simulate the images obtained from the perspectives of cameras at K positions, the feature maps of the images from K perspectives are respectively extracted by VGGNet with shared weights. Through the perspective pooling operation, the maximum value of all perspective results at each position is retained to obtain the features based on the image input.

[0012] As a preferred technical solution, K is selected to be 12, and the intervals between every two of the 12 perspectives are 30°, and they point downward from 30° above the plane towards the grid centroid.

[0013] As a preferred technical solution, the image features pass through three fully connected layers to reduce the dimension of the extracted feature maps to the set dimension, and the outputs of the last three fully connected layers are respectively used as the inputs corresponding to the sizes of the paths of the multi-view convolution input of the multi-modal feature fusion module.

[0014] As a preferred technical solution, when performing point cloud feature extraction, the specific steps are as follows:

[0015] First, the point cloud data is input into an encoder, and the encoder consists of four layers of attention mechanisms;

[0016] Then, the results of the four-layer attention operations are concatenated, and then through linear transformation, batch normalization, non-linear activation, and Dropout layers, the features of the points are obtained;

[0017] Finally, after max pooling and average pooling and then concatenation, an n*1 global feature vector is obtained, and the global feature vector will be used as the second input of the feature fusion module.

[0018] As a preferred technical solution, the global feature vector will follow the subsequent operations of the original model. Through the fully connected layers, the vector changes from n*1 to n / 2*1 and then to n / 4*1, and finally is transformed into a vector with the set dimension. The subsequent three fully connected layers will also be respectively input to the corresponding scale positions of the point cloud Transformer input path of the feature fusion module.

[0019] As a preferred technical solution, the specific steps for performing multi-modal feature fusion are as follows:

[0020] The feature maps extracted by the multi-view convolutional neural network model are reduced in dimension by the first encoder to form a first vector with a dimension of n*1. The point cloud features extracted by the multi-view convolutional neural network model are used as a second vector with a dimension of n*1. The first vector and the second vector are concatenated into a third vector with a dimension of 2n*1;

[0021] The third vector passes through the first decoder to obtain a fourth vector with a dimension of 4n*1 and a fifth vector with a dimension of n / 2*1. The fourth vector serves as the input of the first path, and the fifth vector serves as the input of the second path. In the first path, the fourth vector is first superimposed with the output of the first fully connected layer of the multi-view convolutional neural network model at the corresponding scale, and then becomes a vector with a dimension of n / 2*1 through the second encoder. In the second path, the fifth vector with a dimension of n / 2*1 is first superimposed with the output of the first fully connected layer of the point cloud Transformer model at the corresponding scale;

[0022] The vectors after the first superimposition of the two paths are concatenated to form a sixth vector with a dimension of n*1. The sixth vector is decoded by the second decoder into a seventh vector with a dimension of 4n*1 and an eighth vector with a dimension of n / 4*1. In the first channel, the seventh vector is superimposed with the vector output by the second fully connected layer of the multi-view convolutional neural network model. The resulting vector after superimposition passes through the third encoder to obtain a ninth vector with a dimension of n / 4*1. In the second path, the eighth vector is superimposed with the output of the second fully connected layer of the same dimension of the point cloud Transformer model to obtain a tenth vector with a dimension of n / 4*1;

[0023] The ninth vector and the tenth vector are concatenated to form an eleventh vector with a dimension of n / 2*1. The eleventh vector passes through the third decoder to form two vectors with set dimensions. The two vectors with set dimensions are respectively concatenated with the outputs of the third fully connected layer of the multi-view convolutional neural network model and the third fully connected layer of the point cloud Transformer model, and then uniformly pass through a fully connected layer to obtain a vector with a set dimension;

[0024] The classification task is performed on the vectors with the final set dimensions of the multi-view convolutional neural network model, the vectors with the final set dimensions of the point cloud Transformer model, and the vectors with the set dimensions obtained in the previous paragraph.

[0025] On the other hand, the present invention provides a multi-modal feature fusion point cloud data classification device applied to the multi-modal feature fusion point cloud data classification method described above, including an image feature extraction module, a point cloud feature extraction module, and a multi-modal feature fusion module;

[0026] The image feature extraction module is used to extract image features through a pre-established multi-view convolutional neural network model;

[0027] The point cloud feature extraction module is used to extract point cloud features from point cloud data through a pre-established point cloud Transformer model;

[0028] The multi-modal feature fusion module is used to perform multi-modal feature fusion on the image features and point cloud features, and obtain the classification result of the point cloud data according to the fused features; the feature fusion module includes a first path and a second path. The first path takes the feature map obtained from the multi-view convolutional neural network model as input, and stacks the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module at the intermediate using the corresponding size. The second path is the feature obtained from the point cloud Transformer model, and also stacks the output of the original point cloud Transformer model and the intermediate result at the corresponding scale position in the middle.

[0029] In another aspect of the present invention, there is provided an electronic device, which includes:

[0030] At least one processor; and,

[0031] A memory communicatively connected to the at least one processor; wherein,

[0032] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the multi-modal feature fusion point cloud data classification method described above.

[0033] In yet another aspect of the present invention, there is provided a computer-readable storage medium storing a program, and when the program is executed by a processor, the multi-modal feature fusion point cloud data classification method described above is implemented.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] The present invention uses the idea of feature fusion to fuse the existing model implementations and combines the information extracted by different models. The advantages and disadvantages of the two models are complemented through the designed multi-modal feature fusion module: the multi-view convolutional model pays more attention to the underlying information, and the point cloud Transformer model pays more attention to the high-level semantic information, which improves the final classification effect.

[0036] The present invention overcomes the respective disadvantages of the prior art and improves the classification effect of 3D point cloud data. By combining the features extracted by the two models through the multi-modal feature fusion method, the features at different levels are considered, which promotes the classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0038] Figure 1 It is a flowchart of the method for classifying point cloud data based on multi-modal feature fusion according to the embodiments of the present invention;

[0039] Figure 2 It is a schematic structural diagram of the system for classifying point cloud data based on multi-modal feature fusion provided by the embodiments of the present invention;

[0040] Figure 3 It is a schematic structural diagram of the electronic device according to the embodiments of the present invention. Detailed implementation manners

[0041] In order to enable those skilled in the art of this technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions of the present invention in combination with the embodiments and the accompanying drawings in the present application. It should be understood that the accompanying drawings are only for illustrative purposes and cannot be construed as a limitation to this patent. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0042] Referring to "embodiments" in the present application means that the specific features, structures or characteristics described in combination with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0043] Compared with two-dimensional images, three-dimensional shapes usually have more complex structural information, which makes it difficult for a single modality to completely describe three-dimensional shapes. Although different modalities have different representations, their features should have strong correlations. The present invention uses the feature fusion of different modalities, but there are many differences in the dimensions between the features of different modalities. The encoder is used to reduce the dimension of the high-dimensional information, and the decoder is used to restore the dimension after fusion to reduce the reduction of feature information.

[0044] The multi-view convolutional neural network model MVCNN uses multi-view CNNs to fuse features generated from multiple 2D feature projections in an end-to-end manner. First, the 3D shape is projected into multiple views, and the multiple views are placed into a basic 2D image CNN. Each perspective image of the same 3D shape independently passes through the first-stage CNN1 convolutional network, and "aggregation" is performed in a layer called View-pooling. Then, it is sent into the remaining CNN2 convolutional network. All branches in the first part of the entire network share the same parameters in CNN1. In the View-pooling layer, an element-wise maximum operation is taken. Finally, Softmax is used for classification.

[0045] PCT mainly avoids defining the order of point cloud data by using the order-invariance inherent in the transformer and performs feature learning through the attention mechanism. First, self-attention takes the sum of the input word embeddings and position encodings as input, and calculates three vectors for each word through a trained linear layer: query, key, and value. Then, the attention weights between any two words can be obtained by matching (dot product) the query and key vectors. Finally, the attention feature is defined as the weighted sum of all value vectors and the attention weights. After obtaining the attention feature, a convolutional layer and softmax are used for classification.

[0046] Please refer to Figure 1 , which is a schematic flowchart of a point cloud data classification method based on multi-modal feature fusion provided by an embodiment of the present invention. The method includes the following steps:

[0047] S1. Image feature extraction;

[0048] In this step, a pre-established multi-view convolutional neural network model is used for image feature extraction. Specifically, given 12 projections at different positions to simulate the images obtained from the perspectives of 12 cameras (with an interval of 30° between each of the 12 perspectives and pointing downward from 30° above the plane towards the centroid of the grid), the VGGNet with shared weights is used to extract the features of the 12 perspective images respectively. Through the View Pooling operation on this feature map, the maximum value of all perspective results at each position is retained to obtain the feature F1 based on the image input. This feature map will serve as the first input to the multi-modal feature fusion module. Meanwhile, this feature passes through three fully connected layers, changing the dimension from 25088*1 to 4096*1, 4096*1, and finally 40*1. The outputs of these last three fully connected layers will respectively serve as the inputs of corresponding sizes for the input paths of the multi-view convolutional neural network model in the multi-modal feature fusion module.

[0049] S2. Point cloud feature extraction;

[0050] Use the pre-established point cloud Transformer model to perform point cloud feature extraction on the point cloud data; specifically, it includes the following steps:

[0051] S21. Input the point cloud data into an encoder composed of four layers of attention mechanisms; since the attention operation performs attention scoring pairwise, it can learn regardless of the order.

[0052] S22. Concatenate the results of the four-layer attention operation, and then obtain the features of the points through linear transformation (Linear), batch normalization (BatchNorm), non-linear activation (ReLU), and Dropout layers.

[0053] S33. After maximum pooling and average pooling and then concatenation, a 1024*1 global feature vector F2 is obtained. Similarly, this vector will serve as the second input to the feature fusion module. Meanwhile, this vector will also follow the subsequent operations of the original model, passing through fully connected layers to change from a 1024*1 vector to a 512*1 vector, then a 256*1 vector, and finally transformed into a 40*1 vector. These subsequent three fully connected layers will also be respectively input to the corresponding scale positions of the input path of the point cloud Transformer module in the feature fusion module.

[0054] S3. Multi-modal feature fusion;

[0055] In this step, the above two extracted features are fused in a multi-modal manner to jointly promote the final classification effect. Please refer to again Figure 1, the overall framework consists of two paths. The first path (the path for the input of the multi-view convolutional neural network model) takes the feature map obtained from the above multi-view convolutional model as input, and in the middle, the data of the original model and the intermediate result of the feature fusion module are superimposed using the corresponding size. The second path (the path for the input of the point cloud Transformer model) is the feature obtained from the above point cloud Transformer model. The output of the original model and the intermediate result are also superimposed at the corresponding scale position in the middle.

[0056] Furthermore, the specific steps for feature fusion are as follows:

[0057] S31. The feature map (dimension 512*7*7) extracted by the above multi-view convolutional neural network model is formed into a first vector of 1024*1 through an encoder 1. In this embodiment, the structure of the decoder is the same as that of the encoder, both are multi-layer fully connected layers, and the parameters are self-learned through data-driven. The purpose of dimensionality reduction here is that the feature dimension obtained by the multi-view convolutional model is relatively high, and the feature dimensions of the two different modalities are quite different, so dimensionality reduction is required through the encoder. The feature obtained from the point cloud Transformer is also a second vector of 1024*1.

[0058] Further, the first vector and the second vector obtained from these two paths are concatenated into a third vector of 2048*1. The third vector passes through the decoder 1 to obtain a fourth vector of 4096*1 dimension (the part of the path for the input of the multi-view convolutional neural network model) and a fifth vector of 512*1 dimension (the part of the path for the input of the point cloud Transformer model).

[0059] First, look at the multi-view convolutional input path. The reason for increasing the vector dimension of the multi-view convolutional input path is to reduce information loss and ensure consistent dimensions. After obtaining the fourth vector of 4096*1, the output of the first fully connected layer at the corresponding scale of the above multi-view convolutional neural network model is superimposed with this vector for the first time; then it passes through the encoder 2 to become a vector of 512*1 dimension. For the fifth vector of 512*1 in the point cloud Transformer input path, the output of the first fully connected layer at the corresponding scale of the point cloud Transformer model is also superimposed with this fifth vector for the first time.

[0060] S32. Next, similar to the vector concatenation of the first two paths, here two 512*1 vectors are also concatenated to form a 1024*1 sixth vector. The sixth vector then passes through a decoder 2 to form a 4096*1 seventh vector (the path part of the multi-view convolution input) and a 256*1 eighth vector (the path part of the point cloud Transformer input). For the path part of the multi-view convolution input, this seventh vector is still superimposed with the output vector of the second fully connected layer of the multi-view convolution neural network model of the corresponding dimension, and then passes through an encoder 3 to obtain a 256*1 ninth vector. For the path of the point cloud Transformer input, the outputs of the second fully connected layer of the same dimension of the point cloud Transformer model are also superimposed to form a 256*1 tenth vector.

[0061] S33. The two 256*1 vectors of the two paths are concatenated to form an 512*1 eleventh vector, which passes through a decoder 3 to form two 40*1 vectors (for the two paths respectively). The vectors of the two paths are respectively concatenated with the output vectors of the fully connected layers of their respective original models at the corresponding scales, and then uniformly pass through a fully connected layer to obtain a 40*1 vector.

[0062] S44. The final 40*1 vector of model one, the final 40*1 vector of model two, and the 40*1 vector obtained in the previous paragraph are used for the classification task. The overall process is as Figure 2 shown. By means of feature fusion, the advantages and disadvantages of the two original models are complementary, thereby improving the classification accuracy.

[0063] In another embodiment of the present invention, a point cloud data classification device based on multi-modal feature fusion will be introduced. For related content, please refer to the above method embodiment.

[0064] See Figure 2 , which is a schematic structural diagram of an image classification device based on continuous learning provided in this embodiment. The device includes: an image feature extraction module, a point cloud feature extraction module, and a multi-modal feature fusion module;

[0065] The image feature extraction module is used to extract image features through a pre-established multi-view convolution neural network model;

[0066] The point cloud feature extraction module is used to extract point cloud features from point cloud data through a pre-established point cloud Transformer model;

[0067] The multi-modal feature fusion module is used to perform multi-modal feature fusion on the image features and point cloud features, and obtain the classification result of the point cloud data according to the fused features; the feature fusion module includes a first path and a second path. The first path takes the feature map obtained from the multi-view convolutional neural network model as input, and superimposes the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module at the corresponding size in the middle. The second path is the feature obtained from the point cloud Transformer model, and also superimposes the output of the original point cloud Transformer model and the intermediate result at the corresponding scale position in the middle.

[0068] In the first possible implementation of this example, when the image feature extraction module extracts features:

[0069] Given the projections at K different positions of the input to simulate the images obtained from the perspectives of K cameras, the feature maps of the images at K perspectives are respectively extracted by the VGGNet with shared weights. Through the perspective pooling operation, the maximum value of all perspective results at each position is retained to obtain the features based on the image input.

[0070] Furthermore, the image features pass through three fully connected layers, changing the dimension from 25088*1 to 4096*1, 4096*1, and finally 40*1. The outputs of these last three fully connected layers are respectively used as the inputs at the corresponding sizes of the multi-view convolutional input path of the multi-modal feature fusion module.

[0071] In the second possible implementation of this example, when the point cloud feature extraction model extracts point cloud features:

[0072] First, the point cloud data is input into an encoder, which consists of four layers of attention mechanisms;

[0073] Then, the results of the four-layer attention operations are concatenated, and then passed through linear transformation, batch normalization, non-linear activation, and Dropout layers to obtain the features of the points;

[0074] Finally, through max pooling and average pooling and then concatenation, a global feature vector of 1024*1 is obtained, and the global feature vector will be used as the second input of the feature fusion module.

[0075] Furthermore, the global feature vector will follow the subsequent operations of the original model, passing through fully connected layers to change from a vector of 1024*1 to a vector of 512*1, then to a vector of 256*1, and finally transformed into a vector of 40*1. These subsequent three fully connected layers will also be respectively input to the corresponding scale positions of the point cloud Transformer input path of the feature fusion module.

[0076] In the third possible implementation of this example, the specific steps for the multi-modal feature fusion to perform multi-modal feature fusion are as follows:

[0077] The feature map extracted by the multi-view convolutional neural network model is reduced in dimension by the first encoder to form a first vector of 1024*1 dimension. The point cloud features extracted by the multi-view convolutional neural network model are used as a second vector of 1024*1 dimension. The first vector and the second vector are concatenated into a third vector of 2048*1 dimension;

[0078] The third vector passes through the first decoder to obtain a fourth vector of 4096*1 dimension and a fifth vector of 512*1 dimension. The fourth vector is used as the input of the first path, and the fifth vector is used as the input of the second path; In the first path, the fourth vector is first superimposed with the output of the corresponding scale fully connected layer of the multi-view convolutional neural network model, and then becomes a vector of 512*1 dimension after passing through the second encoder; In the second path, the fifth vector of 512*1 dimension is first superimposed with the output of the corresponding scale fully connected layer of the point cloud Transformer model;

[0079] The vectors after the first superimposition of the two paths are concatenated to form a sixth vector of 1024*1 dimension. The sixth vector is decoded by the second decoder into a seventh vector of 4096*1 dimension and an eighth vector of 256*1 dimension; In the first channel, the seventh vector is superimposed with the vector output by the second fully connected layer of the multi-view convolutional neural network model. The vector obtained after the superimposition passes through the third encoder to obtain a ninth vector of 256*1 dimension; In the second path, the eighth vector is superimposed with the output of the second fully connected layer of the same dimension of the point cloud Transformer model to obtain a tenth vector of 256*1 dimension;

[0080] The ninth vector and the tenth vector are concatenated to form an eleventh vector of 512*1 dimension. The eleventh vector passes through the third decoder to form two vectors of 40*1 dimension; The two vectors of 40*1 dimension are respectively concatenated with the outputs of the third fully connected layer of the multi-view convolutional neural network model and the third fully connected layer of the point cloud Transformer model, and then uniformly pass through a fully connected layer to obtain a vector of 40*1 dimension;

[0081] The final 40*1 dimension vector of the multi-view convolutional neural network model, the final 40*1 dimension vector of the point cloud Transformer model, and the 40*1 dimension vector obtained in the previous paragraph are used for the classification task.

[0082] It should be noted that the multi-modal feature fusion point cloud data classification device of the present invention corresponds one-to-one with the multi-modal feature fusion point cloud data classification method of the present invention. The technical features and their beneficial effects described in the embodiments of the above multi-modal feature fusion point cloud data classification method are applicable to the embodiments of the multi-modal feature fusion point cloud data classification device. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be elaborated here. This is hereby declared.

[0083] In addition, in the implementation manner of the multi-modal feature fusion point cloud data classification device in the above embodiments, the logical division of each program module is only for illustration. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the continuous learning-based image classification device is divided into different program modules to complete all or part of the functions described above.

[0084] As Figure 3 shown, in one embodiment, an electronic device for a multi-modal feature fusion point cloud data classification method is provided. The electronic device 300 may include a first processor 301, a first memory 302, and a bus, and may also include a computer program stored in the first memory 302 and executable on the first processor 301, such as a multi-modal feature fusion point cloud data classification program 303.

[0085] Among them, the first memory 302 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The first memory 302 may be an internal storage unit of the electronic device 300 in some embodiments, such as the mobile hard disk of the electronic device 300. The first memory 302 may also be an external storage device of the electronic device 300 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 300. Further, the first memory 302 may include both an internal storage unit and an external storage device of the electronic device 300. The first memory 302 can be used not only to store application software installed on the electronic device 300 and various types of data, such as the code of the multi-modal feature fusion point cloud data classification program 303, but also to temporarily store data that has been output or will be output.

[0086] In some embodiments, the first processor 301 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including the combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 301 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing programs or modules (such as the federated learning defense program, etc.) stored in the first memory 302, and calling the data stored in the first memory 302, to execute various functions of the electronic device 300 and process data.

[0087] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that, Figure 3 the shown structure does not constitute a limitation on the electronic device 300, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0088] The multi-modal feature fusion point cloud data classification program 303 stored in the first memory 302 of the electronic device 300 is a combination of multiple instructions. When running in the first processor 301, it can achieve:

[0089] Performing image feature extraction using a pre-established multi-view convolutional neural network model;

[0090] Performing point cloud feature extraction on the point cloud data using a pre-established point cloud Transformer model;

[0091] Performing multi-modal feature fusion on the image features and point cloud features using a feature fusion module, and obtaining a point cloud data classification result according to the fused features; the feature fusion module includes a first path and a second path. The first path inputs the feature map obtained from the multi-view convolutional neural network model, and superimposes the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module at the corresponding size in the middle; the second path is the feature obtained from the point cloud Transformer model, and also superimposes the output of the original point cloud Transformer model and the intermediate result at the corresponding scale position in the middle.

[0092] Furthermore, if the modules / units integrated in the electronic device 300 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0093] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memories (ROMs), programmable ROMs (PROMs), electrically programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), or flash memories. Volatile memories can include random access memories (RAMs) or external cache memories. By way of illustration and not limitation, RAMs are available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0094] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0095] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A point cloud data classification method based on multi-modal feature fusion, characterized in that Including the following steps: Using a pre-established multi-view convolutional neural network model for image feature extraction; Using a pre-established point cloud Transformer model for point cloud feature extraction of point cloud data; Performing multi-modal feature fusion on the image features and point cloud features using a feature fusion module, and obtaining a point cloud data classification result according to the fused features; the feature fusion module includes a first path and a second path, the first path takes the feature map obtained from the multi-view convolutional neural network model as input, and at the middle, the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module are superimposed using the corresponding size; the second path is the feature obtained from the point cloud Transformer model, and also at the corresponding scale position in the middle, the output of the original point cloud Transformer model and the intermediate result are superimposed; The specific steps for performing multi-modal feature fusion are as follows: Reducing the feature map extracted by the multi-view convolutional neural network model through a first encoder to form a first vector of n*1 dimension, taking the point cloud feature extracted by the multi-view convolutional neural network model as a second vector of n*1 dimension, and splicing the first vector and the second vector into a third vector of 2n*1 dimension; The third vector passes through a first decoder to obtain a fourth vector of 4n*1 dimension and a fifth vector of n / 2*1 dimension, the fourth vector is used as the input of the first path, and the fifth vector is used as the input of the second path; In the first path, the fourth vector is first superimposed with the output of the first fully connected layer of the multi-view convolutional neural network model at the corresponding scale, and then becomes a vector of n / 2*1 dimension through a second encoder; in the second path, the fifth vector of n / 2*1 dimension is first superimposed with the output of the first fully connected layer of the point cloud Transformer model at the corresponding scale; Splicing the vectors after the first superimposition of the two paths to form a sixth vector of n*1 dimension, the sixth vector is decoded by a second decoder into a seventh vector of 4n*1 dimension and an eighth vector of n / 4*1 dimension; in the first channel, the seventh vector is superimposed with the vector output by the second fully connected layer of the multi-view convolutional neural network model, and the vector obtained after the superimposition passes through a third encoder to obtain a ninth vector of n / 4*1 dimension; in the second path, the eighth vector is superimposed with the output of the second fully connected layer of the same dimension of the point cloud Transformer model to obtain a tenth vector of n / 4*1 dimension; Splicing the ninth vector and the tenth vector to form an eleventh vector of n / 2*1 dimension, the eleventh vector passes through a third decoder to form two vectors of a set dimension; splicing the two vectors of the set dimension with the outputs of the third fully connected layer of the multi-view convolutional neural network model and the third fully connected layer of the point cloud Transformer model respectively, and then uniformly passing through a fully connected layer to obtain a vector of the set dimension; Perform a classification task on the vectors of the last set dimension of the multi-view convolutional neural network model, the vectors of the last set dimension of the point cloud Transformer model, and the vectors of the set dimension obtained in the previous paragraph.

2. The method for classifying point cloud data based on multi-modal feature fusion according to claim 1, wherein When performing image feature extraction, given K projections at different positions of the input to simulate images obtained from the perspectives of K cameras at these positions, the feature maps of the images from K perspectives are respectively extracted using a VGGNet with shared weights. Through a perspective pooling operation on this feature map, the maximum value of all perspective results at each position is retained to obtain the features based on the image input.

3. The method for classifying point cloud data based on multi-modal feature fusion according to claim 2, wherein, K is selected as 12, with a 30° interval between each of the 12 perspectives, and pointing downward from 30° above the plane towards the centroid of the grid.

4. The point cloud data classification method based on multi-modal feature fusion according to claim 1, wherein The image features pass through three fully connected layers to reduce the dimension of the extracted feature map to the set dimension. The outputs of these last three fully connected layers are respectively used as inputs of corresponding sizes for the multi-view convolutional input path of the multi-modal feature fusion module.

5. The method for classifying point cloud data based on multi-modal feature fusion according to claim 1, wherein, When performing point cloud feature extraction, the specific steps are as follows: First, input the point cloud data into an encoder, which consists of four layers of attention mechanisms; Then, concatenate the results of the four-layer attention operations, and then obtain the features of the points through linear transformation, batch normalization, non-linear activation, and a Dropout layer; Finally, after max pooling and average pooling and then concatenation, obtain an n*1 global feature vector, which will be used as the second input of the feature fusion module.

6. The method for classifying point cloud data based on multi-modal feature fusion according to claim 5, wherein, The global feature vector will follow the subsequent operations of the original model. Through fully connected layers, it changes from an n*1 vector to an n / 2*1 vector, then an n / 4*1 vector, and finally is transformed into a vector of the set dimension. The outputs of these subsequent three fully connected layers will also be respectively input to the corresponding scale positions of the point cloud Transformer input path of the feature fusion module.

7. A point cloud data classification device based on multi-modal feature fusion, characterized in that, Applied to the multi-modal feature fusion point cloud data classification method according to any one of claims 1-6, characterized in that it includes an image feature extraction module, a point cloud feature extraction module, and a multi-modal feature fusion module; The image feature extraction module is used to extract image features through a pre-established multi-view convolutional neural network model; The point cloud feature extraction module is used to extract point cloud features from point cloud data through a pre-established point cloud Transformer model; The multi-modal feature fusion module is used to perform multi-modal feature fusion on the image features and the point cloud features, and obtain the point cloud data classification result according to the fused features; the feature fusion module includes a first path and a second path. The first path takes the feature map obtained from the multi-view convolutional neural network model as input, and uses the corresponding size in the middle to superimpose the data of the original multi-view convolutional neural network model and the intermediate result of the feature fusion module; the second path is the feature obtained from the point cloud Transformer model, and also superimposes the output of the original point cloud Transformer model and the intermediate result at the corresponding scale position in the middle.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the multi-modal feature fusion point cloud data classification method according to any one of claims 1-6.

9. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, the multi-modal feature fusion point cloud data classification method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Deep multi-mode cross-layer cross-fusion method, terminal equipment and storage medium

    CN111860425A

  • RGB-D image semantic segmentation method based on extensible 2.5 D convolution and two-way gate fusion

    CN113850262A