3D Object Recognition Method Based on Multi-View Pooling Transformer
Through the Multi-view Pooling Transformer network model, the optimal view set parallel training is built, which solves the problems of multi-view redundancy and information loss, and achieves efficient and accurate 3D object recognition.
Patent Information
- Application Number
- CN202210671530.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-15
AI Technical Summary
The existing 3D object recognition method based on multi-view has problems such as long training time, high redundancy and information loss, which makes it difficult to further improve the recognition accuracy.
The Multi-view Pooling Transformer network model is used to build the best view set through information entropy, and the low-level local features of multi-view are extracted using ResNet and Embedding networks. The parallel training is combined with Pooling Transformer to generate compact 3D global descriptors.
It improves the recognition accuracy and training efficiency of the network model, effectively captures relevant feature information between multiple views, reduces redundancy, and improves the recognition accuracy and training speed.
Smart Images

Figure CN114972794B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of 3D object recognition, and particularly relates to a three-dimensional object recognition method based on a multi-view Pooling Transformer. Background Art
[0002] With the rapid development of 3D acquisition technologies, sensors such as 3D scanners, depth scanners, and 3D cameras have become popular and inexpensive, and the acquisition of 3D data such as point clouds and meshes has become more convenient and accurate. These factors have promoted the wide application of object recognition technologies based on 3D data in fields such as environmental perception in autonomous driving, grasping recognition in robots, and scene understanding in augmented reality. Therefore, 3D object recognition has become a current research hotspot.
[0003] Currently, deep learning-based methods have become the mainstream research technologies for 3D object recognition tasks. Generally speaking, these methods can be divided into three categories according to the types of data input into the deep neural network: voxel-based methods [1-8], point cloud-based methods [9-16], and multi-view-based methods [17-26].
[0004] Voxel-based methods: Voxelize the 3D object represented by the point cloud, and then use a 3D convolutional neural network (3DCNN) to learn the features of the 3D object from voxels of a fixed size. Daniel Maturana et al. proposed VoxNet [1], which uses 3D CNN to extract the features of the voxelized 3D object and processes non-overlapping voxels through max pooling, but it cannot automatically learn the shape distribution information of the 3D object. For this reason, Zhirong Wu et al. proposed 3D ShapeNets [2], which represents the 3D object as a probability distribution of binary variables on a 3D voxel grid and extracts features, so as to automatically discover the part representation of hierarchically composed 3D objects. VoxNet and 3D ShapeNets voxelize 3D objects, which can solve the problem of point cloud unstructuredness to a certain extent, but there are still problems such as the increase in computational cost with the increase in resolution and the inability to compactly represent the structure of 3D objects. Therefore, some works have studied the representation of the 3D object voxelization structure. Among them, OctNet [3] hierarchically divides 3D space into a set of unbalanced octrees by utilizing the sparsity of the input 3D data. Each leaf node of the octree stores a pooled feature representation. This structure well considers the global features of 3D objects, but its ability to process high-resolution voxelized 3D objects needs to be strengthened. Pengshuai Wang et al. proposed the Octree-based Convolutional Neural Network (O-CNN) [4] by restricting the calculation of CNN and the octants occupied by features on the surface of 3D objects, which can effectively analyze high-resolution 3D objects. Although replacing voxels with a fixed resolution with a flexible octree structure can reduce the memory occupancy of 3D object representation, the computational cost is still too high when traversing the octree from the root node in the case of high resolution. Therefore, there is still room for optimization in this tree representation structure. Therefore, Kd-network [5] was proposed. It creates a structure diagram of 3D objects based on the Kd-tree structure and shares learnable parameters when doing transformations, and calculates a sequence of hierarchical representations in a feed-forward bottom-up manner. In this way, its network structure has less memory occupancy and higher computational efficiency, but it does not consider the relationship between global and local background information. 3D ContextNet [6] proposed by Wei Zeng et al. explored this relationship. This network gradually learns representation vectors along the tree structure, identifies local patterns by using adaptive features, and calculates the global pattern as the non-local response of different regions at the same level, thus capturing the correlation between local and global information. Although these methods using Kd-tree can reduce the memory occupancy during calculation based on the indexing and structuring ability of Kd-tree, they may lose the information of local geometric structures.PointGrid [7] proposed by Truc Le et al. can solve this problem. It samples a fixed number of points using point quantization operations within each grid cell, thereby learning a higher-order local approximation function to avoid the loss of local information. However, its ability to adapt to the sparsity of voxel grids still needs to be improved. The VV-NET [8] network considers enhancing the capture of sparse distributions in voxels. It uses an interpolation variational autoencoder structure to encode the local geometry within voxels and then uses radial basis functions to calculate local continuous representations. Although these voxel-based methods well solve the problems of high memory occupancy and long training time in point cloud voxelization, there are still inevitable problems in this process, that is, information loss may occur when the 3D object is voxelized at a low resolution, and high resolution always leads to high computational costs.
[0005] Point cloud-based methods: Since there are inevitable information loss problems during the voxelization process of point clouds, and this lost information may be very important. Some methods consider directly processing point clouds efficiently and accurately without voxelization to complete subsequent object recognition tasks. These methods can be classified into one category, namely point cloud-based methods. This category of methods can be further subdivided into methods based on neighborhood feature pooling, methods based on attention mechanisms, and methods based on graph neural networks. The first category is methods based on neighborhood feature pooling: PointNet [9] proposed by Charles R. Qi et al. is the earliest method to directly process point clouds. It uses T-Net to perform an affine transformation on the input point matrix and extracts per-point features through a multi-layer perceptron (MLP). This can solve the problems of the disorder and permutation invariance of point cloud data, but it cannot capture the local neighborhood information between points. Thus, PointNet++
[10] was further proposed. It constructs local neighborhood subsets by introducing a hierarchical neural network and then extracts local neighborhood features based on PointNet. PointNet++ solves the problem of local neighborhood information extraction of PointNet to a certain extent, but it does not have the ability to achieve both orientation awareness and scale awareness at the same time. In response to this, Mingyang Jiang et al. proposed the PointSIFT
[11] network. It integrates the information of the oriented direction by using orientation-encoded convolution (OEC) and realizes multi-scale representation by stacking orientation-encoded units, but it still cannot adaptively find the connection between points with features. Connect the points in the local neighborhood densely to accurately represent this area and use adaptive feature adjustment (AFA) to construct a local network. Thus, PointWeb
[12] can learn point features from the differences between point pairs, but loses the correlation between global features. The second category is methods based on attention mechanisms: In life, people selectively focus their attention on certain parts of the visual space, which can strengthen the mutual dependence of different visual positions and thus can efficiently learn effective features. Therefore, the attention mechanism has become the direction considered by researchers. The Dual Attention Network (DANet)
[13] adopts dual attention. Its position attention module selectively aggregates local neighborhood features, and the channel attention module integrates the associated features between global channel maps and fuses the outputs of the two attention modules, thereby enhancing the ability of feature representation, but the generalization of its local geometric features in space is still insufficient. For this reason, Mingtao Feng et al. proposed LocalAttention-Edge Convolution (LEA-Conv)
[14] . It constructs a local feature map based on a multi-directional search strategy, then assigns attention coefficients to each edge of the map, and aggregates the central point features as the weighted sum of its adjacent nodes to obtain more fine-grained local geometric features.The third category is the method based on graph neural network: This type of method converts the point cloud into a k-nearest neighbor graph or an overlay graph, and uses the graph theory evolution network to explore the topological structure, which can effectively capture the local geometric structure while maintaining permutation invariance. Dynamic Graph CNN (DGCNN)
[15] constructs a dynamic graph convolutional neural network through EdgeConv for object recognition. EdgeConv can extract local neighborhood feature information, and the features of the local shape of the extracted point cloud can maintain permutation invariance. However, its deep features and neighborhoods may be too similar to provide valuable edge vectors. Linked Dynamic Graph CNN (LDGCNN)
[16] further improves on the basis of DGCNN, combines the hierarchical features of different dynamic graphs, and uses the current index to obtain useful edge vectors from the previous features to learn new features, thereby improving the recognition accuracy. Although the method based on point cloud can directly process the point cloud to reduce information loss, its network model is often very complex, the training time is relatively long, and the final recognition accuracy is not high enough.
[0006] Multi-view based methods render 3D data objects into multiple 2D views. In this way, it no longer needs to rely on complex 3D features, but inputs the rendered multi-views into a mature 2D image classification network to extract efficient and accurate features for object recognition. Especially for the case where 3D objects are occluded, this type of method captures the detailed features of 3D objects that can complement each other according to the views from different viewpoints. Compared with voxel-based methods and point cloud-based methods, this type of method has achieved the best 3D object recognition accuracy currently. Hang Su et al. first proposed Multi-view Convolutional Neural Networks (MVCNN)
[17] . It uses a 2D CNN network to process the rendered multi-views separately, and then combines the information of multiple views into a single and compact shape descriptor through view pooling. However, since it performs pooling on all views, some position information of the viewpoints will be lost. Some methods have started to consider grouping multi-view features to strengthen the capture of multi-view position information. RCPCNN
[18] aggregates information from grouped similar views, and then feeds the aggregated feature vectors into the same layer in a cyclic manner. It captures the information between similar views to a certain extent, but does not consider the distinctiveness of different views. By grouping the view-level descriptors extracted by CNN under different viewpoints and aggregating the features according to their discriminant weights in groups, Multi-view Convolutional Neural Networks (GVCNN)
[19] considers both intra-group similarity and inter-group distinctiveness between views, but it needs to consider all views for inference. By treating the viewpoint labels as latent vectors and training and learning in an unsupervised manner, RotationNet
[20] can also obtain good recognition performance with only a few views, but its limitation is that there will be information loss when processing views separately. Equivariant Multi-View Networks (EMV) proposed by Carlos Esteves et al.
[21] solves this problem. It performs convolutions on discrete subgroups of the rotation group, so it can jointly infer all views in an equivariant manner, but its network model may be a bit complex. Since the attention mechanism can flexibly capture the connection between global and local features, thus optimizing the network model structure, some research works have started to consider adding the attention mechanism. Zhizhong Han proposed 3D to Sequential Views (3D2SeqViews)
[22] . It encodes the content information of each view, aggregates features with hierarchical attention, and simultaneously aggregates the content information of the encoded views and the sequential spatiality between views to strengthen the distinctiveness of the learned features, but it can only aggregate sequential views and is not applicable to unordered views.View N-gram Network (View-gram)
[23] divides the view sequence into a set of visual n-grams, which can capture the spatial information across multiple views and help learn discriminative global embeddings for each 3D object. However, the information of its single-viewpoint image is lost. The Relation Network (RN)
[24] can strengthen the information of single-view images and consider the region-to-region and view-to-view relationships between different views. Since it uses a relation network to effectively connect the corresponding regions from different viewpoints and utilizes the mutual relationships on a set of views, it cannot flexibly simulate different view configurations. Songle Chen et al. believe that aggregating views by using an RNN to select views after treating multiple views as a sequence can also well consider the connections between views. For this, they proposed the View-Enhanced Recurrent Attention Model (VERAM)
[25] network, which is a view-enhanced recurrent attention model. It conducts reinforcement training for view estimation by designing a reward function and can actively select view sequences for high-precision 3D object recognition. However, it cannot perform local feature fusion by adaptively calculating the weights of features. On this basis, Hierarchical multi-view context modelling (HMVCM)
[26] aggregates features into a compact 3D object descriptor by adaptively calculating the weights of features. This is a hierarchical multi-view context modelling method. After using a module that combines a convolutional neural network (CNN) and a bidirectional long short-term memory (Bi-LSTM) network to learn the visual context features of a single view and its neighborhood, it finally achieved an overall recognition accuracy of 94.57% on the ModelNet 40 dataset. However, it cannot consider the local features of all views in parallel during training, losing the relevant information between views. As a result, the aggregated global descriptor is not compact enough, so there is still room for further improvement in its 3D object recognition accuracy.
[0007] The above method based on multi - views has the local feature information of the original 3D object retained in its rendered multiple 2D views. Here, the views from different viewpoints can also complement the detailed features of the 3D object. After subsequent fusion processing, the recognition accuracy of the 3D object can be greatly improved. Therefore, compared with the methods based on voxels and point clouds, the method based on multi - views has higher recognition accuracy. However, there are still problems with this type of method at present, such as the inability to extract feature information from all views at one time during training, the inability to efficiently capture the relevant feature information between multiple views, and the redundancy problem when rendering the 3D object into multi - views. The relevant feature information between multi - views is indispensable for finally aggregating the local features of multi - views into a compact global descriptor. The partial omission of them is the main reason for the difficulty in further improving the recognition accuracy of this type of method. And the view redundancy problem will increase the unnecessary training time of the network model and further affect the final recognition accuracy. Summary of the Invention
[0008] To solve the above problems of multi - view redundancy, complex model training, and omission of feature information, a 3D object recognition method based on multi - views is provided. The present invention adopts the following technical solutions:
[0009] The present invention provides a 3D object recognition method based on Multi - view Pooling Transformer, which is characterized by including the following steps: Step S1, construct a Multi - view Pooling Transformer network model, which has an optimal view set acquisition module, a low - level local feature token sequence generation module, a global descriptor generation module based on Pooling Transformer, and a classifier; Step S2, input the object to be measured into the MVPT model, obtain the corresponding multi - views through the optimal view set acquisition module, and construct an optimal view set according to the information entropy of the multi - views; Step S3, the low - level local feature token sequence generation module extracts the low - level local features of the multi - views in the optimal view set, and generates a corresponding multi - view low - level local feature token sequence based on the low - level local features of the multi - views; Step S4, the global descriptor generation module aggregates the local view information token sequence of the multi - view low - level local feature token sequence and its global feature information sequence to generate a 3D global descriptor of the object to be measured; Step S5, the classifier uses the 3D global descriptor as the input for 3D object recognition, so as to obtain the recognition result of the object to be measured.
[0010] The 3D object recognition method based on multi-view Pooling Transformer provided by the present invention may further have the following technical features. Among them, step S2 includes the following sub-steps: Step S2-1, obtaining a plurality of corresponding 2D views of the object to be measured according to the dodecahedron viewpoints; Step S2-2, calculating the information entropy of each 2D view and sorting them according to the magnitude of the information entropy value; Step S2-3, selecting the views ranked top n in terms of information entropy as the best view set, thereby reducing redundant views.
[0011] The 3D object recognition method based on multi-view Pooling Transformer provided by the present invention may further have the following technical features. Among them, the calculation formula of information entropy is:
[0012]
[0013] P a,b = f(a, b) / W·H
[0014] In the formula, H i represents the information entropy of the i-th view v i . (a, b) is a binary tuple, a represents the gray value at the center within a certain sliding window, and b is the average gray value of the pixels except the center pixel within this window; P a,b represents the probability that (a, b) appears in the entire view v i ; f(a, b) represents the number of times the binary tuple (a, b) appears in the entire view v i ; W and H represent the width and height of the view v i .
[0015] The 3D object recognition method based on multi-view Pooling Transformer provided by the present invention may further have the following technical features. Among them, the low-level local feature token sequence generation module has a ResNet network and an Embedding network. Step S3 includes the following sub-steps: Step S3-1, extracting multi-view low-level local features of the best view set by the ResNet network; Step S3-2, generating local view token sequences of the multi-view low-level local features based on the Embedding network:
[0016] [x1,...x i ...,x n = Emb{Res[v1,...v i ...,v n}
[0017] In the formula, [v i ,…v i …,v n is the best view set, vi represent one of the views; Step S3-3, add an initialized class token x class to the head of the local view token sequence, and splice them with the position encoding E pos respectively, and finally generate a multi-view low-level local feature token sequence:
[0018]
[0019] wherein, X0 is the multi-view low-level local feature token sequence, x class is a randomly initialized value matching the dimension of the local view token sequence, and E pos is used to save the position information from different viewpoints x i of.
[0020] The 3D object recognition method based on multi-view Pooling Transformer provided by the present invention may also have the following technical features. Among them, the global descriptor generation module includes a global feature information generation sub-module based on Transformer and a local view information token sequence aggregation sub-module based on Pooling. The global feature information generation sub-module based on Transformer has a Layer Normalization network, a Multi-Head Multi-View Attention network, a multi-layer perceptron network, and a residual connection.
[0021] The 3D object recognition method based on multi-view Pooling Transformer provided by the present invention may also have the following technical features. Among them, Step S5 includes the following sub-steps: Step S4-1, the Layer Normalization network normalizes the multi-view low-level local feature token sequence: Step S4-2, the Multi-Head Multi-View Attention network completes the MHMVA calculation on the normalized token sequence through linear transformation to generate a token sequence X MHMVA ; Step S4-3, use the residual connection for the token sequence X MHMVA to obtain a token sequence X1 to avoid gradient disappearance, and then input X1 into the Layer Normalization network for normalization processing and then input it into the multi-layer perceptron network; Step S4-4, perform a residual connection on the output result of the multi-layer perceptron network and X1 to obtain a local view information token sequence: wherein, the local view information token sequence is composed of the global class token and the local view information token sequence which consists of a global class token that stores the global feature information of the local view token sequence, that is In step S4-5, the local view information token sequence aggregation sub-module based on Pooling aggregates the local view information token sequence to perform pooling processing to obtain a single best local view information token, and then concatenate and aggregate this best local view information token with the global class token to finally generate the corresponding 3D global descriptor Y:
[0022] The 3D object recognition method based on the multi-view Pooling Transformer provided by the present invention may also have the following technical features. Among them, the Multi-Head Multi-View Attention network consists of multiple Multi-View Attention, and the MHMVA calculation is to perform multiple parallelized Multi-View Attention calculations: In step S4-2-1, the after normalization is first linearly transformed to generate three vectors: Query, Key, and Value
[0023]
[0024] In step S4-2-2, according to the number of heads N, the three vectors in the previous step are evenly divided into the inputs q i 、k i 、v i of multiple Multi-View Attention, which can form multiple subspaces to focus on different parts of the input feature information. Finally, these feature information are concatenated to obtain richer information:
[0025]
[0026] In step S4-2-3, Multi-View Attention calculates MVA according to the input, that is, calculates the product of q i and the transpose of k i to obtain a score, divides it by for normalization to stabilize the gradient, and then uses the normalized result value as the input of the softmax function. The output of this softmax function is dot-multiplied with v i to obtain
[0027]
[0028] where d k is the dimension of k i ; Step S4-2-4, for each of the calculated perform Concat, and then finally complete the MHMVA calculation through a linear transformation:
[0029]
[0030] Functions and effects of the invention
[0031] According to the three-dimensional object recognition method based on multi-view Pooling Transformer of the present invention, this method constructs a Multi-view Pooling Transformer network model, which has an optimal view set acquisition module, a low-level local feature token sequence generation module, a global descriptor generation module based on Pooling Transformer, and a classifier. First, an optimal view set is constructed based on the information entropy of the multi-views of the object to be measured, thereby reducing the redundancy of the multi-views and improving the accuracy of the network model for recognition. Secondly, the ResNet network and the Embedding network are used to extract feature information from all views at one time, and a multi-view low-level local feature token sequence of the optimal view set is obtained, so that it can be input into the Pooling Transformer to complete parallel training. Then, through the Pooling Transformer, the local view information token sequences of the multi-view low-level local feature token sequences are realized, and the multi-view low-level local feature token sequences are aggregated from the global and local respectively into a compact and single 3D global descriptor. Finally, the classifier recognizes the 3D global descriptor to obtain the recognition result of the object to be measured.
[0032] The three-dimensional object recognition method based on multi-view Pooling Transformer of the present invention can efficiently and accurately capture the relevant feature information between multiple views, and greatly improve the recognition accuracy and training efficiency of the network model. Description of the drawings
[0033] Figure 1 is a schematic flowchart of the three-dimensional object recognition method based on multi-view Pooling Transformer in an embodiment of the present invention;
[0034] Figure 2 is a schematic structural diagram of the Multi-view Pooling Transformer network model in an embodiment of the present invention;
[0035] Figure 3 It is a schematic diagram of the viewpoint setting of the regular dodecahedron camera in the embodiment of the present invention;
[0036] Figure 4 It is a schematic diagram of the global descriptor generation module in the embodiment of the present invention;
[0037] Figure 5 It is a schematic structural diagram of the Multi-Head Multi-View Attention network in the embodiment of the present invention;
[0038] Figure 6 It is a schematic structural diagram of the Multi-View Attention network in the embodiment of the present invention;
[0039] Figure 7 It is a schematic diagram of some category objects in the dataset ModelNet40 in the embodiment of the present invention;
[0040] Figure 8 It is a schematic diagram of the construction process of the optimal view set based on information entropy in the embodiment of the present invention. Detailed implementation manners
[0041] In order to improve the recognition accuracy of the current multi-view based 3D object recognition method and reduce the training time of the network model, the present invention proposes a Multi-view Pooling Transformer (abbreviated as MVPT) network framework based on the Transformer model, pooling technology, and information entropy calculation. The MVPT network constructs an optimal view set based on information entropy to reduce the redundancy of multi-views, and extracts the optimal view set as the multi-view low-level local feature token sequence, which is then input into the Pooling Transformer for parallel training. This method extracts feature information from all views at once, thereby efficiently capturing the relevant feature information between multiple views and greatly improving the recognition accuracy and training efficiency of the network model.
[0042] In order to make the technical means, creative features, achieved purposes, and effects realized by the present invention easy to understand, the following specifically elaborates on the 3D object recognition method based on multi-view Pooling Transformer of the present invention in combination with embodiments and drawings.
[0043] <Embodiment>
[0044] Figure 1 It is a schematic flowchart of the 3D object recognition method based on multi-view Pooling Transformer in the embodiment of the present invention.
[0045] As Figure 1As shown in the figure, the three-dimensional object recognition method based on the multi-view Pooling Transformer includes the following steps:
[0046] Step S1, construct a Multi-view Pooling Transformer, that is, a multi-view pooling Transformer network model.
[0047] Figure 2 It is a schematic structural diagram of the Multi-view Pooling Transformer network model in an embodiment of the present invention.
[0048] As Figure 2 shown, the Multi-view Pooling Transformer model has an optimal view set acquisition module, a low-level local feature token sequence generation module, a global descriptor generation module based on Pooling Transformer, and a classifier.
[0049] Step S2, input the object to be measured into the MVPT network model, obtain the corresponding multi-views through the optimal view set acquisition module, and construct an optimal view set according to the information entropy of the multi-views.
[0050] In the input part of the MVPT network, a 3D object represented by point cloud or mesh can be rendered into multiple 2D views. In this embodiment, the mesh representation form with higher 3D object recognition accuracy is selected. Of course, a 3D object in point cloud form can also be reconstructed into mesh form.
[0051] Since the multi-views obtained by common multi-view rendering methods often have the problem of redundancy, which leads to an unnecessary increase in the training time of the network model. Therefore, in this embodiment, a method for constructing an optimal view set based on information entropy is proposed by calculating the information entropy of 2D views. Specifically:
[0052] For a 3D object O, different 2D rendering views V = {v1,... v i ..., v N}, [v1,... v i ..., v N = Render(O) can be obtained by setting camera viewpoints at different positions, where v i represents the view obtained from the i-th viewpoint.
[0053] Figure 3 It is a schematic diagram of the regular dodecahedron camera viewpoint setting in an embodiment of the present invention.
[0054] In this embodiment, a regular dodecahedron camera viewpoint setting is used: place the 3D object at the center of the regular dodecahedron, and then set the camera viewpoints according to the vertices of the regular dodecahedron, asFigure 3 As shown. This method uses a regular dodecahedron for viewpoint setting. Since the number of vertices N of a regular dodecahedron is 20, each vertex is a camera viewpoint. Such a viewpoint setting can make the camera viewpoints be evenly distributed in 3D space, so as to capture as much global spatial information of the 3D object as possible, thereby reducing information loss.
[0055] By observing the 20 2D views rendered by this camera viewpoint setting, although they evenly cover all parts of the 3D object, there will always be overlapping parts between the views. This may lead to redundancy in the features extracted by the deep neural network and increase the training time of the network model, ultimately resulting in a decrease in the recognition accuracy of the 3D object.
[0056] To address the problem of redundancy in the regular dodecahedron camera viewpoint setting, this embodiment uses the information entropy of the 2D views as an evaluation criterion to construct an optimal view set to reduce redundant views. Information entropy can highlight the comprehensive characteristics of the gray information of pixel positions and the gray distribution within the pixel neighborhood on the premise of the amount of information contained in the view. Therefore, information entropy can be used as an effective means to evaluate the quality of views.
[0057] First, calculate the information entropy of the 2D views N (N = 20):
[0058]
[0059] P a,b = f(a, b) / W·H
[0060] In the formula, H i represents the information entropy of the i-th view v i . (a, b) is a binary tuple, a represents the gray value at the center within a certain sliding window, and b is the average gray value of the pixels except the central pixel within this window; P a,b represents the probability of the occurrence of (a, b) in the entire view v i ; f(a, b) represents the number of occurrences of the binary tuple (a, b) in the entire view v i ; W and H represent the width and height of the view v i .
[0061] Then, sort the information entropy values H i (i = 1,..., N, N = 20) in descending order.
[0062] Finally, take the views with the top n information entropy rankings (n < N, n = 6 in this embodiment) as the optimal view set V = {v1,...v i ..., v n}. When n = 1 in the optimal view set, that is, select the single view with the highest information entropy value, which is called the optimal view.
[0063] After this processing process, the MVPT network model no longer needs to rely on complex 3D object features, and can use a mature 2D image classification network to extract efficient and accurate low-level view features, thereby optimizing the complexity of the network model. In addition, if the 3D object collected is occluded, the 2D views from different camera viewpoints can also complement each other with the detailed features of the 3D object, thereby improving the 3D object recognition accuracy of the network model.
[0064] Step S3: The multi-view low-level local feature generation module extracts the multi-view low-level local features of the best view set, and generates the corresponding multi-view low-level local feature token sequence based on the multi-view low-level local features.
[0065] Since Transformer was proposed in natural language processing tasks, its input requirement is a two-dimensional matrix sequence. And the best view set V = {v1,...v i ...,v n}, the dimension of v i is Bn×C×H×W (where B is the batch size, n is the number of views, C is the number of channels, H is the height of the picture, and W is the width of the picture). Therefore, the obtained views cannot be directly input into Transformer for processing, and it is also necessary to extract the low-level local features of each view v i and flatten them into a local view token sequence X = {x1,...x i ...,x n}. x i represents the local view token generated by the i-th view, and its dimension is Bn×D (where D is the dimension after feature extraction and Embedding of view v i ).
[0066] Among them, for view low-level feature extraction, any mature 2D image classification network can be used, such as the ResNet series network. Residual Network (ResNet) first introduced the concept of residual connection to solve problems such as gradient disappearance and information loss, enabling deeper networks to be well trained. This network has been widely used in fields such as image classification or as a backbone network to complete computer vision tasks, and common ones include 18-layer, 34-layer, and 50-layer. Specifically:
[0067] First, the ResNet34 network extracts the multi-view low-level local features of the best view set. Among them, there are 34 layers in the ResNet34 network, and the last fully connected layer is removed after fine-tuning.
[0068] Then, perform an Embedding operation on the multi-view low-level local features to generate a local view token sequence X = {x1,...x i ...,x n}:
[0069] [x1,...x i ...,x n = Emb{Res[v1,...v i ...,v n}.
[0070] Finally, add an initialized class token x class to the head of the local view token sequence, and concatenate them with the position encoding E pos respectively, finally generating a multi-view low-level local feature token sequence:
[0071]
[0072] wherein, X0 is the multi-view low-level local feature token sequence, x class is a randomly initialized value matching the dimension of the local view token sequence, and E pos is used to save the position information from different viewpoints x i .
[0073] Step S4, the global descriptor generation module aggregates the local view information token sequence of the multi-view low-level local feature token sequence with its global feature information sequence to generate a compact and single 3D global descriptor.
[0074] Figure 4 is a schematic diagram of the global descriptor generation module in the embodiment of the present invention.
[0075] As Figure 4 shown, the global descriptor generation module includes a Transformer-based global feature information generation sub-module and a Pooling-based local view information token sequence aggregation sub-module. Among them, the Transformer-based global feature information generation sub-module has a Layer Normalization network, a Multi-Head Multi-View Attention network, a multi-layer perceptron network, and a residual connection.
[0076] This step S4 includes the following sub-steps:
[0077] Step S4-1, use the Layer Normalization network to normalize the multi-view low-level local feature token sequence:
[0078] Step S4-2, based on the Multi-Head Multi-View Attention network, the normalized token sequence completes the MHMVA calculation through linear transformation to generate the token sequence X MHMVA .
[0079] Figure 5 is the schematic structural diagram of the Multi-Head Multi-View Attention network in the embodiment of the present invention, and Figure 6 is the schematic structural diagram of the Multi-View Attention network in the embodiment of the present invention.
[0080] As Figure 5 and Figure 6 shown, the Multi-Head Multi-View Attention network is composed of multiple Multi-View Attention, and performing multiple parallel Multi-View Attention calculations is the MHMVA calculation. Specifically:
[0081] Step S4-2-1, the MHMVA calculation requires three vectors of Query, Key, and Value. Therefore, first, the after normalization is generated into three vectors of Query, Key, and Value through linear transformation:
[0082]
[0083] Step S4-2-2, according to the number of Heads N, the three vectors in the previous step are evenly divided into the inputs q i , k i , v i of multiple Multi-View Attention to form multiple subspaces, so that the Multi-Head Multi-View Attention network can focus on different parts of the input features, and splicing these feature information can obtain richer information:
[0084]
[0085] Step S4-2-3, the Multi-View Attention calculates MVA according to the input, that is, calculates the product of q i and the transpose of k i to obtain a score, and divides it by Normalization is performed to stabilize the gradient, and the normalized result value is used as the input of the softmax function. The output of the softmax function is dot-multiplied with v i to obtain
[0086]
[0087] where d k is the dimension of k i .
[0088] Step S4-2-4, for each calculated , Concat is performed, and then after a linear transformation, the MHMVA calculation is finally completed:
[0089]
[0090] Step S4-3, for the token sequence X MHMVA , a residual connection is used to obtain the token sequence X1 to avoid gradient disappearance: X1 = X MHMVA + X0, and then X1 is input into the LayerNormalization network for normalization and then input into the multi-layer perceptron network.
[0091] Since the Multi-Head Multi-View Attention network has insufficient fitting ability for complex processes, in this embodiment, a multi-layer perceptron MLP is added behind it to enhance the generalization ability of the model. The MLP consists of Linear layers and uses the GELU activation function:
[0092] MLP(X) = GELU(XW1 + b1)W2 + b2
[0093] where W1 and b1 are the weights of the first fully connected layer, W2 and b2 are the weights of the second fully connected layer, and X represents the input feature information.
[0094] Step S4-4, the output result of the multi-layer perceptron is connected with X1 through a residual connection to obtain the local view information token sequence:
[0095]
[0096] where the local view information token sequence consists of the global class token and the local view information token sequence
[0097] After the parallel calculation of MHMVA, the global class token preserves the global feature information of the local view token sequence, but there is a problem of losing the single best local view information token. This part of the information is very effective for aggregating into a 3D global descriptor. For this reason, this embodiment proposes a method for aggregating the local view information token sequence based on Pooling. This method can capture the single best local view information token while retaining the global feature information of the local view token sequence. Specifically:
[0098] Step S4-5, the sub-module for aggregating the local view information token sequence based on Pooling processes the local view information token sequence through pooling to obtain a single best local view information token, and then concatenates and aggregates this best local view information token with the global class token Through these processing steps, we can achieve aggregating the multi-view low-level local feature token sequences from both local and global aspects, and finally generate a more compact 3D global descriptor Y:
[0099] Step S5, the classifier uses the 3D global descriptor Y as the input for 3D object recognition, thereby obtaining the recognition result of the object to be measured.
[0100] In this embodiment, in order to evaluate the performance of the MVPT network model, multiple comparative experiments are carried out on the widely used 3D object recognition dataset ModelNet40. ModelNet40 is widely popular due to its advantages such as diverse categories, clean shapes, and good construction. It consists of 40 categories (such as airplanes, cars, plants, lights), a total of 12,311 CAD models, including 9,843 training samples and 2,468 test samples. The composition of ModelNet40 is as Figure 7 shown.
[0101] In this embodiment, multiple representative 3D object recognition methods are selected for comparative experiments under the same experimental environment settings, and the MVPT network model proposed in this embodiment is quantitatively analyzed. The overall recognition accuracy (OA), average recognition accuracy (AA), and the training time of the entire network model are used as evaluation indicators.
[0102] Among them, the overall recognition accuracy OA represents the ratio of the number of correctly recognized samples in all categories to the total number of samples, and the calculation formula is as follows:
[0103]
[0104] where N is the total number of samples, and x ii is the number of correctly identified samples distributed along the diagonal of the confusion matrix, and C represents the number of classes.
[0105] The average recognition accuracy AA represents the average of the ratios of the number of correctly identified samples for each class to the total number of samples, and its calculation formula is as follows:
[0106]
[0107] where recall represents the ratio of the number of correctly identified samples for each class to the total number of samples, sum represents summation, and C represents the number of classes.
[0108] This embodiment is carried out using PyCharm on a computer with a Windows 10 system. The relevant configurations of this computer are as follows: (1) Central Processing Unit (CPU): Intel(R) Xeon CPU @ 2.80GHz; (2) Graphic Processing Unit (GPU): RTX2080 (3) random access memory (RAM): 64.0GB (4) Pytorch1.6.
[0109] During the experiment, the training is divided into two stages. The first stage only processes a single view to achieve object recognition for fine-tuning the network model. The second stage processes all the input views to complete the training and testing work, where the number of iterations is set to 20 times. To optimize the MVPT network model during training, the learning rate is initialized to 0.0001, and the Adam optimizer is used. The learning rate decay and L2 regularization weight decay can avoid overfitting of the network model.
[0110] Test on the influence of the 2D image classification network on the image recognition performance:
[0111] In the multi-view low-level local feature token sequence generation stage, using different 2D image classification networks will affect the object recognition accuracy and training time of the entire network model. Therefore, this embodiment selects multiple classic image classification networks pre-trained on ImageNe for comparative experiments to evaluate their influence on the recognition accuracy and training time, so as to select the best network for subsequent experiments. In this test, the number of views n of the best view set is set to 6, the number of training times is set to 20 times, and other experimental settings are also kept consistent. Comparing VGG11, DenseNet121, ResNet18, ResNet50, and ResNet34, the experimental results are shown in Table 1 below (the bold values represent the best performance):
[0112]
[0113]
[0114] Table 1
[0115] As can be seen from Table 1 above, as an early proposed 2D image classification network, the overall recognition accuracy (OA) and average recognition accuracy (AA) of VGG11 are the lowest, and its training time performance is also not good. The DenseNet121 network has deep layers and a long training time, but it also fails to achieve the best recognition accuracy. The ResNet series of networks perform the best among these CNN models. Among them, ResNet34 achieves the best OA of 97.32% and AA of 95.95% with the second least training time (149 min). Therefore, in this embodiment, ResNet34 is selected as the multi-view low-level feature extractor for subsequent experiments.
[0116] Test on the influence of the number of view settings on image recognition performance:
[0117] Rendering a 3D object into multiple 2D views, different numbers of views will have different effects on the object recognition accuracy and training time of the network model. We used the optimal view set construction method based on information entropy to select five different numbers of views, namely single view, 3 views, 6 views, 12 views, and 20 views rendered with the icosahedron viewpoint setting, to quantitatively analyze the recognition accuracy and training time of the MVPT method. Among them, the process of constructing the optimal view set based on information entropy is as Figure 8 shown (the constructed optimal view set (n = 6) is in the black bold box).
[0118]
[0119] Table 2
[0120] The object recognition accuracy of the MVPT method under different numbers of views is as shown in Table 2 above. As can be seen from Table 2, when the number of views is set to 20, its overall recognition accuracy (OA) is lower than that of 3 views, 6 views, and 12 views, and it is accompanied by a large increase in training time. This experiment demonstrates the problem of redundancy in the current multi-view based 3D object recognition method when rendering a 3D object into multiple views. And the optimal view set construction method based on information entropy proposed in this embodiment can better solve this problem.
[0121] In this embodiment, the six views with the top information entropy values are constructed into the optimal view set, and the MVPT method achieves the best OA of 97.32% and AA of 95.95%. This result represents an improvement of 0.77% in OA compared to the 96.55% obtained with 20 views and an improvement of 0.67% in AA compared to the 95.28%. In terms of training time, the network model with 6 views takes 149 minutes. This is a reduction of 37.3% and 57.1% compared to the training times of 238 minutes and 348 minutes for 12 views and 20 views respectively. When the number of views is set to 1, the MVPT method can achieve an OA of 95.74% and an AA of 93.78. This result is already better than many current multi-view based 3D object recognition methods (see Table 5), and their number of views is often set to 12. When set to 20 epochs, the training time for a single view is greatly reduced. It only takes 80 minutes to complete the training of the network model, which is the least among the various numbers of views in Table 2. This also verifies the effectiveness of the optimal view set construction method based on information entropy.
[0122] For different numbers of view settings, this embodiment also conducted comparative experiments with other multi-view based 3D object recognition methods, including MVCNN, RCPCNN, 3D2SeqViews, VERAM, MHBN, and RN. The experimental results of various multi-view object recognition methods under different numbers of views are shown in Table 3 below:
[0123]
[0124] Table 3 (The experimental results are represented by the overall recognition accuracy, in %, and the bold values indicate the best performance)
[0125] As can be seen from Table 3, for the view settings of 3 views, 6 views, and 12 views, the MVPT method proposed in this embodiment always shows the most advanced performance, which are 96.88%, 97.32%, and 96.71% respectively. Among the above other methods, the RN method achieves the highest OA. This method reaches an OA of 94.30% in the case of 12 views, while the MVPT method is 3.02% higher than RN. It is worth noting that when the number of views n of most methods increases from 6 to 12, their object recognition accuracy decreases. This is another manifestation of the redundancy problem in the current method of rendering 3D objects as multi-views proposed in the present invention. At the same time, the 6-view optimal view set constructed in this embodiment achieves the best object recognition accuracy. This is because the feature information extracted from views with high information entropy values is also more abundant; in addition, when the number of views n is set to 6, the overlapping parts between each view are reduced, reducing the redundancy of the extracted feature information.
[0126] Regarding the test of the influence of different aggregation methods on object recognition accuracy:
[0127] To verify the effectiveness of the global descriptor generation method based on Pooling Transformer, a comparative experiment was conducted in this embodiment with an aggregation method that only uses maxpool and Transformer. The number of experimental views n was set to 6, and ResNet34 was selected as the multi-view low-level feature extractor. With other experimental environment settings being the same, the experimental results are shown in Table 4 below:
[0128]
[0129] Table 4
[0130] As can be seen from Table 4, the training times of these three methods are basically the same. However, the Pooling Transformer method proposed in this embodiment achieves the best object recognition accuracy. The overall recognition accuracy (OA) is 97.32%, and the average classification accuracy (AA) also reaches 95.95%. Compared with the original Transformer method, OA is increased by 1.35% and AA is increased by 0.99%. There is an even greater improvement compared with the max pool method. This is because the Pooling Transformer method solves the problem of insufficient local feature aggregation ability of the original Transformer. It can aggregate the feature information of all local view token sequences from both local and global aspects.
[0131] Regarding the comparative test with other methods in the object recognition experiment:
[0132] In this embodiment, a 3D object recognition comparative experiment of the MVPT method was conducted with voxel-based methods 3D ShapeNets, VoxNet, and O-CNN, point cloud-based methods PointNet, PointNet++, PointWeb
[16] , and DGCNN, and multi-view-based methods MVCNN, GVCNN, 3D2SeqViews, VERAM, RN, HMVCM, EMV, and MHBN on the ModelNet40 dataset. The experimental results are shown in Table 5 below:
[0133]
[0134]
[0135] Table 5
[0136] As can be seen from Table 5, the performance of the MVPT method is far superior to other current methods, with an overall recognition accuracy of 97.32% and an average recognition accuracy of 95.95%. In addition, for each 3D object, the present invention only requires 6 views to complete the object recognition task. Compared with other multi-view based methods, the number of views is also the least, which will help reduce the training time and computational cost. In addition, most multi-view based methods are superior to point cloud and voxel based methods.
[0137] In summary, a large number of experiments were carried out on the ModelNet40 dataset in this embodiment to verify the performance of the method. The MVPT can achieve an overall recognition accuracy of 97.32% and an average recognition accuracy of 95.95% with only 6 views. Compared with existing deep learning based methods, the MVPT network achieves state-of-the-art performance. And compared with the view point setting of the regular dodecahedron, the proposed information entropy based optimal view set construction method in this embodiment reduces the training time of the network model to 42.8% of the original. When selecting the best view, that is, the single view with the highest information entropy value, the MVPT method can achieve an OA of 95.74%, which is already superior to many current multi-view based 3D object recognition methods, but its training time is shorter, only 80 minutes, and under the same computing hardware conditions, it is reduced to 51% of the training time of existing advanced algorithms.
[0138] Functions and effects of the embodiment
[0139] According to the three-dimensional object recognition method based on multi-view Pooling Transformer provided in this embodiment, the method proposes a Multi-view Pooling Transformer (MVPT) network model. The MVPT model first constructs an optimal view set based on the information entropy of multi-views, and uses the ResNet network and the Embedding network to extract feature information from all views at one time, obtaining a multi-view low-level local feature token sequence, so that it can be input into the Pooling Transformer to complete parallel training. Then, through the Pooling Transformer, the multi-view low-level local feature token sequences are aggregated from the global and local respectively into a compact and single 3D global descriptor. Finally, the classifier recognizes the 3D global descriptor to obtain the recognition result of the object to be measured.
[0140] In the embodiment, since the optimal view set is constructed by selecting the information entropy with the top ranking in multiple views, the problem of redundancy in rendering multiple views from 3D objects currently is solved, and the accuracy of the network model for recognition is improved. Also, since the Transformer is applied to the 3D object recognition task, the problem that the current multi-view based method loses the relevant information between different views is solved. In addition, the global descriptor generation method based on Pooling Transformer also solves the problem of insufficient local feature information aggregation ability of the Transformer.
[0141] Therefore, the 3D object recognition method based on multi-view Pooling Transformer in this embodiment extracts feature information from all views at once, solves the redundancy problem of multi-views and the low efficiency problem of model training, thereby efficiently and accurately capturing the relevant feature information between multiple views, and greatly improving the recognition accuracy and training efficiency of the network model.
[0142] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.
[0143] The above references are as follows:
[0144] [1]Maturana D,Scherer S.Voxnet:A 3d convolutional neural network forreal-time object recognition[C] / / 2015IEEE / RSJ International Conference onIntelligent Robots and Systems(IROS).IEEE,2015:922-928.
[0145] [2]Wu Z,Song S,Khosla A,et al.3d shapenets:A deep representation forvolumetric shapes[C] / / Proceedings of the IEEE conference on computer visionand pattern recognition.2015:1912-1920.
[0146] [3] Riegler G, Osman Ulusoy A, Geiger A. Octnet: Learning deep 3d representations at high resolutions[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017:3577-3586.
[0147] [4] Wang P S, Liu Y, Guo Y X, et al. O-cnn: Octree-based convolutional neural networks for 3d shape analysis[J]. ACM Transactions On Graphics(TOG), 2017, 36(4):1-11.
[0148] [5] Klokov R, Lempitsky V. Escape from cells: Deep kd-networks for the recognition of 3d point cloud models[C] / / Proceedings of the IEEE International Conference on Computer Vision. 2017:863-872.
[0149] [6] Zeng W, Gevers T. 3dcontextnet: Kd tree guided hierarchical learning of point clouds using local and global contextual cues[C] / / Proceedings of the European Conference on Computer Vision(ECCV)Workshops. 2018:0-0.
[0150] [7]Le T, Duan Y. Pointgrid: A deep network for 3d shape understanding[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018:9204-9214.
[0151] [8]Meng H Y, Gao L, Lai Y K, et al. Vv-net: Voxel vae net with group convolutions for point cloud segmentation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:8500-8508.
[0152] [9]Qi C R, Su H, Mo K, et al. Pointnet: Deep learning on point sets for 3d classification and segmentation[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017:652-660.
[0153]
[10] Qi C R, Yi L, Su H, et al. Pointnet++: Deep hierarchical feature learning on point sets in a metric space[J]. arXiv preprint arXiv:1706.02413, 2017.
[0154]
[11] Jiang M, Wu Y, Zhao T, et al. Pointsift: A sift-like network module for 3d point cloud semantic segmentation[J]. arXiv preprint arXiv:1807.00652, 2018.
[0155]
[12] Zhao H, Jiang L, Fu C W, et al. Pointweb: Enhancing local neighborhood features for point cloud processing[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019:5565-5573.
[0156]
[13] Fu J, Liu J, Tian H, et al. Dual attention network for scene segmentation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019:3146-3154.
[0157]
[14] Feng M, Zhang L, Lin X, et al. Point attention network for semantic segmentation of 3D point clouds[J]. Pattern Recognition, 2020, 107:107446.
[0158]
[15] Wang Y, Sun Y, Liu Z, et al. Dynamic graph cnn for learning on point clouds[J]. Acm Transactions On Graphics(tog), 2019, 38(5):1-12.
[0159]
[16] Zhang K, Hao M, Wang J, et al. Linked dynamic graph cnn: Learning on point cloud via linking hierarchical features[J]. arXiv preprint arXiv:1904.10014, 2019.
[0160]
[17] Su H,Maji S,Kalogerakis E,et al.Multi-view convolutional neuralnetworks for 3d shape recognition[C] / / Proceedings of the IEEE internationalconference on computer vision.2015:945-953.
[0161]
[18] Wang C,Pelillo M,Siddiqi K.Dominant set clustering and poolingfor multi-view 3d object recognition[J].arXiv preprint arXiv:1906.01592,2019.
[0162]
[19] Su H,Maji S,Kalogerakis E,et al.Multi-view convolutional neuralnetworks for 3d shape recognition[C] / / Proceedings of the IEEE internationalconference on computer vision.2015:945-953.
[0163]
[20] Kanezaki A,Matsushita Y,Nishida Y.Rotationnet:Joint objectcategorization and pose estimation using multiviews from unsupervisedviewpoints[C] / / Proceedings of the IEEE Conference on Computer Vision andPattern Recognition.2018:5010-5019.
[0164]
[21] Esteves C, Xu Y, Allen-Blanchette C, et al. Equivariant multi-view networks[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:1568-1577.
[0165]
[22] Han Z, Lu H, Liu Z, et al. 3D2SeqViews: Aggregating sequential views for 3D global feature learning by CNN with hierarchical attention aggregation[J]. IEEE Transactions on Image Processing, 2019, 28(8):3986-3999.
[0166]
[23] He X, Huang T, Bai S, et al. View n-gram network for 3d object retrieval[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:7515-7524.
[0167]
[24] He X, Huang T, Bai S, et al. View n-gram network for 3d object retrieval[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:7515-7524.
[0168]
[25] Chen S, Zheng L, Zhang Y, et al. Veram: View-enhanced recurrent attention model for 3d shape classification[J]. IEEE transactions on visualization and computer graphics, 2018, 25(12):3244-3257.
[0169]
[26] Liu A A,Zhou H,Nie W,et al.Hierarchical multi-view contextmodelling for 3D object classification and retrieval[J].Information Sciences,2021,547:984-995.
Claims
1. A three-dimensional object recognition method based on multi-view Pooling Transformer, characterized in that, It includes the following steps: Step S1, construct a Multi-view Pooling Transformer network model, which has an optimal view set acquisition module, a low-level local feature token sequence generation module, a global descriptor generation module based on Pooling Transformer, and a classifier; Step S2, input the object to be measured into the Multi-view Pooling Transformer network model, obtain the corresponding multi-views through the optimal view set acquisition module, and construct an optimal view set according to the information entropy of the multi-views; Step S3, the low-level local feature token sequence generation module extracts the multi-view low-level local features of the optimal view set, and generates a corresponding multi-view low-level local feature token sequence based on the multi-view low-level local features; Step S4, the global descriptor generation module aggregates the local view information token sequence of the multi-view low-level local feature token sequence and its global feature information sequence to generate a 3D global descriptor of the object to be measured; Step S5, the classifier uses the 3D global descriptor as an input for three-dimensional object recognition, so as to obtain the recognition result of the object to be measured, wherein, the global descriptor generation module includes a global feature information generation sub-module based on Transformer and a local view information token sequence aggregation sub-module based on Pooling, the global feature information generation sub-module based on Transformer has a Layer Normalization network, a Multi-Head Multi-View Attention network, a multi-layer perceptron network, and a residual connection, Step S4 includes the following sub-steps: Step S4-1, the Layer Normalization network normalizes the multi-view low-level local feature token sequence: Step S4-2, the Multi-Head Multi-View Attention network performs MHMVA calculation on the normalized token sequence through linear transformation to generate the token sequence X MHMVA , The Multi-Head Multi-View Attention network consists of multiple Multi-View Attention; The MHMVA calculation is to perform multiple parallelized Multi-View Attention calculations: Step S4-2-1, the normalized is first linearly transformed to generate three vectors: Query, Key, and Value. Step S4-2-2: Divide the three vectors from the previous step into multiple inputs q of the Multi-View Attention according to the number N of Heads i , k i , v i , which can form multiple subspaces, focus on the information of different parts of the input features, and finally splice these feature information to obtain more abundant information: Step S4-2-3, Multi-View Attention calculates MVA based on the input, that is, calculates the product of q i and the transpose of k i to obtain a score, divides it by and performs normalization to stabilize the gradient. Then, the normalized result value is used as the input of the softmax function, and the output of the softmax function is dot-multiplied with v i to obtain where d k is the dimension of k i ; Step S4-2-4, for each of the calculated perform Concat, and then finally complete the MHMVA calculation after another linear transformation: Step S4-3, for the token sequence X MHMVA Use a residual connection to obtain the token sequence X1 to avoid gradient vanishing, and then input X1 into the Layer Normalization network for normalization processing and then input it into the multi-layer perceptron network; Step S4-4, perform a residual connection between the output result of the multi-layer perceptron network and X1 to obtain the local view information token sequence: Among them, the local view information token sequence consists of a global and a local view information token sequence The global stores the global feature information of the local view token sequence, that is Step S4-5, the local view information token sequence aggregation sub-module based on Pooling aggregates the local view information token sequence to perform pooling processing to obtain a single best local view information token, and then splice and aggregate this best local view information token with the global to finally generate the corresponding 3D global descriptor Y:
2. The 3D object recognition method based on multi-view Pooling Transformer according to claim 1, It is characterized in that: wherein, Step S2 includes the following sub-steps: Step S2-1, obtain a corresponding plurality of 2D views of the object to be measured according to the regular dodecahedron viewpoints; Step S2-2, calculate the information entropy of each 2D view, and sort them according to the magnitude of the information entropy value; Step S2-3, select the views ranked top n in terms of information entropy as the optimal view set, so as to reduce redundant views.
3. The three-dimensional object recognition method based on multi-view Pooling Transformer according to claim 2, wherein: Among them, The calculation formula of the information entropy is: P a,b = f(a, b) / W·H Where, H i represents the information entropy of the i-th view v i . (a, b) is a binary tuple, where a represents the grayscale value at the center within a certain sliding window, and b is the average grayscale value of the pixels within the window excluding the central pixel; P a,b represents the probability of the occurrence of (a, b) in the entire view v i ; f(a, b) represents the number of occurrences of the binary tuple (a, b) in the entire view v i ; W and H represent the width and height of the view v i .
4. The three-dimensional object recognition method based on multi-view Pooling Transformer according to claim 1, wherein: Among them, The low-level local feature token sequence generation module has a ResNet network and an Embedding network, The step S3 includes the following sub-steps: Step S3-1: Extract multi-view low-level local features of the best view set by the ResNet network; Step S3-2: Generate a local view token sequence of the multi-view low-level local features based on the Embedding network: [x1,...x i ...,x n = Emb{Res[v1,...v i ...,v n} wherein, [v i , … v i …, v n is the set of the optimal views, and v i represents one of the views; Step S3-3, add an initialized class token x class to the head of the local view token sequence, and splice them with the position encoding E pos respectively, and finally generate the multi-view low-level local feature token sequence: where, X0 is a multi-view low-level local feature token sequence, and x class is a randomly initialized value that matches the dimension of the local view token sequence, and E pos is used to save the position information from different viewpoints x i of.