A 3D Object Recognition Method and System Based on Self-Attention Mechanism

By adopting a self-attention mechanism method in 3D object recognition, aggregating the feature information of multi-view diagrams, the problem of redundancy of view information is solved, and the accuracy and efficiency of recognition are improved.

CN115393841BActive Publication Date: 2025-07-01GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210894043.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-07-01
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

The prior art ignores the connection between multi-view diagrams in 3D object recognition, resulting in redundant view information and fails to effectively overcome this problem.

Method used

Using a method based on self-attention mechanism, the self-attention network model allows the multi-view diagram to be linked through correlation, and the feature information of the multi-view diagram is aggregated on a representative view diagram to reduce feature information redundancy.

Benefits of technology

It effectively reduces feature information redundancy, pays attention to information from representative perspective maps, and improves the accuracy and efficiency of 3D object recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393841B_ABST
    Figure CN115393841B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D object recognition method and system based on a self-attention mechanism, relating to the technical field of 3D object recognition, including: acquiring multi-view images and their position information; inputting the multi-view images into a convolutional neural network to obtain feature encodings and classification scores; embedding the position information and view-independent information into the feature encodings of the corresponding view images to obtain first embedded feature encodings, inputting the first embedded feature encodings into a first self-attention network model, and outputting first self-attention feature encodings; sampling the first self-attention feature encodings by using the classification scores, embedding the position information and view-independent information, obtaining second embedded feature encodings, inputting the second embedded feature encodings into a second self-attention network model, and outputting second self-attention feature encodings; constructing a global feature descriptor based on the first and second self-attention feature encodings for classification and retrieval to obtain recognition results. The present invention fully considers the connections between multi-view images, aggregates feature information on representative view images, and reduces feature information redundancy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D object recognition, and more specifically, to a 3D object recognition method and system based on a self-attention mechanism. Background Art

[0002] In recent years, 3D object recognition has become one of the most important research and application directions in artificial intelligence and is also extremely challenging in real application scenarios. Its main research methods are divided into three categories: voxel-based methods, point cloud-based methods, and view-based methods. Due to the maturity of convolutional neural network image feature extraction technology, view-based methods have made the greatest progress. In traditional view-based methods, multi-view images of a 3D object are passed through a convolutional neural network to extract features, and then a max pooling operation is performed to aggregate the information of the multi-view images, thereby obtaining a descriptor representing the multi-view images of the 3D object. The multi-view images of a 3D object are images obtained through mappings from different perspectives. However, there are feature-related connections between adjacent or distant views. For example, adjacent views have more common features, while distant views have greater differences. However, using a simple max pooling operation to aggregate the information of multi-view images ignores the connections between views and will cause redundancy in view information. In recent years, more and more researchers have considered the connections between views and have used various methods to increase the connections between views so that the training model can learn more useful information, but few people have considered the problem of redundant multi-view image information. Since there are many similar features between views, how to reduce information redundancy and pay more attention to the information of representative views is also a problem that needs to be studied currently.

[0003] The prior art discloses an unsupervised 3D object recognition and retrieval method based on deep circle views, including the following steps: Step 1, multi-circle data sampling; Step 2, training a multi-view deep network model based on circle data; Step 3, similarity matching and retrieval; using the trained multi-view deep network model to extract the features of each circle view and calculating the similarity distance for all circle views; optimizing the multi-view deep network model by using max pooling, average pooling, attention pooling, and optimal matching; performing sorting and retrieval based on the similarity distance; Step 4, adopting a circle feature filtering and circle attention strategy to filter out circle features with importance lower than a specified threshold. Although this invention avoids a large amount of manual annotation, it ignores the connections between views and still does not overcome the problem of redundant view information. Summary of the Invention

[0004] To overcome the problem that the above-mentioned existing technologies ignore the connection between multi-view images and redundant feature information during 3D object recognition, the present invention provides a 3D object recognition method and system based on a self-attention mechanism, which fully considers the connection between multi-view images, filters out similar feature information between multi-view images, aggregates the feature information of multi-view images on a representative view image, and reduces redundant feature information.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] The present invention provides a 3D object recognition method based on a self-attention mechanism, including

[0007] S1: Obtain multi-view images of the 3D object to be recognized and the position information of each view image;

[0008] S2: Input the multi-view images of the 3D object to be recognized into a pre-trained convolutional neural network for classification to obtain the feature encoding and classification score of the multi-view images;

[0009] S3: Combine the position information of each view image with view-independent information and embed it into the feature encoding of the corresponding view image to obtain the first embedded feature encoding;

[0010] S4: Input the first embedded feature encoding into the first self-attention network model and output the first self-attention feature encoding;

[0011] S5: Sample the first self-attention feature encoding according to the classification score of the view image to obtain a sampling result;

[0012] S6: Embed the corresponding position information and view-independent information of the view image into the sampling result to obtain the second embedded feature encoding;

[0013] S7: Input the second embedded feature encoding into the second self-attention network model and output the second self-attention feature encoding;

[0014] S8: Construct a global feature descriptor according to the first self-attention feature encoding and the second self-attention feature encoding;

[0015] S9: Perform classification retrieval according to the global feature descriptor to obtain the recognition result of the 3D object.

[0016] Preferably, in the step S1, the specific method for obtaining the position information of each view image is:

[0017] Number each view image according to the view order of the multi-view images of the 3D object to be recognized; encode each view image using the sine-cosine position encoding function, and use the obtained position encoding as the position information of each view image;

[0018] The sine-cosine position encoding function is as follows:

[0019] PE (pos,2i) = sin(pos / 10000 2i / d )

[0020] PE (pos,2i+1) = cos(pos / 10000 2i+1 / d )

[0021] In the formula, pos represents the sequential number of the perspective view, pos = 1, 2, …, n, where n represents the number of multi-perspective views; d represents the length of the feature vector of the position encoding; i represents the i-th element in the feature vector of the position encoding. The even positions are encoded with sin and the odd positions are encoded with cos. Then, PE (pos,2i) represents the position encoding value of the even positions of the pos-th perspective view, and PE (pos,2i+1) represents the position encoding value of the odd positions of the pos-th perspective view. According to PE (pos,2i) and PE (pos,2t+1) , the position encoding E pos of the pos-th perspective view is formed as the position information of this perspective view, and E pos ∈R (n+1)×512 .

[0022] Preferably, in step S2, the pre-trained convolutional neural network is a VggNet network; the multi-perspective views of the 3D object to be recognized are input into the pre-trained VggNet network, and the feature encoding F = [f1, f2, …, f n and the classification score of the multi-perspective views are output. f n represents the feature encoding of the n-th perspective view, and f n ∈R 512 .

[0023] Preferably, in step S3, the specific method for obtaining the first embedded feature encoding is as follows:

[0024] Set the view-agnostic information [class] token, and the view-agnostic information [class] token is a learnable vector with the same dimension as the feature encoding of the perspective view; then the first embedded feature encoding is:

[0025] X0 = [f [class] , f1, f2, …, f n + E pos

[0026] In the formula, X0 represents the first embedded feature encoding, f [class] represents the first feature encoding of the view-agnostic information, and f [class] ∈R 512 .

[0027] Embedding view-agnostic information [class] tokens is to objectively learn the feature information of other perspective views different from the current perspective view; each of the other perspective views carries its own preferred feature information. By embedding view-agnostic information, more feature information can be learned. The purpose of embedding position information is to learn perspective information and increase the comprehensiveness and accuracy of feature information.

[0028] Preferably, in the step S4, the first self-attention network model includes N self-attention networks connected in sequence; each self-attention network includes a first normalization layer, a multi-head attention layer, a first residual connection point, a second normalization layer, a linear mapping layer, and a second residual connection point connected in sequence;

[0029] The first embedded feature encoding X0 is input into the first self-attention network. After being processed by the first normalization layer and the multi-head attention layer in sequence, it is connected with X0 at the first residual connection point to obtain the intermediate feature encoding X'; X' is processed by the second normalization layer and the linear mapping layer in sequence, and is connected with X' at the second residual connection point to obtain the output X1 of the first self-attention network; that is:

[0030] X' = MHA(LN1(X0)) + X0

[0031] X1 = MLP(LN2(X')) + X'

[0032] In the formula, LN1 represents the first normalization operation, MHA represents the multi-head attention operation, LN2 represents the second normalization operation, and MLP represents the linear mapping operation;

[0033] The output X1 of the first self-attention network is input into the second self-attention network. According to the same method, the output X2 of the second self-attention network is output; until after being processed by N self-attention networks, then:

[0034] X' = MHA(LN1(X N-1 )) + X N-1

[0035] X N = MLP(LN2(X')) + X'

[0036] In the formula, X N represents the output of the Nth self-attention network;

[0037] Taking the output of the first self-attention network model as the first self-attention feature encoding X last X last = [f' [class] , f'1, f'2,..., f' n , where f' [class] represents the first self-attention feature encoding value of the view-agnostic information, f' nRepresents the first self-attention feature encoding value of the nth perspective view.

[0038] The self-attention network model can establish connections between multi-perspective views through correlations, aggregate the feature information of multi-perspective views onto several representative perspective views, thereby reducing the redundancy of feature information.

[0039] Preferably, the specific method of step S5 is as follows:

[0040] Sort the perspective views in descending order according to the classification scores of the perspective views, and retain the perspective views whose classification scores are in the top n / 2 positions; use the retained perspective views to sample the first self-attention feature encoding X last Perform sampling, that is, filter and retain the first self-attention feature encoding values corresponding to the perspective views, and record the sampling result as S = [s1, s2,..., s n / 2 , s n / 2 represents the n / 2th element in the sampling result, and n represents the number of multi-perspective views.

[0041] Sampling using the classification scores of each perspective view makes the feature information of multi-perspective views more aggregated and increases the special feature information of a single perspective view.

[0042] Preferably, the specific method of step S6 is as follows:

[0043] Embed the view-independent information [class] token and the corresponding perspective view position information into the sampling result to obtain the second embedded feature encoding S0:

[0044] S0 = [s [class] , s1, s2,..., s n / 2 + E pos

[0045] In the formula, S [class] represents the second feature encoding of the view-independent information.

[0046] In the sampled and retained perspective views, embed the view-independent information and position information again to further learn the feature information and perspective information irrelevant to the perspective view, and increase the comprehensiveness and accuracy of the feature information.

[0047] Preferably, in step S7, the second self-attention network model has exactly the same structure as the first self-attention network model;

[0048] After the second embedded feature encoding S0 is processed by the second self-attention network model, the second self-attention feature encoding S last , S last = [s′ [class] , s′1, s′2,..., s′ n / 2 is output, where s′[class] The second self-attention feature encoding value representing view-independent information, s′ n / 2 The second self-attention feature encoding value representing the n / 2-th element in the sampling result.

[0049] The self-attention network model can establish connections between multi-view graphs through correlations, further aggregating the feature information of the multi-view graphs retained by sampling onto several representative view graphs, and reducing the redundancy of feature information again.

[0050] Preferably, in the step S8, the specific method for constructing the global feature descriptor is as follows:

[0051] After respectively performing max-pooling operations on the first self-attention feature encoding and the second self-attention feature encoding, they are concatenated into a global feature descriptor, and the calculation formula is:

[0052] F global = Cat(Max(X last ), Max(S last ))

[0053] In the formula, F global represents the global feature descriptor, Cat represents the concatenation operation, and Max represents the max-pooling operation.

[0054] 3D object recognition includes a classification task and a retrieval task. When performing the classification task, the global feature descriptor F global of the 3D object to be recognized is followed by a fully connected layer and a normalization layer to obtain the classification result as the recognition result; when performing the retrieval task, the Euclidean distance between the global feature descriptor F global of the 3D object to be recognized and the global feature descriptor F global of the known 3D object is calculated, and the known 3D object with the smallest Euclidean distance is used as the retrieval result to obtain the recognition result of the 3D object to be recognized.

[0055] The present invention also provides a 3D object recognition system based on the self-attention mechanism to implement the above-mentioned 3D object recognition method based on the self-attention mechanism. The system includes:

[0056] A data acquisition module for acquiring multi-view graphs of the 3D object to be recognized and the position information of each view graph;

[0057] A convolutional neural network module for inputting the multi-view graphs of the 3D object to be recognized into a pre-trained convolutional neural network for classification to obtain the feature encoding and classification score of the multi-view graphs;

[0058] A first embedding module for combining the position information of each view graph with view-independent information and embedding it into the feature encoding of the corresponding view graph to obtain the first embedding feature encoding;

[0059] The first self-attention module is used to encode the first embedded feature into the first self-attention network model and output the first self-attention feature encoding;

[0060] The sampling module is used to sample the first self-attention feature encoding according to the classification score of the perspective view map to obtain a sampling result;

[0061] The second embedding module is used to embed the position information and view-independent information of the corresponding perspective view map in the sampling result to obtain the second embedded feature encoding;

[0062] The second self-attention module is used to encode the second embedded feature into the second self-attention network model and output the second self-attention feature encoding;

[0063] The global feature construction module is used to construct a global feature descriptor according to the first self-attention feature encoding and the second self-attention feature encoding;

[0064] The recognition module is used to perform classification retrieval according to the global feature descriptor to obtain the recognition result of the recognized 3D object.

[0065] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0066] The present invention first obtains multi-perspective view maps of the 3D object to be recognized and their position information, classifies the multi-perspective view maps by using a convolutional neural network to obtain feature encoding and classification scores; then embeds view-independent information and position information into the feature encoding, inputs it into the first self-attention network model, uses the view-independent information and position information to increase the comprehensiveness of the feature encoding, learns more feature information and perspective information of other views, and the self-attention network model makes the multi-views related through correlation, aggregates the feature information of the multi-perspective view maps on several representative perspective view maps, thereby reducing the redundancy of feature information; furthermore, samples the first self-attention feature encoding according to the classification score of each perspective view map, makes the feature information of the multi-perspective view maps more aggregated, and increases the special feature information of a single perspective view map; then embeds the view-independent information and position information into the sampling result again, further aggregates the feature information of the multi-perspective view maps retained by the sampling on several representative perspective view maps, reduces the redundancy of feature information again, and pays more attention to the feature information of the representative perspective view maps; finally, constructs a global feature descriptor according to the first self-attention feature encoding and the second self-attention feature encoding for classification retrieval to obtain the recognition result of the recognized 3D object. The present invention fully considers the relationship between multi-perspective view maps, filters out similar feature information between multi-perspective view maps, aggregates the feature information of multi-perspective view maps on representative perspective view maps, and reduces the redundancy of feature information. Description of the Drawings

[0067] Figure 1 Flow chart of a 3D object recognition method based on self-attention mechanism described in Embodiment 1;

[0068] Figure 2 Structural diagram of the self-attention network described in Embodiment 2;

[0069] Figure 3 Flow chart when the global feature descriptor described in Embodiment 2 performs a classification task;

[0070] Figure 4 Structural diagram of a 3D object recognition system based on self-attention mechanism described in Embodiment 3. Detailed implementation manners

[0071] The drawings are only for illustrative purposes and should not be construed as limitations of this patent;

[0072] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;

[0073] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0074] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0075] Embodiment 1

[0076] This embodiment provides a 3D object recognition method based on self-attention mechanism, as Figure 1 shown, including

[0077] S1: Obtain multi-view images of the 3D object to be recognized and the position information of each view image;

[0078] S2: Input the multi-view images of the 3D object to be recognized into a pre-trained convolutional neural network for classification to obtain the feature encoding and classification score of the multi-view images;

[0079] S3: Combine the position information of each view image with view-independent information and embed it into the feature encoding of the corresponding view image to obtain the first embedded feature encoding;

[0080] S4: Input the first embedded feature encoding into the first self-attention network model and output the first self-attention feature encoding;

[0081] S5: Sample the first self-attention feature encoding according to the classification score of the view image to obtain a sampling result;

[0082] S6: Embed the position information and view-independent information of the corresponding perspective view in the sampling result to obtain a second embedded feature encoding;

[0083] S7: Input the second embedded feature encoding into the second self-attention network model to output a second self-attention feature encoding;

[0084] S8: Construct a global feature descriptor based on the first self-attention feature encoding and the second self-attention feature encoding;

[0085] S9: Perform classification retrieval based on the global feature descriptor to obtain the recognition result for identifying the 3D object.

[0086] In the specific implementation process, in this embodiment, first, the multi-perspective views and their position information of the 3D object to be recognized are obtained, and the convolutional neural network is used to classify the multi-perspective views to obtain feature encodings and classification scores; then, the view-independent information and position information are embedded in the feature encodings and input into the first self-attention network model. The view-independent information and position information are used to increase the comprehensiveness of the feature encodings, learn more feature information and perspective information of other views, and the self-attention network model enables the multi-views to be related through correlation, aggregating the feature information of the multi-perspective views on several representative perspective views, thereby reducing the redundancy of feature information; furthermore, the first self-attention feature encoding is sampled using the classification score of each perspective view, making the feature information of the multi-perspective views more aggregated and increasing the special feature information of a single perspective view; then, the view-independent information and position information are embedded in the sampling result again, further aggregating the feature information of the multi-perspective views retained by the sampling on several representative perspective views, reducing the redundancy of feature information again, and paying more attention to the feature information of the representative perspective views; finally, a global feature descriptor is constructed based on the first self-attention feature encoding and the second self-attention feature encoding for classification retrieval to obtain the recognition result for identifying the 3D object. The present invention fully considers the relationship between multi-perspective views, filters out the similar feature information between multi-perspective views, aggregates the feature information of multi-perspective views on representative perspective views, and reduces the redundancy of feature information.

[0087] Embodiment 2

[0088] This embodiment provides a 3D object recognition method based on a self-attention mechanism, including

[0089] S1: Obtain the multi-perspective views of the 3D object to be recognized and the position information of each perspective view;

[0090] The specific method for obtaining the position information of each perspective view is:

[0091] Number each perspective view according to the perspective order of the multi-perspective views of the 3D object to be recognized; encode each perspective view using the sine-cosine positional encoding function, and use the obtained positional encoding as the positional information of each perspective view.

[0092] The sine-cosine positional encoding function is as follows:

[0093] PE (pos,2i) = sin(pos / 10000 2i / d )

[0094] PE (pos,2i+1) = cos(pos / 10000 2i+1 / d )

[0095] where pos represents the sequential number of the perspective view, pos = 1, 2,..., n, n represents the number of multi-perspective views; d represents the length of the feature vector of the positional encoding; i represents the i-th element in the feature vector of the positional encoding, the even positions are encoded with sin, and the odd positions are encoded with cos, then PE (pos,2i) represents the positional encoding value of the even positions of the pos-th perspective view, and PE (pos,2i+1) represents the positional encoding value of the odd positions of the pos-th perspective view. According to PE (pos,2i) and PE (pos,2t+1) , form the positional encoding E pos of the pos-th perspective view, as the positional information of this perspective view, E pos ∈ R (n+1)×512 .

[0096] S2: Input the multi-perspective views of the 3D object to be recognized into a pre-trained convolutional neural network for classification to obtain the feature encoding and classification score of the multi-perspective views;

[0097] The pre-trained convolutional neural network is the VggNet network; input the multi-perspective views of the 3D object to be recognized into the pre-trained VggNet network, and output the feature encoding F = [f1, f2,..., f n and the classification score, where f n represents the feature encoding of the n-th perspective view, and f n ∈ R 512 .

[0098] S3: Combine the positional information of each perspective view with view-agnostic information and embed it into the feature encoding of the corresponding perspective view to obtain the first embedded feature encoding; Set the view-agnostic information [class]token, and the view-agnostic information [class]token is a learnable vector with the same dimension as the feature encoding of the perspective view; then the first embedded feature encoding is:

[0099] X0 = [f [class], f1, f2, …, f n + E pos

[0100] In the formula, X0 represents the first embedded feature encoding, and f [class] represents the first feature encoding of view-independent information, and f [class] ∈R 512 .

[0101] S4: Input the first embedded feature encoding into the first self-attention network model, and output the first self-attention feature encoding;

[0102] The first self-attention network model includes N self-attention networks connected in sequence; as Figure 2 shown, each self-attention network includes a first normalization layer, a multi-head attention layer, a first residual connection point, a second normalization layer, a linear mapping layer, and a second residual connection point connected in sequence; in the figure, the plus sign represents the residual connection point;

[0103] The first embedded feature encoding X0 is input into the first self-attention network. After being processed by the first normalization layer and the multi-head attention layer in sequence, it is connected with X0 at the first residual connection point to obtain the intermediate feature encoding X'; X' is processed by the second normalization layer and the linear mapping layer in sequence, and is connected with X' at the second residual connection point to obtain the output X1 of the first self-attention network; that is:

[0104] X' = MHA(LN1(X0)) + X0

[0105] X1 = MLP(LN2(X')) + X'

[0106] In the formula, LN1 represents the first normalization operation, MHA represents the multi-head attention operation, LN2 represents the second normalization operation, and MLP represents the linear mapping operation;

[0107] Input the output X1 of the first self-attention network into the second self-attention network. According to the same method, output the output X2 of the second self-attention network; until after being processed by N self-attention networks, then:

[0108] X' = MHA(LN1(X N-1 )) + X N-1

[0109] X N = MLP(LN2(X')) + X'

[0110] In the formula, X N represents the output of the Nth self-attention network;

[0111] Take the output of the first self-attention network model as the first self-attention feature encoding X last , X last = [f'[class] , f′1, f′2, …, f′ n , where f′ [class] represents the first self-attention feature encoding value of view-independent information, and f′ n represents the first self-attention feature encoding value of the nth perspective view.

[0112] The self-attention network model can establish connections between multi-perspective views through correlations, aggregate the feature information of multi-perspective views on several representative perspective views, thereby reducing the redundancy of feature information. In this embodiment, the first self-attention network model includes 12 sequentially connected self-attention networks, and the output of the last self-attention network is used as the output of the first self-attention network model.

[0113] S5: Sample the first self-attention feature encoding according to the classification scores of the perspective views to obtain a sampling result;

[0114] Sort the classification scores of the perspective views from large to small, and retain the perspective views whose classification scores are in the top n / 2 positions; use the retained perspective views to sample the first self-attention feature encoding X last That is, filter and retain the first self-attention feature encoding values corresponding to the perspective views. The sampling result is denoted as S = [s1, s2, …, s n / 2 , where s n / 2 represents the n / 2th element in the sampling result, n represents the number of multi-perspective views, and S ∈ R (n+1)×512 .

[0115] Sampling using the classification scores of each perspective view makes the feature information of multi-perspective views more aggregated and increases the special feature information of a single perspective view.

[0116] S6: Embed the position information and view-independent information of the corresponding perspective views in the sampling result to obtain a second embedded feature encoding;

[0117] Embed the view-independent information [class]token and the corresponding perspective view position information into the sampling result to obtain a second embedded feature encoding S0:

[0118] S0 = [s [class] , s1, s2, …, s n / 2 + E pos

[0119] In the formula, s [class] represents the second feature encoding of the view-independent information.

[0120] In the perspective view retained by sampling, view-independent information and position information are embedded again to further learn feature information and perspective information that are irrelevant to this perspective view, enhancing the comprehensiveness and accuracy of the feature information.

[0121] S7: Input the second embedded feature encoding into the second self-attention network model to output the second self-attention feature encoding;

[0122] The second self-attention network model has exactly the same structure as the first self-attention network model;

[0123] After the second embedded feature encoding S0 is processed by the second self-attention network model, the second self-attention feature encoding S is output last , S last =[s′ [class] , s′1, s′2, …, s′ n / 2 , where s′ [class] represents the second self-attention feature encoding value of the view-independent information, and s′ n / 2 represents the second self-attention feature encoding value of the n / 2-th element in the sampling result.

[0124] The self-attention network model can establish connections between multiple perspective views through correlations, further aggregating the feature information of the multiple perspective views retained by sampling onto several representative perspective views, and reducing the redundancy of the feature information again.

[0125] S8: Construct a global feature descriptor based on the first self-attention feature encoding and the second self-attention feature encoding;

[0126] After performing max-pooling operations on the first self-attention feature encoding and the second self-attention feature encoding respectively, they are concatenated into a global feature descriptor. The calculation formula is:

[0127] F global =Cat(Max(X last ), Max(S last ))

[0128] In the formula, F global represents the global feature descriptor, Cat represents the concatenation operation, and Max represents the max-pooling operation

[0129] S9: Perform classification and retrieval based on the global feature descriptor to obtain the recognition result for identifying the 3D object.

[0130] Such as Figure 3As shown, when performing the classification task, the blank rectangular blocks represent the feature encodings of the perspective views, the asterisk rectangular blocks represent the feature encodings of the view-independent information, and the numbered rectangular blocks represent the position encodings of the perspective views; the global feature descriptor is input into the fully connected layer and the normalization layer to obtain the classification result as the recognition result; when performing the retrieval task, the global feature descriptor F of the 3D object to be recognized is calculated global and the global feature descriptor F of the known 3D object global The Euclidean distance is calculated, and the known 3D object with the smallest Euclidean distance is used as the retrieval result to obtain the recognition result of the 3D object to be recognized.

[0131] In the specific implementation process, the performance of the method proposed in this embodiment is verified on three public datasets, namely modelnet40, modelnet10, and shapenet55. Among them, modelnet40 consists of 12,311 3D shapes in 40 categories, including 9,843 training objects and 2,468 test objects for shape classification, with different numbers of shapes in different categories. And modelnet10 is a subset of modelnet40. The Shapenet55 dataset contains 51,162 3D models, which are divided into 55 classes and 204 subclasses. Among the 51,162 3D models, the training set, validation set, and test set are 70% (35,764), 10% (5,133), and 20% (10,265) respectively.

[0132] The method proposed in this embodiment (Ours) is compared with voxel-based methods, point cloud-based methods, and view-based methods. The comparison results on the shapenet55 dataset are shown in the following table: The algorithms for comparison include: Kanezaki algorithm, Zhou algorithm, Tatsuma algorithm, Furuya algorithm, Thermos algorithm, Deng algorithm, Li algorithm, Mk algorithm, SHREC16-Su algorithm, SHREC16-Bai algorithm, RotationNet algorithm, View-GCN algorithm, and CAR algorithm; The comparison results respectively include precision (P@N), recall (R@N), F1 value (F1@N), mean average precision (mAP), and normalized discounted cumulative gain (NDCG@N) in micro-average (Micro) and macro-average (Macro); Macro-average is used to give the unweighted average of the entire dataset, and the scores are averaged with the same weight; In micro-average, the query and retrieval results are treated equally among categories, so the results are averaged without adjusting the weights according to the category size; That is to say, macro-average first calculates the average of each category and then calculates the average. While micro-average directly calculates the overall average. As can be seen from the following table, for the method (Ours) proposed in this embodiment, among the 10 micro and macro indicators, 9 indicators are the highest and one indicator is the second highest, indicating that the method proposed in this embodiment can achieve the optimal query and classification effects in both micro-average and macro-average;

[0133]

[0134]

[0135] The method proposed in this embodiment (Ours) is compared with voxel-based methods, point cloud-based methods, and view-based methods on the ModelNet40 and ModelNet10 datasets. The comparison results are shown in the following table. Among them, voxel-based (Voxels) methods include the VRN-Ensemble algorithm, the VRN (w / o Ensemble) algorithm, and the LP-3DCNN algorithm; point cloud-based (Views) methods include the RS-CNN algorithm, the SO-Net algorithm, the LDGCNN algorithm, the Point2Sequence algorithm, the PointNet++ algorithm, and the PointNet algorithm; view-based (Views) methods include the RotationNet algorithm, the RotationNet algorithm, the MVCNN-New algorithm, the View N-gram algorithm, the Adjacent Views algorithm, the RelationNetwork algorithm, the MHBN algorithm, the MLVCNN algorithm, the Wang et al. algorithm, the 3D2SeqViews

[30] algorithm, the GVCNN

[24] algorithm, the Ma et al.

[37] algorithm, the MVCNN

[11] algorithm, and the CAR-Net algorithm; the comparison results include classification accuracy and retrieval mean average precision (mAP); among them, Ours, 12× means that the method proposed in this embodiment inputs 12 perspective views, and Ours, 20× means that the method proposed in this embodiment inputs 20 perspective views. It can be seen from the table that the method proposed in this embodiment has high classification accuracy and retrieval mean average precision on the ModelNet40 and ModelNet10 datasets. When inputting 12 perspective views, the classification accuracy is only lower than that of the CAR-Net algorithm when inputting 20 perspective views; while the retrieval mean average precision is the highest among all methods; moreover, the more perspective views are input, the higher the classification accuracy and retrieval mean average precision of the method proposed in this embodiment.

[0136]

[0137] Embodiment 3

[0138] This embodiment provides a 3D object recognition system based on the self-attention mechanism to implement the 3D object recognition method based on the self-attention mechanism described in Embodiment 1 or 2, as Figure 4 shown, the system includes:

[0139] A data acquisition module, configured to acquire multi-perspective views of the 3D object to be recognized and the position information of each perspective view;

[0140] A convolutional neural network module for inputting multi-view images of the 3D object to be recognized into a pre-trained convolutional neural network for classification to obtain feature encodings and classification scores of the multi-view images;

[0141] A first embedding module for combining the position information of each view image with view-independent information and embedding it into the feature encoding of the corresponding view image to obtain a first embedded feature encoding;

[0142] A first self-attention module for inputting the first embedded feature encoding into a first self-attention network model and outputting a first self-attention feature encoding;

[0143] A sampling module for sampling the first self-attention feature encoding according to the classification score of the view image to obtain a sampling result;

[0144] A second embedding module for embedding the corresponding position information and view-independent information of the view image into the sampling result to obtain a second embedded feature encoding;

[0145] A second self-attention module for inputting the second embedded feature encoding into a second self-attention network model and outputting a second self-attention feature encoding;

[0146] A global feature construction module for constructing a global feature descriptor according to the first self-attention feature encoding and the second self-attention feature encoding;

[0147] An identification module for performing classification retrieval according to the global feature descriptor to obtain an identification result of the recognized 3D object.

[0148] The same or similar reference numerals correspond to the same or similar components;

[0149] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0150] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A 3D object recognition method based on self-attention mechanism, characterized in that Including the steps: S1: Obtain multi-view images of the 3D object to be recognized and the position information of each view image; S2: Input the multi-view images of the 3D object to be recognized into a pre-trained convolutional neural network for classification to obtain the feature encoding and classification score of the multi-view images; S3: Combine the position information of each view image with view-independent information and embed it into the feature encoding of the corresponding view image to obtain the first embedded feature encoding; S4: Input the first embedded feature encoding into the first self-attention network model and output the first self-attention feature encoding; The first self-attention network model includes N self-attention networks connected in sequence; each self-attention network includes a first normalization layer, a multi-head attention layer, a first residual connection point, a second normalization layer, a linear mapping layer, and a second residual connection point connected in sequence; The first embedded feature encoding X0 is input into the first self-attention network. After being processed by the first normalization layer and the multi-head attention layer in sequence, it is connected to X0 at the first residual connection point to obtain the intermediate feature encoding X'; X' is processed by the second normalization layer and the linear mapping layer in sequence, and is connected to X' at the second residual connection point to obtain the output X1 of the first self-attention network; Input the output X1 of the first self-attention network into the second self-attention network. According to the same method, output the output X2 of the second self-attention network until it passes through the processing of N self-attention networks; Take the output of the first self-attention network model as the first self-attention feature encoding X last ; S5: Sample the first self-attention feature encoding according to the classification score of the view image to obtain a sampling result; S6: Embed the corresponding position information and view-independent information of the view image into the sampling result to obtain the second embedded feature encoding; S7: Input the second embedded feature encoding into the second self-attention network model and output the second self-attention feature encoding; S8: Construct a global feature descriptor according to the first self-attention feature encoding and the second self-attention feature encoding; S9: Perform classification retrieval according to the global feature descriptor to obtain the recognition result of the 3D object to be recognized.

2. The 3D object recognition method based on the self-attention mechanism according to claim 1, characterized in that In the step S1, the specific method for obtaining the position information of each view image is: Number each view image according to the view order of the multi-view images of the 3D object to be recognized; encode each view image using the sine-cosine position encoding function, and use the obtained position encoding as the position information of each view image; The sine-cosine position encoding function is: PE (pos,2i) = sin(pos / 10000 2i / d ) PE (pos,2i+1) = cos(pos / 10000 2i+1 / d ) where pos represents the sequential number of the perspective view, pos = 1, 2, …, n, and n represents the number of multi-perspective views; d represents the length of the feature vector of the position encoding; i represents the i-th element in the feature vector of the position encoding, with the even positions encoded by sin and the odd positions encoded by cos, then PE (pos,2i) represents the position encoding value of the even positions of the pos-th perspective view, and PE (pos,2i+1) represents the position encoding value of the odd positions of the pos-th perspective view. According to PE (pos,2i) and PE (pos,2i+1) to form the position encoding E pos of the pos-th perspective view, which serves as the position information of this perspective view.

3. The 3D object recognition method based on the self-attention mechanism according to claim 2, wherein In the step S2, the pre-trained convolutional neural network is the VggNet network; the multi-view images of the 3D object to be recognized are input into the pre-trained VggNet network, and the feature encoding F = [f1, f2,..., f n and the classification score are output, where f n represents the feature encoding of the nth view image.

4. The 3D object recognition method based on the self-attention mechanism according to claim 3, characterized in that, In the step S3, the specific method for obtaining the first embedded feature encoding is: Set the view-independent information [class]token, and the view-independent information [class]token is a learnable vector with the same dimension as the feature encoding of the view image; then the first embedded feature encoding is: X0 = [f [class] , f1, f2, …, f n + E pos Where X0 represents the first embedded feature encoding, and f [class] represents the first feature encoding of view-independent information.

5. The 3D object recognition method based on the self-attention mechanism according to claim 4, wherein, The specific method of the step S5 is: Sort the perspective views in descending order according to their classification scores, and retain the perspective views whose classification scores are among the top n / 2; use the retained perspective views to encode the first self-attention feature X last Perform sampling, that is, filter the first self-attention feature encoding values corresponding to the retained perspective views, and the sampling result is denoted as S = [s1, s2, …, s n / 2 , s n / 2 represents the n / 2-th element in the sampling result, and n represents the number of multi-perspective views.

6. The 3D object recognition method based on the self-attention mechanism according to claim 5, characterized in that, The specific method of the step S6 is: Embed the view-independent information [class]token and the corresponding position information of the view image into the sampling result to obtain the second embedded feature encoding S0: S0 = [s [class] , s1, s2, …, s n / 2 + E pos where s [class] represents a second feature code for view-independent information.

7. The 3D object recognition method based on the self-attention mechanism according to claim 6, wherein In the step S7, the second self-attention network model has the same structure as the first self-attention network model; After the second embedded feature encoding S0 is input into the second self-attention network model for processing, the second self-attention feature encoding S is output last , S last = [s' [class] , s'1, s'2, …, s' n / 2 , where s' [class] represents the second self-attention feature encoding value of view-agnostic information, and s' n / 2 represents the second self-attention feature encoding value of the n / 2-th element in the sampling result.

8. The 3D object recognition method based on the self-attention mechanism according to claim 7, characterized in that In the step S8, the specific method for constructing the global feature descriptor is: After performing max pooling operations on the first self-attention feature encoding and the second self-attention feature encoding respectively, they are concatenated into a global feature descriptor, and the calculation formula is as follows: F global = Cat(Max(X last ), Max(S last )) where F global represents the global feature descriptor, Cat represents the concatenation operation, and Max represents the max pooling operation.

9. A 3D object recognition system based on self-attention mechanism, characterized in that, Implement the 3D object recognition method based on self-attention mechanism according to any one of claims 1-8, the system includes: A data acquisition module, configured to acquire multi-view images of the 3D object to be recognized and the position information of each view image; A convolutional neural network module, configured to input the multi-view images of the 3D object to be recognized into a pre-trained convolutional neural network for classification, and obtain the feature encoding and classification score of the multi-view images; A first embedding module, configured to combine the position information of each view image with view-independent information and embed it into the feature encoding of the corresponding view image to obtain a first embedded feature encoding; A first self-attention module, configured to input the first embedded feature encoding into a first self-attention network model and output a first self-attention feature encoding; A sampling module, configured to sample the first self-attention feature encoding according to the classification score of the view image to obtain a sampling result; A second embedding module, configured to embed the corresponding position information and view-independent information of the view image into the sampling result to obtain a second embedded feature encoding; A second self-attention module, configured to input the second embedded feature encoding into a second self-attention network model and output a second self-attention feature encoding; A global feature construction module, configured to construct a global feature descriptor according to the first self-attention feature encoding and the second self-attention feature encoding; An identification module, configured to perform classification retrieval according to the global feature descriptor to obtain an identification result for identifying the 3D object.

Citation Information

Patent Citations

  • Texture recognition method based on deep self-attention network and local feature coding

    CN113674334A

  • Case and news correlation analysis method based on multi-view integrated learning

    CN113901990A