Three-dimensional model retrieval method based on multi-view aggregation

By using visual Transformer and multi-view attention module in the three-dimensional model retrieval method, the shortcomings of multi-view acquisition and feature aggregation in the existing methods are solved, and more efficient feature extraction and retrieval performance is achieved.

CN120067376AActive Publication Date: 2025-05-30HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202411924188.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-30
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing three-dimensional model retrieval methods based on view have shortcomings in multi-view acquisition, single-view feature extraction, and multi-view feature aggregation, resulting in limited feature expression capabilities and it is difficult to extract feature representations with high discrimination.

Method used

A three-dimensional model retrieval method based on multi-view aggregation is proposed, using visual Transformer as a single-view feature extraction network, and introducing a multi-view attention module, a channel attention aggregation module and a residual channel attention module to generate a global feature representation through the information fusion of multiple views.

Benefits of technology

The feature expression ability and training optimization efficiency of three-dimensional model retrieval has been significantly improved, and higher retrieval accuracy and better generalization ability have been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067376A_ABST
    Figure CN120067376A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional model retrieval method based on multi-view aggregation, and the method comprises the steps: 1, designing a three-dimensional model retrieval network based on views, and enabling the three-dimensional model retrieval network to extract the features of a three-dimensional model from the view representation of the three-dimensional model for three-dimensional model retrieval; the three-dimensional model retrieval network MVAT comprises a single-view feature extraction network and a multi-view feature aggregation network, the single-view feature extraction network is responsible for independently extracting local features of a three-dimensional model from each view, and the multi-view feature aggregation network is responsible for generating global feature representation through information fusion of multiple views; 2, weight freezing and contrast learning technologies are introduced, and a three-stage training method is adopted to train the three-dimensional model retrieval network; and 3, performing performance evaluation on the trained three-dimensional model retrieval network on a test set. The three-dimensional model retrieval method has the advantages of being high in feature expression ability, efficient in training method, high in precision and good in generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D model retrieval, and particularly to a 3D model retrieval method based on multi-view aggregation. Background Art

[0002] Existing 3D model retrieval methods are mainly divided into two categories. One is the traditional 3D model retrieval method, and the other is the 3D model retrieval method based on deep learning. Traditional 3D model retrieval methods mainly rely on manually designed feature descriptors, and the feature expressiveness is limited, often unable to fully reflect the high-level semantic information of 3D models. The 3D model retrieval method based on deep learning can automatically learn high-level features with richer semantic information, thereby improving the accuracy and robustness of retrieval. Based on the different data representations of the input 3D model, the deep learning-based method can be further divided into four sub-categories: view-based 3D model retrieval method, point cloud-based 3D model retrieval method, voxel-based 3D model retrieval method, and multi-modal 3D model retrieval method.

[0003] Traditional 3D model retrieval methods mainly rely on manually designed feature descriptors, which support the effective retrieval of models by capturing the geometric or visual features of the models. Patent CN119066231A discloses a fast industrial 3D model retrieval method based on binary decomposition, which generates 3D model feature descriptors by recursive segmentation and surface area function. Patent CN118350072A discloses a 3D model retrieval method, device, equipment, medium and product for vehicle parts, which generates a spectrogram through a two-dimensional projection sequence to describe the features of the 3D model. CN117852152A discloses a building model retrieval method, system, equipment and storage medium based on 3D space feature matching, which uses the average voxel color as the 3D model feature descriptor to realize 3D model retrieval.

[0004] The core idea of the view-based 3D model retrieval method is: arranging multiple virtual cameras around the 3D model, taking rendered views from multiple different angles, extracting features for each view, and then aggregating these view features into a global feature representation for the 3D model retrieval task. Patent CN119089000A discloses a multi-view Figure 3 dimensional model retrieval method and system based on dual parameter fusion network, which obtains highly expressive multi-view features through the feature fusion of light multi-views, depth multi-views and thickness multi-views. Patent CN117743616A discloses an image-based 3D model retrieval method, system, equipment and storage medium, which uses an image domain network and a model domain network to fuse multi-view features for 3D model retrieval.

[0005] The 3D model retrieval method based on point cloud is a relatively mature technology, and its general paradigm mainly includes two steps: First, extract features according to the three-dimensional spatial coordinate information of each point in the point cloud, that is, encode each constituent point of the point cloud to capture its spatial geometric structure; Then, effectively aggregate these point-level features to form a global feature descriptor, providing a basis for 3D model retrieval. Patent CN117033686A discloses a 3D model retrieval method for mechanical parts based on feature extraction, using PointNet to extract features from the point cloud representation of the 3D model.

[0006] The 3D model retrieval method based on voxel extracts features from the voxel representation of the 3D model. Patent CN113052298A discloses a 3D model retrieval method based on capsule network, which converts the 3D model into discrete voxels, uses convolutional neural network and capsule network to extract features of the voxel grid, and uses the dynamic routing algorithm to optimize the weight iteration to achieve efficient 3D model retrieval.

[0007] The multi-modal 3D model retrieval method demonstrates excellent performance by fusing 3D model features from different data representations, promoting the in-depth understanding of the high-level semantic features of 3D models, thereby improving the effect of 3D model retrieval. By utilizing the complementary characteristics of multi-modal, better results than single-modal methods can be achieved in collaborative processing and information fusion. CN112270762A discloses a 3D model retrieval method based on multi-modal fusion, using PointNet to extract point cloud features, using MVCNN to extract multi-view and panoramic view features, and fusing them into a multi-modal feature for retrieval. Patent CN110163091A discloses a 3D model retrieval method based on multi-modal information fusion of LSTM network, using two-layer LSTM to fuse skeleton feature domain view features to construct the multi-modal feature of the 3D model.

[0008] Problems and defects of the prior art:

[0009] Traditional 3D model retrieval methods mainly rely on manually designed feature descriptors. The expressive power of these features is limited and cannot fully reflect the high-level semantic information of 3D models. In addition, traditional methods are less robust when faced with problems such as data noise, rotation, and scale changes, and are easily interfered by these factors, resulting in a decrease in retrieval accuracy. For 3D models in new scenarios or different retrieval requirements, traditional methods lack flexibility and usually require redesigning features or methods, which undoubtedly increases the workload and difficulty of retrieval. For example, patents CN119066231A, CN118350072A, CN117852152A, CN107807970A, CN107066547A, CN106599053A, etc. all use manually designed feature descriptors and are difficult to adapt to complex geometric structures and diverse application scenarios.

[0010] In contrast, deep learning-based 3D model retrieval methods can automatically learn and extract high-level features with more semantic information, thereby improving the accuracy and robustness of retrieval. Through training on large-scale datasets, deep learning models can better handle challenges such as noise and transformation and show good generalization ability between different datasets. However, existing deep learning-based retrieval methods still have deficiencies in feature expression ability and are difficult to extract feature representations with high discrimination. For example, the feature extraction networks used in patents CN119089000A, CN117743616A, CN117033686A, CN113052298A, CN112270762A, CN110163091A have significantly lagged behind the current technical level, and their feature extraction ability is not dominant in today's application scenarios.

[0011] The retrieval performance of deep learning-based methods is closely related to their data representation. Among them, view-based 3D model retrieval methods have become a relatively mainstream type of 3D model retrieval algorithm due to their high usage frequency, excellent comprehensive performance, and wide application scenarios. However, existing view-based 3D model retrieval methods still have certain defects. The multi-view Figure 3 3D model retrieval method and system based on a dual-parameter fusion network disclosed in patent CN119089000A extract features and calculate similarities from light multi-views, depth multi-views, and thickness multi-views respectively. This method has the problem of view information redundancy. The depth and thickness information of the 3D model is already contained in the light multi-views. Using these three as inputs simultaneously will not only multiply the computational burden of the network but also cannot effectively improve the expressive ability of the features. Moreover, when obtaining multi-views, this patent sets virtual cameras in an axisymmetric manner, and when the pose of the 3D model changes, its retrieval performance will drop significantly, lacking sufficient rotational robustness. The multi-viewFigure 3 3D model retrieval method and system, which uses the max pooling method to aggregate features of multiple views. This pooling method completely abandons non-maximum feature components, resulting in serious loss of view information and inability to fully utilize the correlation information between multiple views. In addition, these two patents respectively select the outdated FCN and CNN to extract features of a single view, and there is an obvious technological generation gap in their feature extraction capabilities compared with the current advanced ViT. Summary of the Invention

[0012] In order to overcome the deficiencies of the existing view-based 3D model retrieval methods in aspects such as multi-view acquisition, single-view feature extraction, and multi-view feature aggregation, the present invention proposes a 3D model retrieval method based on multi-view aggregation, aiming to achieve highly symmetric multi-view acquisition, efficient single-view feature extraction, and fully utilize the correlation information between each view for multi-view feature aggregation.

[0013] The present invention provides a 3D model retrieval method based on multi-view aggregation, including the following steps:

[0014] Step 1: Design a view-based 3D model retrieval network, which can extract features of a 3D model from the view representation of the 3D model for 3D model retrieval; the 3D model retrieval network MVAT includes a single-view feature extraction network and a multi-view feature aggregation network. The single-view feature extraction network is responsible for independently extracting local features of the 3D model from each view and inputting them into the multi-view feature aggregation network. The multi-view feature aggregation network is responsible for receiving the view information sent by the single-view feature extraction network and generating a global feature representation through information fusion of multiple views;

[0015] Step 2: Introduce the weight freezing and contrast learning technologies, and use a three-stage training method to train the 3D model retrieval network in Step 1;

[0016] Step 3: Perform performance evaluation on the 3D model retrieval network trained in Step 2 on the test set.

[0017] As a further improvement of the present invention, in Step 1, the multi-view feature aggregation network includes a multi-view attention module, a channel attention aggregation module, and a residual channel attention module connected in sequence. The multi-view attention module is used to focus on the relevant information between different perspectives through the self-attention mechanism to enhance the network's perception of important features. The channel attention aggregation module is used to further improve the effect of feature fusion by weighting the features of different channels. The residual channel attention module is used to introduce a residual connection to solve the problem of channel information loss and improve the feature expression ability.

[0018] As a further improvement of the present invention, in the first step, a Vision Transformer is selected as the single-view feature extraction network. At the same time, a weight sharing mechanism is introduced to cope with the ability of view rotation changes. Each view is extracted with features by a group of Vision Transformers with shared weights. Through the self-attention mechanism of the Vision Transformer, the single-view feature extraction network can not only flexibly capture the local detail information in the image, but also pay attention to the overall structure and semantic associations of the image, so as to extract high-level features with global expression ability.

[0019] As a further improvement of the present invention, both the input and output of the multi-view attention module are feature tensors with a shape of N×V×C, which are concatenated by a set of view feature sets with a batch size of N, a view number of V, and a feature channel number of C;

[0020] The multi-view attention module includes a global attention module, a surface attention module, and a vertex attention module connected in sequence. The global attention module, the surface attention module, and the vertex attention module capture the correlation information between features from the global scope, local surface, and vertex level respectively. The global attention module, the surface attention module, and the vertex attention module all adopt an architecture based on the Transformer encoder, and through a hierarchical focusing method, the multi-view attention module obtains comprehensive and detailed feature expressions.

[0021] As a further improvement of the present invention, the global attention module introduces a standard Transformer encoder to process the input view feature set. During this process, the self-attention mechanism is used to calculate the correlation between each view feature, and no mask is set to ensure that the features of all views are treated equally, so as to capture the complete global dependencies. And in this way, the features of a single view are identified and the feature information of other views is fused through adaptive weights, so as to generate a new set of feature tensors;

[0022] When the surface attention module extracts the correlation information between view features, it only calculates the self-attention for the co-planar view features in the same group to capture the deep correlation between the views in the group, while no self-attention calculation is performed between different groups, avoiding introducing irrelevant features while reducing the computational complexity;

[0023] When the vertex attention module calculates the attention, it only performs self-attention calculation within the co-point views in the group, while the cross-group self-attention calculation is omitted.

[0024] As a further improvement of the present invention, the channel attention aggregation module aggregates the features in the view dimension V to generate global features with a shape of N×C. The specific steps are as follows:

[0025] Step S1: Using the feature aggregation method of max pooling and average pooling, aggregate the input features along the view dimension V respectively to obtain two pooled features with the shape of N×C. Max pooling is used to capture the significant values in the features, and average pooling is used to provide global statistical information;

[0026] Step S2: Input the max pooling and average pooling features obtained in Step S1 into the MLP layer with shared weights for processing. After the MLP layer processing, add the features obtained by different pooling methods to perform information fusion; the MLP optimizes the pooled features through the process of compression and restoration;

[0027] Step S3: Perform a non-linear transformation on the fused features in Step S2 in combination with the activation function to further enhance the effective components in the features and suppress the irrelevant components. At the same time, normalize the value range of the feature tensor and adjust the output to a range that is easy to perform similarity measurement. Finally, obtain a feature tensor with the shape of N×C for similarity calculation in subsequent tasks.

[0028] As a further improvement of the present invention, the channel attention aggregation module introduces the tanh function as the activation function. Through the tanh function, the feature distribution is in all the same-sign subspaces in the C-dimensional space, increasing the diversity of the feature distribution, and expanding the included angle range between features to [-180°, 180°], thereby enhancing the discrimination ability of the angle-based similarity measurement method.

[0029] As a further improvement of the present invention, in the residual channel attention module, the output feature z after introducing the residual connection is expressed as:

[0030] z = x + y = x * (1 + tanh[MLP(x)]) (3)

[0031] where x is the input of the module and y is the output of the channel attention;

[0032] Through this design, the change of the feature scale is effectively limited within the interval with a mean of 1 and a value range of (0, 2), thereby ensuring the statistical stability of the feature components in the multi-stage cascade, as shown in Equation (4):

[0033] E(|z i |) ≈ E(|x i |) (4)

[0034] where E(|z i |) and E(|x i |) respectively represent the expectations of the z and x feature component scales.

[0035] As a further improvement of the present invention, in the second step, the three-stage training method includes the following steps:

[0036] Step y1, stage one: Remove the multi-view feature aggregation network, connect a classification head behind the single-view feature extraction network, and train the single-view feature extraction network separately with the single-view classification task.

[0037] Step y2, stage two: Remove the classification head added in the first stage, reconnect the multi-view feature aggregation network behind the single-view feature extraction network. At the same time, freeze the weights of the single-view feature extraction network, and train the multi-view feature aggregation network separately with the contrastive learning method.

[0038] Step y3, stage three: Unfreeze the weights of the single-view feature extraction network frozen in the second stage, and train the complete network using the same contrastive learning method as in the second stage.

[0039] The beneficial effects of the present invention are as follows: 1. Strong feature expression ability - Based on the self-attention mechanism, a multi-view attention module is proposed, including three sub-modules: global attention, surface attention, and vertex attention. These modules capture the correlation information between views from different spatial levels respectively, and adopt a step-by-step refinement strategy from global to local and from coarse-grained to fine-grained, which improves the description ability of multi-view correlation information. Based on the channel attention mechanism, a channel attention aggregation module and a residual channel attention module are proposed. The channel attention aggregation module combines the advantages of max pooling and average pooling, effectively reducing the loss of view information in feature aggregation. At the same time, through the compression and expansion of feature dimensions, noise is further filtered. The residual channel attention module utilizes the characteristics of residual connections, deepens the network depth through multi-stage cascading, and mines the deep semantic information of the aggregated features; 2. Efficient training method - Aiming at the limitations of traditional two-stage training methods, a three-stage training method is designed by combining weight freezing and contrast learning techniques. In the first stage, the single-view feature extraction network is independently trained using single-view classification tasks; in the second stage, the weights of the single-view feature extraction network are frozen, and the multi-view feature aggregation network is independently trained using contrast learning; in the third stage, the weights of the single-view feature extraction network are unfrozen, and the entire network is jointly trained using contrast learning. This three-stage training method ensures the full training of each component of MVAT. Aiming at the characteristics of the 3D model retrieval task, the symmetric cross-entropy loss is adaptively improved, and the absolute cosine similarity is designed. The improved symmetric cross-entropy loss ensures the consistency of the loss scale in the three-stage training; the absolute cosine similarity enhances the discrimination ability of different category features, adjusts the feature distribution to a more stable state, and improves the applicability of contrast learning; 3. High accuracy and good generalization - The retrieval performance of MVAT is evaluated on three large-scale datasets, ModelNet10, ModelNet40, and MCB-B, and visualization results are given. The experimental results show that MVAT is superior to other mainstream 3D model retrieval methods in all indicators, which can prove that the method proposed by the present invention has high accuracy and good generalization. Description of the Drawings

[0040] Figure 1 is the network structure of MVAT of the present invention;

[0041] Figure 2 is the position and pose of the virtual camera of the present invention;

[0042] Figure 3 is the network structure of MVA of the present invention;

[0043] Figure 4 is the network structure of the MVA sub-module of the present invention;

[0044] Figure 5It is the network structure of the FSA of the present invention;

[0045] Figure 6 It is the network structure of the VSA of the present invention;

[0046] Figure 7 It is the network structure of the CAA of the present invention;

[0047] Figure 8 It is the network structure of the RCA of the present invention;

[0048] Figure 9 It is the first stage of the three - stage training method of the present invention;

[0049] Figure 10 It is the second and third stages of the three - stage training method of the present invention;

[0050] Figure 11 It is the test method of the MVAT of the present invention;

[0051] Figure 12 It is the top 10 retrieval results of MVAT on the ModelNet40 dataset of the present invention;

[0052] Figure 13 It is the top 10 retrieval results of MVAT on the MCB - B dataset of the present invention. Detailed implementation manners

[0053] Traditional 3D model retrieval methods mainly rely on manually designed feature descriptors. However, these descriptors are limited by the prior knowledge of the designers and are difficult to effectively handle complex geometric structures and diverse data representations. In contrast, deep - learning - based 3D model retrieval methods break through the bottlenecks of traditional methods through end - to - end feature learning and optimization, get rid of the dependence on manually designed features, and achieve layer - by - layer abstraction and expression of features through multi - layer network structures. Among them, view - based 3D model retrieval methods have become the mainstream retrieval methods due to their high application frequency, excellent comprehensive performance, and wide applicable scenarios. However, existing view - based 3D model retrieval methods fail to fully exploit the correlation information between view features and have certain limitations in improving feature distinctiveness. To solve these problems, the present invention proposes a 3D model retrieval method based on multi - view aggregation, aiming to achieve more distinctive feature expression and more efficient training optimization, thus significantly improving the retrieval performance of 3D models.

[0054] The present invention discloses a 3D model retrieval method based on multi - view aggregation, comprising the following steps:

[0055] Step 1: Design a 3D model retrieval network - MVAT (Multi-View Aggregation Transformer) for 3D model retrieval;

[0056] Step 2: Use a three-stage training method to train the 3D model retrieval network;

[0057] Step 3: Evaluate the performance of the trained 3D model retrieval network on the test set.

[0058] 1. Network Structure of MVAT

[0059] The present invention designs a view-based 3D model retrieval network - MVAT, which can extract the features of 3D models from the view representations of 3D models for 3D model retrieval. Its network structure is as Figure 1 shown. MVAT can be generally divided into two major parts: single-view feature extraction network and multi-view feature aggregation network. The single-view feature extraction network is responsible for independently extracting the local features of 3D models from each view, while the multi-view feature aggregation network generates a global feature representation through information fusion of multiple views, thereby improving the retrieval accuracy. The multi-view feature extraction network can be further divided into three key modules: multi-view attention module (Multi-View Attention, MVA), channel attention aggregation module (Channel Attention Aggregation, CAA), and residual channel attention module (Residual Channel Attention, RCA). MVA focuses on the relevant information between different perspectives through the self-attention mechanism, enhancing the network's perception of important features; CAA further improves the effect of feature fusion by weighting the features of different channels; and RCA introduces residual connections, solving the problem of channel information loss and effectively improving the feature expression ability. Next, according to the above network structure division, the design details of each part and its role in the entire network will be elaborated in detail.

[0060] 1.1 Acquisition of Multi-Views

[0061] The present invention selects a regular dodecahedron as the virtual camera configuration framework to obtain richer symmetric view information. The position and attitude of the virtual camera are related to Figure 2As shown in the figure. A regular dodecahedron has twenty vertices, and each vertex is adjacent to three adjacent edges. To render the view images, the 3D model is enclosed inside a regular dodecahedron, and three virtual cameras are arranged along the adjacent edge directions at each vertex. The front vector of each virtual camera points to the center of the regular dodecahedron, and the up vector is perpendicular to the front vector and coplanar with the front vector and one of the adjacent edges of the vertex. With this configuration, sixty virtual cameras jointly render sixty view images from different perspectives, and these images will be used as the input of MVAT. In addition, in order to ensure clarity while achieving view lightweighting, this method utilizes the symmetry of the regular dodecahedron, which not only ensures the balanced distribution of perspectives but also can effectively capture various view features of the 3D model, enhancing the rotational robustness during retrieval. During the rendering process, a supersampling anti-aliasing algorithm is also adopted: First, render at 16 (4×4) times the resolution, then use the Gaussian smoothing algorithm to remove high-frequency signals, and finally scale the view to the target size through the downsampling algorithm. In this way, the rendered views not only have clear details but also do not impose too much burden on the calculation of MVAT.

[0062] 1.2 Single-View Feature Extraction Network

[0063] Existing 3D model retrieval methods often use Convolutional Neural Network (CNN) to extract the features of a single view. However, CNN has some limitations that cannot be ignored. It mainly focuses on local receptive fields and is difficult to effectively capture global information. In addition, when processing different regions of an image, CNN usually applies the same convolution operation to all positions, unable to distinguish the importance of different image regions, which limits the understanding of the global semantics of the image. To solve the above problems, the present invention selects Vision Transformer as the single-view feature extraction network for its strong representation ability and the advantage of capturing global features in image processing tasks. At the same time, to cope with the ability of view rotation changes, the present invention introduces a weight sharing mechanism. Each view is extracted with features by a group of Vision Transformers sharing weights. Through the self-attention mechanism of Vision Transformer, the network can not only flexibly capture the local detailed information in the image but also pay attention to the overall structure and semantic associations of the image, thus extracting high-level features with global expression ability.

[0064] 1.3 Multi-View Attention Module

[0065] After extracting the features of each single view through the Vision Transformer, a set consisting of sixty view feature tensors is obtained. These feature tensors correspond to different perspectives of the 3D model respectively, comprehensively reflecting the geometric structure of the 3D model in space. However, these single-view features are independent of each other and cannot directly reflect the correlation relationships and global information among multiple views. To overcome the above problems, the present invention proposes a multi-view attention module (Multi-View Attention, MVA), as Figure 3 shown. Its input and output are both feature tensors of shape N×V×C, which are concatenated by a set of view feature sets with a batch size of N, a view number of V, and a feature channel number of C.

[0066] MVA consists of three sub-modules: global attention, face attention, and vertex attention. According to the hierarchical analysis of the view distribution, these sub-modules capture the correlation information between features from the global scope, local surface, and vertex level respectively. All three sub-modules adopt an architecture based on the Transformer encoder, and their common structure is as Figure 4 shown, where MLP is the Multi-Layer Perceptron and Layer Norm is the Layer Normalization. Although the basic architectures of the three are similar, they use different masks during self-attention calculation, respectively focusing on capturing the correlation information at different levels of the 3D model, thus differentiating into three different self-attention mechanisms. Through this hierarchical focusing method, MVA can obtain a more comprehensive and detailed feature representation.

[0067] 1.3.1 Global Attention

[0068] The spatial distribution of views exhibits a high degree of symmetry, which makes there more or less a certain correlation relationship between every two views. To deeply explore the potential correlation information between each view and all other views, a standard Transformer encoder is introduced to process the input view feature set. During this process, the self-attention mechanism is used to calculate the correlation between view features, and no mask is set to ensure that the features of all views are treated equally, so as to capture the complete global dependency relationship. In this way, not only can the features of a single view be identified, but also the feature information of other views can be effectively fused through adaptive weights, thus generating a new and more expressive set of feature tensors. The present invention names this MVA sub-module Global Attention (GA).

[0069] 1.3.2 Face Attention

[0070] After extracting the correlation information between view features at the global level, the feature correlation in the local spatial structure is equally crucial. Utilizing this local correlation can more comprehensively mine the information hidden in view features, thereby enhancing the quality of overall feature representation and retrieval performance. As mentioned before, the spatial distribution of virtual cameras adopts a dodecahedron structure, which has high symmetry and regularity in three-dimensional space. Among them, each face of the dodecahedron is a regular pentagon, and there are three virtual cameras at each vertex of each pentagonal face. And there must be a special virtual camera among them, and the plane determined by its front vector and up vector passes through the central normal of the pentagonal face, thus forming a view distribution method with geometric characteristics. To make full use of this feature, the sixty views on the dodecahedron are divided into twelve groups, each group contains five views, corresponding to a regular pentagonal face of the dodecahedron. The plane determined by the front vector and up vector of the virtual camera of each view within the group intersects at the central normal of the pentagonal face, forming a highly correlated spatial arrangement. When extracting the correlation information between view features, self-attention is only calculated for the view features within the same group to capture the deep correlation between views within the group, while self-attention is not calculated between different groups, reducing the computational complexity and avoiding introducing irrelevant features at the same time, such as Figure 5 shown. In the figure, the regular pentagon represents a face of the dodecahedron, each black dot represents a vertex of the dodecahedron, each arrow represents the up vector of the virtual camera, and the multiplication sign represents matrix multiplication. The self-attention mechanism focusing on the pentagonal face in the present invention is called Face Self-Attention (FSA), and the MVA sub-module designed based on FSA is called Face Attention (FA).

[0071] 1.3.3 Vertex Attention

[0072] On the basis of establishing the correlation information of coplanar view features, the correlation information between view features at other levels is also worthy of in-depth research and exploration. In the view generation method described before, three virtual cameras are placed at each vertex of the dodecahedron. The front vectors of these virtual cameras are the same, but the up vectors are different. Through this setting, three views are rendered at each vertex. There is a certain correlation between these three views, but they are also different due to the perspective differences. Drawing on the design idea of FSA, the sixty view features on the dodecahedron are divided into twenty groups, each group contains three views, corresponding to a vertex of the dodecahedron. The virtual cameras of these three views share the same vertex and naturally have a high degree of correlation. When calculating the attention, self-attention is only performed within the group, and the cross-group self-attention calculation is omitted. This method not only reduces the computational complexity but also effectively avoids the introduction of irrelevant information, such as Figure 6As shown. The present invention refers to this self-attention mechanism based on the co-point vertex view features as vertex self-attention (VSA), and the MVA sub-module designed based on VSA is called vertex attention (VA).

[0073] 1.4 Channel Attention Aggregation Module

[0074] The present invention transfers the channel attention mechanism to the field of 3D model retrieval, and makes adaptive adjustments according to the specific requirements of this field to be compatible with the heterogeneous features of shape N×V×C extracted by vision transformers and MVA. By means of the weight calculation idea in channel attention, the features are aggregated in the view dimension V to generate a global feature of shape N×C. The present invention names this module the Channel Attention Aggregation (CAA) module, and its structure is as Figure 7 shown, and the specific steps are as follows.

[0075] Step S1: First, two classic feature aggregation methods, max pooling and average pooling, are used to aggregate the input features along the view dimension V respectively to obtain two pooled features of shape N×C. Max pooling can capture the significant values in the features, while average pooling can provide global statistical information. The combination of the two can comprehensively retain the information in the input features.

[0076] Step S2: Subsequently, these two pooled features are respectively input into an MLP layer with shared weights for processing. Different from ordinary MLP, the MLP in channel attention has a unique design: the channel dimensions of its input and output remain the same, while the number of channels in the hidden layer is significantly reduced. Its purpose is to first compress the channel dimension, extract key features, and filter out redundant information and noise, and then restore the important information by expanding the channel dimension, so as to ensure that the final output features are both compact and have key expression capabilities. That is to say, the role of this MLP is similar to a filter, and through the process of compression and restoration, the pooled features are efficiently optimized. After being processed by the MLP layer, CAA adds the features obtained by different pooling methods for information fusion. In this way, CAA can integrate the complementary information provided by different pooling methods.

[0077] Step S3: Finally, a non-linear transformation is performed on the fused features in combination with an activation function to further enhance the effective components in the features and suppress the irrelevant components. At the same time, the value range of the feature tensor is normalized to adjust the output to a range that is easy to perform similarity measurement. Finally, a feature tensor of shape N×C is obtained for similarity calculation in subsequent tasks.

[0078] In the above structural design, the choice of activation function plays a crucial role in the expressiveness of the features and the final performance of the model. The Sigmoid function is usually selected as the activation function in channel attention. This function maps the input value to the interval (0, 1), which is convenient for weighting the feature components. However, the Sigmoid function has obvious limitations in the 3D model retrieval scenario, especially for the angle-based similarity measurement method (such as cosine similarity). When using different activation functions, the distribution range of the features will be significantly different. When using the Sigmoid function, since its output is only positive, the features can only be distributed in a subspace of the same sign in the C-dimensional feature space, thereby limiting the range of the angle between the features, so that the angle can only vary between (-90°, 90°). Then, features of different categories are often difficult to distance in the angle space, thereby reducing the distinguishing ability of the angle-based similarity measurement method. To solve this problem, CAA introduces the tanh function as an activation function. Similar to the Sigmoid function, tanh is also a smooth nonlinear transformation function, but its range is extended to (-1, 1). This feature enables tanh to effectively broaden the distribution range of features. h When the function is used, the features can be distributed in all subspaces of the same sign in the C-dimensional space. This not only increases the diversity of feature distribution, but also expands the angle range between features to [-180°, 180°], thereby significantly improving the distinguishing ability of the angle-based similarity measurement method.

[0079] 1.5 Residual Channel Attention Module

[0080] In CAA, efficient feature aggregation is achieved through the adaptively optimized channel attention mechanism, giving full play to the important role of different channels in feature expression. However, the potential of channel attention goes far beyond this. As a dynamic weight allocation mechanism, channel attention can generate channel weights using the internal information of features, thereby enhancing the expressiveness of important features while suppressing redundant information. This feature makes channel attention very suitable for feature optimization tasks. By introducing the channel attention mechanism, the aggregated comprehensive features can be further mapped to an optimized feature space, so that these features show higher discrimination and robustness in similarity measurement tasks, thereby significantly improving the overall performance of 3D model retrieval.

[0081] First, channel attention is applied to the output features of CAA. Let the output feature of CAA (i.e. the input feature of this module) be x, and the output feature after applying channel attention be y, then y is defined as follows:

[0082] y=x*tanh[MLP(x)] (1)

[0083] Among them, * represents element-wise multiplication. It should be noted that there are no max-pooling and average-pooling operations when the channel attention (CA) is unfolded, because feature aggregation has been completed in the CAA, so there is no need for pooling operations again.

[0084] The scale of feature components is an issue that needs to be focused on in 3D model retrieval, which directly affects the discrimination of features in similarity measurement. Examine the scale expectations E(|x i |) and E(|y i |) of each component of features x and y. Since the output range of the tanh function is the open interval (-1, 1), the scale expectation of feature components will decrease, as shown in Equation (2).

[0085] E(|y i |) < E(|x i |) (2)

[0086] The scale of features decreases layer by layer as the network deepens, resulting in features gradually gathering near the origin in the high-dimensional space. When the distribution of features tends to be concentrated, the effect of similarity measurement will be greatly weakened, which will reduce the ability of the network to distinguish different categories of 3D models and have a significant negative impact on the performance of the entire retrieval system. To effectively address this problem and ensure that the network is deep enough to extract complex semantic features, the residual connection mechanism proposed in ResNet is worth learning from. The residual connection introduces a shortcut that directly connects the input and output, enabling the input features to be directly retained and superimposed on the output features. This design not only alleviates the problem of gradient disappearance but also effectively avoids excessive shrinkage of feature scales, thus maintaining the discriminability of features in deeper networks. The output feature z after introducing the residual connection can be expressed as:

[0087] z = x + y = x * (1 + tanh[MLP(x)]) (3)

[0088] Among them, x is the input of the module, and y is the output of the channel attention;

[0089] Through this design, the change of feature scale is effectively limited within the interval with a mean of 1 and a value range of (0, 2), thus ensuring the statistical stability of feature components in multi-level cascades, as shown in Equation (4).

[0090] E(|z i |) ≈ E(|x i |) (4)

[0091] Among them, E(|z i |) and E(|x i |) respectively represent the scale expectations of the z and x feature components.

[0092] Benefiting from this scale stability design, introducing a multi-stage cascaded structure allows the depth of the network to be further increased, enabling it to capture more complex and abstract features. Such deep feature representation is particularly crucial for similarity measurement, as it can more accurately depict the subtle differences between input data, thereby significantly enhancing the network's performance in retrieval tasks. The network module formed by cascading the above structures through multiple stages in the present invention is called a Residual Channel Attention (RCA) module, and its individual constituent unit is called an RCA block. The structures of RCA and its constituent units are as Figure 8 shown.

[0093] 2 Training and Testing Methods of MVAT

[0094] 2.1 Three-Stage Training Method

[0095] The traditional two-stage training method is as follows:

[0096] (1) Stage 1: Remove the multi-view feature aggregation network, connect a classification head behind the single-view feature extraction network, and train the single-view feature extraction network separately with a single-view classification task.

[0097] (2) Stage 2: Remove the classification head added in Stage 1, reconnect the multi-view feature aggregation network behind the single-view feature extraction network, and train the complete network with a multi-view classification task.

[0098] The above two-stage training method has two significant drawbacks. Firstly, in Stage 2, the parameters of the multi-view feature aggregation network are randomly initialized. This disordered initialization method may lead to the generation of low-quality or even incorrect feature representations when the network aggregates multi-view features. These poor features will mislead the overall update direction of the network, further negatively affecting the weights of the single-view feature extraction network that has been finely trained in the first stage, thereby reducing the overall efficiency and effectiveness of feature extraction. Secondly, in Stage 2, the overall network is trained through a classification task. Although the classification task can extract features with strong expressive power, these features are often more suitable for classification requirements and may not necessarily meet the specific requirements of the retrieval task. The retrieval task pays more attention to the intra-class consistency and inter-class distinguishability of features, that is, the feature similarity of samples in the same class should be significantly higher than that between samples of different classes. This property is often difficult to be fully guaranteed in the classification task.

[0099] To overcome the inherent defects of the traditional two-stage training method, the present invention introduces weight freezing and contrast learning techniques and proposes a three-stage training method. The specific steps are as follows:

[0100] Step y1, Phase 1: Remove the multi-view feature aggregation network, connect a classification head behind the single-view feature extraction network, and train the single-view feature extraction network independently with the single-view classification task, as Figure 9 shown.

[0101] Step y2, Phase 2: Remove the classification head added in Phase 1, reconnect the multi-view feature aggregation network behind the single-view feature extraction network. At the same time, freeze the weights of the single-view feature extraction network, and train the multi-view feature aggregation network independently with the contrastive learning method, as Figure 10 shown.

[0102] Step y3, Phase 3: Unfreeze the weights of the single-view feature extraction network frozen in Phase 2, and train the complete network with the same contrastive learning method as in Phase 2, as Figure 10 shown.

[0103] 2.1.1 Contrastive Learning

[0104] The contrastive learning method proposed by the present invention is as Figure 10 shown, and the specific steps are as follows:

[0105] Step 1: In each training, first extract a batch of data of a fixed size from the training set to ensure that the number of samples in the batch is large enough to maintain the diversity and statistical significance of the sample pairs.

[0106] Step 2: Then, use MVAT to extract the feature vectors of each sample. After extracting the features, traverse all possible sample feature pairs (F i , F j ) in this batch of data, calculate the similarity s ij between them, and form a similarity matrix S. At the same time, generate a ground-truth matrix GT with the same shape as S according to the sample categories to label whether each pair of samples belongs to the same category.

[0107] Step 3: Finally, calculate the loss function by comparing the differences between S and GT, and use the optimizer to update the network weights, thus completing one training iteration. Among them, GT is defined as: the matrix elements of the same-class sample pairs take the value of 1, and the matrix elements of the different-class sample pairs take the value of 0, as shown in Equation (5), where C i and C j respectively represent the categories of the i-th and j-th samples in the batch.

[0108]

[0109] 2.1.2 Symmetric Cross-Entropy Loss

[0110] The loss function is one of the key factors determining the optimization of neural network weights, directly affecting the learning efficiency and final performance of the model. In the first stage of the three-stage training method, the single-view feature extraction network is trained separately using a classification task, so the cross-entropy loss L commonly used in classification tasks is adopted as the loss function for optimization. The cross-entropy loss guides the network to learn more accurate feature representations by measuring the difference between the predicted distribution and the true distribution. However, in the second and third stages, the training method changes from a classification task to contrastive learning. At this time, simply relying on the traditional cross-entropy loss is no longer applicable. To better handle the contrastive learning task, the present invention uses the symmetric cross-entropy loss as the loss function. By calculating the cross-entropy loss for the rows and columns of the similarity matrix respectively and taking the average of the two as the final loss value, the bidirectional matching between image features and text features is considered simultaneously during the optimization process. The definition of the symmetric cross-entropy loss L is as follows: ce As the loss function for optimization. The cross-entropy loss guides the network to learn more accurate feature representations by measuring the difference between the predicted distribution and the true distribution. However, in the second and third stages, the training method changes from a classification task to contrastive learning. At this time, simply relying on the traditional cross-entropy loss is no longer applicable. To better handle the contrastive learning task, the present invention uses the symmetric cross-entropy loss as the loss function. By calculating the cross-entropy loss for the rows and columns of the similarity matrix respectively and taking the average of the two as the final loss value, the bidirectional matching between image features and text features is considered simultaneously during the optimization process. The symmetric cross-entropy loss L sce is defined as follows:

[0111]

[0112] where, L ce (GT, S) represents the cross-entropy loss calculated along the row direction of the matrix, and L ce (GT T , S T ) represents the cross-entropy loss calculated along the column direction of the matrix. Since the 3D model retrieval task belongs to a homogeneous modality retrieval task, the true value matrix GT and the similarity matrix S are both symmetric matrices. Therefore, the cross-entropy losses calculated along the rows and columns are necessarily equal and can be combined into a unified loss term.

[0113] 2.1.3 Absolute Cosine Similarity

[0114] Cosine similarity has certain instability during the training process, mainly reflected in the restriction on the geometric distribution of features during the optimization process. Specifically, for cosine similarity, when the similarity between two features is high, its value is close to 1, meaning the angle between the features is close to 0°, that is, the features tend to be collinear in the same direction; on the contrary, when the similarity is low, its value is close to -1, indicating that the angle between the features is close to ±180°, that is, the features tend to be collinear in the opposite direction. Therefore, after the weight optimization is completed, similar features usually tend to be collinear in the same direction, while dissimilar features tend to be collinear in the opposite direction. However, when the number of feature categories is greater than 2, this optimization result cannot be achieved geometrically. In other words, for the multi-category feature distribution, it is impossible for features to satisfy the requirements of multiple pairs of collinearity in the opposite direction simultaneously. Due to the geometric conflict that cannot be completely resolved, the features can only reach a dynamic balance state in the end. Although this dynamic balance can complete the task to a certain extent, it may reduce the convergence efficiency and performance of the model, and at the same time limit the applicability of cosine similarity in multi-category scenarios.

[0115] To address the inherent deficiencies of traditional cosine similarity in specific scenarios, the present invention proposes an improved similarity metric function - absolute cosine similarity, which is defined as shown in Equation (7). By taking the absolute value of the cosine similarity, absolute cosine similarity effectively solves the geometric constraint problem that occurs during the optimization process of cosine similarity. Specifically, under the measurement of absolute cosine similarity, when the similarity between two features is high, the similarity value is close to 1, and the included angle between the corresponding feature vectors is close to 0° or ±180°, that is, these two feature vectors tend to be collinear; while when the similarity between two features is low, the similarity value is close to 0, and at this time the included angle between the feature vectors is close to ±90°, that is, the feature vectors are perpendicular to each other. This geometric relationship has a clear geometric meaning in the high-dimensional feature space, and as long as the condition that the feature dimension is greater than or equal to the number of categories is satisfied, it can be achieved through optimization. Since the feature vectors of the same class tend to be collinear and the feature vectors of different classes tend to be perpendicular to each other, a stable static balance is finally formed between the features, thereby improving the performance in the retrieval task.

[0116]

[0117] 2.2 Testing Method

[0118] After completing the training of the 3D model retrieval network, the next step is to evaluate the performance of the trained network on the test set. Performing the 3D model retrieval task is divided into two stages: the preprocessing stage and the retrieval stage. When testing the 3D model retrieval network, it is also necessary to operate according to these two stages, and its process is as Figure 11 shown. The specific steps are as follows:

[0119] Step 1, preprocessing stage: Before the test starts, first use the trained MVAT to extract features from the multi-views of each 3D model in the test set, and persistently store all the extracted features in a 3D model database, thereby improving the retrieval efficiency and avoiding repeated calculation of features during each retrieval.

[0120] Step 2, retrieval stage: In each test, select a 3D model from the test set as the query. Next, use the 3D model database generated in the preprocessing stage to calculate the absolute cosine similarity between the features of this query and the features of all 3D models in the database. According to the similarity size, sort all the 3D models in the test set in descending order from high to low, and the sorted 3D model sequence is the retrieval result of this query.

[0121] 3 Experimental Results and Analysis

[0122] 3.1 Datasets and Evaluation Metrics

[0123] 3.1.1 ModelNet Dataset and Its Evaluation Metrics

[0124] The ModelNet dataset contains 127,915 high-quality 3D models across 662 categories, which are widely used in fields such as computer vision, computer graphics, robotics, and cognitive science. It plays an important role especially in research related to 3D model analysis and processing. Its comprehensiveness and wide coverage of categories make it one of the important benchmark datasets in the field of 3D model retrieval. However, due to the extremely large scale of the entire ModelNet dataset, directly using the entire set for experiments poses significant challenges in terms of computational resources and time costs. Therefore, to more conveniently evaluate the performance of 3D model retrieval networks, researchers constructed two commonly used subsets from the ModelNet dataset: ModelNet10 and ModelNet40. The ModelNet10 dataset contains 4,899 3D models distributed across 10 categories, with 3,991 for training and 908 for testing. These categories are all common object categories and are highly representative. The ModelNet40 dataset further expands to 40 categories, containing a total of 12,311 3D models, with 9,843 for training and 2,468 for testing. This dataset covers a more diverse range of object categories and is suitable for more complex and extensive 3D model research tasks. In the evaluation system of ModelNet, the core metric for evaluating retrieval performance is mAP.

[0125] 3.1.2 MCB-B Dataset and Its Evaluation Metrics

[0126] The MCB dataset is an important dataset that has received much attention in 3D vision tasks in recent years. This dataset is divided into two subsets: MCB-A and MCB-B. Among them, the MCB-B dataset, with its moderate data volume, diverse rotation postures, and rich evaluation metric system, has become the preferred test benchmark for many 3D model retrieval networks to comprehensively evaluate the performance of 3D model retrieval. The MCB-B dataset focuses on mechanical parts, covering 25 categories of mechanical parts, with a total of 18,038 3D models collected. Among them, 14,451 are used for training and 3,587 are used for testing, providing a good verification environment for the generalization ability of the network under different conditions. In its evaluation system, the core retrieval performance metrics include: F1-score, mAP, and NDCG. In the calculation of the metrics of the MCB-B dataset, each evaluation metric is further divided into two versions: micro and macro, depending on whether it is weighted based on the data scale of each category when calculating the average value. The micro-average treats each query and its retrieval results equally and directly calculates the average value on the entire dataset; while the macro-average first calculates the average value within each category and then takes the inter-category average of the intra-category averages of all categories.

[0127] 3.2 Experimental Details and Hyperparameter Settings

[0128] In the experiment, an advanced vision Transformer - MaxViT was selected as the single-view feature extraction network of MVAT. To ensure that MaxViT can exert its powerful feature extraction ability in the experiment and avoid the problem that it is difficult to converge when the vision Transformer is trained from scratch, in the experiment, MaxViT was pre-trained with the ImageNet-1K dataset. In the experiment, the pixel size of each view was set to 224×224, and the background color was uniformly set to black to reduce the interference of background noise on feature extraction. At the same time, to reduce information redundancy while ensuring the feature extraction ability, the size of the feature dimension output by the single-view feature extraction network was set to 512. This dimension not only serves as the input feature dimension of the multi-view feature aggregation network but also remains unchanged after passing through RCA, CAA, and RCA, and finally becomes the feature dimension output by the network. In RCA, the number of cascaded RCA blocks is an important parameter affecting the network performance, and this value was set to 16 in this experiment.

[0129] During the training process, Adam, which is widely used and has excellent performance in deep learning, is selected as the optimizer. Its stability and efficiency in dealing with complex models have been widely verified. In the three-stage training method proposed in the present invention, the number of training rounds in each stage is set to 50 to achieve a balance between training efficiency and training quality. Regarding the setting of the learning rate, after multiple experimental attempts, a relatively ideal configuration is finally determined: in the first stage, the learning rate is set to 10 -4 ; while in the second and third stages, the learning rate is reduced to 10 -5 . This learning rate adjustment strategy can effectively promote the convergence of the model to obtain experimental results close to the optimal ones.

[0130] 3.3 Performance Evaluation Experiments

[0131] 3.3.1 Performance Evaluation Experiments on the ModelNet Dataset

[0132] According to the experimental settings described above, MVAT was trained and tested on the ModelNet10 dataset and the ModelNet40 dataset respectively to comprehensively evaluate its performance in the 3D model retrieval task. To accurately evaluate the performance of MVAT, it was compared with various 3D model retrieval methods. The performance evaluation results of MVAT and these methods on the ModelNet10 dataset and the ModelNet40 dataset are shown in Table 1 and Table 2 respectively.

[0133] Table 1 Performance Evaluation Experiments on the ModelNet110 Dataset

[0134]

[0135] Table 2 Performance Evaluation Experiments on the ModelNet140 Dataset

[0136]

[0137] It can be seen from the experimental results that on the ModelNet10 dataset, the mAP of MVAT reached 95.1%, which is significantly better than other methods in terms of performance. Compared with the previous method SPNet, the mAP increased by 0.9%; while compared with VAM-IAM which also introduced the attention mechanism, the performance was improved by 1.5% in terms of mAP. On the ModelNet40 dataset, MVAT also achieved remarkable results, with the mAP reaching 93.2%. Compared with the method MVTN, the mAP increased by 0.3%; compared with VAM-IAM which also introduced the attention mechanism, the mAP was improved by 0.4%. These results indicate that MVAT has strong generalization ability and excellent performance in the 3D model retrieval task.

[0138] To further verify the actual performance of MVAT, several samples were randomly selected from the test set of the ModelNet40 dataset, and the top 10 retrieval results were visually displayed, as Figure 12 shown. In the figure, the 3D model used as the query and its class label are shown on the left, and the top 10 sorted retrieval results are shown on the right. Among them, the 3D models marked with a box represent the samples with retrieval errors, and the label below the box represents their true class.

[0139] It can be observed from the figure that MVAT achieved excellent retrieval accuracy. Even though there were a small number of incorrect retrievals, these incorrect results still had a certain semantic correlation with the correct results. For 3D models of some complex classes, MVAT was also able to identify high-level features in their geometric structures, thus showing high robustness and semantic perception ability. Generally speaking, whether comparing the overall performance or the visual analysis of the actual retrieval results, MVAT demonstrated significant advantages in the 3D model retrieval task and performed excellently on complex and large-scale datasets.

[0140] 3.3.2 Performance Evaluation Experiment on the MCB-B Dataset

[0141] According to the experimental settings described above, MVAT was comprehensively trained and tested on the MCB-B dataset to deeply evaluate its performance in the 3D model retrieval task. To accurately evaluate the performance of MVAT, multiple 3D model retrieval methods were selected for comparison, and the experimental results are shown in Table 3.

[0142] Table 3 Performance Evaluation Experiment on the ModelNet140 Dataset

[0143]

[0144] As can be seen from the experimental results, the performance of MVAT on the MCB-B dataset is significantly better than other methods. In terms of micro metrics, the F1-score, mAP, and NDCG of MVAT reached 89.6%, 96.1%, and 96.7% respectively, which are 22.0%, 4.8%, and 4.2% higher than the previous methods in these three metrics. In terms of macro metrics, the F1-score, mAP, and NDCG of MVAT reached 87.6%, 94.7%, and 95.3% respectively, which are 2.2%, 3.8%, and 8.2% higher than the previous methods. Different from the ModelNet dataset with diverse class distributions, the MCB-B dataset is a 3D model database focusing on mechanical parts, and such 3D models usually have highly structured and standardized characteristics. The excellent performance of MVAT on the MCB-B dataset indicates that it has strong capabilities in feature extraction and retrieval of structured models. Especially the significant improvement in micro mAP further reflects the excellent general retrieval performance of MVAT on high-frequency and easy-to-classify classes. This not only proves the superiority of MVAT in complex 3D model retrieval tasks but also shows its application potential in the engineering field.

[0145] To further evaluate the actual performance of MVAT on the MCB-B dataset, according to the same visualization method as before, several samples were randomly selected from the test set of the MCB-B dataset, and the top 10 retrieval results were shown, as Figure 13 shown.

[0146] The results show that the retrieval accuracy of MVAT on the MCB-B dataset is also excellent. Although there are a small number of misclassifications, these misclassified samples still have a high geometric similarity with the correct results. This further verifies that MVAT can demonstrate excellent 3D model retrieval capabilities whether facing the ModelNet dataset with rich classes or the relatively single-class MCB-B dataset, fully reflecting its excellent retrieval performance.

[0147] 3.4 Ablation Experiments

[0148] Detailed ablation experiments were conducted on the structure of MVAT and each part involved in the training process to verify the effectiveness of the structure and method proposed in the present invention. The dataset selected for the experiment is ModelNet40, and a simple network directly concatenated by MaxViT and max pooling, i.e., MaxViT→MaxPool, was constructed as the baseline network for the ablation experiment.

[0149] 3.4.1 Ablation Experiment of Channel Attention Aggregation Module

[0150] First, study the role of the Channel Attention Aggregation module (CAA) in the 3D model retrieval network. CAA is a typical feature aggregation module, whose function is similar to the max-pooling operation in the baseline network, but it is more flexible and efficient in design. To comprehensively evaluate the impact of different aggregation modules on the 3D model retrieval performance, the present invention replaces the max-pooling operation in the baseline network with other aggregation modules and observes the comparison of retrieval performance through experiments. On the ModelNet40 dataset, the retrieval performance of CAA and two traditional aggregation methods is shown in Table 4.

[0151] Table 4 Ablation experiments of CAA on the ModelNet140 dataset

[0152]

[0153] The experimental results show that the retrieval network using CAA as the aggregation method achieves 89.7% in the mAP metric, which is 3.7% and 0.9% higher than the retrieval networks using max-pooling and average-pooling respectively. This fully demonstrates the significant advantage of CAA in feature aggregation. This advantage stems from the unique structure of CAA, which combines the advantages of max-pooling and average-pooling. It can not only give full play to their strengths but also further optimize the feature distribution by introducing the MLP and tanh activation functions, thus achieving more efficient feature expression and aggregation.

[0154] 3.4.2 Ablation experiments on residual channel attention

[0155] The Residual Channel Attention module (RCA) further improves the feature distribution after CAA and is one of the key components of MVAT. RCA is formed by sequentially connecting multiple RCA blocks with the same structure. Each RCA block enhances the feature expression ability of the network through residual connections and channel attention mechanisms. After replacing the max-pooling layer in the baseline network with CAA, different numbers of RCA blocks are further connected in series behind it to evaluate the impact of the introduction of RCA and the number of cascaded RCA blocks on the 3D model retrieval performance. On the ModelNet40 dataset, the detailed settings and experimental results of the relevant ablation experiments are shown in Table 5.

[0156] Table 5 Ablation experiments of RCA on the ModelNet140 dataset

[0157]

[0158] The experimental results show that the introduction of RCA has a significant improvement effect on the retrieval performance. Just introducing 1 RCA block increases the mAP by 0.3%. When the number of cascaded RCA blocks reaches 16, the retrieval performance of the network achieves the most significant improvement. Compared with the baseline network that only introduces CAA, the mAP is increased by 2.2%. This result indicates that the multi-level cascaded RCA blocks can effectively improve the discrimination ability for different category features by continuously optimizing the feature distribution and mapping the features to a more discriminative feature space, which verifies the key role of the RCA module and its multi-level cascaded structure in the 3D model retrieval task.

[0159] 3.4.3 Ablation Experiments of the Multi-View Attention Module

[0160] The multi-view attention module (MVA) is a subsequent module of the single-view feature extraction network, which is used to capture the correlation information between multi-view features to provide richer and more accurate input features for the subsequent feature aggregation module CAA. MVA consists of three sub-modules, namely global attention (GA), face attention (FA), and vertex attention (VA). The design concepts of these sub-modules have different focuses, respectively modeling different levels of feature relationships from global to local. To comprehensively evaluate the individual effects and collaborative effects of each sub-module of MVA, a set of ablation experiments were designed. Based on the baseline network with CAA and RCA added, different combinations and arrangements of sub-modules were introduced respectively, and their retrieval performances on the ModelNet40 dataset were tested. The experimental results are shown in Table 6.

[0161] Table 6 Ablation Experiments of MVA on the ModelNet140 Dataset

[0162]

[0163] The experimental results show that even if any single sub-module is introduced alone, it can significantly improve the retrieval performance, and the increase in mAP is more than 0.8%. Among them, the VA sub-module has the most significant improvement, and the mAP reaches 93.0% when used alone. Among the pairwise combinations of sub-modules, the arrangement of VA→FA performs the best, with the mAP reaching 91.5%, slightly lower than the baseline network by 0.4%. When the three sub-modules are integrated and combined in different orders, the experiment further reveals the significant impact of the order on the performance. Among them, the arrangement order of GA→FA→VA performs the best, with the mAP reaching 93.2%, which is 1.3% higher than the baseline network. This result indicates that the strategy of gradually obtaining feature correlation information from global to local has important value. The introduction of MVA transforms the original 60 single-view features into a new feature set with more complex information expression ability, thus significantly enhancing the representation ability and retrieval performance of the overall features.

[0164] 3.4.4 Ablation Experiments on Contrastive Learning and Symmetric Cross-Entropy Loss

[0165] In the training process of the present invention, a contrastive learning method is used, and symmetric cross-entropy is selected as the loss function. To verify the effectiveness of contrastive learning and symmetric cross-entropy loss, multiple groups of ablation experiments are set on the ModelNet40 dataset. By training the complete MVAT using different training techniques, the influence of the use of contrastive learning and the selection of the loss function on the retrieval performance is analyzed. The experimental results are shown in Table 7.

[0166] Table 7 Ablation Experiments on Contrastive Learning and Symmetric Cross-Entropy Loss on ModelNet140 Data

[0167]

[0168] The experimental results show that compared with the non-contrastive learning method, training using the contrastive learning method can significantly improve the retrieval performance, at least increasing the mAP by 1.1%. In addition, among the contrastive learning methods, different loss functions have different effects on the performance. Among them, the symmetric cross-entropy loss performs best in all experiments. Compared with the non-contrastive learning method, the contrastive learning method using the symmetric cross-entropy loss increases the mAP by 2.0%, further verifying the superiority of the symmetric cross-entropy loss in contrastive learning.

[0169] 3.4.5 Ablation Experiments on Absolute Cosine Similarity

[0170] The selection of the similarity metric function plays a key role in the training and testing of 3D model retrieval networks. To verify the effectiveness of the absolute cosine similarity in the training and testing of 3D model retrieval networks, based on the previous experimental settings, only the similarity metric function is replaced, and a new set of ablation experiments is conducted on the ModelNet40 dataset. The experimental results are shown in Table 8.

[0171] Table 8 Ablation Experiments on Absolute Cosine Similarity on ModelNet140 Data

[0172]

[0173] The experimental results show that, compared with the traditional cosine similarity, using the absolute cosine similarity increases the mAP of MVAT in the 3D model retrieval task by 1.2%. This result fully demonstrates the advantages of the absolute cosine similarity in the training and testing processes of the 3D model retrieval task. In the training stage, the absolute cosine similarity can more effectively guide the features to be optimized in a more discriminative direction and finally form a static and stable feature distribution, thereby enhancing the model's ability to distinguish the features of different 3D models. In the testing stage, the absolute cosine similarity can more accurately measure the similarity between different 3D models, further improving the accuracy and reliability of the retrieval results.

[0174] 3.4.6 Ablation Experiments on the Three-Stage Training Method

[0175] Based on the two-stage training method, the present invention introduces the weight freezing and contrast learning techniques, effectively ensuring that all parts of the single-view feature extraction network and the multi-view feature aggregation network can be fully trained. To deeply explore the reasons behind the excellent performance of the three-stage training method and verify the necessity of each training stage, the present invention designs a new set of ablation experiments on the ModelNet40 dataset, and the experimental settings and results are shown in Table 9.

[0176] Table 9 Ablation Experiments on the Three-Stage Training Method on the ModelNet140 Dataset

[0177]

[0178] The experimental results show that the single-stage training method performs the worst, with an mAP of only 87.1%. The two-stage training method without weight freezing improves this problem to a certain extent, and the mAP is significantly increased by 4.3%. In the two-stage training method with weight freezing, the weight freezing technique is further applied, and the mAP is further increased by 1.1%. In the three-stage training method, the introduction of the weight unfreezing stage effectively improves the parameter adjustment of the network, and the final mAP is increased by 0.7% again. The three-stage training method fully combines the advantages of the traditional two-stage method and solves the deficiencies in the two-stage method by introducing the weight freezing technique. This method not only accelerates the training process but also improves the convergence performance of the network. Finally, through this method, the network achieves better overall performance, demonstrating the effectiveness and superiority of the three-stage training method in the 3D model retrieval task.

[0179] The present invention discloses a 3D model retrieval method based on multi-view aggregation, which has the advantages of achieving more discriminative feature representation and more efficient training optimization. At the same time, a multi-view attention module, a channel attention aggregation module and a residual channel attention module are proposed, as well as a three-stage training method, and high-precision retrieval performance is achieved on three datasets. The specific advantages are as follows:

[0180] 1. Strong feature representation ability - Based on the self-attention mechanism, a multi-view attention module is proposed, including three sub-modules: global attention, face attention and vertex attention. These modules capture the correlation information between views from different spatial levels respectively, and adopt a step-by-step refinement strategy from global to local, from coarse-grained to fine-grained, which improves the description ability of multi-view correlation information. Based on the channel attention mechanism, a channel attention aggregation module and a residual channel attention module are proposed. The channel attention aggregation module combines the advantages of max pooling and average pooling, effectively reducing the loss of view information in feature aggregation. At the same time, through the compression and expansion of the feature dimension, the noise is further filtered. The residual channel attention module utilizes the characteristics of residual connection, deepens the network depth through multi-stage cascading, and mines the deep semantic information of the aggregated features.

[0181] 2. Efficient training method - Aiming at the limitations of the traditional two-stage training method, a three-stage training method is designed by combining weight freezing and contrastive learning techniques. In the first stage, the single-view feature extraction network is independently trained using the single-view classification task; in the second stage, the weights of the single-view feature extraction network are frozen, and the multi-view feature aggregation network is independently trained using contrastive learning; in the third stage, the weights of the single-view feature extraction network are unfrozen, and the entire network is jointly trained using contrastive learning. This three-stage training method ensures the full training of each component of MVAT. Aiming at the characteristics of the 3D model retrieval task, the symmetric cross-entropy loss is adaptively improved, and the absolute cosine similarity is designed. The improved symmetric cross-entropy loss ensures the consistency of the loss scale in the three-stage training; the absolute cosine similarity enhances the discrimination ability of different category features, adjusts the feature distribution to a more stable state, and improves the applicability of contrastive learning.

[0182] 3. High precision and good generalization - The retrieval performance of MVAT is evaluated on three large-scale datasets, ModelNet10, ModelNet40 and MCB-B, and the visualization results are given. The experimental results show that MVAT is superior to other mainstream 3D model retrieval methods in various indicators, which can prove that the method proposed by the present invention has high precision and good generalization.

[0183] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. A 3D model retrieval method based on multi-view aggregation, characterized in that: The following steps are involved: Step 1: Design a view-based 3D model retrieval network, which can extract the features of the 3D model from the view representation of the 3D model for 3D model retrieval; the 3D model retrieval network MVAT includes a single-view feature extraction network and a multi-view feature aggregation network, the single-view feature extraction network is responsible for independently extracting local features of the 3D model from each view and inputting them into the multi-view feature aggregation network, the multi-view feature aggregation network is responsible for receiving the view information sent by the single-view feature extraction network and generating a global feature representation through information fusion of multiple views; Step 2: introducing weight freezing and contrastive learning technology, and adopting a three-stage training method to train the 3D model retrieval network of step 1; Step 3: Evaluate the performance of the 3D model retrieval network trained in step 2 on the test set.

2. The three-dimensional model retrieval method according to claim 1, characterized in that: In the step 1, the multi-view feature aggregation network includes a multi-view attention module, a channel attention aggregation module, and a residual channel attention module connected in sequence. The multi-view attention module is used to focus on relevant information between different perspectives through a self-attention mechanism to enhance the network's perception of important features. The channel attention aggregation module is used to further improve the effect of feature fusion by weighting features of different channels. The residual channel attention module is used to introduce residual connections to solve the problem of channel information loss and improve the expression ability of features. In the step 1, a visual Transformer is selected as the single-view feature extraction network. At the same time, a weight sharing mechanism is introduced to cope with the view rotation change. Each view is extracted features by a set of visual Transformers with shared weights. Through the self-attention mechanism of the visual Transformer, the single-view feature extraction network can not only flexibly capture the local detail information in the image, but also pay attention to the overall structure and semantic association of the image, thereby extracting high-level features with global expression capabilities.

3. The three-dimensional model retrieval method according to claim 2, characterized in that: The input and output of the multi-view attention module are both feature tensors of shape N×V×C, which are concatenated from a set of view feature sets with batch size N, number of views V, and number of feature channels C; The multi-view attention module includes a global attention module, a face attention module, and a vertex attention module which are connected in sequence. The global attention module, the face attention module, and the vertex attention module capture the correlation information between features from the global range, the local surface, and the vertex level, respectively. The global attention module, the face attention module, and the vertex attention module all adopt a Transformer encoder-based architecture, and through a hierarchical focusing method, the multi-view attention module obtains a comprehensive and detailed feature expression; The global attention module introduces a standard Transformer encoder to process the input view feature set. In this process, the self-attention mechanism is used to calculate the correlation between the features of each view, and no mask is set to ensure that the features of all views are treated equally, thereby capturing the complete global dependency. In this way, the features of a single view are identified and the feature information of other views is fused through adaptive weights, thereby generating a new feature tensor set. When extracting the association information between view features, the face attention module only calculates self-attention for the coplanar view features of the same group to capture the deep association between views within the group, while no self-attention calculation is performed between different groups, thereby reducing the computational complexity and avoiding the introduction of irrelevant features. When calculating attention, the vertex attention module only performs self-attention calculation on the common point views within the group, while the self-attention calculation across groups is omitted.

4. The three-dimensional model retrieval method according to claim 2, characterized in that: The channel attention aggregation module aggregates features on the view dimension V to generate global features of shape N×C. The specific steps are as follows: Step S1: Use the feature aggregation methods of maximum pooling and average pooling to aggregate the input features along the view dimension V, respectively, to obtain two pooling features of shape N×C. Maximum pooling is used to capture the salient values ​​in the features, and average pooling is used to provide global statistical information. Step S2: inputting the maximum pooling and average pooling features obtained in step S1 into the MLP layer with shared weights for processing respectively, and after the MLP layer processing, adding the features obtained by different pooling methods to perform information fusion; the MLP optimizes the pooling features through the process of compression and restoration; Step S3: Combining the activation function, the fused features of step S2 are transformed nonlinearly to further enhance the effective components in the features and suppress the irrelevant components. At the same time, the value range of the feature tensor is normalized, and the output is adjusted to a range that is easy to measure the similarity. Finally, a feature tensor with a shape of N×C is obtained for similarity calculation of subsequent tasks; the channel attention aggregation module introduces the tanh function as an activation function, and the tanh function is used to distribute the features in all subspaces with the same sign in the C-dimensional space, thereby increasing the diversity of feature distribution and expanding the angle range between features to [-180°, 180°], thereby improving the distinguishing ability of the angle-based similarity measurement method.

5. The three-dimensional model retrieval method according to claim 2, characterized in that: In the residual channel attention module, the output feature z after the introduction of the residual connection is expressed as: z=x+y=x*(1+tanh[MLP(x)]) (3)where x is the input of the module and y is the output of the channel attention; Through this design, the change of the characteristic scale is effectively limited to an interval with a mean of 1 and a value range of (0, 2), thereby ensuring the statistical stability of the characteristic component in multi-stage series connection, as shown in formula (4): E(|z i |)≈E(|x i |) (4) Among them, E(|z i |) and E(|x i |) represent the expected scales of the z and x feature components, respectively.

6. The three-dimensional model retrieval method according to claim 1, characterized in that: In step 2, the three-stage training method includes the following steps: Step y1, stage 1: remove the multi-view feature aggregation network, connect a classification head behind the single-view feature extraction network, and train the single-view feature extraction network separately using the single-view classification task; Step y2, stage 2: remove the classification head added in the stage 1, reconnect the multi-view feature aggregation network behind the single-view feature extraction network, and at the same time, freeze the weight of the single-view feature extraction network, and use the contrastive learning method to train the multi-view feature aggregation network alone; Step y3, stage three: unfreeze the weights of the single view feature extraction network frozen in stage two, and train the complete network using the same contrastive learning method as in stage two.

7. The three-dimensional model retrieval method according to claim 6, characterized in that: The contrastive learning method comprises the following steps: Step 1: First, extract a fixed-size batch of data from the training set to ensure that the number of samples in the batch is large enough to maintain the diversity and statistical significance of the sample pairs; Step 2: Use the 3D model retrieval network to extract the feature vector of each sample. After extracting the features, traverse all possible sample feature pairs (F i , F j ), calculate the similarity S between them ij , forming a similarity matrix S. At the same time, a true value matrix GT with the same shape as S is generated according to the sample category to mark whether each pair of samples belongs to the same category; Step 3: Calculate the loss function by comparing the difference between S and GT, and use the optimizer to update the network weights to complete a training iteration; Among them, GT is defined as: the matrix elements of the same category sample pairs are 1, and the matrix elements of different category sample pairs are 0, as shown in formula (5), where C i and C j Respectively represent the categories of the i-th and j-th samples in the batch:

8. The three-dimensional model retrieval method according to claim 6, characterized in that: In the three-stage training method, the symmetric cross entropy loss is used as the loss function. The cross entropy loss is calculated for the rows and columns of the similarity matrix respectively, and the average of the two is taken as the final loss value. In this way, the two-way matching between image features and text features is considered in the optimization process. The symmetric cross entropy loss L sce is defined as follows: Among them, L ce (GT, S) represents the cross entropy loss calculated along the row direction of the matrix, L ce (GT T , S T ) represents the cross entropy loss calculated along the column direction of the matrix; In the three-stage training method, the absolute cosine similarity measurement function is used to solve the geometric constraint problem of cosine similarity in the optimization process. Specifically, under the measurement of absolute cosine similarity, when the similarity between two features is high, the similarity value is close to 1, and the corresponding feature vector angle is close to 0° or ±180°, that is, the two feature vectors tend to be collinear, and when the similarity between the two features is low, the similarity value is close to 0, and the angle between the feature vectors is close to ±90°, that is, the feature vectors are perpendicular to each other.

9. The three-dimensional model retrieval method according to claim 1, characterized in that: In step 3, when testing the three-dimensional model retrieval network, operations are performed according to the following preprocessing stage and retrieval stage. The specific steps are as follows: Step 1, preprocessing stage: before the test begins, the trained 3D model retrieval network is first used to extract features from multiple views of each 3D model in the test set, and all the extracted features are persistently stored in a 3D model database, thereby improving retrieval efficiency and avoiding repeated feature calculations during each retrieval; Step 2, retrieval stage: In each test, a 3D model is selected from the test set as a query. Next, the 3D model database generated in the preprocessing stage of step 1 is used to calculate the absolute cosine similarity between the features of the query and the features of all 3D models in the database. All 3D models in the test set are sorted in descending order from high to low according to the similarity. The sorted 3D model sequence is the retrieval result of the query.

10. The three-dimensional model retrieval method according to claim 1, characterized in that: In the step 1, a multi-view acquisition step is also included, which is specifically as follows: A regular dodecahedron is selected as the virtual camera configuration framework, the 3D model is enclosed inside a regular dodecahedron, and three virtual cameras are arranged on each vertex along the adjacent edge direction. The front vector of each virtual camera points to the center of the regular dodecahedron, and the up vector is perpendicular to the front vector and coplanar with the front vector and one of the adjacent edges of the vertex. Multiple view images obtained by this configuration are used as inputs of the 3D model retrieval network. In the step of acquiring multiple views, a supersampling anti-aliasing algorithm is also used in the rendering process of the view image: first, rendering is performed at 16 (4×4) times the resolution, then a Gaussian smoothing algorithm is used to remove high-frequency signals, and finally the view is scaled to the target size through a downsampling algorithm.

Citation Information

Patent Citations

  • Three-dimensional model retrieval method

    CN106599053A

  • LIRE-based three-dimensional model retrieval method

    CN107066547A

  • Weighted matching point-based three-dimensional model retrieval method

    CN107807970A

  • Three-dimensional model retrieval method based on LSTM network multi-modal information fusion

    CN110163091A

  • Multi-view three-dimensional model retrieval method and system based on pairing depth feature learning

    CN111382300A

Cited By

  • Three-dimensional model retrieval method based on shape perception enhancement

    CN120744200A

  • A shape-awareness enhanced 3D model retrieval method

    CN120744200B

  • Three-dimensional CAD model retrieval method and device

    CN120763350A

  • Industrial three-dimensional model adaptive rendering method and system based on semantic driving

    CN121982261A

  • Semantic-driven adaptive rendering method and system for industrial 3d models

    CN121982261B