Three-dimensional model classification based on Swindow-MVGA network multi-view feature fusion
The three-dimensional model is fused through the Swin-MVGA network, and the view features are extracted using SIFT and LBP algorithms, and the RMQF method is designed to solve the problem of insufficient information in the classification of three-dimensional model and achieve higher classification accuracy and efficiency.
Patent Information
- Application Number
- CN202510569652.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-04
- Publication Date
- 2025-08-15
AI Technical Summary
In the existing three-dimensional model classification methods, single-view information is limited, making it difficult to fully and accurately represent the three-dimensional model, while multi-view methods tend to ignore the complementarity and correlation between views, affecting the accuracy and efficiency of classification.
The Swin-MVGA network is adopted to generate multiple two-dimensional views by multi-angle projection of the three-dimensional model, and the view features are extracted in combination with SIFT and LBP algorithms. The global Multi-head Self-Attention module of SwinTransformer is used for feature fusion, and the RMQF method is designed to select representative features to improve classification accuracy.
Effectively capture the global and local information of the three-dimensional model, enhance the view feature representation ability, and significantly improve the accuracy and efficiency of three-dimensional model classification through multi-view feature fusion and RMQF method.
Smart Images

Figure SMS_3 
Figure SMS_4 
Figure SMS_5
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion, which has good application in the field of three-dimensional model classification. Background Art
[0002] In recent years, computer vision and 3D geometric modeling tools have made significant progress. With continuous technological innovation, the amount of available 3D model data has shown an exponential growth trend, covering an increasingly wide range of areas, and the diversity and complexity of models are also continuing to increase. With their powerful ability to express object shape information, 3D models have been widely used in key fields such as autonomous driving, remote sensing, virtual reality, and mechanical manufacturing. This makes how to quickly and accurately identify 3D models and how to efficiently manage and utilize them an important research topic in the field of computer vision. With the rapid development of computer technology and the significant increase in computing power, a large amount of 3D model data has continued to emerge, providing an opportunity for the development of 3D model classification methods based on deep learning.
[0003] The three-dimensional model is projected into a set of two-dimensional views, and then the features of the two-dimensional views are extracted using deep learning technology to finally achieve model classification. This method has certain advantages. Since the dimensionality of the two-dimensional view data is low and there are a large number of annotated image data sets, better classification results can be achieved and it has high feasibility. However, this method also has some problems. The model information contained in a single view is limited, and it is difficult to fully and accurately characterize the three-dimensional model; and the multi-view three-dimensional model classification method easily ignores the complementarity and correlation between multiple views, affecting the accuracy and efficiency of classification. In response to the above problems, the present invention first proposes a multi-view self-attention (MHSA) method based on SwinTransformer by adding a global Multi-head Self-Attention (MHSA) module. Figure 3 The Swin Multi-View Global Attention (Swin-MVGA) method is used to extract view features from the 2D views of a 3D model. These features are then fused with SIFT features and LBP edge texture features. A multi-layer perceptron (MLP) is then used to extract the fused features. Finally, the Root Mean Quotient Feature Method (RMQF) is designed to select representative features, further improving the model's classification accuracy. Summary of the Invention
[0004] In order to solve the problems existing in the field of 3D model classification, the present invention discloses a 3D model classification method based on multi-view feature fusion of Swin-MVGA network.
[0005] To this end, the present invention provides the following technical solutions:
[0006] 1. A 3D model classification method based on Swin-MVGA network multi-view feature fusion, characterized in that the method mainly comprises the following steps:
[0007] Step 1: Preprocess the 3D model dataset Modelnet10 and project it to obtain 2D view features.
[0008] Step 2: Extract SIFT features of the 2D view and use the SIFT algorithm to extract key points and descriptors of the 2D view.
[0009] Step 3: Extract the edge texture features of the 2D view and use the LBP feature vector to represent the edge texture information of the 2D view.
[0010] Step 4: Divide the 3D model set into a 3D model training set and a 3D model test set, design the Swin-MVGA network to extract view features, use the training data to optimize Swin-MVGA, and evaluate the model performance on the test set.
[0011] Step 5: Fuse the view features extracted by Swin-MVGA with SIFT features and edge texture features LBP to generate comprehensive two-dimensional view fusion features, and design a four-layer MLP to extract multi-view fusion features.
[0012] Step 6: Design the RMQF method to extract representative features, input the multi-view fusion features into Softmax, convert them into probability distribution, use RMQF to extract representative features from the probability distribution, and finally classify the representative features extracted by RMQF.
[0013] 2. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, wherein in step 1, the three-dimensional model is pre-processed, specifically comprising the following steps:
[0014] Step 1-1 sets a circle with a fixed tilt angle above the model of the 3D model dataset Modelnet10;
[0015] Steps 1-2 distribute the camera positions evenly around the model, for example, sampling every 60°, and obtaining a total of 6 views. The 2D view set of the model is represented by V t ={V1, V2,…, V6}, where 1≤t≤n, and n=6 is the number of projection views of the three-dimensional model.
[0016] 3. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, wherein in step 2, the SIFT features of the two-dimensional view are extracted, specifically comprising the following steps:
[0017] Step 2-1 Two-dimensional view of the model V t (x t ,y t ) is Gaussian smoothed to obtain a scale space image:
[0018] L t (x t ,y t ,σ)=G t (x t ,y t ,σ)*V t (x t ,y t )
[0019] The Gaussian kernel is σ is the scale parameter;
[0020] Step 2-2 constructs a Gaussian difference (DoG) pyramid and calculates the difference image between adjacent scales:
[0021] D t (x t ,y t ,σ)=L t (x t ,y t ,kσ)-L t (x t ,y t ,σ)
[0022] Where k is the scale factor;
[0023] Steps 2-3 detect local extreme points in the DoG pyramid. These extreme points are used as a preliminary set of SIFT keypoint candidates to eliminate keypoints with low contrast or too strong edge response:
[0024] D t (x t ,y t ,σ) is a local extreme value if
[0025] If D t (x t ,y t ,σ) is greater than the value in all its neighborhoods, then it is a local maximum. If D t (x t ,y t,σ) is smaller than the value in all its neighborhoods, then it is a local minimum;
[0026] Steps 2-4 construct a gradient direction histogram in the neighborhood of the key point, select the main direction θ0 as the key point direction, and calculate the gradient magnitude and direction:
[0027]
[0028] Steps 2-5 generate descriptors. Taking the key point as the center, take a 16×16 neighborhood and divide it into 4×4 sub-regions. For each sub-region, calculate the gradient histogram in 8 directions to form a 128-dimensional descriptor. Normalize the descriptor:
[0029] Descriptor = {h ij |i=1,...,16;j=1,...,8}
[0030] In steps 2-6, for each key point, its two-dimensional position and 128 descriptors are concatenated into a 130-dimensional feature vector.
[0031] 4. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, wherein in step 3, edge texture features LBP of the two-dimensional view are extracted, specifically comprising the following steps:
[0032] Step 3-1 calculates the extended LBP features on the grayscale image. LBP generates a binary pattern by comparing the size relationship between the central pixel and its neighboring pixels:
[0033]
[0034] Among them, (x c ,y c ) is the coordinate of the center pixel, P is the number of neighborhood pixels, R is the radius of the neighborhood, g c is the grayscale value of the center pixel, g p is the grayscale value of the pth neighborhood pixel, and s(z) is the sign function:
[0035]
[0036] Step 3-2 converts the calculated LBP features into a histogram form to better represent the global texture features of the image and normalizes the histogram:
[0037]
[0038] Among them, H(i) is the original frequency of the i-th Bin of the histogram, N is the total number of Bins, ∈ is a small constant to prevent division by zero, and H norm (i) is the normalized histogram value.
[0039] 5. A 3D model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that in step 4, training data and test data are obtained from the 3D model set, a Swin-MVGA network is designed to extract view features, the training data is used to optimize Swin-MVGA, and the model performance is evaluated on the test set, specifically the following steps:
[0040] Step 4-1 projects the 3D model in ModelNet10 to obtain a 2D view set, which is divided into a 3D model training set and a 3D model test set;
[0041] In step 4-2, a Swin-MVGA-based network is designed to extract features from each view. The core modules of SwinTransformer, such as Patch Partition, Linear Embedding, Windowed Self-Attention (W-MSA), Hierarchical Structure, and Patch Merging, are used. SwinBlocks in Stage 1 to Stage 3 focus on local feature learning. Patch Merging gradually reduces the spatial resolution and increases the receptive field. Adding a global MHSA module after Stage 3 can integrate global information and make up for the shortcomings of SwinTransformer's local modeling. The formula of MHSA is as follows:
[0042] MultiHead(Q,K,V)=Concat(head1,head 2, ...,head h )W O
[0043] head i =Attention(Q i ,K i ,V i )
[0044] Among them, head i Denotes the i-th attention head, Concat(head1,...,head h ) represents the concatenation of the outputs of all attention heads, W O Represents the final linear transformation matrix;
[0045] In step 4-3, the Swin-MVGA network is trained end-to-end and its parameters are updated by constructing a suitable loss function and optimization algorithm using the training set data to improve the network's ability to extract view features and evaluate the model performance on the test set.
[0046] 6. The three-dimensional model classification method based on multi-view feature fusion of a Swin-MVGA network according to claim 1, characterized in that in step 5, the view features are fused with SIFT features and edge texture features (LBP) to generate comprehensive two-dimensional view fusion features, and a four-layer MLP is designed to extract the multi-view fusion features, specifically the following steps:
[0047] Step 5-1 fuses the SIFT features, LBP features, and view features extracted in steps 2, 3, and 4 to generate a comprehensive two-dimensional view fusion feature F;
[0048] In step 5-2, the concatenated features F are fed into the designed four-layer MLP network. For each layer, a linear transformation is first performed, followed by Dropout and ReLU activation. The calculation is as follows:
[0049] First layer: h1 = ReLU(Dropout(W1F+b1))
[0050] Second layer: h2=ReLU(Dropout(W2h1+b2))
[0051] Third layer: h3=ReLU(Dropout(W3h2+b3))
[0052] Fourth layer: h4=W4h3+b4
[0053] Among them, W is the weight matrix, h is the output vector, and b is the bias vector.
[0054] 7. A three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that in step 6, a RMQF method is designed to classify representative features extracted from multi-view fusion features, and the specific steps are as follows:
[0055] Step 6-1 converts the feature matrix h4 output by the MLP network into the category probability distribution matrix P through the Softmax layer. The calculation process of P is as follows:
[0056]
[0057] Step 6-2 uses the predicted probability p ij and the true label y ij Calculate the loss. The cross entropy loss function is calculated as follows:
[0058]
[0059] Among them, n represents the number of samples, c represents the number of categories, and p ij represents the probability that the i-th sample belongs to category j, y ij∈{0,1}, indicating one-hot encoding;
[0060] Step 6-3: Design the RMQF method to extract representative features. The specific steps are as follows:
[0061] Step 6-3-1 Normalize the matrix P to obtain the characteristic matrix, h ab norm Represents the normalized value of sample a (a = 1, ..., N) on feature dimension b (b = 1, ..., C):
[0062]
[0063] Step 6-3-2 calculates the column mean vector and obtains the feature distribution vector by taking the mean of the elements in each column b:
[0064]
[0065] Step 6-3-3 Calculate D RMQF Each row h in the quantized normalized feature matrix a norm The degree of difference from the characteristic distribution p, calculate D RMQF When p b Adding a smoothing term ∈ avoids p b The numerical instability caused by approaching 0, D RMQF The calculation process is as follows:
[0066]
[0067] Among them, N is the number of samples, C is the feature dimension, and ∈ is a very small positive number;
[0068] Step 6-3-4 to D RMQF Sort in descending order, select the top k largest indexes, select the corresponding rows from the original feature matrix h4, and then calculate the mean by row to obtain the new feature h rep ;
[0069] Step 6-4 is based on the obtained representative feature h rep Directly determine the prediction category and select h rep The category corresponding to the index with the largest median value is selected as the final predicted category. The selection process is as follows:
[0070]
[0071] Beneficial effects:
[0072] The invention relates to a three-dimensional model classification method based on the fusion of multi-view features of a Swin-MVGA network.
[0073] 1. This paper addresses the problem that 2D view features are insufficient to fully capture the overall 3D shape information in 3D model classification. By integrating multiple features with multi-view information, this paper proposes a classification method that extracts key view features from the 2D views of a 3D model using Swin-MVGA, a visual model based on self-attention and global MHSA that effectively captures both global and local information about different parts of an image. The resulting 2D view features contain rich spatial context and can effectively represent the model's shape and structure.
[0074] 2. To further enhance classification results, this invention enhances the information representation capabilities of the 2D view by fusing features from different sources. By fusing the features extracted by Swin-MVGA with SIFT features and LBP edge texture features, a more comprehensive 2D view fusion feature is obtained. SIFT features extract local geometric features, while LBP features capture texture information. Together, these features provide a more comprehensive 2D view representation and help improve the robustness of the model.
[0075] 3. The present invention designs the RMQF method to extract representative features, which greatly improves the classification efficiency of three-dimensional models. First, MLP is used to further integrate the fused multi-view features. Through nonlinear mapping, MLP can effectively handle the complex relationship between multi-view features and deeply fuse the feature information of different views; then RMQF is used to select representative features of the fused features. Compared with traditional methods, RMQF, as the core method, can specifically select the most representative features, further enhance the feature representation capability, and the model can better mine and utilize the complementary information between multiple views, thereby improving the classification effect of three-dimensional models.
[0076] 4. The proposed 3D model classification method is verified using the publicly available ModelNet10 dataset. The results show that the proposed classification method is effective. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 This is an example diagram of a three-dimensional model toilet to be classified in an embodiment of the present invention.
[0078] Figure 2 A three-dimensional model classification framework diagram in an embodiment of the present invention
[0079] Figure 3 Six angle views of a three-dimensional model obtained by a virtual camera in an embodiment of the present invention
[0080] Figure 4 Swin-MVGA structure diagram in the embodiment of the present invention DETAILED DESCRIPTION
[0081] In order to make the technical solutions in the embodiments of the present invention clearly and completely described, Figure 1 Taking the three-dimensional model of toilet category in the ModelNet10 training set as an example, the present invention is further described in detail in combination with the ModelNet10 training set.
[0082] The Swin-MVGA network multi-view feature fusion 3D model classification framework of the present invention is as follows: Figure 2 As shown, the following steps are included.
[0083] Step 1: Preprocess the 3D model. The specific steps are as follows:
[0084] Step 1-1 sets a circle with a fixed tilt angle above the model of the 3D model dataset Modelnet10;
[0085] Steps 1-2 distribute the camera positions evenly around the model, for example, sampling every 60° to obtain a total of 6 views, such as Figure 3 As shown, the two-dimensional view set of the model is V t ={V1, V2,…, V6}, where 1≤t≤n, and n=6 is the number of projection views of the three-dimensional model.
[0086] Step 2: Extract SIFT features of the 2D view. The specific steps are as follows:
[0087] Step 2-1 Two-dimensional view of the model V t (x t ,y t ) is Gaussian smoothed to obtain a scale space image:
[0088] L t (x t ,y t ,σ)=G t (x t ,y t ,σ)*V t (x t ,y t )
[0089] The Gaussian kernel is σ is the scale parameter;
[0090] Step 2-2 constructs a Gaussian difference (DoG) pyramid and calculates the difference image between adjacent scales:
[0091] D t (x t ,y t ,σ)=L t (x t,y t ,kσ)-L t (x t ,y t ,σ)
[0092] Where k is the scale factor;
[0093] Steps 2-3 detect local extreme points in the DoG pyramid. These extreme points are used as a preliminary set of SIFT keypoint candidates to eliminate keypoints with low contrast or too strong edge response:
[0094] D t (x t ,y t ,σ) is a local extreme value if
[0095] If D t (x t ,y t ,σ) is greater than the value in all its neighborhoods, then it is a local maximum. If D t (x t ,y t ,σ) is smaller than the value in all its neighborhoods, then it is a local minimum;
[0096] Steps 2-4 construct a gradient direction histogram in the neighborhood of the key point, select the main direction θ0 as the key point direction, and calculate the gradient magnitude and direction:
[0097]
[0098] Steps 2-5 generate descriptors. Taking the key point as the center, take a 16×16 neighborhood and divide it into 4×4 sub-regions. For each sub-region, calculate the gradient histogram in 8 directions to form a 128-dimensional descriptor. Normalize the descriptor:
[0099] Descriptor = {h ij |i=1,...,16;j=1,...,8}
[0100] Steps 2-6: For each key point, its two-dimensional position and 128 descriptors are concatenated into a 130-dimensional feature vector. Figure 3 The SIFT features extracted from the 3D model view shown are:
[0101] tensor([[0.0676,0.0917,...,0.0000,0.0000,0.0000],
[0102] [0.0848,0.0902,...,0.0136,0.0006,0.0011],
[0103] ...,
[0104] [0.1135,0.1002,...,0.0000,0.0000,0.0006]]).
[0105] Step 3: Extract edge texture features LBP of the two-dimensional view. The specific steps are as follows:
[0106] Step 3-1 calculates the extended LBP features on the grayscale image. LBP generates a binary pattern by comparing the size relationship between the central pixel and its neighboring pixels:
[0107]
[0108] Among them, (x c ,y c ) is the coordinate of the center pixel, P is the number of neighborhood pixels, R is the radius of the neighborhood, g c is the grayscale value of the center pixel, g p is the grayscale value of the pth neighborhood pixel, and s(z) is the sign function:
[0109]
[0110] Step 3-2 converts the calculated LBP features into a histogram form to better represent the global texture features of the image and normalizes the histogram:
[0111]
[0112] Among them, H(i) is the original frequency of the i-th Bin of the histogram, N is the total number of Bins, ∈ is a small constant to prevent division by zero, and H norm (i) is the normalized histogram value, from Figure 3 The edge texture feature LBP extracted from the 3D model view shown is:
[0113] tensor([[5.5927e-02,2.3142e-03,...,1.2857e-03,3.4585e-02,9.9665e-01],
[0114] [2.7737e-02,3.4672e-03,...,3.4672e-03,3.9525e-02,9.9853e-01],
[0115] ...,
[0116] [4.5961e-01,2.5534e-02,...,7.2956e-03,6.9306e-02,8.7909e-01]]).
[0117] Step 4: Obtain training data and test data from the 3D model set and design the Swin-MVGA network to extract view features. The specific steps are as follows:
[0118] Step 4-1 projects the 3D model in ModelNet10 to obtain a 2D view set, which is divided into a 3D model training set and a 3D model test set;
[0119] In step 4-2, a Swin-MVGA-based network is designed to extract features from each view. The core modules of SwinTransformer, such as Patch Partition, Linear Embedding, Windowed Self-Attention (W-MSA), Hierarchical Structure, and Patch Merging, are used. SwinBlocks in Stage 1 to Stage 3 focus on local feature learning. Patch Merging gradually reduces the spatial resolution and increases the receptive field. Adding a global MHSA module after Stage 3 can integrate global information and make up for the shortcomings of SwinTransformer's local modeling. The formula of MHSA is as follows:
[0120] MultiHead(Q,K,V)=Concat(head 1, head 2,·.·, head h )W O
[0121] head i= Attention(Q i ,K i ,V i )
[0122] Among them, head i Denotes the i-th attention head, Concat(head1,...,head h ) represents the concatenation of the outputs of all attention heads, W O Represents the final linear transformation matrix, from Figure 3 The extracted view features of the 2D view set shown are:
[0123] tensor([[-2.6307e-01,-1.7338e+00,...,8.0071e-02,-5.5511e-01,1.6058e+00],
[0124] [-1.1188e+00,-1.8577e+00,...,-8.5813e-02,-1.2455e+00,4.2884e-01],
[0125] ...,
[0126] [-4.7470e-01,-1.4942e+00,...,4.7115e-02,-1.0320e+00,1.0283e+00]]);
[0127] In step 4-3, the Swin-MVGA network is trained end-to-end and its parameters are updated by constructing a suitable loss function and optimization algorithm using the training set data to improve the network's ability to extract view features and evaluate the model performance on the test set.
[0128] Step 5: Fuse the view features with SIFT features and edge texture features LBP to generate comprehensive two-dimensional view fusion features. Design a four-layer MLP to extract multi-view fusion features. The specific steps are as follows:
[0129] Step 5-1 fuses the SIFT features, LBP features, and view features extracted in steps 2, 3, and 4 to generate a comprehensive two-dimensional view fusion feature F. The fusion feature F is as follows:
[0130] tensor([[-2.6307e-01,-1.7338e+00,...,1.2857e-03,3.4585e-02,9.9665e-01],
[0131] [-1.1188e+00,-1.8577e+00,...,3.4672e-03,3.9525e-02,9.9853e-01],
[0132] ...,
[0133] [-4.7470e-01,-1.4942e+00,...,7.2956e-03,6.9306e-02,8.7909e-01]]);
[0134] In step 5-2, the concatenated features F are fed into the designed four-layer MLP network. For each layer, a linear transformation is first performed, followed by Dropout and ReLU activation. The calculation is as follows:
[0135] First layer: h1 = ReLU(Dropout(W1F+b1))
[0136] The two-dimensional view fusion features after the first layer of MLP network are as follows:
[0137] tensor([[0.1026,0.2008,...,-0.0897,0.1203,-0.1461],
[0138] [-0.0779,-0.0027,...,0.0936,0.0156,-0.0350],
[0139] ...,
[0140] [0.0604,-0.0216,...,-0.0401,0.2689,-0.4058]])
[0141] Second layer: h2 = ReLU (Dropout (W2h1 + b2))
[0142] The two-dimensional view fusion features after the second layer of MLP network are as follows:
[0143] tensor([[0.1368,0.2678,...,0.0000,0.1604,0.0000],
[0144] [0.0000,0.2264,...,0.1247,0.0208,0.0000],
[0145] ...,
[0146] [0.1918,0.0000,...,0.2104,0.0817,0.0000]])
[0147] Third layer: h3 = ReLU (Dropout (W3h2 + b3))
[0148] The two-dimensional view fusion features after the third layer of MLP network are as follows:
[0149] tensor([[-0.0968,0.0726,...,0.0233,-0.0304,0.0074],
[0150] [-0.0778,0.0997,...,-0.0293,-0.0055,0.0252],
[0151] ...,
[0152] [-0.0569,0.0822,...,-0.0211,-0.0192,0.0009]])
[0153] Fourth layer: h4=W4h3+b4
[0154] The two-dimensional view fusion features after the fourth layer of MLP network are as follows:
[0155] tensor([[0.0000,0.0967,...,0.0311,0.0000,0.0099],
[0156] [0.0000,0.0000,...,0.0000,0.0000,0.0336],
[0157] ...,
[0158] [0.0000,0.1096,...,0.0000,0.0000,0.0012]])
[0159] Among them, W is the weight matrix, h is the output vector, and b is the bias vector.
[0160] Step 6: Design the RMQF method to extract representative features from multi-view fusion features and classify them. The specific steps are as follows:
[0161] Step 6-1 converts the feature matrix h4 output by the MLP network into the category probability distribution matrix P through the Softmax layer. The calculation process of P is as follows:
[0162]
[0163] The probability distribution matrix P obtained by the softmax layer is as follows:
[0164] tensor([[5.5017e-04,1.2030e-03,...,2.7718e-04,4.6192e-04,9.9447e-01],
[0165] [5.7809e-04,1.3795e-03,...,3.0021e-04,5.2959e-04,9.9342e-01],
[0166] ...,
[0167] [6.5827e-04,1.3171e-03,...,3.0170e-04,4.6950e-04,9.9438e-01]]);
[0168] Step 6-2 uses the predicted probability p i Calculate the loss with the true label y. The cross entropy loss function is calculated as follows:
[0169]
[0170] Among them, n represents the number of samples, c represents the number of categories, and p ij represents the probability that the i-th sample belongs to category j, y ij∈{0,1};
[0171] Step 6-3: Design the RMQF method to extract representative features. The specific steps are as follows:
[0172] Step 6-3-1 Normalize the matrix P to obtain the characteristic matrix, h ab norm Represents the normalized value of sample a (a = 1, ..., N) on feature dimension b (b = 1, ..., C):
[0173]
[0174] Step 6-3-2 calculates the column mean vector and obtains the feature distribution vector by taking the mean of the elements in each column b:
[0175]
[0176] Step 6-3-3 Calculate D RMQF Each row h in the quantized normalized feature matrix a norm The degree of difference from the characteristic distribution p, calculate D RMQF When p b Adding a smoothing term ∈ avoids p b The numerical instability caused by approaching 0, D RMQF The calculation process is as follows:
[0177]
[0178] Among them, N is the number of samples, C is the feature dimension, and ∈ is a very small positive number;
[0179] Step 6-3-4 to D RMQF Sort in descending order, select the top k largest indexes, select the corresponding rows from the original feature matrix h4, and then calculate the mean by row to obtain the new feature h rep , the representative feature h extracted from multi-view fusion features using RMQF rep is [5.9749e-04,1.3024e-03,5.6231e-04,4.9084e-04,2.3040e-04,7.2576e-04,1.1810e-03,2.9295e-04,4.8384e-04,9.9413e-01];
[0180] Step 6-4 is based on the obtained representative feature h rep Directly determine the prediction category and select h rep The category corresponding to the index with the largest median value is selected as the final predicted category. The selection process is as follows:
[0181]
[0182] h rep9 =9.9413e-01, the 9th category is toilet, Figure 1 The category of the 3D model instance shown is toilet.
[0183] The present invention is implemented on the test set of the 3D model dataset ModelNet10, with a classification accuracy of 91.89%.
[0184] The three-dimensional model classification method based on the Swin-MVGA network multi-view feature fusion implemented in the embodiment of the present invention fuses the view features, SIFT features and edge texture features LBP of the two-dimensional view to obtain the two-dimensional view fusion features, inputs the multi-view fusion features extracted by MLP into Softmax to convert them into probability distribution, designs the RMQF method to extract representative features from the probability distribution, and finally uses the representative features extracted by RMQF to determine the category of the three-dimensional model with high accuracy.
[0185] The above is a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. The specific embodiments herein are only intended to help understand the method of the present invention. For those skilled in the art, according to the concept of the present invention, changes and modifications may be made in the specific embodiments and application scope. Therefore, this specification should not be understood as limiting the present invention.
Claims
1. The 3D model classification method based on Swin-MVGA network multi-view feature fusion is characterized by: The method comprises the following steps: Step 1: Preprocess the 3D model dataset Modelnet10 and project it to obtain 2D view features. Step 2: Extract SIFT features of the 2D view and use the SIFT algorithm to extract key points and descriptors of the 2D view; Step 3: Extract edge texture features of the 2D view and use LBP feature vector to represent the edge texture information of the 2D view; Step 4: Divide the 3D model set into a 3D model training set and a 3D model test set, design a Swin-MVGA network to extract view features, use the training data to optimize Swin-MVGA, and evaluate the model performance on the test set; Step 5: Fuse the view features extracted by Swin-MVGA with SIFT features and edge texture features LBP to generate comprehensive two-dimensional view fusion features, and design a four-layer MLP to extract multi-view fusion features; Step 6: Design the RMQF method to extract representative features, input the multi-view fusion features into Softmax, convert them into probability distribution, use RMQF to extract representative features from the probability distribution, and finally classify the representative features extracted by RMQF.
2. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1 is characterized in that: In step 1, the three-dimensional model is preprocessed, and the specific steps are as follows: Step 1-1 sets a circle with a fixed tilt angle above the model of the 3D model dataset Modelnet10; Steps 1-2 distribute the camera positions evenly around the model, for example, sampling every 60°, and obtaining a total of 6 views. The 2D view set of the model is represented by V t ={V1, V2,…, V6}, where 1≤t≤n, and n=6 is the number of projection views of the three-dimensional model.
3. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that: In step 2, SIFT features of the two-dimensional view are extracted, and the specific steps are as follows: Step 2-1 Two-dimensional view of the model V t (x t ,y t ) is Gaussian smoothed to obtain a scale space image: L t (x t ,and t ,σ)=G t (x t ,and t ,σ)*V t (x t ,and t ) The Gaussian kernel is σ is the scale parameter; Step 2-2 constructs a Gaussian difference (DoG) pyramid and calculates the difference image between adjacent scales: D t (x t ,y t ,σ)=L t (x t ,y t ,kσ)-L t (x t ,y t ,σ) Where k is the scale factor; Steps 2-3 detect local extreme points in the DoG pyramid. These extreme points are used as a preliminary set of SIFT keypoint candidates to eliminate keypoints with low contrast or too strong edge response: D t (x t ,y t ,σ) is a local extreme value If D t (x t ,y t ,σ) is greater than the value in all its neighborhoods, then it is a local maximum. If D t (x t ,y t ,σ) is smaller than the value in all its neighborhoods, then it is a local minimum; Steps 2-4 construct a gradient direction histogram in the neighborhood of the key point, select the main direction θ0 as the key point direction, and calculate the gradient magnitude and direction: Steps 2-5 generate descriptors. Taking the key point as the center, take a 16×16 neighborhood and divide it into 4×4 sub-regions. For each sub-region, calculate the gradient histogram in 8 directions to form a 128-dimensional descriptor. Normalize the descriptor: Descriptor = {h ij |i=1,...,16;j=1,...,8} In steps 2-6, for each key point, its two-dimensional position and 128 descriptors are concatenated into a 130-dimensional feature vector.
4. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that: In step 3, the edge texture feature LBP of the two-dimensional view is extracted, and the specific steps are: Step 3-1 calculates the extended LBP features on the grayscale image. LBP generates a binary pattern by comparing the size relationship between the central pixel and its neighboring pixels: Among them, (x c ,y c ) is the coordinate of the center pixel, P is the number of neighborhood pixels, R is the radius of the neighborhood, g c is the grayscale value of the center pixel, g p is the grayscale value of the pth neighborhood pixel, and s(z) is the sign function: Step 3-2 converts the calculated LBP features into a histogram form to better represent the global texture features of the image and normalizes the histogram: Among them, H(i) is the original frequency of the i-th Bin of the histogram, N is the total number of Bins, ∈ is a small constant to prevent division by zero, and H norm (i) is the normalized histogram value.
5. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that: In step 4, training data and test data are obtained from the 3D model set, a Swin-MVGA network is designed to extract view features, the training data is used to optimize Swin-MVGA, and the model performance is evaluated on the test set. The specific steps are as follows: Step 4-1 projects the 3D model in ModelNet10 to obtain a 2D view set, which is divided into a 3D model training set and a 3D model test set; In step 4-2, a Swin-MVGA-based network is designed to extract features from each view. The core modules of SwinTransformer, such as Patch Partition, Linear Embedding, Windowed Self-Attention (W-MSA), Hierarchical Structure, and PatchMerging, are used. SwinBlocks in Stage 1 to Stage 3 focus on local feature learning. Patch Merging gradually reduces the spatial resolution and increases the receptive field. Adding a global MHSA module after Stage 3 can integrate global information and make up for the shortcomings of SwinTransformer's local modeling. The formula of MHSA is as follows: MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O head i =Attention(Q i ,K i ,V i ) Among them, head i Denotes the i-th attention head, Concat(head1,...,head h ) represents the concatenation of the outputs of all attention heads, W O Represents the final linear transformation matrix; In step 4-3, the Swin-MVGA network is trained end-to-end and its parameters are updated by constructing a suitable loss function and optimization algorithm using the training set data to improve the network's ability to extract view features and evaluate the model performance on the test set.
6. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that: In step 5, the view features are fused with the SIFT features and the edge texture features LBP to generate a comprehensive two-dimensional view fusion feature. A four-layer MLP is designed to extract the multi-view fusion feature. The specific steps are as follows: Step 5-1 fuses the SIFT features, LBP features, and view features extracted in steps 2, 3, and 4 to generate a comprehensive two-dimensional view fusion feature F; In step 5-2, the concatenated features F are fed into the designed four-layer MLP network. For each layer, a linear transformation is first performed, followed by Dropout and ReLU activation. The calculation is as follows: First layer: h1 = ReLU(Dropout(W1F+b1)) Second layer: h2 = ReLU (Dropout (W2h1 + b2)) Third layer: h3 = ReLU (Dropout (W3h2 + b3)) Fourth layer: h4=W4h3+b4 Among them, W is the weight matrix, h is the output vector, and b is the bias vector.
7. The three-dimensional model classification method based on Swin-MVGA network multi-view feature fusion according to claim 1, characterized in that: In step 6, the RMQF method is designed to extract representative features from multi-view fusion features and classify them. The specific steps are as follows: Step 6-1 converts the feature matrix h4 output by the MLP network into the category probability distribution matrix P through the Softmax layer. The calculation process of P is as follows: Step 6-2 uses the predicted probability p ij and the true label y ij Calculate the loss. The cross entropy loss function is calculated as follows: Among them, n represents the number of samples, c represents the number of categories, and p ij represents the probability that the i-th sample belongs to category j, y ij ∈{0,1}, indicating one-hot encoding; Step 6-3: Design the RMQF method to extract representative features. The specific steps are as follows: Step 6-3-1 Normalize the matrix P to obtain the characteristic matrix, h ab norm Represents the normalized value of sample a (a = 1, ..., N) on feature dimension b (b = 1, ..., C): Step 6-3-2 calculates the column mean vector and obtains the feature distribution vector by taking the mean of the elements in each column b: Step 6-3-3 Calculate D RMQF Each row h in the quantized normalized feature matrix a norm The degree of difference from the characteristic distribution p, calculate D RMQF When p b Adding a smoothing term ∈ avoids p b The numerical instability caused by approaching 0, D RMQF The calculation process is as follows: Among them, N is the number of samples, C is the feature dimension, and ∈ is a very small positive number; Step 6-3-4 to D RMQF Sort in descending order, select the top k largest indexes, select the corresponding rows from the original feature matrix h4, and then calculate the mean by row to obtain the new feature h rep ; Step 6-4 is based on the obtained representative feature h rep Directly determine the prediction category and select h rep The category corresponding to the index with the largest median value is selected as the final predicted category. The selection process is as follows:
Citation Information
Cited By
Pathological image diagnosis method based on GDKAN and multi-order context interaction gating
CN121527083A