Cross-modal matching model construction method of three-dimensional model

By learning intra-class and inter-class features in the cross-modal matching of three-dimensional models and introducing adaptive weight balance different modal features, the existing methods have solved the problem of distinguishing ability and feature model length sensitivity, and achieved stronger distinction ability and wider application scenarios.

CN120236104APending Publication Date: 2025-07-01UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510383130.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The cross-modal matching method of existing three-dimensional models has weak distinction capabilities when dealing with different categories of features, and is sensitive to feature modes, making it difficult to process real scene images of complex backgrounds.

Method used

A cross-modal matching model construction method for three-dimensional models is proposed. By learning intra-class and inter-class features in a shared embedding space, and introducing adaptive weights to balance the contribution of different modal features to the optimization goal, we overcome the sensitivity to feature model length.

Benefits of technology

It realizes a strong distinction ability when processing different categories of feature, reduces the sensitivity to feature mode length, can be extended to complex two-dimensional image scenes, and breaks through the limitations of the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236104A_ABST
    Figure CN120236104A_ABST
Patent Text Reader

Abstract

The invention relates to the field of information retrieval, and discloses a cross-modal matching model construction method for a three-dimensional model, which comprises the following steps of: firstly, constructing a training sample set; then, for each input sample, obtaining corresponding features of data of each modal through a corresponding feature extraction network, and mapping the features to a shared embedding space after normalization; and then, on the basis of the intra-class distance, the inter-class distance and a preset distance parameter, center loss and adaptive loss weight are calculated respectively, adaptive cosine center loss is calculated by using the center loss and the adaptive loss weight, and training of the matching model is realized by minimizing the adaptive cosine center loss. Therefore, intra-class and inter-class learning can be realized at the same time; by introducing the self-adaptive loss weight, balancing the contribution of each category and each mode, overcoming the sensitivity to the characteristic mode length, and expanding the cross-mode matching of the three-dimensional model to a complex image scene outside a rendering graph, the method is suitable for various cross-mode matching tasks of the three-dimensional model, such as retrieval and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval, and more particularly to a method for constructing a cross-modal matching model for three-dimensional models. Background Art

[0002] With the development of information technology, in the fields of virtual reality, game development, and industrial product design, three-dimensional data has shown explosive growth. Therefore, how to efficiently and accurately retrieve in the increasingly large-scale three-dimensional model data has become the focus of people's attention.

[0003] Therefore, someone proposed the DSCMR method (the full English name is Deep Supervised Cross-Modal Retrieval, and the full Chinese name is Deep Supervised Cross-Modal Retrieval Method, from: Zhen L, Hu P, Wang X, et al. Deep supervised cross-modal retrieval [C]. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 10394-10403). This method maps different modalities to a unified semantic space through a deep neural network and directly compares the similarities of cross-modal samples in this space.

[0004] However, it should be noted that in the cross-modal retrieval of 3D models, that is, in the cross-modal matching task of 3D models, there are significant differences in the feature distributions between different modalities, such as images, point clouds, meshes, etc. Traditional DSCMR methods are difficult to effectively handle the heterogeneity between modalities. To solve this problem, the CMCL method (full English name: Cross-modal Center Loss, Chinese full name: Cross-modal Center Loss, source: Jing L, Vahdani E, Tan J, et al. Cross-modal center loss for 3D cross-modal retrieval [C]. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021: 3142-3151) has been proposed. In addition to using deep neural networks to map different modalities to a unified semantic space, the CMCL method also generates a central feature point for each category by introducing "cross-modal central features", and minimizes the distance between the features of the same category within different modalities and the central feature, while maximizing the distance between the features of different categories. Therefore, in a multi-modal environment, it can aggregate the same-class features from different modalities, thus achieving cross-modal feature alignment.

[0005] However, the CMCL method mainly reduces the differences between modalities through the convergence of intra-class features, without performing contrastive learning between different categories, which results in weak discrimination ability when dealing with different-category features. At the same time, in the CMCL method, since its optimization objective depends on the norm of the feature vector, the inconsistency of the feature norm will affect the distance calculation. Therefore, currently it is mainly applied to the CAD (Computer-Aided Design) rendering model image scenario where the feature norm changes little, that is, through the matching of the rendering image and the 3D model to achieve cross-modal retrieval. It is difficult to process real-scene images with complex backgrounds, such as images generated by generative models, real photos, etc. With the wide application of generative models, complex two-dimensional image data are constantly emerging. The limitation on query data hinders the application of cross-modal 3D model retrieval technology in the actual 3D model retrieval process. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to propose a method for constructing a cross-modal matching model of 3D models, which performs inter-class learning while learning intra-class features, and introduces an adaptive weight to balance the contributions of different-modal features to the optimization objective, thereby overcoming the sensitivity to the feature norm.

[0007] The technical solution adopted by the present invention to solve the above technical problems is as follows:

[0008] A method for constructing a cross-modal matching model of a three-dimensional model, comprising the following steps:

[0009] A1. Construct a training sample set, where each training sample in the training sample set respectively includes the class label of the training sample and the data of each modality to be matched of the training sample; in the shared embedding space, initialize the class centers of each category included in the training sample set;

[0010] A2. Extract the training samples and input them into the model;

[0011] A3. For each training sample input in step A2, map it to the shared embedding space respectively according to the following steps:

[0012] A31. For the current training sample, the data of each modality to be matched of it are respectively passed through the feature extraction network of the corresponding modality to obtain the features of each modality to be matched, and the dimensions of the obtained features of each modality to be matched are the same, all being d-dimensional;

[0013] A32. For the current training sample, normalize the features of each modality to be matched obtained in step A31, and map them to the shared embedding space in the form of tensors;

[0014] A4. Calculate the loss, and the loss includes the adaptive cosine center loss The adaptive cosine center loss is calculated according to the following formula:

[0015]

[0016] where M is the number of modalities to be matched, N is the number of training samples input into the model in step A2, is the adaptive loss weight of the j-th modality to be matched of the i-th training sample, r i j is the center loss of the j-th modality to be matched of the i-th training sample;

[0017] The adaptive loss weight is calculated according to the following formula:

[0018]

[0019] The center loss r i j , is calculated according to the following formula:

[0020]

[0021] Where K is the number of categories of training samples in the training sample set, and e is the natural constant;

[0022] represents the point of the j-th modality to be matched of the i-th training sample in the shared embedding space and the vector angle between the class center of the category k to which the i-th training sample belongs i ; is a function representing the intra-class distance, m is a preset distance parameter satisfying , and cos(·) represents the cosine similarity;

[0023] represents the point of the j-th modality to be matched of the i-th training sample in the shared embedding space and the vector angle between the class center of other categories except the category k to which the i-th training sample belongs i ; is a function representing the inter-class distance, and cos(·) represents the cosine similarity;

[0024] A5. Using the loss obtained by calculating in step A4, update the parameters of each feature extraction network and the class centers of each category in the shared embedding space based on the gradient;

[0025] A6. Loop through steps A2 to A5 until the training end condition is reached, and obtain a cross-modal matching model that has completed training.

[0026] Preferably, use the unit hypersphere as the shared embedding space; in step A32, normalize the features of each modality to be matched respectively according to the following formula, and map them to the unit hypersphere in tensor form:

[0027]

[0028] where norm represents L2 norm normalization, and f j () represents the feature extraction network of the j-th modality to be matched, represents the data of the j-th modality to be matched of the i-th training sample, represents the point of the j-th modality to be matched of the i-th training sample in the shared embedding space.

[0029] Preferably, the value of the preset distance parameter m is 0.3 to 0.5.

[0030] Furthermore, in step A2, input the training samples in batches;

[0031] The loss in step A4 And it is calculated according to the following formula:

[0032]

[0033] Where λ cma is the weight, is the adaptive cosine center loss, is the cross-modal affinity loss;

[0034] The calculation of the cross-modal affinity loss includes the following steps:

[0035] First, take the points of each to-be-matched modality of each training sample obtained in step A32 in the shared embedding space as the point set S; then divide the point set S into subsets by category and calculate the affinity loss of each subset respectively k p represents the category of the training sample corresponding to the points included in the subset:

[0036]

[0037]

[0038] Where M is the number of to-be-matched modalities, is the number of points included in the subset t is a preset scaling coefficient, z and z′ respectively represent the points in the subset z≠z′ means that z and z′ are different points, θ zz′ represents the vector angle between points z and z′ in the shared embedding space, cos(·) represents the cosine similarity function, and e is the natural constant;

[0039] After that, calculate the cross-modal affinity loss according to the following formula

[0040]

[0041] Where K ′ represents the number of categories with more than 1 training sample included in the input of step A2.

[0042] Preferably, the value of the scaling coefficient t is 1 to 3.

[0043] Furthermore, the loss in step A4 also includes a classification loss and it is calculated according to the following formula:

[0044]

[0045] Where λcma and λ ce are weights, is the adaptive cosine center loss, is the cross-modal affinity loss, is the classification loss;

[0046] The classification loss is calculated according to the following formula:

[0047]

[0048] where M is the number of modalities to be matched, N is the number of training samples input to the model in step A2, and y i represents the class label of the i-th training sample; is the prediction classification of the j-th modality to be matched of the i-th training sample obtained based on the feature of the j-th modality to be matched of the i-th training sample obtained in step A32, using the shared classifier; each modality to be matched is classified and predicted by the same shared classifier, and in step A5, the parameters of the shared classifier are updated based on the gradient using the loss calculated in step A4.

[0049] Preferably, the weight λ cma is preset to be 0.1 to 1, and the weight λ ce is preset to be 0.1 to 1.

[0050] Furthermore, the modalities to be matched of the training samples include at least two types of modalities among 3D point clouds, 3D meshes, and 2D images.

[0051] Specifically, the modalities to be matched of the training samples include 3D point clouds, 3D meshes, and 2D images; in the training sample set, the 3D point cloud data and 3D mesh data included in each training sample are from the same instance, and the 2D image data included in each training sample is of the same category as the 3D point cloud data and 3D mesh data.

[0052] Specifically, the feature extraction networks corresponding to each modality are: the DGCNN network for extracting 3D point cloud data features, the MeshNet network for extracting 3D mesh data features, and the ResNet18 network for extracting 2D image data features.

[0053] The beneficial effects of the present invention are:

[0054] First, in the method of the present invention, before mapping the features to the shared embedding space, the feature vectors are normalized to make their norms consistent.

[0055] Second, the method of the present invention defines Represents the within-class distance, reflecting the compactness of the within-class feature distribution; defined Represents the between-class distance, reflecting the between-class confusion degree of the feature semantics; define a preset distance parameter m for adjusting the learning of the within-class compactness and the between-class confusion degree. Based on the above definitions, the center loss is obtained by calculating and comparing the cosine center error, and by minimizing the loss, the between-class learning is realized while learning the within-class features.

[0056] Thirdly, the method of the present invention is based on the defined within-class distance and the between-class distance to calculate and obtain an adaptive loss weight. By introducing the distance-based adaptive loss weight, the contributions of different categories and different modality features to the optimization objective are balanced, thereby overcoming the sensitivity to the feature norm length.

[0057] Therefore, the method of the present invention has strong discrimination ability when dealing with different category features, is not sensitive to the feature norm length, and can extend the cross-modal matching of 3D models to complex 2D image scenes other than rendered images, such as real image scenes, breaking through the limitations of existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 Shows the sensitivity test results of the hyperparameter m;

[0059] Figure 2 Shows the sensitivity test results of the hyperparameter t;

[0060] Figure 3 Shows the visual comparison of the retrieval results. DETAILED DESCRIPTION OF THE INVENTION

[0061] The present invention aims to propose a method for constructing a cross-modal matching model of 3D models, which also uses a deep neural network to map different modalities to a unified semantic space. The core lies in the proposal of an adaptive cosine center loss

[0062] Furthermore, define to represent the within-class distance, reflecting the compactness of the within-class feature distribution; define to represent the between-class distance, reflecting the between-class confusion degree of the feature semantics; define a preset distance parameter m that satisfies for adjusting the learning of the within-class compactness and the between-class confusion degree.

[0063] Based on the above definitions, according to the following formula, the center loss r is obtained by calculating and comparing the cosine center error i j :[[]]END]]

[0064]

[0065] Meanwhile, based on the defined within-class distance and between-class distance, the adaptive loss weight w is calculated according to the following formula i j :

[0066]

[0067] where K is the number of categories of training samples in the training sample set, and e is the natural constant; represents the point of the j-th to-be-matched modality of the i-th training sample in the shared embedding space and the vector angle between the class center of the category k i to which the i-th training sample belongs; represents the point of the j-th to-be-matched modality of the i-th training sample in the shared embedding space and the vector angle between the class center of other categories except the category k i to which the i-th training sample belongs; cos(·) represents the cosine similarity.

[0068] Based on the obtained center loss r i j and the adaptive loss weight the adaptive cosine center loss is calculated according to the following formula

[0069]

[0070] where M is the number of to-be-matched modalities, and N is the number of input training samples.

[0071] By minimizing the adaptive cosine center loss intra-class learning is achieved while inter-class learning is carried out; by introducing a distance-based adaptive loss weight, the contributions of different categories and different modality features to the optimization objective are balanced, thereby overcoming the sensitivity to the feature modulus length.

[0072] Furthermore, in the cross-modal shared embedding space, each category contains multiple training samples, and each training sample contains multiple normalized feature vectors of different modalities. For the feature vectors within the same category, regardless of which training sample and which modality they come from, the angle between them can be counted as θ zz′ . Based on this, the present invention constructs an affinity function:

[0073]

[0074]

[0075] z≠z ′

[0076] Among them, is a subset composed of points of class k in the total point set S formed by points of each modality to be matched in the shared embedding space for each training sample, z and z p respectively represent the points in the subset ′ z ≠ z means that z and z ′ are different points, θ ′ represents the vector angle between points z and z' in the shared embedding space, cos(·) represents the cosine similarity function, e is the natural constant, and t is a preset scaling factor. zz′ This affinity function has a formal correlation with the cosine distance. In the cross-modal shared embedding space, features with a greater distance are assigned stronger affinity values. Therefore, based on the affinity function

[0077] a cross-modal affinity loss calculated according to the following formula is constructed which can further reduce the cross-modal difference.

[0078]

[0079]

[0080] where M is the number of modalities to be matched, is the number of points contained in the subset K ′ represents the number of classes with more than 1 training samples included in the input.

[0081] In addition, in order to further improve the embedding ability of the feature extraction network, the present invention also introduces a classification loss of cross entropy

[0082]

[0083] where M is the number of modalities to be matched, N is the number of training samples in the input, y i represents the class label of the i-th training sample; is the feature of the j-th modality to be matched of the i-th training sample, and the predicted classification of the j-th modality to be matched of the i-th training sample obtained by using the shared classifier.

[0084] Finally, in the way of the sum of weights, the above three types of losses can be fused according to the following formula to calculate the final total loss:

[0085]

[0086] where λ cma and λ ce are weights, is the adaptive cosine center loss, is the cross-modal affinity loss, is the classification loss.

[0087] Therefore, the present invention can extend the technology of cross-modal 3D model matching to real image scenarios, has strong robustness to complex background images, has strong learning ability in the feature space, breaks through the limitations of existing methods, and achieves better retrieval effects.

[0088] The following is further described in conjunction with embodiments.

[0089] Embodiment

[0090] This embodiment provides a method for constructing a cross-modal matching model of a 3D model, including the following steps:

[0091] A1. Data preparation

[0092] In this step, first, a training sample set is constructed. Each training sample in the training sample set respectively includes the class label of the training sample and the data of each modality to be matched of the training sample.

[0093] Since the retrieval constructed based on the matching model of the present invention can implement tasks in six types of scenarios: inputting a 2D image and outputting a 3D point cloud, inputting a 2D image and outputting a 3D mesh, inputting a 3D point cloud and outputting a 2D image, inputting a 3D mesh and outputting a 2D image, inputting a 3D point cloud and outputting a 3D mesh, and inputting a 3D mesh and outputting a 3D point cloud. Therefore, based on actual needs, the modalities to be matched of the training sample include at least two of 3D point cloud, 3D mesh, and 2D image.

[0094] In this embodiment, for the subsequent verification of the six types of scenarios, the modalities to be matched of the training sample include 3D point cloud, 3D mesh, and 2D image; in the training sample set, the 3D point cloud data and 3D mesh data included in each training sample are from the same instance, and the 2D image data included in each training sample is of the same class as the 3D point cloud data and 3D mesh data. Since the 3D data and 2D image only require the same class and are not limited to the same instance, the samples can be greatly enriched.

[0095] In order to overcome problems such as the origin offset of the original data, inconsistent sizes, non-uniform point cloud densities, and non-uniform numbers of mesh polygon faces, when constructing the training sample set, the same data preprocessing method as the CMCL method is adopted to unify the point cloud density, mesh face number, and image resolution.

[0096] In this step, the class centers of all categories included in the training sample set are also initialized in the shared embedding space. The categories included in the training sample set should be consistent with the categories to be finally matched, so as to ensure that the class centers of all categories can be obtained through training. Any existing method can be used for the initialization of the class centers. In this embodiment, random initialization is adopted.

[0097] The shared embedding space can be any existing method. In this embodiment, the unit hypersphere is used as the shared embedding space. The advantages of the unit hypersphere as the shared embedding space lie in its feature normalization, alignment and uniformity, geometric properties, and computational efficiency. These properties make it perform well in tasks such as feature representation, classification, clustering, and cross-modal retrieval. Especially in 3D modeling and cross-modal retrieval tasks, the unit hypersphere can effectively align the features of different modalities, thereby improving the retrieval performance.

[0098] A2. Input samples

[0099] In this step, the training samples are extracted and input into the model. Considering the subsequent cross-modal affinity loss requirements, the training samples are input in batches in this step.

[0100] A3. Construct the shared embedding space

[0101] In this step, for each training sample input in step A2, it is mapped to the shared embedding space respectively according to the following steps:

[0102] A31. For the current training sample, the data of each modality to be matched are respectively passed through the feature extraction network of the corresponding modality to obtain the features of each modality to be matched, and the dimensions of the features of each modality to be matched obtained are the same, all being d-dimensional.

[0103] The feature extraction network of each modality can adopt any existing neural network according to the input modality. Specifically, in this embodiment, the feature extraction networks corresponding to each modality are respectively: the DGCNN network for extracting the features of 3D point cloud data, the MeshNet network for extracting the features of 3D mesh data, and the ResNet18 network for extracting the features of 2D image data. Considering the trade-off between the complexity of the basic feature extraction network, the video memory occupancy, and the retention of data specificity, in this embodiment, the dimension d of the output features is 512.

[0104] A32. For the current training sample, the features of each modality to be matched obtained in step A31 are respectively normalized and mapped to the shared embedding space in the form of tensors.

[0105] Specifically, in this embodiment, the unit hypersphere is used as the shared embedding space; in this step, the features of each modality to be matched are normalized by the L2 norm according to the following formula and mapped to the unit hypersphere in the form of a tensor:

[0106]

[0107] where norm represents L2 norm normalization, f j () represents the feature extraction network of the j-th modality to be matched, represents the data of the j-th modality to be matched of the i-th training sample, represents the point of the j-th modality to be matched of the i-th training sample in the shared embedding space.

[0108] A4. Calculate the loss

[0109] In this step, the loss is calculated according to the following formula:

[0110]

[0111] where λ cma and λ ce are weights, is the adaptive cosine center loss, is the cross-modal affinity loss, is the classification loss.

[0112] Among them, the adaptive cosine center loss is calculated according to the following formula:

[0113]

[0114] where M is the number of modalities to be matched, N is the number of training samples input to the model in step A2, is the adaptive loss weight of the j-th modality to be matched of the i-th training sample, r i j is the center loss of the j-th modality to be matched of the i-th training sample;

[0115] The adaptive loss weight is calculated according to the following formula:

[0116]

[0117] The center loss r i j , is calculated according to the following formula:

[0118]

[0119] where K is the number of categories of training samples in the training sample set, and e is the natural constant;

[0120] represents the point of the j-th modality to be matched of the i-th training sample in the shared embedding space and the class center of the class k to which the i-th training sample belongs i the vector angle between them; is a function representing the intra-class distance, m is a preset distance parameter satisfying , and cos(·) represents the cosine similarity;

[0121] represents the point of the j-th modality to be matched of the i-th training sample in the shared embedding space and the class center of other classes except the class k to which the i-th training sample belongs i the vector angle between them; is a function representing the inter-class distance, and cos(·) represents the cosine similarity.

[0122] The calculation of the cross-modal affinity loss includes the following steps:

[0123] First, take the points of each modality to be matched of each training sample obtained in step A32 in the shared embedding space as the point set S; then divide the point set S into subsets according to the category and calculate the affinity loss of each subset respectively k p represents the category of the training sample corresponding to the points included in the subset:

[0124]

[0125]

[0126] where M is the number of modalities to be matched, is the number of points included in the subset , t is a preset scaling factor, z and z′ respectively represent the points in the subset , z≠z′ means that z and z′ are different points, and θ zz′ represents the vector angle between points z and z′ in the shared embedding space, cos(·) represents the cosine similarity function, and e is the natural constant;

[0127] After that, calculate the cross-modal affinity loss according to the following formula

[0128]

[0129] where K′ Indicates the number of categories with more than 1 training samples input in step A2.

[0130] The loss in step A4 Also includes classification loss And is calculated according to the following formula:

[0131]

[0132] Where, λ cma and λ ce Are weights, Is the adaptive cosine center loss, Is the cross-modal affinity loss, Is the classification loss;

[0133] The classification loss Is calculated according to the following formula:

[0134]

[0135] Where, M is the number of modalities to be matched, N is the number of training samples input to the model in step A2, y i Represents the class label of the i-th training sample; Is the feature of the j-th modality to be matched of the i-th training sample obtained based on step A32, and the predicted classification of the j-th modality to be matched of the i-th training sample obtained by using the shared classifier; each modality to be matched is classified and predicted through the same shared classifier. In this embodiment, the shared classifier is a multi-layer perceptron.

[0136] According to subsequent experiments, optimally, the value of the preset distance parameter m is 0.3 to 0.5, the value of the scaling coefficient t is 1 to 3, the weight λ cma The value is preset to 0.1 to 1, and the weight λ ce The value is preset to 0.1 to 1.

[0137] A5. Parameter update

[0138] In this step, using the loss calculated in step A4, based on the gradient, the parameters of each feature extraction network, the class centers of each category in the shared embedding space, and the parameters of the shared classifier are updated.

[0139] A6. Iteration

[0140] In this step, based on the completion condition, a condition determination is made. Specifically, steps A2 to A5 are repeatedly executed until the training end condition is reached, and a cross-modal matching model that has completed training is obtained.

[0141] Experimental verification:

[0142] The inventor used the matching model obtained through training in the above embodiments, and verified the effect of the present invention through retrieval using three datasets, namely ModelNet40, MI3DOR, and Pix3D. The verification index adopted the mean Average Precision (mAP). mAP is an evaluation index widely used in information retrieval and machine learning, mainly used to measure the performance of object detection models. Among them, AP, that is, average precision, measures the detection ability of the model for a certain category by calculating the average value of the precision at different recall rates; mAP is the average value of APs for all categories, used to measure the comprehensive detection ability of the model on all categories.

[0143] The above dataset ModelNet40 is a three-dimensional object benchmark dataset, containing 12,311 CAD models belonging to 40 different categories, of which 9,843 are used for training and 2,468 are used for testing. This dataset provides three modalities: 2D images, 3D point clouds, and 3D meshes.

[0144] The above dataset MI3DOR is a large-scale dataset for 2D to 3D tasks, containing 21,000 images and 7,690 models from 21 categories, of which 3,842 are used for training and 3,848 are used for testing. The images in MI3DOR are real-world images with rich backgrounds and colors, having a more complex background compared to the synthetic rendered images in ModelNet40.

[0145] The above dataset Pix3D is a large-scale dataset, containing real images and real models with precise 2D-3D alignment, with a total of 395 models and 16,913 images from 9 categories, of which 313 are used for training and 82 are used for testing.

[0146] Since the MI3DOR and Pix3D datasets only contain 3D models. Therefore, the 3D mesh data in the MI3DOR and Pix3D datasets was sampled to obtain point cloud data including 2,048 points for training and testing.

[0147] Performing retrieval tests using the matching model includes:

[0148] B1. Extract sample data from the dataset to construct a training sample set and a testing sample set; since for the training samples and testing samples, the 3D point cloud and 3D mesh contained in each sample come from the same instance, and the 2D image contained in each sample only needs to come from the same category as the 3D point cloud and 3D mesh, therefore, the 2D images are extracted from the images in the dataset in the file order of the category. The training sample set and the testing sample set both completely contain all categories of the dataset to be retrieved.

[0149] B2. Using the training sample set, obtain a trained cross-modal matching model according to the matching model construction method of the present invention;

[0150] B3. Using the trained cross-modal matching model obtained in step B2, map the data of each modality included in the test sample set to the shared embedding space respectively in the manner of step A3, and obtain a shared embedding space containing the knowledge included in the test sample set;

[0151] B4. Extract data from the data of the test sample set to construct a query, and use the trained cross-modal matching model obtained in step B2 to obtain its point in the shared embedding space in the manner of step A3;

[0152] B5. Based on the point of the query in the shared embedding space, query in the shared embedding space obtained in step B3 to obtain the retrieval target.

[0153] In step B2 and step B3, each training sample and test sample includes 2 two-dimensional images from the same category. When extracting features, first, use the feature extraction network to obtain their features respectively, and then, through matrix calculation, obtain the features that fuse the 2 two-dimensional images as the features of the final two-dimensional image of the sample. In step B4, extract data from the data of the test sample set to construct a query, that is: at this time, there are no samples, but only isolated data, and the query is constructed by random extraction. It should be particularly noted that in step B4, when using a two-dimensional image as the query, its features can come from 1 two-dimensional image, or can be obtained by fusing the features of multiple randomly obtained images of the same category, and are set through the parameter of the number of test images.

[0154] Experiment 1

[0155] The inventors simulated the sensitivity of the hyperparameters m, t, λ cma and λ ce During the simulation, the mAP of six scenarios was tested respectively, and the average value of the mAP of the six scenarios was used as the final score under the condition of this hyperparameter. The six scenarios include: 1. Input two-dimensional image and output three-dimensional point cloud; 2. Input two-dimensional image and output three-dimensional mesh; 3. Input three-dimensional point cloud and output two-dimensional image; 4. Input three-dimensional mesh and output two-dimensional image; 5. Input three-dimensional point cloud and output three-dimensional mesh; 6. Input three-dimensional mesh and output three-dimensional point cloud. The Batchsize during the training of this experiment is 64, and the number of test images during the test is 1.

[0156] For the sensitivity test of the hyperparameter m, two datasets, ModelNet40 and MI3DOR, were used. During the test, the value range of m is 0 to 0.9, and the other hyperparameters are: t = 2, λcma = 1, λ ce = 1, the results are as Figure 1 shown.

[0157] For the sensitivity test of the hyperparameter t, two datasets, ModelNet40 and MI3DOR, were used. During the test, the value range of t was 1 to 10, and the other hyperparameters were: m = 0.35, λ cma = 1, λ ce = 1, the results are as Figure 2 shown.

[0158] Hyperparameter λ cma and λ ce sensitivity test only used the ModelNet40 dataset. During the test, the values of λ cma and λ ce were four scales of 0.01, 0.1, 1, and 10, and the other hyperparameters were: t = 2, m = 0.35, λ cma = 1, λ ce = 1, and the results are shown in Table 1.

[0159] Table 1. Hyperparameter λ cma and λ ce sensitivity test

[0160]

[0161] Experiment 2

[0162] Furthermore, on three datasets, a comparative test was carried out with the cross-modal center loss method (CMCL, Cross-modal CenterLoss) as the comparison model. In the comparative test, the hyperparameters of the matching model of the present invention were m = 0.35, t = 2, λ cma = 1, λ ce = 1. This hyperparameter setting scheme can better balance the retrieval accuracy in various scenarios. The Batchsize during the training of this experiment was 128, and the number of test images during the test was 4, that is, the points of the two-dimensional images obtained as queries in the shared embedding space were the fusion features of four randomly obtained pictures from the same category. The results are shown in Table 2, Table 3, and Table 4.

[0163] Table 2. Retrieval performance on the ModelNet40 dataset

[0164]

[0165] Table 3. Retrieval performance on the MI3DOR dataset

[0166]

[0167] Table 4, Retrieval Results of the Pix3D Dataset

[0168]

[0169] The results show the superiority of the matching model of the present invention in cross-modal 3D retrieval. Its mAP is better than the CMCL method in the cross-modal 3D model retrieval tasks in various scenarios. Especially in the dataset containing real images Pix3D, the mAP retrieved based on the method of the present invention has increased by 7.8% from image to mesh, 5.7% from image to point cloud, 5.3% from mesh to image, 0.6% from mesh to point cloud, 9.2% from point cloud to image, and 3.9% from point cloud to mesh. The overall scene mAP has increased by 5.4%. Except for mesh to point cloud, significant improvements have been achieved. In the cross-modal retrieval between four types of images and 3D models, namely image→mesh, image→point cloud, mesh→image, and point cloud→image, the improvement is more than 5%.

[0170] For further demonstration, the inventor visualized the retrieval results on the MI3DOR dataset, showing the TOP-5 data among all retrieval results of each query, as Figure 3 shown. In the figure, the dashed square indicates that the category and the query do not belong to the same class, that is, the retrieval result is incorrect; the solid square indicates that the category and the query belong to the same class, that is, the retrieval result is correct. It can be seen from the figure that in the TOP-5 data in the real image scenario, the retrieval accuracy rate of the CMCL method is low, while the present invention can retrieve the correct result with a relatively high accuracy rate under the same input.

[0171] Finally, it should be noted that the above embodiments are only preferred embodiments and are not intended to limit the present invention. It should be pointed out that for those of ordinary skill in the art in the technical field, without departing from the spirit and scope protected by the claims of the present invention, several modifications, equivalent replacements, improvements, etc. should be included within the protection scope of the present invention.

Claims

1. A method for constructing a cross-modal matching model of a three-dimensional model, characterized in that: The steps include: A1. Construct a training sample set, wherein each training sample in the training sample set includes a category label of the training sample and data of each modality to be matched for the training sample; initialize the class center of each category included in the training sample set in the shared embedding space; A2. Extract training samples and input them into the model; A3. For each training sample input in step A2, map it to the shared embedding space according to the following steps: A31. For the current training sample, the data of each modality to be matched are respectively passed through the feature extraction network of the corresponding modality to obtain the features of each modality to be matched, and the dimensions of the features of each modality to be matched are consistent, which are all d dimensions; A32, for the current training sample, normalize the features of each modality to be matched obtained in step A31, and map them to the shared embedding space in the form of tensors; A4. Calculate the loss, the loss Including adaptive cosine center loss The adaptive cosine center loss Calculated according to the following formula: Where M is the number of modalities to be matched, N is the number of training samples input into the model in step A2, is the adaptive loss weight of the jth to-be-matched modality of the i-th training sample, is the central loss of the jth to-be-matched modality of the i-th training sample; The adaptive loss weight Calculated as follows: The center loss Calculated as follows: Among them, K is the number of categories of training samples in the training sample set, and e is a natural constant; Represents the point of the jth modality to be matched for the i-th training sample in the shared embedding space The category k to which the i-th training sample belongs i The vector angle between the class centers of ; is a function representing the distance within a class, and m satisfies The preset distance parameter, cos(·) represents the cosine similarity; Represents the point of the jth modality to be matched for the i-th training sample in the shared embedding space The category k to which the i-th training sample belongs i The vector angle between the class centers of other categories except ; is a function representing the distance between classes, cos(·) represents cosine similarity; A5. Using the loss calculated in step A4, based on the gradient, update the parameters of each feature extraction network and the class center of each class in the shared embedding space; A6. Loop through steps A2 to A5 until the training end condition is met, thereby obtaining a trained cross-modal matching model.

2. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 1, characterized in that: The unit hypersphere is used as the shared embedding space; in step A32, the features of each modality to be matched are normalized by L2 norm according to the following formula, and mapped to the unit hypersphere in the form of a tensor: Among them, norm represents L2 norm normalization, f j () represents the feature extraction network of the jth modality to be matched, represents the data of the jth mode to be matched for the i-th training sample, Represents the point of the jth to-be-matched modality of the i-th training sample in the shared embedding space.

3. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 1, characterized in that: The preset distance parameter m has a value of 0.3-0.

5.

4. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 1, characterized in that: In step A2, training samples are input in batches; The loss in step A4 Also includes cross-modal affinity loss And calculated according to the following formula: Among them, λ cma is the weight, is the adaptive cosine center loss, is the cross-modal affinity loss; The cross-modal affinity loss The calculation of includes the following steps: First, the points of each to-be-matched modality of each training sample obtained in step A32 in the shared embedding space are taken as the point set S; then, the point set S is divided into subsets according to category: And calculate each subset separately Loss of affinity k p Indicates the category of the training samples corresponding to the points contained in the subset: Where M is the number of modes to be matched, For subset The number of points contained, t is the preset scaling factor, z and z′ represent the subsets points in the, z≠z′ means z and z′ are different points, θ zz′ represents the vector angle between points z and z′ in the shared embedding space, cos(·) represents the cosine similarity function, and e is a natural constant; Afterwards, the cross-modal affinity loss is calculated according to the following formula: Wherein, K′ represents the number of categories whose number of training samples input in step A2 is greater than 1.

5. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 4, characterized in that: The value of the scaling factor t is 1-3.

6. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 4, characterized in that: The loss in step A4 It also includes classification loss And calculated according to the following formula: Among them, λ cma and λ ce is the weight, is the adaptive cosine center loss, is the cross-modal affinity loss, is the classification loss; The classification loss Calculated as follows: Where M is the number of modalities to be matched, N is the number of training samples of the input model in step A2, and y i Represents the category label of the i-th training sample; Based on the features of the j-th modality to be matched of the i-th training sample obtained in step A32, a shared classifier is used to obtain the predicted classification of the j-th modality to be matched of the i-th training sample; each modality to be matched is classified and predicted by the same shared classifier, and in step A5, the loss calculated in step A4 is used to update the parameters of the shared classifier based on the gradient.

7. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 6, characterized in that: The weight λ cma The value of is preset to 0.1 to 1, and the weight λ ce The value of is preset to 0.1~1.

8. A method for constructing a cross-modal matching model of a three-dimensional model according to any one of claims 1 to 7, characterized in that: The modalities to be matched of the training samples include at least two types of modalities among three-dimensional point cloud, three-dimensional grid and two-dimensional image.

9. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 8, characterized in that: The modalities to be matched of the training samples include three-dimensional point clouds, three-dimensional meshes and two-dimensional images; in the training sample set, the three-dimensional point cloud data and three-dimensional mesh data contained in each training sample are derived from the same instance, and the two-dimensional image data contained in each training sample are derived from the same category as the three-dimensional point cloud data and three-dimensional mesh data.

10. The method for constructing a cross-modal matching model of a three-dimensional model according to claim 9, characterized in that: The feature extraction networks corresponding to each modality are: DGCNN network for extracting features of three-dimensional point cloud data, MeshNet network for extracting features of three-dimensional mesh data, and ResNet18 network for extracting features of two-dimensional image data.

Citation Information

Cited By

  • Reconstruction-guided multi-modal CAD (computer-aided design) retrieval method and system

    CN122364477A