Target retrieval model obtaining method, retrieval method and device of 3D model
Through self-supervised learning and hierarchical classification fine-tuning training, the feature extraction and matching of the 3D model retrieval method is optimized, and the problems of difficulty in retrieval, low efficiency and low accuracy in the existing technology are solved, and efficient, flexible and accurate 3D model retrieval is achieved in the early stage of user business.
Patent Information
- Application Number
- CN202510363708.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-22
AI Technical Summary
The existing 3D model retrieval methods have problems such as difficulty in searching operations, low retrieval efficiency, low accuracy, low reliability and scarce 3D model classification labels. Especially when users cannot provide a complete 3D model, the flexibility and convenience of the traditional methods are insufficient.
The self-supervised learning method is used to train the feature extraction of three-view and multi-view, and combine the SimCLR loss function and the VICREG loss function to optimize the feature representation. The training set and the second target loss function are fine-tuned by hierarchical classification to obtain the trained target retrieval model, which can learn effective feature representation from labelless data and conduct efficient 3D model retrieval.
It improves the accuracy and reliability of 3D model retrieval, lowers the threshold for user use, enables users to efficiently search through simple three-views, meets a wider range of user needs, and reduces dependence on professional 3D modeling capabilities and equipment.
Smart Images

Figure CN120356029A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional model data retrieval and recognition. Specifically, it relates to a method, a retrieval method, and a device for obtaining a target retrieval model of a 3D model. Background Art
[0002] With the rapid development of computer graphics and computer vision technologies, 3D model retrieval has been widely applied in many fields such as industrial design, virtual reality, and augmented reality.
[0003] Traditional 3D model retrieval methods mostly rely on complete 3D model data and achieve retrieval by comparing features such as geometric shapes, textures, and colors between models. Although some existing methods attempt to use 2D three-view drawings for 3D model retrieval, they have strict requirements on the number and orientation of input views. For example, users are required to provide a specific number of 2D three-view drawings, or the three-view drawings must be provided in a specific order and direction. Such strict restrictions increase the operation difficulty for users and reduce the flexibility and convenience of the retrieval method. However, in practical applications, especially in the early stage of users' business, it is difficult to obtain complete 3D models. Users often can only provide simple 2D three-view drawings, such as the three-view drawings in the front-back, left-right, and up-down directions. This severely limits the applicable scope of retrieval methods based on 3D model input and cannot meet the actual needs of users, resulting in low retrieval efficiency. At the same time, the existing feature extraction and feature matching methods based on 2D three-view drawings and multi-views generally have poor accuracy problems: due to the limited information volume of 2D views and being interfered by factors such as viewing angles, lighting, and occlusion, the extracted features are not precise and stable enough, thereby affecting the accuracy and reliability of retrieval results.
[0004] Furthermore, traditional 3D model retrieval methods usually rely on a large number of classification labels for feature extraction and model matching. However, in actual scenarios, it is difficult to obtain classification labels or the number of classification labels is limited, which seriously affects the performance of retrieval methods. Especially in specific application scenarios such as custom model retrieval in industrial design, the lack of classification labels makes it difficult for traditional retrieval methods to operate effectively.
[0005] Therefore, the existing 3D model retrieval methods have problems such as difficult retrieval operations, low retrieval efficiency, low accuracy, low reliability, and scarcity of 3D model classification labels. Summary of the Invention
[0006] The main objective of the present invention is to provide a method for obtaining a target retrieval model of a 3D model, a retrieval method, a device for obtaining a target retrieval model, a retrieval device, a terminal device, and a storage medium, aiming to solve the technical problems that the existing 3D model retrieval methods have difficult retrieval operations, low retrieval efficiency, low accuracy, low reliability, and scarcity of 3D model classification labels.
[0007] In a first aspect, the present invention provides a retrieval method for a 3D model, including:
[0008] Perform orthogonal direction truncation on the 3D model to obtain a three-view training set; wherein, the three-view training set includes the three views of the 3D model in each orthogonal direction.
[0009] Perform two-dimensional projection on the 3D model to obtain a multi-view auxiliary training set; wherein, the multi-view auxiliary training set includes multi-views of the 3D model projected at different angles.
[0010] Based on the three-view training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and a first target loss function, perform feature extraction training on the target retrieval model to be trained.
[0011] Based on a fine-tuning training set with classification labels and a second target loss function, perform hierarchical classification fine-tuning training on the target retrieval model to obtain a trained target retrieval model.
[0012] Further, after performing orthogonal direction truncation on the 3D model to obtain a three-view training set, it further includes:
[0013] Perform grayscale processing and image contrast enhancement on the lines in the three views.
[0014] Further, the first target loss function includes a SimCLR loss function and a VICREG loss function. The performing feature extraction training on the target retrieval model to be trained based on the three-view training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and the first target loss function includes:
[0015] Perform feature extraction on the three views in the three-view training set through the target retrieval model to obtain target three-view feature vectors.
[0016] Perform feature extraction on the multi-views in the multi-view auxiliary training set through the target retrieval model to obtain target multi-view feature vectors.
[0017] Calculate the SimCLR loss function and the VICREG loss function for the target three-view feature vectors and the target multi-view feature vectors, so that the target three-view feature vectors approach the target multi-view feature vectors.
[0018] Further, the target retrieval model includes a convolutional neural network, a view pooling layer, and a fully connected layer connected in sequence.
[0019] The performing feature extraction on the three views in the three-view training set through the target retrieval model to obtain target three-view feature vectors includes:
[0020] Feature extraction is performed on the three - view drawings in the three - view training set through the convolutional neural network to obtain an initial three - view feature vector;
[0021] Feature fusion is performed on the initial three - view feature vector through the view pooling layer and the fully - connected layer to obtain intermediate three - view feature vectors for multiple pairs of opposite faces;
[0022] The SimCLR loss function and the VICREG loss function are calculated for the intermediate three - view feature vectors of each pair of opposite faces to obtain target three - view feature vectors;
[0023] The obtaining of the target multi - view feature vector by performing feature extraction on the multi - view drawings in the multi - view auxiliary training set through the target retrieval model includes:
[0024] Feature extraction is performed on the multi - view drawings in the multi - view auxiliary training set through the convolutional neural network to obtain an initial multi - view feature vector;
[0025] Feature fusion is performed on the initial multi - view feature vector through the view pooling layer and the fully - connected layer to obtain a target multi - view feature vector.
[0026] Furthermore, the second target loss function includes a hierarchical classification loss function, a first triplet center loss function, and a second triplet center loss function, and the fine - tuning training set includes a first - level class vector distribution and a second - level class vector distribution;
[0027] The hierarchical classification fine - tuning training of the target retrieval model based on the fine - tuning training set with classification labels and the second target loss function to obtain a trained target retrieval model includes:
[0028] Feature extraction and hierarchical classification processing are performed on the three - view drawings rendered by multiple second - level class 3D models through the target retrieval model to respectively obtain the first - level class vector distribution and the second - level class vector distribution; wherein, the second - level class 3D models are collected by clustering and labeling according to the model shape features;
[0029] Based on the first - level class vector distribution, the second - level class vector distribution, and the hierarchical classification loss function, hierarchical classification training is performed on the target retrieval model so that the target retrieval model learns the manually calibrated hierarchical classification information;
[0030] Based on the first triplet center loss function, the second triplet center loss function, the first - level class vector distribution, and the second - level class vector distribution, metric learning is performed on the target retrieval model to obtain a trained target retrieval model.
[0031] Further, the feature extraction and hierarchical classification processing of the three-view drawings generated by rendering multiple three-dimensional (3D) models of the secondary classes through the target retrieval model respectively obtain the distribution of the first-class vectors and the distribution of the second-class vectors, including:
[0032] Feature extraction is performed on the three-view drawings generated by rendering multiple 3D models of the secondary classes through the target retrieval model to obtain classified three-view feature vectors;
[0033] The classified three-view feature vectors are subjected to hierarchical classification processing through two fully connected layer branches of the target retrieval model to respectively obtain the distribution of the first-class vectors and the distribution of the second-class vectors.
[0034] In a second aspect, the present invention provides a method for retrieving 3D models, which adopts a method for obtaining a target retrieval model of 3D models as described in the first aspect. The retrieval method includes:
[0035] Obtain the three-view drawings to be retrieved of the 3D model to be retrieved;
[0036] Feature extraction is performed on the three-view drawings of the 3D model to be retrieved through the trained target retrieval model to obtain the feature vectors of the three-view drawings to be retrieved;
[0037] Based on the maximum pooling algorithm or the Hungarian algorithm, feature matching is performed on the feature vectors of the three-view drawings to be retrieved, and the target 3D model matching the 3D model to be retrieved is retrieved.
[0038] In a third aspect, the present invention provides an apparatus for obtaining a target retrieval model of 3D models, including:
[0039] A three-view drawing training set acquisition module, configured to perform orthogonal direction truncation on a 3D model to obtain a three-view drawing training set; wherein, the three-view drawing training set includes the three-view drawings of the 3D model truncated in each orthogonal direction;
[0040] A multi-view auxiliary training set module, configured to perform two-dimensional projection on the 3D model to obtain a multi-view auxiliary training set; wherein, the multi-view auxiliary training set includes multi-views of the 3D model projected at different angles;
[0041] A feature extraction training module, configured to perform feature extraction training on the target retrieval model to be trained based on the three-view drawing training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and a first target loss function;
[0042] A hierarchical classification fine-tuning training module, configured to perform hierarchical classification fine-tuning training on the target retrieval model based on a fine-tuning training set with classification labels and a second target loss function to obtain a trained target retrieval model.
[0043] Fourth aspect, the present invention provides a 3D model retrieval device, which adopts a target retrieval model acquisition device for 3D models as described in the third aspect. The retrieval device includes:
[0044] A module for obtaining three views to be retrieved, configured to obtain the three views to be retrieved of the 3D model to be retrieved;
[0045] A module for feature vectors of three views to be retrieved, configured to extract features from the three views of the 3D model to be retrieved through the trained target retrieval model, and obtain the feature vectors of the three views to be retrieved;
[0046] A feature matching module, configured to perform feature matching on the feature vectors of the three views to be retrieved based on the max pooling algorithm or the Hungarian algorithm, and retrieve the target 3D model that matches the 3D model to be retrieved.
[0047] Fifth aspect, the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a method for obtaining a target retrieval model for 3D models as described in the first aspect or a method for retrieving 3D models as described in the second aspect.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] (1) Aiming at the problem of scarce classification labels for existing 3D models, the present invention adopts a self-supervised learning method to train the feature extraction of three views and multi-views. Through self-supervised learning, the target retrieval model can learn effective feature representations from unlabeled data, thus solving the retrieval problem caused by insufficient classification labels. This innovation not only improves the retrieval accuracy and ensures the accuracy and reliability of the retrieval results, but also makes the present invention more flexible and extensive in practical applications.
[0050] (2) Based on the fine-tuning training set with classification labels and the second target loss function, the present invention performs hierarchical classification fine-tuning training on the target retrieval model to obtain a trained target retrieval model. Through the target retrieval model after hierarchical classification fine-tuning training, even if a complete 3D model cannot be obtained in the early stage of the user's business, efficient 3D model retrieval can be performed through simple three views (such as front and back, left and right, up and down). This feature greatly reduces the user's usage threshold, and the retrieval operation is simple, enabling users to still perform 3D model retrieval conveniently without professional 3D modeling capabilities or equipment, meeting the needs of a wider range of users. Description of the Drawings
[0051] Figure 1 is a schematic flowchart of a method for obtaining a target retrieval model for 3D models provided by an embodiment of the present invention;
[0052] Figure 2 Schematic diagram of the process of self-supervised learning for feature extraction by the target retrieval model based on multi-view and three-view provided in an embodiment of the present invention;
[0053] Figure 3 Schematic diagram of the process of hierarchical classification fine-tuning training for the target retrieval model based on three three-view provided in an embodiment of the present invention;
[0054] Figure 4 Schematic diagram of the process of a 3D model retrieval method provided in an embodiment of the present invention;
[0055] Figure 5 Schematic diagram of the process based on the three three-view feature matching strategy provided in an embodiment of the present invention;
[0056] Figure 6 Schematic diagram of the structure of a 3D model target retrieval model acquisition device provided in an embodiment of the present invention;
[0057] Figure 7 Schematic diagram of the structure of a 3D model retrieval device provided in an embodiment of the present invention;
[0058] Figure 8 Schematic diagram of the structure of a terminal device provided in an embodiment of the present invention.
[0059] Wherein:
[0060] The realization, functional characteristics and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0061] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0062] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0063] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "mounted", "connected" and "coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0064] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.
[0065] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of a method for obtaining a target retrieval model of a 3D model provided by an embodiment of the present invention.
[0066] A method for obtaining a target retrieval model of a 3D model according to an embodiment of the present invention includes the following steps:
[0067] S100. Perform orthogonal direction truncation on the 3D model to obtain a three-view training set; wherein, the three-view training set includes the three views of the 3D model in each orthogonal direction.
[0068] In this embodiment, an open-source tool OCC or the tool Hoops Exchange is used to perform three-view truncation on the 3D model in the front-back, left-right, and up-down (i.e., six orthogonal directions) to obtain a three-view training set.
[0069] In a specific embodiment, after the step S100 of performing orthogonal direction truncation on the 3D model to obtain a three-view training set, the following steps are further included:
[0070] S110. Perform grayscale processing and image contrast enhancement on the lines in the three views.
[0071] In this embodiment, grayscale processing is performed on the lines in the intercepted three views. By reducing the image difference between the strengthened lines and the opposite faces, the image contrast is improved, making the features of the three views more obvious. This step helps to improve the accuracy and robustness of subsequent feature extraction. And the intercepted three views are processed through line enhancement to improve the visibility and contrast of the lines in the image, eliminate the influence of virtual and real lines in the opposite faces, enhance the recognizability of the features, and at the same time reduce the difference between the opposite faces.
[0072] S200. Perform two-dimensional projection on the 3D model to obtain a multi-view auxiliary training set; wherein, the multi-view auxiliary training set includes multi-views of the 3D model projected at different angles.
[0073] In this embodiment, two-dimensional projection is performed on the 3D model to generate an additional 20 multi-views (such as projection diagrams at different angles) as auxiliary information. Specifically, the vtk rendering tool is used to render and generate 20 multi-views from the three-dimensional CAD model (such as cross-sections at four angles of 0°, 45°, 90°, and 135°, and rendering sampling is performed at intervals of 60° for each cross-section).
[0074] Steps S100 and S200 ensure the diversity and integrity of the input data, providing a solid foundation for subsequent feature extraction and retrieval.
[0075] S300. Based on the three-view training set, the multi-view auxiliary training set, the self-supervised learning algorithm, and the first target loss function, perform feature extraction training on the target retrieval model to be trained.
[0076] In the feature extraction stage, the embodiment of the present invention adopts a self-supervised learning method to perform feature extraction training for three views and multiple views. In the context of limited classification labels, self-supervised learning has shown its unique advantages in the field of 3D model retrieval. As an efficient technical means, self-supervised learning can use auxiliary tasks (pretext) to mine its own supervisory information from large-scale unsupervised data, and train the network through this constructed supervisory information, so as to learn valuable representations for downstream tasks without relying on external labeled data. This method can effectively improve the generalization ability and feature extraction effect of the target retrieval model with a small amount of annotation resources. Through self-supervised learning, the target retrieval model can learn useful feature representations from the data itself, and then achieve more accurate recognition and retrieval in 3D model retrieval tasks. Moreover, self-supervised training of massive CAD data can improve the generalization and robustness of geometric feature extraction of the target retrieval model.
[0077] In a specific embodiment, the first objective loss function includes a SimCLR loss function and a VICREG loss function, and the step S300 performs feature extraction training on the target retrieval model to be trained based on the three-view training set, the multi-view auxiliary training set, the self-supervised learning algorithm and the first objective loss function, including the following steps:
[0078] S310, extracting features of the three views in the three-view training set by using the target retrieval model to obtain target three-view feature vectors;
[0079] S320, extracting features from multiple views in the multi-view auxiliary training set by using the target retrieval model to obtain a target multi-view feature vector;
[0080] S330 , performing the SimCLR loss function and the VICREG loss function calculation on the target three-view feature vector and the target multi-view feature vector, so as to make the target three-view feature vector approach the target multi-view feature vector.
[0081] In this embodiment, the embodiment of the present invention adopts the SimCLR loss function and the VICREG loss function to guide the learning process of the target retrieval model. These two loss functions not only make the feature representations extracted from three views and multiple views continuously approach each other in the vector space, but also can shorten the distance between the features of the opposite faces in the three views in the vector space. During the training process, the embodiment of the present invention calculates the loss function for each of the three pairs of opposite faces to optimize the distance metric in the feature space. In this way, the target retrieval model can learn effective feature representations from unlabeled data, thereby improving the distinctiveness of the features and the accuracy of retrieval.
[0082] In a specific embodiment, the target retrieval model includes a Multi-view Convolutional Neural Networks (MVCNN).
[0083] Please refer to Figure 2 , Figure 2 which is a schematic flow chart of self-supervised learning for feature extraction by the target retrieval model based on multi-view and three-view provided in an embodiment of the present invention.
[0084] The self-supervised learning framework of the target retrieval model may include two branches: a three-view feature vector extraction branch and a multi-view feature vector extraction branch. Specifically, the target retrieval model may include a three-view feature extraction neural network and a multi-view feature extraction neural network. The three-view in the three-view training set is subjected to feature extraction through the three-view feature extraction neural network to obtain the target three-view feature vector, and the multi-view in the multi-view auxiliary training set is subjected to feature extraction through the multi-view feature extraction neural network to obtain the target multi-view feature vector.
[0085] Optionally, the self-supervised learning framework can batch process N = 10 3D models.
[0086] Please refer to Figure 3 , in a specific embodiment, the target retrieval model includes a convolutional neural network, a view pooling layer, and a fully connected layer connected in sequence.
[0087] The step S310 extracts features from the three-views in the three-view training set through the target retrieval model to obtain the target three-view feature vector, including the following steps:
[0088] S311. Extract features from the three-views in the three-view training set through the convolutional neural network to obtain an initial three-view feature vector;
[0089] S312. Perform feature fusion on the initial three-view feature vector through the view pooling layer and the fully connected layer to obtain multiple pairs of intermediate three-view feature vectors of opposite faces;
[0090] S313. Calculate the SimCLR loss function and the VICREG loss function for each pair of intermediate three-view feature vectors of opposite faces to obtain the target three-view feature vector;
[0091] The step S320 extracts features from the multi-views in the multi-view auxiliary training set through the target retrieval model to obtain the target multi-view feature vector, including the following steps:
[0092] S321. Extract features from the multi-views in the multi-view auxiliary training set through the convolutional neural network to obtain an initial multi-view feature vector;
[0093] S322. Perform feature fusion on the initial multi-view feature vector through the view pooling layer and the fully connected layer to obtain a target multi-view feature vector.
[0094] In this embodiment, the processed three-view drawings pass through a pre-trained convolutional neural network to obtain 6 initial three-view drawing feature vectors. Then, through the view pooling layer and fully connected layers, the implementation feature fusion strategy is processed to obtain intermediate three-view drawing feature vectors for multiple pairs of opposite faces. For the intermediate three-view drawing feature vectors of each pair of opposite faces, the SimCLR loss function and the VICREG loss function are used to control the intermediate three-view drawing feature vectors of each pair of opposite faces and approach, reducing the distance between them in the vector space, and finally obtaining a 512-dimensional target three-view drawing feature vector Z representing the three-view drawings i .
[0095] In one embodiment, the convolutional neural network includes a VGG11 neural network.
[0096] The multi-views are also respectively processed by the same pre-trained convolutional neural network to obtain 20 initial multi-view feature vectors, and then feature fusion processing is performed through the view pooling layer and fully connected layers to obtain a 512-dimensional target multi-view feature vector h i .
[0097] Finally, the SimCLR loss function and the VICREG loss function are used to control the cross-modal target three-view drawing feature vector Z i and the target multi-view feature vector h i to approach and prevent the expansion of feature parameters.
[0098] In a specific embodiment, both the three-view drawing feature extraction neural network and the multi-view feature extraction neural network include a convolutional neural network, a view pooling layer, and a fully connected layer connected in sequence. The three-view drawings in the three-view drawing training set are sequentially processed by the convolutional neural network, the view pooling layer, and the fully connected layer of the three-view drawing feature extraction neural network to obtain a target three-view drawing feature vector; the multi-views in the multi-view training set are sequentially processed by the convolutional neural network, the view pooling layer, and the fully connected layer of the multi-view feature extraction neural network to obtain a target multi-view feature vector.
[0099] S400. Based on the fine-tuning training set with classification labels and the second objective loss function, perform hierarchical classification fine-tuning training on the target retrieval model to obtain a trained target retrieval model.
[0100] In this embodiment, through the hierarchical classification fine-tuning training of the three views, the present invention can further optimize the performance of the target retrieval model to meet the actual needs of users. This method not only improves the accuracy of 3D model retrieval, but also reduces the dependence on a large amount of labeled data through self-supervised learning, making the target retrieval model more flexible and adaptable, and capable of better meeting various needs in practical applications.
[0101] In a specific embodiment, the second objective loss function includes a hierarchical classification loss function, a first triplet center loss function, and a second triplet center loss function, and the fine-tuning training set includes a first-level class vector distribution and a second-level class vector distribution; the step S400 of performing hierarchical classification fine-tuning training on the target retrieval model based on the fine-tuning training set with classification labels and the second objective loss function to obtain a trained target retrieval model includes the following steps:
[0102] S410. Extract features and perform hierarchical classification processing on the three views rendered by multiple second-level class 3D models through the target retrieval model to respectively obtain the first-level class vector distribution and the second-level class vector distribution; wherein, the second-level class 3D models are collected by clustering and labeling according to the model shape features.
[0103] S420. Based on the first-level class vector distribution, the second-level class vector distribution, and the hierarchical classification loss function, perform hierarchical classification training on the target retrieval model so that the target retrieval model learns the manually calibrated hierarchical classification information.
[0104] S430. Based on the first triplet center loss function, the second triplet center loss function, the first-level class vector distribution, and the second-level class vector distribution, perform metric learning on the target retrieval model to obtain a trained target retrieval model.
[0105] In this embodiment, the embodiment of the present invention takes two key steps in the fine-tuning stage to optimize the feature extraction of the three-view images.
[0106] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of hierarchical classification fine-tuning training for a target retrieval model based on three three-view images provided by an embodiment of the present invention.
[0107] First, the embodiment of the present invention uses the three - view fine - tuning training data with classification labels to adjust the pre - trained MVCNN neural network. In this process, the embodiment of the present invention introduces metric learning, especially Triplet Center Loss, to optimize the distance metric in the feature space. Specifically, in the fine - tuning training, the embodiment of the present invention uses two Triplet Center Loss functions, TCL1 (the first Triplet Center Loss function) and TCL2 (the second Triplet Center Loss function), as part of the metric learning. TCL1 and TCL2 respectively correspond to the two - level classification tasks (the classification tasks of the first - level class and the second - level class) of the target retrieval model. These two loss functions help the target retrieval model learn more accurate feature representations by pulling closer the feature representations of samples of the same class and pushing away the feature representations of samples of different classes.
[0108] Secondly, to further refine the feature extraction ability of the model, the embodiment of the present invention introduces a fine - tuning training process based on the three - view hierarchical classification task. In this step, the embodiment of the present invention constructs a fine - grained model hierarchical classification framework (about 800 second - level classes) according to the clustering analysis of the model shape features. Then, using about 20,000 sets of fine - tuning training data with classification labels and the hierarchical classification loss function, the target retrieval model is subjected to hierarchical classification fine - tuning training, driving the target retrieval model to learn the manually calibrated hierarchical classification information. This process not only deepens the target retrieval model's understanding of the 3D model features, but also improves the target retrieval model's recognition and discrimination ability for 3D models of different classes through the refined classification framework.
[0109] Thus, the core advantage of the present invention is that it allows for efficient 3D model retrieval through simple three - views (such as front - back, left - right, up - down) even when complete 3D models cannot be obtained in the early stage of the user's business. This feature greatly reduces the user's usage threshold, enabling users to still conveniently perform 3D model retrieval without professional 3D modeling capabilities or equipment, meeting the needs of a wider range of users.
[0110] In a specific embodiment, step S410 performs feature extraction and hierarchical classification processing on the three - views rendered by the target retrieval model for multiple second - level class 3D models, respectively obtaining the first - level class vector distribution and the second - level class vector distribution, including the following steps:
[0111] S411: Perform feature extraction on the three - views rendered by the target retrieval model for multiple second - level class 3D models to obtain classification three - view feature vectors;
[0112] S412. Perform hierarchical classification processing on the classified three - view feature vectors through two fully - connected layer branches of the target retrieval model to obtain the first - level class vector distribution and the second - level class vector distribution respectively.
[0113] In this embodiment, the present invention constructs a three - dimensional model data secondary classification framework (C1 first - level classes, C2 second - level classes) based on a real - world industrial resource library. For each second - level class, about 20 3D models are collected for fine - tuning training of the target retrieval model. Please refer to Figure 3 , the three three - views rendered from each 3D model are used to extract 512 - dimensional classified three - view feature vectors f by a target retrieval model pre - trained with self - supervision i , and then through 2 fully - connected layer branches, the first - level class vector distribution C1 and the second - level class vector distribution C2 are obtained respectively. These fine - tuning training data with classification labels can be used to perform hierarchical classification fine - tuning training on the three - view feature extraction neural network of the target retrieval model.
[0114] In summary, compared with the prior art, the beneficial effects of a 3D model retrieval method provided by an embodiment of the present invention are as follows:
[0115] (1) Aiming at the problem of scarce classification labels for existing 3D models, the present invention uses a self - supervised learning method to train the extraction of three - view and multi - view features. Through self - supervised learning, the target retrieval model can learn effective feature representations from unlabeled data, thus solving the retrieval problem caused by insufficient classification labels. This innovation not only improves the retrieval accuracy, ensures the accuracy and reliability of retrieval results, but also makes the present invention more flexible and extensive in practical applications.
[0116] (2) Based on a fine - tuning training set with classification labels and a second target loss function, the present invention performs hierarchical classification fine - tuning training on the target retrieval model to obtain a trained target retrieval model. Through the target retrieval model after hierarchical classification fine - tuning training, even in the early stage of the user's business when a complete 3D model cannot be obtained, efficient 3D model retrieval can be performed through simple three - views (such as front - back, left - right, up - down). This feature greatly reduces the user's usage threshold, and the retrieval operation is simple, enabling users to still perform 3D model retrieval conveniently without professional 3D modeling capabilities or equipment, meeting the needs of a wider range of users.
[0117] Please refer to Figure 4 , Figure 4 is a schematic flowchart of a 3D model retrieval method provided by an embodiment of the present invention.
[0118] A 3D model retrieval method according to an embodiment of the present invention uses a 3D model target retrieval model acquisition method as described in any one of the above - mentioned embodiments. The retrieval method includes the following steps:
[0119] S500. Obtain the three views to be retrieved of the 3D model to be retrieved;
[0120] S600. Extract features from the three views of the 3D model to be retrieved through the trained target retrieval model to obtain the feature vectors of the three views to be retrieved;
[0121] S700. Based on the max pooling algorithm or the Hungarian algorithm, perform feature matching on the feature vectors of the three views to be retrieved, and retrieve the target 3D model that matches the 3D model to be retrieved.
[0122] In this embodiment, features are extracted from the three views of the 3D model to be retrieved through the trained target retrieval model to obtain the feature vectors of the three views to be retrieved. After the feature extraction stage is completed, the present invention provides two flexible feature matching methods to adapt to different retrieval requirements and scenarios:
[0123] The main method adopted is to perform max pooling on the 3*512 feature vectors of the three views to be retrieved to extract the most prominent features for retrieval. This method is preferentially selected due to its efficiency and accuracy.
[0124] As an auxiliary strategy, the Hungarian algorithm is introduced for full mapping matching of features. This algorithm calculates the Euclidean distance between the selected 3*512 feature vectors of the three views to be retrieved and the corresponding 3*512 target three-view feature vectors, and selects the feature pairs with the smallest total distance and no duplicates, providing another option for improving the retrieval accuracy in specific scenarios.
[0125] Exemplarily, please refer to Figure 5 , Figure 5 which is a schematic flowchart of a three-view feature matching strategy provided by an embodiment of the present invention.
[0126] Specifically, the present invention provides two optional feature matching strategies.
[0127] The first strategy is to perform pooling on the feature vectors of the three views to be retrieved and the 3 512-dimensional feature vectors extracted from the target three views respectively through a max pooling layer, and obtain 1 512-dimensional vector representing the three three views respectively. Then calculate the Euclidean distance between the vectors of these three views to be retrieved and the 512-dimensional vectors of the target three views, and perform three-view feature matching in ascending order of the Euclidean distance.
[0128] The second strategy is not to perform max pooling on the 512-dimensional feature vectors of the three-view drawings to be retrieved extracted from the three-view drawings of Zhang San and the 512-dimensional feature vectors of the target three-view drawings, but to retain the originally extracted 3 512-dimensional feature vectors of the three-view drawings to be retrieved and the 512-dimensional feature vectors of the target three-view drawings. The Hungarian algorithm is used to calculate the distance between each feature vector of the three-view drawings to be retrieved and each vector of the target three-view drawing. Under the condition of ensuring that each vector can find a corresponding vector to achieve a full mapping (without repeatedly finding the same vector for correspondence), the shortest distance between the vectors of each pair of drawings is taken. After comparing the total distances, the three-view drawing feature matching is performed in ascending order of the Euclidean distance.
[0129] By combining these two strategies, the present invention can meet the demand for high-precision retrieval while maintaining the retrieval efficiency. At the same time, the implementation of this multi-dimensional strategy enables the system to more flexibly and accurately identify target features when processing complex retrieval tasks, thereby providing a better retrieval service.
[0130] In summary, compared with the prior art, the beneficial effects of a 3D model retrieval method provided by an embodiment of the present invention are as follows:
[0131] (1) Aiming at the problem of the scarcity of classification labels for existing 3D models, the present invention adopts a self-supervised learning method to train the extraction of three-view and multi-view features. Through self-supervised learning, the target retrieval model can learn effective feature representations from unlabeled data, thereby solving the retrieval problem caused by insufficient classification labels. This innovation not only improves the retrieval accuracy but also makes the present invention more flexible and extensive in practical applications.
[0132] (2) Based on a fine-tuning training set with classification labels and a second target loss function, the present invention performs hierarchical classification fine-tuning training on the target retrieval model to obtain a trained target retrieval model. Through the target retrieval model after hierarchical classification fine-tuning training, even if a complete 3D model cannot be obtained in the early stage of the user's business, efficient 3D model retrieval can be performed through simple three-view drawings (such as front and back, left and right, up and down). This feature greatly reduces the user's usage threshold, enabling users to still conveniently perform 3D model retrieval without professional 3D modeling capabilities or equipment, meeting the needs of a wider range of users.
[0133] (3) The present invention optimizes the number of input views. The trained target retrieval model only needs to take one three-view drawing each from up and down, left and right, and front and back to achieve high-precision 3D model retrieval. This flexible input limit minimizes the number of input views while ensuring the retrieval accuracy, reducing the user's operation complexity. At the same time, the present invention does not distinguish the order of front and back (up and down, left and right), further simplifying the user's operation process and improving the retrieval efficiency.
[0134] (4) In the feature matching stage, the present invention adopts diverse feature matching strategies aimed at enhancing the precision and credibility of retrieval results. The core method is to utilize the max-pooling technique to screen features, thereby extracting the most prominent features for efficient retrieval. Additionally, as a supplementary strategy, the Hungarian algorithm is deployed to handle un-fused features and achieve comprehensive mapping matching. This algorithm constructs a one-to-one corresponding feature matching relationship under the condition of the minimum total distance by calculating the Euclidean distance between the feature vectors of 3D models. As an auxiliary means, this strategy can play a role when higher requirements for retrieval accuracy are imposed, ensuring the optimization of retrieval results. Through this multi-dimensional strategy, the present invention can meet the demand for high-precision retrieval while maintaining retrieval efficiency.
[0135] Therefore, the 3D model retrieval strategy of the present invention has the advantages of being all-round, efficient, easy to operate, reliable, and highly accurate.
[0136] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an apparatus for obtaining a target retrieval model of a 3D model provided by an embodiment of the present invention.
[0137] An apparatus for obtaining a target retrieval model of a 3D model according to an embodiment of the present invention includes:
[0138] A three-view training set acquisition module, configured to perform orthogonal direction interception on the 3D model to obtain a three-view training set; wherein, the three-view training set includes the three views of the 3D model intercepted in each orthogonal direction;
[0139] A multi-view auxiliary training set module, configured to perform two-dimensional projection on the 3D model to obtain a multi-view auxiliary training set; wherein, the multi-view auxiliary training set includes the multi-views of the 3D model projected at different angles;
[0140] A feature extraction training module, configured to perform feature extraction training on the target retrieval model to be trained based on the three-view training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and a first target loss function;
[0141] A hierarchical classification fine-tuning training module, configured to perform hierarchical classification fine-tuning training on the target retrieval model based on a fine-tuning training set with classification labels and a second target loss function to obtain a trained target retrieval model.
[0142] The apparatus for obtaining a target retrieval model of a 3D model provided by an embodiment of the present invention can execute all steps and functions of the method for obtaining a target retrieval model of a 3D model provided by any of the above embodiments, and the specific functions of this apparatus will not be elaborated herein.
[0143] Please refer toFigure 7 , Figure 7 is a schematic structural diagram of a 3D model retrieval device provided by an embodiment of the present invention.
[0144] A 3D model retrieval device according to an embodiment of the present invention employs a target retrieval model acquisition device for 3D models as described in any of the above embodiments. The retrieval device includes:
[0145] A to-be-retrieved three-view acquisition module, configured to acquire the to-be-retrieved three-views of the to-be-retrieved 3D model;
[0146] A to-be-retrieved three-view feature vector module, configured to perform feature extraction on the three-views of the to-be-retrieved 3D model through the trained target retrieval model to obtain a to-be-retrieved three-view feature vector;
[0147] A feature matching module, configured to perform feature matching on the to-be-retrieved three-view feature vector based on the max pooling algorithm or the Hungarian algorithm to retrieve a target 3D model that matches the to-be-retrieved 3D model.
[0148] A 3D model retrieval device provided by an embodiment of the present invention can execute all steps and functions of a 3D model retrieval method provided by any of the above embodiments. The specific functions of this device will not be elaborated here.
[0149] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a terminal device provided by an embodiment of the present invention. The terminal device includes:
[0150] A processor 101, a memory 201, and a computer program stored in the memory 201 and configured to be executed by the processor 101. When the processor 101 executes the computer program, it implements a method for acquiring a target retrieval model for 3D models as described in any of the above embodiments or a 3D model retrieval method as described in any of the above embodiments.
[0151] The processor 101 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0152] The memory 201 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 201 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 201 and are called by the processor 101 to execute a method for obtaining a target retrieval model of a 3D model or a retrieval method of a 3D model according to an embodiment of the present invention;
[0153] The input / output interface 301 is used to implement information input and output;
[0154] The communication interface 401 is used to implement communication interaction between this device and other devices. It can communicate through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0155] The bus 501 transmits information between various components of the device (such as the processor 101, the memory 201, the input / output interface 301, and the communication interface 401);
[0156] Among them, the processor 101, the memory 201, the input / output interface 301, and the communication interface 401 are communicatively connected to each other inside the device through the bus 501.
[0157] The embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute a method for obtaining a target retrieval model of a 3D model according to any one of the above embodiments or a retrieval method of a 3D model according to any one of the above embodiments.
[0158] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The embodiments described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.
[0160] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0162] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0163] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0164] It should be understood that in the present invention, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0165] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0166] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0168] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0169] The preferred embodiments of the embodiments of the present invention have been described above with reference to the accompanying drawings, which does not limit the scope of the rights of the embodiments of the present invention. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. A method for obtaining a target retrieval model of a 3D model, characterized in that, Including: Performing orthogonal direction truncation on the 3D model to obtain a three-view training set; wherein, the three-view training set includes the three views of the 3D model in each orthogonal direction; Performing two-dimensional projection on the 3D model to obtain a multi-view auxiliary training set; wherein, the multi-view auxiliary training set includes multi-views of the 3D model projected at different angles; Based on the three-view training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and a first target loss function, performing feature extraction training on the target retrieval model to be trained; Based on a fine-tuning training set with classification labels and a second target loss function, performing hierarchical classification fine-tuning training on the target retrieval model to obtain a trained target retrieval model.
2. The method for obtaining a target retrieval model of a 3D model according to claim 1, wherein After performing orthogonal direction truncation on the 3D model to obtain a three-view training set, it further includes: Performing grayscale processing and image contrast enhancement on the lines in the three views.
3. A method for obtaining a target retrieval model of a 3D model according to claim 1, characterized in that, The first target loss function includes a SimCLR loss function and a VICREG loss function. The performing feature extraction training on the target retrieval model to be trained based on the three-view training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and a first target loss function includes: Performing feature extraction on the three views in the three-view training set through the target retrieval model to obtain target three-view feature vectors; Performing feature extraction on the multi-views in the multi-view auxiliary training set through the target retrieval model to obtain target multi-view feature vectors; Calculating the SimCLR loss function and the VICREG loss function for the target three-view feature vectors and the target multi-view feature vectors to make the target three-view feature vectors approach the target multi-view feature vectors.
4. A method for obtaining a target retrieval model of a 3D model according to claim 3, characterized in that, The target retrieval model includes a convolutional neural network, a view pooling layer, and a fully connected layer connected in sequence; The performing feature extraction on the three views in the three-view training set through the target retrieval model to obtain target three-view feature vectors includes: Performing feature extraction on the three views in the three-view training set through the convolutional neural network to obtain initial three-view feature vectors; Performing feature fusion on the initial three-view feature vectors through the view pooling layer and the fully connected layer to obtain multiple pairs of intermediate three-view feature vectors of opposite faces; Calculating the SimCLR loss function and the VICREG loss function for each pair of intermediate three-view feature vectors of opposite faces to obtain target three-view feature vectors; The performing feature extraction on the multi-views in the multi-view auxiliary training set through the target retrieval model to obtain target multi-view feature vectors includes: Performing feature extraction on the multi-views in the multi-view auxiliary training set through the convolutional neural network to obtain initial multi-view feature vectors; Performing feature fusion on the initial multi-view feature vectors through the view pooling layer and the fully connected layer to obtain target multi-view feature vectors.
5. The method for obtaining a target retrieval model of a 3D model according to claim 3, characterized in that The second target loss function includes a hierarchical classification loss function, a first triplet center loss function, and a second triplet center loss function, and the fine-tuning training set includes a first-level class vector distribution and a second-level class vector distribution; Based on the fine-tuning training set with classification labels and the second target loss function, performing hierarchical classification fine-tuning training on the target retrieval model to obtain a trained target retrieval model, including: Performing feature extraction and hierarchical classification processing on the three-view drawings generated by rendering multiple second-level class 3D models through the target retrieval model to respectively obtain the first-level class vector distribution and the second-level class vector distribution; wherein, the second-level class 3D models are collected by clustering and labeling according to the model shape features; Based on the first-level class vector distribution, the second-level class vector distribution, and the hierarchical classification loss function, performing hierarchical classification training on the target retrieval model so that the target retrieval model learns the manually calibrated hierarchical classification information; Performing metric learning on the target retrieval model based on the first triplet center loss function, the second triplet center loss function, the first-level class vector distribution, and the second-level class vector distribution to obtain a trained target retrieval model.
6. The method for obtaining a target retrieval model of a 3D model according to claim 5, characterized in that The performing feature extraction and hierarchical classification processing on the three-view drawings generated by rendering multiple second-level class 3D models through the target retrieval model to respectively obtain the first-level class vector distribution and the second-level class vector distribution includes: Performing feature extraction on the three-view drawings generated by rendering multiple second-level class 3D models through the target retrieval model to obtain classification three-view feature vectors; Performing hierarchical classification processing on the classification three-view feature vectors through two fully connected layer branches of the target retrieval model to respectively obtain the first-level class vector distribution and the second-level class vector distribution.
7. A retrieval method for 3D models, characterized in that, Adopting a method for obtaining a target retrieval model of a 3D model as described in any one of claims 1 to 6, the retrieval method including: Obtaining the three-view drawings to be retrieved of the 3D model to be retrieved; Performing feature extraction on the three-view drawings of the 3D model to be retrieved through the trained target retrieval model to obtain three-view feature vectors to be retrieved; Based on the maximum pooling algorithm or the Hungarian algorithm, performing feature matching on the three-view feature vectors to be retrieved to retrieve a target 3D model matching the 3D model to be retrieved.
8. An apparatus for obtaining a target retrieval model of a 3D model, characterized in that, Including: A three-view drawing training set acquisition module, configured to perform orthogonal direction truncation on a 3D model to obtain a three-view drawing training set; wherein, the three-view drawing training set includes the three-view drawings of the 3D model truncated in each orthogonal direction; A multi-view auxiliary training set module, configured to perform two-dimensional projection on the 3D model to obtain a multi-view auxiliary training set; wherein, the multi-view auxiliary training set includes multi-views of the 3D model projected at different angles; A feature extraction training module, configured to perform feature extraction training on the target retrieval model to be trained based on the three-view drawing training set, the multi-view auxiliary training set, a self-supervised learning algorithm, and a first target loss function; The hierarchical classification fine-tuning training module is used to perform hierarchical classification fine-tuning training on the target retrieval model based on the fine-tuning training set with classification labels and the second target loss function, so as to obtain a trained target retrieval model.
9. A retrieval device for 3D models, characterized in that, An apparatus for obtaining a target retrieval model of a 3D model according to claim 8, the retrieval apparatus comprising: A module for obtaining three views to be retrieved, configured to obtain the three views to be retrieved of the 3D model to be retrieved; A module for obtaining feature vectors of the three views to be retrieved, configured to extract features from the three views of the 3D model to be retrieved through the trained target retrieval model, so as to obtain feature vectors of the three views to be retrieved; A feature matching module, configured to perform feature matching on the feature vectors of the three views to be retrieved based on the max pooling algorithm or the Hungarian algorithm, so as to retrieve a target 3D model that matches the 3D model to be retrieved.
10. A terminal device, characterized in that, Comprising: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a method for obtaining a target retrieval model of a 3D model according to any one of claims 1 to 6 or a method for retrieving a 3D model according to claim 7.
Citation Information
Patent Citations
Three-dimensional model retrieval system and method based on parts division of hand-drawn draft
CN104850633A
Three-dimensional feature extraction method and device, electronic equipment and readable storage medium
CN117541810A
Multi-modal fusion fine tuning training method and device, electronic equipment and readable storage medium
CN119169405A