Three-dimensional cad model retrieval method and device based on progressive learning and medium

By employing a progressive learning approach and utilizing a shared-weight convolutional neural network and a view attention feature generation module, the problem of contextual information loss in 3D CAD model retrieval is solved, thereby improving the accuracy and precision of retrieval.

CN121434436BActive Publication Date: 2026-03-20SHENZHEN JIALICHUANG TECH DEV CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202512017141.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-20
Estimated Expiration
2045-12-30

AI Technical Summary

Technical Problem

Existing 3D CAD model retrieval methods have shortcomings in feature representation and retrieval accuracy. In particular, multi-view-based methods are prone to losing key contextual information and are difficult to accurately represent the subtle structural differences between model categories, resulting in low retrieval accuracy.

Method used

A progressive learning-based approach is adopted, which uses a convolutional neural network with shared weights, a global feature generation module, and a view attention feature generation module, combined with a classification output head, to process multi-view 2D images, extract 3D shape descriptors, and determine the target retrieval results in the feature library.

Benefits of technology

It improves the accuracy of 3D CAD model retrieval, preserves complete contextual information and local detail features, and enhances the similarity measurement and retrieval accuracy between models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434436B_ABST
    Figure CN121434436B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional CAD model retrieval method and device based on progressive learning and a medium, and the method comprises the following steps: inputting a multi-view two-dimensional image obtained by performing format conversion and multi-angle rendering on a three-dimensional CAD model to be retrieved into a progressive learning network to obtain a query descriptor; determining a target retrieval result based on all target descriptors selected from three-dimensional shape descriptors in a feature library, wherein each three-dimensional shape descriptor in the feature library is uniquely associated with a three-dimensional CAD model. The multi-view image is input into a convolutional neural network with shared weights of the progressive learning network, a global feature generation module, a view attention feature generation module and a classification output head, and then the view image is subjected to progressive learning from global to local, so that the query descriptor retaining complete context information is obtained, the most similar target retrieval result is determined from the feature library, and the accuracy of the three-dimensional CAD model retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of computer vision technology, and in particular to a method, apparatus, and medium for retrieving 3D CAD models based on progressive learning. Background Technology

[0002] Currently, deep learning technology has been introduced into 3D model retrieval research. Leveraging its powerful representation learning capabilities, deep learning can automatically extract multi-level, abstract feature representations from raw data, effectively improving the similarity measurement between models and retrieval accuracy. This technology provides a new research paradigm and technical path for feature representation and retrieval of 3D models. In the process of using deep learning technology to achieve CAD model retrieval, the construction of feature representations and the design of feature extraction models are key factors determining retrieval performance. Researchers have explored various feature extraction methods for 3D data structures. Voxel-based methods model 3D shapes as discrete voxel meshes, integrating internal structural information. Although voxelization simplifies the representation, the increased number of voxels significantly increases computational complexity. In contrast, point clouds depict 3D structures as discrete points, reducing dimensionality and enhancing feature recognition capabilities, but facing challenges due to irregularity and sparsity. Multi-view-based methods employ multi-view projection strategies, effectively mitigating the limitations of other methods by utilizing the pre-training advantages of 2D image data and stable multi-view structures. However, existing multi-view-based methods are prone to losing key contextual information, which weakens the expressive power of the final 3D shape descriptor and makes it difficult to accurately represent the subtle structural differences between model categories, thus failing to guarantee the accuracy of 3D CAD model retrieval results. Summary of the Invention

[0003] This application provides a method, apparatus, and medium for retrieving 3D CAD models based on progressive learning, which can improve the retrieval accuracy of 3D CAD models.

[0004] In a first aspect, embodiments of this application provide a method for retrieving 3D CAD models based on progressive learning, including:

[0005] Upon receiving the 3D CAD model to be retrieved, the model undergoes sequential format conversion and multi-angle rendering to obtain a multi-view result. Figure Two 3D image;

[0006] The multi-view Figure Two The 3D image is input into a trained progressive learning network to obtain a query descriptor corresponding to the 3D CAD model to be retrieved. The progressive learning network includes a convolutional neural network with shared weights, a global feature generation module, a view attention feature generation module, and a classification output head.

[0007] selecting a plurality of target descriptors corresponding to the query descriptor from three-dimensional shape descriptors in a preset feature library, determining a target retrieval result based on CAD models corresponding to all the target descriptors, wherein the number of the target descriptors is a target number, the similarity between the query descriptor and each of the target descriptors exceeds a first threshold, each of the target descriptors in the target retrieval result is sorted in descending order of the corresponding similarity, and each of the three-dimensional shape descriptors in the feature library is uniquely associated with a three-dimensional CAD model.

[0008] In some embodiments, the feature library is constructed according to the following steps:

[0009] obtaining three-dimensional CAD models generated by modeling of industrial parts, converting all the three-dimensional CAD models into STL format, and constructing a model database based on each three-dimensional CAD model after format conversion;

[0010] classifying each three-dimensional CAD model in the model database according to functional similarity and appearance similarity;

[0011] performing multi-angle rendering on each three-dimensional CAD model in the model database by setting a virtual camera to surround the model, to obtain a plurality of corresponding multi-view images of the three-dimensional CAD model; Figure Two

[0012] inputting all the multi-view images into the trained progressive learning network to output corresponding three-dimensional shape descriptors, storing all the three-dimensional shape descriptors, and constructing the feature library. Figure Two

[0013] In some embodiments, each three-dimensional CAD model in the model database is associated with a category identifier, and the progressive learning network is trained according to the following steps:

[0014] taking the multi-view images corresponding to all the three-dimensional CAD models of the same category identifier as an image set, dividing the image set to obtain a training set and a test set; Figure Two

[0015] training an initial network based on a preset joint classification loss function, combining all the training set and the test set, to obtain the progressive learning network, wherein the joint classification loss function is composed of a hinge loss function based on three-dimensional shape descriptors, and a cross-entropy loss function and a contrast loss function based on view features.

[0016] In some embodiments, the multi-view images include a plurality of view images, and the multi-view images are obtained by: Figure Two Figure Two ​​​​inputting the view images into the trained progressive learning network to obtain a query descriptor corresponding to the three-dimensional CAD model to be searched, comprising:

[0017] inputting each of the view images into the convolutional neural network with shared weights for feature extraction, and outputting corresponding first intermediate features;

[0018] inputting all the first intermediate features into the global feature generation module to output second intermediate features capable of representing global features corresponding to the three-dimensional CAD model to be searched;

[0019] inputting the second intermediate features into the view attention feature generation module to perform feature enhancement processing on each of the first intermediate features to obtain corresponding third intermediate features;

[0020] concatenating the second intermediate features and all the third intermediate features to obtain a query descriptor;

[0021] inputting the query descriptor into the classification output head to output a corresponding target category identifier, and associating the target category identifier with the query descriptor.

[0022] In some embodiments, the convolutional neural network with shared weights comprises an input module, a first residual stage, a second residual stage, a third residual stage, and a fourth residual stage. The input module comprises, in sequence, a convolutional layer with a convolution kernel size of 7x7 and a step size of 2, a batch normalization layer, a ReLU activation layer, and a maximum pooling layer with a step size of 2. The first residual stage is obtained by stacking three first residual blocks. Each first residual block comprises three convolutional layers connected in sequence. Each convolutional layer is connected with a batch normalization layer and a ReLU activation layer. The sizes of the three convolutional layers are 1x1, 3x3, and 1x1 in sequence. The input of the first residual block has a shortcut connection mechanism with the ReLU activation layer corresponding to the last convolutional layer in the first residual block. The second residual stage is obtained by stacking four second residual blocks. The first second residual block in the second residual stage can halve the size of the input feature map and double the number of channels. The remaining second residual blocks have the same structure as the first residual blocks. The third residual stage is obtained by stacking six third residual blocks. The first third residual block in the third residual stage can halve the size of the input feature map and double the number of channels. The remaining third residual blocks have the same structure as the first residual blocks. The fourth residual stage is obtained by stacking three fourth residual blocks. The first fourth residual block in the fourth residual stage can halve the size of the input feature map and double the number of channels. The remaining fourth residual blocks have the same structure as the first residual blocks.

[0023] In some embodiments, the first intermediate features are input into the global feature generation module, and second intermediate features capable of representing global features of the three-dimensional CAD model to be retrieved are output, including:

[0024] The first intermediate features are spliced to obtain a reference feature vector;

[0025] The reference feature vector is input into a multi-layer perception (MLP) for nonlinear transformation and information integration processing to obtain the second intermediate features.

[0026] In some embodiments, the second intermediate features are input into the view attention feature generation module for feature enhancement processing of each first intermediate feature to obtain corresponding third intermediate features, including:

[0027] Element-wise subtraction operations are performed on the second intermediate features and each first intermediate feature at a channel level to obtain fourth intermediate features;

[0028] All fourth intermediate features are input into a Sigmoid activation function to output an attention weight vector, wherein the attention weight vector includes multiple weight values, and any weight value is used to indicate a difference attention weight value between a view image with the same dimension as the corresponding first intermediate feature and the three-dimensional CAD model to be retrieved;

[0029] Each first intermediate feature is subjected to attention weighting processing based on each weight value to obtain a fifth intermediate feature;

[0030] Each first intermediate feature and the corresponding fifth intermediate feature are added to obtain the corresponding third intermediate feature.

[0031] In some embodiments, an initial network is trained based on a preset joint classification loss function in combination with the training set and the test set to obtain the progressive learning network, including:

[0032] The convolutional neural network with shared weights in the initial network is parameterized using pre-trained ResNet50 network weights;

[0033] The training set is input into the initial network after parameterization, and a predicted class label is output. A loss value is calculated based on the predicted class label, a corresponding real class label, and the joint classification loss function;

[0034] When the loss value is greater than or equal to a preset second threshold, a network parameter of the initial network is adjusted using an ADAM optimizer, the training set is input into the network after the model parameter is adjusted, retraining is performed, and an iteration training number is recorded until the loss value corresponding to any iteration training number is less than the second threshold, a network corresponding to the loss value less than the second threshold is determined as an intermediate network, or the current iteration training number reaches a preset maximum training period number, a network corresponding to the current iteration training number is determined as the intermediate network.

[0035] Based on the test set, accuracy analysis, recall rate analysis, and comprehensive index analysis are performed on the intermediate network, and an intermediate network meeting a preset condition is determined as the progressive learning network.

[0036] In a second aspect, an embodiment of the present application provides a three-dimensional CAD model retrieval device based on progressive learning, including at least one control processor and a memory in communication connection with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the three-dimensional CAD model retrieval method based on progressive learning as described in the first aspect.

[0037] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer executable instructions for executing the three-dimensional CAD model retrieval method based on progressive learning as described in the first aspect.

[0038] Embodiments of the present application provide a three-dimensional CAD model retrieval method, device and medium based on progressive learning, the method including: when receiving a three-dimensional CAD model to be retrieved, sequentially performing format conversion and multi-angle rendering processing on the three-dimensional CAD model to be retrieved to obtain multi-view images; inputting the multi-view images into a retrieval model to obtain a retrieval result; and outputting the retrieval result. Figure Two Figure Two ​A 3D image is input into a trained progressive learning network to obtain a query descriptor corresponding to the 3D CAD model to be retrieved. The progressive learning network includes a convolutional neural network with shared weights, a global feature generation module, a view attention feature generation module, and a classification output head. Multiple target descriptors corresponding to the query descriptor are selected from 3D shape descriptors in a preset feature library. The target retrieval result is determined based on all the target descriptors. The number of target descriptors is the number of targets. The similarity between the query descriptor and each of the target descriptors exceeds a first threshold. The target descriptors in the target retrieval result are sorted in descending order according to their corresponding similarity. Each 3D shape descriptor in the feature library is uniquely associated with a 3D CAD model. According to the scheme provided in the embodiments of this application, after obtaining multi-view images of the model to be retrieved, the multi-view features of the multi-view images are extracted through a convolutional neural network with shared weights in a progressive learning network. Global features for the model to be retrieved are generated through a global feature generation module. A three-dimensional shape descriptor is obtained by combining the global features and the enhanced view features through a view attention feature generation module. The three-dimensional shape descriptor is then associated with a category identifier through a classification output head. That is, after performing progressive learning on the view images, first globally and then locally, a query descriptor that retains complete contextual information is obtained, which is used to determine the most similar target retrieval result from the feature library and improves the accuracy of three-dimensional CAD model retrieval. Attached Figure Description

[0039] Figure 1 This is a flowchart of the steps of a 3D CAD model retrieval method based on progressive learning provided in one embodiment of this application;

[0040] Figure 2 This is a schematic diagram of a 3D CAD model retrieval method based on progressive learning provided in another embodiment of this application;

[0041] Figure 3 This is a schematic diagram illustrating the display of target retrieval results provided in another embodiment of this application;

[0042] Figure 4 This is a structural diagram of a 3D CAD model retrieval device based on progressive learning provided in another embodiment of this application. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0044] It is to be understood that, while the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in a manner different from the module division in the device or the order in the flowchart. The terms "first", "second", and the like in the specification, claims, or above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence.

[0045] At present, deep learning technology is introduced into three-dimensional model retrieval research. Deep learning technology can automatically extract multi-level and abstract feature representations from raw data due to its strong representation learning ability, thereby effectively improving the similarity measurement and retrieval accuracy between models. The introduction of this technology provides a new research paradigm and technical path for feature representation and retrieval of three-dimensional models. In the process of realizing CAD model retrieval by deep learning technology, the construction of feature representation and the design of feature extraction model are key factors that determine the retrieval performance. Researchers have explored various feature extraction methods for three-dimensional data structures. The voxel-based method models three-dimensional shapes as discrete voxel grids, integrating internal structural information. Although voxelization simplifies the representation, the increase in the number of voxels significantly increases the computational complexity. In contrast, point clouds depict three-dimensional structures as discrete points, which can reduce the dimension and enhance the feature recognition ability, but face the challenges brought by irregularity and sparsity. The multi-view-based method uses a multi-view projection strategy, effectively alleviating the limitations of other methods by utilizing the pre-training advantages of two-dimensional image data and the stable multi-view structure. However, existing multi-view-based methods easily lead to the loss of key context information, which weakens the expression ability of the final three-dimensional shape descriptor and makes it difficult to accurately represent the subtle structural differences between model categories, thereby failing to guarantee the accuracy of three-dimensional CAD model retrieval results.

[0046] To solve the above problems, the embodiments of the present application provide a three-dimensional CAD model retrieval method, device and medium based on progressive learning, which comprises: when receiving a three-dimensional CAD model to be retrieved, sequentially performing format conversion and multi-angle rendering processing on the three-dimensional CAD model to be retrieved to obtain multi-view images; inputting the multi-view images into a pre-trained multi-view feature extraction model to obtain a three-dimensional CAD model feature vector; and inputting the three-dimensional CAD model feature vector into a pre-trained three-dimensional CAD model retrieval model to obtain a three-dimensional CAD model retrieval result. Figure Two Figure Two ​The view image is input to the trained progressive learning network to obtain a query descriptor corresponding to the three-dimensional CAD model to be searched, wherein the progressive learning network comprises a convolutional neural network sharing weights, a global feature generation module, a view attention feature generation module, and a classification output head; a plurality of target descriptors corresponding to the query descriptor are selected from three-dimensional shape descriptors in a preset feature library, and a target search result is determined based on all the target descriptors, wherein the number of the target descriptors is a target number, the similarity between the query descriptor and each target descriptor exceeds a first threshold, each target descriptor in the target search result is sorted in descending order of the corresponding similarity, and each three-dimensional shape descriptor in the feature library is uniquely associated with a three-dimensional CAD model. According to the scheme provided in the embodiments of the present application, after the multi-view view images of the model to be searched are obtained, the multi-view features of the multi-view images are extracted by the convolutional neural network sharing weights of the progressive learning network, the global features for the model to be searched are generated by the global feature generation module, the three-dimensional shape descriptors are obtained by combining the global features and the enhanced view features through the view attention feature generation module, and the three-dimensional shape descriptors are associated with the class identifiers through the classification output head, that is, after the progressive learning of the view images from global to local, the query descriptor retaining complete context information is obtained, which facilitates the subsequent determination of the most similar target search result from the feature library, and improves the accuracy of three-dimensional CAD model search.

[0047] The embodiments of the present application are further described below with reference to the accompanying drawings.

[0048] Reference Figure 1 , Figure 1 is a step flowchart of a three-dimensional CAD model search method based on progressive learning provided by an embodiment of the present application. The embodiments of the present application provide a three-dimensional CAD model search method based on progressive learning, which comprises but is not limited to the following steps:

[0049] Step S10, when receiving a three-dimensional CAD model to be searched, sequentially performing format conversion and multi-angle rendering processing on the three-dimensional CAD model to be searched to obtain a plurality of view images corresponding to the three-dimensional CAD model to be searched. Figure Two View image.

[0050] It can be understood that after obtaining the three-dimensional CAD model to be searched, the embodiment converts the model into STL format regardless of the format of the model, such as STL, STP, STEP, 3MF, etc. After the format conversion is completed, a virtual camera is set to surround the three-dimensional CAD model to be searched in STL format for multi-angle rendering to obtain a plurality of view images corresponding to the three-dimensional CAD model to be searched. Figure TwoThe multi-view images V={Vi, 1≤i≤N} are input into the trained progressive learning network to obtain a query descriptor corresponding to the three-dimensional CAD model to be searched, wherein the progressive learning network comprises a convolutional neural network sharing weights, a global feature generation module, a view attention feature generation module, and a classification output head.

[0051] In step S20, the multi-view images of each three-dimensional CAD model in the model database are input into the initial network to obtain a three-dimensional shape descriptor corresponding to each three-dimensional CAD model. Figure Two In step S20, the multi-view images of each three-dimensional CAD model in the model database are input into the initial network to obtain a three-dimensional shape descriptor corresponding to each three-dimensional CAD model.

[0052] It should be noted that in some embodiments, each three-dimensional CAD model in the model database is associated with a category identifier, and the progressive learning network is trained according to the following steps:

[0053] The multi-view images corresponding to all three-dimensional CAD models corresponding to the same category identifier are input into the initial network to obtain a three-dimensional shape descriptor corresponding to each three-dimensional CAD model. Figure Two The multi-view images corresponding to all three-dimensional CAD models corresponding to the same category identifier are input into the initial network to obtain a three-dimensional shape descriptor corresponding to each three-dimensional CAD model.

[0054] Based on the preset joint classification loss function, the initial network is trained in combination with all the training set and test set to obtain the progressive learning network, wherein the joint classification loss function is composed of a hinge loss function based on the three-dimensional shape descriptor, and a cross-entropy loss function and a contrast loss function based on the view feature.

[0055] Specifically, in the present embodiment, the proportion of the training set and the test set corresponding to the same category identifier for training the progressive learning network is 8:2, the data source of the training set and the test set is each three-dimensional CAD model in the model database, the feature library is constructed based on the model database, and the multi-view images are obtained by multi-view rendering of the corresponding three-dimensional CAD model correctly placed in the Z-axis direction. γ The multi-view images are obtained by multi-view rendering of the corresponding three-dimensional CAD model correctly placed in the Z-axis direction.

[0056] It can be understood that the present embodiment realizes the collaborative optimization of local and global features based on the joint classification loss function combined by the single-view classification loss, the global classification loss, and the contrast loss to the initial network, improves the representation ability of the three-dimensional shape descriptor of the final progressive learning network, and provides effective support for subsequent extraction of accurate query descriptors.

[0057] Specifically, the expression of the joint classification loss function L in the present embodiment is as follows:

[0058] ;

[0059] ;

[0060] ;

[0061] ;

[0062] wherein, L v is a cross-entropy loss function based on view features, and cross-entropy loss of classification is calculated for each enhanced third intermediate feature x v e individually, aiming to improve the discriminative ability of a single view feature; N is the number of samples, V corresponds to the number of view features of each three-dimensional CAD model to be retrieved, and are the true label and the predicted value of the nth sample, respectively; L con is a contrastive loss function based on view features, and L con can enhance the feature similarity of the same class of samples under the same view angle and pull apart the feature distance of different class samples, and the contrastive loss is selectively applied to a part of the view, Q is the number of view images in each sample to which the contrastive loss is applied, and Sim(x q ei , x q ej ) represents the dot product of the features, and L h is a hinge loss function based on three-dimensional shape descriptors, which can use hinge loss for final classification prediction on the initial descriptor composed of global features (i.e., the second intermediate feature) and enhanced view features (i.e., the third intermediate feature), to obtain the final query descriptor, K is the total number of subcategories, b is the boundary value, and c ij is the predicted value of the ith sample belonging to class j, and δ{l i =j} is an indicator function, the true label of the ith sample is j, the function takes the value 1, otherwise the function takes the value -1, P is a norm used in regularization, and L1 norm is selected to prevent overfitting, α , β , Figure 1 are all balance parameters.

[0063] Further, based on the preset joint classification loss function, the initial network is trained in combination with the entire training set and test set to obtain a progressive learning network, including:

[0064] The pre-trained ResNet50 network weight is used to initialize the parameters of the shared weight convolutional neural network in the initial network;

[0065] inputting the training set into the initial network after the parameter initialization, outputting a predicted class label, and calculating a loss value based on the predicted class label, a corresponding real class label, and the joint classification loss function;

[0066] when the loss value is greater than or equal to a preset second threshold value, adjusting the network parameters of the initial network using an ADAM optimizer, retraining the network after the model parameters are adjusted, and recording the number of iterative training times until the loss value corresponding to any iterative training time is less than the second threshold value, determining the network corresponding to the loss value less than the second threshold value as an intermediate network, or determining the network corresponding to the current iterative training time as the intermediate network when the current iterative training time reaches a preset maximum training period;

[0067] performing accuracy analysis, recall rate analysis, and comprehensive index analysis on the intermediate network based on the test set, and determining the intermediate network meeting a preset condition as the progressive learning network.

[0068] It can be understood that the embodiment first uses the pre-trained ResNet50 network weight to initialize the parameters of the shared weight convolutional neural network in the initial network, which makes the convolutional neural network have good image feature extraction ability at the beginning of training, thereby accelerating the convergence speed of the whole network and helping to improve the generalization performance; then, the training set is input into the initial network after parameter initialization, and the initial descriptor is output. Based on the initial descriptor, the corresponding real descriptor and the joint classification loss function, the loss value is calculated, when the loss value is greater than or equal to the preset second threshold, the network parameters of the initial network are adjusted using the ADAM optimizer, the training set is input into the network after adjusting the model parameters for retraining, and the iteration training times are recorded, until the loss value corresponding to any iteration training time is less than the second threshold, the network corresponding to the loss value less than the second threshold is determined as the intermediate network, or the current iteration training time reaches the preset maximum training period, the network corresponding to the current iteration training time is determined as the intermediate network; that is, the joint classification loss function established in the foregoing is used as the optimization target, and the ADAM optimizer is used to update and optimize all network parameters (including the convolutional neural network, the global feature generation module, the attention feature generation module and the classification output head) in the entire initial network in an end-to-end manner. In the embodiment, the specific training hyperparameter setting is: the initial learning rate is 0.00001, the momentum is 0.9, the weight decay is 0, the mini-batch size is 8, and the Dropout rate is 0.5; based on such parameters, the initial network is continuously iteratively optimized, until the loss value corresponding to any iteration training time is less than the second threshold (that is, the corresponding loss function converges), or when the iteration times reach the preset maximum training period (100 periods in the embodiment), the training is terminated, and the trained network model (that is, the intermediate network) is saved for subsequent feature extraction and retrieval tasks; then, the intermediate network is further analyzed for accuracy, recall rate and comprehensive index using the test set, when each index meets the condition, the intermediate network is determined as the progressive learning network, and the training of the progressive learning network is completed.

[0069] In particular, in some embodiments, Figure 2 Step S20 includes but is not limited to the following steps:

[0070] Step S21, input each view image into the shared weight convolutional neural network for feature extraction, and output corresponding each first intermediate feature;

[0071] Step S22, input all first intermediate features into the global feature generation module, and output second intermediate features capable of representing the global features of the three-dimensional CAD model to be retrieved;

[0072] Step S23, inputting the second intermediate feature into a view attention feature generation module to perform feature enhancement processing on each first intermediate feature to obtain a corresponding third intermediate feature;

[0073] Step S24, splicing the second intermediate feature and all third intermediate features to obtain a query descriptor;

[0074] Step S25, inputting the query descriptor into a classification output head to output a corresponding target category identifier, and associating the target category identifier to the query descriptor.

[0075] It can be understood that in the embodiment, the reference Figure 2 , the convolutional neural network sharing the weights completes feature extraction on the N view images corresponding to the three-dimensional CAD model to be retrieved, and the convolutional neural network adopts a ResNet50 architecture; the global feature generation module can complete generation of a global feature representing the three-dimensional CAD model to be retrieved, i.e., obtain the second intermediate feature; the view attention feature generation module completes enhancement of the N view features (i.e., the first intermediate features) by taking the global feature (i.e., the second intermediate feature) as guidance to obtain the enhanced third intermediate features; finally, the second intermediate feature and the N enhanced third intermediate features are spliced to form a query descriptor, and then the query descriptor is classified through the classification output head to complete classification, so as to realize effective extraction of the three-dimensional shape descriptor of the model by the classification constraint network, and the target category identifier output by the classification output head is associated to the query descriptor. In this way, the progressive learning network can not only realize aggregation of multi-view information of the model to be retrieved, but also retain context information and local detail features, and the output query descriptor can guarantee determination of the most similar target retrieval result from the feature library, thereby improving the accuracy of three-dimensional CAD model retrieval.

[0076] Specifically, the convolutional neural network sharing weights in the embodiment includes an input module, a first residual stage, a second residual stage, a third residual stage, and a fourth residual stage, wherein the structures of the respective parts are as follows: the input module includes, in sequence, a convolutional layer with a convolution kernel size of 7x7 and a step of 2, a batch normalization layer, a ReLU activation layer, and a maximum pooling layer with a step of 2, the first residual stage is obtained based on stacking of three first residual blocks, each first residual block includes three convolutional layers connected in sequence, each convolutional layer is connected with a batch normalization layer and a ReLU activation layer, the sizes of the three convolutional layers are 1x1, 3x3, and 1x1 in sequence, the input of the first residual block has a shortcut connection mechanism with the ReLU activation layer corresponding to the last convolutional layer in the first residual block, the second residual stage is obtained based on stacking of four second residual blocks, the first second residual block in the second residual stage can halve the size of the input feature map and double the number of channels, the remaining second residual blocks have the same structure as the first residual block, the third residual stage is obtained based on stacking of six third residual blocks, the first third residual block in the third residual stage can halve the size of the input feature map and double the number of channels, the remaining third residual blocks have the same structure as the first residual block, and the fourth residual stage is obtained based on stacking of three fourth residual blocks, the first fourth residual block in the fourth residual stage can halve the size of the input feature map and double the number of channels, the remaining fourth residual blocks have the same structure as the first residual block.

[0077] It can be understood that the convolutional layer of the input module of the convolutional neural network sharing weights in the embodiment is used for preliminary feature extraction, the batch normalization layer is used for accelerating training and stabilizing the network, the ReLU activation layer is used for nonlinear transformation, and the maximum pooling layer is used for reducing the size of the feature map; after residual learning of the first residual stage, the second residual stage, the third residual stage, and the fourth residual stage, the global average pooling layer of the output module is used to reduce each input feature map to one value to reduce the number of parameters, then the feature is mapped to the final classification dimension through the fully connected layer, and finally the probability distribution of classification is output through the Softmax layer to obtain each view feature corresponding to the N view images of the to-be-retrieved three-dimensional CAD model, i.e., the first intermediate feature 1 , x 2 ,..., x v}, v =N.

[0078] Specifically, step S22 in the embodiment includes but is not limited to the following steps:

[0079] In step S221, the reference feature vector is obtained by splicing all the first intermediate features.

[0080] Step S222, input the reference feature vector into a multi-layer perception machine (MLP) for nonlinear transformation and information integration processing to obtain a second intermediate feature.

[0081] It is to be noted that the progressive learning network in this embodiment simulates the global scanning process of human vision through the global feature generation module, aiming to form an overall impression based on the preliminary features of all views; it performs a splicing operation on the N view features {x 1 , x 2 ,..., x v}, and generates a global feature representing the context information of the complete three-dimensional CAD model to be retrieved, i.e., the second intermediate feature X G , through an MLP function integration. X G is obtained according to the following formula:

[0082] ;

[0083] where [ ] is a splicing operation, and the integration function represented by MLP.

[0084] Specifically, step S23 in this embodiment includes but is not limited to the following steps:

[0085] Step S231, perform an element-by-element subtraction operation on the second intermediate feature and each first intermediate feature at the channel level to obtain a plurality of fourth intermediate features;

[0086] Step S232, input all fourth intermediate features into a Sigmoid activation function to output an attention weight vector, wherein the attention weight vector includes a plurality of weight values, and any weight value is used to indicate the difference attention weight value between the view image with the same dimension as the corresponding first intermediate feature and the three-dimensional CAD model to be retrieved;

[0087] Step S233, perform attention weighting processing on each first intermediate feature based on each weight value to obtain each fifth intermediate feature;

[0088] Step S234, add each first intermediate feature and the corresponding fifth intermediate feature to obtain each third intermediate feature.

[0089] It is to be noted that in this embodiment, the progressive learning network simulates the local fixation process of human vision through the view attention feature generation module, aiming to focus on the specific part of each view feature (i.e., the first intermediate feature) of the three-dimensional CAD model to be retrieved guided by the global impression (i.e., the second intermediate feature output by the global feature generation module). Specifically, first, it utilizes the second intermediate feature X GAs a guide, the difference between each first intermediate feature x v The difference between the single first intermediate feature and the second intermediate feature is obtained by performing an element-wise subtraction operation at the channel level, and the attention weight vector {A 1 , A 2 ,..., A v} is generated by the Sigmoid function, and the respective weight values corresponding to {A 1 , A 2 ,..., A v} are calculated according to the following formula:

[0090] ;

[0091] wherein, is any weight value.

[0092] Then, the respective weight values of the attention weight vector corresponding to the difference attention are multiplied point by point with the corresponding original view features, and the original view features x v are combined with the attention-weighted features through a residual connection to obtain the enhanced fifth intermediate feature x v e While preserving the original information, the significant features in each first intermediate feature are selectively enhanced to obtain the third intermediate feature, and the third intermediate feature x v e is calculated according to the following formula:

[0093] .

[0094] It can be understood that the embodiment processes the view image corresponding to the three-dimensional CAD model to be searched through the aforementioned convolutional neural network with shared weights, the global feature generation module, and the view attention feature generation module, and outputs an initial descriptor. Then, the initial descriptor is sent to the classification output head to complete classification, aiming to realize effective extraction of the final query descriptor by the network through the classification task, and to use the query descriptor to complete the subsequent search of the three-dimensional CAD model.

[0095] In step S30, a plurality of target descriptors corresponding to the query descriptor are selected from the three-dimensional shape descriptors in the preset feature library, and a target search result is determined based on the CAD models corresponding to all the target descriptors, wherein the number of target descriptors is a target number, the similarity between the query descriptor and each target descriptor exceeds a first threshold, each target descriptor in the target search result is sorted in descending order of the corresponding similarity, and each three-dimensional shape descriptor in the feature library is uniquely associated with a three-dimensional CAD model.

[0096] Understandably, after processing in step S20, a query descriptor is obtained that can characterize the contextual information and feature details of the 3D CAD model to be retrieved, as shown in the reference. Figure 3 and Figure 3 At this point, by calculating the similarity between the query descriptor and all 3D shape descriptors in the feature library, and sorting them in descending order of similarity, the CAD model with the highest similarity to the target is returned as the target retrieval result, such as... Figure Two As shown, in this embodiment, the CAD models corresponding to 10 target descriptors are returned as target retrieval results. The similarity of each target descriptor is sorted from high to low (100.00, 94.80, 92.23, 89.94, 89.18, 87.93, 87.77, 87.61) and displayed in the interface of the retrieval platform of the progressive learning-based 3D CAD model retrieval device in this embodiment.

[0097] It should be noted that the construction of the feature library in this embodiment includes, but is not limited to, the following steps:

[0098] The process involves acquiring 3D CAD models generated from industrial part modeling, converting all 3D CAD models to STL format, and building a model database based on the converted 3D CAD models.

[0099] The 3D CAD models in the model database are categorized according to functional similarity and appearance similarity.

[0100] By setting up a virtual camera to surround the model, various 3D CAD models in the model database are rendered from multiple angles to obtain their respective multi-view results. Figure Two 3D image;

[0101] All multi-view Figure Two The 3D image is input into the trained progressive learning network, which outputs the corresponding 3D shape descriptor, stores all the 3D shape descriptors, and constructs a feature library.

[0102] It is understood that the feature library construction method in this embodiment includes: acquiring 3D CAD models generated from industrial part modeling; converting all 3D CAD models into STL format; and constructing a model database based on the converted 3D CAD models; classifying the 3D CAD models in the model database according to functional similarity and appearance similarity; and performing multi-angle rendering on each 3D CAD model in the model database by setting a virtual camera to surround the model, thereby obtaining their respective multi-view results. Figure Two A multi-view image V = {Vi, 1 ≤ i ≤ N}, where each Vi is a view image with a size of 224 × 224 pixels; all multi-view images are represented.Figure 4 The inputted images are inputted into the trained progressive learning network, and corresponding three-dimensional shape descriptors are outputted. All the three-dimensional shape descriptors are stored and a feature library is constructed, thereby providing effective support for accurate model retrieval for a query descriptor obtained by feature extraction on a received three-dimensional CAD model to be retrieved through step S20.

[0103] As shown in Figure 4 , the structure diagram of the three-dimensional CAD model retrieval device based on progressive learning is provided in an embodiment of the present application. The present application further provides a three-dimensional CAD model retrieval device 400 based on progressive learning, comprising: ​ The processor 410 can be implemented in the form of a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0104] The memory 420 can be implemented in the form of a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 420 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 420 and are called and executed by the processor 410 to implement the three-dimensional CAD model retrieval method based on progressive learning of the embodiments of the present application.

[0105] The input / output interface 430 is used to realize information input and output.

[0106] The communication interface 440 is used to realize the communication interaction between the device and other devices. The communication can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0107] The bus 450 transmits information between various components (such as the processor 410, the memory 420, the input / output interface 430, and the communication interface 440) of the device.

[0108] The processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are connected to each other through the bus 450 for internal communication connection in the device.

[0109] ​

[0110] In addition, the embodiment of the present application further provides a storage medium, which is a computer readable storage medium, and stores a computer program. The computer program is executed by a processor to implement the three-dimensional CAD model retrieval method based on progressive learning.

[0111] The memory, as a non-transitory computer readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and the remote memory can be connected to the processor through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The above-described device embodiments are only schematic, and units described as separate units can or can not be physically separate, and can be implemented in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0112] Those skilled in the art can understand that all or some of the steps in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, as known to those skilled in the art, communication media generally includes computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and can include any information delivery medium.

[0113] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the above-described embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.

Claims

1. A method for retrieving 3D CAD models based on progressive learning, characterized in that, include: When a 3D CAD model to be retrieved is received, the model is sequentially converted into a new format and rendered from multiple angles to obtain a multi-view 2D image. The multi-view 2D image is input into a trained progressive learning network to obtain a query descriptor corresponding to the 3D CAD model to be retrieved. The progressive learning network includes a convolutional neural network with shared weights, a global feature generation module, a view attention feature generation module, and a classification output head. Multiple target descriptors corresponding to the query descriptor are selected from the three-dimensional shape descriptors in the preset feature library. The target retrieval result is determined based on the CAD model corresponding to all the target descriptors. The number of target descriptors is the number of targets. The similarity between the query descriptor and each of the target descriptors exceeds a first threshold. The target descriptors in the target retrieval result are sorted in descending order according to their corresponding similarity. Each three-dimensional shape descriptor in the feature library is uniquely associated with a three-dimensional CAD model. The feature library is constructed according to the following steps: Obtain the 3D CAD models generated from industrial part modeling, convert all the 3D CAD models into STL format, and build a model database based on the converted 3D CAD models. The 3D CAD models in the model database are categorized according to functional similarity and appearance similarity. By setting up a virtual camera to surround the model, the various 3D CAD models in the model database are rendered from multiple angles to obtain their respective multi-view 2D images. All the multi-view 2D images are input into the trained progressive learning network, which outputs the corresponding 3D shape descriptors, stores all the 3D shape descriptors, and constructs the feature library. Each 3D CAD model in the model database is associated with a category label, and the progressive learning network is trained according to the following steps: The multi-view two-dimensional images corresponding to all three-dimensional CAD models with the same category identifier are taken as an image set, and the image set is divided to obtain a training set and a test set; Based on a preset joint classification loss function, the initial network is trained by combining all the training sets and the test sets to obtain the progressive learning network. The joint classification loss function consists of a hinge loss function based on a 3D shape descriptor, a cross-entropy loss function based on view features, and a contrast loss function.

2. The 3D CAD model retrieval method based on progressive learning according to claim 1, characterized in that, The multi-view 2D image includes multiple view images. The features of the multi-view 2D image are input into a trained progressive learning network to obtain a query descriptor corresponding to the 3D CAD model to be retrieved, including: Each of the aforementioned view images is input into the shared-weight convolutional neural network for feature extraction, and the corresponding first intermediate features are output. All the first intermediate features are input into the global feature generation module, and the second intermediate features that can characterize the global features corresponding to the 3D CAD model to be retrieved are output. The second intermediate feature is input into the view attention feature generation module to perform feature enhancement processing on each of the first intermediate features to obtain the corresponding third intermediate features; By concatenating the second intermediate feature with all of the third intermediate features, a query descriptor is obtained; The query descriptor is input into the classification output header, the corresponding target category identifier is output, and the target category identifier is associated with the query descriptor.

3. The 3D CAD model retrieval method based on progressive learning according to claim 2, characterized in that, The shared-weight convolutional neural network includes an input module, a first residual stage, a second residual stage, a third residual stage, and a fourth residual stage. The input module sequentially includes a convolutional layer with a kernel size of 7x7 and a stride of 2, a batch normalization layer, a ReLU activation layer, and a max-pooling layer with a stride of 2. The first residual stage is based on stacked first residual blocks. Each first residual block includes three sequentially connected convolutional layers, each followed by a batch normalization layer and a ReLU activation layer. The sizes of the three convolutional layers are 1x1, 3x3, and 1x1, respectively. The input of the first residual block has a shortcut connection to the ReLU activation layer corresponding to the last convolutional layer in the first residual block. The mechanism is as follows: the second residual stage is based on the stacking of four second residual blocks. The first second residual block in the second residual stage can halve the size of the input feature map and double the number of channels. The remaining second residual blocks have the same structure as the first residual block. The third residual stage is based on the stacking of six third residual blocks. The first third residual block in the third residual stage can halve the size of the input feature map and double the number of channels. The remaining third residual blocks have the same structure as the first residual block. The fourth residual stage is based on the stacking of three fourth residual blocks. The first fourth residual block in the fourth residual stage can halve the size of the input feature map and double the number of channels. The remaining fourth residual blocks have the same structure as the first residual block.

4. The 3D CAD model retrieval method based on progressive learning according to claim 2, characterized in that, All the first intermediate features are input into the global feature generation module, which outputs second intermediate features that can characterize the global features corresponding to the 3D CAD model to be retrieved, including: The reference feature vector is obtained by concatenating all the first intermediate features; The reference feature vector is input into a multilayer perceptron (MLP) for nonlinear transformation and information integration processing to obtain the second intermediate feature.

5. The 3D CAD model retrieval method based on progressive learning according to claim 2, characterized in that, The second intermediate feature is input into the view attention feature generation module to perform feature enhancement processing on each of the first intermediate features, thereby obtaining the corresponding third intermediate features, including: The second intermediate feature is subtracted from each of the first intermediate features at the channel level to obtain multiple fourth intermediate features; All of the fourth intermediate features are input into the Sigmoid activation function, and the attention weight vector is output. The attention weight vector includes multiple weight values, and any one of the weight values ​​is used to indicate the difference in attention weight values ​​between the view image with the same dimension as the corresponding first intermediate feature and the 3D CAD model to be retrieved. Based on each of the weight values, attention weighting is applied to each of the corresponding first intermediate features to obtain each of the fifth intermediate features; Each of the first intermediate features is added to the corresponding fifth intermediate feature to obtain the corresponding third intermediate features.

6. The 3D CAD model retrieval method based on progressive learning according to claim 1, characterized in that, Based on a preset joint classification loss function, an initial network is trained using all of the training and test sets to obtain the progressive learning network, which includes: The parameters of the convolutional neural network with shared weights in the initial network are initialized using the pre-trained ResNet50 network weights; The training set is input into the initial network after parameter initialization, and the predicted class label is output. The loss value is calculated based on the predicted class label, the corresponding true class label and the joint classification loss function. When the loss value is greater than or equal to a preset second threshold, the network parameters of the initial network are adjusted using the ADAM optimizer. The training set is then input into the network after the model parameters are adjusted for retraining. The number of iterations is recorded until the loss value corresponding to any iteration is less than the second threshold. The network corresponding to the loss value less than the second threshold is determined as an intermediate network. Alternatively, if the current number of iterations reaches a preset maximum number of training cycles, the network corresponding to the current number of iterations is determined as an intermediate network. Based on the test set, the intermediate network is analyzed for accuracy, recall, and comprehensive indicators. The intermediate network that meets the preset conditions is identified as the progressive learning network.

7. A 3D CAD model retrieval device based on progressive learning, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the progressive learning-based 3D CAD model retrieval method as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the progressive learning-based 3D CAD model retrieval method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Methods and systems for detection in industrial internet of things data collection environment with large data sets

    CN110073301A

  • Multi-view three-dimensional model retrieval method and system based on pairing depth feature learning

    CN111382300A