A method and system for retrieving 3D models from sketches based on cross-modal class center alignment
By obtaining the class center of the 3D model through the teacher model and combining it with the line perception enhancement and distillation learning of the student model, efficient alignment of sketch features and 3D model is achieved, which solves the problem of insufficient retrieval accuracy in cross-modal understanding and improves retrieval accuracy and efficiency.
Patent Information
- Application Number
- CN202511120087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing technologies struggle to effectively understand the multi-view geometric information of sketches and 3D models across modalities, resulting in insufficient retrieval accuracy and failing to meet practical needs.
By using the teacher model to process multi-view images of 3D models to obtain class centers, and combining the line perception enhancement preprocessing and distillation learning of the student model, sketch features are aligned to the class centers of similar 3D models, and cosine similarity is used for retrieval.
It improves the accuracy and efficiency of retrieval from sketches to 3D models, meeting the practical application needs of scenarios such as industrial design.
Smart Images

Figure CN120632135B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sketch-based 3D model retrieval technology, and in particular to a sketch-based 3D model retrieval method and system based on cross-modal class center alignment. Background Technology
[0002] Sketch-based 3D model retrieval achieves precise mapping from abstract lines to concrete models through cross-modal matching of hand-drawn sketches and 3D model features. It is a crucial task in the intersection of computer vision and graphics, widely applied in industrial design, virtual simulation, and other scenarios. With increasing demands for design efficiency, enabling machines to understand the sparse semantics of sketches and efficiently align them with the multi-view geometric information of 3D models has become a core direction for technological breakthroughs.
[0003] Currently, the practicality of sketch-based 3D model retrieval is still limited by the technical bottleneck of cross-modal understanding. Traditional methods often rely on manually designed features (such as edge direction histograms), which are difficult to adapt to the abstractness and diversity of sketch lines, and often neglect the overall structure due to excessive focus on local strokes. In 3D model processing, feature fusion of multi-view images often adopts simple stitching or averaging strategies, resulting in the loss of geometric correlation between viewpoints and failing to fully utilize the spatial information of the model.
[0004] A deeper problem lies in bridging modal differences: sketches convey semantics using black lines on a two-dimensional plane, while three-dimensional models present a three-dimensional structure through multi-view textures, creating a natural gap in feature distribution between the two. Existing cross-modal alignment methods mostly rely on global semantic matching, lacking fine-grained class center guidance mechanisms. This results in poor clustering of similar features in the common space and insufficient separation of dissimilar features, ultimately leading to similar models ranking poorly and accuracy failing to meet practical needs during retrieval. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes a method and system for retrieving 3D models from sketches based on cross-modal class center alignment. The method involves acquiring multi-view images of 3D models by constructing a stable lighting scene, using a teacher model to train class centers to build a common feature space, applying line-aware enhancement preprocessing to the sketches, and then inputting them into a student model. During distillation learning, a cross-modal class center alignment loss function is used to align sketch features with the class centers of similar 3D models. Finally, cosine similarity is used to achieve sketch-to-3D model retrieval, effectively improving retrieval accuracy.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for retrieving 3D models from sketches based on cross-modal class center alignment, comprising:
[0008] Multi-view images of the 3D model are input into the pre-trained teacher model to extract and classify the features of the 3D model. The classifier weight matrix is used as the class center of each category of the 3D model.
[0009] Identify high-density line regions in the sketch to be retrieved; through a selective erasure mechanism, preserve the core outline and generate sketch data with enhanced line features;
[0010] The enhanced sketch data is input into the pre-trained student model, and cross-modal aligned sketch features are obtained based on deep semantic feature capture and global feature filtering.
[0011] The similarity between the sketch features and the 3D model features is calculated, and the sketches are sorted in descending order of similarity to achieve the retrieval of the 3D model.
[0012] The training process of the student model includes:
[0013] The enhanced sketch samples are input into the feature extractor of the weight-initialized student model to extract the sketch sample features;
[0014] Based on distillation learning, the class centers and sketch sample features are input into the classifier of the student model. Based on minimizing the loss function, the sketch sample features are aligned with the class centers of the same type of 3D model.
[0015] Secondly, the present invention provides a sketch retrieval 3D model system based on cross-modal class center alignment, comprising:
[0016] The 3D model classification module is configured to input multi-view images of the 3D model into a pre-trained teacher model, extract 3D model features and classify them, and use the classifier weight matrix as the class center of each category of the 3D model.
[0017] The feature enhancement module is configured to identify high-density line regions in the sketch to be retrieved; and to generate sketch data with enhanced line features by retaining the core outline through a selective erasure mechanism.
[0018] The cross-modal feature alignment module is configured to input the feature-enhanced sketch data into a pre-trained student model, and obtain cross-modal aligned sketch features based on deep semantic feature capture and global feature filtering; wherein, the training process of the student model includes: inputting the feature-enhanced sketch samples into the feature extractor of the student model with weights initialized, and extracting sketch sample features; based on distillation learning, inputting the class centers and sketch sample features into the classifier of the student model, and aligning the sketch sample features to the class centers of the same type of 3D model based on minimizing the loss function;
[0019] The retrieval and matching module is configured to calculate the similarity between the sketch features and the 3D model features, and sort them in descending order of similarity to achieve the retrieval of sketches to 3D models.
[0020] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the sketch retrieval 3D model method based on cross-modal class center alignment described in the first aspect.
[0021] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the sketch retrieval 3D model method based on cross-modal class center alignment described in the first aspect.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0023] This invention obtains class centers by processing multi-view images of 3D models using a teacher model, providing a stable benchmark for cross-modal alignment and solving the problem of insufficient multi-view feature fusion. High-density line detection and selective erasure are applied to the sketches to strengthen core contours and prevent the model from focusing on redundant features. The student model undergoes distillation learning to align sketch features towards class centers, reducing modal differences. Finally, retrieval is achieved through similarity calculation. The overall process forms a closed loop from 3D model classification and feature extraction to cross-modal alignment, preserving multi-view information of the 3D model while enhancing the expression of core sketch features, effectively improving retrieval accuracy and efficiency, and meeting the needs of precise matching in practical applications.
[0024] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0025] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.
[0026] Figure 1 This is a flowchart illustrating the main process of a sketch retrieval method for 3D models based on cross-modal class center alignment, as provided in an embodiment of the present invention.
[0027] Figure 2 This is a flowchart illustrating a method for retrieving 3D models from sketches based on cross-modal class center alignment, provided in an embodiment of the present invention. Detailed Implementation
[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] Example 1
[0030] like Figure 1 As shown, this embodiment discloses a method for retrieving 3D models from sketches based on cross-modal class center alignment, including the following steps:
[0031] S1: Input the multi-view images of the 3D model into the pre-trained teacher model, extract the features of the 3D model and classify them, and use the classifier weight matrix as the class center of each category of the 3D model.
[0032] S2: Identify high-density line regions in the sketch to be retrieved; through a selective erasure mechanism, preserve the core outline and generate sketch data with enhanced line features;
[0033] S3: Input the feature-enhanced sketch data into the pre-trained student model, and obtain cross-modal aligned sketch features based on deep semantic feature capture and global feature filtering;
[0034] S4: Calculate the similarity between the sketch features and the 3D model features, and sort them in descending order of similarity to achieve the retrieval from sketch to 3D model.
[0035] Next, combined Figure 2 This embodiment provides a detailed description of a sketch retrieval method for 3D models based on cross-modal class center alignment.
[0036] S1. First, build a stable, lit 3D scene and obtain the complete visual representation of the 3D model, covering the model features of different categories (such as typical categories like chairs).
[0037] Specifically, a blank 3D scene environment is built using Blender software. Considering the need to maintain stable lighting conditions after the model is imported (to prevent sudden changes in light from affecting subsequent analysis), this embodiment designs an eight-point light source lighting system: with the center of the scene as the reference, point light sources are evenly arranged at each corner of the cube, and smooth light and shadow gradients are formed through symmetrical lighting to ensure that the model presents a consistent visual appearance from different perspectives, laying a stable foundation for multi-view feature acquisition.
[0038] The target model is placed at the origin of the scene coordinate system, and a circular shooting array consisting of 12 virtual cameras is configured. Each camera is aimed at the center of the model at a 60° elevation angle and evenly distributed at 30° intervals along the horizontal direction, thereby acquiring 12 rendered images of the model from different angles. These multi-view images cover the model's omnidirectional visual information, jointly constructing a complete visual expression of the 3D model and providing multi-dimensional input for subsequent cross-modal feature matching.
[0039] Furthermore, the features of the 3D model are extracted and the 3D model is classified.
[0040] In this embodiment, the 3D model classifier is used as a teacher network to guide the classification of sketches.
[0041] First, 12 3D model images belonging to a single sample are input into the teacher model, which is built based on the ViT-L / 16 model. Checkpointing is used to selectively save and reconstruct intermediate calculation results, optimizing memory usage and thus overcoming hardware limitations to increase the training batch size. Simultaneously, the loss function used for classification is AmSoftMaxLoss, with the following formula:
[0042] ;
[0043] In the formula, N represents the number of samples in the batch; s is a scaling factor parameter, which is usually set to a value of not less than 30 to control the hardness of the classification boundary; The feature vector representing the i-th sample and its true class The angle between the corresponding weight vectors; m is the additive margin hyperparameter, which is generally between 0.35 and 0.5 and is used to enhance inter-class separability; This represents the cosine similarity between the i-th sample and the j-th class weight vector. This loss function improves the discriminative power of the features by introducing a margin term m into the cosine similarity of the target classes, making samples of the same class more compact in the feature space and samples of different classes more separated.
[0044] After classification, the classifier's weight matrix serves as the class center for each class in the 3D model. Simultaneously, each class center is surrounded by model features identical to those of its class, thus forming a common feature space. This common feature space is used to construct a unified feature metric for the 3D model and the sketch. Subsequently, cross-modal feature projection and similarity calculation are used to achieve retrieval and matching from the sketch to the 3D model.
[0045] S2, Obtain the sketch to be searched and preprocess it.
[0046] Inspired by the concept of MAE (Masked AutoEncoders), this embodiment proposes a line-aware intelligent sketch occlusion enhancement strategy. Its core technology lies in using dynamic detection and selective erasure mechanisms to erase lines in high-density areas while preserving core contour features, and simulating line-missing scenarios to achieve targeted data enhancement of sketch line features. Specifically:
[0047] S201. Line Density Detection
[0048] The Canny edge detection algorithm is used to calculate the line density of the input sketch in real time. When the number of detected lines exceeds a preset threshold, an erasure operation is triggered to ensure the accuracy of data augmentation.
[0049] Overly dense lines in a sketch may contain redundant details or complex strokes (such as repeated outlines or noise interference). However, these dense lines can cause the model to learn redundant features instead of focusing on the overall structure. Therefore, to enable the model to focus on the core contours, this embodiment incorporates line density detection.
[0050] If the line density exceeds a preset threshold, it indicates that there are high-density lines in the current area. In this embodiment, high-density lines are defined as redundant lines that are not part of the core contour, such as repeated texture lines in mechanical sketches or auxiliary lines in hand-drawn sketches. When their density exceeds the threshold, they are determined to be eraseable details.
[0051] It is important to note that this embodiment only applies to lines in high-density areas, not to random erasure. For example, in a car sketch, the dense lines of the tire tread pattern may be erased, while the body outline, due to its low density, is retained, thus preserving the core outline.
[0052] S202. Selective erasure mechanism
[0053] Within a randomly generated rectangular occlusion area, the black lines (RGB values < 128) are precisely identified and erased only by comparing the three-channel pixel values, thus preserving the integrity of the background area.
[0054] Among these measures, compliance verification of the rectangular occlusion area ensures the legitimacy of the occlusion operation.
[0055] (1) Regional validity verification
[0056] If the randomly generated rectangular area overlaps with the original canvas by more than 50% (e.g., obscuring most of the image), the area is regenerated to avoid excessive occlusion that could lead to feature loss.
[0057] (2) Line Existence Detection
[0058] If the number of lines detected within a rectangular area is 0 (i.e., it's all background), then that area is discarded to prevent invalid erasure (e.g., erasing a blank background is meaningless). It should be understood that those skilled in the art can achieve line detection through methods such as edge detection.
[0059] Furthermore, for compliant rectangular occlusion areas, the erasure area is determined by comparing three-channel pixel values based on RGB thresholds. While erasing redundant lines, the process simulates the situation in real-world sketches where strokes are occluded or worn, such as paper creases covering lines.
[0060] The erased area needs to be verified for compliance. The erased area must contain valid lines (number of lines > 0). The erased area should not exceed the preset maximum scale of the total area of the sketch to avoid damaging core features.
[0061] The specific method for calculating the three-channel pixel value is as follows: for each pixel within the rectangular area, compare the RGB channel value with the value of 128.
[0062] If all RGB values are <128, the line is identified as black (foreground) and is erased (set to white or background color).
[0063] If the RGB value is ≥128, it is determined to be the background and the original color is retained.
[0064] Sketches are typically drawn with black lines on a white background. Erasing non-lined areas will destroy the integrity of the background and cause the model to mislearn background noise.
[0065] This embodiment simulates a real-world scenario of missing strokes by precisely erasing lines rather than random areas, forcing the model to learn more robust feature representations. Intelligent triggering based on line density ensures targeted enhancement and avoids ineffective interference. Selective erasure produces "partially visible" samples, preventing overfitting and promoting global feature learning. Random parameter combinations expand the data representation space. This mechanism aligns with the human cognitive pattern of learning complete concepts from incomplete samples. This intelligent erasure mechanism, based on line density threshold control, retains the computational efficiency advantages of traditional random erasure while introducing line density detection, making the data augmentation process more consistent with the characteristics of sketch data.
[0066] As one implementation method, in order to enable the erasure operation to better adapt to sketch data of different sizes and complexities and flexibly respond to diverse sketch features, an adaptive parameter adjustment strategy is introduced, including dynamically scaling the erase area ratio (default 0.02-0.33) and intelligently matching the aspect ratio (default 0.3-3.3).
[0067] (1) Dynamically scale the erase area: The erase area is automatically adjusted according to the sketch resolution. For example, for a sketch of 224×224 pixels, the minimum erase area is 224×224×0.02=1003 pixels, and the maximum is 224×224×0.33=16748 pixels;
[0068] (2) Intelligent matching aspect ratio: Automatically generates a rectangle with a similar aspect ratio based on the shape of the main body of the sketch. For example, for a vertically elongated sketch (width-to-height ratio 0.5), it generates an occlusion area with a width-to-height ratio of 0.3-0.7 to avoid horizontally wide areas from damaging the main structure.
[0069] In terms of proportion control, the total area of the input image is first calculated, and then the target erasure area is randomly generated within a preset proportion range to ensure that the proportion of the erasure area to the original image area is within a reasonable range. For shape control, a geometric calculation method based on area constraints is adopted: a aspect ratio is first randomly generated, and then the height and width of the actual erasure area are derived through mathematical formulas. Specifically, the height value is equal to the square root of the target erasure area divided by the aspect ratio, and the width value is equal to the square root of the target erasure area multiplied by the aspect ratio. This calculation method ensures that the area of the erasure region always meets the preset requirements regardless of changes in the aspect ratio. Finally, the system checks whether the calculated erasure region size exceeds the image boundary to ensure the effectiveness and safety of the erasure operation. The entire process, through strict mathematical constraints and a random sampling mechanism, achieves intelligent adaptation of the erasure region in terms of size and shape, ensuring effective enhancement effects on sketches of different sizes and complexities.
[0070] After data augmentation with random occlusion, random rotations within ±15 degrees are performed to enhance the model's robustness to orientation changes. Random color dithering is then used to alter the image's brightness, contrast, and saturation to improve color generalization. Finally, a horizontal flip is performed with a 50% probability to increase horizontal symmetry recognition. The processed image is then converted to PyTorch tensor format and normalized.
[0071] This embodiment uses Canny edge detection to identify high-density line regions, selectively erasing black lines while preserving the background through rectangular occlusion. This enhancement method accurately simulates real-world scenarios with missing strokes, forcing the model to learn robust feature representations and avoiding over-focusing on local textures. Density-based intelligent triggering improves the enhancement's targeting and reduces invalid interference. The "partially visible" samples formed by selective erasure prevent overfitting and promote global feature learning. Random parameter combinations expand the data representation space, enabling the model to adapt to sketches of varying complexity. This mechanism aligns with the human tendency to perceive complete concepts from incomplete samples, effectively improving the generalization ability and matching accuracy of sketch features for 3D model retrieval.
[0072] S3. Input the preprocessed sketch to be retrieved into the student model and extract sketch features.
[0073] To standardize feature input, the enhanced sketch samples and the model domain view are forcibly resized to 224×224 pixels to match the pre-training input requirements of the student model. This ensures input consistency in the feature extraction process and lays the foundation for subsequent cross-domain feature alignment and semantic matching. The student model is also built based on the ViT-L / 16 model.
[0074] The enhanced sketch data is input into the student model, and a visual Transformer encoder extracts features through multi-layer attention interactions. During feature extraction, the model freezes the first N attention layers and fully connected layers; that is, the parameters of the frozen bottom layers do not participate in the update. The model mainly relies on the open feature normalization layer to fine-tune and adapt the model based on the sketch data, ultimately outputting the sketch features. This allows the model to align domain differences (there are data distribution differences between sketches and natural images) by only fine-tuning this layer when adapting to sketch data. This design reuses the performance of the pre-trained model while significantly reducing computational complexity by freezing a large number of parameters.
[0075] However, considering the noise differences in outputs at different levels, in order to accurately capture purer and more focused global semantic deep features and reduce the interference of shallow attention noise on effective features, this embodiment connects a normalization layer at the end of the encoder. By utilizing its aggregation characteristics of high-level semantic features, it can capture deep semantic features purified by multiple layers of attention in a targeted manner.
[0076] Furthermore, to avoid feature confusion, the feature space formed by deep semantic features is explicitly labeled by using "deep" key-value pairs. This differs from the traditional method of directly extracting from the [CLS] token. The [CLS] token is prone to mixing with shallow attention noise (such as local texture interference), while the output of the final normalization layer focuses more on global semantics, thereby achieving noise filtering.
[0077] Next, dynamic feature compression is performed. The acquired deep semantic features still have spatial dimensional redundancy, such as the 196×1024 dimension of the Transformer output, which contains 196 spatial location codes. In order to remove redundant features at local locations, suppress the interference of redundant spatial attention on the model, and reduce the feature processing scale to improve computational efficiency, this embodiment implements feature filtering through the [:,0,:] global semantic vector slicing operation.
[0078] Here, the "0" index corresponds to the global context representation vector, and "[:,0,:]" means selecting all batches of patches with a spatial location index of 0, along with all feature dimensions of that patch. This discards features from all spatial location dimensions except for the one with index 0, retaining only the global context representation vector corresponding to index 0, thus eliminating redundant features at local locations. Through this operation, the feature dimension is compressed from 196×1024 to 1×1024, suppressing the interference of redundant spatial attention on the model, preventing the model from over-focusing on irrelevant local details, and significantly reducing the feature processing scale, directly improving subsequent computational efficiency.
[0079] The training process of the student model includes:
[0080] The enhanced sketch samples are input into the feature extractor of the weight-initialized student model to extract the sketch sample features;
[0081] Based on distillation learning, the class centers and sketch sample features are input into the classifier of the student model. Based on minimizing the loss function, the sketch sample features are aligned with the class centers of the same type of 3D model.
[0082] Specifically, for the feature extractor design of the student model, to reduce training costs, a transfer learning framework is constructed using a ViT-L / 16 pre-trained model with SWAG-based linearly fine-tuned weight initialization. The current implementation directly loads the complete pre-trained parameters for end-to-end feature extraction without implementing any network layer freezing strategy. The model obtains the global feature representation through the final normalization layer, specifically extracting the feature vector corresponding to the [CLS] label as the output. In this architecture, all pre-trained parameters (including the feature normalization layer) remain trainable, without employing a hierarchical parameter tuning mechanism, thus fully preserving the feature extraction capabilities of the pre-trained model.
[0083] SWAG (Stochastic Weight Averaging in Gaussian) is a model weighting technique that can improve model generalization ability. When using a pre-trained ViT-L / 16 model, weights that have been linearly fine-tuned using the SWAG method are loaded into the model as initial weights. With SWAG-tuned weights, the model can have a better parameter distribution during initialization. Compared to random initialization or ordinary pre-trained weights, this is more conducive to subsequent transfer learning on sketch data, speeds up model convergence, improves the model's performance in sketch feature extraction, enhances model generalization ability, and reduces the risk of getting trapped in poor local optima during training.
[0084] For the knowledge distillation process, due to the severe class imbalance problem in the SHREC13 dataset, bidirectional data distillation can lead to problems such as misalignment of batch sampling between the two models, missing gradient signals of minority class samples, and noise-dominated gradients of majority class. Therefore, this embodiment adopts a phased training strategy to replace the synchronous bidirectional distillation method.
[0085] That is, in the teacher stage, the weight matrix of each class center of the 3D model is obtained through the classifier. Then, in the student stage, the weight matrix is introduced into the training and used as the target of sketch classification, thereby achieving the process of the 3D model guiding the sketch.
[0086] This part is no longer a simple classification, but rather it guides the features of the sketch to approach the class center of similar 3D models. The loss function TransferLoss, which achieves this function, is as follows:
[0087] ;
[0088] In the formula, N represents the number of samples in the batch, C represents the total number of categories, and It is an adjustable hyperparameter (default value is 0.5) used to balance the strength of intra-class aggregation and inter-class separation. The larger the value, the more the model focuses on amplifying the differences between different categories. For each sample... , It is the extracted d-dimensional feature vector. This corresponds to the category label (integer index). The center vector of all categories is stored in a matrix. Among them Let d represent the central feature of the j-th class. During the calculation process... This represents the Euclidean distance (L2 norm), while The operator is used for vector dot product.
[0089] In addition, the differential classification of the sketch features themselves is also considered, so the cross-entropy function is also introduced:
[0090] ;
[0091] In the formula, N represents the number of samples in the current batch, used to calculate the average loss; C represents the total number of categories in the classification task. is the true label of the i-th sample in category c, using one-hot encoding (i.e., 1 if it belongs to this category, and 0 otherwise); This represents the probability that the model predicts the i-th sample belongs to class c. It is normalized using the softmax function to ensure that the sum of the predicted probabilities for all classes is 1. The loss function calculates the negative log-likelihood of the true label and the predicted probability, sums these probabilities over all samples and classes, and then averages the results. The smaller the final value, the closer the model's prediction is to the true distribution.
[0092] The final loss function formula is as follows, where Hyperparameters to be set:
[0093] ;
[0094] S4, Retrieval Phase. Using cosine similarity as the matching metric, the sketch-based 3D model retrieval process is as follows:
[0095] First, the sketch features and 3D model features are obtained. Then, the cosine similarity value S between the feature vectors of the two is calculated. The models are sorted in descending order according to the similarity score S, where the model with a higher S value indicates a stronger similarity to the input sketch and ranks higher accordingly; conversely, the model with a lower S value indicates a lower similarity to the sketch and ranks lower in the search results.
[0096] Table 1 illustrates the significant advantages of the method in this embodiment on the SHREC 2013 dataset. Looking at the model evolution trend: early methods (CMDR, SBR-VC) had mAP of only 25.0 and 11.6 respectively; AlexNet-based methods (such as TCL) improved to 80.7; the ResNet50 method (DCA) reached 81.3; and the ViT-B series methods (VTS, DPSML, CGN, CFTTSL) further improved to 87.9-89.7. This embodiment, employing the Vit-L architecture, achieves a decisive lead across all metrics: NN (89.4), FT (90.9), ST (94.2), EE (43.5), DCG (93.9), and mAP (91.9). Specifically, mAP is 2.2% higher than the best ViT-B method (CFTTSL) and 10.6% higher than the best CNN method (DCA), with an innovative EE score of 43.5. This result not only verifies the superiority of larger-scale ViT models, but also demonstrates the breakthrough progress of this embodiment in model design and training strategies. In particular, while maintaining the efficient feature extraction capability of the ViT architecture, it significantly improves fine-grained retrieval performance (EE index exceeds 43), showcasing the current state-of-the-art retrieval performance level.
[0097] Table 1 Performance Comparison Table;
[0098]
[0099] In this specific embodiment, sketch feature processing employs line density detection and selective erasure to accurately preserve the core contours, avoiding the semantic structure loss caused by excessive focus on local strokes in traditional methods. This allows the model to capture more representative abstract line features. The first model processes multi-view images of the 3D model, constructing class centers using a classifier weight matrix. This overcomes the limitations of insufficient geometric information utilization caused by simple splicing or averaging of multi-view features, providing a stable semantic benchmark for cross-modal alignment. The second model combines distillation learning and class center guidance to achieve deep adaptation between sketch features and 3D model class centers, alleviating the feature matching challenges caused by modal differences. Retrieval is based on cosine similarity matching, improving accuracy while maintaining efficiency, making sketch-to-3D model retrieval more accurate and efficient, effectively supporting practical application needs in scenarios such as industrial design.
[0100] Example 2
[0101] This embodiment provides a sketch retrieval 3D model system based on cross-modal class center alignment, including:
[0102] The 3D model classification module is configured to input multi-view images of the 3D model into a pre-trained teacher model, extract 3D model features and classify them, and use the classifier weight matrix as the class center of each category of the 3D model.
[0103] The feature enhancement module is configured to identify high-density line regions in the sketch to be retrieved; and to generate sketch data with enhanced line features by retaining the core outline through a selective erasure mechanism.
[0104] The cross-modal feature alignment module is configured to input the feature-enhanced sketch data into a pre-trained student model, and obtain cross-modal aligned sketch features based on deep semantic feature capture and global feature filtering; wherein, the training process of the student model includes: inputting the feature-enhanced sketch samples into the feature extractor of the student model with weights initialized, and extracting sketch sample features; based on distillation learning, inputting the class centers and sketch sample features into the classifier of the student model, and aligning the sketch sample features to the class centers of the same type of 3D model based on minimizing the loss function;
[0105] The retrieval and matching module is configured to calculate the similarity between the sketch features and the 3D model features, and sort them in descending order of similarity to achieve the retrieval of sketches to 3D models.
[0106] Example 3
[0107] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the sketch retrieval 3D model method based on cross-modal class center alignment as described in Embodiment 1 above.
[0108] Example 4
[0109] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the sketch retrieval 3D model method based on cross-modal class center alignment as described in Embodiment 1 above.
[0110] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0111] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for retrieving 3D models from sketches based on cross-modal class center alignment, characterized in that, include: Multi-view images of the 3D model are input into the pre-trained teacher model to extract and classify the features of the 3D model. The classifier weight matrix is used as the class center of each category of the 3D model. Identify high-density line regions in the sketch to be retrieved; Based on line density detection, high-density line areas in the sketch to be retrieved are identified; through a selective erasure mechanism, the black line areas are accurately identified and only erased, while keeping the background area intact. Preserving the core outline, sketch data with enhanced line features is generated, specifically including: For areas with high-density lines, randomly generate rectangular occlusion areas; Perform compliance checks on rectangular occlusion areas: If the randomly generated rectangular occlusion area overlaps with the sketch border by more than a preset value, then regenerate the rectangular occlusion area; if the number of lines in the rectangular occlusion area is 0, then discard the area. For compliant rectangular occlusion areas, the erasure area is determined by comparing the three-channel pixel values based on the RGB threshold; compliance verification is performed on the erasure area: the erasure area contains lines, and the erasure area does not exceed the preset maximum proportion of the total area of the sketch. Erase redundant lines in compliant erasure areas, retaining the core outline; The enhanced sketch data is input into the pre-trained student model, and cross-modal aligned sketch features are obtained based on deep semantic feature capture and global feature filtering. The similarity between the sketch features and the 3D model features is calculated, and the sketches are sorted in descending order of similarity to achieve the retrieval of the 3D model. The training process of the student model includes: The enhanced sketch samples are input into the feature extractor of the weight-initialized student model to extract the sketch sample features; Based on distillation learning, the class centers and sketch sample features are input into the classifier of the student model. Based on minimizing the loss function, the sketch sample features are aligned with the class centers of the same type of 3D model. The loss function TransferLoss is: ; In the formula, N represents the number of samples in the batch, and C represents the total number of categories. It is an adjustable hyperparameter used to balance the strength of intra-class aggregation and inter-class separation; for each sample , It is the extracted d-dimensional feature vector. These are the corresponding category labels; the center vectors of all categories are stored in a matrix. middle, and Representing categories and the d-dimensional central features of the j-th class; during the calculation process Represents Euclidean distance. Used for vector dot product.
2. The sketch retrieval method for 3D models based on cross-modal class center alignment as described in claim 1, characterized in that, The teacher model includes a feature extractor built on a Transformer encoder for extracting features from the 3D model; It also includes a classifier for classifying 3D models.
3. The sketch retrieval method for 3D models based on cross-modal class center alignment as described in claim 1, characterized in that, The Canny edge detection algorithm is used to calculate the number of lines in the sketch; when the number of lines exceeds a preset threshold, it is determined to be a high-density line area.
4. The sketch retrieval method for 3D models based on cross-modal class center alignment as described in claim 1, characterized in that, The process of inputting the feature-enhanced sketch data into a pre-trained student model, and obtaining cross-modal aligned sketch features based on deep semantic feature capture and global feature filtering, specifically includes: The student model includes a feature extractor and a classifier built on a Transformer encoder; a normalization layer is added to the end of the Transformer encoder to capture deep semantic features; Based on the global semantic vector slicing operation of [:,0,:], global feature filtering is performed on deep semantic features; [:,0,:] means retaining the global context representation vector corresponding to index 0 and discarding features of other spatial location dimensions except for index 0.
5. The sketch retrieval method for 3D models based on cross-modal class center alignment as described in claim 1, characterized in that, The enhanced sketch data is input into the pre-trained student model. At this time, the feature extractor of the student model freezes the first N attention layers and fully connected layers, and opens the feature normalization layer for fine-tuning; where N≥1.
6. The sketch retrieval method for 3D models based on cross-modal class center alignment as described in claim 1, characterized in that, The method of minimizing the loss function to align the features of sketch samples to the class center of similar 3D models specifically includes: ; ; ; in, For the total loss function, For hyperparameters; For cross-modal class center alignment loss function, The cross-entropy loss function is used; N represents the number of samples in the batch, and C represents the total number of classes. It is an adjustable hyperparameter used to balance the strength of intra-class aggregation and inter-class separation; for each sample , It is the extracted d-dimensional feature vector. These are the corresponding category labels; the center vectors of all categories are stored in a matrix. Among them and Representing categories and the d-dimensional central features of the j-th class; during the calculation process Represents the Euclidean distance, and T is used for the vector dot product; is the true label of the i-th sample in category C, using one-hot encoding; This represents the probability that the model predicts the i-th sample belongs to class c.
7. A sketch retrieval system for 3D models based on cross-modal class center alignment, employing the sketch retrieval method for 3D models as described in any one of claims 1-6, characterized in that, include: The 3D model classification module is configured to input multi-view images of the 3D model into a pre-trained teacher model, extract 3D model features and classify them, and use the classifier weight matrix as the class center of each category of the 3D model. The feature enhancement module is configured to identify high-density line regions in the sketch to be retrieved; and to generate sketch data with enhanced line features by retaining the core outline through a selective erasure mechanism. The cross-modal feature alignment module is configured to input the feature-enhanced sketch data into a pre-trained student model, and obtain cross-modal aligned sketch features based on deep semantic feature capture and global feature filtering; wherein, the training process of the student model includes: inputting the feature-enhanced sketch samples into the feature extractor of the student model with weights initialized, and extracting sketch sample features; based on distillation learning, inputting the class centers and sketch sample features into the classifier of the student model, and aligning the sketch sample features to the class centers of the same type of 3D model based on minimizing the loss function; The retrieval and matching module is configured to calculate the similarity between the sketch features and the 3D model features, and sort them in descending order of similarity to achieve the retrieval of sketches to 3D models.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the sketch retrieval 3D model method based on cross-modal class center alignment as described in any one of claims 1-6.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the sketch retrieval 3D model method based on cross-modal class center alignment as described in any one of claims 1-6.
Citation Information
Patent Citations
Deep learning-based freehand sketch image retrieval method
CN106126581A
Cross-modal three-dimensional model retrieval method based on noise data cleaning
CN115080778A