Zero-sample sketch retrieval method and system based on multi-modal semantic information interaction
By adopting multimodal semantic information interaction method in sketch retrieval, text semantic embedding and large language model are used to fusion of cross-modal information, and training through semantic self-distillation loss function, the problem of excessive difference between modal and knowledge in zero-sample sketch retrieval is solved, and the retrieval performance is improved.
Patent Information
- Application Number
- CN202510268986.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing sketch search method has poor retrieval performance due to excessive differences in modality and knowledge under zero sample situation.
Using a method based on multimodal semantic information interaction, a neural network is constructed by constructing a feature extraction of teachers and students, and a text semantic information of the training set is expanded by text semantic embedding module and large language model, cross-modal information fusion is performed, multimodal semantic information interaction features are obtained, and training is carried out through semantic self-distillation loss function.
It effectively alleviates the problem of excessive difference between modal and knowledge in zero-sample sketch retrieval, improves retrieval performance, reduces the modal gap, and realizes knowledge transfer learning.
Smart Images

Figure CN120196780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sketch retrieval method in the field of image retrieval, and in particular to a zero-shot sketch retrieval method based on multi-modal semantic information interaction, and also to a zero-shot sketch retrieval system based on multi-modal semantic information interaction. Background Art
[0002] Both hand-drawn sketches and natural images are carriers of information. Sketches and natural images of the same category usually express the same semantic information. This meaning hidden in the external manifestation forms (such as colors, textures) is called semantics. Similarly, texts also contain rich semantic information. Semantics is a common concept running through sketches, real photos and texts, and is the key bridge connecting multi-modal information. Therefore, semantic-based information guidance is of great significance for improving the performance of the sketch retrieval task.
[0003] Existing sketch retrieval methods directly use text semantics as the anchor points for modality alignment, which has obvious limitations. Due to the large differences between the text and image modalities, this difference may be manifested as one modality containing information not provided by the other modality (for example, for the category "pigeon", the image may convey visual details such as the shape of the beak and the form of the wings, while the text may contain the abstract semantic "peace" represented by the pigeon). Therefore, direct alignment may be too strict, resulting in the loss of some discriminative information, and there is a problem of too large differences between modalities and knowledge in zero-shot sketch retrieval. Summary of the Invention
[0004] To solve the technical problem that existing sketch retrieval methods have too large differences between modalities and knowledge in zero-shot sketch retrieval, the present invention provides a zero-shot sketch retrieval method and system based on multi-modal semantic information interaction.
[0005] The present invention is implemented by the following technical solutions: A zero-shot sketch retrieval method based on multi-modal semantic information interaction, which includes the following steps:
[0006] S1: Construct a data set divided into a training set and a test set;
[0007] S2: Construct feature extraction neural networks for the teacher and the student;
[0008] S3: Expand the semantic information of the class name texts in the training set through a text semantic embedding module and a large language model to obtain text prompt features;
[0009] S4: First, input the images of the training set into the teacher feature extraction neural network. Then, through channel interaction and spatial interaction, enable cross-modal information fusion between the text prompt features and the image features of the training set to obtain multi-modal semantic information interaction features. Finally, train the student feature extraction neural network.
[0010] S5: Input the images to be retrieved in the test set into the student feature extraction neural network trained in step S4 to obtain the features of the images to be retrieved.
[0011] S6: First, input the query sketch image in the test set into the student feature extraction neural network trained in step S4 to obtain the features of the query sketch image. Then, compare and retrieve the query sketch image and the images to be retrieved based on the features of the query sketch image and the features of the images to be retrieved.
[0012] By effectively utilizing text semantic information, the present invention alleviates the problem of excessive modality and knowledge differences in zero-shot sketch retrieval, improves the retrieval performance, and the semantic self-distillation loss can complete knowledge transfer learning without additionally obtaining a pre-trained large model on a large-scale dataset as the teacher network. It uses a large language model to semantically expand the class name text, sends the text prompt features and image features into the text-image semantic information interaction module for sufficient semantic information interaction, effectively improves the feature extraction effect of the image feature extraction neural network, achieves the purpose of reducing the modality gap, and solves the technical problem of excessive modality and knowledge differences in existing sketch retrieval methods in zero-shot sketch retrieval.
[0013] As a further improvement of the above solution, in step S1, first obtain the labeled basic image data, and then preprocess the basic image data to construct the dataset; in step S2, the feature extraction neural networks of the teacher and the student adopt the Vision Transformer structure. The teacher feature extraction neural network embeds a semantic information interaction module in the benchmark network, and both feature extraction neural networks are initialized with the same pre-trained parameters; in step S6, calculate the distance between the features of the query sketch image and the features of the images to be retrieved to obtain a distance ranking as the retrieval result ranking, and perform zero-shot sketch retrieval based on multi-modal semantic information interaction.
[0014] As a further improvement of the above solution, in step S3, the expanded semantic information includes the following steps:
[0015] S31: Combine the class names of the training set with the question template to obtain question words embedded with class names.
[0016] S32: Input the question words into the large language question-answering model to obtain corresponding multiple related semantically rich description statements.
[0017] S33: Feed the described descriptive statement into the CLIP text encoder to obtain text features of multiple descriptions;
[0018] S34: Feed the text features into the adapter to obtain text prompt features adapted to the large language question - answering model.
[0019] As a further improvement of the above solution, the step S4 includes:
[0020] S41: Perform a chunking operation on the images in the training set to obtain multiple image chunks, and feed them into the first layer of the vision transformer to obtain corresponding multiple image chunk features;
[0021] S42: Feed the semantic prompt features and the features of the image chunks into the text - image semantic information interaction module to obtain image features of semantic prompts;
[0022] S43: Feed the image features of semantic prompts into the subsequent layers of the vision transformer to obtain the multi - modal semantic information interaction features;
[0023] S44: Train the student feature extraction neural network through the N - pair loss function;
[0024] S45: Train the student feature extraction neural network through the semantic self - distillation loss function.
[0025] Further, the step S42 includes:
[0026] S421: Average the features of the image chunks to obtain average image features;
[0027] S422: Horizontally concatenate the average image features and the semantic prompt features to obtain text - image concatenated features;
[0028] S423: Feed the text - image concatenated features into a multi - layer perceptron to obtain channel interaction features;
[0029] S424: Add the channel interaction features and the image chunk features point - by - point respectively to obtain channel semantic information interaction features;
[0030] S425: Vertically concatenate the channel semantic information interaction features and the semantic prompt features to obtain multi - modal information input features;
[0031] S426: Multiply the multi - modal information input features with multiple query weight matrices, key weight matrices, and value weight matrices to obtain multiple groups of query vectors, key vectors, and value vectors;
[0032] S427: Calculate attention weights through multiple sets of query vectors, key vectors, and value vectors;
[0033] S428: Multiply the attention weight and the value vector, perform horizontal splicing, and send them to the linear layer to obtain text semantic hint features.
[0034] Furthermore, the calculation formula of the attention weight is:
[0035]
[0036] Where h is the number of heads, a is the attention weight, q is the query vector, k is the key vector, and d k is the scaling factor, softmax(·) is the normalized exponential function, q h k h T To calculate q h and k h The dot product of .
[0037] Further, the step S44 includes:
[0038] S441: Obtain image data in the sketch image domain and the natural image domain respectively, and pair them up and send them into the feature extraction neural network;
[0039] S442: guiding the student feature extraction neural network to bring the sketch image and the natural image of the same category closer to each other and push different categories away from each other through the N-pair loss function, and projecting the images of the two domains into the same feature space;
[0040] Among them, the N-pair loss function is:
[0041]
[0042] Where B is the number of data sets in a training batch, d(·) is the cosine distance metric, and f i S is the i-th sketch feature, and They are The positive and negative photo features of , i,j∈{1,2,…,B}, margin is the edge hyperparameter.
[0043] Furthermore, the step S45 includes:
[0044] Based on the teacher feature extraction neural network and the student feature extraction neural network, through the self-distillation model, the student feature extraction neural network retains the features containing rich semantic interaction information learned in the teacher feature extraction neural network; wherein, the formula of the self-distillation model is:
[0045]
[0046] In the formula, is the feature obtained through the teacher feature extraction neural network, is the feature obtained through the student feature extraction neural network; KL(·) is the KL divergence, expressed as:
[0047]
[0048] In the formula, is the i-th feature obtained through the teacher feature extraction neural network, is the i-th feature obtained through the student feature extraction neural network.
[0049] As a further improvement of the above solution, the step S1 includes:
[0050] S11: Obtain the labeled basic image data, where the basic image data includes sketch images and natural images, compress and crop the basic image data, and convert the image format of the sketch images;
[0051] S12: Divide the data set into a training set and a test set with non-overlapping categories, and both the training set and the test set contain natural images and hand-drawn images.
[0052] The present invention also provides a zero-shot sketch retrieval system based on multi-modal semantic information interaction, which includes:
[0053] A data set division module, which is used to construct a data set divided into a training set and a test set;
[0054] An initialization module, which is used to construct the feature extraction neural networks of the teacher and the student;
[0055] An expansion module, which is used to expand the semantic information of the class name text of the training set through a text semantic embedding module and a large language model to obtain semantic prompt features;
[0056] A training module, which is used to first send the images of the training set into the feature extraction neural network, and then let the semantic prompt features and the image features of the training set perform cross-modal information fusion through channel interaction and spatial interaction to obtain multi-modal semantic information interaction features, and finally train the student feature extraction neural network;
[0057] A feature extraction module, which is used to input the image to be retrieved in the test set into the student feature extraction neural network trained by the training module to obtain the features of the image to be retrieved;
[0058] A sketch retrieval module, which is used to first input the query sketch image in the test set into the student feature extraction neural network trained by the training module to obtain the features of the query sketch image, and then compare and retrieve the query sketch image and the image to be retrieved according to the features of the query sketch image and the features of the image to be retrieved.
[0059] Compared with the existing sketch retrieval methods and systems, the zero-shot sketch retrieval method and system based on multi-modal semantic information interaction of the present invention have the following beneficial effects:
[0060] 1. The zero-shot sketch retrieval method based on multi-modal semantic information interaction effectively utilizes text semantic information to alleviate the problem of excessive modality and knowledge differences in zero-shot sketch retrieval, improve retrieval performance, and the semantic self-distillation loss can complete knowledge transfer learning without additionally obtaining a pre-trained large model as a teacher network on a large-scale dataset. It uses a large language model to semantically expand the class name text, and sends the text prompt features and image features into the text-image semantic information interaction module for sufficient semantic information interaction, effectively improving the feature extraction effect of the image feature extraction neural network, achieving the purpose of reducing the modality gap, and solving the technical problem of excessive modality and knowledge differences in zero-shot sketch retrieval existing in the existing sketch retrieval methods.
[0061] 2. The zero-shot sketch retrieval method based on multi-modal semantic information interaction designs a semantic self-distillation module to transfer the knowledge of multi-modal interaction to the target feature extractor. This sketch retrieval method uses two vision transformers with the same initialization as the teacher network and the student network respectively. In the training stage, by injecting semantic embeddings into the teacher network and using the distillation loss to align the knowledge representations of the student network and the teacher network, the student network can retain the semantic knowledge of the teacher network without injecting semantic information in the inference stage.
[0062] 3. The zero-shot sketch retrieval method based on multi-modal semantic information interaction effectively utilizes the semantic information of the text and the image features for sufficient information interaction, improves the semantic representation ability of the image features, and thus improves the performance of zero-shot sketch retrieval. This sketch retrieval method uses a semantic self-distillation structure to transfer the knowledge of the teacher feature extraction neural network to the student feature extraction neural network, reducing the semantic gap in zero-shot sketch retrieval. Description of the Drawings
[0063] Figure 1Flowchart of the zero-shot sketch retrieval method based on multi-modal semantic information interaction in Embodiment 1 of the present invention;
[0064] Figure 2 System framework diagram of the zero-shot sketch retrieval method based on multi-modal semantic information interaction in Embodiment 1;
[0065] Figure 3 Schematic diagram of channel information interaction of the zero-shot sketch retrieval method based on multi-modal semantic information interaction in Embodiment 1;
[0066] Figure 4 Schematic diagram of spatial information interaction of the zero-shot sketch retrieval method based on multi-modal semantic information interaction in Embodiment 1. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0068] Embodiment 1
[0069] Please refer to Figures 1 to 4 , this embodiment provides a zero-shot sketch retrieval method based on multi-modal semantic information interaction. This sketch retrieval method can handle the technical problem of lacking category labels of unseen classes in the retrieval stage under the zero-shot setting. Specifically, the sketch retrieval method of this embodiment includes the following steps, namely Step S1 to Step S6.
[0070] Step S1: Construct a data set divided into a training set and a test set. In this embodiment, first obtain labeled basic image data, and then preprocess the basic image data to construct a data set. Specifically, Step S1 includes: S11: Obtain labeled basic image data. The basic image data includes sketch images and natural images. Compress and crop the basic image data, and convert the image format of the sketch image; S12: Divide the data set into a training set and a test set with non-overlapping categories. Both the training set and the test set contain natural images and hand-drawn images. Among them, the preprocessing includes compressing and cropping the basic image to conform to the pixel size of 256*256, and converting the image format of the sketch image to the png format.
[0071] Step S2: Construct feature extraction neural networks for the teacher and the student. In this embodiment, the feature extraction neural networks for the teacher and the student adopt the Vision Transformer structure. The teacher feature extraction neural network embeds a semantic information interaction module in the benchmark network, and both feature extraction neural networks are initialized with the same pre-trained parameters.
[0072] Step S3: Through the text semantic embedding module and the large language model, expand the semantic information of the class name text in the training set to obtain text prompt features. In the embodiment, the class names of the training set are fed into the text semantic embedding module to obtain semantically rich class descriptions, and further obtain semantic prompt embeddings. The expansion of semantic information includes the following steps: S31: Combine the class names of the training set with question templates, such as "Describe the appearance of {class}", "{What does class} look like?", to obtain question words embedded with class names; S32: Feed the question words into the large language question answering model, and for {class}, obtain multiple corresponding semantically rich description statements; S33: Feed the description statements into the CLIP text encoder to obtain text features Z of multiple descriptions p ; S34: Feed the text features into the adapter to obtain text prompt features adapted to the large language question answering model
[0073] Step S4: First, feed the images of the training set into the teacher feature extraction neural network, and then through channel interaction and spatial interaction, let the text prompt features and the image features of the training set perform cross-modal information fusion to obtain multi-modal semantic information interaction features, and finally train the student feature extraction neural network. In this embodiment, specifically, step S4 includes: S41: Perform a chunking operation on the images of the training set to obtain multiple image chunks (defined as N image chunks), and feed them into the first layer of the Vision Transformer (ViT) to obtain corresponding multiple image chunk features; S42: Feed the semantic prompt features and the features of the image chunks into the text-image semantic information interaction module to obtain image features with semantic prompts; S43: Feed the image features with semantic prompts into the subsequent layers of the Vision Transformer (ViT) to obtain multi-modal semantic information interaction features; S44: Train the student feature extraction neural network through the N-pair loss function; S45: Train the student feature extraction neural network through the semantic self-distillation loss function.
[0074] In this embodiment, step S42 includes: S421: Averaging the features of the image patches to obtain an average image feature; S422: Horizontally concatenating the average image feature and the semantic prompt feature to obtain a text-image concatenated feature; S423: Feeding the text-image concatenated feature into a multi-layer perceptron (MLP) to obtain a channel interaction feature; S424: Pointwise adding the channel interaction feature and the image patch feature respectively to obtain a channel-semantic information interaction feature; S425: Vertically concatenating the channel-semantic information interaction feature and the semantic prompt feature to obtain a multi-modal information input feature; S426: Multiplying the multi-modal information input feature by multiple query weight matrices, key weight matrices, and value weight matrices to obtain multiple groups of query vectors, key vectors, and value vectors; S427: Calculating attention weights through the multiple groups of query vectors, key vectors, and value vectors; S428: Horizontally concatenating the results of multiplying the multiple attention weights by the value vectors and feeding them into a linear layer to obtain a text semantic prompt feature.
[0075] Among them, the calculation formula for the attention weight is:
[0076]
[0077] Among them, h is the number of heads, a is the attention weight, q is the query vector, k is the key vector, d k is the scaling factor, softmax(·) is the normalization exponential function, v is the value vector, q h k h T is used to calculate the dot product of q h and k h .
[0078] In this embodiment, in ViT, the input image (S and P represent the sketch and the real image) will be divided into a series of image patch features after passing through the first layer of ViT where m is the number of image patches, and is input into ViT together with the corresponding semantic prompt vector . Through cross-modal information interaction, the semantic relevance of the image features is enhanced, so as to generate a more discriminative and cross-modal consistent feature representation, effectively improving the performance and generalization ability of the sketch retrieval task. Cross-modal information interaction includes two key steps: "channel information interaction" ( Figure 3 ) and "spatial information interaction" ( Figure 4 ).
[0079] In the channel information interaction stage, for a single sample, first calculate the average value of all the image patches in the image patch feature Z to obtain a global visual context vector This vector is used to capture the global feature information of the image. Subsequently, the global visual context vector With the semantic prompt vector Perform channel-wise concatenation Input it into a two-layer multi-layer perceptron (MLP) module to further adjust and enhance the visual features. After the transformation operation of the MLP, an enhanced visual feature vector is generated
[0080]
[0081] where σ is the activation operation in the MLP, W1, b1, W2, b2 are the parameters of the MLP, and avg(z) is the average value of the image patch sequence
[0082] After the channel information interaction, the semantic prompt and the enhanced image patch sequence are concatenated into a new sequence and sent together to the multi-head attention for spatial information interaction:
[0083]
[0084]
[0085] where is the weight of the h-th head in the multi-head attention is the scaling factor to avoid gradient vanishing. Multiply the obtained a h and v h to get the weighted output O h , vertically concatenate to get the attention weight O, and send it to the linear layer to obtain the features of semantic information interaction:
[0086]
[0087] W out is the parameter of the output linear layer
[0088] Subsequently, send to the subsequent ViT layer to obtain the image feature f S|P containing rich semantic information
[0089] In this embodiment, step S44 includes: S441: Obtain the image data of the sketch image domain and the natural image domain respectively, and pair them and send them into the feature extraction neural network; S442: Through the N-pair loss function, guide the student feature extraction neural network to pull the sketch image and the natural image of the same category closer to each other, push the different categories farther apart from each other, and project the images of the two domains into the same feature space; where the N-pair loss function is:
[0090]
[0091] Wherein, B is the number of data groups in a training batch, d(·) is the cosine distance metric, and f i s is the i-th sketch feature, and are respectively the positive example photo feature and the negative example photo feature of, i, j ∈ {1, 2, …, B}, and margin is the margin hyperparameter.
[0092] In this embodiment, step S45 includes: based on the teacher feature extraction neural network and the student feature extraction neural network, through the self-distillation model, enabling the student feature extraction neural network to retain the features containing rich semantic interaction information learned in the teacher feature extraction neural network; wherein, the formula of the self-distillation model is:
[0093]
[0094] Wherein, is the feature obtained through the teacher feature extraction neural network, is the feature obtained through the student feature extraction neural network; KL(·) is the KL divergence, expressed as:
[0095]
[0096] Wherein, is the i-th feature obtained through the teacher feature extraction neural network, is the i-th feature obtained through the student feature extraction neural network.
[0097] Step S5: Input the image to be retrieved in the test set into the student feature extraction neural network trained in step S4 to obtain the feature of the image to be retrieved.
[0098] Step S6: First, input the query sketch image in the test set into the student feature extraction neural network trained in step S4 to obtain the feature of the query sketch image, and then compare and retrieve the query sketch image and the image to be retrieved according to the feature of the query sketch image and the feature of the image to be retrieved. In this embodiment, the distance between the feature of the query sketch image and the feature of the image to be retrieved is calculated to obtain the distance ranking as the retrieval result ranking, and zero-shot sketch retrieval based on multi-modal semantic information interaction is performed.
[0099] Table 1: Comparative experiment result table on three dataset settings
[0100]
[0101]
[0102] The following is an example to illustrate the technical effects of this embodiment: Table 1 shows the experimental results of this embodiment compared with other mainstream methods. As can be seen from the table, on the two mainstream datasets of TU-Belin and Sketchy, in the two evaluation metrics of mean average precision at N (mAP@N) and precision at N (Prec@N), this embodiment has achieved the best results.
[0103] In summary, compared with the existing sketch retrieval methods and systems, the zero-shot sketch retrieval method based on multi-modal semantic information interaction in this embodiment has the following beneficial effects:
[0104] 1. The zero-shot sketch retrieval method based on multi-modal semantic information interaction effectively utilizes text semantic information to alleviate the problem of excessive modality and knowledge differences in zero-shot sketch retrieval, improves retrieval performance, and the semantic self-distillation loss can complete knowledge transfer learning without additionally obtaining a pre-trained large model on a large-scale dataset as the teacher network. It uses a large language model to semantically expand the class name text, and sends the text prompt features and image features into the text-image semantic information interaction module for sufficient semantic information interaction, effectively improving the feature extraction effect of the image feature extraction neural network, achieving the purpose of reducing the modality gap, and solving the technical problem of excessive modality and knowledge differences in zero-shot sketch retrieval existing in the existing sketch retrieval methods.
[0105] 2. The zero-shot sketch retrieval method based on multi-modal semantic information interaction designs a semantic self-distillation module to transfer the knowledge of multi-modal interaction to the target feature extractor. This sketch retrieval method uses two vision transformers with the same initialization as the teacher network and the student network respectively. In the training stage, by injecting semantic embeddings into the teacher network and using the distillation loss to align the knowledge representations of the student network and the teacher network, the student network can retain the semantic knowledge of the teacher network without injecting semantic information in the inference stage.
[0106] 3. The zero-shot sketch retrieval method based on multi-modal semantic information interaction effectively utilizes the semantic information of the text and the image features for sufficient information interaction, improves the semantic representation ability of the image features, and thus improves the performance of zero-shot sketch retrieval. This sketch retrieval method uses the semantic self-distillation structure to transfer the knowledge of the teacher feature extraction neural network to the student feature extraction neural network, reducing the semantic gap in zero-shot sketch retrieval.
[0107] Embodiment 2
[0108] This embodiment provides a zero-shot sketch retrieval system based on multimodal semantic information interaction. The system includes a dataset partitioning module, an initialization module, an augmentation module, a training module, a feature extraction module, and a sketch retrieval module. The dataset partitioning module is used to construct a dataset partitioned into a training set and a test set. The initialization module is used to construct feature extraction neural networks for the teacher and the student. The augmentation module is used to augment the semantic information of the class name texts in the training set through a text semantic embedding module and a large language model to obtain semantic prompt features. The training module is used to first input the images in the training set into the feature extraction neural network, and then let the semantic prompt features and the image features in the training set perform cross-modal information fusion through channel interaction and spatial interaction to obtain multimodal semantic information interaction features. Finally, the student feature extraction neural network is trained. The feature extraction module is used to input the images to be retrieved in the test set into the student feature extraction neural network trained by the training module to obtain the features of the images to be retrieved. The sketch retrieval module is used to first input the query sketch images in the test set into the student feature extraction neural network trained by the training module to obtain the features of the query sketch images, and then compare and retrieve the query sketch images and the images to be retrieved based on the features of the query sketch images and the images to be retrieved.
[0109] In this embodiment, the dataset partitioning module, the initialization module, the augmentation module, the training module, the feature extraction module, and the sketch retrieval module respectively implement steps S1 to S6 in the zero-shot sketch retrieval method based on multimodal semantic information interaction in Embodiment 1. The specific method steps to be executed have been described in Embodiment 1 and will not be elaborated here. Of course, in some other embodiments, these modules can also execute some of the steps in the retrieval method in Embodiment 1 as long as the basic functions of the modules can be realized.
[0110] Embodiment 3
[0111] This embodiment provides a computer terminal, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the zero-shot sketch retrieval method based on multimodal semantic information interaction in Embodiment 1.
[0112] When the method in Embodiment 1 is applied, it can be applied in the form of software, such as designed as an independently running program and installed on a computer terminal. The computer terminal can be a computer, a smart phone, a control system, and other Internet of Things devices, etc. The method in Embodiment 1 or Embodiment 2 can also be designed as an embedded running program and installed on a computer terminal, such as installed on a single-chip microcomputer.
[0113] Embodiment 4
[0114] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the zero-shot sketch retrieval method based on multimodal semantic information interaction in Embodiment 1 are implemented.
[0115] When the method of Embodiment 1 or Embodiment 2 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program on a computer-readable storage medium. The computer-readable storage medium can be a USB flash drive, designed as a USB key, and the program for starting the entire method by external triggering can be designed through the USB flash drive.
[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A zero-shot sketch retrieval method based on multimodal semantic information interaction, characterized in that: It includes the following steps: S1: Construct a dataset divided into training set and test set; S2: Constructing feature extraction neural networks for teachers and students; S3: Expanding the semantic information of the class name text of the training set through a text semantic embedding module and a large language model to obtain text prompt features; S4: firstly, the image of the training set is sent to the teacher feature extraction neural network, and then the text prompt feature and the image feature of the training set are cross-modal information fused by channel interaction and space interaction to obtain multi-modal semantic information interaction features, and finally the student feature extraction neural network is trained; S5: inputting the image to be retrieved in the test set into the student feature extraction neural network trained in step S4 to obtain features of the image to be retrieved; S6: First, input the query sketch image in the test set into the student feature extraction neural network trained in step S4 to obtain the query sketch image features, and then compare and retrieve the query sketch image and the image to be retrieved based on the query sketch image features and the image features to be retrieved.
2. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 1, characterized in that: In the step S1, basic image data with labels are first obtained, and then the basic image data are preprocessed to construct the data set; in the step S2, the feature extraction neural networks of the teacher and the student adopt a visual Transformer structure, the teacher feature extraction neural network embeds a semantic information interaction module on the baseline network, and both feature extraction neural networks are initialized with the same pre-trained parameters; in the step S6, the distance between the query sketch image features and the image features to be retrieved is calculated, the distance ranking is obtained as the retrieval result ranking, and zero-sample sketch retrieval based on multimodal semantic information interaction is performed.
3. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 1, characterized in that: The step S3, expanding the semantic information comprises the following steps: S31: combining the training set class name with the question template to obtain a question word embedded with the class name; S32: sending the question words into a large language question answering model to obtain a corresponding plurality of relevant semantically rich description sentences; S33: sending the description sentence to a CLIP text encoder to obtain text features of multiple descriptions; S34: Send the text features to an adapter to obtain text prompt features suitable for the large language question-answering model.
4. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 1, characterized in that: The step S4 comprises: S41: performing a block operation on the image of the training set to obtain multiple image blocks, and sending the blocks to the first layer of the visual transformer to obtain corresponding multiple image block features; S42: sending the semantic hint feature and the feature of the image block to a text image semantic information interaction module to obtain the image feature of the semantic hint; S43: sending the image features of the semantic hint to the subsequent layer of the visual transformer to obtain the multimodal semantic information interaction features; S44: training the student feature extraction neural network through an N-pair loss function; S45: Training the student feature extraction neural network through a semantic self-distillation loss function.
5. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 4, characterized in that: The step S42 comprises: S421: Averaging the features of the image blocks to obtain average image features; S422: horizontally splicing the average image feature and the semantic hint feature to obtain a text image splicing feature; S423: Sending the text image splicing features into a multi-layer perceptron to obtain channel interaction features; S424: Add the channel interaction feature and the image block feature point by point to obtain a channel semantic information interaction feature; S425: vertically splicing the channel semantic information interaction feature and the semantic prompt feature to obtain a multimodal information input feature; S426: multiplying the multimodal information input feature with a plurality of query weight matrices, key weight matrices, and value weight matrices to obtain a plurality of groups of query vectors, key vectors, and value vectors; S427: Calculate attention weights through multiple sets of query vectors, key vectors, and value vectors; S428: Multiply the attention weight and the value vector, perform horizontal splicing, and send them to the linear layer to obtain text semantic hint features.
6. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 5, characterized in that: The calculation formula of the attention weight is: Where h is the number of heads, a is the attention weight, q is the query vector, k is the key vector, and d k is the scaling factor, softmax(·) is the normalized exponential function, q h k h T To calculate q h and k h The dot product of .
7. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 4, characterized in that: The step S44 comprises: S441: Obtain image data in the sketch image domain and the natural image domain respectively, and pair them up and send them into the feature extraction neural network; S442: guiding the student feature extraction neural network to bring the sketch image and the natural image of the same category closer to each other and push different categories away from each other through the N-pair loss function, and projecting the images of the two domains into the same feature space; Among them, the N-pair loss function is: Where B is the number of data sets in a training batch, d(·) is the cosine distance metric, and f i S is the i-th sketch feature, and f i S The positive and negative photo features of , i,j∈{1,2,…,B}, margin is the edge hyperparameter.
8. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 6, characterized in that: The step S45 comprises: Based on the teacher feature extraction neural network and the student feature extraction neural network, the student feature extraction neural network retains the features containing rich semantic interaction information learned in the teacher feature extraction neural network through a self-distillation model; wherein the formula of the self-distillation model is: In the formula, is the feature obtained by the teacher feature extraction neural network, is the feature obtained by the student feature extraction neural network; KL(·) is the KL divergence, expressed as: In the formula, is the i-th feature obtained by the teacher feature extraction neural network, is the i-th feature obtained by the student feature extraction neural network.
9. The zero-shot sketch retrieval method based on multimodal semantic information interaction as claimed in claim 1, characterized in that: The step S1 comprises: S11: acquiring labeled basic image data, the basic image data including a sketch image and a natural image, compressing and cropping the basic image data, and converting the sketch image into an image format; S12: Divide the data set into a training set and a test set with non-overlapping categories, wherein both the training set and the test set contain natural images and hand-drawn images.
10. A zero-shot sketch retrieval system based on multimodal semantic information interaction, characterized in that: It includes: A data set partitioning module is used to construct a data set divided into a training set and a test set; Initialization module, which is used to build the feature extraction neural network of teachers and students; An expansion module, which is used to expand the semantic information of the class name text of the training set through a text semantic embedding module and a large language model to obtain semantic prompt features; A training module, which is used to first send the images of the training set to the feature extraction neural network, then perform cross-modal information fusion of the semantic prompt features and the image features of the training set through channel interaction and spatial interaction to obtain multi-modal semantic information interaction features, and finally train the student feature extraction neural network; A feature extraction module, which is used to input the image to be retrieved in the test set into the student feature extraction neural network trained by the training module to obtain the features of the image to be retrieved; The sketch retrieval module is used to first input the query sketch image in the test set into the student feature extraction neural network trained by the training module to obtain the query sketch image features, and then compare and retrieve the query sketch image and the image to be retrieved based on the query sketch image features and the image features to be retrieved.
Citation Information
Cited By
Multi-modal CAD model retrieval method and device based on text and sketch
CN122019818A