A Cross-modal Retrieval Method, System and Related Devices Enhanced by Large Models

Through large-model enhancement technology, combined with fine-grained semantic segmentation, text summary extraction and multi-branch encoder, the retrieval accuracy problem caused by cross-modal data deviation is solved, and a more efficient image-text retrieval effect is achieved.

CN119719451BActive Publication Date: 2025-05-27UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510220410.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-27
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing cross-modal search technology lacks in-depth analysis of cross-modal data bias, resulting in low retrieval accuracy.

Method used

Using the method of large-model enhancement, the image is fine-grained semantic segmentation through the SAM segmentation model, and the text summary is obtained using a large language model. Combined with the multi-branch encoder of the pre-trained CLIP model, a multi-level collaborative alignment loss function is constructed to coordinate the image mode and text mode in the common semantic space.

Benefits of technology

By enhancing collaborative learning of information, the accuracy of image-text retrieval is improved and the model's understanding of cross-modal data is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719451B_ABST
    Figure CN119719451B_ABST
Patent Text Reader

Abstract

The present invention discloses a large model-enhanced cross-modal retrieval method, system and related devices. The cross-modal retrieval method includes: obtaining image-text pairs; based on the image-text pairs, obtaining large model enhancement information of the image-text pairs, combining the original images, texts and enhancement information, and using a multi-branch encoder of a pre-trained CLIP model to obtain multiple feature vectors, constructing a multi-level collaborative alignment loss function, and performing collaborative alignment on the image modality and the text modality in a common semantic space; training the model through the multi-level collaborative alignment loss function and a pre-constructed training database, and performing retrieval through the trained model. The present invention performs collaborative learning on the image and text features obtained by the encoder, adds auxiliary semantic enhancement information, and performs collaborative alignment on the image modality and the text modality in a common semantic space to train a better retrieval network so as to improve the accuracy of image-text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cross-modal retrieval, and specifically to a large model enhanced cross-modal retrieval method, system and related devices. Background Art

[0002] With the rapid increase in the quantity of multi-modal data (such as text, pictures and videos) from various sources, the demand for multi-modal data analysis is also growing. Cross-modal retrieval aims to search for corresponding text given a large-scale image as a query, or search for corresponding pictures given text as a query.

[0003] In recent years, great progress has been made in cross-modal image-text retrieval by advancing the structures of convolutional neural networks and LSTMs to Transformers and pre-trained vision and language models. Using better model structures and models pre-trained on large-scale data, the performance of retrieval has been greatly improved. However, these works rarely focus on the fundamental aspect of a task, namely cross-modal data bias. Specifically, the text and its corresponding image may contain inconsistent information with each other, which has a potential harm to the retrieval performance.

[0004] On the one hand, the content described in the text does not necessarily exactly match the content in the image, that is, there is an inherent semantic gap between the two. Therefore, the text may contain redundant information that has nothing to do with the visual appearance of the corresponding image. On the other hand, images usually also contain information irrelevant to the text, such as background and noise information, etc.

[0005] Generally speaking, although the existing technologies have made certain progress, they lack in-depth analysis of the cross-modal data bias characteristics of the task, thus restricting the further development of retrieval. Therefore, how to effectively analyze and solve the existing cross-modal data bias is a technical problem that cross-modal retrieval urgently needs to solve. Summary of the Invention

[0006] In this embodiment, a large model enhanced cross-modal retrieval method, system, electronic device and storage medium are provided to solve the problem of low accuracy in cross-modal data bias and image-text retrieval in related technologies.

[0007] In a first aspect, an embodiment of the present invention provides a large model enhanced cross-modal retrieval method, and the cross-modal retrieval method includes:

[0008] Obtain image-text pairs;

[0009] Based on the image-text pair, obtain the large model enhancement information of the image-text pair, including using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein; using the large language model to obtain the summary of the text and extract the key semantic fragments;

[0010] Combine the original image, text, and enhancement information, and adopt the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors. The multiple feature vectors include image features, text features, fine-grained image entity features, and text summary features;

[0011] Construct a multi-level collaborative alignment loss function to perform collaborative alignment on the image modality and text modality in the common semantic space;

[0012] Train the model through the multi-level collaborative alignment loss function and the pre-constructed training database, use the gradient descent algorithm to train the multi-branch encoder, and perform retrieval through the trained model.

[0013] In an optional embodiment, using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein includes:

[0014] Use the SAM segmentation model to perform semantic segmentation on the image, use the uniformly sampled points of the image as prompts to perform image segmentation suggestions, and obtain all sets of segmentation masks;

[0015] Filter all the segmentation masks to obtain the set of fine-grained semantic segmentation instances of the filtered image;

[0016] Supplement the enhancement with partial image missing. Randomly crop a quarter of the original image area as a fine-grained instance to supplement the length of the set of fine-grained semantic segmentation instances of the image.

[0017] In an optional embodiment, using the large language model to obtain the summary of the text and extract the key semantic fragments includes:

[0018] Obtain key information of the input long text through the Llama2 large language model, including the title, ingredients, and introduction;

[0019] Generate a text summary based on the title, ingredients, and introduction.

[0020] In an optional embodiment, the construction of the multi-branch encoder of the pre-trained CLIP model includes:

[0021] Add an adaptation layer to the Transformer structure in the text encoder and image encoder of the CLIP model; the adaptation layer is obtained through downsampling projection, upsampling projection, and a residual connection;

[0022] Construct a feature encoder for different input information based on the encoder of the CLIP model with an adaptation layer;

[0023] The feature encoder includes an image encoder, a text encoder, an instance encoder, and a text summary encoder.

[0024] In an optional embodiment, combine the original image, text, and enhancement information, and use the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors, including:

[0025] Use the image encoder with an adaptation layer to extract features from the input image to obtain image features;

[0026] Extract features from the title, ingredients, and introduction of the input text respectively through the text encoder of the CLIP model with an adaptation layer to obtain three partial feature vectors, and then input the parallel feature vectors into a two-layer Transformer structure for fusion;

[0027] Concatenate the three partial feature vectors, and pass through a fully connected layer and a non-linear activation function to obtain the final text features;

[0028] Encode the text summary through the text encoder of the CLIP model with an adaptation layer added to obtain text summary features;

[0029] Encode the fine-grained segmentation entities of the image through the image encoder of the CLIP model to obtain fine-grained image entity features.

[0030] In an optional embodiment, construct a multi-level collaborative alignment loss function to perform collaborative alignment on the image modality and text modality in the common semantic space, including:

[0031] Define the matching image and text as a positive sample pair, and the non-matching ones as a negative sample pair, and use the cosine similarity as the standard to measure the relationship between sample pairs;

[0032] Use the circular loss function to increase the similarity between positive sample pairs and decrease the similarity between negative sample pairs;

[0033] Construct an alignment loss based on multiple feature vectors. The alignment loss includes the semantic alignment loss between the image and text, the semantic alignment loss between the text and the fine-grained segmentation of the image, the semantic alignment loss between the image and the text summary, and the semantic alignment loss within the long text;

[0034] Sum the alignment losses with weights to obtain the multi-level collaborative alignment loss function.

[0035] In an optional embodiment, the model is trained through the multi-level collaborative alignment loss function and a pre-constructed database, and retrieval is performed through the trained model, including:

[0036] For samples in the database that have no corresponding images but only contain text, semantic alignment within the text is performed using the semantic alignment loss within the long text; for samples containing image-text, combined with the generated large model enhanced information, training is carried out using the multi-level collaborative alignment loss function.

[0037] The model parameters are updated by alternately training in two ways for the samples.

[0038] Compared with the prior art, the beneficial effects of the large model enhanced cross-modal retrieval method of the present invention are as follows:

[0039] In the present invention, through the SAM segmentation model, fine-grained semantic targets in the image are segmented; through the Llama2 large language model, key information in the long text is summarized and a key semantic summary text is generated; based on the lightweight fine-tuned image-text multi-branch encoder of CLIP, the pre-trained model CLIP is fine-tuned with a lightweight adaptation layer to obtain feature vectors such as images and texts; and a multi-level circular loss is designed to perform collaborative learning on the image and text features obtained by the encoder. By adding auxiliary semantic enhancement information, the image modality and the text modality are collaboratively aligned in the common semantic space to train a better retrieval network, thereby improving the accuracy of image-text retrieval.

[0040] In a second aspect, an embodiment of the present invention provides a large model enhanced cross-modal retrieval system, including:

[0041] An acquisition module, configured to acquire image-text pairs;

[0042] An enhanced information acquisition module, configured to obtain the large model enhanced information of the image-text pair based on the image-text pair, including performing fine-grained semantic segmentation on the image using the SAM segmentation model and selecting the semantic entities therein; using the large language model to obtain the summary of the text and extract the key semantic segments;

[0043] A feature extraction module, configured to combine the original image, text, and enhanced information, and adopt the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors, where the multiple feature vectors include image features, text features, fine-grained image entity features, and text summary features;

[0044] A loss function design module, configured to construct a multi-level collaborative alignment loss function to perform collaborative alignment on the image modality and the text modality in the common semantic space;

[0045] A retrieval module, configured to train a model through the multi-level collaborative alignment loss function and a pre-constructed training database, train a multi-branch encoder using a gradient descent algorithm, and perform retrieval through the trained model.

[0046] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the bus. The processor can call the logical instructions in the memory to execute the steps of the method provided in the first aspect.

[0047] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large model enhanced cross-modal retrieval method described in the first aspect.

[0048] Compared with the prior art, the beneficial effects of the large model enhanced cross-modal retrieval system, electronic device, and storage medium of the present invention are the same as those of the large model enhanced cross-modal retrieval method described in the first aspect, so they will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of the large model enhanced cross-modal retrieval method in the embodiment of the present invention;

[0051] Figure 2 It is a schematic diagram of the principle of the large model enhanced cross-modal retrieval method in the embodiment of the present invention;

[0052] Figure 3 It is a schematic diagram of large model enhanced content generation of an image in the embodiment of the present invention;

[0053] Figure 4 It is a schematic diagram of large model enhanced content generation of text in the embodiment of the present invention;

[0054] Figure 5 It is a schematic diagram of an image-text multi-branch encoder for synergistically enhancing semantic information in the embodiment of the present invention;

[0055] Figure 6 It is a structural block diagram of the large model enhanced cross-modal retrieval system in the embodiment of the present invention;

[0056] Figure 7This is a structural block diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0057] To more clearly understand the purpose, technical solution, and advantages of the present application, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments.

[0058] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products, or devices. The terms "connected", "coupled", etc. involved in the present application do not limit to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The "multiple" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in the present application only distinguish similar objects and do not represent a specific order for the objects.

[0059] In an embodiment of the present invention, a large model enhanced cross-modal retrieval method is provided. Figure 1 This is a flowchart of the large model enhanced cross-modal retrieval method of the present invention. Figure 2 This is a schematic diagram of the large model enhanced cross-modal retrieval method in an embodiment of the present invention; as Figure 1 and Figure 2 shown, the process includes the following steps:

[0060] S100. Obtain image-text pairs;

[0061] Specifically, paired image-text pairs are extracted from the database. For the case of multiple-to-one or one-to-many, a random selection method is adopted to learn better model robustness.

[0062] S200. Based on the image-text pair, obtain the large model enhancement information of the image-text pair, including using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein; using the large language model to obtain the summary of the text and extract the key semantic fragments;

[0063] First of all, it should be noted that by introducing the basic large model and utilizing its general semantic extraction ability, semantic enhancement operations are performed on the image and the text respectively. Use the image segmentation model (SAM segmentation model) to perform semantic segmentation on the image, and then screen all the segmentation instances to obtain the fine-grained semantic instances of the image; use the large language model to summarize and extract the key semantic information of the text to generate the summary text of the long text.

[0064] In this embodiment, using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein includes:

[0065] Use the SAM segmentation model to perform semantic segmentation on the image, use the evenly sampled points of the image as prompts to make image segmentation suggestions, and obtain all the sets of segmentation masks;

[0066] Screen all the segmentation masks to obtain the set of fine-grained semantic segmentation instances of the image after screening;

[0067] Supplement the enhancement of some missing images. Randomly crop a quarter of the original image area as a fine-grained instance to supplement the length of the set of fine-grained semantic segmentation instances of the image.

[0068] Specifically, first, introduce the SAM segmentation model, use the "everything" mode to segment the input image, realize zero-shot segmentation object suggestion, for the input image For example, on the image, use a point grid of 64*64 pixel size to uniformly sample the image. The sampled points are used as prompt vectors to predict the generation of segmentation masks. For the large number of masks returned, the loU prediction module of the model selects the masks with high confidence. After obtaining the stable masks, use non-maximum suppression (NMS) to filter out duplicate masks. The threshold of NMS is set to 0.9. For a food image, about 900 masks can be generated on average. Sort according to the mean value of the confidence and stability scores to obtain the mask set :

[0069] ;

[0070] For all the masks in the set , first sort them from largest to smallest according to the area, filter out the mask instances with too small area and loss of semantic information, and update the set There are only 16 instances. Further, the instances are input into the image encoder of CLIP in parallel to obtain a set of instance features:

[0071] ;

[0072] To further screen out the fine-grained semantic instances of the image, as Figure 3 shown, the text encoder of the CLIP model is used to obtain the feature vector of the text "an image of {topic}", calculate the similarity between the features in the instance set and this text feature, screen out the instances that meet the similarity threshold, and select the top 4 with the largest similarity scores as the key fine-grained semantic instances. For images with less than 4 fine-grained semantic instances reaching the threshold, randomly crop a quarter of the original image area as instance supplement.

[0073] Further, as Figure 4 shown, a large language model is used to obtain the summary of the text and extract the key semantic fragments, including:

[0074] Obtain key information from the input long text through the Llama2 large language model, including the title, ingredients, and introduction;

[0075] Generate a text summary based on the title, ingredients, and introduction.

[0076] Specifically, for the large language model Llama2, the model of is adopted, and through designing prompt engineering, the text information is input into the large model.

[0077] Exemplarily, first, let the large model play the role of a helpful assistant and input the system prompt: You are a helpful assistant, try to provide help as much as possible when answering questions, and the answer should be within 30 words. Second, input the text summary prompt: Provide a text as follows: Title: {title}, Ingredients: {ingredients}, Introduction: {introduction}. Return the key summary information of the long text, and the summary should be informative and as objective as possible. Among them, {title}, {ingredients}, and {introduction} are replaced with the content in the long text, and let the long text be , including the title ingredients , introduction , then there is the text summary as:

[0078] );

[0079] Among them, the maximum value of the number of input tokens is set to 2048, the maximum output token is 256, and the model temperature coefficient is 0.6.

[0080] S300. Combine the original image, text, and enhanced information, and use the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors, where the multiple feature vectors include image features, text features, fine-grained image entity features, and text summary features;

[0081] Combine Figure 5 As shown, it should be noted that the construction of the multi-branch encoder of the pre-trained CLIP model includes:

[0082] Add an adaptation layer to the Transformer structure in the text encoder and image encoder of the CLIP model; the adaptation layer is obtained through downsampling projection, upsampling projection, and a residual connection;

[0083] Specifically, for the pre-trained CLIP model, add a lightweight adaptation layer, keep all the pre-trained parameters of CLIP frozen, and only fine-tune and update the newly added adaptation layer. Specifically, in each layer of the Transformer structure in the text encoder and image encoder of CLIP, after its multi-head attention layer and MLP layer, add an adaptation layer respectively. The adaptation layer includes a downsampling projection Project the hidden layer to a low dimension , and after passing through an upsampling projection Project back to the original dimension, and a residual connection , in the following form:

[0084] ;

[0085] is the activation function, and here the rectified linear unit function is adopted. Among them, the projection dimension is controlled by setting the scaling factor , the dimension of is 512, and the scaling factor is set to 4:

[0086] ;

[0087] Construct feature encoders for different input information according to the encoder of the CLIP model with an adaptation layer;

[0088] The feature encoders include an image encoder, a text encoder, an instance encoder, and a text summary encoder.

[0089] Combine the original image, text, and enhanced information, and use the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors, including:

[0090] Use the image encoder with an adaptation layer to extract features from the input image to obtain image features;

[0091] The title, ingredients, and introduction of the input text are respectively subjected to feature extraction through the text encoder of the CLIP model with an adaptation layer to obtain three partial feature vectors, and the parallel feature vectors are then input into a two-layer Transformer structure for fusion;

[0092] The three partial feature vectors are concatenated and passed through a fully connected layer and a non-linear activation function to obtain the final text feature;

[0093] The text summary is encoded through the text encoder of the CLIP model with an adaptation layer to obtain the text summary feature;

[0094] The fine-grained segmentation entities of the image are subjected to feature encoding through the image encoder of the CLIP model to obtain the fine-grained image entity features.

[0095] Specifically, as shown in Figure 5 for the input image , feature extraction is performed using the image encoder with an adaptation layer, and the image features are:

[0096] ;

[0097] For the input text , it includes three parts: the title, ingredients, and introduction, that is , and each of the three parts is encoded separately to extract independent features, using the text encoder of the CLIP model with an adaptation layer.

[0098] Furthermore, for the title , there are:

[0099] ;

[0100] Since the text of the ingredients and description is relatively long, for these two parts, they are first segmented into sentences and then encoded in parallel, with a maximum length of 20 tokens per sentence and a maximum of 15 sentences.

[0101] ;

[0102] ;

[0103] The parallel feature vectors are then input into a two-layer Transformer structure for fusion, and there are:

[0104] ;

[0105] ;

[0106] Finally, the three partial feature vectors are concatenated, and after passing through a fully connected layer and a non-linear activation function , here the hyperbolic tangent function is used to obtain the final text feature ( The dimension of which is uniformly 512 dimensions):

[0107] ;

[0108] For the introduced fine-grained image segmentation enhanced information ( ), the image encoder of the fully frozen CLIP model is used for feature encoding, and then the average of all segmentation enhanced features in the set is taken to obtain (512 dimensions), there is:

[0109] ;

[0110] S400. Construct a multi-level collaborative alignment loss function to collaboratively align the image modality and the text modality in the common semantic space;

[0111] In this embodiment, constructing a multi-level collaborative alignment loss function to collaboratively align the image modality and the text modality in the common semantic space includes:

[0112] Define the matching images and texts as positive sample pairs, and the non-matching ones as negative sample pairs, and use cosine similarity as the standard to measure the relationship between sample pairs;

[0113] Use the circular loss function to increase the similarity between positive sample pairs and decrease the similarity between negative sample pairs;

[0114] Construct an alignment loss based on multiple feature vectors. The alignment loss includes the semantic alignment loss between images and texts, the semantic alignment loss between texts and fine-grained image segmentation, the semantic alignment loss between images and text summaries, and the semantic alignment loss within long texts;

[0115] Sum the weighted alignment losses to obtain the multi-level collaborative alignment loss function.

[0116] Specifically, a loss function for collaboratively aligning multiple feature vectors in the text and visual modality spaces is designed. During the retrieval process, the matching images and texts are defined as positive sample pairs, and the non-matching ones are defined as negative sample pairs. Cosine similarity is used as the standard to measure the relationship between sample pairs, and it is represented by :

[0117] ;

[0118] The similarity between positive sample pairs is denoted as , and the similarity between negative sample pairs is denoted as , the circular loss function is used to increase the similarity between positive sample pairs and decrease the similarity between negative sample pairs, as follows:

[0119] ;

[0120] where , are the margins of positive and negative samples respectively. is the scaling factor, and are independent weighting factors, which are linear functions of and . Then let , be the relaxation factor, with a value of 0.25. Further, the bidirectional circular loss is adopted:

[0121] ;

[0122] Next, based on the semantic enhancement vector and the original text and image vectors, the semantic alignment loss between the image and the text , the semantic alignment loss between the text and the fine-grained segmentation of the image , and the semantic alignment loss between the image and the text summary are designed. In addition, for the three components within the long text, the semantic alignment loss is constructed between each pair, forming the semantic alignment loss within the long text . Finally, the four groups of losses are combined to construct a multi-level large model enhanced collaborative alignment loss function :

[0123] ;

[0124] S500. The model is trained through the multi-level collaborative alignment loss function and the pre-constructed training database, and the multi-branch encoder is trained using the gradient descent algorithm, and retrieval is performed through the trained model.

[0125] The model is trained through the multi-level collaborative alignment loss function and the pre-constructed database, and retrieval is performed through the trained model, including:

[0126] For samples in the database that have no corresponding image but only contain text, the semantic alignment loss within the long text is used to perform semantic alignment within the text; for samples containing image-text, combined with the generated large model enhanced information, the multi-level collaborative alignment loss function is used for training;

[0127] The model parameters are updated alternately through the two sample training methods.

[0128] Specifically, in the training phase, for samples in the database that have no corresponding images but only contain text, only the semantic alignment loss within the long text is adopted. Semantic alignment is performed within the text, while for samples containing image-text, combined with the enhanced information of the generated large model, the collaborative alignment loss function is used for training, and the two sample training methods are alternately used to update the model parameters.

[0129] In the testing phase, using the trained multi-branch encoder, the trained image features I and text features are obtained. Calculate the cosine similarity of the feature vectors between the two. Sort the candidate texts / images according to the similarity score, and select the sorted samples according to the retrieval needs.

[0130] To illustrate the effectiveness of the present invention, the following experiments were conducted for verification.

[0131] Experiments were conducted on the Recipe1M dataset. Using the commonly used recall metrics in the retrieval field, Recall@1, Recall@5, Recall@10, and the median position medR of the correct retrieval, the size of the test model's database is 1000 pairs and 10000 pairs. Compare the model performance with other baseline models on these two test scales to prove the superior performance of the large model enhanced cross-modal retrieval method. Finally, ablation experiments were conducted on the proposed image fine-grained segmentation enhancement, text summary enhancement, multi-branch feature encoder, and collaborative alignment loss to verify the effectiveness of each step design.

[0132] As shown in Table 1 and Table 2, the experiments show the model performance comparison experiments with other methods on the 1000-pair and 10000-pair test data scales of Recipe1M, and show the results of retrieving text from images by the model.

[0133]

[0134] Table 1 Performance comparison of the model on the 1000-pair sample test scale of Recipe1M

[0135]

[0136] Table 2 Performance comparison of the model on the 10000-pair sample test scale of Recipe1M

[0137] As can be seen from Table 1 and Table 2, on the test data of two scales, the large model enhanced cross-modal retrieval method shows excellent performance. In addition to the retrieval performance exceeding other methods, compared with other methods that fine-tune CLIP in full (such as VLPCOOK, etc.), the parameter update in the training stage of this method is only about 14% of theirs, reflecting the excellent implementation of this method.

[0138] An embodiment of the present invention also provides a large model enhanced cross-modal retrieval system, which is used to implement the above method embodiments. Those that have been described will not be repeated here. The following terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated.

[0139] As Figure 6 shown, Figure 6 is a structural block diagram of the large model enhanced cross-modal retrieval system in the present invention. The system includes:

[0140] An acquisition module 101, configured to acquire image-text pairs;

[0141] An enhanced information acquisition module 102, configured to acquire large model enhanced information of the image-text pair based on the image-text pair, including using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein; using the large language model to obtain the summary of the text and extract the key semantic fragments;

[0142] A feature extraction module 103, configured to combine the original image, text, and enhanced information, and use the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors. The multiple feature vectors include image features, text features, fine-grained image entity features, and text summary features;

[0143] A loss function design module 104, configured to construct a multi-level collaborative alignment loss function to perform collaborative alignment on the image modality and the text modality in the common semantic space;

[0144] A retrieval module 105, configured to train the model through the multi-level collaborative alignment loss function and a pre-constructed training database, train the multi-branch encoder using the gradient descent algorithm, and perform retrieval through the trained model.

[0145] Figure 7 is a structural block diagram of the electronic device provided by the embodiment of the present invention. As Figure 7As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the following method:

[0146] Obtain an image-text pair;

[0147] Based on the image-text pair, obtain the large model enhancement information of the image-text pair, including using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein; using the large language model to obtain the summary of the text and extract the key semantic fragments;

[0148] Combine the original image, text, and enhancement information, and adopt the multi-branch encoder of the pre-trained CLIP model to obtain multiple feature vectors. The multiple feature vectors include image features, text features, fine-grained image entity features, and text summary features;

[0149] Construct a multi-level collaborative alignment loss function to perform collaborative alignment on the image modality and the text modality in the common semantic space;

[0150] Train the model through the multi-level collaborative alignment loss function and the pre-constructed training database, use the gradient descent algorithm to train the multi-branch encoder, and perform retrieval through the trained model.

[0151] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0152] The embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the methods provided in the above-mentioned embodiments, for example, including:

[0153] Obtain image-text pairs;

[0154] Based on the image-text pairs, obtain large model enhancement information of the image-text pairs, including using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select semantic entities therein; using a large language model to obtain a summary of the text and extract key semantic fragments;

[0155] Combine the original image, text, and enhancement information, and adopt a multi-branch encoder of a pre-trained CLIP model to obtain multiple feature vectors, where the multiple feature vectors include image features, text features, fine-grained image entity features, and text summary features;

[0156] Construct a multi-level collaborative alignment loss function to perform collaborative alignment on the image modality and the text modality in a common semantic space;

[0157] Train the model through the multi-level collaborative alignment loss function and a pre-constructed training database, use the gradient descent algorithm to train the multi-branch encoder, and perform retrieval through the trained model.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model enhanced cross-modal retrieval method, characterized in that: The cross-modal retrieval method comprises: Get image-text pairs; Based on the image-text pair, we obtain the large model enhancement information of the image-text pair, including using the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities therein; using the large language model to obtain the summary of the text and extract the key semantic fragments; Combining the original image, text and enhanced information, a multi-branch encoder of the pre-trained CLIP model is used to obtain multiple feature vectors, including image features, text features, fine-grained image entity features and text summary features; The multi-branch encoder construction of the pre-trained CLIP model includes: An adaptation layer is added to the text encoder of the CLIP model and the Transformer structure in the image encoder; the adaptation layer is obtained by downsampling projection and upsampling projection and a residual connection; Based on the encoder of the CLIP model with an adaptation layer, construct feature encoders for different input information; Feature encoders include image encoders, text encoders, instance encoders, and text summary encoders; Construct a multi-level collaborative alignment loss function to collaboratively align image and text modalities in a common semantic space, including: The matching images and texts are defined as positive sample pairs, the mismatched ones are defined as negative sample pairs, and the cosine similarity is used as the standard to measure the relationship between sample pairs; The ring loss function is used to increase the similarity between positive sample pairs and reduce the similarity between negative sample pairs; Alignment loss is constructed based on multiple feature vectors. The alignment loss includes semantic alignment loss between image and text, semantic alignment loss between text and image fine-grained segmentation, semantic alignment loss between image and text summary, and semantic alignment loss within long text. The weighted sum of the alignment losses is used to obtain the multi-level collaborative alignment loss function; The model is trained using a multi-level collaborative alignment loss function and a pre-built training database, and retrieval is performed using the trained model.

2. The large model enhanced cross-modal retrieval method according to claim 1, characterized in that: Use the SAM segmentation model to perform fine-grained semantic segmentation on the image and select the semantic entities, including: The SAM segmentation model is used to perform semantic segmentation on the image, and points uniformly sampled from the image are used as prompts to make image segmentation suggestions and obtain a set of all segmentation masks; Filter all segmentation masks to obtain a set of filtered image fine-grained semantic segmentation instances; To supplement the missing enhancements of some images, a quarter of the original image is randomly cropped and sampled as a fine-grained instance, which supplements the length of the image fine-grained semantic segmentation instance set.

3. The large model enhanced cross-modal retrieval method according to claim 1, characterized in that: Use a large language model to obtain a summary of the text and extract key semantic fragments, including: The Llama2 large language model is used to obtain key information from the input long text, including title, ingredients and introduction; Generate a text summary based on the title, components, and introduction.

4. The large model enhanced cross-modal retrieval method according to claim 1, characterized in that: Combining the original image, text, and enhanced information, a multi-branch encoder from the pre-trained CLIP model is used to obtain multiple feature vectors, including: Use an image encoder with an adaptation layer to extract features from the input image to obtain image features; The text encoder of the CLIP model with an adaptation layer extracts features from the title, components, and introduction of the input text respectively to obtain three partial feature vectors. The parallel feature vectors are then input into a two-layer Transformer structure for fusion. The three partial feature vectors are concatenated and passed through a fully connected layer and a nonlinear activation function to obtain the final text feature; The text summary is encoded by adding the text encoder of the CLIP model with an adaptation layer to obtain the text summary features; The image encoder of the CLIP model is used to encode the features of the fine-grained image segmentation entities to obtain fine-grained image entity features.

5. The large model enhanced cross-modal retrieval method according to claim 1, characterized in that: The model is trained by using the multi-level collaborative alignment loss function and the pre-built database, and retrieval is performed by using the trained model, including: For samples in the database that have no corresponding images but only contain text, the semantic alignment loss within the long text is used to perform semantic alignment within the text; for samples containing image-text, the multi-level collaborative alignment loss function is used for training in combination with the generated large model enhancement information; The model parameters are updated alternately through two sample training methods.

6. A large model-enhanced cross-modal retrieval system, characterized in that: include: An acquisition module, used to acquire image-text pairs; An enhanced information acquisition module is used to acquire large model enhanced information of the image-text pair based on the image-text pair, including using a SAM segmentation model to perform fine-grained semantic segmentation on the image and select semantic entities therein; using a large language model to acquire a summary of the text and extract key semantic fragments; A feature extraction module, for combining the original image, text and enhanced information, using a multi-branch encoder of a pre-trained CLIP model to obtain a plurality of feature vectors, wherein the plurality of feature vectors include image features, text features, fine-grained image entity features and text summary features; The multi-branch encoder construction of the pre-trained CLIP model includes: An adaptation layer is added to the text encoder of the CLIP model and the Transformer structure in the image encoder; the adaptation layer is obtained by downsampling projection and upsampling projection and a residual connection; Based on the encoder of the CLIP model with an adaptation layer, construct feature encoders for different input information; Feature encoders include image encoders, text encoders, instance encoders, and text summary encoders; The loss function design module is used to construct a multi-level collaborative alignment loss function to collaboratively align image and text modalities in a common semantic space, including: The matching images and texts are defined as positive sample pairs, the mismatched ones are defined as negative sample pairs, and the cosine similarity is used as the standard to measure the relationship between sample pairs; The ring loss function is used to increase the similarity between positive sample pairs and reduce the similarity between negative sample pairs; Construct alignment loss based on multiple feature vectors. The alignment loss includes semantic alignment loss between image and text, semantic alignment loss between text and image fine-grained segmentation, semantic alignment loss between image and text summary, and semantic alignment loss within long text. The weighted sum of the alignment losses is used to obtain the multi-level collaborative alignment loss function; The retrieval module is used to train the model through the multi-level collaborative alignment loss function and the pre-constructed training database, train the multi-branch encoder using the gradient descent algorithm, and perform retrieval through the trained model.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the large model enhanced cross-modal retrieval method according to any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the large model enhanced cross-modal retrieval method as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Cross-modal retrieval model and method based on anti-fact reasoning and computer equipment

    CN115146100A

  • Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism

    CN118861327A