Image cross-modal retrieval method and device based on large language model and medium
Through the cross-modal image retrieval method based on large language model, text descriptions of images are generated and image query descriptions are optimized, which solves the problem that traditional image retrieval technology is difficult to capture high-level semantic information, and achieves more efficient and accurate image retrieval and improves user experience.
Patent Information
- Application Number
- CN202510010670.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional image retrieval technology is difficult to effectively capture high-level semantic information, resulting in the search results that are inconsistent with user needs and the target recognition performance is inefficient in complex contexts.
Using a cross-modal image retrieval method based on a large language model, the text description of the image is generated through the BLIP model and converted into an image vector, the mapping relationship between the image and text is established, the image query description is optimized, and the matching degree with the image to be selected is calculated to filter out the target image.
It enhances the system's understanding and description ability of image content, improves the accuracy and user experience of image retrieval, and significantly narrows the semantic gap between text description and image content.
Smart Images

Figure CN119938972A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language models, and in particular to a method, device and medium for cross-modal image retrieval based on a large language model. Background Art
[0002] Image retrieval technology aims to achieve efficient search by analyzing and matching image content. Traditional image retrieval technology mainly relies on basic visual features, such as color, texture, and shape, to compare the similarity between images. However, this method is often inefficient and inaccurate when processing large-scale data sets, because it is difficult to capture high-level semantic information, which may lead to significant deviations between retrieval results and users' actual needs. In addition, traditional methods have limited understanding of objects and scenes in images, especially when facing complex backgrounds, the effectiveness of target recognition is greatly reduced.
[0003] With the advancement of large language model technology, it has gradually been applied to the field of image retrieval. By integrating natural language processing and deep learning technology, it not only enhances the system's ability to understand and describe image content, but also significantly improves the accuracy and stability of cross-modal retrieval. Despite this, image retrieval technology based on large language models still faces many challenges, such as solving the semantic gap between text descriptions and image content, and improving the ability to capture and express subtle features of images. The existence of these problems not only restricts the overall performance of the image retrieval system, but also affects the user experience to a certain extent. Summary of the invention
[0004] In order to solve the above problems, this application proposes an image cross-modal retrieval method based on a large language model, including:
[0005] Generate a text description corresponding to the original image through the BLIP model, and convert the text description into a corresponding image vector;
[0006] Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database;
[0007] Obtaining an image query description submitted by a user, optimizing the image query description through a preset language model, and converting the optimized image query description into a corresponding query vector;
[0008] Determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector;
[0009] The matching degree between the image query description and the plurality of candidate images is calculated, and according to the matching degree, a target image satisfying the image query description is screened out from the plurality of candidate images.
[0010] In one implementation of the present application, a text description corresponding to the original image is generated by using a BLIP model, specifically including:
[0011] The BLIP model consists of a visual encoder and a text decoder;
[0012] Extracting image features from the original image through the visual encoder;
[0013] Based on the attention mechanism, the image features are mapped to the corresponding text description space through the text decoder to generate a text description corresponding to the original image.
[0014] In an implementation of the present application, determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector specifically includes:
[0015] According to the first similarity between each image vector in the vector database and the query vector, a specified image vector whose corresponding first similarity is greater than a preset threshold is selected from the vector database;
[0016] According to the specified image vector, a number of candidate images corresponding to the image query description are determined.
[0017] In one implementation of the present application, before selecting from the vector database a designated image vector whose corresponding first similarity is greater than a preset threshold, the method further includes:
[0018] Obtaining a text description corresponding to the image vector, and matching the text description with the image query description to determine matching description words in the text description and the image query description;
[0019] Determining the importance of the descriptive word in text description according to the feature type corresponding to the descriptive word, and determining the correction coefficient corresponding to the first similarity according to the importance;
[0020] Based on the correction coefficient, the first similarity between the image vector and the query vector is corrected to obtain a corrected first similarity.
[0021] In one implementation of the present application, determining the importance of the descriptive word in text description according to the feature type corresponding to the descriptive word specifically includes:
[0022] Determine the feature type corresponding to the descriptor; wherein the feature type includes a functional description feature and a key description feature;
[0023] The importance of the descriptive word in text description is determined according to the association between the feature type and the preset importance; wherein the importance corresponding to the key description feature is greater than the importance corresponding to the functional description feature.
[0024] In an implementation of the present application, determining a correction coefficient corresponding to the first similarity according to the importance specifically includes:
[0025] For different feature types, determining the ratio of the descriptive words corresponding to the feature types in all the descriptive words;
[0026] The product of the ratio and the importance is calculated, and the product is used as a correction coefficient corresponding to the first similarity.
[0027] In one implementation of the present application, calculating the degree of matching between the image query description and the plurality of candidate images specifically includes:
[0028] Preprocessing the plurality of images to be selected to convert the plurality of images to be selected into processed images with specified pixels and specified image formats;
[0029] Encoding the processed image and the image query description respectively through a visual encoder in the CLIP model to convert the processed image and the image query description into corresponding target vectors respectively;
[0030] A second similarity between the processed image and the target vectors respectively corresponding to the image query description is calculated, and the second similarity is used as a matching degree between the image query description and the plurality of candidate images.
[0031] In one implementation of the present application, determining a number of candidate images corresponding to the image query description according to the specified image vector specifically includes:
[0032] Obtaining a specified mapping relationship corresponding to the specified image vector from the vector database;
[0033] Based on the image retrieval path corresponding to the specified mapping relationship, a number of candidate images corresponding to the image query description are obtained.
[0034] The present application embodiment provides an image cross-modal retrieval device based on a large language model, the device comprising:
[0035] at least one processor;
[0036] and, a memory communicatively coupled to the at least one processor;
[0037] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:
[0038] Generate a text description corresponding to the original image through the BLIP model, and convert the text description into a corresponding image vector;
[0039] Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database;
[0040] Obtaining an image query description submitted by a user, optimizing the image query description through a preset language model, and converting the optimized image query description into a corresponding query vector;
[0041] Determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector;
[0042] The matching degree between the image query description and the plurality of candidate images is calculated, and according to the matching degree, a target image satisfying the image query description is screened out from the plurality of candidate images.
[0043] The embodiment of the present application provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as follows:
[0044] Generate a text description corresponding to the original image through the BLIP model, and convert the text description into a corresponding image vector;
[0045] Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database;
[0046] Obtaining an image query description submitted by a user, optimizing the image query description through a preset language model, and converting the optimized image query description into a corresponding query vector;
[0047] Determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector;
[0048] The matching degree between the image query description and the plurality of candidate images is calculated, and according to the matching degree, a target image satisfying the image query description is screened out from the plurality of candidate images.
[0049] The image cross-modal retrieval method based on a large language model proposed in this application can bring the following beneficial effects:
[0050] Generating an image text description of the original image and converting it into an image vector effectively builds a bridge between image and text, enhancing the system's ability to understand and describe image content. This not only improves the accuracy of image retrieval, but also makes the retrieval results more in line with the actual needs of users.
[0051] Secondly, the large language model is used to optimize image query descriptions, further narrowing the semantic gap between text descriptions and image content. By accurately capturing user intent, relevant images can be returned more accurately, significantly improving the stability of cross-modal retrieval and user experience.
[0052] Finally, deep matching is performed between the selected images and their corresponding image query descriptions, and the consistency of comprehensive visual features and semantic information is evaluated, which helps to further improve search accuracy and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0054] Figure 1 A schematic diagram of a flow chart of a method for cross-modal image retrieval based on a large language model provided in an embodiment of the present application;
[0055] Figure 2 A schematic diagram of the structure of a large language model-based image cross-modal retrieval device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0057] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0058] like Figure 1 As shown, an image cross-modal retrieval method based on a large language model provided in an embodiment of the present application includes:
[0059] S101: Generate a text description corresponding to the original image through the BLIP model, and convert the text description into a corresponding image vector.
[0060] The BLIP model is trained to understand and generate the correspondence between images and text. The original image is input into the BLIP model, which analyzes the content of the image and generates one or more text descriptions based on the key information and features in the image. The bge-m3 model can capture the semantic information in the text through a multi-layer neural network structure and convert the text description into a dense vector form. Therefore, after generating the text description, the text can be mapped to a vector representation in a high-dimensional space through this efficient embedded vector model. During the mapping process, the text description will first be segmented. Each word obtained after segmentation will be converted into a vector through the embedding layer in the bge-m3 model, and then abstracted through multiple layers of neural networks. Finally, the model will output an image vector that integrates the entire text semantics. The image vector is equivalent to the digital fingerprint of the original image, which can effectively represent the semantic information of the image and support subsequent retrieval operations.
[0061] The BLIP model adopts a multimodal hybrid architecture based on encoder-decoder, which consists of a visual encoder and a text decoder. It can flexibly process image and text data and support both understanding and generation tasks. Among them, the visual encoder is a Transformer-based ViT architecture, which is used to extract high-level semantic features of the image. It divides the input image into multiple patches and encodes these patches into a series of image embeddings. Therefore, the image features in the input original image can be extracted through the visual encoder. The image-based text decoder is used to generate text descriptions conditioned on the image. It replaces Bi Self-Attention with Casual Self-Attention to achieve the generation task. After receiving the image features output by the visual encoder, the text decoder maps the image features to the corresponding text description space, and then generates the text description corresponding to the original image. Assuming there is an image containing a cat, the visual encoder processes the image into multiple patches, encodes these patches into a series of image embeddings, and extracts the image features. After receiving the image features output by the visual encoder, the text decoder maps these features to the corresponding text description space. Then, a text description corresponding to the image is generated based on these features, such as "This is a cute cat with soft fur and bright eyes."
[0062] It should be noted that before generating text descriptions for the original images, in order to ensure the accuracy of the description words and improve the efficiency of generating text descriptions, each original image in the dataset needs to be preprocessed to ensure that they meet the requirements of subsequent processing. Preprocessing includes but is not limited to resizing (for example, unifying to 256x256 pixels), format conversion (such as conversion to JPEG format), and possible color mode adjustment (such as conversion from RGB to grayscale). In addition, some advanced processing, such as noise removal and contrast enhancement, can also be performed to improve image quality.
[0063] S102: Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database.
[0064] After the original image has its corresponding image vector, it is necessary to propose a mapping relationship between the image vector and the original image. The mapping relationship exists in the form of a key-value pair, where the key is the unique identifier of the image vector (such as the image ID or hash value), and the value is the image retrieval path of the original image. The corresponding image retrieval path is recorded for each vector in the database, so that the corresponding image can be found efficiently and accurately in the vector retrieval. Both the image vector and the established mapping relationship need to be stored in the preset milvus vector database. Milvus is an open source vector search engine suitable for the management and retrieval of large-scale vector data. In milvus, image vectors are organized by index structures to quickly perform similarity search operations.
[0065] S103: Obtain the image query description submitted by the user, optimize the image query description through a preset language model, and convert the optimized image query description into a corresponding query vector.
[0066] When a user submits an image query description for the desired image, the qwen2-7b-chat large language model is used to understand and optimize the description, converting it into a more precise and easy-to-process form to enhance its expressiveness and accuracy. For example, the Chinese description "Help me find a picture of a girl holding a cat, she is sitting on the sofa" is converted into a description of the picture and then converted into English: "A girl is sitting on a sofa, holding a kitten in her arms." This is because the BLIP model and bge-m3 model used later are both based on English text. The optimized text is then converted into a query vector through the bge-m3 model, ready for the subsequent retrieval process.
[0067] S104: Determine a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector.
[0068] In the vector database, image vectors of a large number of images are stored. In order to find images similar to the query image, it is necessary to calculate the first similarity between the query vector and each image vector in the database. Among them, the calculation of the first similarity can be carried out by means of cosine similarity or Euclidean distance. According to the calculated first similarity, the image vectors in the database can be sorted. At this time, it is necessary to select the first N images corresponding to the specified image vectors whose first similarity is greater than the preset threshold as the selected images. Here, N can be adjusted according to actual needs to balance the accuracy and efficiency of the retrieval. The preset threshold can also be set according to the actual retrieval needs. If the number of similar images retrieved by matching is large, in order to improve the accuracy of the screening results, it is necessary to expand the image range of the selected images as much as possible. Therefore, the preset threshold can be set relatively smaller to screen out a larger number of selected images as candidates.
[0069] When screening the candidate images, it is first necessary to obtain the specified mapping relationship corresponding to the specified image vector from the vector database, and then obtain several candidate images corresponding to the image query description based on the image retrieval path corresponding to the specified mapping relationship.
[0070] In one embodiment, the first similarity is calculated only based on the text similarity between the image vector and the query vector. However, when selecting images for selection, text similarity ignores the importance of descriptive words in retrieving images. For example, the image query description is "Please help me select two images of cats with yellow hair color." For this image query description, two, hair color, yellow, and cat are all key descriptive words with clear feature orientation, while please help me, select, and so on are functional descriptive words that form a complete sentence structure. Therefore, when selecting images for selection based on the similarity between vectors, the feature type of each descriptive word must also be considered.
[0071] Specifically, obtain the text description corresponding to the image vector, match the text description with the image query description, and determine the matching descriptive words in the text description and the image query description. It should be noted that the descriptive words screened out at this time are not only words with completely consistent descriptions, but also approximate words with similar semantics. According to the feature type corresponding to the descriptive word, determine the importance of the descriptive word when performing text description. Feature types include functional description features and key description features. After determining the feature type corresponding to each descriptive word, the importance of the descriptive word when performing text description can be determined based on the correlation between the feature type and the preset importance. Among them, the importance corresponding to the key description feature is greater than the importance corresponding to the functional description feature.
[0072] According to the importance, a correction coefficient corresponding to the first similarity is determined. The correction coefficient is a value between 0 and 1, which is used to indicate the influence of the descriptor on the overall similarity. The higher the importance of the descriptor, the larger the correction coefficient should be. Based on the correction coefficient, the first similarity between the image vector and the query vector is corrected to obtain the corrected first similarity.
[0073] When calculating the correction coefficient, the ratio of feature types and the importance of the descriptor are combined to determine the correction coefficient of the first similarity. Specifically, the ratio of descriptors of different feature types in all descriptors is counted, and the importance of the descriptor is multiplied by the ratio of the feature type to which it belongs to obtain the correction coefficient of the descriptor. Similarity correction is achieved by multiplying the correction coefficient by the first similarity. The corrected first similarity will more accurately reflect the correlation between the image and the query, while taking into account the influence of feature types and the importance of descriptors.
[0074] S105: Calculate the matching degree between the image query description and the plurality of candidate images, and select a target image that meets the image query description from the plurality of candidate images according to the matching degree.
[0075] After screening out several candidate images, the CLIP model is needed to further evaluate the degree of match between the image query description and the candidate images, and then screen out the best match that meets the user's retrieval needs, namely the target image.
[0076] The CLIP model also consists of two parts: a visual encoder and a text encoder. It learns the complex relationship between images and texts through pre-training, conducts in-depth analysis of each candidate image and the user's query image description, and comprehensively considers the consistency of visual features and semantic information. It can accurately locate the image that best meets user needs, which helps to further improve search accuracy and user experience.
[0077] The CLIP model has zero-sample transfer capability, that is, it can accurately match images and texts without specific domain training. Therefore, the model can achieve accurate matching between text descriptions and candidate images. First, several candidate images need to be preprocessed to convert them into processed images with specified pixels and specified image formats. The specified pixels here are generally 224x224 pixels, and the specified image format is RGB format. Then, the processed image and image query description are encoded respectively through the visual encoder in the CLIP model to convert the processed image and image query description into corresponding target vectors. Finally, the second similarity between the target vectors corresponding to the processed image and the image query description is calculated, and the second similarity is used as the matching degree between the image query description and several candidate images. After obtaining the matching degree between the image query description and several candidate images, the candidate image with the largest corresponding matching degree is selected as the target image finally retrieved by the user.
[0078] The above are embodiments of the method proposed in this application. Based on the same idea, some embodiments of this application also provide devices and non-volatile computer storage media corresponding to the above methods.
[0079] Figure 2 A schematic diagram of the structure of a large language model-based image cross-modal retrieval device provided in an embodiment of the present application. Figure 2 As shown, including:
[0080] at least one processor; and,
[0081] at least one processor is communicatively connected to a memory; wherein,
[0082] The memory stores instructions executable by at least one processor, the instructions being executed by at least one processor to enable the at least one processor to:
[0083] Generate the text description corresponding to the original image through the BLIP model, and convert the text description into the corresponding image vector;
[0084] Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database;
[0085] Obtain the image query description submitted by the user, optimize the image query description through the preset language model, and convert the optimized image query description into a corresponding query vector;
[0086] Determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector;
[0087] The matching degree between the image query description and several candidate images is calculated, and according to the matching degree, the target image that meets the image query description is screened out from the several candidate images.
[0088] The embodiment of the present application provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as follows:
[0089] Generate the text description corresponding to the original image through the BLIP model, and convert the text description into the corresponding image vector;
[0090] Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database;
[0091] Obtain the image query description submitted by the user, optimize the image query description through the preset language model, and convert the optimized image query description into a corresponding query vector;
[0092] Determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector;
[0093] The matching degree between the image query description and several candidate images is calculated, and the target image that meets the image query description is screened out from the several candidate images according to the matching degree.
[0094] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0095] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0096] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0097] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0100] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0101] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0102] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0103] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0104] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A cross-modal image retrieval method based on a large language model, characterized in that: The method comprises: Generate a text description corresponding to the original image through the BLIP model, and convert the text description into a corresponding image vector; Establishing a mapping relationship between the image vector and the original image, and storing the mapping relationship and the image vector in a preset vector database; Obtaining an image query description submitted by a user, optimizing the image query description through a preset language model, and converting the optimized image query description into a corresponding query vector; Determining a number of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector; The matching degree between the image query description and the plurality of candidate images is calculated, and according to the matching degree, a target image satisfying the image query description is screened out from the plurality of candidate images.
2. The image cross-modal retrieval method based on a large language model according to claim 1, characterized in that: Generate the text description corresponding to the original image through the BLIP model, including: The BLIP model consists of a visual encoder and a text decoder; Extracting image features from the original image through the visual encoder; Based on the attention mechanism, the image features are mapped to the corresponding text description space through the text decoder to generate a text description corresponding to the original image.
3. The image cross-modal retrieval method based on a large language model according to claim 1, characterized in that: Determining a plurality of candidate images corresponding to the image query description according to a first similarity between each image vector in the vector database and the query vector, specifically includes: According to the first similarity between each image vector in the vector database and the query vector, a specified image vector whose corresponding first similarity is greater than a preset threshold is selected from the vector database; According to the specified image vector, a number of candidate images corresponding to the image query description are determined.
4. The image cross-modal retrieval method based on a large language model according to claim 3, characterized in that: Before selecting the designated image vector whose corresponding first similarity is greater than a preset threshold from the vector database, the method further includes: Obtaining a text description corresponding to the image vector, and matching the text description with the image query description to determine matching description words in the text description and the image query description; Determining the importance of the descriptive word in text description according to the feature type corresponding to the descriptive word, and determining the correction coefficient corresponding to the first similarity according to the importance; Based on the correction coefficient, the first similarity between the image vector and the query vector is corrected to obtain a corrected first similarity.
5. The image cross-modal retrieval method based on a large language model according to claim 4, characterized in that: Determining the importance of the descriptive word in text description according to the feature type corresponding to the descriptive word specifically includes: Determine the feature type corresponding to the descriptor; wherein the feature type includes a functional description feature and a key description feature; The importance of the descriptive word in text description is determined according to the association between the feature type and the preset importance; wherein the importance corresponding to the key description feature is greater than the importance corresponding to the functional description feature.
6. The image cross-modal retrieval method based on a large language model according to claim 4, characterized in that: Determining a correction coefficient corresponding to the first similarity according to the importance specifically includes: For different feature types, determining the ratio of the descriptive words corresponding to the feature types in all the descriptive words; The product of the ratio and the importance is calculated, and the product is used as a correction coefficient corresponding to the first similarity.
7. The image cross-modal retrieval method based on a large language model according to claim 1, characterized in that: Calculating the matching degree between the image query description and the plurality of candidate images specifically includes: Preprocessing the plurality of images to be selected to convert the plurality of images to be selected into processed images with specified pixels and specified image formats; Encoding the processed image and the image query description respectively through a visual encoder in the CLIP model to convert the processed image and the image query description into corresponding target vectors respectively; A second similarity between the processed image and the target vectors respectively corresponding to the image query description is calculated, and the second similarity is used as a matching degree between the image query description and the plurality of candidate images.
8. The image cross-modal retrieval method based on a large language model according to claim 3, characterized in that: Determining a number of candidate images corresponding to the image query description according to the specified image vector specifically includes: Obtaining a specified mapping relationship corresponding to the specified image vector from the vector database; Based on the image retrieval path corresponding to the specified mapping relationship, a number of candidate images corresponding to the image query description are obtained.
9. An image cross-modal retrieval device based on a large language model, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image cross-modal retrieval method based on a large language model as described in any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: A cross-modal image retrieval method based on a large language model as described in any one of claims 1 to 8.
Citation Information
Cited By
Geographic positioning method and device based on multi-modal large model, equipment and medium
CN120256489A
Geolocation method and apparatus based on multi-modal large model, device and medium
CN120256489B
Image query method and device, computing equipment, storage medium and program product
CN120541254A
Urban management case image recognition method and system based on multi-modal large model
CN121074769A
Image retrieval method and device, equipment and storage medium
CN121350296A