Image search engine generation method and image retrieval method and system
Through the visual large language model, the image is converted into text descriptions and an image search engine is generated, which solves the problem of insufficient image retrieval accuracy and response time in the prior art, and realizes accurate recognition and real-time retrieval of various object categories and scenes.
Patent Information
- Application Number
- CN202411950903.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
The existing image retrieval technology is based on embedded representation, and there are insufficient search capabilities for general object types and semantic requests, low matching accuracy and long response time, making it difficult to meet the needs of real-time application scenarios.
The Visual Large Language Model (VLM) is used to convert images into text descriptions, and an image search engine is generated through indexing and sorting, supporting real-time retrieval and large-scale concurrent request processing.
It realizes accurate recognition of various object categories and scenes, without the need for additional image model training, improves retrieval accuracy and response speed, and is suitable for real-time and large-scale application scenarios.
Smart Images

Figure CN119988665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image retrieval and mining technology, and in particular to a method for generating an image search engine, an image retrieval method and a system. Background Art
[0002] Current image retrieval is mainly based on embedded representations of images; the embedded representations of all images in the index image library are indexed, and when a user searches through text, the text request is mapped into the embedded representation, and then matched with the image in the index library; this method has insufficient retrieval capabilities for general object types and insufficient retrieval capabilities for semantic requests, that is, image retrieval based on embedded representations of images has the problem of insufficient matching accuracy; in addition, since the embedded representation has a large amount of computational complexity during matching, the request response time is long, which is not conducive to real-time application scenarios. Summary of the invention
[0003] In view of the above problems, the first object of the present invention is to provide a method for generating an image search engine, which converts images into accurate text descriptions based on a visual language model (VLM), and then indexes and sorts the text descriptions to generate an image search engine based on the VLM. The image search engine can be used to explore new object categories and new scenes without the need for additional image model training.
[0004] The second object of the present invention is to provide a system for generating the above-mentioned image search engine.
[0005] The third object of the present invention is to provide an image retrieval method, which implements real-time retrieval based on the above-mentioned image search engine, can return sorting results within milliseconds, and supports large-scale concurrent request processing.
[0006] The fourth object of the present invention is to provide the above-mentioned image retrieval system.
[0007] The first technical solution adopted by the present invention is: a method for generating an image search engine, comprising the following steps:
[0008] S100: Build and train a visual large language model to obtain a trained visual large language model;
[0009] S200: construct an image library, and obtain information text corresponding to each image in the image library;
[0010] S300: Input each image and its corresponding information text into the trained visual language model to obtain a text description corresponding to each image;
[0011] S400: Segmenting the text description and extracting keywords corresponding to each image;
[0012] S500: Create an inverted index for the keywords corresponding to each image and write it into a search engine to generate an image search engine based on a visual large language model.
[0013] Preferably, the step S100 includes the following sub-steps:
[0014] S110: Acquire visual question answering data, and generate a training set based on the visual question answering data;
[0015] S120: constructing a visual large language model based on a multimodal Transformer; and pre-training the visual large language model based on the training set to obtain a pre-trained visual large language model;
[0016] S130: Acquire domain-specific visual question-answering data, and fine-tune the pre-trained visual large language model based on the domain-specific visual question-answering data to obtain a trained visual large language model.
[0017] Preferably, the visual question answering data includes an original image in an open field and a number of question answering text data corresponding to the original image;
[0018] The domain-specific visual question-answering data includes domain-specific images and a set of question-answering text data corresponding to the images.
[0019] Preferably, the visual large language model comprises a visual feature extractor, a fully connected layer and a large language model connected in sequence;
[0020] The pre-training of the visual large language model based on the training set includes:
[0021] The original images in the training set and the question-and-answer text data corresponding to the original images are input into the visual large language model, and the original images are processed through a visual feature extractor to obtain block-by-block image features; the block-by-block image features are mapped into text space vectors through a fully connected layer; the text space vectors and the question-and-answer text data corresponding to the original images are input into the large language model for pre-training.
[0022] Preferably, obtaining the information text corresponding to each image in the image library in step S200 includes:
[0023] Describe the overall scene, key vehicles or moving objects, and adverse conditions for each image in the image library, so as to obtain the information text corresponding to each image;
[0024] The text description includes one or more of scene information, main objects in the image and their attributes, relationships between objects, and potential semantic information.
[0025] Preferably, the step S400 includes:
[0026] The text description is decomposed into several fine-grained words based on a word segmenter; and the weight of each word for the text description is obtained based on TF-IDF, and several words with weights greater than a set threshold are used as keywords of the text description, that is, keywords of the image corresponding to the text description are obtained.
[0027] The second technical solution adopted by the present invention is: a system for generating an image search engine, comprising a visual large language model building module and an image search engine generating module;
[0028] The visual large language model construction module is used to construct and train a visual large language model to obtain a trained visual large language model;
[0029] The image search engine generation module is used to: build an image library and obtain the information text of each image in the image library; input each image and its corresponding information text into the trained visual large language model to obtain the text description corresponding to each image; segment the text description and extract the keywords corresponding to each image; and create an inverted index for the keywords corresponding to each image and write them into the search engine to generate an image search engine based on the visual large language model.
[0030] The third technical solution adopted by the present invention is: an image retrieval method, comprising the following steps:
[0031] S10: Obtaining a user search request;
[0032] S20: parsing the user search request to obtain keywords;
[0033] S30: Input the keyword into the image search engine generated by the method in the first technical solution to perform a search and obtain a search result.
[0034] Preferably, the image retrieval method further comprises step S40: presenting the retrieval results to the user, and the user further filters, sorts or modifies the query based on the retrieval results.
[0035] The fourth technical solution adopted by the present invention is: an image retrieval system, comprising a search request acquisition module, a parsing module and a retrieval module;
[0036] The search request acquisition module is used to acquire user search requests;
[0037] The parsing module is used to parse the user search request to obtain keywords;
[0038] The retrieval module is used to input the keywords into the image search engine generated by the method in the first technical solution to perform retrieval and obtain retrieval results.
[0039] Beneficial effects of the above technical solution:
[0040] (1) An image search engine provided by the present invention indexes images in various fields and scenes (including but not limited to autonomous driving) through a visual language model (VLM), and can identify various object categories, rather than being limited to dozens of pre-set categories. This design is directly applicable to mining new object categories (i.e., recognition of new target types) and new scenes without the need for additional image model training.
[0041] (2) By introducing a large language model (LLM) in the sorting stage, the present invention can not only perform in-depth semantic understanding of user requests, but also dynamically adjust query results. For example, LLM can accurately understand the semantic logic implicit in complex natural language requests (such as "find small objects against a white background") and perform semantic matching scores on retrieval results. In addition, it can filter out misunderstood or low-relevant results, significantly improving retrieval accuracy and user experience. At the same time, the image search engine supports efficient processing of multi-task requests (such as querying multiple scenes and object combinations at the same time), greatly enhancing the flexibility and generalization ability of the system.
[0042] (3) By segmenting the text description of each image and establishing an efficient inverted index, the image retrieval system can quickly match keywords or semantic features in user requests in a massive image library. The real-time retrieval algorithm based on the image search engine can return sorting results within milliseconds and support large-scale concurrent request processing. In order to further improve efficiency, the image retrieval system also adopts a distributed architecture and caching mechanism to optimize high-frequency queries and commonly used indexes. In addition, combined with a domain-related feature weighting strategy, it can prioritize the retrieval of results that meet the needs of specific scenarios, thereby ensuring response speed while maintaining high retrieval quality and domain adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of a flow chart of a method for generating an image search engine provided by an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of the structure of a system for generating an image search engine provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following detailed description of the embodiments of the present invention is further described in detail in conjunction with the accompanying drawings and examples. The detailed description of the following embodiments and the accompanying drawings are used to exemplarily illustrate the principles of the present invention, but cannot be used to limit the scope of the present invention, that is, the present invention is not limited to the preferred embodiments described, and the scope of the present invention is defined by the claims.
[0046] In the description of the present invention, it should be noted that, unless otherwise specified, “plurality” means two or more than two; the terms “first”, “second”, etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance; for ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0047] Embodiment 1
[0048] like Figure 1 As shown, an embodiment of the present invention provides a method for generating an image search engine, comprising the following steps:
[0049] S100: Build and train a visual large language model to obtain a trained visual large language model;
[0050] Building a visual language model (VLM) and training it includes the following sub-steps:
[0051] S110: Acquire visual question answering data, and generate a training set based on the visual question answering data;
[0052] Collect visual question-answering data from public datasets, including common open-domain data (such as COCO-VQA, GQA, etc.); the visual question-answering data includes one or more original images in the open domain, and a number of question-answering text data corresponding to the original images (i.e., question text data and answer text data); the number of question-answering text data corresponding to the original images includes multi-round dialogue data, for example: ① Question: What is the weather shown in the picture? Answer: It is a sunny day; ② Question: Are there any road signs in the picture? Answer: There is a license plate with a speed limit of 60 kilometers on the right side of the road; ③ Which lane is the gray car in? Answer: In the first lane from the left, it has just entered the lane.
[0053] A training set is generated based on the visual question answering data.
[0054] S120: constructing a visual large language model based on a multimodal Transformer; and pre-training the visual large language model based on the training set to obtain a pre-trained visual large language model;
[0055] A visual large language model based on a multimodal Transformer architecture is constructed, wherein the visual large language model includes a visual feature extractor (e.g., CLIP), a fully connected layer, and a large language model (LLM) connected in sequence; the visual feature extractor and the large language model both include a Transformer architecture.
[0056] Pre-training the visual large language model based on the training set includes:
[0057] The original image in the training set and a number of question and answer text data corresponding to the original image are input into the visual large language model, and the original image is processed through a visual feature extractor to obtain block-by-block image features; and the block-by-block image features are mapped into text space vectors through a fully connected layer; the text space vectors and a number of question and answer text data corresponding to the original image are input into a large language model (LLM) for pre-training to learn the correspondence between the original image and its corresponding several questions and answers (i.e., the text description of the original image).
[0058] During pre-training, the parameters of the large language model are first fixed, and the fully connected layer mapping from the original image to the text space vector is trained. At the same time, the large language model (LLM) is trained based on the text space vector and several question and answer text data corresponding to the original image.
[0059] S130: Acquire domain-specific visual question-answering data, and fine-tune the pre-trained visual large language model based on the domain-specific visual question-answering data to obtain a trained visual large language model;
[0060] Acquiring domain-specific visual question-and-answer data includes: collecting relevant domain-specific visual question-and-answer data according to application fields (such as medical imaging, industrial inspection, etc.); the domain-specific visual question-and-answer data includes a domain-specific image and a question-and-answer text data set corresponding to the image.
[0061] Fine-tuning the pre-trained visual large language model based on the domain-specific visual question-answering data includes:
[0062] The visual question and answer data are mixed with the domain-specific visual question and answer data in a certain proportion to obtain a fine-tuning training set; when training the visual large language model, the large language model (LLM) is fine-tuned based on a number of question and answer text data corresponding to the original pictures in the fine-tuning training set to obtain a trained visual large language model.
[0063] By fine-tuning the pre-trained visual big language model, it is ensured that the visual big language model has domain adaptability; for example, the pre-trained visual big language model is fine-tuned by using visual question-answering data exclusive to the field of autonomous driving to optimize the performance of the visual big language model for specific tasks; during the fine-tuning process, a contrastive learning strategy can be used to enhance the visual big language model's understanding of the image-text matching relationship.
[0064] Furthermore, in one embodiment, it also includes: performing performance verification on the trained visual large language model, specifically: evaluating the question-answering accuracy and the quality of generated description of the trained visual large language model on the domain data through an independent verification set.
[0065] S200: construct an image library, and obtain information text corresponding to each image in the image library;
[0066] Images of the corresponding field are collected according to the needs of the field, and an image library is built based on the images of the corresponding field. For example, in the field of autonomous driving, the images of the on-board visual camera can be sampled evenly along time, or the road sections and scenes can be sampled purposefully and focused to obtain images of the autonomous driving field, and an image library is built based on the images of the autonomous driving field.
[0067] Obtaining the information text (i.e., question text) corresponding to each image in the image library includes but is not limited to: performing an overall scene description, a key vehicle or moving object description, and an adverse condition description for each image in the image library, so as to obtain the information text corresponding to each image; the overall scene description includes describing the scene of the vehicle driving on the road, including the surrounding environment, weather, road conditions, etc., for example, not more than 150 words; the key vehicle or moving object description includes listing all vehicles or moving objects that may affect driving decisions, especially nearby objects, and describing them in detail, including location, moving direction, speed, etc.; the adverse condition description includes describing any special circumstances, such as strong light reflection, bad weather, or bad road conditions. If there are no special circumstances, this part is omitted.
[0068] S300: Input each image and its corresponding information text into the trained visual language model to obtain a text description corresponding to each image;
[0069] The text description includes, but is not limited to: scene information, main objects in the image and their attributes, relationships between objects, and potential semantic information; the scene information is, for example, a description of weather, lighting and other conditions; the main objects in the image and their attributes are, for example, the color, shape, absolute and relative position of the vehicle; the relationship between objects is, for example, a red car is driving on the left hand side; the potential semantic information is, for example, "the car in front wants to merge" or "the car in front is very close."
[0070] S400: Segmenting the text description and extracting keywords corresponding to each image;
[0071] The text description is segmented and the keywords corresponding to each image are extracted, including:
[0072] A dynamic programming-based or open-source tokenizer decomposes the text description into several fine-grained words to enhance the retrieval performance of subsequent search engines; for example, after tokenizing the sentence "We are driving on the road, this image is taken from our dashcam", the tokens obtained are: We are driving on the road, this image is taken from our dashcam;
[0073] The weight of each word for the text description is obtained based on TF-IDF, and several words with weights greater than a set threshold are used as keywords of the text description, that is, the keywords of the image corresponding to the text description are obtained.
[0074] S500: creating an inverted index for the keywords corresponding to each image and writing it into a search engine to generate an image search engine based on a visual large language model;
[0075] Inverted index construction: Create an inverted index for the keywords corresponding to each image; each keyword in the inverted index points to the image set related to it, and records the weight of the keyword in the text description corresponding to each image;
[0076] The inverted index is stored in an efficient text search engine (such as ElasticSearch) to generate an image search engine based on a visual large language model.
[0077] Embodiment 2
[0078] like Figure 2 As shown, an embodiment of the present invention provides a system for generating an image search engine, including a visual large language model building module and an image search engine generating module;
[0079] The visual large language model construction module is used to construct and train a visual large language model to obtain a trained visual large language model; the visual large language model construction module performs the following operations:
[0080] S110: Acquire visual question answering data, and generate a training set based on the visual question answering data;
[0081] S120: constructing a visual large language model based on a multimodal Transformer; and pre-training the visual large language model based on the training set to obtain a pre-trained visual large language model;
[0082] S130: Acquire domain-specific visual question-answering data, and fine-tune the pre-trained visual large language model based on the domain-specific visual question-answering data to obtain a trained visual large language model.
[0083] The image search engine generation module is used to: build an image library and obtain the information text of each image in the image library; input each image and its corresponding information text into the trained visual large language model to obtain the text description corresponding to each image; segment the text description and extract the keywords corresponding to each image; and create an inverted index for the keywords corresponding to each image and write them into the search engine to generate an image search engine based on the visual large language model.
[0084] Embodiment 3
[0085] An embodiment of the present invention provides an image retrieval method, comprising the following steps:
[0086] S10: Obtaining a user search request;
[0087] Users enter search requests through natural language or send structured search requests through the API; user search requests can be single questions or complex combinations of multiple questions.
[0088] S20: parsing the user search request to obtain keywords;
[0089] Parsing user search requests includes:
[0090] Based on the word segmenter, the user search request is decomposed into several fine-grained words, and the weight of each word for the user search request is obtained based on TF-IDF, and several words with weights greater than the set threshold are used as keywords for the user search request.
[0091] S30: inputting the keyword into the image search engine for searching and obtaining search results;
[0092] The keywords obtained by parsing the user's search request are input into the image search engine based on the visual large language model generated in Example 1, and a quick search is performed and the search results are output. The search results are the most relevant image set, and the most relevant image set is sorted from high to low according to the weight of the keyword in each image, and the text description and keyword weight information matched by each image are attached.
[0093] Furthermore, in one embodiment, the method further includes: displaying the search results to the user, and supporting the user to further filter, sort or modify the search results, thereby forming an iterative search process.
[0094] Furthermore, in one embodiment, it also includes: performing data mining based on the image search engine based on the visual large language model, specifically: for complex analysis needs, correlation analysis can be performed on the retrieval results of multiple queries by users, potential image collections can be mined, and high-level reports or statistical information can be generated.
[0095] Embodiment 4
[0096] An embodiment of the present invention provides an image retrieval system, comprising a search request acquisition module, a parsing module and a retrieval module;
[0097] The search request acquisition module is used to acquire user search requests;
[0098] The parsing module is used to parse the user search request to obtain keywords;
[0099] The retrieval module is used to input the keyword into the image search engine for retrieval to obtain retrieval results.
[0100] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0101] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0102] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0103] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0104] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical disks.
[0105] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for generating an image search engine, characterized in that: The following steps are involved: S100: construct and train a visual large language model to obtain a trained visual large language model; S200: construct an image library, and obtain information text corresponding to each image in the image library; S300: Input each image and its corresponding information text into the trained visual language model to obtain a text description corresponding to each image; S400: Segmenting the text description and extracting keywords corresponding to each image; S500: Create an inverted index for the keywords corresponding to each image and write it into a search engine to generate an image search engine based on a visual large language model.
2. The method for generating an image search engine according to claim 1, characterized in that: The step S100 includes the following sub-steps: S110: Acquire visual question answering data, and generate a training set based on the visual question answering data; S120: constructing a visual large language model based on a multimodal Transformer; and pre-training the visual large language model based on the training set to obtain a pre-trained visual large language model; S130: Acquire domain-specific visual question-answering data, and fine-tune the pre-trained visual large language model based on the domain-specific visual question-answering data to obtain a trained visual large language model.
3. The method for generating an image search engine according to claim 2, characterized in that: The visual question answering data includes an original image in an open field and a number of question answering text data corresponding to the original image; The domain-specific visual question-answering data includes domain-specific images and a set of question-answering text data corresponding to the images.
4. The method for generating an image search engine according to claim 2, characterized in that: The visual large language model includes a visual feature extractor, a fully connected layer and a large language model connected in sequence; The pre-training of the visual large language model based on the training set includes: The original images in the training set and the question-and-answer text data corresponding to the original images are input into the visual large language model, and the original images are processed through a visual feature extractor to obtain block-by-block image features; the block-by-block image features are mapped into text space vectors through a fully connected layer; the text space vectors and the question-and-answer text data corresponding to the original images are input into the large language model for pre-training.
5. The method for generating an image search engine according to claim 1, characterized in that: The step S200 of obtaining the information text corresponding to each image in the image library includes: Describe the overall scene, key vehicles or moving objects, and adverse conditions for each image in the image library, so as to obtain the information text corresponding to each image; The text description includes one or more of scene information, main objects in the image and their attributes, relationships between objects, and potential semantic information.
6. The method for generating an image search engine according to claim 1, characterized in that: The step S400 includes: The text description is decomposed into several fine-grained words based on a word segmenter; and the weight of each word for the text description is obtained based on TF-IDF, and several words with weights greater than a set threshold are used as keywords of the text description, that is, keywords of the image corresponding to the text description are obtained.
7. A system for generating an image search engine, characterized in that: It includes a visual large language model building module and an image search engine generation module; The visual large language model construction module is used to construct and train a visual large language model to obtain a trained visual large language model; The image search engine generation module is used to: build an image library and obtain the information text of each image in the image library; input each image and its corresponding information text into the trained visual large language model to obtain the text description corresponding to each image; segment the text description and extract the keywords corresponding to each image; and create an inverted index for the keywords corresponding to each image and write them into the search engine to generate an image search engine based on the visual large language model.
8. An image retrieval method, characterized in that: The following steps are involved: S10: Obtaining a user search request; S20: parsing the user search request to obtain keywords; S30: Input the keyword into the image search engine generated by any one of the methods of claims 1 to 6 to perform a search and obtain a search result.
9. The image retrieval method according to claim 8, characterized in that: The method further includes step S40: displaying the search results to the user, and the user further filters, sorts or modifies the query based on the search results.
10. An image retrieval system, characterized in that: It includes a search request acquisition module, a parsing module and a retrieval module; The search request acquisition module is used to acquire user search requests; The parsing module is used to parse the user search request to obtain keywords; The retrieval module is used to input the keyword into the image search engine generated by any method of claims 1-6 to perform retrieval and obtain retrieval results.