Picture retrieval method and device based on understanding
By combining the semantic understanding of the big model and the picture encoder, image vectors are generated and fine-tuned to identify and understand relevant elements in the picture, the problems of insufficient key information and lack of generalization ability in traditional image retrieval methods are solved, and the image retrieval effect with high accuracy and flexibility is achieved.
Patent Information
- Application Number
- CN202510116932.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional image search methods have problems such as insufficient key information, inability to search for local harmful/sensitive content, requiring standard basic image libraries, loss of semantic information, low recall and lack of generalization ability.
The understanding-based image retrieval method is adopted, and the semantic understanding of the big model is used. Through the combination of the picture encoder and the large language model, the picture vector is generated and fine-tuned, the relevant elements in the picture and the semantic text are identified, the target recognition and matching are performed, and the picture is judged whether the picture is abnormal.
It improves the generalization ability of image retrieval, can identify more abnormal images, continuously improve image coverage ability, has flexibility and high accuracy, and is suitable for a variety of scenarios, especially suitable for accurate retrieval of sensitive images.
Smart Images

Figure CN119988663A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image retrieval technology, and in particular relates to an understanding-based image retrieval method, device, computer-readable storage medium, and electronic device. Background Art
[0002] The core technology of traditional image retrieval methods is to search for images by images, but there are the following problems: (1) The search is based on the overall picture, without highlighting the key information / key parts, especially the harmful / sensitive parts, and cannot search for harmful / sensitive content in local pictures.
[0003] (2) It is necessary to maintain a standard basic searchable image library. Images that are not in the library cannot be searched.
[0004] (3) The semantic information of the image will be lost during the retrieval process.
[0005] (4) The recall rate of image retrieval is low and it lacks generalization ability.
[0006] (5) For images that need to be understood, it is difficult to understand and identify them as a whole. For example, all the elements of a picture are normal, but the meaning of the whole picture needs to be understood more. Summary of the invention
[0007] In order to address the above problems of traditional image retrieval methods, this application proposes a new understanding-based image retrieval method. This method combines the semantic understanding of a large model with a large-scale training data set to improve the generalization ability of the retrieval method, identify more abnormal images, and continuously improve the image coverage capability.
[0008] In order to achieve the above objectives, this application provides the following technical solutions: The first aspect of the present application provides an understanding-based image retrieval method, the method comprising: Input the image to be retrieved, use the image encoder (such as the CLIP framework) to process the image to be retrieved, and output the image vector required for the large language model to understand; Fine-tune the response format and input format of the large language model according to the needs of the retrieval scenario; The large language model understands the input image vector and outputs the analysis result, which includes the relevant elements in the image and the semantic text of the image. According to the semantic text of the picture, the relevant elements in the picture are identified, the type of the identified target is identified (such as people, objects, etc.), and then the target is matched; For the matched target, the corresponding image vector is extracted, and whether the image is abnormal is determined based on the text description information of the image.
[0009] Optionally, the method of the present application further includes: inputting text related to the image to be retrieved while inputting the image to be retrieved; The text is the context text at the location near (indicated by) the image or the entire text in the file where the image is located; When the input text is the entire text in the file where the image is located, the image location is provided at the same time, and the image location is encoded into the long text of the large language model, which is automatically discovered by the large language model; The text is used as annotated text or context prompt (prompt word / instruction).
[0010] Optionally, the method of the present application also includes: for text, first using the basic CLIP model to retrieve relevant text, then summarizing the retrieval results, and outputting the text summary required for the large language model to understand.
[0011] Optionally, the present application method also includes: Train an image understanding model using the image encoder and the large language model; Use the image encoder to extract different image codes, then concatenate the extracted code vectors and input them into the Q-Former network; Adding an instruction dimension to the Q-Former network incorporates the context of the text in which the image is located (e.g., web page content) and the corresponding structural features (e.g., the website where the web page is located, the time, the publishing organization, etc.); The Q-Former network outputs features that enter the large language model for understanding and outputs analysis results, which include relevant elements in the image and semantic text of the image.
[0012] Optionally, the method of the present application also includes: using the visual and text projection layers to fuse the image and text information, and outputting a multimodal fused vector, so that the large language model can understand the image more easily.
[0013] Optionally, the method of the present application further includes: comparing the image through a comparison library according to the understanding text description information of the image, and then judging whether the image is abnormal through a discrimination model; judging whether the image is abnormal includes judging the whole image and judging the part, and the judging method is as follows: (1) Whole-image judgment (a) Using the content of the text and using natural language processing (NLU) to determine whether the entire image is abnormal based on the semantic text as a whole; (b) Determine whether the entire image is abnormal through the entire image vector retrieval library; (c) When the confidence level is low, manual intervention is performed; (2) Local judgment (a) Based on the content of the text understanding, use target detection (such as YOLO) to identify key targets and extract their vector features; (b) Automatically compare the image with the comparison library and use NLU to judge the semantics of the understood text to determine whether it is an abnormal image; (c) When the confidence level is low, manual intervention is performed; (d) Combine text and target comparison to determine the sensitive and non-sensitive parts of the image.
[0014] Optionally, the method of the present application further includes: tokenizing the image in an adaptive manner through a Vision Tokenizer to form a visual encoding of the image; and judging key harmful information based on different image encoders (such as CLIP, etc.) or pooling encoders (such as RMAC, GeM, CROW); The visual and text projection layers understand the entire image by combining the image's relevant contextual information and output it to the large language model for text understanding. The large language model generation layer uses the language generation ability and world knowledge of the large language model to understand the image and generate text description information of the understanding of the image.
[0015] Optionally, the method of the present application also includes further vectorization or judgment of the image understanding text, including: Mark whether the image is abnormal and convert the image annotation judgment into text annotation judgment; Use Bert / BGE to embed key text content; Collect the target speech text through machine generation or manual construction, and then directly use the search method to determine whether the image has the target value.
[0016] A second aspect of the present application provides an understanding-based image retrieval device, the device comprising: Preprocessing module: used to input the image to be retrieved, use the image encoder (such as CLIP framework) to process the image to be retrieved, and output the image vector required for the large language model to understand; Model fine-tuning module: used to fine-tune the response format and input format of the large language model according to the needs of the retrieval scenario; Image understanding module: used to understand the input image vector through a large language model and output the analysis results, which include the relevant elements in the image and the semantic text of the image; Object detection module: used to identify relevant elements in the image based on the semantic text of the image, identify the type of the identified target (such as people, objects, etc.), and then perform target matching; Result judgment module: used to extract the corresponding image vector for the matching target and judge whether the image is abnormal based on the text description information of the image.
[0017] The device implements the steps of the aforementioned understanding-based image retrieval method when running.
[0018] A third aspect of the present application provides an electronic device, comprising: a memory and a processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the aforementioned understanding-based image retrieval method.
[0019] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the aforementioned understanding-based image retrieval method are implemented.
[0020] In summary, this application proposes a new understanding-based image retrieval method, which has the following characteristics: (1) This method is different from the traditional image search. It proposes a new retrieval architecture from image to semantics and from semantics to retrieval, which can directly perform local vectorized retrieval of certain images. This solution first associates the image with the text semantics, and then uses semantics to understand and weight the image properties, key attributes, and types. It is more flexible and can retrieve richer and more accurate content.
[0021] (2) This method does not rely on retrieving image libraries. For some scenes that cannot be enumerated, this method has extremely strong generalization capabilities.
[0022] (3) This method converts images into semantic understanding, truly simulating the way people understand images. The application architecture can be applied to all scenarios, connecting images with NLP and having strong versatility.
[0023] (4) This method performs appropriate small-scale training for specific scenarios and can be embedded in different vertical scenarios without relying on a large amount of labeled data.
[0024] (5) This method can accurately retrieve sensitive images, especially those with local sensitive information.
[0025] Other features and advantages of the present application will be described in the following description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the techniques indicated in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 This is a basic design framework diagram of the embodiment method of this application.
[0028] Figure 2 This is a schematic diagram of the templated results output in the method of the embodiment of the present application.
[0029] Figure 3 This is a basic design framework diagram of the image understanding step in the embodiment method of this application.
[0030] Figure 4 This is a schematic diagram of the judgment and annotation of pictures in the embodiment method of this application.
[0031] Figure 5 This is the overall implementation flow chart of the understanding-based image retrieval method of this application.
[0032] Figure 6 This is a schematic diagram of the core process of automatic labeling in an embodiment of the present application.
[0033] Figure 7 This is a schematic diagram of determining the sensitive portion of an image in an embodiment of the present application.
[0034] Figure 8 This is a schematic diagram for determining the non-sensitive portion of a picture in an embodiment of the present application.
[0035] Fig. 9 This is a structural diagram of the composition of the image retrieval device based on understanding of this application.
[0036] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.
[0038] As used herein, the term “including” and its variations are open inclusions, ie, “including but not limited to”; the term “based on” means “based at least in part on”; and the term “one embodiment” means “at least one embodiment”.
[0039] It should be noted that the modifications of "one" and "plurality" mentioned in the present application are illustrative rather than restrictive, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0040] Terminology explanation: Vision Tokenizer: is a tool that converts image data into discrete token sequences. The generated token sequences can be processed by large language models (LLMs), etc., so that the model can understand and generate vision-related tasks (this conversion process is similar to the word segmentation process in natural language processing, but for visual information).
[0041] Tokenizer: It is a basic tool in natural language processing (NLP) that is used to segment continuous text strings into discrete token units, which are usually words, subwords, or characters. Tokenizer plays a vital role in the text preprocessing stage, laying the foundation for subsequent tasks such as text analysis, model training, and natural language understanding.
[0042] Input and output description of this method: 1. Input: text + picture; the text part can be used as annotated text or context prompt (prompt word / instruction); the picture is used as input; in the actual reasoning process, if there is context, the following text can be used as part of the prompt (instruction) for reasoning.
[0043] 2. Output: A piece of text that understands the semantics of the image, whether the image is abnormal, and the location of the abnormality for local anomalies.
[0044] 1. Basic framework of this method Figure 1 The figure shows the basic design framework of this method, including image understanding, fixed template output, target detection, vector extraction and other steps.
[0045] The basic framework is as follows: Step 1: Use an image encoder (such as the CLIP framework) and a large language model to train an image understanding model. This step requires relevant fine-tuning data (SFT) to optimize the Q-Former network. The overall network is as follows: In this model block, multiple image encoders are used to extract different image encodings, and then these encoding vectors are concatenated and input into the Q-Former network. The entire Q-Former output features enter the large language model.
[0046] The instruction dimension is added to Q-Former, which incorporates the context of the text in which the image is located (such as the web page content) and other contexts, such as corresponding structural features (the website where the web page is located, the time, the publishing organization, etc.).
[0047] Step 2: If Figure 2 As shown, according to the first step, a templated result is output, which includes relevant elements in the image and the understood semantic text of the image.
[0048] Step 3: Based on the semantic text output in step 2, identify the target in the output "object" and at the same time, identify the type of the identified target, such as person or object, so as to perform target matching.
[0049] Step 4: For the required target, extract the corresponding image vector and use relevant models such as CNN.
[0050] 2. Basic framework of image understanding Figure 3 The following is the basic design framework of the image understanding step in this method, including: (1) Vision Tokenizer (Vision Tokenizer) In this method, the image is tokenized using Vision Tokenizer to form a visual encoding of the image. In this process, considering that the meanings represented by different angles of the image may be different, the image is tokenized in an adaptive way (highlighting the important parts or the parts of interest). Figure 3 The figure shows the judgment of key harmful information based on different image encoders (such as CLIP, etc.) or pooling encoders with special important parts (such as RMAC, GeM, CROW).
[0051] (2) Visual and text projection layer This layer understands the entire image based on the image and its related context information, and outputs it to the LLM for text understanding output.
[0052] a. This layer connects LLM with the image relationship part, encoding prompt, picture context, picture embedding and other information as the final PEFT (such as lora, P-tuning v2, etc.) original model.
[0053] b. This layer can introduce textual context or other requirements of the image.
[0054] (3) LLM generation layer This layer uses the language generation ability and world knowledge of the big model to generate an understanding of the picture, and the big model ultimately generates a text description of the picture and an extraction of objects or things; the content format is as follows:
[0055] 3. Further vectorization or judgment of image content understanding like Figure 4 As shown, through this step, whether the image is harmful is marked, and the image annotation judgment is converted into a text annotation judgment, which conforms to the first principle.
[0056] Use Bert / BGE etc. to embed key text content.
[0057] At the same time, because a large amount of target speech text can be collected (generated by machine or constructed manually), the retrieval method can be used directly to determine whether the image has target value.
[0058] Complete the result output of the retrieval or discrimination model.
[0059] Figure 5 The overall implementation process of the understanding-based image retrieval method provided by this application is shown, including the following steps: Input the image to be retrieved, use the image encoder (such as the CLIP framework) to process the image to be retrieved, and output the image vector required for the large language model to understand; Fine-tune the response format and input format of the large language model according to the needs of the retrieval scenario; The large language model understands the input image vector and outputs the analysis result, which includes the relevant elements in the image and the semantic text of the image. According to the semantic text of the picture, the relevant elements in the picture are identified, the type of the identified target is identified (such as people, objects, etc.), and then the target is matched; For the matched target, the corresponding image vector is extracted, and whether the image is abnormal is determined based on the text description information of the image.
[0060] In order to better understand the technical solution of the present application, it is further described with reference to the following embodiments.
[0061] Figure 6 The figure shows the core process of automatic labeling in this embodiment.
[0062] 1. Input part: It is divided into picture input and text input.
[0063] a. There may be multiple pictures, and you need to find the text paragraph that is most relevant to this picture.
[0064] The context related to the location near the image (indicator).
[0065] Provide all text and the location of the image, encoded into the LLM long text, and automatically discovered by LLM.
[0066] b. Text is an optional input.
[0067] 2. Image and text processing: Output the image vector and text summary required for LLM understanding.
[0068] a. For images, the encoding format of the image is obtained through multiple finetuning image encoders, such as CLIP.
[0069] b. For text, very long text contains a lot of extra information that does not need to be paid attention to. First, we use the basic CLIP model to retrieve relevant text and then summarize it.
[0070] 3. Project Layer (visual and text projection layer): outputs multimodal fusion vectors.
[0071] a. As an intermediate layer associated with the LLM model, it integrates image and text information, making it easier for LLM to understand images while reducing training costs.
[0072] b. Reduce unnecessary LLM forgetting disasters.
[0073] 4. LLM after scene fine-tuning.
[0074] a. For open source models, fine-tune the response format and input format based on the needs of the scenario, so that the understood output better meets the usage requirements.
[0075] b. Output reference.
[0076]
[0077] 5. Post-processing of understood text a. Whether the text description information is sensitive: whether the content directly understood by the output text is sensitive
[0078] Compare through the comparison library (sensitive content can be adjusted in real time) Because of the important figures at that time, the spoof XXX needs to be paid attention to Through discriminant models: such as text classification (NLU related, sentiment related), etc. b. Whole picture judgment: (a) Using the content of the text, we use natural language processing (NLU) to determine whether the entire image is sensitive based on the semantic text as a whole.
[0079] (b) Determine whether it is sensitive through the whole image vector retrieval library.
[0080] (c) Manual check labeling (when confidence is low, manual intervention labeling).
[0081] Output: The entire image is sensitive.
[0082] c. Local judgment (in the above comic, the character to be extracted is "XXX"): (a) Based on the content of the text understanding, use target detection such as YOLO to identify key (sensitive) targets (such as XXX) and extract their vector features.
[0083] (b) Automatically check whether the person is sensitive (by comparing the database, determine whether the person is a sensitive person) and NLU’s judgment on the semantics of the understood text.
[0084] (c) Manually check whether the output annotation targets and the mapped semantic understanding parts are sensitive (when confidence is low, manual intervention is required for annotation).
[0085] Output: Combine text and character targets to determine sensitive and non-sensitive parts, such as Figure 7 and Figure 8 shown.
[0086] Fig. 9 The figure shows an understanding-based image retrieval device proposed in this application, comprising: Preprocessing module: used to input the image to be retrieved, use the image encoder (such as CLIP framework) to process the image to be retrieved, and output the image vector required for the large language model to understand; Model fine-tuning module: used to fine-tune the response format and input format of the large language model according to the needs of the retrieval scenario; Image understanding module: used to understand the input image vector through a large language model and output the analysis results, which include the relevant elements in the image and the semantic text of the image; Object detection module: used to identify relevant elements in the image based on the semantic text of the image, identify the type of the identified target (such as people, objects, etc.), and then perform target matching; Result judgment module: used to extract the corresponding image vector for the matching target and judge whether the image is abnormal based on the text description information of the image.
[0087] When the above device is running, the steps of the understanding-based image retrieval method disclosed in this application are implemented.
[0088] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the device, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the specified function or operation, or can be realized by a combination of special hardware and computer instructions.
[0089] like Fig.10 As shown, the embodiment of the present application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing a computer program executable by the processor, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 runs the executable computer program to implement the steps of the above-mentioned understanding-based image retrieval method.
[0090] It is understandable that, in addition to the memory and the processor, the electronic device may also include an input device such as a keyboard, an output device such as a display, and other communication modules. The input device, the output device, and other communication modules communicate with the processor via an I / O interface (i.e., an input / output interface).
[0091] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0092] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can perform the various steps of the understanding-based image retrieval method disclosed in the present application.
[0093] In the context of the present application, a computer-readable storage medium may be a tangible medium, more specific examples of which would include a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] In particular, according to an embodiment of the present application, the process described in the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the understanding-based image retrieval method disclosed in the present application. When the computer program is executed by a processing device, the above functions defined in the method of the embodiment of the present application are executed.
[0095] Although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present application. The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by a specific combination of the above technical features, but also should cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above public concept.
[0096] Those skilled in the art should also understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for image retrieval based on understanding, characterized in that: The method comprises: Input the image to be retrieved, use the image encoder to process the image to be retrieved, and output the image vector required for the large language model to understand; Fine-tune the response format and input format of the large language model according to the needs of the retrieval scenario; The large language model understands the input image vector and outputs the analysis result, which includes the relevant elements in the image and the semantic text of the image. According to the semantic text of the picture, the relevant elements in the picture are identified, the type of the identified target is identified, and then the target is matched; For the matched target, the corresponding image vector is extracted, and whether the image is abnormal is determined based on the image's understanding text description information.
2. The method according to claim 1, characterized in that The method further comprises: inputting text related to the picture to be retrieved while inputting the picture to be retrieved; The text is the context text near the image or the entire text in the file where the image is located; When the input text is the entire text in the file where the image is located, the image location is provided at the same time, and the image location is encoded into the long text of the large language model, which is automatically discovered by the large language model; The text is used as the marked text or context prompt.
3. The method according to claim 2, characterized in that The method further includes: for the text, firstly using the basic CLIP model to retrieve relevant text, then summarizing the retrieval results, and outputting the text summary required for the large language model to understand.
4. The method according to claim 2, characterized in that: The method further comprises: Train an image understanding model using the image encoder and the large language model; Use the image encoder to extract different image codes, then concatenate the extracted code vectors and input them into the Q-Former network; Adding an instruction dimension to the Q-Former network incorporates the context of the text in which the image is located and the corresponding structural features; The Q-Former network outputs features that enter the large language model for understanding and outputs analysis results, which include relevant elements in the image and semantic text of the image.
5. The method according to claim 2, characterized in that: The method also includes: using the visual and text projection layers to fuse the image and text information, and outputting a multimodal fused vector, so that the large language model can more easily understand the image.
6. The method according to claim 2, characterized in that The method further includes: comparing the image through a comparison library according to the understanding text description information of the image, and then judging whether the image is abnormal through a discrimination model; judging whether the image is abnormal includes whole-image judgment and local judgment, and the judgment method is as follows: (1) Whole-image judgment (a) Using the content of the text and using text understanding to determine whether the entire image is abnormal based on the semantic text as a whole; (b) Determine whether the entire image is abnormal through the entire image vector retrieval library; (c) When the confidence level is low, manual intervention is performed; (2) Local judgment (a) Based on the content of text understanding, use object detection to identify key objects and extract their vector features; (b) Automatically compare the image with the comparison library and use NLU to judge the semantics of the understood text to determine whether it is an abnormal image; (c) When the confidence level is low, manual intervention is performed; (d) Combine text and target comparison to determine the sensitive and non-sensitive parts of the image.
7. The method according to claim 2, characterized in that The method further includes: tokenizing the image in an adaptive manner by using a Vision Tokenizer to form a visual encoding of the image; and determining key harmful information based on different image encoders or pooling encoders; The visual and text projection layers understand the entire image by combining the image's relevant contextual information and output it to the large language model for text understanding. The large language model generation layer uses the language generation ability and world knowledge of the large language model to understand the image and generate text description information of the understanding of the image.
8. The method according to claim 2, characterized in that: The method also includes further vectorization or judgment of the image understanding text, including: Mark whether the image is abnormal and convert the image annotation judgment into text annotation judgment; Use Bert / BGE to embed key text content; Collect the target speech text through machine generation or manual construction, and then directly use the search method to determine whether the image has the target value.
9. An image retrieval device based on understanding, characterized in that: The device comprises: Preprocessing module: used to input the image to be retrieved, use the image encoder to process the image to be retrieved, and output the image vector required for the large language model to understand; Model fine-tuning module: used to fine-tune the response format and input format of the large language model according to the needs of the retrieval scenario; Image understanding module: used to understand the input image vector through a large language model and output the analysis results, which include the relevant elements in the image and the semantic text of the image; Object detection module: used to identify the relevant elements in the image based on the semantic text of the image, identify the type of the identified target, and then perform target matching; Result judgment module: used to extract the corresponding image vector for the matching target and judge whether the image is abnormal based on the text description information of the image.
10. An electronic device, characterized in that: include: Memory and processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the understanding-based image retrieval method as described in any one of claims 1-8.