Image retrieval method and device
By using word and semantic matching algorithms to screen candidate texts in image and text datasets and using a large language model to determine the target image, the problems of retrieval limitations and low accuracy in image retrieval methods are solved, and more efficient image retrieval is achieved.
Patent Information
- Application Number
- CN202510881942.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-26
AI Technical Summary
The existing image retrieval methods are difficult to perform effective retrieval through text due to the high complexity of images, resulting in large retrieval limitations and low accuracy.
By determining the text to be retrieved and the related image and text datasets, the word matching algorithm and semantic matching algorithm are used to screen candidate texts in the image and text dataset, and the large language model is used to screen the target text from the word and semantic matching candidate texts, and then the target image is determined.
It improves the generalization ability and accuracy of image retrieval and enhances the effect of target image retrieval.
Smart Images

Figure CN120705344A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to image retrieval methods and devices. Background Art
[0002] In the context of text-based image retrieval, images contain many elements, resulting in high complexity and making them difficult to retrieve through text. In the prior art, images are generally manually annotated with functional labels and component labels, and fixed labels are used for retrieval during the image retrieval process. However, this image retrieval method has limited categories of manually annotated labels and does not take into account characteristics such as the structure and orientation of the image. This leads to significant limitations in image retrieval, high search difficulty, and low search accuracy. Therefore, there is an urgent need for a more effective image retrieval method to address the above-mentioned issues. Summary of the Invention
[0003] In view of this, embodiments of this specification provide an image retrieval method. One or more embodiments of this specification also relate to an image retrieval apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0004] According to a first aspect of the embodiments of this specification, there is provided an image retrieval method, comprising: Determining a text to be retrieved and a graphic and text dataset associated with the text to be retrieved; Determining candidate word matching texts that match the text to be retrieved in the image-text dataset using a word matching algorithm, and determining candidate semantic matching texts that match the text to be retrieved in the image-text dataset using a semantic matching algorithm; A large language model is used to screen a target text that matches the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts, and a target image corresponding to the target text is determined in the image-text dataset.
[0005] Optionally, the construction of any image-text data pair in the image-text data set includes: Determining a target image, and determining graphic and text elements contained in the target image, as well as image category, element layout information, and image structure information corresponding to the target image; Generate graphic element description information corresponding to the graphic element, and construct an image description text based on the graphic element description information, the image category, the element layout information and the image structure information; The image-text data pair is generated based on the target image and the image description text.
[0006] Optionally, the determining, in the image-text dataset, a candidate word matching text that matches the text to be retrieved using a word matching algorithm includes: Determining at least two candidate texts in the image-text dataset, performing word segmentation processing on the at least two candidate texts to obtain at least two groups of candidate word segments, and performing word segmentation processing on the text to be retrieved to obtain word segments to be retrieved; Using the word matching algorithm, construct an inverted index based on the at least two candidate texts and at least two groups of candidate segmentation words, and calculate the text matching degree between the to-be-retrieved text and each candidate text based on the to-be-retrieved segmentation words and the inverted index; The term matching candidate text is determined from the at least two candidate texts based on a text matching degree between the to-be-retrieved text and each candidate text.
[0007] Optionally, performing word segmentation processing on the text to be searched to obtain the word segments to be searched includes: Perform word segmentation processing on the text to be searched based on the target word segmentation library to obtain initial word segmentation to be searched; Stop words are determined in the initial participles to be searched, and the stop words are deleted from the initial participles to be searched to obtain the participles to be searched.
[0008] Optionally, the determining, in the image-text dataset, semantic matching candidate texts that match the text to be retrieved using a semantic matching algorithm includes: Determining at least two candidate texts in the image-text dataset, and converting the at least two candidate texts into at least two candidate text vectors using the semantic matching algorithm, and converting the to-be-retrieved text into the to-be-retrieved text vector; A vector index is constructed based on the at least two candidate texts and the at least two candidate text vectors, and the semantically matching candidate text that is similar to the text vector to be retrieved is retrieved based on the vector index.
[0009] Optionally, the converting the at least two candidate texts into at least two candidate text vectors by using the semantic matching algorithm includes: Segmenting the at least two candidate texts respectively to obtain candidate text segments corresponding to the at least two candidate texts respectively; The candidate text segments corresponding to each candidate text are encoded to obtain the at least two candidate text vectors.
[0010] Optionally, constructing a vector index based on the at least two candidate texts and the at least two candidate text vectors includes: The at least two candidate texts and the at least two candidate text vectors are stored in a similarity search library, and the vector index is constructed in the similarity search library.
[0011] Optionally, the using of the large language model to screen a target text matching the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts includes: Constructing a model input boost word and a model output prompt word for the text to be retrieved; The text to be retrieved, the term matching candidate text, the semantic matching candidate text, the model input boost word and the model output prompt word are input into the large language model to obtain the target text.
[0012] Optionally, inputting the to-be-retrieved text, the term matching candidate text, the semantic matching candidate text, the model input boost word, and the model output prompt word into the large language model to obtain the target text includes: Based on the model input prompt word, the large language model is used to calculate at least two text matching degrees between the text to be retrieved and the term matching candidate text and the semantic matching candidate text respectively; Sorting the term matching candidate texts and the semantic matching candidate texts according to the at least two text matching degrees, and generating model output information according to the sorting results and the model output prompt words; The target text is determined based on the model output information.
[0013] According to a second aspect of the embodiments of this specification, there is provided an image retrieval apparatus, comprising: A determination module configured to determine a text to be retrieved and a graphic and text dataset associated with the text to be retrieved; a matching module configured to determine, in the image-text dataset, a term matching candidate text that matches the text to be retrieved using a term matching algorithm, and to determine, in the image-text dataset, a semantic matching candidate text that matches the text to be retrieved using a semantic matching algorithm; The screening module is configured to use a large language model to screen target texts matching the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts, and determine a target image corresponding to the target text in the image-text dataset.
[0014] According to a third aspect of the embodiments of this specification, a computing device is provided, including: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned image retrieval method are implemented.
[0015] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned image retrieval method are implemented.
[0016] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program or instructions, which implement the steps of the above-mentioned image retrieval method when executed by a processor.
[0017] An image retrieval method provided by an embodiment of the present specification determines a text to be retrieved and a graphic-text data set associated with the text to be retrieved. A word matching algorithm is used to determine a word matching candidate text that matches the text to be retrieved in the graphic-text data set, and a word matching candidate text with a high degree of matching between the word dimension matching and the text to be retrieved. A semantic matching algorithm is used to determine a semantic matching candidate text that matches the text to be retrieved in the graphic-text data set, and a semantic matching candidate text with a high degree of matching between the semantic dimension matching and the text to be retrieved. A large language model is used to screen a target text that matches the text to be retrieved from the word matching candidate texts and the semantic matching candidate texts, and by rearranging the word matching candidate texts and the semantic matching candidate texts, determining the target text, and determining the target image corresponding to the target text in the graphic-text data set, the generalization ability of image retrieval can be improved, and the accuracy of target image retrieval can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of an image retrieval method provided by one embodiment of this specification; Figure 2 This is a retrieval schematic diagram of an image retrieval method provided by one embodiment of this specification; Figure 3 This is a schematic diagram of retrieval results of an image retrieval method provided in one embodiment of this specification; Figure 4 This is a flowchart of a processing process of an image retrieval method provided by one embodiment of this specification; Figure 5 This is a schematic diagram of the structure of an image retrieval device provided by one embodiment of this specification; Figure 6 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0019] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0020] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0021] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0022] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0023] First, the terms involved in one or more embodiments of this specification are explained.
[0024] Design draft: During the product development process, the designer creates a visual draft for the product, which consists of a series of different components arranged in a customized layout.
[0025] Large language model (LLM): A language model consisting of an artificial neural network with many parameters (typically billions of weights or more) that is trained on large amounts of unlabeled text using self-supervised or semi-supervised learning. Large language models are capable of understanding and generating natural language and other types of content to perform a variety of tasks.
[0026] The BM25 algorithm is a probabilistic information retrieval model used to assess the relevance of documents to queries. It improves on the traditional TF-IDF algorithm by introducing document length normalization and parameter adjustments, significantly improving the quality of search results. BM25 is widely used in search engines, document retrieval, and other scenarios, and remains a strong baseline model in the information retrieval field.
[0027] TF-IDF (Term Frequency-Inverse Document Frequency): A statistical method used to assess the importance of a term in a document or corpus. This technique is often used as a basis for keyword extraction in information retrieval and text mining. It consists of two components: term frequency (TF) and inverse document frequency (IDF). Term frequency refers to the frequency with which a term appears in a document. Inverse document frequency measures the general importance of a term.
[0028] FAISS (Facebook AI Similarity Search): An open-source and efficient similarity search library, it is primarily used to quickly find the top-K vectors most similar to a query vector in large-scale vector data. It is widely used in recommendation systems, image retrieval, semantic search, natural language processing, and other fields.
[0029] FLAT: A basic indexing mode for precise nearest neighbor retrieval. When using the FLAT index, FAISS stores the original vector data and performs an exhaustive linear search to find the most similar vectors during retrieval. This approach ensures accurate retrieval results because it determines the nearest neighbors by directly comparing the query vector with each vector in the database and calculating similarity (e.g., using inner product, Euclidean distance, etc.).
[0030] The inner product, also known as the dot product, is a fundamental concept in linear algebra. It defines the multiplication of two vectors, resulting in a scalar (a single number). For two vectors a and b, their inner product is a·b.
[0031] Jieba: A very popular Chinese word segmentation library that supports Python. It can segment Chinese text and provides multiple modes and functions to meet different needs.
[0032] BGE-M3 model: A general semantic vector model with leading multilingual and cross-lingual retrieval capabilities, providing comprehensive and high-quality support for input texts of different granularities, such as "sentences," "paragraphs," "chapters," and "documents."
[0033] Token: In the field of natural language processing (NLP), a token refers to the basic unit of text data, which can be a word, a punctuation mark, a number, or any meaningful character sequence.
[0034] In this specification, an image retrieval method is provided. This specification also relates to an image retrieval device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0035] See also Figure 1 , Figure 1 A flowchart of an image retrieval method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0036] Step 102: Determine the text to be retrieved and the image and text dataset associated with the text to be retrieved.
[0037] Specifically, the text to be retrieved can be an image requirement description text used for image search, which is used to retrieve images that match the semantics and word meanings of the text to be retrieved. The image-text dataset contains multiple image-text data pairs consisting of text and images that match the text. The image in the image-text data pair can be a design draft image, and accordingly, the text in the image-text data pair is the design draft image description text. The design draft image description text includes a description of the image and text elements in the design draft image, as well as a description of the image category, element layout, and image structure corresponding to the image and text elements.
[0038] Based on this, the text to be retrieved for image retrieval and the image-text dataset that can be retrieved based on the associated text to be retrieved are determined. Subsequently, text retrieval can be performed in the image-text dataset based on the image description corresponding to the text to be retrieved.
[0039] In actual applications, when the image in the image-text data pair is a design draft image, the design draft image description text includes but is not limited to the application name corresponding to the design draft image, application category, design draft function category, components included in the design draft, the theme mode of the design draft, the structure and layout information of the design draft, and a detailed description of the design draft.
[0040] Furthermore, the image-text data set includes multiple image-text data pairs, each of which includes a target image and an image description text describing the target image. The image description text can be obtained by describing image features such as elements and element layout contained in the target image. The specific implementation is as follows: Determine a target image, and determine the graphic elements contained in the target image, as well as the image category, element layout information and image structure information corresponding to the target image; generate graphic element description information corresponding to the graphic elements, and construct an image description text based on the graphic element description information, the image category, the element layout information and the image structure information; generate the graphic data pair based on the target image and the image description text.
[0041] Specifically, the target image can be a design draft image, that is, an image corresponding to a graphic user interface, and specifically can be an interface image in an application, such as an image of a shopping cart interface, an image of a music player interface, an image and text editing interface, and other visual images. The graphic elements contained in the target image can be personalized elements such as controls, buttons, icons, selection boxes, and labels. Graphic and text elements can contain only images or only text, or they can contain both images and text. The image category corresponding to the target image can be the application category corresponding to the image, and the element layout information can be information such as the position, size, color, and relationship between elements in the target image; the image structure information can be the combination of pictures and text in the target image, the theme mode (color mode), and the screen orientation (landscape or portrait).
[0042] Based on this, when constructing an image-text data pair, the target image can be first determined. An image description text for the target image is generated by analyzing the target image in terms of structure, color, layout, and theme patterns. The target image's image and text elements are determined, along with the corresponding image category, element layout information, and image structure information. Graphic and text element description information corresponding to the graphic and text elements is generated, and an image description text is constructed based on the graphic and text element description information, image category, element layout information, and image structure information. A graphic and text data pair is generated based on the target image and image description text.
[0043] For example, in the text-image retrieval scenario of a design draft, the design draft retrieval can be performed based on the design draft retrieval text input by the user. The user provides the design draft retrieval text, that is, the text to be retrieved. By searching in the image-text dataset based on the text to be retrieved, you can retrieve image-text data pairs that have a high degree of match with the description content of the text to be retrieved in the image-text dataset that stores a large number of image-text data pairs, and then determine the target image that matches the text to be retrieved. The image-text dataset can be constructed in advance based on a large number of design draft images. Take the design draft image as the target image, describe the target image as detailed as possible in terms of structure, color, layout, theme pattern and other dimensions, and obtain image-text description information. Construct an image-text data pair based on the target image and the image-text description information.
[0044] For an order screenshot in an item recycling app, the image description text could be: "This screenshot is from the item recycling app. The app category is shopping. This screenshot category is order. The components in this screenshot include checkbox, checkbox + text, drop-down trigger, and tab. 1. Summary: This screenshot shows a page for evaluating and recycling wallets for used luxury goods. Users can choose to recycle in-store or by courier, and then fill in the relevant information to submit the order. 2. Content: The top of the page displays an image of a wallet and detailed information, including its name, color, and size. The image is on the left, and text information is on the right. Below this, there's a "Follow" button that users can click to follow the item. Next is a status selection area where users can select the wallet's status (unused, minor marks, obvious marks, or severe marks). These options are arranged horizontally as tabs. A prompt appears in the center, informing users that they can choose to in-store or mail the item for an appraisal, along with a button for obtaining an accurate appraisal. At the bottom of the page are two tabs for recycling methods, with in-store recycling selected by default. Below the in-store recycling tab is a row of icons and text describing the recycling process: making an appointment to visit the store, taking photos of the item, bidding for 15 minutes, having it authenticated by an appraiser, and receiving payment from the platform. Below the in-store recycling tab are two buttons: "Please select a recycling store" and "Please select an appointment time." At the bottom of the page is a checkbox and a button to submit the order. Users must agree to the recycling agreement before submitting the order. 3. Cards with a combination of images and text: At the top of the page, there's a wallet image on the left, with text on the right. In the center, there's a small icon on the left, with text on the right. 4. Page Theme and Screen Orientation: Theme: Light Mode, Screen Orientation: Portrait.
[0045] In summary, by describing the target image in multiple dimensions and generating image description information, the target image can be described in detail and fully, thereby improving the accuracy of the text description.
[0046] Step 104: using a word matching algorithm to determine a word matching candidate text that matches the text to be retrieved in the image-text dataset, and using a semantic matching algorithm to determine a semantic matching candidate text that matches the text to be retrieved in the image-text dataset.
[0047] Specifically, after determining the text to be retrieved and the image-text dataset associated with the text to be retrieved, a term matching algorithm can be used to identify candidate term matching texts that match the text to be retrieved in the image-text dataset, and a semantic matching algorithm can be used to identify candidate semantic matching texts that match the text to be retrieved in the image-text dataset. The term matching algorithm is used to identify text content in the image-text dataset that has a high degree of term similarity with the text to be retrieved based on term similarity. A candidate term matching text can be at least one image description text determined based on term overlap in the image-text dataset. The number of candidate term matching texts can be set based on actual needs. The semantic matching algorithm is used to identify text content in the image-text dataset that has a high degree of semantic similarity with the text to be retrieved based on semantic similarity. A candidate semantic matching text can be at least one image description text determined based on semantic similarity in the image-text dataset. The number of candidate semantic matching texts can be set based on actual needs. The term matching algorithm can be the BM25 algorithm, and the semantic matching algorithm can be an algorithm corresponding to the BGE-M3 model.
[0048] Based on this, after determining the text to be retrieved and the image-text dataset associated with the text to be retrieved, a word matching algorithm is used to determine a set number of at least one word matching candidate texts in the image-text dataset that match the text to be retrieved, and a semantic matching algorithm is used to determine at least one semantic matching candidate text in the image-text dataset that matches the text to be retrieved. The number of word matching candidate texts can be the same as the number of semantic matching candidate texts.
[0049] Furthermore, when searching the image-text dataset based on the text to be detected, the image description text in the image-text data pair contained in the image-text dataset can be segmented respectively, and then text matching can be performed based on the segmented words. The specific implementation is as follows: At least two candidate texts are determined in the image-text dataset, word segmentation processing is performed on the at least two candidate texts to obtain at least two groups of candidate word segments, and word segmentation processing is performed on the text to be retrieved to obtain the word segments to be retrieved; using the word matching algorithm, an inverted index is constructed based on the at least two candidate texts and the at least two groups of candidate word segments, and a text matching degree between the text to be retrieved and each candidate text is calculated based on the word segments to be retrieved and the inverted index; and the word matching candidate text is determined from the at least two candidate texts based on the text matching degree between the text to be retrieved and each candidate text.
[0050] Specifically, all candidate texts in the image and text dataset can be used as at least two candidate texts, that is, a comprehensive search can be performed on the image and text dataset. Jieba word segmentation library can be used to segment text sentences into words and to segment continuous Chinese characters into meaningful terms. The words obtained by segmenting the candidate text are the candidate analysis, and the words obtained by segmenting the text to be searched are the words to be searched. The inverted index is a data structure used to quickly query the candidate texts corresponding to the word segmentation.
[0051] Based on this, at least two candidate texts are determined in the image-text dataset, and the at least two candidate texts are segmented based on the Jieba word segmentation library to obtain at least two groups of candidate word segments. The text to be retrieved is also segmented based on the Jieba word segmentation library to obtain the word segments to be retrieved. Using a word matching algorithm, an inverted index is constructed based on the at least two candidate texts and the at least two groups of candidate word segments, and a text matching degree between the text to be retrieved and each candidate text is calculated based on the word segments to be retrieved and the inverted index. According to the calculated text matching degree, the at least two candidate texts are sorted in descending order of matching degree, and a set number of word matching candidate texts are determined from the at least two candidate texts based on the text matching degree between the text to be retrieved and each candidate text.
[0052] Using the above example, Figure 2 As shown, the search interface allows users to enter search text and perform image searches. The design function within the search interface allows for page settings and operations. By adjusting the percentage corresponding to EDEEF0, the interface color can be adjusted. Local variables and local styles can also be selected, as well as content export. Users can create and save layers within the search interface. To perform an image search, users can enter the search text "shopping cart page" in the search interface. The search will display a selection of retrieved design drafts, which can be viewed by clicking the "View Results" control. At this point, images matching "shopping cart page" need to be searched within the image-text dataset. The image-text dataset contains images and their corresponding image descriptions. First, word segmentation is performed on the search text and each image description in the image-text dataset. The word segmentation corresponding to "shopping cart page" is obtained, along with the word segmentation corresponding to each image description. An inverted index is constructed based on the candidate text and the word segmentation corresponding to the candidate text. The inverted index is then used to calculate the similarity between the search text and each candidate text, i.e., the text match. After sorting at least two candidate texts in descending order of text matching degree, the top 20 candidate texts may be selected as term matching candidate texts.
[0053] In summary, after tokenizing the candidate texts, an inverted index is constructed. Based on the word matching algorithm and the inverted index, the text matching degree between the text to be retrieved and each candidate text is calculated. Then, a set number of word matching candidate texts are selected according to the calculated text matching degree. Furthermore, word matching candidate texts that match the text to be retrieved are selected in the dimension of word similarity, improving the efficiency and accuracy of determining word matching candidate texts.
[0054] Furthermore, considering that the text to be retrieved is a descriptive text provided by the user, the initial tokens obtained by tokenizing the text to be retrieved do not all have meanings. After determining the initial tokens to be retrieved, stop words need to be removed and meaningful words need to be retained. The specific implementation is as follows: Tokenize the text to be retrieved based on the target token library to obtain initial tokens to be retrieved; determine stop words in the initial tokens to be retrieved, and delete the stop words in the initial tokens to be retrieved to obtain the tokens to be retrieved.
[0055] Based on this, tokenize the text to be retrieved based on the target token library to obtain initial tokens to be retrieved. Determine meaningless stop words in the initial tokens to be retrieved, and delete the stop words in the initial tokens to be retrieved, then the tokens to be retrieved can be obtained.
[0056] Continuing with the above example, words such as "de" (的) and "shi" (是) in the initial tokens to be retrieved have no practical meaning, and such words can be deleted as stop words to obtain the tokens to be retrieved.
[0057] In summary, determine stop words in the initial tokens to be retrieved, and delete the stop words in the initial tokens to be retrieved to obtain the tokens to be retrieved, achieving the removal of stop words and avoiding negative impacts of stop words on the accuracy of subsequent text retrieval.
[0058] Furthermore, in the semantic dimension, by vectorizing the candidate texts, a vector index can be constructed. Then, based on the vector index for retrieval, semantic matching candidate texts with a relatively high similarity to the text to be retrieved can be accurately determined. The specific implementation is as follows: Determine at least two candidate texts in the graphic and text dataset, and use the semantic matching algorithm to convert the at least two candidate texts into at least two candidate text vectors, and convert the text to be retrieved into a text to be retrieved vector; construct a vector index based on the at least two candidate texts and the at least two candidate text vectors, and retrieve the semantic matching candidate texts similar to the text to be retrieved vector based on the vector index.
[0059] Specifically, a candidate text vector refers to a feature vector obtained by vectorizing the image description text in a text-image dataset. A searchable text vector refers to a feature vector obtained by vectorizing the searchable text. The vector index can be a FAISS index (FLAT mode), which stores the candidate text vectors directly without compression and constructs a vector index based on the candidate text's text identifier.
[0060] Based on this, at least two candidate texts are determined in the image-text dataset, and a semantic matching algorithm is used to convert the at least two candidate texts into at least two candidate text vectors, and the text to be retrieved is also converted into a text vector to be retrieved. By storing the text identifiers of the candidate texts in correspondence with the candidate text vectors, a vector index corresponding to the at least two candidate texts and the at least two candidate text vectors is constructed. Based on the vector index, semantically matching candidate texts similar to the text vector to be retrieved are retrieved, and the similarity between each candidate text vector and the text vector to be retrieved is calculated. Based on the calculated similarity, a semantically matching candidate text is determined from the at least two candidate texts.
[0061] In summary, searching the text to be retrieved based on vector index can accurately determine the semantic matching candidate texts with high similarity to the text to be retrieved, thereby improving the text retrieval accuracy.
[0062] Furthermore, before converting the candidate text into a candidate text vector, the candidate text needs to be preprocessed by semantic parsing, etc. The specific implementation is as follows: The at least two candidate texts are segmented respectively to obtain candidate text segments corresponding to the at least two candidate texts; and the candidate text segment corresponding to each candidate text is encoded to obtain the at least two candidate text vectors.
[0063] Specifically, the vectorization processing of the candidate text fragment corresponding to the candidate text can be performed by first encoding the candidate text fragment, then marking the hidden state as a global semantic representation, and generating a dense vector after normalization, that is, the candidate text vector. The candidate text can be segmented using the dynamic block strategy provided by the BGE-M3 model. The dynamic block strategy can segment very long text into processable text fragments while retaining contextual coherence. For Chinese text, the built-in word segmentation module automatically completes semantic analysis at the word level and obtains several tokens corresponding to the text. Several tokens can be used as input to the model for vector retrieval.
[0064] Based on this, the at least two candidate texts are segmented to obtain candidate text segments corresponding to the at least two candidate texts. The candidate text segments corresponding to the at least two candidate texts retain contextual coherence. The candidate text segments corresponding to each candidate text are encoded and then normalized to obtain at least two candidate text vectors.
[0065] Continuing with the above example, the BGE-M3 model can be used for vectorization processing of the text to be retrieved and at least two candidate texts. The BGE-M3 model supports long text inputs of up to 8192 tokens, and uses a dynamic chunking strategy to split very long texts into processable segments while retaining contextual coherence. For Chinese text, the built-in word segmentation module automatically completes semantic parsing at the word level and obtains several tokens corresponding to the text as input to the model. After the text to be retrieved and at least two candidate texts pass through the Transformer encoding layer, the hidden state of the [CLS] (the default first token of the Transformer model) marker position is used as the global semantic representation. After normalization, a 1024-dimensional dense vector is generated to obtain the text vectors corresponding to the text to be retrieved and at least two candidate texts.
[0066] In summary, at least two candidate texts are segmented respectively to obtain candidate text segments corresponding to at least two candidate texts; the candidate text segment corresponding to each candidate text is encoded, thereby improving the accuracy of candidate text vector generation.
[0067] Furthermore, after converting the candidate text into a candidate text vector, the candidate text vector can be directly stored, and then a vector index can be constructed based on the stored candidate text vector for text retrieval based on the text to be retrieved. The specific implementation is as follows: The at least two candidate texts and the at least two candidate text vectors are stored in a similarity search library, and the vector index is constructed in the similarity search library.
[0068] Based on this, at least two candidate texts and at least two candidate text vectors are stored in a similarity retrieval library, and a vector index is constructed in the similarity retrieval library based on the candidate texts and the candidate text vectors corresponding to the candidate texts to ensure the text retrieval effect.
[0069] Continuing with the previous example, in the design draft retrieval scenario, each screenshot of a design draft generates a text description, and each text segment is processed into a text vector. Therefore, a large amount of vector data needs to be stored and quickly retrieved. Vector storage and retrieval can be performed using the open-source framework FAISS. This framework is an open-source, high-performance similarity retrieval library designed specifically for high-dimensional vectors and widely used in fields such as recommendation systems, image retrieval, and natural language processing. Its core goal is to quickly find the data most similar to the query vector from massive amounts of data. FAISS supports multiple similarity retrieval strategies, including using inner products as vector similarity for retrieval. To ensure effective retrieval, the basic FLAT mode is used when building the index. FLAT mode uses brute-force search, which offers higher accuracy than other indexing methods. In the index created by FAISS, the text vector to be retrieved is compared to the text vectors describing all design drafts, and the 20 candidate texts with the highest similarity are selected.
[0070] In summary, storing at least two candidate texts and at least two candidate text vectors in a similarity retrieval library, constructing a vector index in the similarity retrieval library, and searching the text to be retrieved based on the vector index can improve retrieval accuracy.
[0071] Step 106: Utilize the large language model to screen target texts matching the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts, and determine a target image corresponding to the target text in the image-text dataset.
[0072] Specifically, after using the word matching algorithm to determine the word matching candidate texts that match the text to be retrieved in the image-text dataset, and using the semantic matching algorithm to determine the semantic matching candidate texts that match the text to be retrieved in the image-text dataset, the large language model can be used to screen the target text that matches the text to be retrieved from the word matching candidate texts and the semantic matching candidate texts, and determine the target image corresponding to the target text in the image-text dataset, wherein the large language model can be a reordering model, which can determine the target text with a higher degree of matching or relevance to the text to be retrieved by reordering the word matching candidate texts and the semantic matching candidate texts according to the input prompt words constructed based on the text to be retrieved.
[0073] Based on this, after using the word matching algorithm to determine the word matching candidate texts that match the text to be retrieved in the image and text dataset, and using the semantic matching algorithm to determine the semantic matching candidate texts that match the text to be retrieved in the image and text dataset, the large language model is used to reorder the word matching candidate texts and the semantic matching candidate texts, and the target text with a higher matching degree is determined based on the matching degree between the word matching candidate texts and the semantic matching candidate texts and the text to be retrieved respectively, and the target image corresponding to the target text is determined in the image and text dataset.
[0074] Furthermore, the text processing of the large language model needs to be guided by the model input boost words and model output prompt words. Therefore, it is necessary to construct the model input boost words and model output prompt words to guide the large language model in text processing. The specific implementation is as follows: A model input boost word and a model output prompt word are constructed for the text to be retrieved; the text to be retrieved, the term matching candidate text, the semantic matching candidate text, the model input boost word and the model output prompt word are input into the large language model to obtain the target text.
[0075] Specifically, model input prompts are input instructions for the large language model to perform text reranking tasks. These prompts provide constraints and guidance for the large language model. Model input prompts include an overview of the prediction task and constraints on the prediction process. Model output prompts are used to restrict the format of the large language model's output content. These prompts include intent analysis, text relevance analysis, keyword analysis, image design analysis, and text ranking for the text to be tested.
[0076] Based on this, a set of model input boosting words is constructed for the text to be retrieved, used to guide the large language model in text processing, and model output prompting words are constructed to constrain the format of the text output by the large language model. The text to be retrieved, candidate texts for term matching, candidate texts for semantic matching, model input boosting words, and model output prompting words are input into the large language model. By reordering the candidate texts for term matching and semantic matching, a set number of target texts are selected.
[0077] Continuing with the previous example, the large language model re-ranks term and semantic match candidates based on the relevance between the searched text and the candidate texts, and the relevance between the searched text and the candidate texts with semantic matches. To ensure the accuracy of the large language model's predictions, model input and output prompts are constructed and fed into the large language model along with the searched text, the candidate texts with term matches, and the candidate texts with semantic matches.
[0078] Model input prompts can include a prediction task overview and prediction process constraints. The prediction task overview can be, "You are an experienced designer. You need to learn design ideas from other UI designs. Please complete the relevant requirements based on the specific scenario and carefully consider the designer's search intent to find the most relevant UI. Designers may search for dimensions such as: UI functions, local functions, components included in the UI, and the name and category of the app to which the UI belongs. Please evaluate and rank based on the following aspects: 1. Content relevance: The direct relevance and match degree of the candidate text to the query. 2. Design details: Whether the candidate text provides specific design details and specifications related to the query." The prediction process constraints can be "Please output the final results according to the following requirements: 1. Please output the document numbers in order of relevance from high to low, and use > to separate the numbers. 2. The output format is a JSON string: {"explanation": "order":}. Do not output other additional content and do not omit paragraphs. 3. When explaining, please first analyze the meaning of the query. The designer's query may not be clear enough or contain redundant words. Please first analyze the designer's true intention, remove redundant words, and then analyze the importance order of each keyword in the query, and then analyze its relevance to each paragraph. 4. The designer's query may contain multiple keywords. Please prioritize text descriptions containing more keywords, then sort text descriptions with high-priority keywords, and finally sort text descriptions with low priority. 5. Keywords with higher priority are: component name, screen content direction, app name, app category, screen category. 6. As long as the query contains app keywords, it means searching on the app category dimension."
[0079] In summary, the large language model is fed with the text to be searched, candidate texts for term matching, candidate texts for semantic matching, model input boosting words, and model output prompt words. The large language model then reorders the candidate texts for term matching and semantic matching to select a set number of target texts. During the large language model prediction process, the model input boosting words and model output prompt words constrain and guide the large language model's accuracy in determining the target texts.
[0080] Furthermore, considering that it is necessary to screen out target texts with a higher degree of matching with the text to be detected by re-ranking the candidate texts for word matching and the candidate texts for semantic matching, the candidate texts for word matching and the candidate texts for semantic matching can be re-ranked by respectively calculating the degree of matching between each candidate text for word matching and each candidate text for semantic matching and the text to be detected. The specific implementation is as follows: Based on the model input prompt, use the large language model to calculate at least two text matching degrees between the text to be retrieved and the word matching candidate text and the semantic matching candidate text respectively; sort the word matching candidate text and the semantic matching candidate text according to the at least two text matching degrees, and generate model output information according to the sorting result and the model output prompt; determine the target text based on the model output information.
[0081] Based on this, based on the model input prompt, use the large language model to calculate at least two text matching degrees between the text to be retrieved and the word matching candidate text and the semantic matching candidate text respectively, that is, calculate the text matching degree between the text to be retrieved and each word matching candidate text, and calculate the text matching degree between the text to be retrieved and each semantic matching candidate text. Sort the word matching candidate text and the semantic matching candidate text according to the at least two text matching degrees, and generate model output information according to the sorting result and the model output prompt. Select a set number of candidate texts as the target text according to the sorting of the word matching candidate text and the semantic matching candidate text in the model output information.
[0082] Continuing with the previous example, the large language model rearranges the word matching candidate text and the semantic matching candidate text under the prompt of the model input prompt, and outputs the prediction content in the format of the model output prompt. When the text to be retrieved is "shopping cart page", the prediction content output by the large language model can be "{\"explanation\":\"1. Analyze the query intention: Find the UI interface that displays the shopping cart page. The keywords are'shopping cart' and 'page'. 'Page' is a redundant word and can be removed. The real intention of the designer is to find the UI interface of the shopping cart function. 2. Keyword importance:'shopping cart' is the most important. 3. Text relevance analysis: Texts that directly mention'shopping cart' in the title or previous description and provide detailed descriptions of UI interface components and layouts should be selected first. 4. Descriptions of specific components and design specifications related to the shopping cart page (such as the layout of the shopping cart product list, product card design, bottom operation bar, etc.) are more relevant than texts that only simply mention the shopping cart.\",\"order\":\"41>46>26>38>19>9>31>49>44>50>40>2>4>12>42>14>1>5>29>45>13>15>18>23>25>28>32>36>3>7>8>10>16>20>27>34>43>47>48>6>11>17>21>22>24>30>33>35>37>39\"}". Each serial number in the order corresponds to a word matching candidate text or a semantic matching candidate text. According to the sorting of the text, the first 20 texts can be selected as the target text. For example Figure 3As shown, the search interface will display the retrieved shopping cart image, and you can view multiple retrieved design drafts by scrolling down the page.
[0083] To sum up, by calculating the text matching degree of each word matching candidate text and each semantic matching candidate text with the text to be detected respectively, the word matching candidate texts and semantic matching candidate texts are reordered to improve the accuracy of the target text and achieve accurate retrieval of the text to be retrieved.
[0084] An image retrieval method provided by an embodiment of the present specification determines a text to be retrieved and a graphic-text data set associated with the text to be retrieved. A word matching algorithm is used to determine a word matching candidate text that matches the text to be retrieved in the graphic-text data set, and a word matching candidate text with a high degree of matching between the word dimension matching and the text to be retrieved. A semantic matching algorithm is used to determine a semantic matching candidate text that matches the text to be retrieved in the graphic-text data set, and a semantic matching candidate text with a high degree of matching between the semantic dimension matching and the text to be retrieved. A large language model is used to screen a target text that matches the text to be retrieved from the word matching candidate texts and the semantic matching candidate texts, and by rearranging the word matching candidate texts and the semantic matching candidate texts, determining the target text, and determining the target image corresponding to the target text in the graphic-text data set, the generalization ability of image retrieval can be improved, and the accuracy of target image retrieval can be improved.
[0085] The following combined Figure 4 , taking the application of the image retrieval method provided in this specification in design draft retrieval as an example, the image retrieval method is further explained. Figure 4 A flowchart of a processing process of an image retrieval method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0086] Step 402: Determine the text to be retrieved and the image-text dataset associated with the text to be retrieved.
[0087] Step 404: determining at least two candidate texts in the image-text dataset, performing word segmentation processing on the at least two candidate texts to obtain at least two groups of candidate word segments, and performing word segmentation processing on the text to be retrieved to obtain word segments to be retrieved.
[0088] Step 406: Using a word matching algorithm, construct an inverted index based on at least two candidate texts and at least two groups of candidate segmentations, and calculate the text matching degree between the text to be retrieved and each candidate text based on the segmentation to be retrieved and the inverted index.
[0089] Step 408: Determine a word matching candidate text from at least two candidate texts based on the text matching degree between the to-be-retrieved text and each candidate text.
[0090] Step 410: Determine at least two candidate texts in the image-text dataset, and convert the at least two candidate texts into at least two candidate text vectors using a semantic matching algorithm, and convert the to-be-retrieved text into a to-be-retrieved text vector.
[0091] Step 412: construct a vector index based on at least two candidate texts and at least two candidate text vectors, and retrieve semantically matching candidate texts similar to the text vector to be retrieved based on the vector index.
[0092] Step 414: Construct a model input boosting word and a model output prompt word for the text to be retrieved.
[0093] Step 416: Based on the model input prompt word, the large language model is used to calculate at least two text matching degrees between the text to be retrieved and the term matching candidate text and the semantic matching candidate text.
[0094] Step 418: Sort the word matching candidate texts and the semantic matching candidate texts according to at least two text matching degrees, and generate model output information based on the sorting results and the model output prompt words.
[0095] Step 420: Determine the target text based on the model output information, and determine the target image corresponding to the target text in the image and text dataset.
[0096] In summary, the text to be retrieved and the image-text dataset associated with the text to be retrieved are determined. A word matching algorithm is used to determine candidate word matching texts that match the text to be retrieved in the image-text dataset, and the word matching candidate texts with a high degree of matching between the word dimension matching and the text to be retrieved are determined. A semantic matching algorithm is used to determine candidate semantic matching texts that match the text to be retrieved in the image-text dataset, and the semantic matching candidate texts with a high degree of matching between the semantic dimension matching and the text to be retrieved are determined. A large language model is used to screen the target text that matches the text to be retrieved from the word matching candidate texts and the semantic matching candidate texts. By rearranging the word matching candidate texts and the semantic matching candidate texts, the target text is determined, and the target image corresponding to the target text is determined in the image-text dataset. This can improve the generalization ability of image retrieval and the accuracy of target image retrieval.
[0097] Corresponding to the above method embodiment, this specification also provides an image retrieval device embodiment, Figure 5 FIG. 1 shows a schematic diagram of the structure of an image retrieval device provided by an embodiment of this specification. Figure 5 As shown, the device includes: A determination module 502 is configured to determine a text to be retrieved and a graphic and text dataset associated with the text to be retrieved; The matching module 504 is configured to determine a word matching candidate text that matches the text to be retrieved in the image-text dataset using a word matching algorithm, and to determine a semantic matching candidate text that matches the text to be retrieved in the image-text dataset using a semantic matching algorithm; The screening module 506 is configured to screen target texts matching the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts using a large language model, and determine a target image corresponding to the target text in the image-text dataset.
[0098] In an optional embodiment, the determining module 502 is further configured to: Determining a target image, and determining graphic and text elements contained in the target image, as well as image category, element layout information, and image structure information corresponding to the target image; Generate graphic element description information corresponding to the graphic element, and construct an image description text based on the graphic element description information, the image category, the element layout information and the image structure information; The image-text data pair is generated based on the target image and the image description text.
[0099] In an optional embodiment, the matching module 504 is further configured to: Determining at least two candidate texts in the image-text dataset, performing word segmentation processing on the at least two candidate texts to obtain at least two groups of candidate word segments, and performing word segmentation processing on the text to be retrieved to obtain word segments to be retrieved; Using the word matching algorithm, construct an inverted index based on the at least two candidate texts and at least two groups of candidate segmentation words, and calculate the text matching degree between the to-be-retrieved text and each candidate text based on the to-be-retrieved segmentation words and the inverted index; The term matching candidate text is determined from the at least two candidate texts based on a text matching degree between the to-be-retrieved text and each candidate text.
[0100] In an optional embodiment, the matching module 504 is further configured to: Perform word segmentation processing on the text to be searched based on the target word segmentation library to obtain initial word segmentation to be searched; Stop words are determined in the initial participles to be searched, and the stop words are deleted from the initial participles to be searched to obtain the participles to be searched.
[0101] In an optional embodiment, the matching module 504 is further configured to: Determining at least two candidate texts in the image-text dataset, and converting the at least two candidate texts into at least two candidate text vectors using the semantic matching algorithm, and converting the to-be-retrieved text into the to-be-retrieved text vector; A vector index is constructed based on the at least two candidate texts and the at least two candidate text vectors, and the semantically matching candidate text that is similar to the text vector to be retrieved is retrieved based on the vector index.
[0102] In an optional embodiment, the matching module 504 is further configured to: Segmenting the at least two candidate texts respectively to obtain candidate text segments corresponding to the at least two candidate texts respectively; The candidate text segments corresponding to each candidate text are encoded to obtain the at least two candidate text vectors.
[0103] In an optional embodiment, the matching module 504 is further configured to: The at least two candidate texts and the at least two candidate text vectors are stored in a similarity search library, and the vector index is constructed in the similarity search library.
[0104] In an optional embodiment, the screening module 506 is further configured to: Constructing a model input boost word and a model output prompt word for the text to be retrieved; The text to be retrieved, the term matching candidate text, the semantic matching candidate text, the model input boost word and the model output prompt word are input into the large language model to obtain the target text.
[0105] In an optional embodiment, the screening module 506 is further configured to: Based on the model input prompt word, the large language model is used to calculate at least two text matching degrees between the text to be retrieved and the term matching candidate text and the semantic matching candidate text respectively; Sorting the term matching candidate texts and the semantic matching candidate texts according to the at least two text matching degrees, and generating model output information according to the sorting results and the model output prompt words; The target text is determined based on the model output information.
[0106] An image retrieval device provided by an embodiment of the present specification determines a text to be retrieved and a graphic-text data set associated with the text to be retrieved. A word matching algorithm is used to determine a word matching candidate text that matches the text to be retrieved in the graphic-text data set, and a word matching candidate text with a high degree of matching between the word dimension matching and the text to be retrieved. A semantic matching algorithm is used to determine a semantic matching candidate text that matches the text to be retrieved in the graphic-text data set, and a semantic matching candidate text with a high degree of matching between the semantic dimension matching and the text to be retrieved. A large language model is used to screen a target text that matches the text to be retrieved from the word matching candidate texts and the semantic matching candidate texts, and by rearranging the word matching candidate texts and the semantic matching candidate texts, determining the target text, and determining the target image corresponding to the target text in the graphic-text data set, the generalization ability of image retrieval can be improved, and the accuracy of target image retrieval can be improved.
[0107] The above is a schematic diagram of an image retrieval device according to this embodiment. It should be noted that the technical solution of the image retrieval device and the technical solution of the above-mentioned image retrieval method are based on the same concept. For details not described in detail in the technical solution of the image retrieval device, please refer to the description of the technical solution of the above-mentioned image retrieval method.
[0108] Figure 6 6 shows a block diagram of a computing device 600 according to one embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0109] Computing device 600 also includes an access device 640 that enables computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0110] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0111] Computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 can also be a mobile or stationary server.
[0112] The processor 620 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned image retrieval method when executed by the processor.
[0113] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above-mentioned image retrieval method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned image retrieval method.
[0114] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned image retrieval method when executed by a processor.
[0115] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the aforementioned image retrieval method are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned image retrieval method.
[0116] An embodiment of the present specification further provides a computer program product, including a computer program or instructions, which implements the steps of the above-mentioned image retrieval method when executed by a processor.
[0117] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the aforementioned image retrieval method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned image retrieval method.
[0118] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0119] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0120] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0121] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0122] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An image retrieval method, characterized in that: include: Determining a text to be retrieved and a graphic and text dataset associated with the text to be retrieved; Determining, in the image-text dataset, a word matching candidate text that matches the text to be retrieved using a word matching algorithm, and determining, in the image-text dataset, a semantic matching candidate text that matches the text to be retrieved using a semantic matching algorithm; A large language model is used to screen target texts matching the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts, and a target image corresponding to the target text is determined in the image-text dataset.
2. The image retrieval method according to claim 1, wherein: The construction of any image-text data pair in the image-text data set includes: Determining a target image, and determining graphic and text elements contained in the target image, as well as image category, element layout information, and image structure information corresponding to the target image; Generate graphic element description information corresponding to the graphic element, and construct image description text based on the graphic element description information, the image category, the element layout information and the image structure information; The image-text data pair is generated based on the target image and the image description text.
3. The image retrieval method according to claim 1, wherein: The step of using a word matching algorithm to determine a word matching candidate text that matches the text to be retrieved in the image and text dataset includes: Determining at least two candidate texts in the image-text dataset, performing word segmentation processing on the at least two candidate texts to obtain at least two groups of candidate word segments, and performing word segmentation processing on the text to be retrieved to obtain word segments to be retrieved; Using the word matching algorithm, construct an inverted index based on the at least two candidate texts and at least two groups of candidate segmentation words, and calculate the text matching degree between the to-be-retrieved text and each candidate text based on the to-be-retrieved segmentation words and the inverted index; The term matching candidate text is determined from the at least two candidate texts based on a text matching degree between the to-be-retrieved text and each candidate text.
4. The image retrieval method according to claim 3, wherein: The step of performing word segmentation processing on the text to be searched to obtain the word segments to be searched includes: Perform word segmentation processing on the text to be searched based on the target word segmentation library to obtain initial word segmentation to be searched; Stop words are determined in the initial participles to be searched, and the stop words are deleted from the initial participles to be searched to obtain the participles to be searched.
5. The image retrieval method according to claim 1, wherein: The determining of semantic matching candidate texts that match the text to be retrieved in the image-text dataset using a semantic matching algorithm includes: Determining at least two candidate texts in the image-text dataset, and converting the at least two candidate texts into at least two candidate text vectors using the semantic matching algorithm, and converting the to-be-retrieved text into the to-be-retrieved text vector; A vector index is constructed based on the at least two candidate texts and the at least two candidate text vectors, and the semantically matching candidate text that is similar to the text vector to be retrieved is retrieved based on the vector index.
6. The image retrieval method according to claim 5, characterized in that: The converting the at least two candidate texts into at least two candidate text vectors by using the semantic matching algorithm includes: Segmenting the at least two candidate texts respectively to obtain candidate text segments corresponding to the at least two candidate texts respectively; The candidate text segments corresponding to each candidate text are encoded to obtain the at least two candidate text vectors.
7. The image retrieval method according to claim 5, characterized in that: The constructing a vector index based on the at least two candidate texts and the at least two candidate text vectors includes: The at least two candidate texts and the at least two candidate text vectors are stored in a similarity search library, and the vector index is constructed in the similarity search library.
8. The image retrieval method according to claim 1, wherein: The step of using the large language model to select a target text that matches the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts includes: Constructing a model input boost word and a model output prompt word for the text to be retrieved; The text to be retrieved, the term matching candidate text, the semantic matching candidate text, the model input boost word and the model output prompt word are input into the large language model to obtain the target text.
9. The image retrieval method according to claim 8, wherein: The step of inputting the text to be retrieved, the term matching candidate text, the semantic matching candidate text, the model input boost word, and the model output prompt word into the large language model to obtain the target text includes: Based on the model input prompt word, the large language model is used to calculate at least two text matching degrees between the text to be retrieved and the term matching candidate text and the semantic matching candidate text respectively; Sorting the term matching candidate texts and the semantic matching candidate texts according to the at least two text matching degrees, and generating model output information according to the sorting results and the model output prompt words; The target text is determined based on the model output information.
10. An image retrieval device, characterized in that: include: A determination module configured to determine a text to be retrieved and a graphic and text dataset associated with the text to be retrieved; a matching module configured to determine, in the image-text dataset, a term matching candidate text that matches the text to be retrieved using a term matching algorithm, and to determine, in the image-text dataset, a semantic matching candidate text that matches the text to be retrieved using a semantic matching algorithm; The screening module is configured to use a large language model to screen target texts matching the text to be retrieved from the term matching candidate texts and the semantic matching candidate texts, and determine a target image corresponding to the target text in the image-text dataset.
11. A computing device, characterized in that include: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the image retrieval method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium, characterized in that It stores computer-executable instructions, which, when executed by a processor, implement the steps of the image retrieval method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The method comprises a computer program or instructions, which, when executed by a processor, implements the steps of the image retrieval method according to any one of claims 1 to 9.