Picture searching method and device, electronic equipment and storage medium

By comparing the similarity between the image search text and the image target description text generated by the generative language model, the problem of low image search efficiency in the prior art is solved, and more efficient image retrieval is achieved.

CN120067375APending Publication Date: 2025-05-30GUANGZHOU WERIDE TECH LTD CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411916105.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the image search method that calculates the similarity degree after converting text and pictures into vector representations is less efficient.

Method used

The method of comparing the similarity of the image search text with the predetermined target description text of multiple pictures is adopted. The target description text is generated through a generative language model, and vectorization processing and similarity calculation are used to improve the search efficiency.

Benefits of technology

By comparing text to text, the efficiency of image search is improved, the demand for computing resources is reduced, and faster image retrieval is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067375A_ABST
    Figure CN120067375A_ABST
Patent Text Reader

Abstract

The invention provides a picture searching method and device, electronic equipment and a storage medium, and relates to the technical field of computer processing. Performing similarity comparison on the picture search text and a predetermined target description text of each picture in the plurality of pictures to obtain the similarity between the picture search text and the target description text of each picture; and based on the similarity between the picture search text and the target description text of each picture, a first target picture in the plurality of pictures is displayed, and the similarity corresponding to the first target picture is higher than that of a second target picture which is not displayed in the plurality of pictures. And the comparison efficiency between the texts is higher than the comparison efficiency between the texts and the pictures, so that the picture searching efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer processing technologies, and particularly to a method, apparatus, electronic device, and storage medium for image search. Background Art

[0002] Currently, users can find images by entering text. In related technologies, text and images can be converted into vector representations that can be understood by a computer through appropriate algorithms. Then, the similarity between the vector representations of the text and the images is calculated to find images related to the text.

[0003] However, in related technologies, by converting text and images into vector representations that can be understood by a computer, and then calculating the similarity between the vector representations of the text and the images, the efficiency of image search is low. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide a method, apparatus, electronic device, and storage medium for image search to improve the efficiency of image search.

[0005] In a first aspect, an embodiment of the present invention provides an image search method, including: obtaining an image search text; comparing the similarity between the image search text and the pre-determined target description text of each image in a plurality of images to obtain the similarity between the image search text and the target description text of each image, where the target description text of the image is used to describe the content of the image, and the target description text of the image is the text output by a trained generative language model after inputting the image and a prompt text, and the prompt text is used to prompt the generative language model to output the target description text; based on the similarity between the image search text and the target description text of each image, displaying a first target image among the plurality of images, and the similarity corresponding to the first target image is higher than that of a second target image not displayed among the plurality of images.

[0006] In a possible implementation manner, before comparing the similarity between the image search text and the pre-determined target description text of each image in a plurality of images to obtain the similarity between the image search text and the target description text of each image, it further includes: preprocessing the image search text, and the preprocessing includes at least one of converting the text in the image search text into standard text, removing invalid text in the image search text, or performing synonym replacement on the text in the image search text; comparing the similarity between the image search text and the pre-determined target description text of each image in a plurality of images includes: comparing the similarity between the preprocessed image search text and the pre-determined target description text of each image in a plurality of images to obtain the similarity between the image search text and the target description text of each image.

[0007] In a possible implementation, the similarity between the picture search text and the pre-determined target description text of each picture among multiple pictures is compared to obtain the similarity between the picture search text and the target description text of each picture, including: respectively performing vectorization processing on the picture search text and the target description text of each picture to obtain a first text vector corresponding to the picture search text and a second text vector corresponding to the target description text of each picture; based on the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description text of each picture.

[0008] In a possible implementation, respectively performing vectorization processing on the picture search text and the target description text of each picture to obtain a first text vector corresponding to the picture search text and a second text vector corresponding to the target description text of each picture, including: using a trained Embedding model to respectively perform vectorization processing on the picture search text and the target description text of each picture to obtain a first text vector corresponding to the picture search text and a second text vector corresponding to the target description text of each picture.

[0009] In a possible implementation, the step of using a trained Embedding model to perform vectorization processing on the target description text of each picture to obtain a second text vector corresponding to the target description text of each picture includes: concatenating the target description texts of each picture to obtain a concatenated description text; using a trained Embedding model to perform vectorization processing on the concatenated description text to obtain a text vector corresponding to the concatenated description text; dividing the text vector corresponding to the concatenated description text according to the concatenation order and length of the target description texts of each picture to obtain a second text vector corresponding to the target description text of each picture.

[0010] In a possible implementation, based on the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description text of each picture includes: determining the vector distance between the first text vector and the second text vectors corresponding to the target description texts of each picture, where the vector distance is used to represent the similarity between the first text vector and the second text vectors; based on the vector distance between the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description text of each picture.

[0011] In a possible implementation, the vector distance includes at least two terms. The at least two vector distances include at least two of the Euclidean distance, dot product similarity distance, cosine similarity, or maximum inner product similarity. Based on the vector distance between the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description texts of each picture includes: normalizing the at least two vector distances to obtain the normalized at least two vector distances; fusing the normalized at least two vector distances to obtain the fused vector distance; and obtaining the similarity between the picture search text and the target description texts of each picture based on the fused vector distance.

[0012] In a possible implementation, before comparing the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures, it includes: identifying the search scenario of the picture search text to obtain search scenario information; comparing the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures includes: comparing the similarity between the picture search text and the pre-determined target description texts of each picture in the multiple pictures that match the search scenario information; after displaying the first target picture in the multiple pictures, it further includes: in response to an instruction to switch the displayed picture, comparing the similarity between the picture search text and the target description texts of each picture in the second target picture.

[0013] In a second aspect, an embodiment of the present invention provides a picture search device, including: an acquisition module, configured to acquire a picture search text; a similarity comparison module, configured to compare the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures to obtain the similarity between the picture search text and the target description texts of each picture, where the target description text of the picture is used to describe the content of the picture, and the target description text of the picture is the text output by the trained generative language model after inputting the picture and the prompt text, and the prompt text is used to prompt the generative language model to output the target description text; a picture search module, configured to display the first target picture in the multiple pictures based on the similarity between the picture search text and the target description texts of each picture, and the similarity corresponding to the first target picture is higher than that of the second target picture not displayed in the multiple pictures.

[0014] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method in the first aspect.

[0015] Fourthly, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the method of the first aspect.

[0016] The embodiments of the present invention bring the following beneficial effects: finding text by obtaining a picture; comparing the similarity between the text found from the picture and the pre-determined target description text of each picture among multiple pictures to obtain the similarity between the text found from the picture and the target description text of each picture, where the target description text of a picture is used to describe the content of the picture, and the target description text of a picture is the text output by a trained generative language model after inputting the picture and a prompt text, and the prompt text is used to prompt the generative language model to output the target description text; based on the similarity between the text found from the picture and the target description text of each picture, displaying a first target picture among multiple pictures, and the similarity corresponding to the first target picture is higher than that of a second target picture not displayed among multiple pictures. Since it is the comparison between texts to find pictures, and the comparison efficiency between texts is higher than that between text and picture, the efficiency of picture finding can be improved accordingly.

[0017] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0018] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of a picture finding method provided by an embodiment of the present invention;

[0021] Figure 2 It is a schematic structural diagram of a picture finding device provided by an embodiment of the present invention;

[0022] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0023] For the purpose of making the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Currently, the training of autonomous driving requires a large number of pictures as training data. Therefore, it is desirable to search for pictures through text. Currently, text and pictures can be converted into vector representations that can be understood by a computer through appropriate algorithms. Then, the similarity between the vector representations of the text and the pictures is calculated to find pictures related to the text. However, the computing power resources required for converting pictures into vector representations and then calculating the similarity between the vector representations of the text and the pictures are relatively high, and the efficiency is relatively low.

[0025] In view of this, a picture search method, device, electronic device and storage medium provided by the embodiments of the present invention can improve the efficiency of picture search and reduce the computing power resources required for picture search.

[0026] For the convenience of understanding this embodiment, a picture search method disclosed in the embodiments of the present invention will be introduced in detail first. The picture search method in one of the embodiments of the present disclosure can run on a local terminal device or a server. When the picture search method runs on the server, the method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device.

[0027] As Figure 1 shown, the picture search method is applied to a terminal device, where the terminal device can be the aforementioned local terminal device or the aforementioned client device. The method includes the following steps:

[0028] S110. Obtain a picture search text.

[0029] Among them, the picture search text may refer to the text used to search for pictures. In this embodiment, the picture search text may be the text input by the user. For example, the picture search text may be, for example, "a picture of a vehicle on a bridge", etc., which is not limited herein.

[0030] S120. Compare the similarity between the image search text and the pre-determined target description text of each image among multiple images to obtain the similarity between the image search text and the target description text of each image. Herein, the target description text of an image is used to describe the content of the image. The target description text of an image is the text output by the trained generative language model after inputting the image and the prompt text. The prompt text is used to prompt the generative language model to output the target description text.

[0031] Among them, the generative language model is an important model in natural language processing, and its core task is to generate natural language text. The prompt text can be the input text or instruction provided by the user to the model, which is used to guide the model to generate a specific type of response, such as guiding the model to generate a response that meets expectations or complete a specific task. In this embodiment, the prompt text is used to prompt the generative language model to output the description text of the image, so the target description text can reflect the content of the image. It can also be understood that the prompt text is used to prompt the generative language model to describe the content of the image, and the text output by the generative language model can be the description text of the image.

[0032] S130. Based on the similarity between the image search text and the target description text of each image, display the first target image among multiple images. The similarity corresponding to the first target image is higher than that of the second target image that is not displayed among multiple images.

[0033] In this embodiment, the similarity can be the similarity between texts, such as the similarity between the image search text and the target description text. In this embodiment, based on the size of the similarity, those with higher similarity can be selected for display, while those with lower similarity are not displayed first.

[0034] In this embodiment, since images are searched by comparing text with text, and the efficiency of comparing text with text is higher than that of comparing text with image, the efficiency of image search can be improved accordingly.

[0035] It should be noted that, in this embodiment, the number of target description texts for calculating similarity is greater than the number of first target pictures that can be displayed at one time. For example, the number of target description texts for calculating similarity is num_candidates, and the number of first target pictures is K. Exemplarily, the k-nearest neighbor (KNN) algorithm can be used. For example, find num_candidates approximate nearest neighbor candidates on each shard (one shard includes multiple pictures, and the pictures included in different shards may be the same or different, determined according to the actual sharding result). Search and calculate the similarity between these candidate vectors (second text vectors) and the query vector (first text vector), and select k most similar results from each shard. Then, search and merge the results of each shard to return the global k nearest neighbors (second text vectors). The pictures corresponding to the target description texts of the k nearest neighbors are the first target pictures.

[0036] In a possible implementation manner, before comparing the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures to obtain the similarity between the picture search text and the target description texts of each picture, it further includes:

[0037] Preprocess the picture search text, and the preprocessing includes at least one of converting the text in the picture search text into standard text, removing invalid text in the picture search text, or performing synonym replacement on the text in the picture search text.

[0038] Correspondingly, comparing the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures includes:

[0039] Compare the preprocessed picture search text with the pre-determined target description texts of each picture in multiple pictures to obtain the similarity between the picture search text and the target description texts of each picture.

[0040] It should be noted that, due to different expression habits of users, for the same thing, different users may use different expressions. Therefore, in this embodiment, the text in the picture search text can be converted into standard text, which can improve the accuracy of text similarity comparison.

[0041] It should be noted that there may be some invalid texts in the picture search text input by the user, such as prepositions. In this embodiment, the invalid text in the picture search text can be removed, which can improve the efficiency of text similarity comparison.

[0042] It should be noted that due to different expression habits of users, users may be accustomed to using negative statements, for example, "not dark roads, etc.", which can be converted into "bright roads", etc., which can improve the accuracy of text similarity comparison.

[0043] In another possible implementation, it is not necessary to pre-process the image search text, which can improve the efficiency of text similarity matching.

[0044] In a possible implementation, the target description text output by the generative language model may be filtered. The filtering method may refer to the description of the above preprocessing, and the similarity between the filtered target description text and the image search text may be calculated.

[0045] In another possible implementation, the target description text output by the generative language model may be directly subjected to similarity calculation with the image search text.

[0046] In a possible implementation, performing a similarity comparison between the image search text and a predetermined target description text of each of the multiple images to obtain the similarity between the image search text and the target description text of each image includes:

[0047] The image search text and the target description text of each image are vectorized respectively to obtain a first text vector corresponding to the image search text and a second text vector corresponding to the target description text of each image; based on the first text vector and the second text vector corresponding to the target description text of each image, the similarity between the image search text and the target description text of each image is obtained.

[0048] In this embodiment, the image search text and the target description text may be converted into text vectors respectively and then compared, so that the potential content of the text can be mined, thereby improving the accuracy of text similarity calculation.

[0049] In another possible implementation, the image search text and the target description text may be matched with keywords.

[0050] In a possible implementation, vectorization is performed on the image search text and the target description text of each image respectively to obtain a first text vector corresponding to the image search text and a second text vector corresponding to the target description text of each image, including:

[0051] The trained embedding model is used to vectorize the image search text and the target description text of each image, respectively, to obtain a first text vector corresponding to the image search text and a second text vector corresponding to the target description text of each image.

[0052] Among them, the Embedding model refers to the process of mapping high-dimensional data (such as text, pictures, videos) to a low-dimensional space. Simply put, an Embedding vector is an N-dimensional real-valued vector that represents the input data as a point in a continuous numerical space. Specifically, the Embedding model learns a mapping function to map the input high-dimensional data into a low-dimensional vector space, which is called the embedding space or feature space. This process is based on the distributional hypothesis and can capture the potential relationships and structures of the original data. Optionally, the Embedding model can include but is not limited to BERT, GPT, Word2Vec, GloVe, etc.

[0053] In this embodiment, the picture search text and the target description text of each picture can be vectorized by the same Embedding model, which can improve the accuracy of text similarity calculation.

[0054] In a possible implementation manner, the step of vectorizing the target description text of each picture by using the trained Embedding model to obtain the second text vector corresponding to the target description text of each picture includes:

[0055] Concatenate the target description texts of each picture to obtain the concatenated description text; use the trained Embedding model to vectorize the concatenated description text to obtain the text vector corresponding to the concatenated description text; divide the text vector corresponding to the concatenated description text according to the concatenation order and length of the target description texts of each picture to obtain the second text vector corresponding to the target description text of each picture.

[0056] In this embodiment, the concatenation of the target description texts can be that multiple target description texts are connected end to end in sequence to obtain the concatenated description text. Optionally, some delimiters (such as full stops, line breaks, or special markers) can be added between the description texts to distinguish different descriptions. The Embedding model converts the concatenated description text into one or more vectors. For example, if a model like BERT is used, it may generate a vector for each word or sub-word, and it can also choose to obtain the vector representation of the entire sentence (such as by averaging, max pooling, or a specific [CLS] marker). Then, according to the concatenation order and length of the target description texts of each picture, determine the start and end positions of each target description text, and then extract the vectors within the corresponding range, from which the second text vector corresponding to the target description text of each picture can be obtained. In addition, when generating vectors, the delimiters can be retained, and then the second text vector can be divided by the delimiters.

[0057] It should be noted that when dividing the second text vector by a delimiter, during the training of the Embedding model, the text vector carrying the delimiter is used as the output of the Embedding model for training. In this way, the text vector output by the Embedding model can retain the delimiter, so as to use the delimiter to divide the text vectors corresponding to different pictures.

[0058] In the technical solution of this embodiment, by splicing the target description texts of each picture, the spliced description text is obtained; the trained Embedding model is used to perform vectorization processing on the spliced description text to obtain the text vector corresponding to the spliced description text; the text vector corresponding to the spliced description text is divided according to the splicing order and length of the target description texts of each picture to obtain the second text vector corresponding to the target description text of each picture. In this way, through one processing of the Embedding model, the second text vectors corresponding to multiple target description texts can be obtained respectively, which can improve the efficiency of text similarity calculation.

[0059] In another possible implementation, it can also be that the Embedding model performs vectorization processing on one target description text each time.

[0060] In a possible implementation, based on the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description texts of each picture includes:

[0061] Determining the vector distance between the first text vector and the second text vectors corresponding to the target description texts of each picture, where the vector distance is used to represent the similarity between the first text vector and the second text vectors; based on the vector distance between the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description texts of each picture.

[0062] In this embodiment, the vector distance can be one or more items.

[0063] In a possible implementation, the vector distance includes at least two items. The at least two vector distances include at least two of the Euclidean distance, dot product similarity distance, cosine similarity, or maximum inner product similarity. Based on the vector distance between the first text vector and the second text vectors corresponding to the target description texts of each picture, obtaining the similarity between the picture search text and the target description texts of each picture includes:

[0064] Normalize at least two vector distances to obtain the normalized at least two vector distances; fuse the normalized at least two vector distances to obtain the fused vector distance; based on the fused vector distance, obtain the similarity between the picture search text and the target description texts of each picture.

[0065] It should be understood that whether the larger the vector distance, the greater the similarity, or the smaller the vector distance, the greater the similarity is related to the normalization method and is not limited here.

[0066] In this embodiment, by fusing at least two vector distances, the similarity between texts is determined, which can improve the accuracy of text similarity comparison.

[0067] In another possible implementation, the similarity can also be calculated through one vector distance. For example, the cosine similarity distance is used to calculate the similarity, which can improve the efficiency of text similarity calculation.

[0068] In one possible implementation, before comparing the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures, it includes:

[0069] Identify the search scenario of the picture search text to obtain the search scenario information; compare the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures, including: comparing the similarity between the picture search text and the pre-determined target description texts of each picture in multiple pictures that match the search scenario information; after displaying the first target picture among multiple pictures, it further includes: in response to the instruction to switch the displayed picture, comparing the similarity between the picture search text and the target description texts of each picture in the second target picture.

[0070] Among them, the search scenario can be one of multiple preset scenarios. Exemplarily, the multiple preset scenarios can include but are not limited to obstacle scenarios and road scenarios, etc. For example, if the search scenario is an obstacle scenario, calculate the similarity between the picture search text and the target description texts of each picture related to the obstacle scenario; if the search scenario is a road scenario, calculate the similarity between the picture search text and the target description texts of each picture related to the road scenario. Optionally, the search scenario of the picture search text can be determined by extracting the keywords in the picture search text and matching them with the keywords corresponding to each preset scenario, and the preset scenario that is hit is used as the search scenario of the picture search text. It should be noted that the scenario to which each picture belongs can be pre-marked.

[0071] In this embodiment, first, the search scenario of the picture search text is identified, and then the similarity between the target description texts previously determined for each of the multiple pictures that match the picture search text and the search scenario information is compared, that is, the similarity is matched with the target description texts of some pictures. When new pictures that match the picture search text need to be displayed, the similarity is only matched with the remaining pictures that have not been displayed. This can improve the efficiency of similarity matching and save unnecessary resource waste.

[0072] It should be noted that to identify the search scenario of the picture search text, the trained generative language model can be used to output the search scenario of the picture search text. For example, the picture search text and the prompt text for the generative language model to output the search scenario are provided, so that the generative language model outputs the search scenario based on the picture search text and the prompt. Optionally, it can also be implemented by keyword matching. For example, multiple keywords are preset, and each keyword corresponds to a search scenario. If the picture search text hits at least one of the keywords, the scenario corresponding to the hit keyword is used as the search scenario of the picture search text.

[0073] In another possible implementation, it can also be to perform similarity matching between the picture search text and all target description texts.

[0074] The above embodiments illustrate the method embodiments. The following embodiments illustrate the product embodiments by way of example.

[0075] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a picture search device provided by an embodiment of the present invention. As Figure 2 shown, this picture search device is applied to a terminal device. The terminal device can be the local terminal device described above or the client device described above. As Figure 2 shown, the device may include:

[0076] An acquisition module 210 is configured to acquire text for image search; a similarity comparison module 220 is configured to compare the text for image search with pre-determined target description texts of each of multiple images to obtain the similarity between the text for image search and the target description texts of each image, wherein the target description text of an image is used to describe the content of the image, and the target description text of the image is the text output by a trained generative language model after the image and a prompt text are input into the generative language model, and the prompt text is used to prompt the generative language model to output the target description text; an image search module 230 is configured to display a first target image among the multiple images based on the similarity between the text for image search and the target description texts of each image, and the similarity corresponding to the first target image is higher than that of a second target image not displayed among the multiple images.

[0077] The image search device provided by an embodiment of the present invention has the same technical features as the image search method provided by the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0078] This embodiment further provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above image search method. The electronic device can be a server or a terminal device.

[0079] See Figure 3 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100, and the processor 100 executes the computer-executable instructions to implement the above image search method.

[0080] Furthermore, Figure 3 the electronic device shown further includes a bus 102 and a communication interface 103, and the processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0081] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is established between the system network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0082] Processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in memory 101, and processor 100 reads the information in memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0083] The processor in the above electronic device can implement the steps in the above picture search method by executing computer-executable instructions.

[0084] This embodiment also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the above picture search method.

[0085] The computer-executable instructions stored in the above computer-readable storage medium can implement the steps in the above picture search method by executing the computer-executable instructions.

[0086] This embodiment also provides a computer program product, including program code. The instructions included in the program code can be used to execute the method in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated here.

[0087] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0088] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0089] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0090] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0091] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for searching an image, characterized in that: include: Get the image search text; Comparing the image search text with a predetermined target description text of each of the multiple images to obtain a similarity between the image search text and the target description text of each image, wherein the target description text of the image is used to describe the content of the image, the target description text of the image is the text output by the generative language model after the image and the prompt text are input into a trained generative language model, and the prompt text is used to prompt the generative language model to output the target description text; Based on the similarity between the picture search text and the target description text of each picture, a first target picture among the multiple pictures is displayed, and the similarity corresponding to the first target picture is higher than that of a second target picture among the multiple pictures that is not displayed.

2. The method according to claim 1, characterized in that Before comparing the image search text with a predetermined target description text of each of the multiple images to obtain the similarity between the image search text and the target description text of each image, the method further includes: Preprocessing the image search text, wherein the preprocessing includes at least one of converting the text in the image search text into standard text, removing invalid text in the image search text, or replacing the text in the image search text with synonyms; The comparing the similarity between the image search text and a predetermined target description text of each of the multiple images includes: The preprocessed image search text is compared with a predetermined target description text of each of the multiple images to obtain a similarity between the image search text and the target description text of each image.

3. The method according to claim 1 or 2, characterized in that: The comparing the image search text with a predetermined target description text of each of the multiple images to obtain the similarity between the image search text and the target description text of each image includes: Performing vectorization processing on the image search text and the target description text of each image respectively, to obtain a first text vector corresponding to the image search text and a second text vector corresponding to the target description text of each image; Based on the first text vector and the second text vector corresponding to the target description text of each picture, the similarity between the picture search text and the target description text of each picture is obtained.

4. The method according to claim 3, characterized in that The vectorization processing is performed on the image search text and the target description text of each image to obtain a first text vector corresponding to the image search text and a second text vector corresponding to the target description text of each image, including: The image search text and the target description text of each image are respectively vectorized using the trained embedding model to obtain a first text vector corresponding to the image search text and a second text vector corresponding to the target description text of each image.

5. The method according to claim 4, characterized in that The step of vectorizing the target description text of each image by using the trained embedding model to obtain a second text vector corresponding to the target description text of each image includes: splicing the target description texts of each image to obtain a spliced ​​description text; Using the trained embedding model to vectorize the concatenated description text, to obtain a text vector corresponding to the concatenated description text; The text vectors corresponding to the spliced ​​description texts are divided according to the splicing order and length of the target description texts of each picture to obtain second text vectors corresponding to the target description texts of each picture.

6. The method according to claim 3, characterized in that The obtaining the similarity between the image search text and the target description text of each image based on the first text vector and the second text vector corresponding to the target description text of each image includes: Determine a vector distance between the first text vector and a second text vector corresponding to the target description text of each picture, wherein the vector distance is used to represent a similarity between the first text vector and the second text vector; Based on the vector distance between the first text vector and the second text vector corresponding to the target description text of each picture, the similarity between the picture search text and the target description text of each picture is obtained.

7. The method according to claim 6, characterized in that The vector distance includes at least two items, and the at least two vector distances include at least two items of Euclidean distance, dot product similarity distance, cosine similarity or maximum inner product similarity. The similarity between the image search text and the target description text of each image is obtained based on the vector distance between the first text vector and the second text vector corresponding to the target description text of each image, including: Normalizing at least two of the vector distances to obtain normalized at least two of the vector distances; Fusing at least two normalized vector distances to obtain a fused vector distance; Based on the fused vector distance, the similarity between the image search text and the target description text of each image is obtained.

8. The method according to claim 1 or 2, characterized in that: Before comparing the image search text with the predetermined target description text of each of the multiple images for similarity, the method includes: Identify the search scene of the image search text and obtain the search scene information; The comparing the similarity between the image search text and a predetermined target description text of each of the multiple images includes: Performing a similarity comparison between the image search text and a predetermined target description text of each image in a plurality of images matching the search scene information; After displaying the first target picture among the multiple pictures, the method further includes: In response to an instruction to switch the displayed picture, the picture search text is compared with the target description text of each picture in the second target picture for similarity.

9. A picture search device, characterized in that: include: Acquisition module, used to obtain image search text; A similarity comparison module is used to compare the image search text with a predetermined target description text of each image in a plurality of images to obtain the similarity between the image search text and the target description text of each image, wherein the target description text of the image is used to describe the content of the image, the target description text of the image is the text output by the generative language model after the image and the prompt text are input into the trained generative language model, and the prompt text is used to prompt the generative language model to output the target description text; The image search module is used to display a first target image among the multiple images based on the similarity between the image search text and the target description text of each image, wherein the similarity corresponding to the first target image is higher than that of a second target image that is not displayed among the multiple images.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 8.