Image search method, device, intelligent agent, electronic device and storage medium
Through a multimodal image search method, a large text analysis model is used to generate descriptive information of various semantic granularities, which solves the problems of high system complexity and poor user experience in existing technologies and realizes flexible image search in cross-task scenarios.
Patent Information
- Application Number
- CN202411764535.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Existing image search methods based on large models have high system complexity and maintenance costs due to the diverse input forms of different tasks. They are unable to flexibly respond to cross-task and complex dynamic image search needs, and the user experience is poor.
A multimodal image search method is adopted, and a large text analysis model is used to perform text analysis on the input text information and the description information of the reference image, generating description information of multiple semantic granularities to determine the target image and uniformly process input information of multiple tasks and modalities.
It reduces system complexity and maintenance costs, improves the flexibility and accuracy of image search, and meets users' image search needs in various scenarios.
Smart Images

Figure CN119597954B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly computer vision, deep learning, large models, image search, and other technical fields, and can be applied to scenarios such as AIGC-based content generation based on artificial intelligence. Specifically, it relates to an image search method, device, intelligent agent, electronic device, and storage medium. Background Art
[0002] With the continuous development of artificial intelligence technology, large model technology has also been applied in various fields. For example, large models are used to achieve image search.
[0003] However, when performing image search based on large models, image search with multiple modal inputs corresponds to multiple processing methods, resulting in high system complexity and maintenance costs. It is also difficult to implement image search in complex scenarios, such as flexibly switching between multiple modal inputs for image search. Summary of the Invention
[0004] The present disclosure provides an image search method, device, intelligent agent, electronic device and storage medium.
[0005] According to one aspect of the present disclosure, an image search method is provided, comprising: determining multimodal search information based on input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image; performing text analysis on the second text information and first description information for describing the second reference image using a text analysis large model to generate at least one second description information; and determining at least one target image based on the at least one second description information, wherein each target image is determined based on the at least one second description information.
[0006] According to another aspect of the present disclosure, an image search device is provided, including: a first determination module for determining multimodal search information based on input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image; a generation module for performing text analysis on the second text information and the first description information for describing the second reference image using a text analysis large model to generate at least one second description information; and a second determination module for determining at least one target image based on the at least one second description information, wherein each target image is determined based on the at least one second description information.
[0007] According to another aspect of the present disclosure, an artificial intelligence agent is provided, which is configured to execute the method provided by the embodiment of the present disclosure.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the above method.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above method when executed by a processor.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0013] Figure 1 Schematically illustrates an exemplary system architecture to which the image search method and apparatus according to an embodiment of the present disclosure can be applied;
[0014] Figure 2 The following schematically shows a flow chart of an image search method according to an embodiment of the present disclosure;
[0015] Figure 3 A schematic diagram of a scenario for determining search information according to an embodiment of the present disclosure is shown schematically;
[0016] Figure 4 Schematically shows a scenario diagram of obtaining input information according to an embodiment of the present disclosure;
[0017] Figure 5 A schematic diagram of a scenario for determining at least one target image according to an embodiment of the present disclosure is shown schematically;
[0018] Figure 6A A schematic diagram schematically shows a scenario of generating second description information according to an embodiment of the present disclosure;
[0019] Figure 6B Schematically shows a scenario diagram of generating second description information according to another embodiment of the present disclosure;
[0020] Figure 7A schematic diagram of a scenario of determining a target image from a search image library according to a specific embodiment of the present disclosure is shown schematically;
[0021] Figure 8 A schematic diagram of a scene for determining a target image according to a specific embodiment of the present disclosure is schematically shown;
[0022] Figure 9 A block diagram of an image search device according to a specific embodiment of the present disclosure is schematically shown;
[0023] Figure 10 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and
[0024] Figure 11 A block diagram of an electronic device 1100 suitable for implementing an image search method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Currently, large-scale model-based image search tasks include text-based image search tasks, composed image retrieval (CIR) tasks, and conversational image retrieval (Chat-IR) tasks. Their inputs are: text input in pure language mode, multimodal input of reference image + text instructions, and text input obtained by combining multiple rounds of conversations.
[0027] On the one hand, due to the diverse input forms required by different tasks, traditional single-modality image search methods require the design of different model architectures and optimization strategies for each task, increasing system complexity and computational costs. Furthermore, this separate design increases system complexity and requires the system to adapt to different tasks, adding additional development and maintenance costs.
[0028] On the other hand, as user needs diversify, image search usage scenarios are often no longer limited to a single task model. In complex application scenarios, users' image search needs may be cross-task and dynamically changing. For example, at the beginning, users may only want to perform a simple image search based on text, but as the interaction deepens, they may require the system to combine a reference image or further refine the search results through dialogue. Existing image search methods, due to their design separation of tasks, are unable to flexibly cope with such cross-task, complex and dynamic needs, resulting in a poor image search experience in complex scenarios.
[0029] In order to at least partially solve the above-mentioned technical problems, an embodiment of the present disclosure provides an image search method, comprising: determining multimodal search information based on input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image. Using a large text analysis model, text analysis is performed on the second text information and the first description information used to describe the second reference image to generate at least one second description information. Based on the at least one second description information, at least one target image is determined, wherein each target image is determined based on the at least one second description information. Through the embodiments of the present disclosure, the technical problems of high system complexity and maintenance costs and poor search experience can be at least partially solved, and the technical effects of reducing system complexity and maintenance costs and improving search experience can be achieved.
[0030] Figure 1 An exemplary system architecture to which the image search method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0031] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the image search method and apparatus may be applied may include a terminal device, but the terminal device may implement the method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0032] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0033] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0034] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0035] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0036] The server can be a cloud server, also known as a cloud computing server or cloud host. It is a hosting product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server can also be a distributed system server or a server integrated with blockchain.
[0037] It should be noted that the image search method provided in the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Accordingly, the image search apparatus provided in the embodiments of the present disclosure can also be provided in the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0038] Alternatively, the image search method provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the image search apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The image search method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the image search apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0039] For example, a user inputs input information for image search through the first terminal device 101, the second terminal device 102, and the third terminal device 103, and the first terminal device 101, the second terminal device 102, and the third terminal device 103 determine multimodal search information based on the input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image; using a text analysis large model, text analysis is performed on the second text information and the first description information for describing the second reference image to generate at least one second description information; and based on the at least one second description information, at least one target image is determined, wherein each target image is determined based on the at least one second description information.
[0040] Alternatively, input information for image search is sent to the server 105 through the first terminal device 101, the second terminal device 102, and the third terminal device 103, and the above-mentioned image search method is performed with the help of the server 105 to determine and return at least one searched target image to the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0041] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0042] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0043] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0044] Figure 2The flowchart of the image search method according to the embodiment of the present disclosure is schematically shown.
[0045] like Figure 2 As shown, the method 200 includes operations S210 to S230.
[0046] In operation S210, multimodal search information is determined based on input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image.
[0047] The input information can be in the form of a single modality, such as first text information in a single language modality, or a first reference image in a single visual modality. Alternatively, the input information can be multimodal input information, such as first text information and a first reference image. For example, the input information can be pure text input for a text-image search task, a reference image and text instruction input for a combined image search task, or a combined text input for multiple rounds of dialogue in a conversational image search task.
[0048] For example, the input information includes first text information, such as "a girl is playing with a white cat"; or, the input information includes a first reference image, such as an image representing "a girl is playing with a black cat"; or, the input information includes the first text information and the first reference image, the first reference image is an image representing "a girl is playing with a black cat", and the first text information is "change the black cat to a white cat".
[0049] The search information is in a multimodal form and includes both the second text information and the second reference image. It is understood that the search information may be information obtained by normalizing input information in multiple input forms. Normalization may include unification of modality, form, size, and other aspects. For example, in terms of modality unification, a single or multimodal input form is processed into a multimodal form including the second text information and the second reference image, thereby obtaining multimodal search information.
[0050] For example, the first text information may be directly used as the second text information, the first text information may be processed to obtain the second text information, or the second text information may be of a predetermined type. For example, the first text information may be modified, converted into a sentence structure, segmented, or information extracted to obtain the second text information.
[0051] For another example, the first reference image can be directly used as the second reference image, the first reference image can be processed to obtain the second reference image, or a predetermined type of second reference image can be used. For example, the first reference image can be subjected to processing such as cropping, rotation, and color correction to obtain the second reference image.
[0052] In operation S220, a text analysis model is used to perform text analysis on the second text information and the first description information used to describe the second reference image, thereby generating at least one second description information.
[0053] The first description information of the second reference image is used to describe the content of the second reference image. For example, the first description information is used to describe linguistically meaningful information such as elements, attributes, and spatial relationships within the second reference image. Elements can be objects, such as people, animals, plants, and objects. Attributes can be characteristics that distinguish elements, such as color, shape, and size. Spatial relationships can be understood as the relative positional relationships between elements.
[0054] The text analysis large model may be a large language model (LLM) for processing language modalities. In this embodiment, the text analysis large model may be a pre-trained large language model.
[0055] By using the large text analysis model, the first description information and the second text information can be semantically analyzed from the text perspective to determine the second description information by integrating the semantics of the two parts. The second description information is also the description information of the image that meets the user's image search needs.
[0056] In one embodiment, a large text analysis model can be used to perform a single text analysis on the second text information and the first description information to obtain a single second description information. Alternatively, the large text analysis model can be used to perform multiple text analyses on the same information to obtain a second description information that meets the user's image search requirements and has multiple expression forms.
[0057] In another embodiment, the text analysis model may be used to perform a single text analysis on the second text information and the first description information, generating at least one second description information in at least one form. Alternatively, at least one second description information in at least one semantic granularity may be generated.
[0058] For example, taking the first reference image as an image representing "a girl playing with a black cat" and the first text information as "change the black cat to a white cat", the second description information output by the text analysis model can be "a girl playing with a white cat".
[0059] In operation S230 , at least one target image is determined based on the at least one second description information, wherein each target image is determined based on the at least one second description information.
[0060] The at least one second description information may be regarded as description information of an image that meets the user's image search requirements. Therefore, all or part of the at least one second description information may be combined to search for a target image.
[0061] For example, a candidate image is searched for according to each second description information, and at least one target image is determined according to the number of second description information hit by the searched candidate images.
[0062] Alternatively, the similarity between each second description information and the candidate image is calculated, and the similarities between one or more second description information and the same candidate image are combined to determine whether to use the candidate image as the target image.
[0063] In the embodiment of the present disclosure, since input information in various forms is unified into multimodal search information, this embodiment can perform image search on input information of various tasks, cross-task scenarios, and dynamic changes in modality. There is no need to design multiple processing flows in a targeted manner, which simplifies the complexity of the system and can flexibly respond to different tasks to meet the needs of users in various image search scenarios. In addition, the second text information and the first description information are semantically analyzed using a large text analysis model to generate at least one second description information, and each target image is determined using at least one second description information. This can fully utilize the rich second description information output by the large text analysis model, thereby improving the accuracy of image search.
[0064] Figure 3 A schematic diagram of a scenario for determining search information according to an embodiment of the present disclosure is schematically shown.
[0065] like Figure 3 As shown, embodiment 300 includes operations S301 to S306, which can be used as a specific embodiment of operation S210.
[0066] In operation S301, it is determined whether the input information includes the first text information. After the input information is acquired, whether the input information includes the first text information or not, operation S302 is performed to further determine whether the input information includes the first reference image.
[0067] In operation S302, it is determined whether the input information includes the first reference image. If the input information includes the first text information but does not include the first reference image, operation S303 is performed; if the input information includes both the first text information and the first reference image, operation S304 is performed; if the input information does not include the first text information but includes the first reference image, operation S305 is performed; if the input information does not include the first text information or the first reference image, operation S306 is performed.
[0068] In operation S303 , the first text information is determined as second text information, and a second reference image is acquired, where the second reference image includes a blank reference image.
[0069] The input information only includes the first text information. In this case, the input information is input in a single language mode. The first text information can be used as the second text information and together with the predetermined second reference image constitute multimodal search information.
[0070] In one embodiment, the blank reference image may be an image that does not include any content. For example, the first description information of the blank reference image may be “blank image”.
[0071] In operation S304, the first text information is determined as the second text information, and the first reference image is determined as the second reference image.
[0072] The input information includes the first text information and the first reference image at the same time. In this case, the input information is multimodal input and can be directly used as search information.
[0073] For example, if the input information includes a first reference image representing "a girl playing with a black cat" and a first text message "change the black cat to a white cat," the multimodal search information would be: the image representing "a girl playing with a black cat" and "change the black cat to a white cat."
[0074] In operation S305 , the first reference image is determined as a second reference image, and second text information is acquired, where the second text information includes blank text information.
[0075] The input information only includes the first reference image. In this case, the input information is input in a single visual modality. The first reference image can be used as the second reference image and together with the predetermined second text information constitute multimodal search information.
[0076] In one embodiment, the blank text information may be empty information including no text, or may include predetermined text information representing a no-text instruction. For example, the blank text information may be “empty”, “no-text instruction”, etc.
[0077] In operation S306, an error is reported. If the input information does not include the first text information and the first reference image, it can be regarded as an erroneous input and an error is reported.
[0078] In one embodiment, the prompt information of the text analysis model may include two placeholders for language modality and visual modality. The first text information and / or first reference information included in the input information is first added to the placeholder as the second text information and / or second reference image; for the placeholder with missing content, the blank reference image and / or blank text information is added as the second reference image and / or second text information.
[0079] In an embodiment of the present disclosure, for input text information of a single language modality or a single visual modality, by filling in the blank second reference image or second text information, the image search task is changed from a single modality input form to a unified multimodal input without affecting the semantic content, so that the text analysis model can be used to perform semantic analysis under the same input, thereby achieving compatibility of multiple image search tasks and reducing system complexity and maintenance costs.
[0080] In a specific embodiment, the input information may also be information in other modalities besides language and visual modalities. The other modal information may be converted into first text information in the language modality and / or a first reference image in the visual modality, thereby implementing an image search task compatible with other modalities.
[0081] For example, for voice information in audio mode, the voice information can be converted into first text information through a voice-to-text processing algorithm; then, the search information is determined based on the input information.
[0082] In the embodiments of the present disclosure, by converting other modal information input by the user into input information, the image search method of the above embodiment is not only applicable to a wider range of usage scenarios, but also does not require system process transformation for the newly added modal information. The processing process of multi-modal search information can be reused, thereby improving the scalability and flexibility of the system.
[0083] According to an embodiment of the present disclosure, for operation S210, the image search method further includes: if the input information includes an indication image library, determining the indication image library as the search image library to determine the target image from the search image library. If the input information does not include the indication image library, determining the reference image library as the search image library.
[0084] For example, the indicated image library may be one of the image libraries included in the system for performing the image search method. The user may limit the search scope by adding the indicated image library to the input information. If the input information does not include the indicated image library, the system may use a default reference image library.
[0085] For example, in the language-guided image search scenario, the input of the image search is a scoring function ,in, The second text information representing the language mode, A second reference image representing the visual modality, To search the image library, , such as the search image library can include N candidate images. In the text image search task, Can be , in the combined image search task, Can be , in the conversational image search task, Can be , , represents multiple rounds of dialogue, The text information that constitutes a round of dialogue.
[0086] In the embodiments of the present disclosure, by determining whether the input information includes a reference image library, the need to limit the search scope can be met, expanding the use cases and providing a better user experience. In addition, for input without a limited search scope, a predetermined reference image library can also be used for subsequent image searches, meeting the user's various image search needs.
[0087] According to an embodiment of the present disclosure, the input information is obtained based on at least one of the following methods: determining the first text information and / or the first reference image based on the input text and input image in at least one round of conversation; determining the first reference image and / or the first text information based on at least one output image and / or input text in at least one round of conversation, wherein the output image is a target image determined and output based on the input information before the output image.
[0088] For ease of understanding, four embodiments are used as examples below to illustrate input information obtained in a language-guided image search scenario.
[0089] Figure 4 Schematically shows a scenario diagram of obtaining input information according to an embodiment of the present disclosure. Figure 4 As shown, it includes embodiment 400A, embodiment 400B, embodiment 400C and embodiment 400D.
[0090] In embodiment 400A, input text 411 is obtained through a conversation, such as "A woman walks on a road covered with fallen leaves and bathed in sunlight," and input text 411 is determined as first text information. In this embodiment, the above-described image search method is used to determine target image 1 412 based on the input information including the first text information.
[0091] Similarly, an input image can be obtained through a round of dialogue and used as the first reference image in the input information to search for the corresponding target image based on the current input.
[0092] In embodiment 400B, the specific requirements for the image search are clarified through two rounds of dialogue. For example, the specific requirements for the image search are clarified through input text 421, system output question 422, and input text 423. In this case, the first text information can be obtained by combining input text 421 and input text 423, such as "A woman walks on a road. There are fallen leaves on the road, and the sun shines on the road." The system output question 422 can be a question generated using the conversational macro model to clarify the requirements for the road, such as "What kind of road is it?" In this embodiment, the above-mentioned image search method can be used to determine target image 2 424 based on "A woman walks on a road. There are fallen leaves on the road, and the sun shines on the road."
[0093] In embodiment 400C, input information including first text information and a first reference image can be obtained from input text 431 and input image 432 during a conversation. In this case, the first text information is also input text 431, such as "Change the road in the image below to one covered with fallen leaves and bathed in sunlight." In this embodiment, target image 3 433 is determined based on the input information including the first text information and the first reference image.
[0094] In embodiment 400D, during a conversation, a first text message, such as "A woman walks on a road covered with fallen leaves," is determined based on input text 441. The aforementioned image search method is used to determine a target image based on the input information including the first text message and output it as a result, such as output image 442. In this embodiment, the user originally only wishes to conduct a simple image search based on text, but as the interaction deepens, the user initiates a second conversation, hoping to further optimize the image search results by combining a reference image. In this second conversation, input text 443 may be "Based on the output image above, the road should be filled with sunshine." At this point, based on input text 443, output image 442 can be used as the first reference image. The aforementioned image search method can be used again to determine a new target image 444, such as target image 4, based on the input information including the first text message and the first reference image.
[0095] In the embodiments of the present disclosure, for various image search tasks, cross-task scenarios, and input information with dynamically changing modalities, the input information can be unified into multimodal search information to be applicable to a wider range of scenarios, meet more user image search needs, and improve the user experience.
[0096] According to an embodiment of the present disclosure, for operation S220, the text analysis big model is used to perform the first description generation task, and the text analysis big model is used to perform text analysis on the second text information and the first description information used to describe the second reference image to generate at least one second description information, including: using the text analysis big model to perform the first description generation task to perform semantic analysis on the second text information and the first description information to generate second description information of multiple semantic granularities; wherein the semantic granularity is related to the number and / or attributes of elements in the second description information, and the elements are extracted from the second text information and / or the first description information.
[0097] The first description generation task includes: understanding the second text information, and generating second description information of multiple semantic granularities that meet the user's image search requirements based on the first description information.
[0098] A single user input text may contain multiple semantics, such as operations on multiple elements, multiple operations on an element, and multiple descriptions of an element. Therefore, a large text analysis model can be used to perform the first description generation task and generate second description information at multiple semantic granularities, thereby generating clear and rich second description information.
[0099] Second description information of multiple semantic granularities can be understood as multiple second description information including various amounts of information, where the amount of information can be the number of elements and / or attributes included. The fewer the number of elements and / or attributes, the smaller the semantic granularity; conversely, the greater the number, the larger the semantic granularity.
[0100] For example, taking three semantic granularities as an example, the first semantic granularity, the second semantic granularity, and the third semantic granularity are Core Elements (CE), Enhanced Details (ED), and Comprehensive Synthesis (CS), respectively. The first semantic granularity may include only elements appearing in the second text information, without the use of attributes, where attributes can be regarded as adjectives, adverbials, attributives, etc. in the second text information. The second semantic granularity may include only elements appearing in the second text information, and use necessary attributes from the second description information. The third semantic granularity may include elements appearing in the second text information and related elements in the first description information, and use necessary attributes. It can be understood that the first semantic granularity, the second semantic granularity, and the third semantic granularity gradually increase.
[0101] For example, in one example, the second text information is "The person should be holding a baby," and the first description information of the second reference image is "A woman with black hair is smiling under a gray umbrella with a white flower hanging from it." The second description information of the first semantic granularity generated by the text analysis model is: "A woman holding a baby," the second description information of the second semantic granularity is: "A woman with black hair is holding a baby under an umbrella," and the second description information of the third semantic granularity is: "A woman with black hair is holding a baby, smiling under a gray umbrella." The second description information of the first semantic granularity only includes elements in the second text information, such as "person" and "baby"; the second description information of the second semantic granularity only includes the elements "person" and "baby," and the corresponding attributes "black hair" and "under an umbrella"; the second description information of the third semantic granularity includes the elements "person" and "baby," as well as the attribute "gray" of the related element "umbrella" in the first description information.
[0102] In an embodiment of the present disclosure, a large text analysis model is used to perform a first description generation task, and second description information of multiple semantic granularities is generated for the same input, thereby enriching the semantic hierarchy of the second description information, facilitating the subsequent use of the second description information of multiple semantic granularities, and determining the target image in a wider range, thereby helping to improve the accuracy of image search.
[0103] Figure 5 A schematic diagram of a scene for determining at least one target image according to an embodiment of the present disclosure is schematically shown.
[0104] like Figure 5 As shown, in embodiment 500, second text information 502 and second reference image 503 are determined according to whether input information 501 includes first text information and / or first reference image information.
[0105] After obtaining first description information 504 describing second reference image 503, text analysis is performed on first description information 504 and second text information 502 using text analysis model M1 to obtain at least one second description information, such as second description information 1 505-1…second description information M 505-M. Using second description information 1 505-1…second description information M 505-M, at least one target image, such as target image 1 507-1…, can be obtained from search image library 506.
[0106] According to an embodiment of the present disclosure, for operation S220, a first description generation task is performed using a text analysis big model to perform semantic analysis on the second text information and the first description information to generate second description information of multiple semantic granularities, including: obtaining prompt information, wherein the prompt information includes multiple semantic granularities to be generated, and explanation information for each semantic granularity; and performing the first description generation task using a text analysis big model to perform semantic analysis on the second text information and the first description information based on each semantic granularity to generate second description information under each semantic granularity.
[0107] For example, prompt information A may be:
[0108] # Task description: You will be provided with a description about image search. The first description generation task is to combine the second text information with the information in the second reference image or the first description information to accurately search for images. # First description generation task: Based on the text instructions and reference image analysis, describe what the target image should look like. Provide three sentences to describe the target image, each focusing on a different semantic granularity: (1) The first semantic granularity is core elements: it can be only elements that appear in the second text information, without using attributes. (2) The second semantic granularity is enhanced details: only elements that appear in the second text information are included, and necessary attributes are used from the second description information. (3) The third semantic granularity is comprehensive synthesis: including elements that appear in the second text information and related elements in the first description information, and using necessary attributes. I will give you a second text information and a first description information, and you need to complete the task based on this information. ### Query: Second text information [[instructions]], first description information [[reference image description]].
[0109] In this embodiment, [[Instructions]] and [[Reference Image Description]] are placeholders. "Core Elements," "Enhanced Details," and "Fully Synthesized" are the names of the semantic granularities, and "May only include elements appearing in the second text information, without the use of attributes" is the explanation information for the corresponding semantic granularity.
[0110] For another example, the prompt information may include not only the above content, but also task examples, so that the text analysis model can better perform the first description generation task.
[0111] In an embodiment of the present disclosure, by writing semantic granularity and explanation information of the semantic granularity in the prompt information, the first description generation task is performed using a text analysis model to perform semantic analysis on the second text information and the first description information based on each semantic granularity, and generate second description information under each semantic granularity, thereby generating rich second description information.
[0112] by Figure 5For example, the text analysis model M1 can perform the first description generation task, and the obtained second description information 1 507 - 1 . . . belongs to multiple semantic granularities.
[0113] According to an embodiment of the present disclosure, for operation S220, the text analysis big model is used to sequentially execute the operation generation task and the second description generation task; using the text analysis big model, the second text information and the first description information used to describe the second reference image are subjected to text analysis to generate at least one second description information, including: using the text analysis big model to execute the operation generation task to generate operation prompt information based on the difference between the first description information and the second text information, the operation prompt information including at least one operation of at least one operation type; and using the text analysis big model to execute the second description generation task to generate at least one second description information based on the operation prompt information and the first description information.
[0114] The text entered by the user may contain multiple semantics. For example, if the second text information includes multiple elements, multiple attributes, and multiple operations, using a large text analysis model to directly generate the second description information based on the entire second text information may result in inaccurate second description information. Therefore, in this embodiment, the process of generating the second description information is split into two tasks: the operation generation task and the second description generation task.
[0115] The operation generation task may be a classification task, which is used to compare the first description information and the second text information, and determine the operation to be performed on the first description information and the type of operation based on the difference between the two.
[0116] The operation prompt information can be a benchmark for executing the second description generation task, and is used to prompt the text analysis model to execute the second description generation task and generate at least one second description information.
[0117] The operation type in the operation prompt information may include: addition type, removal type, modification type, comparison type, and retention type.
[0118] In one embodiment, when the text analysis large model sequentially executes the operation generation task and the second description generation task, the prompt information used may include the operation type and explanation information of the operation type.
[0119] For the second reference image described by the second descriptive information, the explanatory information of the operation type can be: addition can be understood as introducing new elements or attributes into the second reference image; deletion can be understood as deleting certain elements or attributes from the second reference image; modification can be understood as changing the attributes of existing elements in the second reference image; comparison uses words such as "different", "same", "more" or "less" to compare the elements in the second reference image and the first text information; retention indicates that certain existing elements or attributes in the second reference image remain unchanged.
[0120] The operation prompt information obtained after executing the operation to generate the character may include second text information, in which the operation type is newly added.
[0121] For example, in one embodiment, the second textual information is "The character should be holding a baby," and the first description of the second reference image is "A dark-haired woman smiling under a gray umbrella with a white flower hanging from it." The action prompt is "Add: Woman holding a baby," and the action involves adding "Woman holding a baby" to the second reference image. The action type is "Add." The action prompt includes "Woman holding a baby," which includes the second textual information "The character should be holding a baby," and is a more specific and accurate description.
[0122] The second description generation task is similar to the first description generation task, but performs a specific operation according to the operation type specified in the operation prompt information, which will not be repeated here.
[0123] In the embodiments of the present disclosure, a large text analysis model is used to sequentially execute an operation generation task and a second description generation task to generate semantically more explicit operation prompt information. At least one second description information is then generated based on the operation prompt information and the second text information. Because these two tasks extract more specific and fine-grained operations and operation types from the complex second text information, the generated second description information better meets image search requirements, helping to improve image search accuracy.
[0124] Figure 6A The following schematically shows a scenario diagram of generating the second description information according to an embodiment of the present disclosure. Figure 6A As shown, in embodiment 600A, the text analysis model M1 sequentially executes the operation generation task T1 and the second description generation task T2.
[0125] For example, the text analysis model M1 is used to perform operation generation task T1 to generate operation prompt information 603 based on the second text information 601 and the first description information 602. The operation prompt information 603 may include at least one operation of at least one operation type, such as operation 1 6031-1 ... operation n 6031-n of operation type 1 6031.
[0126] Taking operation 1 6031 - 1 as an example, the second description generation task T2 is performed using the text analysis model M1 to obtain second description information 1 6041 based on operation 1 6031 - 1 and the first description information 602 .
[0127] In other embodiments, the operation prompt information may include only one operation of one operation type.
[0128] According to an embodiment of the present disclosure, when the second reference image is a blank reference image, the operation type of the at least one operation included in the operation prompt information is all an adding type.
[0129] For a single-language image search task, the input information consists solely of the first textual information, and the second reference image for the search information determined based on the input information is a blank reference image. In this case, the difference between the second textual information and the first description of the second reference image constitutes the entire content of the second textual information. Therefore, the elements, attributes, and operations on elements and / or attributes in the second textual information are all of the add type.
[0130] For example, the second textual information could be "A person holding a baby, the baby is wearing pink clothes," and the second reference image could be a blank reference image. The operation prompt information generated by executing the operation generation task could be: 1. Add: A person holding a baby; 2. Add: The baby is wearing pink clothes. The second description information generated based on the operation prompt information and the first description information could be "A person holding a baby wearing pink clothes."
[0131] If the second text information is blank, the operation type of the operation included in the operation prompt information can be a retain type and / or a comparison type. For example, if the second text information is blank, the first description information of the second reference image may be "A person holding a baby, the baby is wearing pink clothes." The operation prompt information obtained by executing the operation generation task may be: 1. Retain: A person holding a baby; 2. Retain: A baby is wearing pink clothes. The second description information generated based on the operation prompt information and the first description information may also be "A person holding a baby wearing pink clothes."
[0132] In the embodiments of the present disclosure, for a blank second reference image or second text information, the image search task can be performed normally using a large text analysis model, and a target image with semantics similar to the non-blank second text information or second reference image can be obtained, which has a wide range of usage scenarios.
[0133] According to an embodiment of the present disclosure, the text analysis big model is used to sequentially execute the operation generation task and the third description generation task, and the text analysis big model is used to perform text analysis on the second text information and the first description information used to describe the second reference image to generate at least one second description information, including: using the text analysis big model to execute the operation generation task to generate operation prompt information based on the difference between the first description information and the second text information; and using the text analysis big model to execute the third description generation task to generate second description information of multiple semantic granularities based on the operation prompt information and the first description information.
[0134] The operation generation task is similar to the operation generation task above, and the third description generation task is similar to the second description generation task above. It generates second description information based on the operation prompt information and the first description information. Similarly, when generating the second description information, the third description generation task, like the first description generation task, generates second description information at multiple semantic granularities. The specific operations are described above and will not be repeated here.
[0135] When the text analysis large model sequentially executes the operation generation task and the third description generation task, the prompt information used may include multiple semantic granularities to be generated and explanation information for each semantic granularity; at the same time, it also includes the operation type and explanation information of the operation type.
[0136] For example, the prompt information B used by the text analysis large model to sequentially execute the operation generation task and the third description generation task can be:
[0137] # Task description: You will be given a description about image search. The task is to combine the second text information with the information in the second reference image or the first description information to accurately search for the image. You need to follow two steps to infer "what the target image looks like". # Step 1: Classify the given second text information into the following operation types and determine how it affects the second reference image. For each operation type, determine the elements or attributes of the second reference image that are affected. Operation types include: (1) Addition: Introducing new elements or attributes into the second reference image. Determine which existing element the addition is related to or where it should be placed. (2) Deletion: Delete certain elements from the second reference image. Determine which existing element is deleted. (3) Modification: Change the attributes of existing elements in the second reference image. Determine which elements are being modified and how they are being modified. (4) Comparison: Contrast elements in the second reference image using terms such as "different", "same", "more", or "less". Determine the elements and attributes being compared. (5) Preservation: Specify that certain existing elements in the second reference image remain unchanged. Make sure these elements are marked as included in the target image. # Step 2: ... . Among them, "#Step 2:..." is similar to the #First description generation task in the prompt information A above, and will not be repeated here.
[0138] In an embodiment of the present disclosure, by sequentially executing the operation generation task and the third description generation task, the text analysis large model can generate more accurate and semantically richer second description information through hierarchical execution of tasks, thereby improving the accuracy of the target image subsequently determined using the second description information with multiple semantic granularities.
[0139] Figure 6B Schematically shows a scenario diagram of generating second description information according to another embodiment of the present disclosure. Figure 6B As shown, in embodiment 600B, the text analysis model M1 sequentially executes the operation generation task T1 and the third description generation task T3, wherein the manner of generating operation 1 6031-1…operation n 6031-n is similar to that of embodiment 600A and will not be repeated here.
[0140] Taking operation 1 6031-1 as an example, by using the text analysis model M1 to perform the third description generation task T3, multiple second description information of multiple semantic dimensions can be obtained, such as second description information 1 6042-1...second description information m 6042-m.
[0141] According to an embodiment of the present disclosure, for an operation generation task, generating operation prompt information based on the difference between the first description information and the second text information includes: extracting at least one operation from the second text information and determining the operation type based on the difference between the first description information and the second text information; or, extracting at least one initial operation from the second text information and determining the operation type based on the difference between the first description information and the second text information; and rewriting the initial operation using the first description information to obtain the operation.
[0142] The operation generation task may include an operation extraction task and a classification task. The operation extraction task is used to compare the first description information and the second text information, and extract one or more operations to be executed based on the differences between the two; then classify the extracted operations, generate the operation type of the corresponding operation, and obtain operation prompt information.
[0143] The operation generation task may also include an operation extraction task, a classification task, and a rewriting task. By executing the above operation extraction task and classification task, the corresponding initial operation and operation type are generated; then, the initial operation is rewritten according to the first description information to obtain an operation.
[0144] For example, if the second textual information is "The person should be holding a baby" and the first description information of the second reference image is "A woman with dark hair is smiling under a gray umbrella with a white flower hanging from it," the operation prompt information before rewriting is: 1. Add: Person holding a baby. Rewriting the initial operation "Person holding a baby" with the person "woman" in the first description information results in the following operation prompt information: 1. Add: Woman holding a baby.
[0145] In an embodiment of the present disclosure, the text analysis model can generate operation prompt information with clearer instructions by rewriting the initial operation, so as to subsequently generate more accurate second description information based on the operation prompt information and the first description information.
[0146] According to an embodiment of the present disclosure, for operation S230, at least one target image is determined based on at least one second description information, including: determining the similarity between the second description information and third description information used to describe the candidate image, wherein the search image library includes multiple candidate images; for each candidate image, determining a first comprehensive similarity based on the similarity between the second description information and each second description information; and determining at least one target image from the multiple candidate images based on the first comprehensive similarity.
[0147] The user-specified or default search image library includes multiple candidate images. At least one target image can be screened out from the multiple candidate images in the search image library according to the second description information.
[0148] For example, the search image library can store multiple candidate images and third description information used to describe each candidate image. For the second description information and the third description information of the language modality, the similarity can be calculated using an existing similarity algorithm. Then, for each candidate image, the similarity between the third description information and each second description information is weighted and summed to obtain the first comprehensive similarity of each candidate image. Alternatively, the first comprehensive similarity of each candidate image is obtained based on the number of similarities that reach a certain threshold. Thereafter, based on the first comprehensive similarity, at least one target image is directly determined from the multiple candidate images, for example, a candidate image with a high first comprehensive similarity is selected as the target image.
[0149] In the embodiment of the present disclosure, since the third description information of the language modality is adopted and the similarity between the third description information and the second description information is calculated, the first comprehensive similarity obtained by combining the above similarities can determine a target image with more accurate semantics under the same modality, and the search accuracy is higher.
[0150] According to an embodiment of the present disclosure, for operation S230, at least one target image is determined based on at least one second description information, including: obtaining image coding features of each candidate image in the search image library; for each candidate image, determining a second comprehensive similarity based on the similarity between the image coding features and the text coding features of each second description information; and determining at least one target image from multiple candidate images based on the second comprehensive similarity.
[0151] For example, a text encoding model can be used to encode the second description information of the language modality to obtain text encoding features, such as text vectors; and an image encoding model can be used to encode the candidate image of the visual modality to obtain image encoding features, such as image vectors.
[0152] Alternatively, a large multimodal model can be used to encode the second description and candidate images separately, obtaining text encoding features and image encoding features, respectively. For example, a Vision-Language Model (VLM) consisting of a visual encoder and a language encoder can be used to encode the second description and candidate images using the internal language encoder and image encoder, respectively.
[0153] An existing vector similarity algorithm can be used to calculate the similarity between each image encoding feature and the text encoding feature of each second description information. For a single candidate image, the similarities between the image encoding feature and each second description information are weighted and summed to obtain a second comprehensive similarity for each candidate image. Alternatively, the second comprehensive similarity for each candidate image is obtained based on the number of similarities that reach a certain threshold. Subsequently, at least one target image is directly determined from the multiple candidate images based on the second comprehensive similarity. For example, the candidate image with the highest second comprehensive similarity is selected as the target image.
[0154] In an embodiment of the present disclosure, when determining the target image, the similarity between the image coding features of the candidate image and the text coding features of the second description information is used. Therefore, the selected target image can meet the image search requirements from a visual perspective, thereby ensuring the accuracy of the target image from a visual perspective.
[0155] According to an embodiment of the present disclosure, the method also includes: for each candidate image, determining a third comprehensive similarity based on the similarity between the third description information and each second description information, and the similarity between the image coding feature and the text coding feature of each second description information; and determining at least one target image from multiple candidate images based on the third comprehensive similarity.
[0156] Referring to the above calculation method, the similarity between the third description information and the second description information, and the similarity between the image coding feature and the text coding feature can be obtained.
[0157] For ease of description, the two similarities described above are referred to as text-to-text similarity and text-to-image similarity. For each candidate image, the text-to-text similarity and text-to-image similarity between the candidate image and each piece of second description information can be combined to obtain a third comprehensive similarity. Subsequently, at least one target image is directly selected from the multiple candidate images based on the third comprehensive similarity.
[0158] In one embodiment, the text-to-text similarity and text-to-image similarity corresponding to each second description information can be weighted and summed according to the semantic granularity; then the weighted summation results of multiple semantic granularities are added and divided by the number of semantic granularities to obtain the third comprehensive similarity. For example, when weighting the text-to-text similarity and text-to-image similarity corresponding to each second description information, a parameter To control the weight between the two so that the sum of their weights is 1.
[0159] In an embodiment of the present disclosure, a third comprehensive similarity is obtained by integrating text-text similarity and text-image similarity, so that the target image determined according to the third comprehensive similarity can better meet the image search requirements from both text and image perspectives, enhance robustness, and improve the accuracy of image search.
[0160] Figure 7 The following schematically shows a scene diagram of determining a target image from a search image library according to a specific embodiment of the present disclosure. Figure 7 As shown, in embodiment 700, second text information 702 and a second reference image 703 are determined based on whether input information 701 includes first text information and / or first reference image information. After obtaining first description information 704 for describing second reference image 703, a text analysis model M1 is used to perform text analysis on first description information 704 and second text information 702 to obtain at least one second description information, such as second description information 1 705-1 ... second description information M 705-M.
[0161] Search image library 706 includes multiple candidate images, such as candidate images 1 706-1. Taking candidate image 1 706-1 as an example, the similarity between third description information 708-1 and each second description information describing candidate image 1 706-1 can be calculated. Based on the similarity between third description information 708-1 and each second description information, a comprehensive similarity 709-1 corresponding to candidate image 1 706-1 can be obtained. The above method can be used to obtain the corresponding comprehensive similarity for each candidate image.
[0162] Afterwards, at least one target image, such as target image 1 707 - 1 . . . , is determined from the search image library 706 based on the comprehensive similarity of each candidate image.
[0163] The comprehensive similarity may be a first comprehensive similarity, a second comprehensive similarity, or a third comprehensive similarity.
[0164] According to an embodiment of the present disclosure, the method further includes: converting the second reference image into first description information using the description generation large model; and converting the candidate image into third description information using the description generation large model.
[0165] The large model for description generation can be a large multimodal model (LMM), which is used to convert visual modality input into language modality output. For example, it can convert the second input image into a first description and the candidate image into a third description. The large model for description generation can be a pre-trained large multimodal model.
[0166] In the embodiment of the present disclosure, by generating description information of the corresponding image through the description generation large model, the operation can be simplified and the processing flow of the entire image search task is more intelligent.
[0167] To facilitate understanding of the present disclosure, a specific embodiment will be taken as an example to illustrate the image search process. Figure 8 A schematic diagram of a scene for determining a target image according to a specific embodiment of the present disclosure is schematically shown.
[0168] like Figure 8 As shown, the input information may include the first text information "-reference image description / previous round of dialogue information: a black dog and a brown dog are sitting next to a gray refrigerator in the kitchen. -text instructions / text feedback: only black dogs are needed; a woman in a black shirt is cooking on the stove"; the input information also includes a first reference image, such as image A. In this embodiment, the input image has the first reference image, so the second reference image in the search information, that is, image A, can also be obtained by express.
[0169] You can use the description to generate a large model Captioner to convert the second reference image Convert to the first description information , the above generation process can be expressed by formula (1):
[0170] (1)
[0171] like Figure 8As shown in the figure, the text analysis model Reasoner is used to sequentially perform the operation generation task and the third description generation task based on the prompt information Prompt1, generating three second description information of semantic granularity 1 to semantic granularity 3. For example, in the second step, "Semantic Granularity 1: A woman in a black shirt is cooking on the stove, and there is a black dog sitting next to her. Semantic Granularity 2: A woman in a black shirt is cooking on the stove in the kitchen, and there is a black dog sitting next to her. Semantic Granularity 3: A woman in a black shirt is cooking on the stove in the kitchen, and there is a black dog sitting next to the gray refrigerator." Figure 8 In the example, the operation prompt information generated by the operation generation task can be: "1. Add: Add a woman in a black shirt. 2. Add: The woman is cooking. 3. Delete: Delete the brown dog. 4. Keep: Keep the black dog." for the first step.
[0172] The process of generating at least one second description information using the text analysis model Reasoner can be expressed by the following formula (2):
[0173] (2)
[0174] in, is the second text message, Representing M operations A collection of They are the second description information of the three semantic granularities CE, ED, and CS respectively.
[0175] The second description information is generated using the visual language model VLM After vectorization, the obtained text encoding features can be vectorized Indicates that, and , d is the vector length, where .
[0176] In addition, in another branch, for the M candidate images in the search image library, such as image 1, image 2...image M, the third description information of each candidate image can be generated in advance by the description generation model Captioner. For example, given the text input "Please describe this image" and the search image library , describing the third description information of the corresponding image output by the large model Captioner, After this, use The third description information of each candidate image is vectorized to obtain text encoding features (text vectors), and each candidate image is vectorized to obtain image encoding features (image vectors). The above conversion process can be seen in formulas (3) and (4):
[0177] (3)
[0178] (4)
[0179] in, Indicates that the length of each third description information is The text vector, Indicates that the length of each candidate image is The image vector of Represents the matrix form of the text vector of all third description information, and Matrix representation of all image vectors.
[0180] After the text analysis model Reasoner generates at least one second description information, the similarity between the text encoding features of the second description information and the image encoding features of the candidate image and the text encoding features of the third description information of the candidate image is calculated to obtain the third comprehensive similarity of each candidate image. The above process can be expressed by formula (5):
[0181] (5)
[0182] in, is the vector form of the third comprehensive similarity of multiple candidate images, , represents the cosine similarity function, It is the abbreviation of text encoding features at three semantic granularities. For adjustment and The weight between .
[0183] The process of determining at least one target image according to the third comprehensive similarity is as shown in formula (6):
[0184] (6)
[0185] in, For N candidate images Image search sequence after sorting in descending order.
[0186] In one embodiment, you can select Output as the final target image.
[0187] Figure 9 The block diagram of an image search device according to a specific embodiment of the present disclosure is schematically shown.
[0188] like Figure 9 As shown, the image search apparatus 900 includes: a first determination module 910 , a generation module 920 , and a second determination module 930 .
[0189] The first determination module 910 is configured to determine multimodal search information based on input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image.
[0190] The generating module 920 is configured to perform text analysis on the second text information and the first description information used to describe the second reference image using a large text analysis model to generate at least one second description information.
[0191] The second determining module 930 is configured to determine at least one target image according to at least one second description information, wherein each target image is determined according to at least one second description information.
[0192] According to an embodiment of the present disclosure, the first determination module 910 includes: a first determination submodule for determining the first text information as second text information when the input information includes first text information, and obtaining a second reference image, wherein the second reference image includes a blank reference image. A second determination submodule for determining the first reference image as a second reference image when the input information includes the first reference image, and obtaining second text information, wherein the second text information includes blank text information. A third determination submodule for determining the first text information and the first reference image as second text information and a second reference image, respectively, when the input information includes the first text information and the first reference image.
[0193] According to an embodiment of the present disclosure, the image search apparatus 900 further includes: a third determination module configured to, when the input information includes an indication image library, determine the indication image library as the search image library, so as to determine the target image from the search image library; and a fourth determination module configured to, when the input information does not include the indication image library, determine the reference image library as the search image library.
[0194] According to an embodiment of the present disclosure, the input information is obtained based on at least one of the following methods: determining the first text information and / or the first reference image based on the input text and input image in at least one round of conversation; determining the first reference image and / or the first text information based on at least one output image and / or input text in at least one round of conversation, wherein the output image is a target image determined and output based on the input information before the output image.
[0195] According to an embodiment of the present disclosure, the text analysis big model is used to perform the first description generation task, and the generation module 920 includes: a first generation sub-module, used to use the text analysis big model to perform the first description generation task, so as to perform semantic analysis on the second text information and the first description information, and generate second description information of multiple semantic granularities; wherein the semantic granularity is related to the number and / or attributes of elements in the second description information, and the elements are extracted from the second text information and / or the first description information.
[0196] According to an embodiment of the present disclosure, the first generation submodule includes: an acquisition unit for acquiring prompt information, wherein the prompt information includes multiple semantic granularities to be generated, and explanatory information for each semantic granularity; a first generation unit for performing a first description generation task using a text analysis large model to perform semantic analysis on the second text information and the first description information based on each semantic granularity, and generate at least one second description information under each semantic granularity.
[0197] According to an embodiment of the present disclosure, the text analysis big model is used to sequentially execute the operation generation task and the second description generation task; the generation module 920 includes: a second generation sub-module, used to use the text analysis big model to execute the operation generation task, so as to generate operation prompt information according to the difference between the first description information and the second text information, and the operation prompt information includes at least one operation of at least one operation type; and a third generation sub-module, used to use the text analysis big model to execute the second description generation task, so as to generate at least one second description information according to the operation prompt information and the first description information.
[0198] According to an embodiment of the present disclosure, it further includes: when the second reference image is a blank reference image, the operation type of at least one operation included in the operation prompt information is an adding type.
[0199] According to an embodiment of the present disclosure, the text analysis big model is used to sequentially execute the operation generation task and the third description generation task, and the generation module 920 includes: a fourth generation sub-module, used to use the text analysis big model to execute the operation generation task, so as to generate operation prompt information according to the difference between the first description information and the second text information; and a fifth generation sub-module, used to use the text analysis big model to execute the third description generation task, so as to generate second description information of multiple semantic granularities according to the operation prompt information and the first description information.
[0200] According to an embodiment of the present disclosure, the fourth generation submodule includes: a second generation unit, used to extract at least one operation from the second text information and determine the operation type based on the difference between the first description information and the second text information; or, a third generation unit, used to extract at least one initial operation from the second text information and determine the operation type based on the difference between the first description information and the second text information; and rewrite the initial operation using the first description information to obtain the operation.
[0201] According to an embodiment of the present disclosure, the second determination module 930 includes: a fourth determination submodule, used to determine the similarity between the second description information and the third description information used to describe the candidate image, wherein the search image library includes multiple candidate images; a fifth determination submodule, used to determine, for each candidate image, a first comprehensive similarity based on the similarity between the second description information and each second description information; and a sixth determination submodule, used to determine at least one target image from the multiple candidate images based on the first comprehensive similarity.
[0202] According to an embodiment of the present disclosure, the second determination module 930 includes: an acquisition submodule for acquiring image coding features of each candidate image in the search image library; a seventh determination submodule for determining, for each candidate image, a second comprehensive similarity based on the similarity between the image coding features and the text coding features of each second description information; and an eighth determination submodule for determining at least one target image from multiple candidate images based on the second comprehensive similarity.
[0203] According to an embodiment of the present disclosure, the second determination module 930 also includes: a ninth determination submodule, for determining, for each candidate image, a third comprehensive similarity based on the similarity between the third description information and each second description information, and the similarity between the image coding feature and the text coding feature of each second description information; and a tenth determination submodule, for determining at least one target image from multiple candidate images based on the third comprehensive similarity.
[0204] According to an embodiment of the present disclosure, the image search device 900 further includes: a first conversion module for converting the second reference image into first description information using the description generation model; and a second conversion module for converting the candidate image into third description information using the description generation model.
[0205] Figure 10 The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.
[0206] In the embodiments of the present disclosure, inspired by the von Neumann structure in modern computer theory, such as Figure 10As shown, the AI agent 1000 may include five core modules: an input module 1010 , a control module 1020 , a storage module 1030 , a calculation module 1040 and an output module 1050 .
[0207] Input module 1010 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that AI agent 1000 can understand and process. Input module 1010 is the primary link for AI agent 1000 to interact with the outside world. It enables AI agent 1000 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0208] In an example, the input module 1010 may input the input information described above, the first text information, and the first reference image.
[0209] In this example, the control module 1020 is the core support for the AI agent 1000 to handle complex tasks. The control module 1020 can execute the image search method described above.
[0210] In the example, the control module 1020 will continuously interact with the storage module 1030, the computing module 1040, and / or the output module 1050 during operation. However, it should be noted that in the embodiment of the present disclosure, the control module 1020 acts as a single initiator to initiate communication with the storage module 1030, the computing module 1040, and / or the output module 1050, and there is no communication coupling between the storage module 1030, the computing module 1040, and the output module 1050.
[0211] In this example, the performance of control module 1020 may be closely related to the large model underlying AI agent 1000. To fully leverage the capabilities of the large model, the internal structure of control module 1020 may be designed to be highly configurable and extensible to handle a variety of tasks and requirements in real-world scenarios.
[0212] The storage module 1030 may be responsible for memorizing information such as historical conversations, event flows, etc. The aforementioned prompt information, the third description information of the candidate images, etc. may be included in the storage module 1030 .
[0213] The operation module 1040 can be regarded as a predefined tool library. As described above, the controls for text encoding and image encoding can be included in the operation module 1040.
[0214] In an example, the output module 1050 may output at least one target image described above.
[0215] The AI agent 1000 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0216] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0217] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0218] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0219] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0220] Figure 11 A block diagram of an electronic device 1100 suitable for implementing an image search method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0221] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. RAM 1103 may also store various programs and data required for the operation of device 1100. Computing unit 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.
[0222] Various components in device 1100 are connected to an input / output (I / O) interface 1105, including an input unit 1106, such as a keyboard and mouse; an output unit 1107, such as various types of displays and speakers; a storage unit 1108, such as a magnetic disk and optical disk; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0223] Computing unit 1101 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1101 performs the various methods and processes described above, such as the image search method. For example, in some embodiments, the image search method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by computing unit 1101, one or more steps of the image search method described above may be performed. Alternatively, in other embodiments, computing unit 1101 may be configured to perform the image search method via any other suitable means (e.g., via firmware).
[0224] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0225] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0226] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0227] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0228] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0229] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0230] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0231] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An image search method, comprising: Determining multimodal search information according to input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image; Performing text analysis on the second text information and the first description information describing the second reference image using a large text analysis model to generate at least one second description information, the second description information being a description of the image that meets the user's image search requirement and is obtained by combining the second text information and the first description information; determining at least one target image according to at least one piece of the second description information, wherein each of the target images is determined according to at least one piece of the second description information; The determining of multimodal search information based on input information for image search includes: In a case where the input information includes the first text information, determining the first text information as the second text information, and acquiring the second reference image, where the second reference image includes a blank reference image; In a case where the input information includes the first reference image, determining the first reference image as the second reference image, and acquiring the second text information, where the second text information includes blank text information; In a case where the input information includes the first text information and the first reference image, the first text information and the first reference image are determined as the second text information and the second reference image, respectively.
2. The method according to claim 1, further comprising: In a case where the input information includes an indication image library, determining the indication image library as a search image library, so as to determine the target image from the search image library; In a case where the input information does not include the indication image library, a reference image library is determined as the search image library.
3. The method according to claim 1, wherein the input information is obtained based on at least one of the following methods: Determining the first text information and / or the first reference image based on input text and input image in at least one round of conversation; The first reference image and / or first text information is determined based on at least one output image and / or the input text in at least one round of the conversation, wherein: The output image is the target image determined and output according to the input information before the output image.
4. The method according to any one of claims 1 to 3, wherein The text analysis large model is used to perform the first description generation task, and the text analysis large model is used to perform text analysis on the second text information and the first description information used to describe the second reference image to generate at least one second description information, including: Executing the first description generation task using a large text analysis model to perform semantic analysis on the second text information and the first description information to generate the second description information at multiple semantic granularities; The semantic granularity is related to the number and / or attributes of elements in the second description information, and the elements are extracted from the second text information and / or the first description information.
5. The method according to claim 4, wherein The step of using the text analysis model to perform the first description generation task to perform semantic analysis on the second text information and the first description information to generate the second description information at multiple semantic granularities includes: Acquiring prompt information, wherein the prompt information includes a plurality of the semantic granularities to be generated and explanation information for each of the semantic granularities; and The first description generation task is performed using the text analysis large model to perform semantic analysis on the second text information and the first description information based on each semantic granularity to generate second description information at each semantic granularity.
6. The method according to claim 1, wherein The text analysis large model is used to sequentially execute the operation generation task and the second description generation task; the text analysis large model is used to perform text analysis on the second text information and the first description information used to describe the second reference image to generate at least one second description information, including: Executing the operation generation task using a text analysis model to generate operation prompt information based on a difference between the first description information and the second text information, the operation prompt information including at least one operation of at least one operation type; and The second description generation task is performed using the text analysis model to generate at least one second description information based on the operation prompt information and the first description information.
7. The method according to claim 6, further comprising: In a case where the second reference image is a blank reference image, the operation type of at least one operation included in the operation prompt information is an adding type.
8. The method according to claim 1, wherein the text analysis model is used to sequentially execute an operation generation task and a third description generation task, wherein the text analysis model is used to perform text analysis on the second text information and the first description information used to describe the second reference image to generate at least one second description information, including: Utilizing the text analysis model to execute the operation generation task, so as to generate operation prompt information according to the difference between the first description information and the second text information; as well as The third description generation task is performed using a large text analysis model to generate the second description information of multiple semantic granularities based on the operation prompt information and the first description information.
9. The method according to any one of claims 6 to 8, wherein The generating operation prompt information according to the difference between the first description information and the second text information includes: extracting at least one operation from the second text information and determining the operation type according to the difference between the first description information and the second text information; or According to the difference between the first description information and the second text information, at least one initial operation is extracted from the second text information and the operation type is determined; and the initial operation is rewritten using the first description information to obtain the operation.
10. The method according to claim 1, wherein The determining at least one target image according to at least one of the second description information includes: determining a similarity between the second description information and third description information used to describe a candidate image, wherein the search image library includes a plurality of the candidate images; For each candidate image, determine a first comprehensive similarity based on the similarity between the candidate image and each piece of the second description information; and At least one target image is determined from the plurality of candidate images according to the first comprehensive similarity.
11. The method according to claim 1, wherein The determining at least one target image according to at least one of the second description information includes: Obtain image encoding features of each candidate image in the search image library; For each candidate image, determining a second comprehensive similarity based on a similarity between the image encoding feature and each text encoding feature of the second description information; and At least one target image is determined from the plurality of candidate images according to the second comprehensive similarity.
12. The method according to claim 10 or 11, further comprising: For each candidate image, determine a third comprehensive similarity based on a similarity between the third description information and each piece of the second description information, and a similarity between an image encoding feature and a text encoding feature of each piece of the second description information; as well as At least one target image is determined from the plurality of candidate images according to the third comprehensive similarity.
13. The method according to claim 1, further comprising: Converting the second reference image into the first description information using a description generation model; as well as The description is used to generate a large model to convert the candidate image into third description information.
14. An image search device, comprising: a first determining module, configured to determine multimodal search information based on input information for image search, wherein the input information includes first text information and / or a first reference image, and the search information includes second text information and a second reference image; a generation module configured to perform text analysis on the second text information and the first description information used to describe the second reference image using a large text analysis model to generate at least one second description information, wherein the second description information is a description information of the image that meets the user's image search requirement and is obtained by combining the second text information and the first description information; and a second determining module, configured to determine at least one target image according to at least one piece of the second description information, wherein each of the target images is determined according to at least one piece of the second description information; The first determining module includes: a first determining submodule, configured to, when the input information includes the first text information, determine the first text information as the second text information, and obtain the second reference image, where the second reference image includes a blank reference image; a second determining submodule, configured to, when the input information includes the first reference image, determine the first reference image as the second reference image, and obtain the second text information, where the second text information includes blank text information; The third determining submodule is configured to, when the input information includes the first text information and the first reference image, determine the first text information and the first reference image as the second text information and the second reference image, respectively.
15. The apparatus according to claim 14, further comprising: a third determining module, configured to, when the input information includes an indication image library, determine the indication image library as a search image library, so as to determine the target image from the search image library; The fourth determining module is configured to determine a reference image library as the search image library when the input information does not include the indication image library.
16. The apparatus according to claim 14, wherein the input information is obtained based on at least one of the following methods: Determining the first text information and / or the first reference image based on input text and input image in at least one round of conversation; The first reference image and / or first text information is determined based on at least one output image and / or the input text in at least one round of the conversation, wherein: The output image is the target image determined and output according to the input information before the output image.
17. The device according to any one of claims 14 to 16, wherein: The text analysis model is used to perform the first description generation task, and the generation module includes: a first generating submodule, configured to execute the first description generating task by using a large text analysis model, so as to perform semantic analysis on the second text information and the first description information, and generate the second description information of multiple semantic granularities; The semantic granularity is related to the number and / or attributes of elements in the second description information, and the elements are extracted from the second text information and / or the first description information.
18. The device according to claim 17, wherein The first generation submodule includes: an acquiring unit, configured to acquire prompt information, wherein the prompt information includes a plurality of the semantic granularities to be generated and interpretation information for each of the semantic granularities; and The first generation unit is used to use the text analysis model to execute the first description generation task, so as to perform semantic analysis on the second text information and the first description information based on each semantic granularity, and generate at least one second description information under each semantic granularity.
19. The device according to claim 14, wherein The text analysis model is used to sequentially execute the operation generation task and the second description generation task; the generation module includes: a second generation submodule, configured to perform the operation generation task using a text analysis macromodel to generate operation prompt information based on a difference between the first description information and the second text information, the operation prompt information including at least one operation of at least one operation type; and The third generation submodule is used to use the text analysis model to perform the second description generation task to generate at least one second description information based on the operation prompt information and the first description information.
20. The apparatus according to claim 19, further comprising: In a case where the second reference image is a blank reference image, the operation type of at least one operation included in the operation prompt information is an adding type.
21. The apparatus according to claim 14, wherein the text analysis model is used to sequentially execute an operation generation task and a third description generation task, and the generation module comprises: a fourth generation submodule, configured to execute the operation generation task using the text analysis model to generate operation prompt information based on the difference between the first description information and the second text information; as well as The fifth generation submodule is used to use the text analysis model to perform the third description generation task to generate the second description information of multiple semantic granularities based on the operation prompt information and the first description information.
22. The device according to any one of claims 19 to 21, wherein: The fourth generation submodule includes: a second generating unit configured to extract at least one operation from the second text information and determine an operation type based on a difference between the first description information and the second text information; or The third generating unit is configured to extract at least one initial operation from the second text information and determine an operation type based on a difference between the first description information and the second text information; and rewrite the initial operation using the first description information to obtain the operation.
23. The apparatus according to claim 14, wherein The second determining module includes: a fourth determining submodule, configured to determine a similarity between the second description information and third description information used to describe a candidate image, wherein the search image library includes a plurality of the candidate images; a fifth determining submodule, configured to determine, for each candidate image, a first comprehensive similarity based on a similarity between the candidate image and each piece of the second description information; and The sixth determination submodule is configured to determine at least one target image from the plurality of candidate images according to the first comprehensive similarity.
24. The apparatus according to claim 14, wherein The second determining module includes: An acquisition submodule is used to obtain image coding features of each candidate image in the search image library; a seventh determining submodule, configured to determine, for each candidate image, a second comprehensive similarity based on a similarity between the image encoding feature and a text encoding feature of each second description information; and An eighth determination submodule is configured to determine at least one target image from the plurality of candidate images according to the second comprehensive similarity.
25. The apparatus according to claim 23 or 24, wherein the second determining module further comprises: a ninth determination submodule, configured to determine, for each candidate image, a third comprehensive similarity based on a similarity between the third description information and each piece of the second description information, and a similarity between an image coding feature and a text coding feature of each piece of the second description information; as well as The tenth determining submodule is configured to determine at least one target image from the plurality of candidate images according to the third comprehensive similarity.
26. The apparatus of claim 14, further comprising: A first conversion module, configured to convert the second reference image into the first description information by using a description generation model; as well as The second conversion module is used to use the description to generate a large model to convert the candidate image into third description information.
27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.
28. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 13.
29. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Searching method and device, electronic equipment and storage medium
CN115858941A
Multi-modal image retrieval method and system based on self-training Chinese CLIP model
CN117786155A